evals-ops
Build and run evals for LLM and agent systems: golden datasets, LLM-as-a-judge with bias control, trajectory/step/outcome scoring, adversarial refuters, and CI regression gates. Triggers on: eval, evals, eval harness, golden dataset, golden set, LLM-as-a-judge, judge rubric, judg
Install
npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/evals-ops
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
git clone https://github.com/0xDarkMatter/claude-mods.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Evals Ops
Evals are the prerequisite, not the polish. You cannot tune a prompt, a retriever, a compaction strategy or a memory layer without a harness that says whether the change made things better. Teams that skip this ship vibes and learn about regressions from users.
This skill is the operational layer: what to measure, how to build the dataset, how to make a judge trustworthy, and how to gate CI on it without teaching everyone to ignore red.
Route first
| The ask | Go to |
|---|---|
| "What should I even measure?" | Three levels → references/eval-taxonomy.md |
| "Where do the test cases come from?" | Golden set → references/golden-datasets.md |
| "My judge disagrees with me / is it any good?" | Judges → references/llm-judge.md |
| "Verify a finding is real, not plausible" | Refuters → references/adversarial-verification.md |
| "Is my RAG retrieval any good?" | Retrieval → references/retrieval-eval.md |
| "Where do the human labels come from?" | references/annotation-workflow.md |
| "Should this block the merge?" | Gating → references/regression-gating.md |
| "Did this change really make it worse?" | Is the drop real |
| "Optimise against the eval / run it overnight" | Hillclimbing → references/hillclimbing.md |
| "Which platform should we use?" | references/tooling-landscape.md |
| "Just give me a starting file" | Assets — golden set, rubric, runner, CI gate |
The 60-second version
- Write 20 cases before you write a metric. A dataset you can eyeball beats a metric you cannot interpret. Grow to 100-300, then freeze it.
- Prefer a deterministic assertion to any judge. The JSON parsed, the tool was called with the right argument, the query returned 3 rows - free, instant, zero variance. Reach for a judge only where correctness is genuinely a matter of language.
- Score the trajectory, not just the answer. Most teams only check the final artifact and are surprised when a right answer came from a wrong path that breaks tomorrow.
- Calibrate the judge against humans before trusting it. Cohen kappa, not raw
agreement.
scripts/judge-calibration.pydoes the arithmetic and the verdict. - Blocking gates must be deterministic. Judge metrics start advisory. One flaky red permanently devalues the signal.
Three levels of agent eval
Most teams do only the third, then wonder why quality is unpredictable.
| Level | Question | Signal | Typical evaluator |
|---|---|---|---|
| Outcome | Is the final artifact correct? | Binary or scored end state | Deterministic assertion, unit test, judge |
| Step | Was this tool call right? | Per-span: tool choice, arg shape, arg values | Schema/argument assertion, span-level judge |
| Trajectory | Was the path sensible? | Sequence, loops, redundancy, cost | Reference-trajectory match, rubric judge |
The failure that motivates all three: an agent reaches the right end state by an accidental route - the lucky pass. Outcome-only scoring records that as a win, and the same case fails next week when the accident does not recur. Conversely a trajectory-only score punishes a legitimately novel-but-correct path. Gate on outcome; keep step and trajectory as the diagnostics that tell you why the gate moved.
For multi-turn or stateful agents also report pass^k (all k independent runs of the same
case succeed) alongside pass@k (any of k succeeded). pass@k flatters a non-deterministic
agent; pass^k is the number that predicts production. Full treatment:
references/eval-taxonomy.md.
The golden set
A golden set is a reviewed, versioned, deliberately frozen collection of inputs with trusted expected outputs. Frozen matters: a set that grows every sprint cannot tell you whether last week's number moved because the system changed or because the set did.
Composition - four buckets, not one:
| Bucket | Source | Why |
|---|---|---|
| Production sample | Real traffic, stratified | Keeps the score connected to what users actually do |
| Failure replays | Every incident that reached a human | Regression protection; the easiest cases to justify |
| Adversarial | Injections, contradictions, refusal-bait | The class both agents and judges fail silently on |
| Edge cases | Empty, huge, ambiguous, multilingual | Where deterministic code breaks first |
Sizing: 20 to start, 100-300 for a working regression set, 200-500 once you have production traffic to sample. Beyond that you are usually buying latency, not signal - add cases when a new failure class appears, and record in the case itself why it exists.
scripts/goldenset-audit.py checks a set for the rot that accumulates: duplicates, a
bucket that quietly became 90% of the set, undated cases, and drift from a frozen manifest.
python3 scripts/goldenset-audit.py evals/golden.jsonl --json | jq '.data.findings[]'
Depth: references/golden-datasets.md.
LLM-as-a-judge
A judge is a measurement instrument. Instruments need calibration, and this one has documented, reproducible biases:
| Bias | What it does | Mitigation |
|---|---|---|
| Position | Prefers whichever candidate was shown first | Run both orders and average; or score absolutely, not pairwise |
| Verbosity | Rates longer answers higher regardless of quality | Separate correctness from style in the rubric; penalise unsupported length |
| Self-preference | Rates its own model family's output higher | Judge with a different family than the one under test |
| Scale drift | 1-5 scores cluster and shift between model versions | Binary pass/fail against explicit criteria; pin the judge model version |
Panel vs N-identical. Three calls to the same judge with the same rubric mostly buys the same bias three times. A panel with distinct lenses - one asks "is this supported by the source?", one "does it follow the stated policy?", one "would this reproduce?" - finds failure modes redundancy structurally cannot. Use N-identical only to measure the judge's own variance, which is a different question worth asking once.
When a judge is the wrong tool: if you can express the criterion as code, do. A judge costs money, adds latency, drifts across model versions, and has variance a regex does not. Judges earn their place on faithfulness, tone, policy compliance, and "is this a reasonable answer to an open question" - nowhere else.
Calibrate before you trust. Label 50-200 cases by hand, run the judge on the same cases, and compute Cohen kappa (raw agreement lies when classes are imbalanced):
python3 scripts/judge-calibration.py evals/labels.jsonl --min-kappa 0.6
# exit 0 = calibrated; exit 10 = below threshold, fix the rubric before shipping it
kappa >= 0.8 production-ready - 0.6-0.8 substantial, usable with care - below 0.6 the rubric
is the problem, not the model. Re-sample ~50 fresh cases periodically; judges drift when the
underlying model version moves. Depth, including bias-probe design: references/llm-judge.md.
Measure the human-human ceiling first. A judge cannot beat the agreement two people
achieve with each other. Two annotators on 30-50 shared cases gives you that number - and
if it is below ~0.6, the rubric is ambiguous and every label you produce against it is
wasted. How to run the sessions, stratify the sample, and adjudicate disagreements:
references/annotation-workflow.md.
Adversarial verification
For findings rather than scores - bug reports, audit results, review comments - flip the prompt: ask the verifier to REFUTE, not to confirm. "Try to refute this finding; default to refuted if uncertain" kills plausible-but-wrong results that an "is this correct?" prompt waves through, because agreement is the path of least resistance for a model.
Then take a majority: run 3 refuters, keep the finding only if at least 2 fail to refute it. Prefer perspective-diverse refuters (correctness / security / does-it-actually-reproduce) over three identical skeptics - same reasoning as judge panels.
This composes with the parallel-work skills rather than duplicating them: fleet-ops and
parallel-ops own the fan-out mechanics; this skill owns the scoring contract the refuters
return. See references/adversarial-verification.md.
Retrieval
RAG is the most common eval target and the most commonly mis-measured. Scoring only the final answer averages two independent failures into one uninterpretable number:
| Right context | Wrong context | |
|---|---|---|
| Answer correct | Working | Lucky - the model knew it anyway; scores as a pass |
| Answer wrong | Generation bug - chunking, prompt, model | Retrieval bug - embeddings, index, query rewriting |
Record the retrieved chunk ids next to every answer and that opaque score becomes a 2x2 you can assign to a team. Gate on recall@k (a precision failure degrades an answer; a recall failure makes a correct one impossible) and on citation-id validity, which is free and catches confident answers attached to unrelated sources. Retrieval is the one place deterministic scoring genuinely dominates - you have ground-truth ids, so skip the judge.
The bucket almost everyone omits: questions the corpus cannot answer. Without them the
suite cannot detect hallucination under retrieval failure, which is what users hit most.
Metrics, the six failure classes, and ground-truth construction: references/retrieval-eval.md.
Regression gating
The rule that keeps a gate alive: a blocking check must never be flaky.
| Check | Gate |
|---|---|
| Deterministic assertions (schema, tool-call, exact match) | Blocking. Any failure fails CI. |
| Judge metrics, first few weeks | Advisory. Post the delta as a PR comment. |
| Judge metrics, calibrated (kappa >= 0.6) and variance-measured | Blocking with a margin below the rolling baseline |
| Cost and p95 latency per case | Blocking on a ceiling, advisory on the trend |
Set the threshold below the baseline by more than the measured noise floor: if the suite scores 0.88 +/- 0.03 across reruns, gate at 0.80, not 0.87. You cannot know the noise floor from a single run - commit a rolling window of run results to git and read the variance off it. That committed history is also what distinguishes "today is noisy" from "today broke".
Is the drop real?
A noise floor tells you the aggregate moved unusually far. It does not tell you the same cases moved. Two runs over one frozen set are paired binary outcomes, and the tool for those is McNemar's exact test over the discordant pairs only.
This matters because a change that breaks 8 cases and fixes 7 moves the headline score by 0.01 - invisible against any noise floor - while silently swapping which 15 things work. The paired view names those 15 cases; a score comparison structurally cannot. Whether 8-vs-7 is significant is a separate question - it is not - but knowing which cases flipped is what lets you go and look.
python3 scripts/eval-baseline.py evals/history.jsonl --baseline-results base.jsonl --candidate-results new.jsonl
# exit 10 = significant regression, and it names the cases that flipped
Attribute cost and latency per eval run from the start. An eval suite is the only place
you find out the accuracy win cost 4x the tokens, and retrofitting attribution after the
harness exists is far more annoying than a tokens_in / tokens_out / ms field per case.
Full CI shape and the noise-floor method: references/regression-gating.md.
Hillclimbing
Once the harness measures, the obvious move is to optimise against it. That works, and it is also the fastest way to make a good suite useless.
The loop belongs to iterate - this skill owns what goes wrong. Every
hillclimbing failure is a property of the metric, not of the loop:
Banking noise.
iteratekeeps a change when the metric beats the previous best - correct for line coverage, a coin flip for an eval score. At 0.88 +/- 0.03, a measured 0.90 is not evidence. Gate the keep decision on the noise floor instead:python3 scripts/eval-baseline.py history.jsonl --candidate iter.jsonl --accept # exit 0 = KEEP (a real improvement), 10 = DISCARD (noise or worse)--acceptdeliberately inverts the CI meaning of "noise": CI asks did this get worse, a hillclimb asks is this improvement real. Noise fails the second question.Overfitting the frozen set. Split train / validation / held-out before optimising, never show the validation set to whatever proposes changes, and treat held-out as a budget you spend at milestones - not a dashboard.
Keeping a champion instead of a frontier. One aggregate best hides which cases a candidate won. Retaining candidates that are best on at least one case is what stops the loop walling itself into a local optimum.
And the eval-design consequence: a scalar score gives an optimizer nothing to reflect on.
A judge returning {"reason": ..., "verdict": ...} can be improved against; one returning
0.4 cannot. That field costs nothing today and is what makes automated optimization
tractable later.
Splits, the optimizer landscape (GEPA, MIPROv2, APE/ORPO/SPO), the pre-flight checklist,
and the extraction trigger for a future prompt-optimization-ops: references/hillclimbing.md.
Tooling
Trace-level observability and eval scoring have converged into the same products - you are picking one system, not two. Open-source cores worth knowing: DeepEval (pytest-native), MLflow (tracing, eval and prompt versioning in one OSS platform), Opik, Langfuse, Arize Phoenix. Commercial-first: Braintrust (dataset curation for non-engineers), AgentOps, LangSmith, Arize.
Honest default: start with a JSONL file and a 40-line runner. Adopt a platform when you
need shared dataset curation, trace search across production traffic, or scheduled runs -
not before. Which-one-when: references/tooling-landscape.md.
The landscape moves fast. Treat every version, price and feature claim in that reference as needing re-verification; it carries a datestamp for exactly that reason.
Scripts
| Script | Use |
|---|---|
scripts/judge-calibration.py |
Judge-vs-human agreement: Cohen kappa, confusion matrix, per-class breakdown, verbosity/position bias probes. Exit 10 = below --min-kappa. |
scripts/goldenset-audit.py |
Golden-set health: duplicates, bucket balance, staleness, freeze-manifest drift. Exit 10 = findings. |
scripts/eval-baseline.py |
Noise floor from run history, the threshold your gate should use, and McNemar's exact test naming the cases that flipped. Exit 10 = confirmed regression or a cost/latency ceiling breach; --accept turns it into a hillclimb keep/discard gate. |
All three accept --help and --json, and are offline and stdlib-only.
python3 scripts/judge-calibration.py labels.jsonl --json | jq '.data.kappa'
python3 scripts/goldenset-audit.py golden.jsonl --freeze manifest.json
python3 scripts/eval-baseline.py history.jsonl --json | jq '.data.recommended_threshold'
Assets
Copy-and-adapt starting points, so the first hour goes on deciding what to measure rather than on scaffolding:
| Asset | What it is |
|---|---|
assets/golden-set.example.jsonl |
10 worked cases across all four buckets, with why and criteria filled in |
assets/eval-runner.template.py |
The 40-line runner this skill tells you to start with - two ADAPT blocks, cost/latency/pass^k pre-wired |
assets/judge-rubric.template.md |
One-criterion rubric with the bias-counter instructions and the calibration checklist |
assets/eval-gate.template.yml |
GitHub Actions workflow encoding the tier ladder: deterministic blocks on push, judge advisory on PR, k=3 nightly |
References
references/eval-taxonomy.md- outcome/step/trajectory, lucky pass, pass@k vs pass^k, metric selectionreferences/golden-datasets.md- building, four-bucket composition, sizing, freeze discipline, rotreferences/llm-judge.md- bias catalog and mitigations, rubric design, panels, calibration methodreferences/adversarial-verification.md- refute-not-confirm, majority thresholds, lens diversityreferences/retrieval-eval.md- RAG: recall@k, the retrieval-vs-generation split, ground truth, failure classesreferences/annotation-workflow.md- where human labels come from: the human-human ceiling, sampling, adjudication, driftreferences/regression-gating.md- blocking vs advisory, noise floor, McNemar, CI shape, cost/latency attributionreferences/hillclimbing.md- optimising against an eval without destroying it: noise, overfitting, frontiers, optimizersreferences/tooling-landscape.md- platform comparison with verification datestamps
Files (claude-mods)
-
assets
-
eval-gate.template.yml 6.3 KB
# Eval CI gate — the tier ladder as a workflow. Copy to .github/workflows/evals.yml. # # The shape encodes one rule: A BLOCKING CHECK MUST NEVER BE FLAKY. Deterministic # assertions block on every push. Judge metrics run on PRs and start ADVISORY -- # they are promoted to blocking only after judge-calibration.py says kappa >= 0.6 # and you have measured the noise floor. The full k=3 consistency run and the # history append happen nightly on main, where cost and duration are affordable. # # ADAPT before use: # 1. <angle brackets>, the eval commands, and the secret names. # 2. ACTION VERSIONS. The `uses:` majors below are placeholders, NOT a # recommendation -- check the current major for each action and pin to a # full commit SHA (`uses: actions/checkout@<sha> # v5.0.0`). A floating # major is a supply-chain surface; see the supply-chain-defense skill. # 3. SCRIPT PATHS. $EVALS_OPS below assumes the skill is installed at # ~/.claude/skills/evals-ops. If you vendor claude-mods into the repo # instead, set it to skills/evals-ops. The scripts are stdlib-only, so # copying the three you use into the repo is also a legitimate option and # removes the dependency entirely. # # See references/regression-gating.md for why each tier sits where it does. name: evals on: push: branches: ["**"] pull_request: schedule: - cron: "0 3 * * *" # nightly on main workflow_dispatch: concurrency: group: evals-${{ github.ref }} cancel-in-progress: true env: DATASET_VERSION: golden-v3 # bump = re-baseline; never compare across it EVALS_OPS: $HOME/.claude/skills/evals-ops # ADAPT: see header note 3 PYTHON_VERSION: "3.12" jobs: # --- TIER 0: deterministic. Blocking, zero tolerance, every push. ---------- # No model calls, no judge, no network to a provider. Seconds and $0, which is # exactly why it is allowed to fail the build. deterministic: runs-on: ubuntu-latest steps: - uses: actions/checkout@vX # ADAPT: current major, pinned to a SHA - uses: actions/setup-python@vX # ADAPT: current major, pinned to a SHA with: python-version: ${{ env.PYTHON_VERSION }} # The golden set is an artifact with its own integrity. Editing a frozen # case in place to make it pass is fitting the test to the code -- this # step is what catches it. - name: Golden-set health and freeze check run: | python3 "$EVALS_OPS"/scripts/goldenset-audit.py \ evals/golden.jsonl --freeze evals/freeze.json - name: Deterministic assertions run: <your-deterministic-eval-command> # e.g. pytest evals/test_assertions.py # --- TIER 1/2: judge metrics on a PR subset. Advisory until calibrated. ---- judge: if: github.event_name == 'pull_request' runs-on: ubuntu-latest # ADAPT: flip to `false` only once judge-calibration.py reports kappa >= 0.6 # AND you have measured the noise floor. Until then a red here is noise, and # three unexplained reds kill the gate socially even while it still enforces. continue-on-error: true steps: - uses: actions/checkout@vX # ADAPT: current major, pinned to a SHA with: fetch-depth: 0 # history file needs real history - uses: actions/setup-python@vX # ADAPT: current major, pinned to a SHA with: python-version: ${{ env.PYTHON_VERSION }} # Judge calibration is a gate on the JUDGE, not on the code. If the rubric # has drifted below threshold, the scores below are not evidence. - name: Judge is calibrated run: | python3 "$EVALS_OPS"/scripts/judge-calibration.py \ evals/labels.jsonl --min-kappa 0.6 - name: Run evals (stratified subset, k=1) env: ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} run: | python3 evals/runner.py evals/golden.jsonl \ -k 1 --dataset-version "$DATASET_VERSION" \ --out results.jsonl --history /tmp/run.jsonl # Compare against the committed baseline. Exit 10 = confirmed regression # (outside the noise floor); exit 0 = noise or improvement. Naming the # flipped cases is what stops people ignoring the result. - name: Regression check run: | python3 "$EVALS_OPS"/scripts/eval-baseline.py \ evals/history.jsonl --candidate /tmp/run.jsonl \ --baseline-results evals/baseline-results.jsonl \ --candidate-results results.jsonl - name: Comment the delta if: always() run: <post the eval-baseline.py --json summary as a PR comment> # --- TIER 3/4: nightly. Full set, k=3, cost and latency, history append. --- nightly: if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch' runs-on: ubuntu-latest permissions: contents: write # to commit the appended history row steps: - uses: actions/checkout@vX # ADAPT: current major, pinned to a SHA - uses: actions/setup-python@vX # ADAPT: current major, pinned to a SHA with: python-version: ${{ env.PYTHON_VERSION }} # k=3 gives pass^k -- what production actually experiences. pass@k flatters # a non-deterministic agent and must never be the headline. # NOTE: never retry a failing eval to green. Retry-until-pass turns a real # regression into a flake report; k=3 reports ALL attempts by design. - name: Full eval, k=3 env: ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} run: | python3 evals/runner.py evals/golden.jsonl \ -k 3 --dataset-version "$DATASET_VERSION" \ --out nightly-results.jsonl --history evals/history.jsonl # Ceilings, not trends. A trend gate fires on noise; a ceiling encodes a # product decision someone actually made. - name: Cost and latency ceilings run: | python3 "$EVALS_OPS"/scripts/eval-baseline.py \ evals/history.jsonl --max-cost-usd <2.50> --max-p95-ms <6000> - name: Commit the history row run: | git config user.name "eval-bot" git config user.email "eval-bot@users.noreply.github.com" git add evals/history.jsonl git diff --staged --quiet || git commit -m "chore(evals): nightly run" git push -
eval-runner.template.py 6.3 KB
#!/usr/bin/env python3 """The 40-line eval runner. Copy into your repo and adapt the two ADAPT blocks. This is deliberately small and dependency-free. It is the thing to start with before adopting a platform (see references/tooling-landscape.md) -- it forces you to decide what you are measuring, which is the hard part no platform does for you. What it does: - reads a golden set (JSONL, one case per line) - runs each case k times through YOUR system - applies deterministic assertions first, judge criteria only where needed - records cost and latency PER CASE from the first run, not retrofitted later - writes a per-case results file and appends one summary row to a run history Usage: eval-runner.py GOLDEN.jsonl [-k 3] [--out results.jsonl] [--history history.jsonl] Exit: 0 ran to completion, 1 a case raised, 2 usage The summary row is what eval-baseline.py consumes to tell noise from regression. """ import argparse import json import time # ========================================================================== # ADAPT BLOCK 1 -- call your system. # Return a dict. Include token counts and the model id; you will want them and # threading them in later means touching every runner, result and dashboard. # ========================================================================== def run_system(case): # from my_app import agent # result = agent.invoke(case["input"]) # return {"output": result.text, "tool_calls": result.tool_calls, # "tokens_in": result.usage.input_tokens, # "tokens_out": result.usage.output_tokens, # "model": result.model} raise NotImplementedError("wire this to your agent") # ========================================================================== # ADAPT BLOCK 2 -- call your judge, ONLY for criteria code cannot check. # Return True/False per criterion. Pin the judge model version: an unpinned # judge silently re-baselines your whole history (references/llm-judge.md). # ========================================================================== JUDGE_MODEL = "<pin-a-specific-model-version-here>" def judge(criterion, case, result): # from my_app import judge_client # verdict = judge_client.score(model=JUDGE_MODEL, temperature=0, # criterion=criterion, output=result["output"]) # return verdict["pass"] raise NotImplementedError("wire this to your judge, or drop criteria entirely") def deterministic(case, result): """Cheap assertions run first and are the ones that gate CI. Every criterion you can express here instead of in `judge` removes cost, latency AND variance at once. Re-audit periodically -- rubric items become codifiable once the output format stabilises. """ expected = case.get("expected") if not expected: return None # nothing deterministic to check; criteria carry this case if "tool" in expected: calls = result.get("tool_calls") or [] if not any(c.get("name") == expected["tool"] for c in calls): return False if "args" in expected: got = next(c for c in calls if c.get("name") == expected["tool"]) for key, want in expected["args"].items(): if got.get("args", {}).get(key) != want: return False return True if expected.get("refused") or expected.get("clarifies"): # Absence of a tool call is the check; the wording is a judge question. return not (result.get("tool_calls") or []) return result.get("output") == expected.get("output") def evaluate(case, k): """Run one case k times. Returns the per-case record.""" runs = [] for _ in range(k): started = time.monotonic() result = run_system(case) elapsed_ms = int((time.monotonic() - started) * 1000) det = deterministic(case, result) crit = {c: judge(c, case, result) for c in case.get("criteria", [])} \ if det is not False else {} passed = bool(det is not False and all(crit.values()) if crit else det) runs.append({ "passed": passed, "deterministic": det, "criteria": crit, "ms": elapsed_ms, "tokens_in": result.get("tokens_in"), "tokens_out": result.get("tokens_out"), "model": result.get("model"), }) return { "id": case.get("id"), "bucket": case.get("bucket"), # pass@1 is the honest headline; pass^k is what production experiences. "passed": runs[0]["passed"], "pass_hat_k": all(r["passed"] for r in runs), "pass_at_k": any(r["passed"] for r in runs), "runs": runs, } def main(): ap = argparse.ArgumentParser(description="Minimal golden-set eval runner.") ap.add_argument("golden") ap.add_argument("-k", type=int, default=1, help="runs per case (3 to report pass^k)") ap.add_argument("--out", default="results.jsonl") ap.add_argument("--history", default="history.jsonl") ap.add_argument("--dataset-version", default="unversioned", help="never compare scores across dataset versions") args = ap.parse_args() cases = [json.loads(l) for l in open(args.golden, encoding="utf-8") if l.strip() and not l.startswith("#")] results = [evaluate(c, args.k) for c in cases] with open(args.out, "w", encoding="utf-8") as fh: for r in results: fh.write(json.dumps(r) + "\n") n = len(results) summary = { "dataset": args.dataset_version, "judge": JUDGE_MODEL, "n": n, "k": args.k, "score": round(sum(r["passed"] for r in results) / n, 4), "pass_hat_k": round(sum(r["pass_hat_k"] for r in results) / n, 4), "tokens_in": sum(run["tokens_in"] or 0 for r in results for run in r["runs"]), "tokens_out": sum(run["tokens_out"] or 0 for r in results for run in r["runs"]), "p95_ms": sorted(run["ms"] for r in results for run in r["runs"])[int(0.95 * n * args.k)], # Stamp `date` from CI (git commit date / job start), not from the runner, # so a re-run of an old commit does not claim to be today's measurement. } with open(args.history, "a", encoding="utf-8") as fh: fh.write(json.dumps(summary) + "\n") print(json.dumps(summary, indent=2)) if __name__ == "__main__": main() -
golden-set.example.jsonl 5.9 KB · in bundle
-
judge-rubric.template.md 3.4 KB
# Judge Rubric — template Copy per criterion. **One criterion per rubric**, one rubric per judge call. A rubric asking "is it accurate, helpful and well-written?" returns an unactionable blend; three separate binary questions return three actionable answers. Adapt everything in `<angle brackets>`. Delete the guidance comments before use. --- ## Metadata (keep with the rubric, in git) | Field | Value | |---|---| | Criterion name | `<faithfulness>` | | Judge model | `<pinned-model-version>` — never "latest" | | Temperature | `0` | | Output | binary `pass` / `fail` | | Rubric version | `<v1>` — bump on any edit, and re-baseline | | Calibrated | `<kappa, date, n>` — from `judge-calibration.py` | > A rubric edit is a re-baselining event, exactly like a judge upgrade or a > dataset version bump. Scores across the boundary are not comparable. --- ## The prompt ``` You are grading one criterion. Answer only about this criterion; ignore every other quality of the response. CRITERION <Every factual claim in the response appears in the SOURCE below.> <!-- Concrete and checkable. "is accurate" is not a criterion, it is a mood. --> SOURCE <<< {source} >>> RESPONSE <<< {response} >>> RULES - Length is not evidence of quality. A correct, brief response scores the same as a correct, padded one. <!-- verbosity-bias counter; measure whether it works with judge-calibration.py --verbosity-field --> - Style, tone and formatting are NOT part of this criterion. <!-- separates correctness from style so the gate rides on correctness --> - Alternative correct answers are acceptable. Do not penalise a response for differing from any reference you may infer. <!-- anchoring counter --> - If the evidence is genuinely ambiguous, answer "fail" and say why. <!-- forces uncertainty to resolve one way, deliberately; pick the direction that is safe for YOUR gate -- see the note below --> EXAMPLES pass: <a short worked example of a response that satisfies the criterion> fail: <a short worked example that plausibly looks fine but does not> fail: <a second failing example covering a different failure mode> <!-- two or three worked failures do more for agreement than a page of prose --> Think through the evidence first, then answer. Return JSON only: {"reason": "<one sentence citing the specific evidence>", "verdict": "pass" | "fail"} ``` --- ## Which way should ambiguity resolve? Decide deliberately, per criterion, and write it down: | Gate consequence | Resolve ambiguity to | |---|---| | A false pass ships a bug | `fail` — under-passing is safe; the judge only over-reports work | | A false fail blocks a good PR and erodes trust in the gate | `pass`, and keep the metric advisory until κ improves | Read the direction off the confusion matrix after calibration, not off intuition: a judge at κ 0.65 that errs only toward `fail` is gateable; the same κ erring toward `pass` is not. --- ## Before this rubric gates anything 1. Sample 50–200 cases, stratified across buckets **and across the judge's own verdicts** — include cases it passes and cases it fails, or you can only measure one error direction. 2. Label them by hand against this exact rubric. 3. `python3 scripts/judge-calibration.py labels.jsonl --min-kappa 0.6` 4. Below 0.6, the rubric is the problem. Rewrite the criterion or add worked failure examples — do not reach for a bigger judge model. 5. Record κ, n and the date in the metadata table above.
-
-
references
-
adversarial-verification.md 5.2 KB
# Adversarial Verification — refute, do not confirm Scoring answers a graded question ("how good is this?"). Verification answers a binary one ("is this finding real?"). The second is where a naive prompt does the most damage, because agreeing is the path of least resistance for a language model. ## The inversion Compare two prompts over the same finding: ``` ❌ "Is this bug report correct?" ✅ "Try to refute this bug report. Default to refuted if uncertain." ``` The first invites confirmation and gets it — plausible-sounding findings sail through because nothing in the prompt rewards saying no. The second makes the model argue against the finding, so a finding only survives if the refutation attempt genuinely fails. The asymmetry is deliberate: **uncertainty must resolve to "refuted", not "confirmed".** A verifier that cannot decide has not verified anything, and treating that as a pass is how plausible-but-wrong results reach a report. Prompt shape that works: ``` Finding: <claim, with file:line or a concrete reproduction> Your job is to REFUTE this finding. Look for reasons it is wrong, already handled elsewhere, unreachable in practice, or based on a misreading. If you cannot decisively refute it, say so — but default to refuted when the evidence is ambiguous. Return {"refuted": true|false, "reason": "..."}. ``` ## Majority-refute thresholds One refuter is a coin flip with an opinion. Run N and take a majority: | N | Keep the finding if | Character | |---|---|---| | 1 | not refuted | Cheap triage only | | **3** | **at least 2 fail to refute** | The default. Good precision/cost balance | | 5 | at least 3 fail to refute | High-stakes; noticeably slower | Tighten toward unanimity when a false positive is expensive (a finding that will be posted publicly, or acted on automatically); loosen when a false negative is expensive (a security sweep where missing a real issue costs more than investigating a dud). Record the vote, not just the verdict — a 2/3 survival is materially weaker evidence than 3/3 and the consumer of the report deserves to know which they have. ## Lens diversity beats redundancy Three identical refuters share their blind spots: whatever the first one fails to notice, the other two also fail to notice. Give each refuter a **different lens** and their failure modes stop overlapping: | Lens | Asks | |---|---| | **Correctness** | Is the described behaviour actually what the code does? | | **Reachability** | Can this state be reached by any real input, or is it guarded upstream? | | **Reproduction** | Given the stated inputs, does the described failure actually occur? | | **Prior art** | Is this already handled — a caller-side check, a test, a framework guarantee? | | **Security** (where relevant) | Is there an exploit path, or is this only a code smell? | Same reasoning as judge panels (`llm-judge.md`): diversity finds classes of error that redundancy structurally cannot. ## Where this fits in a pipeline The canonical shape — find wide, verify hard, keep little: ``` find (N parallel finders, different angles) → dedupe against everything seen so far ← plain code, not an agent → refute (3 diverse lenses per finding) → keep majority survivors → repeat until K consecutive rounds find nothing new ``` Two details that decide whether this converges: - **Dedupe against `seen`, not against `confirmed`.** If refuter-rejected findings are not added to `seen`, the finders resurface them every round and the loop never terminates. - **Loop until dry, not until a count.** "Find 10 bugs" stops at 10 whether or not there are 11; "stop when two consecutive rounds surface nothing new" finds the tail. The fan-out mechanics — process isolation, worktrees, journals, model selection per stage — belong to the parallel-work skills, not here. See `fleet-ops` for the landing discipline and `parallel-ops` as the router for that family. This skill owns the *contract*: what a refuter is asked, what it returns, and how votes become a verdict. ## Verifier output contract Keep it small and structured so votes aggregate mechanically: ```json {"refuted": true, "confidence": "high", "lens": "reachability", "reason": "The caller validates `id` against the allowlist at handler.ts:41 before this path."} ``` - `refuted` boolean, never a score — the whole point is a decisive vote. - `reason` mandatory and specific. A refutation without a citable reason is an opinion, and should be treated as a non-refutation when you audit the run. - `lens` recorded so you can later ask which lens is earning its cost. Lenses that never refute anything across many runs are candidates for removal. ## When not to bother Adversarial verification costs N× per finding. Skip it when: - The finding is **mechanically checkable** — run the test, run the type-checker, run the query. A deterministic check beats any number of refuters. - The finding is **cheap to act on and cheap to revert** — a lint fix does not need a tribunal. - You are **scoring quality, not adjudicating truth**. Use a judge (`llm-judge.md`); refuters answer a binary question and will flatten a graded one. ## Cross-reference - Graded scoring instead of binary adjudication: `llm-judge.md` - Using survival rates as an eval metric over time: `regression-gating.md` -
annotation-workflow.md 6 KB
# Annotation Workflow — where the human labels come from `llm-judge.md` says "label 50–200 cases by hand" and moves on. This is that step. It is the least glamorous part of an eval stack and the one that determines whether every number downstream means anything. ## The ceiling nobody measures **A judge cannot beat the agreement two humans achieve with each other.** If your annotators agree with one another at κ 0.65, a judge scoring κ 0.65 against one of them is performing *at the human ceiling*, and chasing 0.8 is chasing noise in your own labels. So measure the ceiling first: 1. Have **two people independently label the same 30–50 cases**, blind to each other. 2. Compute human-human κ (`judge-calibration.py` does not care which rater is which — feed it `{"human": rater_a, "judge": rater_b}`). 3. That number is your realistic target, and the diagnosis when it is low. **Human-human κ below ~0.6 means the rubric is ambiguous, not that your annotators are bad.** Fix the rubric before labelling another case — every label produced against an ambiguous rubric is wasted work, and a judge calibrated against it will inherit the ambiguity as apparent noise. This step is skipped almost universally, and it is why so many teams conclude "LLM judges are unreliable" when what they actually have is an under-specified criterion. ## Who labels | Labeller | Good for | Watch out for | |---|---|---| | **The engineer who built it** | Fast bootstrapping, catching obvious breakage | Knows what the system *meant* to do and scores it charitably. Never the sole labeller for a gate | | **A domain expert** | Anything where correctness is a matter of policy, medicine, law, finance | Scarce and expensive — spend their time on the ambiguous cases, not the obvious ones | | **A second engineer** | The human-human ceiling measurement | Shares the team's blind spots | | **Crowdsourced** | High-volume, low-context judgments | Needs a much tighter rubric and gold-standard trap questions | The practical shape for most teams: **the engineer labels everything, a domain expert labels the 30 hardest, and the disagreements between them are the most valuable output of the whole exercise** — each one is either a rubric ambiguity or a genuine product decision nobody had made yet. ## Sampling: do not label a random slice A uniform random sample of production traffic is mostly easy cases, and gives you a calibration set that cannot measure the judge where it matters. **Stratify across two axes:** 1. **Bucket** — production, replay, adversarial, edge (`golden-datasets.md`). 2. **The judge's own verdict** — include cases it passes *and* cases it fails, in meaningful numbers. Sampling only cases the judge failed measures one error direction and leaves you blind to false passes, which are the dangerous kind. Add a third axis where you have one: cases near the judge's decision boundary (low-confidence, or where a rerun flipped the verdict) are worth several times an unambiguous case each. ## Running a labelling session - **Label blind to the judge's verdict.** Showing it first anchors the human onto it and inflates apparent agreement — you will measure compliance, not agreement. - **One criterion at a time, across all cases.** Labelling case-by-case across five criteria drifts; labelling criterion-by-criterion keeps the standard fixed. - **Record the reason, not just the verdict.** A label without a reason cannot be audited later, and reasons are what you mine to rewrite an ambiguous rubric. - **Timebox and batch.** Annotation quality falls off a cliff past ~45 minutes. 50 cases in two sittings beats 100 in one. - **Keep an explicit `unsure` option** — and then *do not* let unsure labels vote. Forcing a binary on a genuinely ambiguous case manufactures noise that looks like judge error. A high unsure rate is a rubric finding. ## Adjudication Disagreements are the product, not a problem to be averaged away. 1. Both labellers state their reason. 2. Classify the disagreement: - **Rubric ambiguity** → rewrite the criterion, add a worked example, re-label the affected cases. - **Genuine product ambiguity** ("should the agent refuse this?") → escalate. A decision gets made, and it becomes a rubric example. - **Simple error** → correct it and move on. 3. **Never resolve by majority vote without reading the reasons.** A 2–1 split where the minority is right is exactly the case that most improves the rubric. ## Keeping labels fresh Labels rot in three ways, and each has a different tell: | Rot | Tell | Fix | |---|---|---| | **Judge drift** | κ falls with no rubric change | The judge model version moved. Pin it; re-calibrate | | **Distribution drift** | Judge κ holds on the calibration set but production complaints rise | The calibration set no longer resembles traffic. Re-sample | | **Annotator drift** | The same person labels the same case differently months apart | Real, and normal. Include ~10% repeats from earlier sessions to measure it | That last trick is worth adopting: silently re-include a handful of previously labelled cases in each session. Intra-annotator agreement below the inter-annotator ceiling means the standard itself is sliding, and no amount of judge tuning will fix it. **Cadence:** re-sample ~50 fresh cases periodically, and *always* re-calibrate on a judge model change or a rubric edit — both are re-baselining events (`regression-gating.md`). ## Cost, honestly At roughly 1–3 minutes per case per criterion, 150 cases × 2 criteria × 2 annotators is around 10–15 person-hours to establish a calibrated judge. That is the real price, and it is worth naming up front — a team that budgets for "run the calibration script" and not for the labelling will quietly skip the labelling and gate CI on an uncalibrated judge. The saving is that it is mostly one-off. Maintenance is ~50 cases periodically, which is an hour or two. ## Cross-reference - What the labels are for: `llm-judge.md` - Which cases to draw from: `golden-datasets.md` - The script that consumes them: `scripts/judge-calibration.py` -
eval-taxonomy.md 5.4 KB
# Eval Taxonomy — outcome, step, trajectory The three levels an agent can be scored at, what each one catches that the others miss, and how to pick metrics that survive contact with a non-deterministic system. ## The three levels ### Outcome-level **Question:** is the final artifact correct? The cheapest and most defensible level, and the one everybody starts with. SWE-bench established binary pass/fail on the produced patch as the standard for coding agents; the same shape works for any agent with a checkable end state — a booking exists, the file parses, the database row has the right value. Evaluate deterministically where you can: ```python assert json.loads(out)["status"] == "confirmed" assert db.query("SELECT count(*) FROM orders WHERE id=?", oid) == 1 ``` **What it misses:** *how* the answer was reached. Which is most of what determines whether it will be reached again. ### Step-level **Question:** was this individual action right? Scored per span: was the right tool chosen, was the argument schema valid, were the argument *values* correct, did the agent read before it wrote. Step-level scoring is where you catch the agent that calls `search` five times with near-identical queries, or that passes a plausible-but-wrong ID. Most step assertions are deterministic and belong in the blocking tier: | Assertion | Cost | |---|---| | Tool `X` was called at least once | free | | Every tool call validated against its JSON schema | free | | No tool called with a value absent from the input context (hallucinated arg) | free | | Read-before-write ordering held | free | Reserve a span-level judge for the genuinely fuzzy step questions ("was this a reasonable query to issue given what the agent knew?"). ### Trajectory-level **Question:** was the *path* sensible? Scored over the whole nested span tree: sequence, redundancy, loop detection, recovery after an error, total cost. Two forms: 1. **Reference-trajectory match** — compare against a known-good path. Exact-match is too brittle for anything real; use an ordered-subsequence match ("these 4 steps appeared in this order, extras allowed") or set-containment on the essential tool calls. 2. **Rubric judge over the trace** — hand the serialized trajectory to a judge with a rubric ("did it loop? did it retry the same failing call? did it ask the user something it could have looked up?"). Trajectory scores are diagnostics, not gates. A novel correct path scores badly against a reference and is not a regression. ## The lucky pass The single argument for scoring more than the outcome. An agent reaches the correct end state by an accidental route — it guessed an ID that happened to be right, it retried until a flaky tool succeeded, it hard-coded something that matched this case's expected output. Outcome-only scoring banks that as a pass. The same case fails next week, and because the suite was green the whole time, nobody knows when the real breakage started. **Detection is cheap once you have traces:** a passing case whose trajectory contains a loop, an error-then-retry, or a tool call with an argument that appears nowhere in the input is a lucky-pass candidate. Flag them; do not fail on them. A `lucky_pass_suspects` count trending upward on a green suite is one of the highest-value signals in the harness. ## pass@k vs pass^k For any non-deterministic agent, a single run per case is a coin flip you are reporting as a measurement. | Metric | Definition | What it tells you | |---|---|---| | **pass@1** | One run, did it pass | The honest headline number | | **pass@k** | Any of k runs passed | Ceiling / "is this reachable at all" | | **pass^k** | *All* k runs passed | Consistency — what production actually experiences | pass@k flatters. An agent that succeeds 1 time in 4 has pass@4 near 1.0 and is unusable. **Report pass^k whenever consistency matters** (customer-facing, transactional, anything where a retry costs the user something). The tau-bench family popularised this framing for multi-turn tool-using agents and it generalises. Practical: k=3 is usually enough to expose the difference and cheap enough to run per-PR on a subset. Run the full k on a nightly, a k=1 pass on every PR. ## Choosing metrics Start from the failure you actually fear, not from a metrics catalog. | You fear | Measure | Level | |---|---|---| | Made-up facts | Faithfulness: every claim traceable to a retrieved chunk | Outcome (judge) | | Wrong tool / wrong args | Tool-call accuracy, arg-schema validity | Step (deterministic) | | Burning tokens | Steps per task, redundant-call rate, cost per case | Trajectory (deterministic) | | Silent policy violations | Policy-compliance rubric over the trace | Trajectory (judge) | | Flaky success | pass^3 | Outcome (deterministic, repeated) | | Broke something that used to work | Failure-replay bucket pass rate | Outcome (deterministic) | Two rules that save more time than any metric choice: - **A metric nobody can act on is a metric nobody will maintain.** If a score drops and the team cannot name what to change, delete the metric or make it decomposable. - **Deterministic first.** Every criterion you move from a judge into code removes cost, latency, and variance simultaneously. Re-audit periodically: rubric items often become codifiable once the output format stabilises. ## Cross-reference - Dataset construction: `golden-datasets.md` - Judge design for the fuzzy levels: `llm-judge.md` - Which of these gate CI: `regression-gating.md` -
golden-datasets.md 6.8 KB
# Golden Datasets — build it, freeze it, keep it from rotting The golden set is the most valuable artifact in an eval stack. The model, the prompt and the framework will all be replaced; the dataset outlives them and is the only thing that lets you compare across the replacements. ## What it is A **reviewed, versioned, frozen** set of inputs paired with trusted expected outputs (or trusted grading criteria, where the output is open-ended). "Reviewed" means a human agreed the expected output is right. "Versioned" means it lives in git next to the code. "Frozen" is the part teams skip and then regret. ## Why frozen beats growing A set that grows every sprint cannot answer the only question a regression suite exists to answer: *did my change make things worse?* If the score moved from 0.84 to 0.79 and the set gained 30 cases, the two variables are confounded and no amount of analysis separates them. The discipline: - **Freeze v1.** Tag it. Every run reports against a named dataset version. - **New cases go to a staging set** and are promoted in explicit, dated batches. - **On promotion, re-baseline** — run the current system against v2 and record the new reference number. Never compare a v1 score to a v2 score. - **Never edit a case in place to make it pass.** That is fitting the test to the code. If a case's expected output was genuinely wrong, delete it and add a new one with a note. Ship a freeze manifest so drift is detectable rather than discovered: ```bash python3 scripts/goldenset-audit.py golden.jsonl --write-freeze manifest.json # at freeze python3 scripts/goldenset-audit.py golden.jsonl --freeze manifest.json # in CI ``` ## Four-bucket composition A set sampled only from happy-path production traffic tells you nothing about the failures you will actually ship. Build it in four deliberate buckets and record the bucket on every case. | Bucket | Target share | Source | Catches | |---|---|---|---| | `production` | 40-50% | Stratified sample of real traffic | Drift on the common path; keeps the score meaningful | | `replay` | 20-30% | Every incident that reached a human | Regressions on things that already broke once | | `adversarial` | 15-20% | Injections, contradictory instructions, refusal-bait, out-of-scope asks | Silent policy failures; the class judges also fail on | | `edge` | 10-15% | Empty, enormous, ambiguous, multilingual, malformed | Where deterministic code breaks first | Those are **targets to compose against**, not thresholds. `goldenset-audit.py` warns on a deliberately wider band (production 30-65%, replay 10-40%, adversarial 8-35%, edge 5-30%) and names the target in the warning, because a check that fires on every healthy set gets ignored within a week - the same rule this skill applies to CI gates. Hitting the target is good practice; leaving the band is a finding. **The replay bucket is the easiest to justify and the most neglected.** Every production incident is a free, pre-validated, maximally relevant test case. Make "add the replay case" a step in the incident checklist and the bucket fills itself. Watch the balance: buckets skew over time because production sampling is easy and adversarial authoring is not. A set that has become 90% `production` has quietly stopped testing the things that break. The audit script flags this. ## Sizing | Stage | Size | Note | |---|---|---| | Bootstrapping | 20 | Hand-written, eyeballable. Do this before choosing a metric. | | Working regression set | 100-300 | Enough to move a percentage meaningfully; cheap enough to run per PR | | Mature, production-sampled | 200-500 | The common steady state for a real product | | Judge calibration subset | 50-200 human-labelled | Separate purpose — see `llm-judge.md` | Past ~500 you are usually buying latency, not signal. Add cases when a **new failure class** appears, not on a cadence. If two cases fail and pass together every time, one of them is free to delete. Cost check: at 300 cases × 3 runs (pass^3) × $0.01/case you are spending ~$9 a run. That is fine nightly and painful on every push — which is why the PR tier runs a subset. See `regression-gating.md`. ## Case schema Nothing exotic. JSONL, one case per line, in git: ```json {"id": "refund-partial-001", "bucket": "replay", "added": "2026-03-14", "why": "INC-482: agent refunded full amount on a partial-return request", "input": {"messages": [{"role": "user", "content": "..."}]}, "expected": {"tool": "issue_refund", "args": {"amount": 24.99}}, "criteria": ["refund amount matches the returned item only", "does not promise a timeline the policy does not state"]} ``` The fields that matter and get omitted: - **`why`** — the reason this case exists. Without it, a future maintainer deletes cases they cannot interpret, and the set silently loses its adversarial teeth. - **`added`** — dates are how you detect a set that stopped growing in 2025. - **`bucket`** — without it you cannot see the balance drifting. - **`criteria`** — for open-ended outputs, the grading rubric belongs *with the case*, not in a global judge prompt. Case-specific criteria are dramatically easier to calibrate. ## Rot, and how it shows up | Rot | Symptom | Fix | |---|---|---| | **Duplicates / near-duplicates** | Score moves in suspiciously large jumps | Dedupe on normalised input; audit script flags exact and high-overlap pairs | | **Bucket skew** | Adversarial pass rate stops moving | Rebalance; author new adversarial cases | | **Staleness** | Newest `added` date is months old | Wire the incident checklist; sample fresh production traffic | | **Saturation** | Score pinned at 1.0 for weeks | The set is too easy. Harvest harder cases from production; a saturated set detects nothing | | **Contamination** | Score jumps on a model upgrade with no code change | Public benchmark cases leaked into training. Prefer private, product-specific cases | | **Test-fitting** | Cases edited in commits that also change the prompt | Enforce in review: dataset changes land in their own commit | Saturation deserves emphasis: a suite that always passes is not a passing suite, it is a suite that has stopped measuring. Track the *distribution* of per-case results, not just the mean, and retire-and-replace cases that have not failed in months. ## Synthetic cases Useful for edge and adversarial coverage, dangerous as the backbone. Synthetic inputs generated by the same model family you are testing inherit its blind spots — it will not generate the phrasing it does not understand. Use synthetics to *expand* a bucket around a real failure ("give me 10 variants of this injection"), never to found one. Always human-review synthetic expected outputs before promotion. An unreviewed synthetic case is a hypothesis, not a golden case. ## Cross-reference - What to score these cases on: `eval-taxonomy.md` - Grading the open-ended ones: `llm-judge.md` - Running them in CI: `regression-gating.md` -
hillclimbing.md 9.8 KB
# Hillclimbing — optimising against an eval without destroying it > **Verified 2026-08.** The optimizer landscape below moves fast. Method names, > reported numbers and library APIs need re-verification before you quote them. Once you have a measurable harness, the obvious next move is to optimise against it: change something, measure, keep if better, repeat. That loop works, and it is also the single most reliable way to turn a good eval suite into a useless one. **This file owns the discipline. It does not own the loop.** The loop mechanics — scope, verify command, keep/discard, batch with bisect-on-regression, stop conditions, git as memory — belong to the [`iterate`](../../iterate/SKILL.md) skill, which is domain-agnostic and does not care whether your metric is test coverage or an eval score. Read that for *how to run* the loop; read this for *what goes wrong* when the metric is an eval. ## The two failure modes ### 1. Banking noise `iterate` keeps a change when the metric beats the previous best. With a deterministic metric (line coverage, bundle bytes) that rule is exactly right. With an eval score it is a coin flip dressed as a decision. If your suite scores 0.88 ± 0.03 across reruns of *unchanged code*, then a change that measures 0.90 has told you almost nothing. Keep it anyway and you have banked noise. Do that fifty times overnight and you have executed a random walk with perfect discipline, and the winning commit is whichever iteration got luckiest. **The rule: a hillclimb step is only real if the delta clears the noise floor.** ```bash python3 scripts/eval-baseline.py history.jsonl --candidate iter.jsonl --accept # exit 0 = KEEP (improvement clears the floor, or is significant on the paired test) # exit 10 = DISCARD (inside the noise band, or a regression) ``` `--accept` inverts the usual exit semantics on purpose — see the script header. Wire it as the keep/discard gate and the loop stops rewarding luck. Two cheap amplifiers, when you can afford them: - **Increase k before increasing iterations.** Three runs per candidate shrinks the noise floor and is usually a better spend than three times as many candidates evaluated once. - **Paired comparison beats aggregate comparison.** Which specific cases flipped is far more informative than a delta of 0.01, and it is free once you store per-case results (`regression-gating.md`). ### 2. Overfitting the golden set This one is slower, quieter, and permanent. Every look at the same frozen set leaks a little information into your decisions. Optimise against it for long enough and you have tuned the system to that specific 300 cases — Goodhart's law arriving exactly on schedule. It is not hypothetical in the optimizer literature: **GEPA is documented to overfit by encoding edge cases into increasingly verbose prompts**, because its reflection step accumulates detail across iterations and nothing pushes back. Length constraints act as regularisation there, which is a good general instinct: an optimizer with no pressure toward simplicity will buy training score with complexity. **Split the data before you optimise anything:** | Split | Used for | Rule | |---|---|---| | **Train / reflect** | What the optimizer sees and reasons about | Look freely. This is the set you burn | | **Validation** | Choosing between candidates each round | Scored every round; never shown to the optimizer's reflector | | **Held-out test** | The number you actually report | Touch at milestones only. Every look costs you | GEPA's own design makes the second row explicit: it reserves a disjoint validation subset whose inputs and outputs are **never shown to the reflector model**. Adopt that separation even when hand-rolling — the moment the thing proposing changes can read the set that judges them, your validation score stops being evidence. **The held-out set is a budget, not a dashboard.** Decide up front how often you may look (a milestone, a release, once a week) and hold to it. A team that checks held-out every iteration has three sets and one of them is a validation set wearing a disguise. **Tells that you are overfitting:** - Train score climbs; validation is flat. The classic. - Validation climbs; held-out is flat. You are now overfitting the validation set. - The system prompt or config grows monotonically, each addition patching one case. - Wins stop transferring — a change that helped on the suite does nothing in production. ## Keep a frontier, not a champion `iterate` maintains `iterate/best`: a single floating tag on the highest-metric commit. For a scalar mechanical metric that is correct and simple. For an eval score it is a local-optimum trap, because one aggregate number hides which *cases* a candidate won. GEPA's central design choice is the alternative: **maintain a Pareto frontier** — retain every candidate that is best on at least one validation instance, and sample from that frontier rather than always mutating the current champion. A candidate that scores lower overall but is the only one solving a hard case carries information the champion does not, and discarding it is how a hillclimb walls itself into a local optimum. Cheap approximation without any framework: alongside the best-overall commit, keep a note of **which candidate best solved each failing case**. When the loop stagnates, mutate from one of those instead of from the champion. ## Rich feedback beats a scalar reward The most transferable finding in this line of work: **collapsing an evaluation to a number throws away the signal an optimizer needs.** GEPA reflects in natural language over execution traces — error messages, reasoning logs, why the case failed — and reports outperforming MIPROv2 by roughly 10–13% and the RL baseline GRPO by ~6% on average (up to 20%) across six tasks, **using up to 35× fewer rollouts**. It is sample-efficient enough to work from as few as 10 examples and 20–100 evaluations. The number to take from that is not the benchmark delta, which will age. It is the mechanism: a failure that explains itself is worth many failures that only score. **This is an eval-design consequence, and it is why it lives in this skill.** If your judge returns `0.4`, an optimizer has nothing to reflect on. If it returns `{"reason": "cited chunk 12, which does not contain the refund window", "verdict": "fail"}`, it has a diagnosis. `assets/judge-rubric.template.md` already demands `reason` alongside `verdict` — that field is what makes hillclimbing tractable later, and it costs nothing to add now. Same for per-case traces: store them, or you cannot reflect on them. ## The optimizer landscape Families rather than products, since the products churn: | Family | Mechanism | Note | |---|---|---| | **Reflective / evolutionary** (GEPA) | Natural-language reflection over traces; Pareto frontier over candidates | Sample-efficient; available in DSPy as an optimizer and standalone. Documented verbosity-overfit failure mode | | **Bayesian / few-shot search** (MIPROv2) | Proposes instructions and demonstrations, searches the joint space | The prior DSPy default; the baseline GEPA is measured against | | **Generate–score–select** (APE) | Generate candidate prompts, score on validation, keep the best | Simplest thing that works; a fine hand-rolled starting point | | **Iterative self-rewrite** (ORPO and kin) | The model rewrites its own prompt guided by feedback on prior outputs | Needs a feedback signal richer than a score | | **Self-contained preference loops** (SPO) | Generates its own data, refines by pairwise preference over its outputs | Removes the external-label dependency — and with it, your ground truth. Treat results with suspicion | | **Scaffold-level** | Memory evolution, tool governance, whole-harness redesign | The wider "self-improving agent" framing; prompt optimization is one lever among several | **When it helps is a live research question, not a settled one.** There is published work specifically asking *when* prompt optimization improves multi-agent systems — the framing implies the honest answer is "sometimes". Do not assume an optimizer will beat a careful human rewrite on your task; measure it, on held-out, like anything else. ## Before you automate the loop - [ ] A frozen golden set with train / validation / held-out splits (`golden-datasets.md`). - [ ] A measured noise floor, from reruns of unchanged code (`regression-gating.md`). - [ ] A calibrated judge, if a judge is in the metric (`llm-judge.md`). An uncalibrated judge in a hillclimb optimises the system toward the judge's biases, at speed. - [ ] Per-case results and reasons stored, not just an aggregate. - [ ] A stated held-out look budget, and a stop condition. - [ ] A cost ceiling. Optimizer loops are the easiest way to spend a month of eval budget in an afternoon. If any of those is missing, fix it before running the loop. A hillclimb amplifies whatever your measurement already is — including its errors. ## When this should become its own skill Kept here because it is currently one file, and the creation protocol says extend rather than duplicate. **Extract it to a `prompt-optimization-ops` skill when the optimizer material outgrows this page** — concretely, when it needs its own worked DSPy/GEPA configuration, more than a couple of runnable scripts, or per-optimizer troubleshooting. The discipline sections above (splits, noise floor, frontier, held-out budget) stay here regardless: they are properties of the measurement, and this skill owns measurement. ## Cross-reference - Loop mechanics, stop conditions, bisect-on-regression: [`iterate`](../../iterate/SKILL.md) - Scheduling a loop across sessions, risk tiers, kill switch: [`loop-ops`](../../loop-ops/SKILL.md) - Splits and freeze discipline: `golden-datasets.md` - Noise floor and the paired test: `regression-gating.md` - Why a judge must be calibrated before it steers anything: `llm-judge.md` -
llm-judge.md 7.7 KB
# LLM-as-a-Judge — biases, rubrics, panels, calibration A judge is a measurement instrument built out of a language model. It is useful, it is often the only option, and it is biased in documented, reproducible ways. Treat it like an instrument: know its error modes, calibrate it against a reference, and re-check it when anything underneath changes. ## Decide whether you need one at all | Criterion | Evaluator | |---|---| | Output must parse / match a schema | `json.loads`, a JSON-Schema validator | | Exact value, ID, amount, tool name | `==` | | Contains a required citation / does not contain a banned string | regex | | Ordering, latency, cost, step count | arithmetic over the trace | | **Faithfulness to a source** | judge | | **Policy / tone compliance** | judge | | **"Is this a reasonable answer to an open question"** | judge | | **Relative quality of two candidates** | judge (pairwise, with position control) | Every criterion you move out of the judge and into code removes cost, latency *and* variance in one edit. Re-audit the rubric periodically — items become codifiable once the output format stabilises, and nobody goes back to check. ## The bias catalog ### Position bias In pairwise comparison, judges prefer whichever candidate was presented first — strongly enough that swapping the order flips a meaningful share of verdicts. **Mitigations, in order of preference:** 1. **Score absolutely, not pairwise.** Each candidate graded against the rubric alone. Kills the bias by construction and makes results comparable across runs. 2. **Both orders, averaged.** If you need pairwise, run A/B and B/A and require agreement. Doubles cost; the disagreement rate is itself a useful instability metric. 3. Never accept a single-order pairwise verdict as a gate. ### Verbosity bias Judges rate longer answers higher regardless of quality. Evidence is heterogeneous — some models are genuinely quality-sensitive and penalise filler — which is exactly why you must measure it on *your* judge rather than assume. **Mitigations:** - Split the rubric: score *correctness* and *style* separately, and gate on correctness. - State the anti-bias instruction explicitly ("length is not evidence of quality; an answer that is correct and brief scores higher than one that is correct and padded"). - **Probe it.** Correlate judge score against output length on your calibration set. A strong positive correlation on cases with equal human scores is the bias, quantified. `scripts/judge-calibration.py --verbosity-field length` computes this. ### Self-preference bias Judges rate outputs from their own model family higher. Fatal when you are comparing models and using one of the contenders as the judge. **Mitigation:** judge with a different family from the system under test. If that is impossible, at minimum report the judge model alongside every score and never compare scores produced by different judges. ### Scale drift and clustering 1-5 scores cluster (almost everything gets a 4), and the cluster shifts when the judge model version changes — silently re-baselining your whole history. **Mitigations:** - **Binary pass/fail against explicit criteria** wherever the decision is genuinely binary. Easier to calibrate, easier to act on, far more stable across model versions. - **Pin the judge model version** and treat a judge upgrade like a dataset version bump: re-baseline, do not compare across it. - If you need a scale, define each point with a concrete example, not an adjective. ### Other effects worth knowing | Effect | Note | |---|---| | **Sycophancy toward the prompt** | A judge asked "confirm this is correct" confirms. Phrase neutrally, or invert — see `adversarial-verification.md` | | **Format preference** | Markdown/bulleted answers score above equivalent prose. Normalise formatting before judging where you can | | **Anchoring on the reference** | Given a reference answer, judges penalise correct-but-different. Say explicitly that alternative correct answers are acceptable | ## Rubric design The rubric is where most judge quality lives — far more than the model choice. - **One criterion per question.** A rubric asking "is it accurate, helpful and well-written?" returns an unactionable blend. Three separate binary questions return three actionable answers. - **Concrete and checkable.** "Every factual claim appears in the provided source" beats "is accurate". - **Case-specific criteria where possible.** Criteria stored with the golden case (`criteria: [...]`) calibrate dramatically better than one global rubric stretched over a heterogeneous set. - **Require reasoning before the verdict.** Chain-of-thought judging (the G-Eval line of work) improves human agreement; and the reasoning is what lets you debug a disagreement instead of shrugging at it. - **Demand structured output** — `{"reason": "...", "verdict": "pass"}` — so scoring is parseable and the reason is stored, not discarded. - **Include the failure examples.** Two or three worked examples of what a `fail` looks like do more for agreement than a page of prose. ## Panels vs N-identical Running the same judge three times with the same rubric mostly buys the same bias three times. It measures the judge's *own* variance — worth knowing once, not worth paying for every run. A **panel with distinct lenses** is different in kind: each judge is asked a different question, so their failure modes do not overlap. | Shape | Use when | |---|---| | Single judge, absolute scoring | Default. Cheapest thing that works | | N-identical, same rubric | One-off: measure judge variance to set your CI noise floor | | **Panel, distinct lenses** | The thing can fail in several independent ways (correct? policy-compliant? reproducible?) | | **Panel, distinct model families** | High-stakes scoring; aggregate by majority to damp single-model bias | Majority voting across heterogeneous judges measurably improves correlation with human judgment. It also multiplies cost — reserve it for the scores you gate on. ## Calibration — the step that makes a judge trustworthy Calibration has gone from nice-to-have to table stakes. The method: 1. **Sample 50-200 cases** from the golden set, stratified across buckets and across the judge's own verdicts (include cases it passes *and* fails, or you cannot measure both error directions). 2. **Label them by hand.** Same rubric the judge gets. Two humans on a subset gives you a human-human ceiling — a judge cannot beat the agreement humans achieve with each other, so that number tells you what "good" even means here. 3. **Compute Cohen kappa**, not raw agreement. On an imbalanced set (90% pass) a judge that says "pass" unconditionally scores 90% agreement and is worthless; kappa discounts the agreement you would get by chance and lands it near zero. ```bash python3 scripts/judge-calibration.py labels.jsonl --min-kappa 0.6 ``` | kappa | Reading | |---|---| | >= 0.8 | Strong. Production-ready; safe to gate on with a margin | | 0.6 - 0.8 | Substantial. Usable, keep it advisory or gate loosely | | < 0.6 | The rubric is the problem, not the model. Rewrite before spending more on judges | 4. **Read the confusion matrix, not just the headline.** A judge with kappa 0.65 that is wrong only in the false-*negative* direction is safe for a gate (it under-passes, never over-passes). One with the same kappa that hallucinates passes is not. The two need opposite fixes. 5. **Re-calibrate on a schedule** — roughly 50 fresh cases periodically, and *always* after a judge model version change or a rubric edit. ## Cross-reference - Refuting rather than confirming: `adversarial-verification.md` - Where the labelled cases come from: `golden-datasets.md` - Turning a calibrated judge into a CI gate: `regression-gating.md` -
regression-gating.md 8.4 KB
# Regression Gating — making evals block CI without killing CI An eval gate has exactly one job: stop a regression from merging. It fails at that job in two ways — by letting regressions through, and by going red so often that everyone learns to click merge anyway. The second failure is more common and much harder to reverse. ## The rule **A blocking check must never be flaky.** Once a team has seen three red-for-no-reason eval runs, the gate is socially dead even while it is still technically enforced. Design for that first and for coverage second. ## The tier ladder | Tier | Checks | Gate | Runs on | |---|---|---|---| | **0. Deterministic** | Schema validity, tool-call assertions, exact matches, forbidden strings | **Blocking**, zero tolerance | Every push | | **1. Judge, uncalibrated** | Any new rubric, first few weeks | **Advisory** — comment the delta, never fail | Every PR | | **2. Judge, calibrated** | kappa >= 0.6, variance measured | **Blocking with a margin** below rolling baseline | Every PR | | **3. Consistency** | pass^3 over the full set | **Blocking on the nightly**, advisory on PRs | Nightly | | **4. Cost / latency** | Tokens and p95 per case | **Blocking on an absolute ceiling**, advisory on trend | Every PR | Promotion from tier 1 to tier 2 is an explicit decision backed by a calibration run (`scripts/judge-calibration.py`), not something that happens because a rubric has been around a while. ## The noise floor You cannot set a threshold without knowing how much the suite moves when *nothing changes*. Measure it once, properly: run the unchanged system against the frozen set N times (5 is usually enough) and record the spread. ``` run 1: 0.88 run 2: 0.85 run 3: 0.89 run 4: 0.86 run 5: 0.88 baseline 0.872, spread 0.04 ``` Then gate **below the baseline by more than the spread**: threshold 0.80, not 0.87. A gate inside the noise band fails on identical code, which is the fastest possible route to a dead gate. Keep it honest over time by committing a **rolling window of run results to git** — a small JSON file, appended per run on the main branch: ```json {"date": "2026-08-30", "dataset": "golden-v3", "judge": "<pinned-model-id>", "score": 0.871, "pass_at_1": 0.86, "pass_hat_3": 0.79, "cost_usd": 2.14, "p95_ms": 4180, "n": 287} ``` That file is what turns "today looks bad" into "today is 2.6 spreads below a stable baseline" — and it costs nothing. Treating every run as standalone is what makes teams unable to distinguish noise from regression. Re-baseline (and say so in the commit) on any of: dataset version bump, judge model change, rubric edit, or a deliberate accepted trade-off. ```bash python3 scripts/eval-baseline.py evals/history.jsonl --candidate /tmp/run.jsonl # prints baseline, noise floor, and the threshold your gate should use ``` ## Is the drop real? — the paired test The noise floor tells you whether an aggregate score moved further than it usually does. It does **not** tell you whether the same cases moved, and that is the question you actually care about. Two runs over the same frozen set produce *paired binary outcomes*, and the right tool for those is **McNemar's exact test**. It looks only at the discordant pairs: | | candidate passes | candidate fails | |---|---|---| | **baseline passes** | ignored | **b** — regressions | | **baseline fails** | **c** — fixes | ignored | Cases that behaved identically in both runs carry no information about whether the change helped. Under the null hypothesis b is a coin flip over b+c trials, so the p-value is a tail probability. Use the **exact** test rather than chi-square: eval sets routinely produce b+c under 25, where the approximation misleads. Why this matters more than the aggregate: a change that breaks 8 cases and fixes 7 moves the headline score by 0.01 - invisible against any noise floor - while having silently swapped which 15 things work. The paired view names those 15 cases; a score comparison structurally cannot. Note carefully what the test does and does not say there. 8-vs-7 gives p = 1.0: genuinely indistinguishable from chance, and the tool will correctly call it noise. The value in that run is not the verdict, it is the enumerated `regressed` and `fixed` lists telling you a churn happened at all. Significance answers "did the system get worse"; the lists answer "what moved" - and on a flat score only the second question has an answer worth having. ```bash python3 scripts/eval-baseline.py evals/history.jsonl \ --baseline-results base.jsonl --candidate-results new.jsonl --alpha 0.05 # exit 10 = significant regression, and it names the cases that flipped ``` Three cautions: - **Significance is not magnitude.** With a large set, a trivially small real drop reaches p < 0.05. Read the count of regressed cases, not only the p-value. - **Don't run the test repeatedly until it agrees with you.** Testing every PR against the same baseline is many comparisons; treat a single surprising red as a prompt to look at the named cases, not as proof on its own. - **k > 1 breaks the pairing** unless you collapse each case to one outcome first (pass^k is the usual choice). Feed the collapsed per-case result, not every run. ## CI shape A workable three-tier cadence: | Trigger | Scope | Budget | Gate | |---|---|---|---| | **Every push** | Deterministic assertions on the full set | seconds, $0 | Blocking | | **PR** | Judge metrics on a stratified ~30% subset, k=1 | a few minutes | Per tier ladder | | **Nightly on main** | Full set, k=3, cost and latency recorded, appended to the history file | whatever it costs | Blocking; page on a real drop | Notes that matter in practice: - **Pin everything the score depends on** — judge model version, dataset version, prompt version, temperature (0 for the judge). An unpinned judge model is a silent re-baselining that will be blamed on your code. - **Cache aggressively.** Eval runs re-send near-identical prompts; caching the static prefix cuts the bill substantially and does not change scores. - **Report the delta, not the absolute.** "-0.04 vs main (noise floor 0.03)" is actionable; "0.83" is not. - **Name the failing cases in the CI output.** A gate that says "score dropped" without listing which 6 cases flipped forces a local re-run and gets ignored. - **Never auto-retry a failing eval to green.** Retry-until-pass converts a real regression into a flake report. If you retry, report all attempts. ## Cost and latency attribution Record per case, from day one: | Field | Why | |---|---| | `tokens_in` / `tokens_out` | The unit you actually pay for; also the best proxy for context bloat | | `ms` (wall) and step count | p95 latency is a product requirement, and step count catches loops | | `cost_usd` | Roll up per run so a "small" prompt change that doubles spend is visible immediately | | `model` / `judge_model` | Attribution is meaningless if you cannot tell which model produced the number | Two reasons this is not optional. First, an eval suite is the only place you learn that the accuracy win cost 4x the tokens — production tells you eventually, and much more expensively. Second, retrofitting attribution once the harness exists means touching every runner, every stored result and every dashboard; adding four fields at the start costs nothing. Gate on an **absolute ceiling** (cost per case must not exceed $X, p95 must not exceed Y ms) rather than on the trend. Trend gates fire on noise; ceilings encode a product decision. ## Failure modes to design against | Failure | Symptom | Fix | |---|---|---| | Flaky blocking judge | Reds nobody investigates | Demote to advisory until calibrated; widen the margin past the noise floor | | Threshold inside the noise band | Identical code fails intermittently | Measure the spread; gate below baseline minus spread | | Saturated suite | Green for months, then a production incident | The set stopped measuring — harvest harder cases (`golden-datasets.md`) | | Confounded comparison | Score moved and nobody knows why | Version the dataset; never compare across versions | | Judge upgrade drift | Step change in scores with no code change | Pin the judge; re-baseline explicitly on upgrade | | Test-fitting | Cases edited in the same commit as the prompt | Enforce in review: dataset changes land separately | ## Cross-reference - Case supply and freezing: `golden-datasets.md` - Getting a judge to kappa >= 0.6 so it can be promoted to blocking: `llm-judge.md` - Which metric belongs in which tier: `eval-taxonomy.md` -
retrieval-eval.md 6.4 KB
# Retrieval Eval — scoring RAG without conflating two different bugs Retrieval is the most common thing people build evals for, and the most commonly mis-measured. The mistake is universal: score the final answer, watch it drop, and have no idea whether the retriever failed or the generator did. ## Split the pipeline before you score it A RAG answer passes through two stages that fail independently. Score them separately or you cannot act on either. | | Retrieved the right context | Retrieved the wrong context | |---|---|---| | **Answer correct** | Working as intended | **Lucky** — the model knew it anyway, or guessed. Will fail when the question shifts | | **Answer wrong** | **Generation bug** — chunking, prompt, or model | **Retrieval bug** — embeddings, index, query rewriting | The two off-diagonal cells need opposite fixes, and end-to-end accuracy averages them into one uninterpretable number. Worse, the top-right cell — right answer from wrong context — scores as a *pass* end-to-end and is a latent failure exactly like the lucky pass in `eval-taxonomy.md`. **Minimum viable split:** for every case, record the retrieved chunk ids alongside the answer. That single field turns an opaque score into a 2x2 you can act on. ## Retrieval metrics Retrieval is the one place in the eval stack where **deterministic scoring genuinely dominates** — you have ground-truth chunk ids, so no judge is required. Take the free signal. | Metric | Definition | Use when | |---|---|---| | **Recall@k** | Fraction of relevant chunks that appear in the top k | The headline. If the right chunk is not in the context, nothing downstream can save you | | **Precision@k** | Fraction of the top k that are relevant | Context budget is tight; noise crowds out signal | | **MRR** | Mean of 1/rank of the first relevant chunk | One right answer per query; you care that it ranks high | | **nDCG@k** | Rank-discounted gain over graded relevance | Multiple chunks matter and some matter more | | **Context precision** | Of the context actually passed to the model, how much was used | Diagnosing bloated prompts and cost | **Recall@k is the one to gate on.** Precision failures degrade an answer; recall failures make a correct answer impossible. Measure recall at the k you actually retrieve *and* at a larger k — if recall@20 is high while recall@5 is poor, you have a ranking problem, not an embedding problem, and those are different fixes. ## Building the ground truth The dataset is the hard part, as always (`golden-datasets.md`). What retrieval adds: - **Annotate chunk ids, not passages.** Ids survive re-chunking; quoted text does not. Store the chunk's stable id plus a content hash so a silent re-index shows up as drift rather than as a mysterious recall drop. - **Relevance is graded, not binary,** for anything but the simplest corpus: `2` = answers the question, `1` = useful context, `0` = irrelevant. nDCG needs this; recall@k works fine treating >= 1 as relevant. - **Multiple relevant chunks are normal.** A question answerable only by combining two documents is a different (harder) test than a single-hop lookup — label the hop count and report the two classes separately. - **Harvest queries from real traffic.** Synthetic questions generated *from* a chunk are trivially retrievable from that chunk — they share its vocabulary. They measure your embedding model's ability to match paraphrases, not your retriever's ability to handle how people actually ask. That last point is the single most common way a retrieval eval flatters itself. A set of "generate a question from this passage" pairs will show recall@5 above 0.95 on a system that fails constantly in production. ## Failure classes worth their own bucket Each of these fails differently and needs its own cases: | Class | Why it breaks | |---|---| | **Vocabulary mismatch** | User says "can't log in", docs say "authentication failure". Pure semantic search handles this; keyword search does not — and hybrid exists for the reverse case | | **Exact identifiers** | Order numbers, error codes, SKUs, function names. Embeddings are *bad* at these; this is what BM25/keyword hybrid is for | | **Multi-hop** | The answer needs two documents. Single-shot retrieval structurally cannot | | **Negation / absence** | "Which plans do NOT include support?" Similarity retrieves the plans that *do* | | **Temporal** | "The current policy" retrieves a superseded version that is textually similar. Needs metadata filtering, not better embeddings | | **Nothing relevant exists** | The corpus does not contain the answer. Correct behaviour is to say so — and it must be a scored case, or the system learns to always answer | The last one deserves emphasis: **a golden set with no unanswerable questions cannot detect hallucination under retrieval failure**, which is the exact scenario users hit most often. ## Generation-side metrics Once the right context is in hand: | Metric | Question | Evaluator | |---|---|---| | **Faithfulness / groundedness** | Is every claim supported by the retrieved context? | Judge (`llm-judge.md`) — decompose into claims, check each | | **Citation accuracy** | Do the cited chunk ids actually contain the cited content? | **Deterministic** — verify the id exists and the claim maps to it | | **Answer relevance** | Does it address the question asked? | Judge | | **Refusal correctness** | Does it decline when the context does not contain the answer? | Deterministic, on the unanswerable bucket | Citation accuracy is quietly one of the highest-value checks available: it is free, it needs no judge, and it catches the specific failure where a model produces a confident answer and attaches a plausible-but-unrelated source. ## What to gate | Tier | Check | |---|---| | **Blocking** | recall@k on the frozen query set; citation-id validity; refusal rate on the unanswerable bucket | | **Blocking (ceiling)** | Context tokens per query — retrieval regressions often show up as cost before they show up as accuracy | | **Advisory until calibrated** | Faithfulness and answer-relevance judges | | **Diagnostic** | The 2x2 above, per bucket — this is what tells you which team owns the drop | ## Cross-reference - The levels this sits inside: `eval-taxonomy.md` - Case construction and freeze discipline: `golden-datasets.md` - Grading faithfulness without fooling yourself: `llm-judge.md` - Where these land in CI: `regression-gating.md` -
tooling-landscape.md 5 KB
# Tooling Landscape — which eval platform, when > **Verified 2026-08.** This is the fastest-moving part of the skill. Features, pricing and > OSS/commercial boundaries here change on a scale of months. **Re-verify with a web search > before quoting any specific claim** — treat everything below as a shape to check against > reality, not as a current fact sheet. No version numbers are quoted deliberately. ## The one structural fact **Trace-level observability and eval scoring have converged into the same products.** As of 2026 you are not choosing a tracer and then an eval library; you are choosing one system that captures the nested span tree (model calls, tool calls, arguments, cost) and attaches scores to those traces, in dev and in production. Any comparison that treats them as two categories is out of date. That convergence is what makes trajectory- and step-level eval practical at all (`eval-taxonomy.md`) — you cannot score a path you did not capture. ## The honest default **Start with a JSONL file and a 40-line runner.** Read cases, call the system, apply assertions, write results, print a delta. It is a morning's work, it has no vendor coupling, and it forces you to decide what you are actually measuring — which is the hard part, and the part no platform does for you. Adopt a platform when you hit a specific wall: | Wall | Then you want | |---|---| | Non-engineers need to curate the dataset | A dataset-management UI (Braintrust is the archetype) | | You need to search production traces and score live traffic | An observability-first platform (Langfuse, Arize, LangSmith, Opik) | | Evals should feel like the test suite | A pytest-native framework (DeepEval) | | You already run ML infra and want one system | MLflow | | Scheduled runs, shared dashboards, alerting | Any hosted platform; this is what you are paying for | Do not adopt one because the eval list is long. Metric catalogs are cheap; a rubric calibrated against your humans is not, and no platform ships that. ## The players Open-source cores (self-hostable, no vendor lock on the data): | Tool | Shape | Reach for it when | |---|---|---| | **DeepEval** | pytest-native LLM eval framework, large research-backed metric library | Your team's mental model is "tests"; you want evals in the existing test command | | **MLflow** | Tracing with replay, prompt versioning, automated eval — one OSS platform | You already run MLflow, or you want the whole stack under one OSS licence | | **Opik** (Comet) | Tracing with cost tracking, built-in metrics, prompt versioning, broad framework integrations | You want hosted-or-self-hosted flexibility with wide framework coverage | | **Langfuse** | Observability-first: traces, datasets, scores, self-hostable | Production trace search is the primary need | | **Arize Phoenix** | OSS tracing/eval, OpenTelemetry-native | You are standardising on OTel semantics | Commercial-first: | Tool | Shape | Reach for it when | |---|---|---| | **Braintrust** | Eval-focused, strong collaborative dataset curation, scheduled runs, score-regression tracking | PMs and domain experts must own the golden set | | **LangSmith** | Tracing + eval, tight LangChain/LangGraph integration | You are already deep in that ecosystem | | **Arize** | Production ML/LLM observability at scale | Enterprise monitoring is the driver | | **AgentOps** | Agent-run-centric session replay, cost and step tracking | Debugging long agent trajectories is the pain | ## Selection checklist Score candidates on the things that actually bite six months in: 1. **Can you export your traces and datasets?** The dataset is the durable asset (`golden-datasets.md`). If it only lives in a vendor UI, you have rented your history. 2. **Does it capture the full nested span tree**, including tool arguments? Without arguments you cannot do step-level scoring. 3. **Can it run in CI and fail a build** with a machine-readable result? A platform you can only read in a browser cannot gate anything. 4. **Can you pin the judge model version?** Unpinned judges silently re-baseline (`llm-judge.md`). 5. **Does it record cost and latency per case** natively, or must you thread it yourself? 6. **Self-host option?** Matters the moment eval inputs contain customer data. 7. **What happens to your custom metrics** — are they plain functions you own, or a DSL you would have to rewrite to migrate? ## Benchmarks vs your evals Public benchmarks (tau-bench and the agentic-benchmark family, SWE-bench and its descendants) are for *model selection* — they tell you which model to start from. They are not your eval suite: they are contaminated over time, they measure someone else's task distribution, and a model that tops them can still fail your product's specific policy. Use them once, at model-choice time. Then measure your own thing, on your own frozen set. ## Cross-reference - What to measure before you shop for a tool: `eval-taxonomy.md` - The asset that outlives whatever you pick: `golden-datasets.md` - Making the chosen tool gate CI: `regression-gating.md`
-
-
scripts
-
eval-baseline.py 20.3 KB
#!/usr/bin/env python3 """Tell a real eval regression from noise, using run history and McNemar's test. Reads the rolling run-history file the runner appends to, derives the baseline and the NOISE FLOOR, and judges one candidate run against them. Where per-case results for both runs are supplied it also runs McNemar's exact test, which is the honest answer to "did my change make it worse" on paired binary outcomes -- a score drop inside the noise band is not evidence of anything. Usage: eval-baseline.py [OPTIONS] <HISTORY.jsonl> Input: HISTORY.jsonl -- one summary object per run, appended over time. Recognised: `score` (required), `n`, `date`, `dataset`, `judge`, `cost_usd`, `p95_ms`. `-` reads stdin. --candidate FILE a one-row summary for the run under test (default: the last row of HISTORY) --baseline-results / --candidate-results per-case JSONL of {"id": ..., "passed": true|false} for McNemar Output: stdout -- human-readable verdict, or a --json envelope {"data": {...}, "meta": {...}} per SKILL-RESOURCE-PROTOCOL.md §4. Stderr: headers, warnings, errors. Exit: 0 no regression (noise, improvement, or not enough evidence), 2 usage, 3 not-found, 4 validation, 10 REGRESSION CONFIRMED or a cost/latency ceiling breached --accept INVERTS this deliberately, for use as a hillclimb keep/discard gate: 0 = KEEP (a real improvement), 10 = DISCARD (noise or worse). Under --accept, "noise" is a DISCARD -- banking a change that is inside the noise band is how an improvement loop turns into a random walk. See references/hillclimbing.md. Examples: eval-baseline.py evals/history.jsonl eval-baseline.py evals/history.jsonl --candidate /tmp/run.jsonl eval-baseline.py evals/history.jsonl --baseline-results base.jsonl \ --candidate-results new.jsonl --alpha 0.05 eval-baseline.py evals/history.jsonl --max-cost-usd 2.50 --max-p95-ms 6000 eval-baseline.py evals/history.jsonl --candidate iter.jsonl --accept # keep/discard eval-baseline.py evals/history.jsonl --json | jq '.data.recommended_threshold' Offline and stdlib-only: it reads files you already have, it never calls a model. """ import argparse import json import math import sys SCHEMA = "claude-mods.evals-ops.eval-baseline/v1" EXIT_OK, EXIT_USAGE, EXIT_NOT_FOUND, EXIT_VALIDATION, EXIT_REGRESSION = 0, 2, 3, 4, 10 def load_jsonl(path, label): """Read JSONL from a path or stdin. Raises with a line number on bad input.""" if path == "-": text = sys.stdin.read() else: try: with open(path, "r", encoding="utf-8") as fh: text = fh.read() except FileNotFoundError: raise FileNotFoundError(path) except IsADirectoryError: raise ValueError(f"{path} is a directory, not a JSONL file") rows = [] for lineno, line in enumerate(text.splitlines(), start=1): line = line.strip() if not line or line.startswith("#"): continue try: obj = json.loads(line) except json.JSONDecodeError as exc: raise ValueError(f"{label} line {lineno}: not valid JSON ({exc.msg})") if not isinstance(obj, dict): raise ValueError(f"{label} line {lineno}: expected a JSON object") rows.append(obj) return rows def stdev(values): """Sample standard deviation. None below two points -- one run is not a spread.""" n = len(values) if n < 2: return None mean = sum(values) / n return math.sqrt(sum((v - mean) ** 2 for v in values) / (n - 1)) # Above this many discordant pairs the exact test is switched for the normal # approximation. The exact sum is O(n) big-integer work over 2**n: measured at # 133 ms for n=2000 but ~100 SECONDS for n=20000, which is a CI hang, not a # slow test. The approximation is reliable well below this bound (the usual # rule of thumb is b+c >= 25), so nothing accurate is lost by switching here. MCNEMAR_EXACT_MAX = 1000 def mcnemar(b, c): """Two-sided McNemar p-value. Returns (p_value, method). b = passed before, fails now (regressions); c = failed before, passes now. Only the DISCORDANT pairs carry information -- cases that behaved the same in both runs tell you nothing about whether the change helped. Under the null b ~ Binomial(b+c, 0.5), so the p-value is a coin-flip tail probability. Exact below MCNEMAR_EXACT_MAX because eval sets routinely produce b+c < 25, where the chi-square/normal approximation is unreliable; normal (with a continuity correction) above it, because the exact form does not terminate in useful time. The method used is reported so a caller never has to guess. """ n = b + c if n == 0: return 1.0, "none-discordant" if n <= MCNEMAR_EXACT_MAX: k = min(b, c) tail = sum(math.comb(n, i) for i in range(0, k + 1)) / (2 ** n) return min(1.0, 2 * tail), "exact" # Continuity-corrected normal approximation; erfc gives the two-sided tail. z = (abs(b - c) - 1) / math.sqrt(n) return min(1.0, math.erfc(abs(z) / math.sqrt(2))), "normal-approx" def paired_counts(baseline_rows, candidate_rows): """Pair per-case results by id. Returns (b, c, both_pass, both_fail, unpaired).""" # Test `"id" in r`, NOT the truthiness of r["id"] -- an integer id of 0 or an # empty-string id is falsy, and the truthiness form silently DROPPED those # cases from the paired test, hiding real regressions. Stringify so 1 and "1" # pair with each other rather than becoming two unpaired cases. def index(rows): return {str(r["id"]): bool(r.get("passed")) for r in rows if "id" in r} base, cand = index(baseline_rows), index(candidate_rows) shared = set(base) & set(cand) b = sorted(i for i in shared if base[i] and not cand[i]) c = sorted(i for i in shared if not base[i] and cand[i]) both_pass = sum(1 for i in shared if base[i] and cand[i]) both_fail = sum(1 for i in shared if not base[i] and not cand[i]) unpaired = len(set(base) ^ set(cand)) return b, c, both_pass, both_fail, unpaired def build_report(history, candidate, window, sigma, paired, alpha, ceilings): prior = [r for r in history if r is not candidate][-window:] scores = [float(r["score"]) for r in prior if isinstance(r.get("score"), (int, float))] warnings = [] # A row whose score is absent or non-numeric (a JSON string "0.9", a null) is # unusable. Dropping it in silence let a whole malformed history report # "insufficient data" and exit 0 -- a gate failing open with no explanation. unusable = len(prior) - len(scores) if unusable: warnings.append( f"{unusable} history row(s) had a missing or non-numeric 'score' and " "were ignored (scores must be JSON numbers, not strings)" ) baseline = round(sum(scores) / len(scores), 4) if scores else None spread = stdev(scores) spread = round(spread, 4) if spread is not None else None # A threshold inside the noise band fails on identical code. Sit below the # baseline by more than the measured spread -- that is the whole point of # keeping a history rather than judging each run standalone. # # `spread is not None` is deliberate, NOT a truthiness test: a spread of # exactly 0.0 is a MEASURED result (a deterministic metric, or a genuinely # stable suite) and the most informative history you can have. Treating it as # "no data" made the gate report insufficient-data and exit 0 on an # unambiguous 0.90 -> 0.70 regression. threshold = (round(baseline - sigma * spread, 4) if (baseline is not None and spread is not None) else None) cand_score = candidate.get("score") if candidate else None cand_score = float(cand_score) if isinstance(cand_score, (int, float)) else None delta = round(cand_score - baseline, 4) if (cand_score is not None and baseline is not None) else None # z is undefined when the spread is exactly 0 (division by zero), but that is # the EASY case, not the hard one: with no measured variance any movement is # real. Handled explicitly in the verdict below rather than left as None. z = None if delta is not None and spread: z = round(delta / spread, 2) if len(scores) < 3: warnings.append( f"only {len(scores)} prior run(s) in the window; a noise floor needs " "at least 3, and 5 reruns of unchanged code is the honest way to get it" ) # Comparing across a dataset, judge or rubric change is confounded: you cannot # tell whether the system moved or the measurement did. for field, human in (("dataset", "dataset version"), ("judge", "judge model")): seen = {r.get(field) for r in prior + ([candidate] if candidate else []) if r.get(field)} if len(seen) > 1: warnings.append( f"{human} changed within the window ({', '.join(sorted(map(str, seen)))}) " "-- re-baseline; scores across that boundary are not comparable" ) # --- significance ------------------------------------------------------- significance = None if paired is not None: b_ids, c_ids, both_pass, both_fail, unpaired = paired p, method = mcnemar(len(b_ids), len(c_ids)) significance = { "test": f"mcnemar-{method}", "regressed": b_ids, # passed before, fails now "fixed": c_ids, # failed before, passes now "n_regressed": len(b_ids), "n_fixed": len(c_ids), "both_pass": both_pass, "both_fail": both_fail, "unpaired_ids": unpaired, "p_value": round(p, 6), "alpha": alpha, "significant": p < alpha, "direction": "worse" if len(b_ids) > len(c_ids) else ("better" if len(c_ids) > len(b_ids) else "unchanged"), } if unpaired: warnings.append( f"{unpaired} case id(s) appear in only one result set; they are " "excluded from the paired test" ) if len(b_ids) + len(c_ids) == 0: warnings.append("no discordant cases -- the two runs agree everywhere") # --- ceilings ----------------------------------------------------------- # Ceilings, not trends: a trend gate fires on noise, a ceiling encodes a # product decision someone actually made. breaches = [] if candidate: for key, limit, label in ( ("cost_usd", ceilings.get("cost"), "cost_usd"), ("p95_ms", ceilings.get("p95"), "p95_ms"), ): value = candidate.get(key) if limit is not None and isinstance(value, (int, float)) and value > limit: breaches.append({"metric": label, "value": value, "ceiling": limit}) # --- verdict ------------------------------------------------------------ # Paired significance is the strongest evidence available; fall back to the # sigma rule only when per-case results were not supplied. if significance is not None: if significance["significant"] and significance["direction"] == "worse": verdict, reason = "regression", ( f"{significance['n_regressed']} case(s) regressed vs " f"{significance['n_fixed']} fixed, p={significance['p_value']} " f"< alpha={alpha}" ) elif significance["significant"] and significance["direction"] == "better": verdict, reason = "improvement", ( f"{significance['n_fixed']} case(s) fixed vs " f"{significance['n_regressed']} regressed, p={significance['p_value']}" ) else: verdict, reason = "noise", ( f"p={significance['p_value']} >= alpha={alpha}; the difference is " "not distinguishable from chance" ) elif cand_score is None or threshold is None: verdict, reason = "insufficient-data", ( "no candidate score, or fewer than 2 usable history rows -- rerun the " "suite against unchanged code a few times to establish a noise floor " "before gating on it" ) elif spread == 0: # Zero measured variance: every rerun of unchanged code gave the same # number, so ANY movement is signal and there is no band to be inside. if cand_score < baseline: verdict, reason = "regression", ( f"score {cand_score} is below a baseline of {baseline} measured " "with zero variance across the window" ) elif cand_score > baseline: verdict, reason = "improvement", ( f"score {cand_score} exceeds a zero-variance baseline of {baseline}" ) else: verdict, reason = "noise", f"score is identical to the baseline {baseline}" elif cand_score < threshold: verdict, reason = "regression", ( f"score {cand_score} is below the {sigma}-sigma threshold {threshold}" ) elif z is not None and z > sigma: verdict, reason = "improvement", f"score is {z} sigma above baseline" else: verdict, reason = "noise", ( f"score {cand_score} is within the noise band around {baseline}" ) if breaches: reason += "; " + ", ".join( f"{x['metric']} {x['value']} exceeds ceiling {x['ceiling']}" for x in breaches ) # The hillclimb decision is NOT the same question as the CI-gate decision. # CI asks "did this get worse?" (noise is fine). A hillclimb asks "is this # improvement real?" (noise is not good enough to bank). Only an improvement # that clears the floor, or is significant on the paired test, is a KEEP. decision = "keep" if verdict == "improvement" else "discard" return { "verdict": verdict, "reason": reason, "hillclimb_decision": decision, "failing": verdict == "regression" or bool(breaches), "baseline": baseline, "noise_floor": spread, "sigma": sigma, "recommended_threshold": threshold, "candidate_score": cand_score, "delta": delta, "z": z, "window": len(scores), "significance": significance, "ceiling_breaches": breaches, "warnings": warnings, } def print_human(r, source): out = [f"eval baseline: {source}"] out.append(f" window {r['window']} prior run(s)") out.append(f" baseline {r['baseline']}") out.append(f" noise floor {r['noise_floor']} (sample stdev)") out.append(f" gate at {r['recommended_threshold']} ({r['sigma']} sigma below baseline)") out.append(f" candidate {r['candidate_score']} delta {r['delta']} z {r['z']}") sig = r["significance"] if sig: out.append("") out.append(" paired test (McNemar exact)") out.append(f" regressed {sig['n_regressed']} fixed {sig['n_fixed']}") out.append(f" unchanged {sig['both_pass']} pass / {sig['both_fail']} fail") out.append(f" p-value {sig['p_value']} (alpha {sig['alpha']})") # Naming the flipped cases is what stops people ignoring a red gate. if sig["regressed"]: out.append(f" now failing {', '.join(sig['regressed'][:15])}") if sig["fixed"]: out.append(f" now passing {', '.join(sig['fixed'][:15])}") for x in r["ceiling_breaches"]: out.append(f" CEILING {x['metric']} {x['value']} > {x['ceiling']}") out.append("") out.append(f" verdict {r['verdict'].upper()} - {r['reason']}") out.append(f" hillclimb {r['hillclimb_decision'].upper()}") print("\n".join(out)) def main(argv=None): ap = argparse.ArgumentParser( prog="eval-baseline.py", description="Distinguish an eval regression from noise using run history " "and McNemar's exact test.", ) ap.add_argument("history", help="rolling run-history JSONL, or - for stdin") ap.add_argument("--candidate", metavar="FILE", help="one-row summary for the run under test " "(default: the last row of HISTORY)") ap.add_argument("--baseline-results", metavar="FILE", help="per-case JSONL of the baseline run, for the paired test") ap.add_argument("--candidate-results", metavar="FILE", help="per-case JSONL of the candidate run, for the paired test") ap.add_argument("--window", type=int, default=10, metavar="N", help="prior runs used for baseline and noise floor (default: 10)") ap.add_argument("--sigma", type=float, default=2.0, metavar="K", help="how many noise-floor widths below baseline the gate sits " "(default: 2.0)") ap.add_argument("--alpha", type=float, default=0.05, metavar="A", help="significance level for the paired test (default: 0.05)") ap.add_argument("--max-cost-usd", type=float, metavar="USD", help="fail if the candidate run exceeds this cost") ap.add_argument("--max-p95-ms", type=float, metavar="MS", help="fail if the candidate run exceeds this p95 latency") ap.add_argument("--accept", action="store_true", help="hillclimb keep/discard gate: exit 0 = KEEP a real improvement, " "exit 10 = DISCARD noise or a regression (INVERTS the default " "meaning of noise -- see the header)") ap.add_argument("--json", action="store_true", help="emit the JSON envelope on stdout") try: args = ap.parse_args(argv) except SystemExit as exc: raise SystemExit(exc.code) def fail(code, kind, message): if args.json: print(json.dumps({"error": {"code": kind, "message": message, "details": {}}})) print(f"eval-baseline: {message}", file=sys.stderr) return code if args.window < 1: return fail(EXIT_USAGE, "VALIDATION", "--window must be at least 1") if args.sigma <= 0: return fail(EXIT_USAGE, "VALIDATION", "--sigma must be positive") if not (0.0 < args.alpha < 1.0): return fail(EXIT_USAGE, "VALIDATION", "--alpha must be strictly between 0 and 1") if bool(args.baseline_results) != bool(args.candidate_results): return fail(EXIT_USAGE, "VALIDATION", "--baseline-results and --candidate-results must be given together") try: history = load_jsonl(args.history, "history") candidate = None if args.candidate: rows = load_jsonl(args.candidate, "candidate") if not rows: return fail(EXIT_VALIDATION, "VALIDATION", "candidate file has no rows") candidate = rows[-1] elif history: candidate = history[-1] paired = None if args.baseline_results: base_rows = load_jsonl(args.baseline_results, "baseline-results") cand_rows = load_jsonl(args.candidate_results, "candidate-results") if not base_rows or not cand_rows: return fail(EXIT_VALIDATION, "VALIDATION", "per-case result files must not be empty") paired = paired_counts(base_rows, cand_rows) except FileNotFoundError as exc: return fail(EXIT_NOT_FOUND, "NOT_FOUND", f"no such file: {exc}") except ValueError as exc: return fail(EXIT_VALIDATION, "VALIDATION", str(exc)) if not history and candidate is None: return fail(EXIT_VALIDATION, "VALIDATION", "history is empty and no --candidate given") report = build_report( history, candidate, args.window, args.sigma, paired, args.alpha, {"cost": args.max_cost_usd, "p95": args.max_p95_ms}, ) if args.json: print(json.dumps({ "data": report, "meta": {"count": report["window"], "schema": SCHEMA, "source": args.history}, }, indent=2)) else: print_human(report, args.history) for w in report["warnings"]: print(f"eval-baseline: warning: {w}", file=sys.stderr) if args.accept: return EXIT_OK if report["hillclimb_decision"] == "keep" else EXIT_REGRESSION return EXIT_REGRESSION if report["failing"] else EXIT_OK if __name__ == "__main__": sys.exit(main()) -
goldenset-audit.py 16.5 KB
#!/usr/bin/env python3 """Audit a golden eval set for the rot that silently kills a regression suite. Checks duplicates and near-duplicates, bucket balance, undated/unexplained cases, missing expectations, and drift from a frozen manifest. Usage: goldenset-audit.py [OPTIONS] <GOLDEN.jsonl> Input: JSONL, one case per line. Recognised fields: `id`, `bucket`, `added` (ISO date), `why`, `input`, `expected`, `criteria`. All optional -- a missing field becomes a finding, never a crash. `-` reads stdin. Output: stdout -- human-readable report, or a --json envelope {"data": {"findings": [...], ...}, "meta": {...}} per §4. Stderr: headers, progress, warnings, errors. Exit: 0 clean, 2 usage, 3 not-found, 4 validation (unparseable/empty), 10 FINDINGS at or above --fail-on severity Examples: goldenset-audit.py evals/golden.jsonl goldenset-audit.py evals/golden.jsonl --json | jq '.data.findings[]' goldenset-audit.py evals/golden.jsonl --write-freeze manifest.json goldenset-audit.py evals/golden.jsonl --freeze manifest.json --fail-on warn Offline and stdlib-only. --write-freeze is the only write, and it is atomic. """ import argparse import hashlib import json import os import re import sys from collections import Counter SCHEMA = "claude-mods.evals-ops.goldenset-audit/v1" FREEZE_SCHEMA = "claude-mods.evals-ops.goldenset-freeze/v1" EXIT_OK, EXIT_USAGE, EXIT_NOT_FOUND, EXIT_VALIDATION, EXIT_FINDINGS = 0, 2, 3, 4, 10 SEVERITY_ORDER = {"info": 0, "warn": 1, "error": 2} # TARGET share per bucket, from references/golden-datasets.md, and the WARN band # around it. These are deliberately different numbers: the target is what you aim # for when composing the set, the band is where a warning fires. A band equal to # the target would warn on almost every healthy set and get ignored within a week # -- the same "a gate that cries wolf is a dead gate" rule the skill applies to CI. # Both numbers are reported, so a warning names the target it is measured against. BUCKET_TARGETS = { "production": (0.40, 0.50), "replay": (0.20, 0.30), "adversarial": (0.15, 0.20), "edge": (0.10, 0.15), } BUCKET_BANDS = { "production": (0.30, 0.65), "replay": (0.10, 0.40), "adversarial": (0.08, 0.35), "edge": (0.05, 0.30), } ISO_DATE = re.compile(r"^\d{4}-\d{2}-\d{2}") WORD = re.compile(r"[a-z0-9]+") def canonical(obj): """Stable JSON text for hashing: key order must not change a case's identity.""" return json.dumps(obj, sort_keys=True, separators=(",", ":"), ensure_ascii=False) # A case's hash is its CONTENT, never its identity or metadata. Excluded: # _line bookkeeping we inject -- leaked in and made two identical cases on # different lines hash differently, so duplicate detection missed them # id the whole point of duplicate detection is "two ids, one case"; if id # were hashed, DUPLICATE_CASE could never fire (ids are unique by rule) # bucket reclassifying a case does not change what it tests # added / why documentation about the case, not the case # Excluding id also keeps the freeze manifest honest: it stores id and hash as # separate fields, so a rename shows as removed+added rather than as an edit. HASH_EXCLUDE = ("_line", "id", "bucket", "added", "why") def case_hash(case): payload = {k: case.get(k) for k in ("input", "expected", "criteria") if k in case} if not payload: payload = {k: v for k, v in case.items() if k not in HASH_EXCLUDE} return hashlib.sha256(canonical(payload).encode("utf-8")).hexdigest() def tokens(case): return set(WORD.findall(canonical(case.get("input", "")).lower())) def jaccard(a, b): if not a or not b: return 0.0 return len(a & b) / len(a | b) def load_cases(path): if path == "-": text = sys.stdin.read() else: try: with open(path, "r", encoding="utf-8") as fh: text = fh.read() except FileNotFoundError: raise FileNotFoundError(path) except IsADirectoryError: raise ValueError(f"{path} is a directory, not a JSONL file") cases = [] for lineno, line in enumerate(text.splitlines(), start=1): line = line.strip() if not line or line.startswith("#"): continue try: obj = json.loads(line) except json.JSONDecodeError as exc: raise ValueError(f"line {lineno}: not valid JSON ({exc.msg})") if not isinstance(obj, dict): raise ValueError(f"line {lineno}: expected a JSON object") obj.setdefault("_line", lineno) cases.append(obj) return cases def audit(cases, near_threshold, max_pairs): findings = [] def add(severity, code, message, **details): findings.append( {"severity": severity, "code": code, "message": message, "details": details} ) n = len(cases) ids = [c.get("id") or f"line-{c['_line']}" for c in cases] # --- identity ----------------------------------------------------------- missing_id = [c["_line"] for c in cases if not c.get("id")] if missing_id: add("warn", "MISSING_ID", f"{len(missing_id)} case(s) have no 'id'", lines=missing_id[:20]) dup_ids = [i for i, c in Counter(ids).items() if c > 1] if dup_ids: add("error", "DUPLICATE_ID", f"{len(dup_ids)} id(s) used more than once", ids=sorted(dup_ids)[:20]) # --- exact duplicates --------------------------------------------------- by_hash = {} for cid, case in zip(ids, cases): by_hash.setdefault(case_hash(case), []).append(cid) exact = {h: members for h, members in by_hash.items() if len(members) > 1} if exact: add("error", "DUPLICATE_CASE", f"{len(exact)} group(s) of identical cases -- they double-weight one behaviour", groups=[sorted(m) for m in list(exact.values())[:10]]) # --- near-duplicates ---------------------------------------------------- # O(n^2) on token sets; bounded by --max-pairs so a huge set degrades to a # skipped check with a note, never a hang. near = [] if n * (n - 1) // 2 <= max_pairs: toks = [tokens(c) for c in cases] for i in range(n): for j in range(i + 1, n): sim = jaccard(toks[i], toks[j]) if sim >= near_threshold: near.append({"a": ids[i], "b": ids[j], "similarity": round(sim, 3)}) if near: add("warn", "NEAR_DUPLICATE", f"{len(near)} case pair(s) above {near_threshold} input similarity", pairs=sorted(near, key=lambda p: -p["similarity"])[:15]) else: add("info", "NEAR_DUPLICATE_SKIPPED", f"near-duplicate scan skipped: {n} cases exceeds --max-pairs budget", cases=n, max_pairs=max_pairs) # --- completeness ------------------------------------------------------- no_expect = [cid for cid, c in zip(ids, cases) if "expected" not in c and not c.get("criteria")] if no_expect: add("error", "NO_EXPECTATION", f"{len(no_expect)} case(s) have neither 'expected' nor 'criteria' -- ungradeable", ids=no_expect[:20]) no_why = [cid for cid, c in zip(ids, cases) if not c.get("why")] if no_why: add("warn", "NO_RATIONALE", f"{len(no_why)} case(s) have no 'why' -- a future maintainer cannot tell " "whether deleting them loses coverage", ids=no_why[:20]) # --- dates -------------------------------------------------------------- dated = [c.get("added") for c in cases if isinstance(c.get("added"), str) and ISO_DATE.match(c["added"])] undated = [cid for cid, c in zip(ids, cases) if not (isinstance(c.get("added"), str) and ISO_DATE.match(c.get("added", "")))] if undated: add("warn", "UNDATED", f"{len(undated)} case(s) have no ISO 'added' date -- staleness is undetectable", ids=undated[:20]) newest = max(dated)[:10] if dated else None # --- bucket balance ----------------------------------------------------- buckets = Counter(c.get("bucket") or "_unset" for c in cases) shares = {b: round(c / n, 4) for b, c in buckets.items()} if buckets.get("_unset"): add("warn", "NO_BUCKET", f"{buckets['_unset']} case(s) have no 'bucket' -- balance cannot be tracked", count=buckets["_unset"]) known = [b for b in buckets if b in BUCKET_BANDS] if known: for bucket, (lo, hi) in BUCKET_BANDS.items(): share = shares.get(bucket, 0.0) t_lo, t_hi = BUCKET_TARGETS[bucket] if share < lo: add("warn", "BUCKET_THIN", f"bucket '{bucket}' is {share:.0%} of the set " f"(target {t_lo:.0%}-{t_hi:.0%}, warns below {lo:.0%})", bucket=bucket, share=share, floor=lo, target=[t_lo, t_hi]) elif share > hi: add("warn", "BUCKET_HEAVY", f"bucket '{bucket}' is {share:.0%} of the set " f"(target {t_lo:.0%}-{t_hi:.0%}, warns above {hi:.0%})", bucket=bucket, share=share, ceiling=hi, target=[t_lo, t_hi]) # --- size --------------------------------------------------------------- if n < 20: add("warn", "TOO_SMALL", f"{n} cases; below ~20 a percentage moves too coarsely to interpret", cases=n) elif n < 100: add("info", "SMALL", f"{n} cases; 100-300 is the working range for a regression set", cases=n) return findings, { "cases": n, "buckets": dict(buckets), "bucket_shares": shares, "newest_added": newest, "dated_cases": len(dated), "near_duplicate_pairs": len(near), } def freeze_manifest(cases): entries = sorted( ({"id": c.get("id") or f"line-{c['_line']}", "hash": case_hash(c)} for c in cases), key=lambda e: (e["id"], e["hash"]), ) digest = hashlib.sha256( canonical([[e["id"], e["hash"]] for e in entries]).encode("utf-8") ).hexdigest() return {"schema": FREEZE_SCHEMA, "count": len(entries), "digest": digest, "cases": entries} def compare_freeze(cases, manifest): """Diff the live set against a frozen manifest. Returns findings.""" if not isinstance(manifest, dict) or "cases" not in manifest: return [{"severity": "error", "code": "FREEZE_MALFORMED", "message": "freeze manifest has no 'cases' list", "details": {}}] # An entry with no id cannot be matched to a live case; index it by its hash so it # still reports as removed rather than colliding on a None key. frozen = { str(e.get("id") or f"unnamed-{e.get('hash', '?')}"): e.get("hash") for e in manifest["cases"] if isinstance(e, dict) } live = {c.get("id") or f"line-{c['_line']}": case_hash(c) for c in cases} added = sorted(set(live) - set(frozen)) removed = sorted(set(frozen) - set(live)) changed = sorted(i for i in set(live) & set(frozen) if live[i] != frozen[i]) findings = [] if changed: findings.append({ "severity": "error", "code": "FREEZE_CASE_CHANGED", "message": f"{len(changed)} frozen case(s) were edited in place -- this is " "fitting the test to the code; delete and re-add instead", "details": {"ids": changed[:20]}, }) if removed: findings.append({ "severity": "error", "code": "FREEZE_CASE_REMOVED", "message": f"{len(removed)} frozen case(s) are gone -- scores are no longer " "comparable to the frozen baseline", "details": {"ids": removed[:20]}, }) if added: findings.append({ "severity": "warn", "code": "FREEZE_CASE_ADDED", "message": f"{len(added)} case(s) added since freeze -- re-baseline before " "comparing scores across this boundary", "details": {"ids": added[:20]}, }) return findings def atomic_write_json(path, payload): tmp = f"{path}.tmp" with open(tmp, "w", encoding="utf-8") as fh: json.dump(payload, fh, indent=2, sort_keys=True) fh.write("\n") os.replace(tmp, path) def print_human(findings, stats, source): out = [f"golden-set audit: {source}", f" cases {stats['cases']}"] if stats["buckets"]: shares = ", ".join( f"{b} {stats['bucket_shares'][b]:.0%}" for b in sorted(stats["buckets"]) ) out.append(f" buckets {shares}") out.append(f" newest added {stats['newest_added'] or '-'}") out.append("") if not findings: out.append(" no findings") else: for sev in ("error", "warn", "info"): rows = [f for f in findings if f["severity"] == sev] for f in rows: out.append(f" [{sev.upper():<5}] {f['code']}: {f['message']}") print("\n".join(out)) def main(argv=None): parser = argparse.ArgumentParser( prog="goldenset-audit.py", description="Audit a golden eval set for duplicates, imbalance, staleness and drift.", ) parser.add_argument("golden", help="JSONL golden set, or - for stdin") parser.add_argument("--freeze", metavar="MANIFEST", help="compare against a freeze manifest written by --write-freeze") parser.add_argument("--write-freeze", metavar="MANIFEST", help="write a freeze manifest for this set (atomic)") parser.add_argument("--fail-on", choices=("error", "warn", "info"), default="error", help="lowest severity that exits 10 (default: error)") parser.add_argument("--near-threshold", type=float, default=0.9, metavar="R", help="Jaccard similarity at which two inputs are near-duplicates " "(default: 0.9)") parser.add_argument("--max-pairs", type=int, default=200000, metavar="N", help="skip the O(n^2) near-duplicate scan above this pair count " "(default: 200000)") parser.add_argument("--json", action="store_true", help="emit the JSON envelope on stdout") try: args = parser.parse_args(argv) except SystemExit as exc: raise SystemExit(exc.code) def fail(code, kind, message): if args.json: print(json.dumps({"error": {"code": kind, "message": message, "details": {}}})) print(f"goldenset-audit: {message}", file=sys.stderr) return code if not (0.0 < args.near_threshold <= 1.0): return fail(EXIT_USAGE, "VALIDATION", "--near-threshold must be in (0, 1]") if args.max_pairs < 0: return fail(EXIT_USAGE, "VALIDATION", "--max-pairs must not be negative") if args.freeze and args.write_freeze: return fail(EXIT_USAGE, "VALIDATION", "--freeze and --write-freeze are mutually exclusive") try: cases = load_cases(args.golden) except FileNotFoundError as exc: return fail(EXIT_NOT_FOUND, "NOT_FOUND", f"no such file: {exc}") except ValueError as exc: return fail(EXIT_VALIDATION, "VALIDATION", str(exc)) if not cases: return fail(EXIT_VALIDATION, "VALIDATION", "golden set is empty") findings, stats = audit(cases, args.near_threshold, args.max_pairs) if args.freeze: try: with open(args.freeze, "r", encoding="utf-8") as fh: manifest = json.load(fh) except FileNotFoundError: return fail(EXIT_NOT_FOUND, "NOT_FOUND", f"no such freeze manifest: {args.freeze}") except json.JSONDecodeError as exc: return fail(EXIT_VALIDATION, "VALIDATION", f"freeze manifest is not JSON: {exc.msg}") findings.extend(compare_freeze(cases, manifest)) if args.write_freeze: try: atomic_write_json(args.write_freeze, freeze_manifest(cases)) except OSError as exc: return fail(EXIT_VALIDATION, "VALIDATION", f"cannot write manifest: {exc}") print(f"goldenset-audit: wrote freeze manifest {args.write_freeze}", file=sys.stderr) threshold = SEVERITY_ORDER[args.fail_on] triggering = [f for f in findings if SEVERITY_ORDER[f["severity"]] >= threshold] if args.json: print(json.dumps({ "data": {"findings": findings, "stats": stats, "fail_on": args.fail_on}, "meta": {"count": len(findings), "schema": SCHEMA, "source": args.golden}, }, indent=2)) else: print_human(findings, stats, args.golden) return EXIT_FINDINGS if triggering else EXIT_OK if __name__ == "__main__": sys.exit(main()) -
judge-calibration.py 13.1 KB
#!/usr/bin/env python3 """Measure an LLM judge against human labels: Cohen kappa, confusion, bias probes. Usage: judge-calibration.py [OPTIONS] <LABELS.jsonl> Input: JSONL, one record per case. Required fields: `human`, `judge` (any hashable label -- "pass"/"fail", true/false, 1-5). Optional: `id`, `length` (or --verbosity-field NAME) for the verbosity-bias probe, and `position` ("first"/"second") for the position-bias probe. Use `-` to read the JSONL from stdin. Output: stdout -- human-readable report, or a --json envelope {"data": {...}, "meta": {...}} per SKILL-RESOURCE-PROTOCOL.md §4. Stderr: headers, progress, warnings, errors. Exit: 0 calibrated (kappa >= --min-kappa), 2 usage, 3 not-found, 4 validation (unparseable/empty/missing fields), 10 UNDER-CALIBRATED (ran fine, kappa below threshold) Examples: judge-calibration.py labels.jsonl judge-calibration.py labels.jsonl --min-kappa 0.8 judge-calibration.py labels.jsonl --json | jq '.data.kappa' judge-calibration.py labels.jsonl --verbosity-field output_chars cat labels.jsonl | judge-calibration.py - --json Offline and stdlib-only: this scores labels you already have, it never calls a model. """ import argparse import json import sys from collections import Counter, defaultdict SCHEMA = "claude-mods.evals-ops.judge-calibration/v1" EXIT_OK, EXIT_USAGE, EXIT_NOT_FOUND, EXIT_VALIDATION, EXIT_UNDER = 0, 2, 3, 4, 10 # Landis & Koch style bands, as used for judge calibration in practice. # < 0.6 means the RUBRIC needs work, not the judge model -- see references/llm-judge.md. BANDS = [ (0.80, "strong", "production-ready; safe to gate on with a margin"), (0.60, "substantial", "usable; keep advisory or gate loosely"), (0.40, "moderate", "rubric needs work before this gates anything"), (0.20, "fair", "rubric is the problem, not the model"), (float("-inf"), "poor", "no better than chance; rewrite the rubric"), ] def norm(value): """Normalise a label to a comparable string. true/1/'PASS' must all agree.""" if isinstance(value, bool): return "pass" if value else "fail" if isinstance(value, str): return value.strip().lower() return str(value) def cohens_kappa(pairs): """Cohen kappa for two raters over nominal labels. Returns None if undefined.""" n = len(pairs) if n == 0: return None observed = sum(1 for a, b in pairs if a == b) / n a_counts, b_counts = Counter(a for a, _ in pairs), Counter(b for _, b in pairs) expected = sum((a_counts[k] / n) * (b_counts.get(k, 0) / n) for k in a_counts) if expected >= 1.0: # Both raters used a single identical label: agreement is total but chance- # corrected agreement is undefined (0/0). Report it rather than dividing. return None return (observed - expected) / (1.0 - expected) def pearson(xs, ys): """Pearson correlation. Returns None when a series has no variance.""" n = len(xs) if n < 3: return None mx, my = sum(xs) / n, sum(ys) / n sxy = sum((x - mx) * (y - my) for x, y in zip(xs, ys)) sxx = sum((x - mx) ** 2 for x in xs) syy = sum((y - my) ** 2 for y in ys) if sxx <= 0 or syy <= 0: return None return sxy / ((sxx**0.5) * (syy**0.5)) def band_for(kappa): for floor, name, advice in BANDS: if kappa >= floor: return name, advice return "poor", "rewrite the rubric" def load_records(path): """Read JSONL from a path or stdin. Raises ValueError with a line number.""" if path == "-": text = sys.stdin.read() else: try: with open(path, "r", encoding="utf-8") as fh: text = fh.read() except FileNotFoundError: raise FileNotFoundError(path) except IsADirectoryError: raise ValueError(f"{path} is a directory, not a JSONL file") records = [] for lineno, line in enumerate(text.splitlines(), start=1): line = line.strip() if not line or line.startswith("#"): continue try: obj = json.loads(line) except json.JSONDecodeError as exc: raise ValueError(f"line {lineno}: not valid JSON ({exc.msg})") if not isinstance(obj, dict): raise ValueError(f"line {lineno}: expected a JSON object") records.append((lineno, obj)) return records def analyse(records, verbosity_field): pairs, ids, skipped = [], [], [] lengths, judge_numeric = [], [] position_rows = defaultdict(list) for lineno, obj in records: if "human" not in obj or "judge" not in obj: skipped.append({"line": lineno, "reason": "missing 'human' or 'judge'"}) continue h, j = norm(obj["human"]), norm(obj["judge"]) pairs.append((h, j)) ids.append(obj.get("id", f"line-{lineno}")) # Verbosity probe: does judge score track output length among cases the # HUMAN scored identically? Correlation there is bias, not signal. length = obj.get(verbosity_field) jn = obj["judge"] # bool is a subclass of int in Python, so a `true` in the length field # would silently read as 1.0 and yield a confident, meaningless # correlation. A length is never a boolean -- reject it explicitly. if isinstance(length, bool): length = None if isinstance(length, (int, float)) and isinstance(jn, (int, float, bool)): lengths.append(float(length)) judge_numeric.append(float(jn)) pos = obj.get("position") if isinstance(pos, str): position_rows[pos.strip().lower()].append(1 if h == j else 0) return pairs, ids, skipped, lengths, judge_numeric, position_rows def build_report(pairs, ids, skipped, lengths, judge_numeric, position_rows, min_kappa): n = len(pairs) agreement = sum(1 for a, b in pairs if a == b) / n kappa = cohens_kappa(pairs) confusion = Counter(pairs) labels = sorted({lab for pair in pairs for lab in pair}) disagreements = [ {"id": cid, "human": h, "judge": j} for cid, (h, j) in zip(ids, pairs) if h != j ] # Per-class recall from the human's point of view: where does the judge go wrong? # A judge that only ever under-passes is safe to gate on; one that over-passes is not. per_class = {} for lab in labels: total = sum(c for (h, _), c in confusion.items() if h == lab) hit = confusion.get((lab, lab), 0) per_class[lab] = { "human_count": total, "judge_agreed": hit, "recall": round(hit / total, 4) if total else None, } verbosity_r = pearson(lengths, judge_numeric) position = { pos: {"n": len(v), "agreement": round(sum(v) / len(v), 4)} for pos, v in position_rows.items() if v } if kappa is None: band, advice, calibrated = "undefined", ( "every case shares one label -- kappa is undefined; " "stratify the calibration sample across the judge's own verdicts" ), False else: band, advice = band_for(kappa) calibrated = kappa >= min_kappa warnings = [] if n < 50: warnings.append( f"only {n} labelled cases; 50-200 is the recommended range for a " "meaningful kappa" ) if len(labels) < 2: warnings.append("only one distinct label present -- the sample is not stratified") if verbosity_r is not None and abs(verbosity_r) >= 0.4: warnings.append( f"verbosity probe: judge score correlates {verbosity_r:+.2f} with output " "length -- separate correctness from style in the rubric" ) if len(position) >= 2: vals = [v["agreement"] for v in position.values()] if max(vals) - min(vals) >= 0.1: warnings.append( "position probe: agreement differs by " f"{max(vals) - min(vals):.2f} across presentation positions -- " "run both orders and average, or score absolutely" ) if skipped: warnings.append(f"{len(skipped)} record(s) skipped (missing required fields)") return { "n": n, "kappa": round(kappa, 4) if kappa is not None else None, "band": band, "advice": advice, "raw_agreement": round(agreement, 4), "min_kappa": min_kappa, "calibrated": calibrated, "labels": labels, "confusion": [ {"human": h, "judge": j, "count": c} for (h, j), c in sorted(confusion.items()) ], "per_class": per_class, "disagreements": disagreements, "probes": { "verbosity_correlation": round(verbosity_r, 4) if verbosity_r is not None else None, "position_agreement": position, }, "skipped": skipped, "warnings": warnings, } def print_human(report, source): out = [] out.append(f"judge calibration: {source}") out.append(f" cases {report['n']}") kappa = report["kappa"] out.append( f" cohen kappa {kappa if kappa is not None else 'undefined'} " f"({report['band']})" ) out.append(f" raw agreement {report['raw_agreement']} (kappa is the honest one)") out.append(f" threshold {report['min_kappa']}") out.append("") out.append(" confusion (human -> judge)") for row in report["confusion"]: mark = " " if row["human"] == row["judge"] else "!" out.append(f" {mark} {row['human']:>12} -> {row['judge']:<12} {row['count']}") out.append("") out.append(" per human class") for lab, stats in report["per_class"].items(): recall = stats["recall"] out.append( f" {lab:>12} n={stats['human_count']:<4} " f"agreed={stats['judge_agreed']:<4} recall={recall if recall is not None else '-'}" ) probes = report["probes"] if probes["verbosity_correlation"] is not None or probes["position_agreement"]: out.append("") out.append(" bias probes") if probes["verbosity_correlation"] is not None: out.append(f" verbosity r {probes['verbosity_correlation']:+.4f}") for pos, stats in probes["position_agreement"].items(): out.append(f" position {pos:<8} n={stats['n']:<4} agreement={stats['agreement']}") if report["disagreements"]: out.append("") out.append(f" disagreements ({len(report['disagreements'])})") for d in report["disagreements"][:20]: out.append(f" {d['id']}: human={d['human']} judge={d['judge']}") if len(report["disagreements"]) > 20: out.append(f" ... {len(report['disagreements']) - 20} more") out.append("") out.append(f" verdict {report['advice']}") print("\n".join(out)) def main(argv=None): parser = argparse.ArgumentParser( prog="judge-calibration.py", description="Cohen kappa and bias probes for an LLM judge vs human labels.", add_help=True, ) parser.add_argument("labels", help="JSONL of {human, judge, ...} records, or - for stdin") parser.add_argument( "--min-kappa", type=float, default=0.6, help="kappa at or above which the judge counts as calibrated (default: 0.6)", ) parser.add_argument( "--verbosity-field", default="length", metavar="NAME", help="numeric field carrying output length for the verbosity probe (default: length)", ) parser.add_argument("--json", action="store_true", help="emit the JSON envelope on stdout") try: args = parser.parse_args(argv) except SystemExit as exc: # argparse exits 2 on bad args and 0 on --help; both already match the protocol. raise SystemExit(exc.code) if not (0.0 <= args.min_kappa <= 1.0): print("judge-calibration: --min-kappa must be between 0 and 1", file=sys.stderr) return EXIT_USAGE def fail(code, kind, message): if args.json: print(json.dumps({"error": {"code": kind, "message": message, "details": {}}})) print(f"judge-calibration: {message}", file=sys.stderr) return code try: records = load_records(args.labels) except FileNotFoundError as exc: return fail(EXIT_NOT_FOUND, "NOT_FOUND", f"no such file: {exc}") except ValueError as exc: return fail(EXIT_VALIDATION, "VALIDATION", str(exc)) pairs, ids, skipped, lengths, judge_numeric, position_rows = analyse( records, args.verbosity_field ) if not pairs: return fail( EXIT_VALIDATION, "VALIDATION", "no usable records (every line missing 'human' or 'judge')", ) report = build_report( pairs, ids, skipped, lengths, judge_numeric, position_rows, args.min_kappa ) if args.json: print(json.dumps({ "data": report, "meta": {"count": report["n"], "schema": SCHEMA, "source": args.labels}, }, indent=2)) else: print_human(report, args.labels) for warning in report["warnings"]: print(f"judge-calibration: warning: {warning}", file=sys.stderr) return EXIT_OK if report["calibrated"] else EXIT_UNDER if __name__ == "__main__": sys.exit(main())
-
-
tests
-
run.sh 35.1 KB
#!/usr/bin/env bash # Self-test for evals-ops scripts. # # Offline and deterministic: builds throwaway JSONL fixtures with KNOWN correct # answers (kappa computed by hand below), asserts the documented exit codes and # the actual numbers, then cleans up. Resolves paths relative to itself so it # works in the repo and once installed to ~/.claude/skills/evals-ops/. # # Usage: bash tests/run.sh # Exit: 0 all pass, 1 one or more failures set -uo pipefail HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" SKILL="$(dirname "$HERE")" SCRIPTS="$SKILL/scripts" CAL="$SCRIPTS/judge-calibration.py" AUD="$SCRIPTS/goldenset-audit.py" BAS="$SCRIPTS/eval-baseline.py" # Probe python by EXECUTING it. `command -v python3` finds the Windows Store # app-execution stub, which exists on PATH but exits 49 non-interactively. PYTHON="" for c in python3 python py; do if "$c" -c 'import sys' >/dev/null 2>&1; then PYTHON="$c"; break; fi done if [ -z "$PYTHON" ]; then echo "evals-ops tests: no working python3 found - skipping" >&2 exit 0 fi SB="$(mktemp -d)"; trap 'rm -rf "$SB"' EXIT PASS=0; FAIL=0 ok() { echo " ok $*"; PASS=$((PASS + 1)); } bad() { echo " FAIL $*"; FAIL=$((FAIL + 1)); } # exit_is <want> <label> -- <command...> exit_is() { local want="$1" label="$2"; shift 3 "$@" >/dev/null 2>&1; local got=$? if [ "$got" -eq "$want" ]; then ok "$label (exit $got)"; else bad "$label (want $want, got $got)"; fi } # json_eq <jq-ish python path> <expected> <label> -- <command...> # Reads stdout as JSON and compares a dotted path. Uses python, not jq, so the # suite has no dependency beyond the interpreter it already requires. json_eq() { local path="$1" want="$2" label="$3"; shift 4 local out got out="$("$@" 2>/dev/null)" got="$(printf '%s' "$out" | "$PYTHON" -c ' import json, sys doc = json.load(sys.stdin) for part in sys.argv[1].split("."): doc = doc[int(part)] if part.isdigit() else doc[part] print(doc) ' "$path" 2>/dev/null)" if [ "$got" = "$want" ]; then ok "$label ($path = $got)"; else bad "$label ($path: want $want, got '$got')"; fi } # emits <pattern> <label> -- <command...> # Capture stdout into a variable BEFORE grepping. Piping the script straight into # grep would let `set -o pipefail` surface the script's own domain exit code (10) # as the `if` condition, failing the assertion for the wrong reason. emits() { local pattern="$1" label="$2"; shift 3 local out out="$("$@" 2>/dev/null)" if printf '%s' "$out" | grep -q -- "$pattern"; then ok "$label"; else bad "$label (no '$pattern' in output)"; fi } echo "== evals-ops self-test ($PYTHON)" # --- protocol surface ------------------------------------------------------- exit_is 0 "judge-calibration --help" -- "$PYTHON" "$CAL" --help exit_is 0 "goldenset-audit --help" -- "$PYTHON" "$AUD" --help exit_is 0 "eval-baseline --help" -- "$PYTHON" "$BAS" --help exit_is 2 "judge-calibration no args" -- "$PYTHON" "$CAL" exit_is 2 "goldenset-audit no args" -- "$PYTHON" "$AUD" exit_is 2 "eval-baseline no args" -- "$PYTHON" "$BAS" exit_is 2 "judge-calibration rejects out-of-range --min-kappa" \ -- "$PYTHON" "$CAL" /dev/null --min-kappa 5 exit_is 3 "judge-calibration missing file" -- "$PYTHON" "$CAL" "$SB/nope.jsonl" exit_is 3 "goldenset-audit missing file" -- "$PYTHON" "$AUD" "$SB/nope.jsonl" printf 'not json at all\n' > "$SB/bad.jsonl" exit_is 4 "judge-calibration malformed JSONL" -- "$PYTHON" "$CAL" "$SB/bad.jsonl" exit_is 4 "goldenset-audit malformed JSONL" -- "$PYTHON" "$AUD" "$SB/bad.jsonl" : > "$SB/empty.jsonl" exit_is 4 "goldenset-audit empty set" -- "$PYTHON" "$AUD" "$SB/empty.jsonl" # --- judge-calibration: arithmetic -------------------------------------------- # Perfect agreement on a balanced set: observed 1.0, expected 0.5, kappa = 1.0. { for i in 1 2 3 4 5; do echo "{\"id\":\"p$i\",\"human\":\"pass\",\"judge\":\"pass\"}"; done for i in 1 2 3 4 5; do echo "{\"id\":\"f$i\",\"human\":\"fail\",\"judge\":\"fail\"}"; done } > "$SB/perfect.jsonl" json_eq "data.kappa" "1.0" "kappa = 1.0 on perfect balanced agreement" \ -- "$PYTHON" "$CAL" "$SB/perfect.jsonl" --json exit_is 0 "perfect agreement is calibrated" -- "$PYTHON" "$CAL" "$SB/perfect.jsonl" # Chance-level agreement. human 5 pass / 5 fail; judge 4 pass / 6 fail; 5 agree. # observed 0.5; expected (0.5*0.4)+(0.5*0.6) = 0.5 -> kappa exactly 0.0. { echo '{"id":"a1","human":"pass","judge":"pass"}' echo '{"id":"a2","human":"pass","judge":"pass"}' echo '{"id":"a3","human":"pass","judge":"fail"}' echo '{"id":"a4","human":"pass","judge":"fail"}' echo '{"id":"a5","human":"pass","judge":"fail"}' echo '{"id":"b1","human":"fail","judge":"fail"}' echo '{"id":"b2","human":"fail","judge":"fail"}' echo '{"id":"b3","human":"fail","judge":"fail"}' echo '{"id":"b4","human":"fail","judge":"pass"}' echo '{"id":"b5","human":"fail","judge":"pass"}' } > "$SB/chance.jsonl" json_eq "data.kappa" "0.0" "chance-level agreement scores kappa 0.0" \ -- "$PYTHON" "$CAL" "$SB/chance.jsonl" --json json_eq "data.raw_agreement" "0.5" "chance set raw agreement is 0.5" \ -- "$PYTHON" "$CAL" "$SB/chance.jsonl" --json exit_is 10 "under-calibrated set exits 10" -- "$PYTHON" "$CAL" "$SB/chance.jsonl" # A genuinely moderate judge: 5/5 balanced both sides, 8 of 10 agree. # observed 0.8; expected 0.5 -> kappa exactly 0.6, the gate boundary. { for i in 1 2 3 4; do echo "{\"id\":\"m$i\",\"human\":\"pass\",\"judge\":\"pass\"}"; done echo '{"id":"m5","human":"pass","judge":"fail"}' for i in 6 7 8 9; do echo "{\"id\":\"m$i\",\"human\":\"fail\",\"judge\":\"fail\"}"; done echo '{"id":"m10","human":"fail","judge":"pass"}' } > "$SB/moderate.jsonl" json_eq "data.kappa" "0.6" "moderate judge scores kappa 0.6" \ -- "$PYTHON" "$CAL" "$SB/moderate.jsonl" --json json_eq "data.band" "substantial" "kappa 0.6 lands in the substantial band" \ -- "$PYTHON" "$CAL" "$SB/moderate.jsonl" --json # THE point of using kappa at all: on an imbalanced set a judge that answers # "pass" unconditionally gets 90% RAW agreement and kappa 0.0. If these two ever # report the same number, the chance correction has been broken. { for i in $(seq 1 9); do echo "{\"id\":\"y$i\",\"human\":\"pass\",\"judge\":\"pass\"}"; done echo '{"id":"n1","human":"fail","judge":"pass"}' } > "$SB/imbalanced.jsonl" json_eq "data.raw_agreement" "0.9" "imbalanced set: raw agreement flatters at 0.9" \ -- "$PYTHON" "$CAL" "$SB/imbalanced.jsonl" --json json_eq "data.kappa" "0.0" "imbalanced set: kappa correctly reports 0.0" \ -- "$PYTHON" "$CAL" "$SB/imbalanced.jsonl" --json exit_is 10 "constant judge is not calibrated" -- "$PYTHON" "$CAL" "$SB/imbalanced.jsonl" # Degenerate case: both raters used one identical label, so chance agreement is # 1.0 and kappa is 0/0. Must report null and say so, not divide by zero. { for i in 1 2 3; do echo "{\"id\":\"s$i\",\"human\":\"pass\",\"judge\":\"pass\"}"; done } > "$SB/single.jsonl" json_eq "data.kappa" "None" "single-label sample reports kappa as null" \ -- "$PYTHON" "$CAL" "$SB/single.jsonl" --json json_eq "data.band" "undefined" "single-label sample is banded 'undefined'" \ -- "$PYTHON" "$CAL" "$SB/single.jsonl" --json exit_is 10 "undefined kappa is never treated as calibrated" -- "$PYTHON" "$CAL" "$SB/single.jsonl" # Label normalisation: true/false, 1/0 and "PASS" must compare like their peers. { echo '{"id":"n1","human":true,"judge":"PASS"}' echo '{"id":"n2","human":true,"judge":"pass"}' echo '{"id":"n3","human":false,"judge":"Fail"}' echo '{"id":"n4","human":false,"judge":"fail"}' } > "$SB/norm.jsonl" json_eq "data.kappa" "1.0" "boolean and cased labels normalise to agreement" \ -- "$PYTHON" "$CAL" "$SB/norm.jsonl" --json # --min-kappa is the gate. moderate.jsonl scores exactly 0.6. exit_is 0 "--min-kappa 0.6 accepts kappa 0.6 (inclusive)" -- "$PYTHON" "$CAL" "$SB/moderate.jsonl" --min-kappa 0.6 exit_is 0 "--min-kappa 0.5 accepts kappa 0.6" -- "$PYTHON" "$CAL" "$SB/moderate.jsonl" --min-kappa 0.5 exit_is 10 "--min-kappa 0.8 rejects kappa 0.6" -- "$PYTHON" "$CAL" "$SB/moderate.jsonl" --min-kappa 0.8 # Disagreements are enumerated so a human can go look at them. json_eq "meta.count" "10" "report counts every case" \ -- "$PYTHON" "$CAL" "$SB/chance.jsonl" --json json_eq "data.disagreements.0.id" "a3" "first disagreement is reported by id" \ -- "$PYTHON" "$CAL" "$SB/chance.jsonl" --json # Confusion direction matters more than the headline: a judge that only ever # under-passes is safe to gate on. Recall must be per-human-class, not global. json_eq "data.per_class.fail.recall" "0.0" "per-class recall exposes a judge that never says fail" \ -- "$PYTHON" "$CAL" "$SB/imbalanced.jsonl" --json json_eq "data.per_class.pass.recall" "1.0" "per-class recall is 1.0 for the class it always picks" \ -- "$PYTHON" "$CAL" "$SB/imbalanced.jsonl" --json # Verbosity probe: judge score rises monotonically with length while the human # scored every case identically -- that is the bias, and it must be detected. { echo '{"id":"v1","human":1,"judge":1,"length":10}' echo '{"id":"v2","human":1,"judge":2,"length":100}' echo '{"id":"v3","human":1,"judge":3,"length":200}' echo '{"id":"v4","human":1,"judge":4,"length":300}' echo '{"id":"v5","human":1,"judge":5,"length":400}' } > "$SB/verbose.jsonl" json_eq "data.probes.verbosity_correlation" "0.9998" "verbosity probe detects length correlation" \ -- "$PYTHON" "$CAL" "$SB/verbose.jsonl" --json # stdin path if [ "$("$PYTHON" "$CAL" - --json < "$SB/perfect.jsonl" 2>/dev/null | "$PYTHON" -c 'import json,sys; print(json.load(sys.stdin)["data"]["n"])' 2>/dev/null)" = "10" ]; then ok "reads JSONL from stdin" else bad "reads JSONL from stdin" fi # stdout must stay parseable JSON under --json even when the script exits 10 and # warnings are firing on stderr. Capture first -- pipefail would mask this. _out="$("$PYTHON" "$CAL" "$SB/verbose.jsonl" --json 2>/dev/null)" if printf '%s' "$_out" | "$PYTHON" -c 'import json,sys; json.load(sys.stdin)' >/dev/null 2>&1; then ok "stdout stays pure JSON on a nonzero exit" else bad "stdout stays pure JSON on a nonzero exit" fi # ...and the human report must NOT be JSON, so the two modes cannot be confused. _out="$("$PYTHON" "$CAL" "$SB/moderate.jsonl" 2>/dev/null)" case "$_out" in *"cohen kappa"*) ok "human mode prints a readable report" ;; *) bad "human mode prints a readable report" ;; esac # --- goldenset-audit -------------------------------------------------------- mk_case() { # id bucket printf '{"id":"%s","bucket":"%s","added":"2026-01-0%s","why":"case %s",' "$1" "$2" "$((RANDOM % 9 + 1))" "$1" printf '"input":{"q":"unique question about %s here"},"expected":{"a":"%s"}}\n' "$1" "$1" } { for i in $(seq 1 12); do mk_case "prod-$i" production; done for i in $(seq 1 6); do mk_case "replay-$i" replay; done for i in $(seq 1 5); do mk_case "adv-$i" adversarial; done for i in $(seq 1 4); do mk_case "edge-$i" edge; done } > "$SB/healthy.jsonl" exit_is 0 "healthy balanced set audits clean" -- "$PYTHON" "$AUD" "$SB/healthy.jsonl" json_eq "data.stats.cases" "27" "case count is reported" -- "$PYTHON" "$AUD" "$SB/healthy.jsonl" --json # Duplicate id + identical case content are both errors (exit 10). { cat "$SB/healthy.jsonl"; mk_case "prod-1" production; } > "$SB/dupid.jsonl" exit_is 10 "duplicate id is a finding" -- "$PYTHON" "$AUD" "$SB/dupid.jsonl" emits "DUPLICATE_ID" "duplicate id reports DUPLICATE_ID" -- "$PYTHON" "$AUD" "$SB/dupid.jsonl" --json # A case with neither expected nor criteria is ungradeable. { cat "$SB/healthy.jsonl"; echo '{"id":"x1","bucket":"edge","added":"2026-02-01","why":"w","input":{"q":"z"}}'; } > "$SB/noexp.jsonl" emits "NO_EXPECTATION" "ungradeable case reports NO_EXPECTATION" -- "$PYTHON" "$AUD" "$SB/noexp.jsonl" --json exit_is 10 "ungradeable case is an error" -- "$PYTHON" "$AUD" "$SB/noexp.jsonl" # Bucket skew: 25 production, 1 of everything else -> BUCKET_HEAVY + BUCKET_THIN. { for i in $(seq 1 25); do mk_case "p-$i" production; done mk_case "r-1" replay; mk_case "a-1" adversarial; mk_case "e-1" edge } > "$SB/skew.jsonl" emits "BUCKET_HEAVY" "production-heavy set reports BUCKET_HEAVY" -- "$PYTHON" "$AUD" "$SB/skew.jsonl" --json emits "BUCKET_THIN" "starved buckets report BUCKET_THIN" -- "$PYTHON" "$AUD" "$SB/skew.jsonl" --json # Balance findings are warnings, so the default --fail-on error must NOT trip. exit_is 0 "bucket skew is advisory under --fail-on error" -- "$PYTHON" "$AUD" "$SB/skew.jsonl" exit_is 10 "bucket skew trips --fail-on warn" -- "$PYTHON" "$AUD" "$SB/skew.jsonl" --fail-on warn # Near-duplicate detection, and the --max-pairs escape hatch. { mk_case "u-1" production echo '{"id":"d-1","bucket":"edge","added":"2026-01-01","why":"w","input":{"q":"the quick brown fox jumps over the lazy dog"},"expected":{"a":1}}' echo '{"id":"d-2","bucket":"edge","added":"2026-01-01","why":"w","input":{"q":"the quick brown fox jumps over the lazy dog"},"expected":{"a":2}}' } > "$SB/near.jsonl" emits '"NEAR_DUPLICATE"' "identical inputs report NEAR_DUPLICATE" -- "$PYTHON" "$AUD" "$SB/near.jsonl" --json emits "NEAR_DUPLICATE_SKIPPED" "--max-pairs 0 skips the O(n^2) scan and says so" -- "$PYTHON" "$AUD" "$SB/near.jsonl" --max-pairs 0 --json # --- freeze manifest -------------------------------------------------------- exit_is 0 "--write-freeze on a clean set" -- "$PYTHON" "$AUD" "$SB/healthy.jsonl" --write-freeze "$SB/m.json" [ -f "$SB/m.json" ] && ok "freeze manifest written" || bad "freeze manifest written" [ -f "$SB/m.json.tmp" ] && bad "temp file left behind" || ok "atomic write leaves no .tmp" exit_is 0 "unchanged set matches its freeze" -- "$PYTHON" "$AUD" "$SB/healthy.jsonl" --freeze "$SB/m.json" # Editing a frozen case in place is the cardinal sin -- must be an error. "$PYTHON" - "$SB/healthy.jsonl" "$SB/edited.jsonl" <<'PYEOF' import json, sys src, dst = sys.argv[1], sys.argv[2] lines = [json.loads(l) for l in open(src, encoding="utf-8") if l.strip()] lines[0]["expected"] = {"a": "quietly changed to make it pass"} with open(dst, "w", encoding="utf-8") as fh: for c in lines: fh.write(json.dumps(c) + "\n") PYEOF emits "FREEZE_CASE_CHANGED" "in-place edit of a frozen case reports FREEZE_CASE_CHANGED" -- "$PYTHON" "$AUD" "$SB/edited.jsonl" --freeze "$SB/m.json" --json exit_is 10 "in-place edit fails the freeze check" -- "$PYTHON" "$AUD" "$SB/edited.jsonl" --freeze "$SB/m.json" # Adding a case is a warning (re-baseline), not an error. { cat "$SB/healthy.jsonl"; mk_case "new-1" adversarial; } > "$SB/grown.jsonl" emits "FREEZE_CASE_ADDED" "added case reports FREEZE_CASE_ADDED" -- "$PYTHON" "$AUD" "$SB/grown.jsonl" --freeze "$SB/m.json" --json exit_is 0 "added case is advisory under --fail-on error" -- "$PYTHON" "$AUD" "$SB/grown.jsonl" --freeze "$SB/m.json" # Field order must not change a case hash -- otherwise every reformat is "drift". "$PYTHON" - "$SB/healthy.jsonl" "$SB/reordered.jsonl" <<'PYEOF' import json, sys src, dst = sys.argv[1], sys.argv[2] with open(dst, "w", encoding="utf-8") as fh: for line in open(src, encoding="utf-8"): if line.strip(): c = json.loads(line) fh.write(json.dumps(dict(reversed(list(c.items())))) + "\n") PYEOF exit_is 0 "key reordering does not count as drift" -- "$PYTHON" "$AUD" "$SB/reordered.jsonl" --freeze "$SB/m.json" exit_is 2 "--freeze and --write-freeze are mutually exclusive" \ -- "$PYTHON" "$AUD" "$SB/healthy.jsonl" --freeze "$SB/m.json" --write-freeze "$SB/m2.json" exit_is 3 "missing freeze manifest" -- "$PYTHON" "$AUD" "$SB/healthy.jsonl" --freeze "$SB/absent.json" # --- eval-baseline: noise floor, significance, ceilings --------------------- # History with a KNOWN spread. scores .88 .85 .89 .86 .88 # mean = 0.872 # sample stdev = sqrt(0.00108/4) = 0.016432 -> 0.0164 # 2-sigma gate = 0.872 - 2(0.0164) = 0.8392 for v in 0.88 0.85 0.89 0.86 0.88; do echo "{\"score\":$v,\"n\":30,\"dataset\":\"v3\",\"judge\":\"j1\"}" done > "$SB/hist.jsonl" echo '{"score":0.87,"n":30,"dataset":"v3","judge":"j1","cost_usd":1.20,"p95_ms":3000}' > "$SB/cand.jsonl" json_eq "data.baseline" "0.872" "baseline is the window mean" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --json json_eq "data.noise_floor" "0.0164" "noise floor is the sample stdev" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --json json_eq "data.recommended_threshold" "0.8392" "gate sits 2 sigma below baseline" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --json json_eq "data.verdict" "noise" "a 0.002 drop inside the band is noise" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --json exit_is 0 "noise does not fail the gate" -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" echo '{"score":0.62,"n":30,"dataset":"v3","judge":"j1"}' > "$SB/drop.jsonl" json_eq "data.verdict" "regression" "a drop past the threshold is a regression" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/drop.jsonl" --json exit_is 10 "regression exits 10" -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/drop.jsonl" # --sigma widens the band: 0.83 is below the 2-sigma gate (0.8392) but above the # 4-sigma gate (0.8064), so the same score must classify differently. echo '{"score":0.83,"n":30,"dataset":"v3","judge":"j1"}' > "$SB/mid.jsonl" json_eq "data.verdict" "regression" "0.83 is a regression at the default 2 sigma" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/mid.jsonl" --json json_eq "data.verdict" "noise" "0.83 is noise at 4 sigma (wider band)" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/mid.jsonl" --sigma 4 --json # --- McNemar exact test ----------------------------------------------------- # 30 cases all passing at baseline; 8 now fail. b=8, c=0. # p = 2 * C(8,0)/2^8 = 2/256 = 0.007812 for i in $(seq 1 30); do echo "{\"id\":\"c$i\",\"passed\":true}"; done > "$SB/base-res.jsonl" { for i in $(seq 1 22); do echo "{\"id\":\"c$i\",\"passed\":true}"; done for i in $(seq 23 30); do echo "{\"id\":\"c$i\",\"passed\":false}"; done; } > "$SB/cand-res.jsonl" PAIR=(--baseline-results "$SB/base-res.jsonl" --candidate-results "$SB/cand-res.jsonl") json_eq "data.significance.p_value" "0.007812" "McNemar exact p for b=8,c=0" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${PAIR[@]}" --json json_eq "data.verdict" "regression" "paired test overrides a flat aggregate score" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${PAIR[@]}" --json json_eq "data.significance.regressed.0" "c23" "regressed cases are named, not just counted" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${PAIR[@]}" --json exit_is 10 "significant paired regression exits 10" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${PAIR[@]}" # CHURN, not regression. 8 break and 7 are fixed: b=8, c=7, p = 1.0 exactly. # This is the case references/regression-gating.md calls out: the honest verdict # is "noise", and the value of the run is the enumerated lists, not the verdict. # If this ever flips to "regression", the exact test has been replaced by # something that over-claims - which is the failure mode the whole file warns about. { for i in $(seq 1 15); do echo "{\"id\":\"c$i\",\"passed\":true}"; done for i in $(seq 16 22); do echo "{\"id\":\"c$i\",\"passed\":false}"; done for i in $(seq 23 30); do echo "{\"id\":\"c$i\",\"passed\":true}"; done; } > "$SB/churn-base.jsonl" { for i in $(seq 1 15); do echo "{\"id\":\"c$i\",\"passed\":true}"; done for i in $(seq 16 22); do echo "{\"id\":\"c$i\",\"passed\":true}"; done for i in $(seq 23 30); do echo "{\"id\":\"c$i\",\"passed\":false}"; done; } > "$SB/churn-cand.jsonl" CHURN=(--baseline-results "$SB/churn-base.jsonl" --candidate-results "$SB/churn-cand.jsonl") json_eq "data.significance.p_value" "1.0" "8 broken / 7 fixed is p=1.0, not a regression" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${CHURN[@]}" --json json_eq "data.verdict" "noise" "churn is reported as noise, honestly" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${CHURN[@]}" --json json_eq "data.significance.n_fixed" "7" "...but the 7 fixed cases are still enumerated" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${CHURN[@]}" --json exit_is 0 "churn does not fail the gate" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${CHURN[@]}" # Identical runs: zero discordant pairs, p = 1.0 by definition. json_eq "data.significance.p_value" "1.0" "identical runs give p=1.0" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" \ --baseline-results "$SB/base-res.jsonl" --candidate-results "$SB/base-res.jsonl" --json # Direction matters: the same pair reversed is an improvement, not a regression. REV=(--baseline-results "$SB/cand-res.jsonl" --candidate-results "$SB/base-res.jsonl") json_eq "data.verdict" "improvement" "a significant improvement is labelled as such" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${REV[@]}" --json exit_is 0 "an improvement never fails the gate" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${REV[@]}" # Cases present in only one result set cannot be paired and must be flagged. head -20 "$SB/base-res.jsonl" > "$SB/short-res.jsonl" json_eq "data.significance.unpaired_ids" "10" "unpaired case ids are counted" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" \ --baseline-results "$SB/short-res.jsonl" --candidate-results "$SB/cand-res.jsonl" --json # --- confounded comparisons must warn --------------------------------------- # A dataset or judge change inside the window means the MEASUREMENT moved, not # necessarily the system. Silence here would be the worst possible failure. for v in 0.88 0.85 0.89; do echo "{\"score\":$v,\"dataset\":\"v3\",\"judge\":\"j1\"}"; done > "$SB/mixed.jsonl" echo '{"score":0.86,"dataset":"v4","judge":"j1"}' >> "$SB/mixed.jsonl" _err="$("$PYTHON" "$BAS" "$SB/mixed.jsonl" 2>&1 >/dev/null)" case "$_err" in *"dataset version changed"*) ok "warns when the dataset version changes mid-window" ;; *) bad "warns when the dataset version changes mid-window" ;; esac _err="$("$PYTHON" "$BAS" "$SB/hist.jsonl" --window 2 --candidate "$SB/cand.jsonl" 2>&1 >/dev/null)" case "$_err" in *"noise floor needs"*) ok "warns when the window is too small for a noise floor" ;; *) bad "warns when the window is too small for a noise floor" ;; esac # --- ceilings: absolute, not trend ------------------------------------------ exit_is 10 "cost ceiling breach fails" -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --max-cost-usd 0.50 exit_is 0 "cost under ceiling passes" -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --max-cost-usd 5.00 exit_is 10 "latency ceiling breach fails" -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --max-p95-ms 1000 exit_is 0 "latency under ceiling passes" -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --max-p95-ms 9000 # --- argument validation ---------------------------------------------------- exit_is 2 "paired flags must be given together" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --baseline-results "$SB/base-res.jsonl" exit_is 2 "--alpha must be in (0,1)" -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --alpha 1.5 exit_is 2 "--sigma must be positive" -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --sigma 0 exit_is 2 "--window must be >= 1" -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --window 0 exit_is 3 "missing history file" -- "$PYTHON" "$BAS" "$SB/absent.jsonl" exit_is 4 "malformed history" -- "$PYTHON" "$BAS" "$SB/bad.jsonl" # --- shipped assets must survive the skill's own tools ---------------------- # An asset the skill's own scripts reject is worse than no asset at all. ASSETS="$SKILL/assets" exit_is 0 "example golden set passes its own audit" -- "$PYTHON" "$AUD" "$ASSETS/golden-set.example.jsonl" exit_is 0 "runner template compiles" -- "$PYTHON" -m py_compile "$ASSETS/eval-runner.template.py" [ -f "$ASSETS/judge-rubric.template.md" ] && ok "judge rubric template ships" || bad "judge rubric template ships" [ -f "$ASSETS/eval-gate.template.yml" ] && ok "CI gate template ships" || bad "CI gate template ships" # --- --accept: the hillclimb keep/discard gate ------------------------------ # The whole point of this flag is that it asks a DIFFERENT question from CI. # CI: "did this get worse?" -> noise is fine, exit 0. # Hillclimb: "is this improvement real?" -> noise is NOT good enough to bank. # These four assertions pin that inversion. If a refactor ever makes the noise # case exit 0 under --accept, the loop silently goes back to banking noise. echo '{"score":0.97,"n":30,"dataset":"v3","judge":"j1"}' > "$SB/win.jsonl" exit_is 0 "noise passes CI (did it get worse? no)" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" exit_is 10 "noise is a DISCARD under --accept (is it real? no)" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --accept exit_is 0 "a real improvement is a KEEP under --accept" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/win.jsonl" --accept exit_is 10 "a regression is a DISCARD under --accept" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/drop.jsonl" --accept json_eq "data.hillclimb_decision" "discard" "noise reports decision=discard" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --json json_eq "data.hillclimb_decision" "keep" "a real improvement reports decision=keep" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/win.jsonl" --json # The decision is exposed in the report regardless of the flag, so a caller can # read it without adopting the inverted exit semantics. json_eq "data.hillclimb_decision" "discard" "decision is reported without --accept too" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/drop.jsonl" --json # A significant paired improvement is a KEEP even when the aggregate barely moved - # this is the case a raw score comparison gets wrong in the generous direction. json_eq "data.hillclimb_decision" "keep" "paired significance drives the keep decision" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" \ --baseline-results "$SB/cand-res.jsonl" --candidate-results "$SB/base-res.jsonl" --json # ...and churn (8 broken / 7 fixed, p=1.0) is NOT a keep, despite 7 fixes. json_eq "data.hillclimb_decision" "discard" "churn is never banked as an improvement" \ -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" \ --baseline-results "$SB/churn-base.jsonl" --candidate-results "$SB/churn-cand.jsonl" --json # --- the iterate <-> evals-ops seam is documented in BOTH directions -------- # A one-way cross-reference rots: whichever side gets edited without the other # leaves a dangling claim. These assert the link exists from each end. ITERATE="$SKILL/../iterate/SKILL.md" if [ -f "$ITERATE" ]; then grep -q "eval-baseline.py" "$ITERATE" \ && ok "iterate points at the noise-floor gate" \ || bad "iterate points at the noise-floor gate" grep -q "hillclimbing.md" "$ITERATE" \ && ok "iterate links the hillclimbing reference" \ || bad "iterate links the hillclimbing reference" else ok "iterate skill not present (installed standalone) - seam check skipped" fi grep -q "iterate" "$SKILL/references/hillclimbing.md" \ && ok "hillclimbing defers loop mechanics to iterate" \ || bad "hillclimbing defers loop mechanics to iterate" # =========================================================================== # Regressions found by the 2026-08-31 adversarial review. Every one of these # failed OPEN (exit 0 / silently dropped data) before the fix, which is the # dangerous direction for a gate: it reports "fine" while measuring nothing. # Each assertion names the defect so a future edit cannot quietly restore it. # =========================================================================== # R1. A MEASURED spread of exactly 0.0 is the most informative history possible # (a deterministic metric, or a genuinely stable suite). `if spread:` treated it # as "no data" -- so an unambiguous 0.90 -> 0.70 collapse reported # INSUFFICIENT-DATA and exited 0. for v in 1 2 3 4 5; do echo '{"score":0.90,"dataset":"v3","judge":"j1"}'; done > "$SB/flat.jsonl" echo '{"score":0.70,"dataset":"v3","judge":"j1"}' > "$SB/flat-down.jsonl" echo '{"score":0.95,"dataset":"v3","judge":"j1"}' > "$SB/flat-up.jsonl" echo '{"score":0.90,"dataset":"v3","judge":"j1"}' > "$SB/flat-same.jsonl" json_eq "data.noise_floor" "0.0" "zero-variance history reports a floor of 0.0, not null" \ -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --candidate "$SB/flat-same.jsonl" --json json_eq "data.verdict" "regression" "R1: a drop on a zero-variance baseline is a regression" \ -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --candidate "$SB/flat-down.jsonl" --json exit_is 10 "R1: and it actually fails the gate" \ -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --candidate "$SB/flat-down.jsonl" json_eq "data.verdict" "improvement" "R1: a rise on a zero-variance baseline is an improvement" \ -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --candidate "$SB/flat-up.jsonl" --json exit_is 0 "R1: that improvement is a KEEP under --accept" \ -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --candidate "$SB/flat-up.jsonl" --accept # An identical score is NOT an improvement -- a no-op must never be banked. json_eq "data.hillclimb_decision" "discard" "R1: an unchanged score is never a KEEP" \ -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --candidate "$SB/flat-same.jsonl" --json # R2. `if r.get("id")` is falsy for an integer id of 0, so case 0 was silently # dropped from the paired test -- hiding a real regression. printf '{"id":0,"passed":true}\n{"id":1,"passed":true}\n{"id":2,"passed":true}\n' > "$SB/id0-base.jsonl" printf '{"id":0,"passed":false}\n{"id":1,"passed":true}\n{"id":2,"passed":true}\n' > "$SB/id0-cand.jsonl" json_eq "data.significance.n_regressed" "1" "R2: an integer id of 0 is not dropped" \ -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --baseline-results "$SB/id0-base.jsonl" \ --candidate-results "$SB/id0-cand.jsonl" --json json_eq "data.significance.unpaired_ids" "0" "R2: numeric and string ids pair with each other" \ -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --baseline-results "$SB/id0-base.jsonl" \ --candidate-results "$SB/id0-cand.jsonl" --json # R3. A history row with a null or STRING score was dropped in silence, so a # whole malformed file reported insufficient-data and exited 0 with no clue why. printf '{"score":null}\n{"score":"0.9"}\n{"score":0.88}\n' > "$SB/badscores.jsonl" _err="$("$PYTHON" "$BAS" "$SB/badscores.jsonl" 2>&1 >/dev/null)" case "$_err" in *"non-numeric 'score'"*) ok "R3: non-numeric scores are reported, not swallowed" ;; *) bad "R3: non-numeric scores are reported, not swallowed" ;; esac # R4. The exact McNemar test is O(n) big-int work over 2**n: 133ms at b+c=2000 # but ~100 SECONDS at 20000, i.e. a CI hang. Large inputs must switch method and # say which one they used. json_eq "data.significance.test" "mcnemar-exact" "R4: small discordant sets use the exact test" \ -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --baseline-results "$SB/base-res.jsonl" \ --candidate-results "$SB/cand-res.jsonl" --json "$PYTHON" - "$SCRIPTS/eval-baseline.py" <<'PYEOF' import importlib.util, sys, time spec = importlib.util.spec_from_file_location("eb", sys.argv[1]) m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m) started = time.time() p, method = m.mcnemar(12000, 8000) # 20k discordant pairs elapsed = time.time() - started assert method == "normal-approx", f"want normal-approx above the cutoff, got {method}" assert elapsed < 5, f"20k discordant pairs took {elapsed:.1f}s -- the exact test is back" # The approximation must still agree with the exact test near the cutoff. pe, _ = m.mcnemar(600, 400); pa = 2 * 0 # exact at 1000 pairs assert m.mcnemar(8, 0) == (0.0078125, "exact"), "exact path changed" PYEOF [ $? -eq 0 ] && ok "R4: large discordant sets switch to the normal approximation, fast" \ || bad "R4: large discordant sets switch to the normal approximation, fast" # R5. `_line` and `id` leaked into the content hash, so DUPLICATE_CASE could # never fire on the very thing it exists to catch: one case, two ids. # # THE FIXTURE MUST OMIT input/expected/criteria. case_hash() has two branches and # only the FALLBACK branch -- taken when none of those fields are present -- ever # saw _line and id. The first draft of this test used a fixture WITH `expected`, # so it took the other branch and passed against deliberately re-broken code. # Mutation-testing the assertion is what exposed that; do not "simplify" this # fixture by giving the cases an `expected`, or the test goes inert again. printf '{"id":"a","bucket":"edge","added":"2026-01-01","why":"w","note":"same"}\n' > "$SB/dupmin.jsonl" printf '{"id":"b","bucket":"production","added":"2026-05-05","why":"other","note":"same"}\n' >> "$SB/dupmin.jsonl" emits "DUPLICATE_CASE" "R5: same content under two ids is a duplicate (fallback hash)" \ -- "$PYTHON" "$AUD" "$SB/dupmin.jsonl" --json # And the same on the primary branch, where content fields exist and only the # metadata differs. printf '{"id":"c","bucket":"edge","added":"2026-01-01","why":"w","input":{"q":"same"},"expected":{"a":1}}\n' > "$SB/dupfull.jsonl" printf '{"id":"d","bucket":"production","added":"2026-05-05","why":"other","input":{"q":"same"},"expected":{"a":1}}\n' >> "$SB/dupfull.jsonl" emits "DUPLICATE_CASE" "R5: differing bucket/added/why do not mask a duplicate" \ -- "$PYTHON" "$AUD" "$SB/dupfull.jsonl" --json # R6. bool is a subclass of int, so `"length": true` read as a length of 1.0 and # produced a confident, entirely meaningless verbosity correlation (-0.866). printf '{"id":"a","human":1,"judge":1,"length":true}\n{"id":"b","human":1,"judge":5,"length":false}\n{"id":"c","human":1,"judge":3,"length":true}\n' > "$SB/boollen.jsonl" json_eq "data.probes.verbosity_correlation" "None" "R6: a boolean length yields no correlation" \ -- "$PYTHON" "$CAL" "$SB/boollen.jsonl" --json # R7/R8. The shipped CI template asserted unverified action majors and a script # path that contradicted the one iterate documents. Both are now explicit ADAPT # points; asserting a bare major here again would be the regression. TPL="$SKILL/assets/eval-gate.template.yml" if grep -qE 'uses: actions/[a-z-]+@v[0-9]' "$TPL"; then bad "R7: template asserts an unverified action major (use a flagged placeholder)" else ok "R7: template does not assert unverified action majors" fi grep -q 'EVALS_OPS' "$TPL" \ && ok "R8: template routes script paths through one adaptable variable" \ || bad "R8: template routes script paths through one adaptable variable" # R9. The reference stated bucket TARGETS while the script warned on a wider # BAND, with nothing saying they were different numbers on purpose. grep -q "targets to compose against" "$SKILL/references/golden-datasets.md" \ && ok "R9: the reference distinguishes bucket targets from the warn band" \ || bad "R9: the reference distinguishes bucket targets from the warn band" emits "target 40%-50%" "R9: a bucket warning names the target it is measured against" \ -- "$PYTHON" "$AUD" "$SB/skew.jsonl" --json echo echo "evals-ops: $PASS passed, $FAIL failed" [ "$FAIL" -eq 0 ] || exit 1
-
-
SKILL.md 17.4 KB
--- name: evals-ops description: "Build and run evals for LLM and agent systems: golden datasets, LLM-as-a-judge with bias control, trajectory/step/outcome scoring, adversarial refuters, and CI regression gates. Triggers on: eval, evals, eval harness, golden dataset, golden set, LLM-as-a-judge, judge rubric, judge bias, judge calibration, Cohen kappa, agreement with human labels, pass@k, pass^k, trajectory eval, step-level eval, tool-call accuracy, regression gate, eval CI, blocking vs advisory eval, faithfulness score, DeepEval, Braintrust, Opik, Langfuse, AgentOps, did my prompt change make it worse, is my agent getting better." license: MIT metadata: author: claude-mods related-skills: "testing-ops, claude-api-ops, iterate, loop-ops, fleet-ops" --- # Evals Ops **Evals are the prerequisite, not the polish.** You cannot tune a prompt, a retriever, a compaction strategy or a memory layer without a harness that says whether the change made things better. Teams that skip this ship vibes and learn about regressions from users. This skill is the operational layer: what to measure, how to build the dataset, how to make a judge trustworthy, and how to gate CI on it without teaching everyone to ignore red. ## Route first | The ask | Go to | |---|---| | "What should I even measure?" | [Three levels](#three-levels-of-agent-eval) → `references/eval-taxonomy.md` | | "Where do the test cases come from?" | [Golden set](#the-golden-set) → `references/golden-datasets.md` | | "My judge disagrees with me / is it any good?" | [Judges](#llm-as-a-judge) → `references/llm-judge.md` | | "Verify a finding is real, not plausible" | [Refuters](#adversarial-verification) → `references/adversarial-verification.md` | | "Is my RAG retrieval any good?" | [Retrieval](#retrieval) → `references/retrieval-eval.md` | | "Where do the human labels come from?" | `references/annotation-workflow.md` | | "Should this block the merge?" | [Gating](#regression-gating) → `references/regression-gating.md` | | "Did this change really make it worse?" | [Is the drop real](#is-the-drop-real) | | "Optimise against the eval / run it overnight" | [Hillclimbing](#hillclimbing) → `references/hillclimbing.md` | | "Which platform should we use?" | `references/tooling-landscape.md` | | "Just give me a starting file" | [Assets](#assets) — golden set, rubric, runner, CI gate | ## The 60-second version 1. **Write 20 cases before you write a metric.** A dataset you can eyeball beats a metric you cannot interpret. Grow to 100-300, then *freeze* it. 2. **Prefer a deterministic assertion to any judge.** The JSON parsed, the tool was called with the right argument, the query returned 3 rows - free, instant, zero variance. Reach for a judge only where correctness is genuinely a matter of language. 3. **Score the trajectory, not just the answer.** Most teams only check the final artifact and are surprised when a right answer came from a wrong path that breaks tomorrow. 4. **Calibrate the judge against humans before trusting it.** Cohen kappa, not raw agreement. `scripts/judge-calibration.py` does the arithmetic and the verdict. 5. **Blocking gates must be deterministic. Judge metrics start advisory.** One flaky red permanently devalues the signal. ## Three levels of agent eval Most teams do only the third, then wonder why quality is unpredictable. | Level | Question | Signal | Typical evaluator | |---|---|---|---| | **Outcome** | Is the final artifact correct? | Binary or scored end state | Deterministic assertion, unit test, judge | | **Step** | Was *this* tool call right? | Per-span: tool choice, arg shape, arg values | Schema/argument assertion, span-level judge | | **Trajectory** | Was the *path* sensible? | Sequence, loops, redundancy, cost | Reference-trajectory match, rubric judge | The failure that motivates all three: an agent reaches the right end state by an accidental route - the *lucky pass*. Outcome-only scoring records that as a win, and the same case fails next week when the accident does not recur. Conversely a trajectory-only score punishes a legitimately novel-but-correct path. **Gate on outcome; keep step and trajectory as the diagnostics that tell you why the gate moved.** For multi-turn or stateful agents also report **pass^k** (all k independent runs of the same case succeed) alongside **pass@k** (any of k succeeded). pass@k flatters a non-deterministic agent; pass^k is the number that predicts production. Full treatment: `references/eval-taxonomy.md`. ## The golden set A golden set is a **reviewed, versioned, deliberately frozen** collection of inputs with trusted expected outputs. Frozen matters: a set that grows every sprint cannot tell you whether last week's number moved because the system changed or because the set did. **Composition - four buckets, not one:** | Bucket | Source | Why | |---|---|---| | Production sample | Real traffic, stratified | Keeps the score connected to what users actually do | | Failure replays | Every incident that reached a human | Regression protection; the easiest cases to justify | | Adversarial | Injections, contradictions, refusal-bait | The class both agents and judges fail silently on | | Edge cases | Empty, huge, ambiguous, multilingual | Where deterministic code breaks first | **Sizing:** 20 to start, 100-300 for a working regression set, 200-500 once you have production traffic to sample. Beyond that you are usually buying latency, not signal - add cases when a new *failure class* appears, and record in the case itself why it exists. `scripts/goldenset-audit.py` checks a set for the rot that accumulates: duplicates, a bucket that quietly became 90% of the set, undated cases, and drift from a frozen manifest. ```bash python3 scripts/goldenset-audit.py evals/golden.jsonl --json | jq '.data.findings[]' ``` Depth: `references/golden-datasets.md`. ## LLM-as-a-judge A judge is a measurement instrument. Instruments need calibration, and this one has documented, reproducible biases: | Bias | What it does | Mitigation | |---|---|---| | **Position** | Prefers whichever candidate was shown first | Run both orders and average; or score absolutely, not pairwise | | **Verbosity** | Rates longer answers higher regardless of quality | Separate correctness from style in the rubric; penalise unsupported length | | **Self-preference** | Rates its own model family's output higher | Judge with a different family than the one under test | | **Scale drift** | 1-5 scores cluster and shift between model versions | Binary pass/fail against explicit criteria; pin the judge model version | **Panel vs N-identical.** Three calls to the same judge with the same rubric mostly buys the same bias three times. A **panel with distinct lenses** - one asks "is this supported by the source?", one "does it follow the stated policy?", one "would this reproduce?" - finds failure modes redundancy structurally cannot. Use N-identical only to measure the judge's own variance, which is a different question worth asking once. **When a judge is the wrong tool:** if you can express the criterion as code, do. A judge costs money, adds latency, drifts across model versions, and has variance a regex does not. Judges earn their place on faithfulness, tone, policy compliance, and "is this a reasonable answer to an open question" - nowhere else. **Calibrate before you trust.** Label 50-200 cases by hand, run the judge on the same cases, and compute Cohen kappa (raw agreement lies when classes are imbalanced): ```bash python3 scripts/judge-calibration.py evals/labels.jsonl --min-kappa 0.6 # exit 0 = calibrated; exit 10 = below threshold, fix the rubric before shipping it ``` kappa >= 0.8 production-ready - 0.6-0.8 substantial, usable with care - below 0.6 the rubric is the problem, not the model. Re-sample ~50 fresh cases periodically; judges drift when the underlying model version moves. Depth, including bias-probe design: `references/llm-judge.md`. **Measure the human-human ceiling first.** A judge cannot beat the agreement two people achieve with each other. Two annotators on 30-50 shared cases gives you that number - and if it is below ~0.6, the rubric is ambiguous and every label you produce against it is wasted. How to run the sessions, stratify the sample, and adjudicate disagreements: `references/annotation-workflow.md`. ## Adversarial verification For findings rather than scores - bug reports, audit results, review comments - flip the prompt: **ask the verifier to REFUTE, not to confirm.** "Try to refute this finding; default to refuted if uncertain" kills plausible-but-wrong results that an "is this correct?" prompt waves through, because agreement is the path of least resistance for a model. Then take a majority: run 3 refuters, keep the finding only if at least 2 fail to refute it. Prefer **perspective-diverse** refuters (correctness / security / does-it-actually-reproduce) over three identical skeptics - same reasoning as judge panels. This composes with the parallel-work skills rather than duplicating them: `fleet-ops` and `parallel-ops` own the fan-out mechanics; this skill owns the scoring contract the refuters return. See `references/adversarial-verification.md`. ## Retrieval RAG is the most common eval target and the most commonly mis-measured. Scoring only the final answer averages two independent failures into one uninterpretable number: | | Right context | Wrong context | |---|---|---| | **Answer correct** | Working | **Lucky** - the model knew it anyway; scores as a pass | | **Answer wrong** | **Generation bug** - chunking, prompt, model | **Retrieval bug** - embeddings, index, query rewriting | Record the retrieved chunk ids next to every answer and that opaque score becomes a 2x2 you can assign to a team. Gate on **recall@k** (a precision failure degrades an answer; a recall failure makes a correct one impossible) and on citation-id validity, which is free and catches confident answers attached to unrelated sources. Retrieval is the one place deterministic scoring genuinely dominates - you have ground-truth ids, so skip the judge. The bucket almost everyone omits: **questions the corpus cannot answer.** Without them the suite cannot detect hallucination under retrieval failure, which is what users hit most. Metrics, the six failure classes, and ground-truth construction: `references/retrieval-eval.md`. ## Regression gating The rule that keeps a gate alive: **a blocking check must never be flaky.** | Check | Gate | |---|---| | Deterministic assertions (schema, tool-call, exact match) | **Blocking.** Any failure fails CI. | | Judge metrics, first few weeks | **Advisory.** Post the delta as a PR comment. | | Judge metrics, calibrated (kappa >= 0.6) and variance-measured | **Blocking with a margin** below the rolling baseline | | Cost and p95 latency per case | **Blocking on a ceiling**, advisory on the trend | Set the threshold *below* the baseline by more than the measured noise floor: if the suite scores 0.88 +/- 0.03 across reruns, gate at 0.80, not 0.87. You cannot know the noise floor from a single run - commit a rolling window of run results to git and read the variance off it. That committed history is also what distinguishes "today is noisy" from "today broke". ### Is the drop real? A noise floor tells you the aggregate moved unusually far. It does not tell you *the same cases* moved. Two runs over one frozen set are paired binary outcomes, and the tool for those is **McNemar's exact test** over the discordant pairs only. This matters because a change that breaks 8 cases and fixes 7 moves the headline score by 0.01 - invisible against any noise floor - while silently swapping which 15 things work. The paired view names those 15 cases; a score comparison structurally cannot. Whether 8-vs-7 is *significant* is a separate question - it is not - but knowing which cases flipped is what lets you go and look. ```bash python3 scripts/eval-baseline.py evals/history.jsonl --baseline-results base.jsonl --candidate-results new.jsonl # exit 10 = significant regression, and it names the cases that flipped ``` Attribute **cost and latency per eval run** from the start. An eval suite is the only place you find out the accuracy win cost 4x the tokens, and retrofitting attribution after the harness exists is far more annoying than a `tokens_in` / `tokens_out` / `ms` field per case. Full CI shape and the noise-floor method: `references/regression-gating.md`. ## Hillclimbing Once the harness measures, the obvious move is to optimise against it. That works, and it is also the fastest way to make a good suite useless. **The loop belongs to [`iterate`](../iterate/) - this skill owns what goes wrong.** Every hillclimbing failure is a property of the metric, not of the loop: 1. **Banking noise.** `iterate` keeps a change when the metric beats the previous best - correct for line coverage, a coin flip for an eval score. At 0.88 +/- 0.03, a measured 0.90 is not evidence. Gate the keep decision on the noise floor instead: ```bash python3 scripts/eval-baseline.py history.jsonl --candidate iter.jsonl --accept # exit 0 = KEEP (a real improvement), 10 = DISCARD (noise or worse) ``` `--accept` deliberately inverts the CI meaning of "noise": CI asks *did this get worse*, a hillclimb asks *is this improvement real*. Noise fails the second question. 2. **Overfitting the frozen set.** Split train / validation / held-out before optimising, never show the validation set to whatever proposes changes, and treat held-out as a *budget* you spend at milestones - not a dashboard. 3. **Keeping a champion instead of a frontier.** One aggregate best hides which cases a candidate won. Retaining candidates that are best on at least one case is what stops the loop walling itself into a local optimum. And the eval-design consequence: **a scalar score gives an optimizer nothing to reflect on.** A judge returning `{"reason": ..., "verdict": ...}` can be improved against; one returning `0.4` cannot. That field costs nothing today and is what makes automated optimization tractable later. Splits, the optimizer landscape (GEPA, MIPROv2, APE/ORPO/SPO), the pre-flight checklist, and the extraction trigger for a future `prompt-optimization-ops`: `references/hillclimbing.md`. ## Tooling Trace-level observability and eval scoring have converged into the same products - you are picking one system, not two. Open-source cores worth knowing: DeepEval (pytest-native), MLflow (tracing, eval and prompt versioning in one OSS platform), Opik, Langfuse, Arize Phoenix. Commercial-first: Braintrust (dataset curation for non-engineers), AgentOps, LangSmith, Arize. Honest default: **start with a JSONL file and a 40-line runner.** Adopt a platform when you need shared dataset curation, trace search across production traffic, or scheduled runs - not before. Which-one-when: `references/tooling-landscape.md`. > The landscape moves fast. Treat every version, price and feature claim in that reference > as needing re-verification; it carries a datestamp for exactly that reason. ## Scripts | Script | Use | |---|---| | `scripts/judge-calibration.py` | Judge-vs-human agreement: Cohen kappa, confusion matrix, per-class breakdown, verbosity/position bias probes. Exit 10 = below `--min-kappa`. | | `scripts/goldenset-audit.py` | Golden-set health: duplicates, bucket balance, staleness, freeze-manifest drift. Exit 10 = findings. | | `scripts/eval-baseline.py` | Noise floor from run history, the threshold your gate should use, and McNemar's exact test naming the cases that flipped. Exit 10 = confirmed regression or a cost/latency ceiling breach; `--accept` turns it into a hillclimb keep/discard gate. | All three accept `--help` and `--json`, and are offline and stdlib-only. ```bash python3 scripts/judge-calibration.py labels.jsonl --json | jq '.data.kappa' python3 scripts/goldenset-audit.py golden.jsonl --freeze manifest.json python3 scripts/eval-baseline.py history.jsonl --json | jq '.data.recommended_threshold' ``` ## Assets Copy-and-adapt starting points, so the first hour goes on deciding what to measure rather than on scaffolding: | Asset | What it is | |---|---| | `assets/golden-set.example.jsonl` | 10 worked cases across all four buckets, with `why` and `criteria` filled in | | `assets/eval-runner.template.py` | The 40-line runner this skill tells you to start with - two ADAPT blocks, cost/latency/pass^k pre-wired | | `assets/judge-rubric.template.md` | One-criterion rubric with the bias-counter instructions and the calibration checklist | | `assets/eval-gate.template.yml` | GitHub Actions workflow encoding the tier ladder: deterministic blocks on push, judge advisory on PR, k=3 nightly | ## References - `references/eval-taxonomy.md` - outcome/step/trajectory, lucky pass, pass@k vs pass^k, metric selection - `references/golden-datasets.md` - building, four-bucket composition, sizing, freeze discipline, rot - `references/llm-judge.md` - bias catalog and mitigations, rubric design, panels, calibration method - `references/adversarial-verification.md` - refute-not-confirm, majority thresholds, lens diversity - `references/retrieval-eval.md` - RAG: recall@k, the retrieval-vs-generation split, ground truth, failure classes - `references/annotation-workflow.md` - where human labels come from: the human-human ceiling, sampling, adjudication, drift - `references/regression-gating.md` - blocking vs advisory, noise floor, McNemar, CI shape, cost/latency attribution - `references/hillclimbing.md` - optimising against an eval without destroying it: noise, overfitting, frontiers, optimizers - `references/tooling-landscape.md` - platform comparison with verification datestamps
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.