Claude Skill

evals-ops

Build and run evals for LLM and agent systems: golden datasets, LLM-as-a-judge with bias control, trajectory/step/outcome scoring, adversarial refuters, and CI regression gates. Triggers on: eval, evals, eval harness, golden dataset, golden set, LLM-as-a-judge, judge rubric, judg

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download 0xdarkmatter-claude-mods-skills_evals-ops-3dfaf0b.zip · 74 KB
Part of 0xdarkmatter/claude-mods — 94 skills

Install

skills CLI npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/evals-ops
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
Git git clone https://github.com/0xDarkMatter/claude-mods.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Evals Ops

Evals are the prerequisite, not the polish. You cannot tune a prompt, a retriever, a compaction strategy or a memory layer without a harness that says whether the change made things better. Teams that skip this ship vibes and learn about regressions from users.

This skill is the operational layer: what to measure, how to build the dataset, how to make a judge trustworthy, and how to gate CI on it without teaching everyone to ignore red.

Route first

The ask Go to
"What should I even measure?" Three levels → references/eval-taxonomy.md
"Where do the test cases come from?" Golden set → references/golden-datasets.md
"My judge disagrees with me / is it any good?" Judges → references/llm-judge.md
"Verify a finding is real, not plausible" Refuters → references/adversarial-verification.md
"Is my RAG retrieval any good?" Retrieval → references/retrieval-eval.md
"Where do the human labels come from?" references/annotation-workflow.md
"Should this block the merge?" Gating → references/regression-gating.md
"Did this change really make it worse?" Is the drop real
"Optimise against the eval / run it overnight" Hillclimbing → references/hillclimbing.md
"Which platform should we use?" references/tooling-landscape.md
"Just give me a starting file" Assets — golden set, rubric, runner, CI gate

The 60-second version

  1. Write 20 cases before you write a metric. A dataset you can eyeball beats a metric you cannot interpret. Grow to 100-300, then freeze it.
  2. Prefer a deterministic assertion to any judge. The JSON parsed, the tool was called with the right argument, the query returned 3 rows - free, instant, zero variance. Reach for a judge only where correctness is genuinely a matter of language.
  3. Score the trajectory, not just the answer. Most teams only check the final artifact and are surprised when a right answer came from a wrong path that breaks tomorrow.
  4. Calibrate the judge against humans before trusting it. Cohen kappa, not raw agreement. scripts/judge-calibration.py does the arithmetic and the verdict.
  5. Blocking gates must be deterministic. Judge metrics start advisory. One flaky red permanently devalues the signal.

Three levels of agent eval

Most teams do only the third, then wonder why quality is unpredictable.

Level Question Signal Typical evaluator
Outcome Is the final artifact correct? Binary or scored end state Deterministic assertion, unit test, judge
Step Was this tool call right? Per-span: tool choice, arg shape, arg values Schema/argument assertion, span-level judge
Trajectory Was the path sensible? Sequence, loops, redundancy, cost Reference-trajectory match, rubric judge

The failure that motivates all three: an agent reaches the right end state by an accidental route - the lucky pass. Outcome-only scoring records that as a win, and the same case fails next week when the accident does not recur. Conversely a trajectory-only score punishes a legitimately novel-but-correct path. Gate on outcome; keep step and trajectory as the diagnostics that tell you why the gate moved.

For multi-turn or stateful agents also report pass^k (all k independent runs of the same case succeed) alongside pass@k (any of k succeeded). pass@k flatters a non-deterministic agent; pass^k is the number that predicts production. Full treatment: references/eval-taxonomy.md.

The golden set

A golden set is a reviewed, versioned, deliberately frozen collection of inputs with trusted expected outputs. Frozen matters: a set that grows every sprint cannot tell you whether last week's number moved because the system changed or because the set did.

Composition - four buckets, not one:

Bucket Source Why
Production sample Real traffic, stratified Keeps the score connected to what users actually do
Failure replays Every incident that reached a human Regression protection; the easiest cases to justify
Adversarial Injections, contradictions, refusal-bait The class both agents and judges fail silently on
Edge cases Empty, huge, ambiguous, multilingual Where deterministic code breaks first

Sizing: 20 to start, 100-300 for a working regression set, 200-500 once you have production traffic to sample. Beyond that you are usually buying latency, not signal - add cases when a new failure class appears, and record in the case itself why it exists.

scripts/goldenset-audit.py checks a set for the rot that accumulates: duplicates, a bucket that quietly became 90% of the set, undated cases, and drift from a frozen manifest.

python3 scripts/goldenset-audit.py evals/golden.jsonl --json | jq '.data.findings[]'

Depth: references/golden-datasets.md.

LLM-as-a-judge

A judge is a measurement instrument. Instruments need calibration, and this one has documented, reproducible biases:

Bias What it does Mitigation
Position Prefers whichever candidate was shown first Run both orders and average; or score absolutely, not pairwise
Verbosity Rates longer answers higher regardless of quality Separate correctness from style in the rubric; penalise unsupported length
Self-preference Rates its own model family's output higher Judge with a different family than the one under test
Scale drift 1-5 scores cluster and shift between model versions Binary pass/fail against explicit criteria; pin the judge model version

Panel vs N-identical. Three calls to the same judge with the same rubric mostly buys the same bias three times. A panel with distinct lenses - one asks "is this supported by the source?", one "does it follow the stated policy?", one "would this reproduce?" - finds failure modes redundancy structurally cannot. Use N-identical only to measure the judge's own variance, which is a different question worth asking once.

When a judge is the wrong tool: if you can express the criterion as code, do. A judge costs money, adds latency, drifts across model versions, and has variance a regex does not. Judges earn their place on faithfulness, tone, policy compliance, and "is this a reasonable answer to an open question" - nowhere else.

Calibrate before you trust. Label 50-200 cases by hand, run the judge on the same cases, and compute Cohen kappa (raw agreement lies when classes are imbalanced):

python3 scripts/judge-calibration.py evals/labels.jsonl --min-kappa 0.6
# exit 0 = calibrated;  exit 10 = below threshold, fix the rubric before shipping it

kappa >= 0.8 production-ready - 0.6-0.8 substantial, usable with care - below 0.6 the rubric is the problem, not the model. Re-sample ~50 fresh cases periodically; judges drift when the underlying model version moves. Depth, including bias-probe design: references/llm-judge.md.

Measure the human-human ceiling first. A judge cannot beat the agreement two people achieve with each other. Two annotators on 30-50 shared cases gives you that number - and if it is below ~0.6, the rubric is ambiguous and every label you produce against it is wasted. How to run the sessions, stratify the sample, and adjudicate disagreements: references/annotation-workflow.md.

Adversarial verification

For findings rather than scores - bug reports, audit results, review comments - flip the prompt: ask the verifier to REFUTE, not to confirm. "Try to refute this finding; default to refuted if uncertain" kills plausible-but-wrong results that an "is this correct?" prompt waves through, because agreement is the path of least resistance for a model.

Then take a majority: run 3 refuters, keep the finding only if at least 2 fail to refute it. Prefer perspective-diverse refuters (correctness / security / does-it-actually-reproduce) over three identical skeptics - same reasoning as judge panels.

This composes with the parallel-work skills rather than duplicating them: fleet-ops and parallel-ops own the fan-out mechanics; this skill owns the scoring contract the refuters return. See references/adversarial-verification.md.

Retrieval

RAG is the most common eval target and the most commonly mis-measured. Scoring only the final answer averages two independent failures into one uninterpretable number:

Right context Wrong context
Answer correct Working Lucky - the model knew it anyway; scores as a pass
Answer wrong Generation bug - chunking, prompt, model Retrieval bug - embeddings, index, query rewriting

Record the retrieved chunk ids next to every answer and that opaque score becomes a 2x2 you can assign to a team. Gate on recall@k (a precision failure degrades an answer; a recall failure makes a correct one impossible) and on citation-id validity, which is free and catches confident answers attached to unrelated sources. Retrieval is the one place deterministic scoring genuinely dominates - you have ground-truth ids, so skip the judge.

The bucket almost everyone omits: questions the corpus cannot answer. Without them the suite cannot detect hallucination under retrieval failure, which is what users hit most. Metrics, the six failure classes, and ground-truth construction: references/retrieval-eval.md.

Regression gating

The rule that keeps a gate alive: a blocking check must never be flaky.

Check Gate
Deterministic assertions (schema, tool-call, exact match) Blocking. Any failure fails CI.
Judge metrics, first few weeks Advisory. Post the delta as a PR comment.
Judge metrics, calibrated (kappa >= 0.6) and variance-measured Blocking with a margin below the rolling baseline
Cost and p95 latency per case Blocking on a ceiling, advisory on the trend

Set the threshold below the baseline by more than the measured noise floor: if the suite scores 0.88 +/- 0.03 across reruns, gate at 0.80, not 0.87. You cannot know the noise floor from a single run - commit a rolling window of run results to git and read the variance off it. That committed history is also what distinguishes "today is noisy" from "today broke".

Is the drop real?

A noise floor tells you the aggregate moved unusually far. It does not tell you the same cases moved. Two runs over one frozen set are paired binary outcomes, and the tool for those is McNemar's exact test over the discordant pairs only.

This matters because a change that breaks 8 cases and fixes 7 moves the headline score by 0.01 - invisible against any noise floor - while silently swapping which 15 things work. The paired view names those 15 cases; a score comparison structurally cannot. Whether 8-vs-7 is significant is a separate question - it is not - but knowing which cases flipped is what lets you go and look.

python3 scripts/eval-baseline.py evals/history.jsonl   --baseline-results base.jsonl --candidate-results new.jsonl
# exit 10 = significant regression, and it names the cases that flipped

Attribute cost and latency per eval run from the start. An eval suite is the only place you find out the accuracy win cost 4x the tokens, and retrofitting attribution after the harness exists is far more annoying than a tokens_in / tokens_out / ms field per case. Full CI shape and the noise-floor method: references/regression-gating.md.

Hillclimbing

Once the harness measures, the obvious move is to optimise against it. That works, and it is also the fastest way to make a good suite useless.

The loop belongs to iterate - this skill owns what goes wrong. Every hillclimbing failure is a property of the metric, not of the loop:

  1. Banking noise. iterate keeps a change when the metric beats the previous best - correct for line coverage, a coin flip for an eval score. At 0.88 +/- 0.03, a measured 0.90 is not evidence. Gate the keep decision on the noise floor instead:

    python3 scripts/eval-baseline.py history.jsonl --candidate iter.jsonl --accept
    # exit 0 = KEEP (a real improvement), 10 = DISCARD (noise or worse)
    

    --accept deliberately inverts the CI meaning of "noise": CI asks did this get worse, a hillclimb asks is this improvement real. Noise fails the second question.

  2. Overfitting the frozen set. Split train / validation / held-out before optimising, never show the validation set to whatever proposes changes, and treat held-out as a budget you spend at milestones - not a dashboard.

  3. Keeping a champion instead of a frontier. One aggregate best hides which cases a candidate won. Retaining candidates that are best on at least one case is what stops the loop walling itself into a local optimum.

And the eval-design consequence: a scalar score gives an optimizer nothing to reflect on. A judge returning {"reason": ..., "verdict": ...} can be improved against; one returning 0.4 cannot. That field costs nothing today and is what makes automated optimization tractable later.

Splits, the optimizer landscape (GEPA, MIPROv2, APE/ORPO/SPO), the pre-flight checklist, and the extraction trigger for a future prompt-optimization-ops: references/hillclimbing.md.

Tooling

Trace-level observability and eval scoring have converged into the same products - you are picking one system, not two. Open-source cores worth knowing: DeepEval (pytest-native), MLflow (tracing, eval and prompt versioning in one OSS platform), Opik, Langfuse, Arize Phoenix. Commercial-first: Braintrust (dataset curation for non-engineers), AgentOps, LangSmith, Arize.

Honest default: start with a JSONL file and a 40-line runner. Adopt a platform when you need shared dataset curation, trace search across production traffic, or scheduled runs - not before. Which-one-when: references/tooling-landscape.md.

The landscape moves fast. Treat every version, price and feature claim in that reference as needing re-verification; it carries a datestamp for exactly that reason.

Scripts

Script Use
scripts/judge-calibration.py Judge-vs-human agreement: Cohen kappa, confusion matrix, per-class breakdown, verbosity/position bias probes. Exit 10 = below --min-kappa.
scripts/goldenset-audit.py Golden-set health: duplicates, bucket balance, staleness, freeze-manifest drift. Exit 10 = findings.
scripts/eval-baseline.py Noise floor from run history, the threshold your gate should use, and McNemar's exact test naming the cases that flipped. Exit 10 = confirmed regression or a cost/latency ceiling breach; --accept turns it into a hillclimb keep/discard gate.

All three accept --help and --json, and are offline and stdlib-only.

python3 scripts/judge-calibration.py labels.jsonl --json | jq '.data.kappa'
python3 scripts/goldenset-audit.py golden.jsonl --freeze manifest.json
python3 scripts/eval-baseline.py history.jsonl --json | jq '.data.recommended_threshold'

Assets

Copy-and-adapt starting points, so the first hour goes on deciding what to measure rather than on scaffolding:

Asset What it is
assets/golden-set.example.jsonl 10 worked cases across all four buckets, with why and criteria filled in
assets/eval-runner.template.py The 40-line runner this skill tells you to start with - two ADAPT blocks, cost/latency/pass^k pre-wired
assets/judge-rubric.template.md One-criterion rubric with the bias-counter instructions and the calibration checklist
assets/eval-gate.template.yml GitHub Actions workflow encoding the tier ladder: deterministic blocks on push, judge advisory on PR, k=3 nightly

References

  • references/eval-taxonomy.md - outcome/step/trajectory, lucky pass, pass@k vs pass^k, metric selection
  • references/golden-datasets.md - building, four-bucket composition, sizing, freeze discipline, rot
  • references/llm-judge.md - bias catalog and mitigations, rubric design, panels, calibration method
  • references/adversarial-verification.md - refute-not-confirm, majority thresholds, lens diversity
  • references/retrieval-eval.md - RAG: recall@k, the retrieval-vs-generation split, ground truth, failure classes
  • references/annotation-workflow.md - where human labels come from: the human-human ceiling, sampling, adjudication, drift
  • references/regression-gating.md - blocking vs advisory, noise floor, McNemar, CI shape, cost/latency attribution
  • references/hillclimbing.md - optimising against an eval without destroying it: noise, overfitting, frontiers, optimizers
  • references/tooling-landscape.md - platform comparison with verification datestamps
Files (claude-mods)
  • assets
    • eval-gate.template.yml 6.3 KB
      # Eval CI gate — the tier ladder as a workflow. Copy to .github/workflows/evals.yml.
      #
      # The shape encodes one rule: A BLOCKING CHECK MUST NEVER BE FLAKY. Deterministic
      # assertions block on every push. Judge metrics run on PRs and start ADVISORY --
      # they are promoted to blocking only after judge-calibration.py says kappa >= 0.6
      # and you have measured the noise floor. The full k=3 consistency run and the
      # history append happen nightly on main, where cost and duration are affordable.
      #
      # ADAPT before use:
      #   1. <angle brackets>, the eval commands, and the secret names.
      #   2. ACTION VERSIONS. The `uses:` majors below are placeholders, NOT a
      #      recommendation -- check the current major for each action and pin to a
      #      full commit SHA (`uses: actions/checkout@<sha>  # v5.0.0`). A floating
      #      major is a supply-chain surface; see the supply-chain-defense skill.
      #   3. SCRIPT PATHS. $EVALS_OPS below assumes the skill is installed at
      #      ~/.claude/skills/evals-ops. If you vendor claude-mods into the repo
      #      instead, set it to skills/evals-ops. The scripts are stdlib-only, so
      #      copying the three you use into the repo is also a legitimate option and
      #      removes the dependency entirely.
      #
      # See references/regression-gating.md for why each tier sits where it does.
      
      name: evals
      
      on:
        push:
          branches: ["**"]
        pull_request:
        schedule:
          - cron: "0 3 * * *"   # nightly on main
        workflow_dispatch:
      
      concurrency:
        group: evals-${{ github.ref }}
        cancel-in-progress: true
      
      env:
        DATASET_VERSION: golden-v3     # bump = re-baseline; never compare across it
        EVALS_OPS: $HOME/.claude/skills/evals-ops   # ADAPT: see header note 3
        PYTHON_VERSION: "3.12"
      
      jobs:
        # --- TIER 0: deterministic. Blocking, zero tolerance, every push. ----------
        # No model calls, no judge, no network to a provider. Seconds and $0, which is
        # exactly why it is allowed to fail the build.
        deterministic:
          runs-on: ubuntu-latest
          steps:
            - uses: actions/checkout@vX      # ADAPT: current major, pinned to a SHA
            - uses: actions/setup-python@vX  # ADAPT: current major, pinned to a SHA
              with:
                python-version: ${{ env.PYTHON_VERSION }}
      
            # The golden set is an artifact with its own integrity. Editing a frozen
            # case in place to make it pass is fitting the test to the code -- this
            # step is what catches it.
            - name: Golden-set health and freeze check
              run: |
                python3 "$EVALS_OPS"/scripts/goldenset-audit.py \
                  evals/golden.jsonl --freeze evals/freeze.json
      
            - name: Deterministic assertions
              run: <your-deterministic-eval-command>   # e.g. pytest evals/test_assertions.py
      
        # --- TIER 1/2: judge metrics on a PR subset. Advisory until calibrated. ----
        judge:
          if: github.event_name == 'pull_request'
          runs-on: ubuntu-latest
          # ADAPT: flip to `false` only once judge-calibration.py reports kappa >= 0.6
          # AND you have measured the noise floor. Until then a red here is noise, and
          # three unexplained reds kill the gate socially even while it still enforces.
          continue-on-error: true
          steps:
            - uses: actions/checkout@vX      # ADAPT: current major, pinned to a SHA
              with:
                fetch-depth: 0          # history file needs real history
            - uses: actions/setup-python@vX  # ADAPT: current major, pinned to a SHA
              with:
                python-version: ${{ env.PYTHON_VERSION }}
      
            # Judge calibration is a gate on the JUDGE, not on the code. If the rubric
            # has drifted below threshold, the scores below are not evidence.
            - name: Judge is calibrated
              run: |
                python3 "$EVALS_OPS"/scripts/judge-calibration.py \
                  evals/labels.jsonl --min-kappa 0.6
      
            - name: Run evals (stratified subset, k=1)
              env:
                ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
              run: |
                python3 evals/runner.py evals/golden.jsonl \
                  -k 1 --dataset-version "$DATASET_VERSION" \
                  --out results.jsonl --history /tmp/run.jsonl
      
            # Compare against the committed baseline. Exit 10 = confirmed regression
            # (outside the noise floor); exit 0 = noise or improvement. Naming the
            # flipped cases is what stops people ignoring the result.
            - name: Regression check
              run: |
                python3 "$EVALS_OPS"/scripts/eval-baseline.py \
                  evals/history.jsonl --candidate /tmp/run.jsonl \
                  --baseline-results evals/baseline-results.jsonl \
                  --candidate-results results.jsonl
      
            - name: Comment the delta
              if: always()
              run: <post the eval-baseline.py --json summary as a PR comment>
      
        # --- TIER 3/4: nightly. Full set, k=3, cost and latency, history append. ---
        nightly:
          if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch'
          runs-on: ubuntu-latest
          permissions:
            contents: write             # to commit the appended history row
          steps:
            - uses: actions/checkout@vX      # ADAPT: current major, pinned to a SHA
            - uses: actions/setup-python@vX  # ADAPT: current major, pinned to a SHA
              with:
                python-version: ${{ env.PYTHON_VERSION }}
      
            # k=3 gives pass^k -- what production actually experiences. pass@k flatters
            # a non-deterministic agent and must never be the headline.
            # NOTE: never retry a failing eval to green. Retry-until-pass turns a real
            # regression into a flake report; k=3 reports ALL attempts by design.
            - name: Full eval, k=3
              env:
                ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
              run: |
                python3 evals/runner.py evals/golden.jsonl \
                  -k 3 --dataset-version "$DATASET_VERSION" \
                  --out nightly-results.jsonl --history evals/history.jsonl
      
            # Ceilings, not trends. A trend gate fires on noise; a ceiling encodes a
            # product decision someone actually made.
            - name: Cost and latency ceilings
              run: |
                python3 "$EVALS_OPS"/scripts/eval-baseline.py \
                  evals/history.jsonl --max-cost-usd <2.50> --max-p95-ms <6000>
      
            - name: Commit the history row
              run: |
                git config user.name  "eval-bot"
                git config user.email "eval-bot@users.noreply.github.com"
                git add evals/history.jsonl
                git diff --staged --quiet || git commit -m "chore(evals): nightly run"
                git push
      
    • eval-runner.template.py 6.3 KB
      #!/usr/bin/env python3
      """The 40-line eval runner. Copy into your repo and adapt the two ADAPT blocks.
      
      This is deliberately small and dependency-free. It is the thing to start with
      before adopting a platform (see references/tooling-landscape.md) -- it forces you
      to decide what you are measuring, which is the hard part no platform does for you.
      
      What it does:
        - reads a golden set (JSONL, one case per line)
        - runs each case k times through YOUR system
        - applies deterministic assertions first, judge criteria only where needed
        - records cost and latency PER CASE from the first run, not retrofitted later
        - writes a per-case results file and appends one summary row to a run history
      
      Usage:   eval-runner.py GOLDEN.jsonl [-k 3] [--out results.jsonl] [--history history.jsonl]
      Exit:    0 ran to completion, 1 a case raised, 2 usage
      
      The summary row is what eval-baseline.py consumes to tell noise from regression.
      """
      
      import argparse
      import json
      import time
      
      
      # ==========================================================================
      # ADAPT BLOCK 1 -- call your system.
      # Return a dict. Include token counts and the model id; you will want them and
      # threading them in later means touching every runner, result and dashboard.
      # ==========================================================================
      def run_system(case):
          # from my_app import agent
          # result = agent.invoke(case["input"])
          # return {"output": result.text, "tool_calls": result.tool_calls,
          #         "tokens_in": result.usage.input_tokens,
          #         "tokens_out": result.usage.output_tokens,
          #         "model": result.model}
          raise NotImplementedError("wire this to your agent")
      
      
      # ==========================================================================
      # ADAPT BLOCK 2 -- call your judge, ONLY for criteria code cannot check.
      # Return True/False per criterion. Pin the judge model version: an unpinned
      # judge silently re-baselines your whole history (references/llm-judge.md).
      # ==========================================================================
      JUDGE_MODEL = "<pin-a-specific-model-version-here>"
      
      
      def judge(criterion, case, result):
          # from my_app import judge_client
          # verdict = judge_client.score(model=JUDGE_MODEL, temperature=0,
          #                              criterion=criterion, output=result["output"])
          # return verdict["pass"]
          raise NotImplementedError("wire this to your judge, or drop criteria entirely")
      
      
      def deterministic(case, result):
          """Cheap assertions run first and are the ones that gate CI.
      
          Every criterion you can express here instead of in `judge` removes cost,
          latency AND variance at once. Re-audit periodically -- rubric items become
          codifiable once the output format stabilises.
          """
          expected = case.get("expected")
          if not expected:
              return None  # nothing deterministic to check; criteria carry this case
          if "tool" in expected:
              calls = result.get("tool_calls") or []
              if not any(c.get("name") == expected["tool"] for c in calls):
                  return False
              if "args" in expected:
                  got = next(c for c in calls if c.get("name") == expected["tool"])
                  for key, want in expected["args"].items():
                      if got.get("args", {}).get(key) != want:
                          return False
              return True
          if expected.get("refused") or expected.get("clarifies"):
              # Absence of a tool call is the check; the wording is a judge question.
              return not (result.get("tool_calls") or [])
          return result.get("output") == expected.get("output")
      
      
      def evaluate(case, k):
          """Run one case k times. Returns the per-case record."""
          runs = []
          for _ in range(k):
              started = time.monotonic()
              result = run_system(case)
              elapsed_ms = int((time.monotonic() - started) * 1000)
      
              det = deterministic(case, result)
              crit = {c: judge(c, case, result) for c in case.get("criteria", [])} \
                  if det is not False else {}
              passed = bool(det is not False and all(crit.values()) if crit else det)
      
              runs.append({
                  "passed": passed,
                  "deterministic": det,
                  "criteria": crit,
                  "ms": elapsed_ms,
                  "tokens_in": result.get("tokens_in"),
                  "tokens_out": result.get("tokens_out"),
                  "model": result.get("model"),
              })
      
          return {
              "id": case.get("id"),
              "bucket": case.get("bucket"),
              # pass@1 is the honest headline; pass^k is what production experiences.
              "passed": runs[0]["passed"],
              "pass_hat_k": all(r["passed"] for r in runs),
              "pass_at_k": any(r["passed"] for r in runs),
              "runs": runs,
          }
      
      
      def main():
          ap = argparse.ArgumentParser(description="Minimal golden-set eval runner.")
          ap.add_argument("golden")
          ap.add_argument("-k", type=int, default=1, help="runs per case (3 to report pass^k)")
          ap.add_argument("--out", default="results.jsonl")
          ap.add_argument("--history", default="history.jsonl")
          ap.add_argument("--dataset-version", default="unversioned",
                          help="never compare scores across dataset versions")
          args = ap.parse_args()
      
          cases = [json.loads(l) for l in open(args.golden, encoding="utf-8")
                   if l.strip() and not l.startswith("#")]
          results = [evaluate(c, args.k) for c in cases]
      
          with open(args.out, "w", encoding="utf-8") as fh:
              for r in results:
                  fh.write(json.dumps(r) + "\n")
      
          n = len(results)
          summary = {
              "dataset": args.dataset_version,
              "judge": JUDGE_MODEL,
              "n": n,
              "k": args.k,
              "score": round(sum(r["passed"] for r in results) / n, 4),
              "pass_hat_k": round(sum(r["pass_hat_k"] for r in results) / n, 4),
              "tokens_in": sum(run["tokens_in"] or 0 for r in results for run in r["runs"]),
              "tokens_out": sum(run["tokens_out"] or 0 for r in results for run in r["runs"]),
              "p95_ms": sorted(run["ms"] for r in results for run in r["runs"])[int(0.95 * n * args.k)],
              # Stamp `date` from CI (git commit date / job start), not from the runner,
              # so a re-run of an old commit does not claim to be today's measurement.
          }
          with open(args.history, "a", encoding="utf-8") as fh:
              fh.write(json.dumps(summary) + "\n")
      
          print(json.dumps(summary, indent=2))
      
      
      if __name__ == "__main__":
          main()
      
    • golden-set.example.jsonl 5.9 KB · in bundle
    • judge-rubric.template.md 3.4 KB
      # Judge Rubric — template
      
      Copy per criterion. **One criterion per rubric**, one rubric per judge call. A
      rubric asking "is it accurate, helpful and well-written?" returns an unactionable
      blend; three separate binary questions return three actionable answers.
      
      Adapt everything in `<angle brackets>`. Delete the guidance comments before use.
      
      ---
      
      ## Metadata (keep with the rubric, in git)
      
      | Field | Value |
      |---|---|
      | Criterion name | `<faithfulness>` |
      | Judge model | `<pinned-model-version>` — never "latest" |
      | Temperature | `0` |
      | Output | binary `pass` / `fail` |
      | Rubric version | `<v1>` — bump on any edit, and re-baseline |
      | Calibrated | `<kappa, date, n>` — from `judge-calibration.py` |
      
      > A rubric edit is a re-baselining event, exactly like a judge upgrade or a
      > dataset version bump. Scores across the boundary are not comparable.
      
      ---
      
      ## The prompt
      
      ```
      You are grading one criterion. Answer only about this criterion; ignore every
      other quality of the response.
      
      CRITERION
      <Every factual claim in the response appears in the SOURCE below.>
      
      <!-- Concrete and checkable. "is accurate" is not a criterion, it is a mood. -->
      
      SOURCE
      <<<
      {source}
      >>>
      
      RESPONSE
      <<<
      {response}
      >>>
      
      RULES
      - Length is not evidence of quality. A correct, brief response scores the same
        as a correct, padded one.
        <!-- verbosity-bias counter; measure whether it works with
             judge-calibration.py --verbosity-field -->
      - Style, tone and formatting are NOT part of this criterion.
        <!-- separates correctness from style so the gate rides on correctness -->
      - Alternative correct answers are acceptable. Do not penalise a response for
        differing from any reference you may infer.
        <!-- anchoring counter -->
      - If the evidence is genuinely ambiguous, answer "fail" and say why.
        <!-- forces uncertainty to resolve one way, deliberately; pick the direction
             that is safe for YOUR gate -- see the note below -->
      
      EXAMPLES
      pass: <a short worked example of a response that satisfies the criterion>
      fail: <a short worked example that plausibly looks fine but does not>
      fail: <a second failing example covering a different failure mode>
      <!-- two or three worked failures do more for agreement than a page of prose -->
      
      Think through the evidence first, then answer.
      Return JSON only: {"reason": "<one sentence citing the specific evidence>",
                         "verdict": "pass" | "fail"}
      ```
      
      ---
      
      ## Which way should ambiguity resolve?
      
      Decide deliberately, per criterion, and write it down:
      
      | Gate consequence | Resolve ambiguity to |
      |---|---|
      | A false pass ships a bug | `fail` — under-passing is safe; the judge only over-reports work |
      | A false fail blocks a good PR and erodes trust in the gate | `pass`, and keep the metric advisory until κ improves |
      
      Read the direction off the confusion matrix after calibration, not off intuition:
      a judge at κ 0.65 that errs only toward `fail` is gateable; the same κ erring
      toward `pass` is not.
      
      ---
      
      ## Before this rubric gates anything
      
      1. Sample 50–200 cases, stratified across buckets **and across the judge's own
         verdicts** — include cases it passes and cases it fails, or you can only
         measure one error direction.
      2. Label them by hand against this exact rubric.
      3. `python3 scripts/judge-calibration.py labels.jsonl --min-kappa 0.6`
      4. Below 0.6, the rubric is the problem. Rewrite the criterion or add worked
         failure examples — do not reach for a bigger judge model.
      5. Record κ, n and the date in the metadata table above.
      
  • references
    • adversarial-verification.md 5.2 KB
      # Adversarial Verification — refute, do not confirm
      
      Scoring answers a graded question ("how good is this?"). Verification answers a binary one
      ("is this finding real?"). The second is where a naive prompt does the most damage, because
      agreeing is the path of least resistance for a language model.
      
      ## The inversion
      
      Compare two prompts over the same finding:
      
      ```
      ❌  "Is this bug report correct?"
      ✅  "Try to refute this bug report. Default to refuted if uncertain."
      ```
      
      The first invites confirmation and gets it — plausible-sounding findings sail through
      because nothing in the prompt rewards saying no. The second makes the model argue against
      the finding, so a finding only survives if the refutation attempt genuinely fails.
      
      The asymmetry is deliberate: **uncertainty must resolve to "refuted", not "confirmed".** A
      verifier that cannot decide has not verified anything, and treating that as a pass is how
      plausible-but-wrong results reach a report.
      
      Prompt shape that works:
      
      ```
      Finding: <claim, with file:line or a concrete reproduction>
      Your job is to REFUTE this finding. Look for reasons it is wrong, already
      handled elsewhere, unreachable in practice, or based on a misreading.
      If you cannot decisively refute it, say so — but default to refuted when
      the evidence is ambiguous.
      Return {"refuted": true|false, "reason": "..."}.
      ```
      
      ## Majority-refute thresholds
      
      One refuter is a coin flip with an opinion. Run N and take a majority:
      
      | N | Keep the finding if | Character |
      |---|---|---|
      | 1 | not refuted | Cheap triage only |
      | **3** | **at least 2 fail to refute** | The default. Good precision/cost balance |
      | 5 | at least 3 fail to refute | High-stakes; noticeably slower |
      
      Tighten toward unanimity when a false positive is expensive (a finding that will be posted
      publicly, or acted on automatically); loosen when a false negative is expensive (a security
      sweep where missing a real issue costs more than investigating a dud).
      
      Record the vote, not just the verdict — a 2/3 survival is materially weaker evidence than
      3/3 and the consumer of the report deserves to know which they have.
      
      ## Lens diversity beats redundancy
      
      Three identical refuters share their blind spots: whatever the first one fails to notice,
      the other two also fail to notice. Give each refuter a **different lens** and their failure
      modes stop overlapping:
      
      | Lens | Asks |
      |---|---|
      | **Correctness** | Is the described behaviour actually what the code does? |
      | **Reachability** | Can this state be reached by any real input, or is it guarded upstream? |
      | **Reproduction** | Given the stated inputs, does the described failure actually occur? |
      | **Prior art** | Is this already handled — a caller-side check, a test, a framework guarantee? |
      | **Security** (where relevant) | Is there an exploit path, or is this only a code smell? |
      
      Same reasoning as judge panels (`llm-judge.md`): diversity finds classes of error that
      redundancy structurally cannot.
      
      ## Where this fits in a pipeline
      
      The canonical shape — find wide, verify hard, keep little:
      
      ```
      find (N parallel finders, different angles)
        → dedupe against everything seen so far   ← plain code, not an agent
        → refute (3 diverse lenses per finding)
        → keep majority survivors
        → repeat until K consecutive rounds find nothing new
      ```
      
      Two details that decide whether this converges:
      
      - **Dedupe against `seen`, not against `confirmed`.** If refuter-rejected findings are not
        added to `seen`, the finders resurface them every round and the loop never terminates.
      - **Loop until dry, not until a count.** "Find 10 bugs" stops at 10 whether or not there are
        11; "stop when two consecutive rounds surface nothing new" finds the tail.
      
      The fan-out mechanics — process isolation, worktrees, journals, model selection per stage —
      belong to the parallel-work skills, not here. See `fleet-ops` for the landing discipline and
      `parallel-ops` as the router for that family. This skill owns the *contract*: what a refuter
      is asked, what it returns, and how votes become a verdict.
      
      ## Verifier output contract
      
      Keep it small and structured so votes aggregate mechanically:
      
      ```json
      {"refuted": true,
       "confidence": "high",
       "lens": "reachability",
       "reason": "The caller validates `id` against the allowlist at handler.ts:41 before this path."}
      ```
      
      - `refuted` boolean, never a score — the whole point is a decisive vote.
      - `reason` mandatory and specific. A refutation without a citable reason is an opinion, and
        should be treated as a non-refutation when you audit the run.
      - `lens` recorded so you can later ask which lens is earning its cost. Lenses that never
        refute anything across many runs are candidates for removal.
      
      ## When not to bother
      
      Adversarial verification costs N× per finding. Skip it when:
      
      - The finding is **mechanically checkable** — run the test, run the type-checker, run the
        query. A deterministic check beats any number of refuters.
      - The finding is **cheap to act on and cheap to revert** — a lint fix does not need a
        tribunal.
      - You are **scoring quality, not adjudicating truth**. Use a judge (`llm-judge.md`); refuters
        answer a binary question and will flatten a graded one.
      
      ## Cross-reference
      
      - Graded scoring instead of binary adjudication: `llm-judge.md`
      - Using survival rates as an eval metric over time: `regression-gating.md`
      
    • annotation-workflow.md 6 KB
      # Annotation Workflow — where the human labels come from
      
      `llm-judge.md` says "label 50–200 cases by hand" and moves on. This is that step.
      It is the least glamorous part of an eval stack and the one that determines whether
      every number downstream means anything.
      
      ## The ceiling nobody measures
      
      **A judge cannot beat the agreement two humans achieve with each other.** If your
      annotators agree with one another at κ 0.65, a judge scoring κ 0.65 against one of
      them is performing *at the human ceiling*, and chasing 0.8 is chasing noise in your
      own labels.
      
      So measure the ceiling first:
      
      1. Have **two people independently label the same 30–50 cases**, blind to each
         other.
      2. Compute human-human κ (`judge-calibration.py` does not care which rater is
         which — feed it `{"human": rater_a, "judge": rater_b}`).
      3. That number is your realistic target, and the diagnosis when it is low.
      
      **Human-human κ below ~0.6 means the rubric is ambiguous, not that your annotators
      are bad.** Fix the rubric before labelling another case — every label produced
      against an ambiguous rubric is wasted work, and a judge calibrated against it will
      inherit the ambiguity as apparent noise.
      
      This step is skipped almost universally, and it is why so many teams conclude "LLM
      judges are unreliable" when what they actually have is an under-specified criterion.
      
      ## Who labels
      
      | Labeller | Good for | Watch out for |
      |---|---|---|
      | **The engineer who built it** | Fast bootstrapping, catching obvious breakage | Knows what the system *meant* to do and scores it charitably. Never the sole labeller for a gate |
      | **A domain expert** | Anything where correctness is a matter of policy, medicine, law, finance | Scarce and expensive — spend their time on the ambiguous cases, not the obvious ones |
      | **A second engineer** | The human-human ceiling measurement | Shares the team's blind spots |
      | **Crowdsourced** | High-volume, low-context judgments | Needs a much tighter rubric and gold-standard trap questions |
      
      The practical shape for most teams: **the engineer labels everything, a domain
      expert labels the 30 hardest, and the disagreements between them are the most
      valuable output of the whole exercise** — each one is either a rubric ambiguity or
      a genuine product decision nobody had made yet.
      
      ## Sampling: do not label a random slice
      
      A uniform random sample of production traffic is mostly easy cases, and gives you a
      calibration set that cannot measure the judge where it matters.
      
      **Stratify across two axes:**
      
      1. **Bucket** — production, replay, adversarial, edge (`golden-datasets.md`).
      2. **The judge's own verdict** — include cases it passes *and* cases it fails, in
         meaningful numbers. Sampling only cases the judge failed measures one error
         direction and leaves you blind to false passes, which are the dangerous kind.
      
      Add a third axis where you have one: cases near the judge's decision boundary
      (low-confidence, or where a rerun flipped the verdict) are worth several times an
      unambiguous case each.
      
      ## Running a labelling session
      
      - **Label blind to the judge's verdict.** Showing it first anchors the human onto
        it and inflates apparent agreement — you will measure compliance, not agreement.
      - **One criterion at a time, across all cases.** Labelling case-by-case across five
        criteria drifts; labelling criterion-by-criterion keeps the standard fixed.
      - **Record the reason, not just the verdict.** A label without a reason cannot be
        audited later, and reasons are what you mine to rewrite an ambiguous rubric.
      - **Timebox and batch.** Annotation quality falls off a cliff past ~45 minutes.
        50 cases in two sittings beats 100 in one.
      - **Keep an explicit `unsure` option** — and then *do not* let unsure labels vote.
        Forcing a binary on a genuinely ambiguous case manufactures noise that looks like
        judge error. A high unsure rate is a rubric finding.
      
      ## Adjudication
      
      Disagreements are the product, not a problem to be averaged away.
      
      1. Both labellers state their reason.
      2. Classify the disagreement:
         - **Rubric ambiguity** → rewrite the criterion, add a worked example, re-label
           the affected cases.
         - **Genuine product ambiguity** ("should the agent refuse this?") → escalate. A
           decision gets made, and it becomes a rubric example.
         - **Simple error** → correct it and move on.
      3. **Never resolve by majority vote without reading the reasons.** A 2–1 split
         where the minority is right is exactly the case that most improves the rubric.
      
      ## Keeping labels fresh
      
      Labels rot in three ways, and each has a different tell:
      
      | Rot | Tell | Fix |
      |---|---|---|
      | **Judge drift** | κ falls with no rubric change | The judge model version moved. Pin it; re-calibrate |
      | **Distribution drift** | Judge κ holds on the calibration set but production complaints rise | The calibration set no longer resembles traffic. Re-sample |
      | **Annotator drift** | The same person labels the same case differently months apart | Real, and normal. Include ~10% repeats from earlier sessions to measure it |
      
      That last trick is worth adopting: silently re-include a handful of previously
      labelled cases in each session. Intra-annotator agreement below the
      inter-annotator ceiling means the standard itself is sliding, and no amount of
      judge tuning will fix it.
      
      **Cadence:** re-sample ~50 fresh cases periodically, and *always* re-calibrate on a
      judge model change or a rubric edit — both are re-baselining events
      (`regression-gating.md`).
      
      ## Cost, honestly
      
      At roughly 1–3 minutes per case per criterion, 150 cases × 2 criteria × 2
      annotators is around 10–15 person-hours to establish a calibrated judge. That is
      the real price, and it is worth naming up front — a team that budgets for "run the
      calibration script" and not for the labelling will quietly skip the labelling and
      gate CI on an uncalibrated judge.
      
      The saving is that it is mostly one-off. Maintenance is ~50 cases periodically,
      which is an hour or two.
      
      ## Cross-reference
      
      - What the labels are for: `llm-judge.md`
      - Which cases to draw from: `golden-datasets.md`
      - The script that consumes them: `scripts/judge-calibration.py`
      
    • eval-taxonomy.md 5.4 KB
      # Eval Taxonomy — outcome, step, trajectory
      
      The three levels an agent can be scored at, what each one catches that the others miss,
      and how to pick metrics that survive contact with a non-deterministic system.
      
      ## The three levels
      
      ### Outcome-level
      
      **Question:** is the final artifact correct?
      
      The cheapest and most defensible level, and the one everybody starts with. SWE-bench
      established binary pass/fail on the produced patch as the standard for coding agents; the
      same shape works for any agent with a checkable end state — a booking exists, the file
      parses, the database row has the right value.
      
      Evaluate deterministically where you can:
      
      ```python
      assert json.loads(out)["status"] == "confirmed"
      assert db.query("SELECT count(*) FROM orders WHERE id=?", oid) == 1
      ```
      
      **What it misses:** *how* the answer was reached. Which is most of what determines whether
      it will be reached again.
      
      ### Step-level
      
      **Question:** was this individual action right?
      
      Scored per span: was the right tool chosen, was the argument schema valid, were the
      argument *values* correct, did the agent read before it wrote. Step-level scoring is where
      you catch the agent that calls `search` five times with near-identical queries, or that
      passes a plausible-but-wrong ID.
      
      Most step assertions are deterministic and belong in the blocking tier:
      
      | Assertion | Cost |
      |---|---|
      | Tool `X` was called at least once | free |
      | Every tool call validated against its JSON schema | free |
      | No tool called with a value absent from the input context (hallucinated arg) | free |
      | Read-before-write ordering held | free |
      
      Reserve a span-level judge for the genuinely fuzzy step questions ("was this a reasonable
      query to issue given what the agent knew?").
      
      ### Trajectory-level
      
      **Question:** was the *path* sensible?
      
      Scored over the whole nested span tree: sequence, redundancy, loop detection, recovery
      after an error, total cost. Two forms:
      
      1. **Reference-trajectory match** — compare against a known-good path. Exact-match is too
         brittle for anything real; use an ordered-subsequence match ("these 4 steps appeared in
         this order, extras allowed") or set-containment on the essential tool calls.
      2. **Rubric judge over the trace** — hand the serialized trajectory to a judge with a
         rubric ("did it loop? did it retry the same failing call? did it ask the user something
         it could have looked up?").
      
      Trajectory scores are diagnostics, not gates. A novel correct path scores badly against a
      reference and is not a regression.
      
      ## The lucky pass
      
      The single argument for scoring more than the outcome. An agent reaches the correct end
      state by an accidental route — it guessed an ID that happened to be right, it retried until
      a flaky tool succeeded, it hard-coded something that matched this case's expected output.
      Outcome-only scoring banks that as a pass. The same case fails next week, and because the
      suite was green the whole time, nobody knows when the real breakage started.
      
      **Detection is cheap once you have traces:** a passing case whose trajectory contains a
      loop, an error-then-retry, or a tool call with an argument that appears nowhere in the
      input is a lucky-pass candidate. Flag them; do not fail on them. A `lucky_pass_suspects`
      count trending upward on a green suite is one of the highest-value signals in the harness.
      
      ## pass@k vs pass^k
      
      For any non-deterministic agent, a single run per case is a coin flip you are reporting as
      a measurement.
      
      | Metric | Definition | What it tells you |
      |---|---|---|
      | **pass@1** | One run, did it pass | The honest headline number |
      | **pass@k** | Any of k runs passed | Ceiling / "is this reachable at all" |
      | **pass^k** | *All* k runs passed | Consistency — what production actually experiences |
      
      pass@k flatters. An agent that succeeds 1 time in 4 has pass@4 near 1.0 and is unusable.
      **Report pass^k whenever consistency matters** (customer-facing, transactional, anything
      where a retry costs the user something). The tau-bench family popularised this framing for
      multi-turn tool-using agents and it generalises.
      
      Practical: k=3 is usually enough to expose the difference and cheap enough to run per-PR on
      a subset. Run the full k on a nightly, a k=1 pass on every PR.
      
      ## Choosing metrics
      
      Start from the failure you actually fear, not from a metrics catalog.
      
      | You fear | Measure | Level |
      |---|---|---|
      | Made-up facts | Faithfulness: every claim traceable to a retrieved chunk | Outcome (judge) |
      | Wrong tool / wrong args | Tool-call accuracy, arg-schema validity | Step (deterministic) |
      | Burning tokens | Steps per task, redundant-call rate, cost per case | Trajectory (deterministic) |
      | Silent policy violations | Policy-compliance rubric over the trace | Trajectory (judge) |
      | Flaky success | pass^3 | Outcome (deterministic, repeated) |
      | Broke something that used to work | Failure-replay bucket pass rate | Outcome (deterministic) |
      
      Two rules that save more time than any metric choice:
      
      - **A metric nobody can act on is a metric nobody will maintain.** If a score drops and the
        team cannot name what to change, delete the metric or make it decomposable.
      - **Deterministic first.** Every criterion you move from a judge into code removes cost,
        latency, and variance simultaneously. Re-audit periodically: rubric items often become
        codifiable once the output format stabilises.
      
      ## Cross-reference
      
      - Dataset construction: `golden-datasets.md`
      - Judge design for the fuzzy levels: `llm-judge.md`
      - Which of these gate CI: `regression-gating.md`
      
    • golden-datasets.md 6.8 KB
      # Golden Datasets — build it, freeze it, keep it from rotting
      
      The golden set is the most valuable artifact in an eval stack. The model, the prompt and
      the framework will all be replaced; the dataset outlives them and is the only thing that
      lets you compare across the replacements.
      
      ## What it is
      
      A **reviewed, versioned, frozen** set of inputs paired with trusted expected outputs (or
      trusted grading criteria, where the output is open-ended). "Reviewed" means a human agreed
      the expected output is right. "Versioned" means it lives in git next to the code. "Frozen"
      is the part teams skip and then regret.
      
      ## Why frozen beats growing
      
      A set that grows every sprint cannot answer the only question a regression suite exists to
      answer: *did my change make things worse?* If the score moved from 0.84 to 0.79 and the set
      gained 30 cases, the two variables are confounded and no amount of analysis separates them.
      
      The discipline:
      
      - **Freeze v1.** Tag it. Every run reports against a named dataset version.
      - **New cases go to a staging set** and are promoted in explicit, dated batches.
      - **On promotion, re-baseline** — run the current system against v2 and record the new
        reference number. Never compare a v1 score to a v2 score.
      - **Never edit a case in place to make it pass.** That is fitting the test to the code. If
        a case's expected output was genuinely wrong, delete it and add a new one with a note.
      
      Ship a freeze manifest so drift is detectable rather than discovered:
      
      ```bash
      python3 scripts/goldenset-audit.py golden.jsonl --write-freeze manifest.json  # at freeze
      python3 scripts/goldenset-audit.py golden.jsonl --freeze manifest.json        # in CI
      ```
      
      ## Four-bucket composition
      
      A set sampled only from happy-path production traffic tells you nothing about the failures
      you will actually ship. Build it in four deliberate buckets and record the bucket on every
      case.
      
      | Bucket | Target share | Source | Catches |
      |---|---|---|---|
      | `production` | 40-50% | Stratified sample of real traffic | Drift on the common path; keeps the score meaningful |
      | `replay` | 20-30% | Every incident that reached a human | Regressions on things that already broke once |
      | `adversarial` | 15-20% | Injections, contradictory instructions, refusal-bait, out-of-scope asks | Silent policy failures; the class judges also fail on |
      | `edge` | 10-15% | Empty, enormous, ambiguous, multilingual, malformed | Where deterministic code breaks first |
      
      Those are **targets to compose against**, not thresholds. `goldenset-audit.py` warns on a
      deliberately wider band (production 30-65%, replay 10-40%, adversarial 8-35%, edge 5-30%)
      and names the target in the warning, because a check that fires on every healthy set gets
      ignored within a week - the same rule this skill applies to CI gates. Hitting the target is
      good practice; leaving the band is a finding.
      
      **The replay bucket is the easiest to justify and the most neglected.** Every production
      incident is a free, pre-validated, maximally relevant test case. Make "add the replay case"
      a step in the incident checklist and the bucket fills itself.
      
      Watch the balance: buckets skew over time because production sampling is easy and
      adversarial authoring is not. A set that has become 90% `production` has quietly stopped
      testing the things that break. The audit script flags this.
      
      ## Sizing
      
      | Stage | Size | Note |
      |---|---|---|
      | Bootstrapping | 20 | Hand-written, eyeballable. Do this before choosing a metric. |
      | Working regression set | 100-300 | Enough to move a percentage meaningfully; cheap enough to run per PR |
      | Mature, production-sampled | 200-500 | The common steady state for a real product |
      | Judge calibration subset | 50-200 human-labelled | Separate purpose — see `llm-judge.md` |
      
      Past ~500 you are usually buying latency, not signal. Add cases when a **new failure class**
      appears, not on a cadence. If two cases fail and pass together every time, one of them is
      free to delete.
      
      Cost check: at 300 cases × 3 runs (pass^3) × $0.01/case you are spending ~$9 a run. That is
      fine nightly and painful on every push — which is why the PR tier runs a subset. See
      `regression-gating.md`.
      
      ## Case schema
      
      Nothing exotic. JSONL, one case per line, in git:
      
      ```json
      {"id": "refund-partial-001",
       "bucket": "replay",
       "added": "2026-03-14",
       "why": "INC-482: agent refunded full amount on a partial-return request",
       "input": {"messages": [{"role": "user", "content": "..."}]},
       "expected": {"tool": "issue_refund", "args": {"amount": 24.99}},
       "criteria": ["refund amount matches the returned item only",
                    "does not promise a timeline the policy does not state"]}
      ```
      
      The fields that matter and get omitted:
      
      - **`why`** — the reason this case exists. Without it, a future maintainer deletes cases
        they cannot interpret, and the set silently loses its adversarial teeth.
      - **`added`** — dates are how you detect a set that stopped growing in 2025.
      - **`bucket`** — without it you cannot see the balance drifting.
      - **`criteria`** — for open-ended outputs, the grading rubric belongs *with the case*, not
        in a global judge prompt. Case-specific criteria are dramatically easier to calibrate.
      
      ## Rot, and how it shows up
      
      | Rot | Symptom | Fix |
      |---|---|---|
      | **Duplicates / near-duplicates** | Score moves in suspiciously large jumps | Dedupe on normalised input; audit script flags exact and high-overlap pairs |
      | **Bucket skew** | Adversarial pass rate stops moving | Rebalance; author new adversarial cases |
      | **Staleness** | Newest `added` date is months old | Wire the incident checklist; sample fresh production traffic |
      | **Saturation** | Score pinned at 1.0 for weeks | The set is too easy. Harvest harder cases from production; a saturated set detects nothing |
      | **Contamination** | Score jumps on a model upgrade with no code change | Public benchmark cases leaked into training. Prefer private, product-specific cases |
      | **Test-fitting** | Cases edited in commits that also change the prompt | Enforce in review: dataset changes land in their own commit |
      
      Saturation deserves emphasis: a suite that always passes is not a passing suite, it is a
      suite that has stopped measuring. Track the *distribution* of per-case results, not just the
      mean, and retire-and-replace cases that have not failed in months.
      
      ## Synthetic cases
      
      Useful for edge and adversarial coverage, dangerous as the backbone. Synthetic inputs
      generated by the same model family you are testing inherit its blind spots — it will not
      generate the phrasing it does not understand. Use synthetics to *expand* a bucket around a
      real failure ("give me 10 variants of this injection"), never to found one.
      
      Always human-review synthetic expected outputs before promotion. An unreviewed synthetic
      case is a hypothesis, not a golden case.
      
      ## Cross-reference
      
      - What to score these cases on: `eval-taxonomy.md`
      - Grading the open-ended ones: `llm-judge.md`
      - Running them in CI: `regression-gating.md`
      
    • hillclimbing.md 9.8 KB
      # Hillclimbing — optimising against an eval without destroying it
      
      > **Verified 2026-08.** The optimizer landscape below moves fast. Method names,
      > reported numbers and library APIs need re-verification before you quote them.
      
      Once you have a measurable harness, the obvious next move is to optimise against it:
      change something, measure, keep if better, repeat. That loop works, and it is also the
      single most reliable way to turn a good eval suite into a useless one.
      
      **This file owns the discipline. It does not own the loop.** The loop mechanics — scope,
      verify command, keep/discard, batch with bisect-on-regression, stop conditions, git as
      memory — belong to the [`iterate`](../../iterate/SKILL.md) skill, which is domain-agnostic
      and does not care whether your metric is test coverage or an eval score. Read that for
      *how to run* the loop; read this for *what goes wrong* when the metric is an eval.
      
      ## The two failure modes
      
      ### 1. Banking noise
      
      `iterate` keeps a change when the metric beats the previous best. With a deterministic
      metric (line coverage, bundle bytes) that rule is exactly right. With an eval score it is
      a coin flip dressed as a decision.
      
      If your suite scores 0.88 ± 0.03 across reruns of *unchanged code*, then a change that
      measures 0.90 has told you almost nothing. Keep it anyway and you have banked noise. Do
      that fifty times overnight and you have executed a random walk with perfect discipline,
      and the winning commit is whichever iteration got luckiest.
      
      **The rule: a hillclimb step is only real if the delta clears the noise floor.**
      
      ```bash
      python3 scripts/eval-baseline.py history.jsonl --candidate iter.jsonl --accept
      # exit 0  = KEEP   (improvement clears the floor, or is significant on the paired test)
      # exit 10 = DISCARD (inside the noise band, or a regression)
      ```
      
      `--accept` inverts the usual exit semantics on purpose — see the script header. Wire it as
      the keep/discard gate and the loop stops rewarding luck.
      
      Two cheap amplifiers, when you can afford them:
      
      - **Increase k before increasing iterations.** Three runs per candidate shrinks the noise
        floor and is usually a better spend than three times as many candidates evaluated once.
      - **Paired comparison beats aggregate comparison.** Which specific cases flipped is far
        more informative than a delta of 0.01, and it is free once you store per-case results
        (`regression-gating.md`).
      
      ### 2. Overfitting the golden set
      
      This one is slower, quieter, and permanent. Every look at the same frozen set leaks a
      little information into your decisions. Optimise against it for long enough and you have
      tuned the system to that specific 300 cases — Goodhart's law arriving exactly on schedule.
      
      It is not hypothetical in the optimizer literature: **GEPA is documented to overfit by
      encoding edge cases into increasingly verbose prompts**, because its reflection step
      accumulates detail across iterations and nothing pushes back. Length constraints act as
      regularisation there, which is a good general instinct: an optimizer with no pressure
      toward simplicity will buy training score with complexity.
      
      **Split the data before you optimise anything:**
      
      | Split | Used for | Rule |
      |---|---|---|
      | **Train / reflect** | What the optimizer sees and reasons about | Look freely. This is the set you burn |
      | **Validation** | Choosing between candidates each round | Scored every round; never shown to the optimizer's reflector |
      | **Held-out test** | The number you actually report | Touch at milestones only. Every look costs you |
      
      GEPA's own design makes the second row explicit: it reserves a disjoint validation subset
      whose inputs and outputs are **never shown to the reflector model**. Adopt that separation
      even when hand-rolling — the moment the thing proposing changes can read the set that
      judges them, your validation score stops being evidence.
      
      **The held-out set is a budget, not a dashboard.** Decide up front how often you may look
      (a milestone, a release, once a week) and hold to it. A team that checks held-out every
      iteration has three sets and one of them is a validation set wearing a disguise.
      
      **Tells that you are overfitting:**
      
      - Train score climbs; validation is flat. The classic.
      - Validation climbs; held-out is flat. You are now overfitting the validation set.
      - The system prompt or config grows monotonically, each addition patching one case.
      - Wins stop transferring — a change that helped on the suite does nothing in production.
      
      ## Keep a frontier, not a champion
      
      `iterate` maintains `iterate/best`: a single floating tag on the highest-metric commit.
      For a scalar mechanical metric that is correct and simple. For an eval score it is a
      local-optimum trap, because one aggregate number hides which *cases* a candidate won.
      
      GEPA's central design choice is the alternative: **maintain a Pareto frontier** — retain
      every candidate that is best on at least one validation instance, and sample from that
      frontier rather than always mutating the current champion. A candidate that scores lower
      overall but is the only one solving a hard case carries information the champion does not,
      and discarding it is how a hillclimb walls itself into a local optimum.
      
      Cheap approximation without any framework: alongside the best-overall commit, keep a note
      of **which candidate best solved each failing case**. When the loop stagnates, mutate from
      one of those instead of from the champion.
      
      ## Rich feedback beats a scalar reward
      
      The most transferable finding in this line of work: **collapsing an evaluation to a number
      throws away the signal an optimizer needs.** GEPA reflects in natural language over
      execution traces — error messages, reasoning logs, why the case failed — and reports
      outperforming MIPROv2 by roughly 10–13% and the RL baseline GRPO by ~6% on average (up to
      20%) across six tasks, **using up to 35× fewer rollouts**. It is sample-efficient enough to
      work from as few as 10 examples and 20–100 evaluations.
      
      The number to take from that is not the benchmark delta, which will age. It is the
      mechanism: a failure that explains itself is worth many failures that only score.
      
      **This is an eval-design consequence, and it is why it lives in this skill.** If your judge
      returns `0.4`, an optimizer has nothing to reflect on. If it returns
      `{"reason": "cited chunk 12, which does not contain the refund window", "verdict": "fail"}`,
      it has a diagnosis. `assets/judge-rubric.template.md` already demands `reason` alongside
      `verdict` — that field is what makes hillclimbing tractable later, and it costs nothing to
      add now. Same for per-case traces: store them, or you cannot reflect on them.
      
      ## The optimizer landscape
      
      Families rather than products, since the products churn:
      
      | Family | Mechanism | Note |
      |---|---|---|
      | **Reflective / evolutionary** (GEPA) | Natural-language reflection over traces; Pareto frontier over candidates | Sample-efficient; available in DSPy as an optimizer and standalone. Documented verbosity-overfit failure mode |
      | **Bayesian / few-shot search** (MIPROv2) | Proposes instructions and demonstrations, searches the joint space | The prior DSPy default; the baseline GEPA is measured against |
      | **Generate–score–select** (APE) | Generate candidate prompts, score on validation, keep the best | Simplest thing that works; a fine hand-rolled starting point |
      | **Iterative self-rewrite** (ORPO and kin) | The model rewrites its own prompt guided by feedback on prior outputs | Needs a feedback signal richer than a score |
      | **Self-contained preference loops** (SPO) | Generates its own data, refines by pairwise preference over its outputs | Removes the external-label dependency — and with it, your ground truth. Treat results with suspicion |
      | **Scaffold-level** | Memory evolution, tool governance, whole-harness redesign | The wider "self-improving agent" framing; prompt optimization is one lever among several |
      
      **When it helps is a live research question, not a settled one.** There is published work
      specifically asking *when* prompt optimization improves multi-agent systems — the framing
      implies the honest answer is "sometimes". Do not assume an optimizer will beat a careful
      human rewrite on your task; measure it, on held-out, like anything else.
      
      ## Before you automate the loop
      
      - [ ] A frozen golden set with train / validation / held-out splits (`golden-datasets.md`).
      - [ ] A measured noise floor, from reruns of unchanged code (`regression-gating.md`).
      - [ ] A calibrated judge, if a judge is in the metric (`llm-judge.md`). An uncalibrated
            judge in a hillclimb optimises the system toward the judge's biases, at speed.
      - [ ] Per-case results and reasons stored, not just an aggregate.
      - [ ] A stated held-out look budget, and a stop condition.
      - [ ] A cost ceiling. Optimizer loops are the easiest way to spend a month of eval budget
            in an afternoon.
      
      If any of those is missing, fix it before running the loop. A hillclimb amplifies whatever
      your measurement already is — including its errors.
      
      ## When this should become its own skill
      
      Kept here because it is currently one file, and the creation protocol says extend rather
      than duplicate. **Extract it to a `prompt-optimization-ops` skill when the optimizer
      material outgrows this page** — concretely, when it needs its own worked DSPy/GEPA
      configuration, more than a couple of runnable scripts, or per-optimizer troubleshooting.
      The discipline sections above (splits, noise floor, frontier, held-out budget) stay here
      regardless: they are properties of the measurement, and this skill owns measurement.
      
      ## Cross-reference
      
      - Loop mechanics, stop conditions, bisect-on-regression: [`iterate`](../../iterate/SKILL.md)
      - Scheduling a loop across sessions, risk tiers, kill switch: [`loop-ops`](../../loop-ops/SKILL.md)
      - Splits and freeze discipline: `golden-datasets.md`
      - Noise floor and the paired test: `regression-gating.md`
      - Why a judge must be calibrated before it steers anything: `llm-judge.md`
      
    • llm-judge.md 7.7 KB
      # LLM-as-a-Judge — biases, rubrics, panels, calibration
      
      A judge is a measurement instrument built out of a language model. It is useful, it is
      often the only option, and it is biased in documented, reproducible ways. Treat it like an
      instrument: know its error modes, calibrate it against a reference, and re-check it when
      anything underneath changes.
      
      ## Decide whether you need one at all
      
      | Criterion | Evaluator |
      |---|---|
      | Output must parse / match a schema | `json.loads`, a JSON-Schema validator |
      | Exact value, ID, amount, tool name | `==` |
      | Contains a required citation / does not contain a banned string | regex |
      | Ordering, latency, cost, step count | arithmetic over the trace |
      | **Faithfulness to a source** | judge |
      | **Policy / tone compliance** | judge |
      | **"Is this a reasonable answer to an open question"** | judge |
      | **Relative quality of two candidates** | judge (pairwise, with position control) |
      
      Every criterion you move out of the judge and into code removes cost, latency *and*
      variance in one edit. Re-audit the rubric periodically — items become codifiable once the
      output format stabilises, and nobody goes back to check.
      
      ## The bias catalog
      
      ### Position bias
      
      In pairwise comparison, judges prefer whichever candidate was presented first — strongly
      enough that swapping the order flips a meaningful share of verdicts.
      
      **Mitigations, in order of preference:**
      
      1. **Score absolutely, not pairwise.** Each candidate graded against the rubric alone. Kills
         the bias by construction and makes results comparable across runs.
      2. **Both orders, averaged.** If you need pairwise, run A/B and B/A and require agreement.
         Doubles cost; the disagreement rate is itself a useful instability metric.
      3. Never accept a single-order pairwise verdict as a gate.
      
      ### Verbosity bias
      
      Judges rate longer answers higher regardless of quality. Evidence is heterogeneous — some
      models are genuinely quality-sensitive and penalise filler — which is exactly why you must
      measure it on *your* judge rather than assume.
      
      **Mitigations:**
      
      - Split the rubric: score *correctness* and *style* separately, and gate on correctness.
      - State the anti-bias instruction explicitly ("length is not evidence of quality; an answer
        that is correct and brief scores higher than one that is correct and padded").
      - **Probe it.** Correlate judge score against output length on your calibration set. A
        strong positive correlation on cases with equal human scores is the bias, quantified.
        `scripts/judge-calibration.py --verbosity-field length` computes this.
      
      ### Self-preference bias
      
      Judges rate outputs from their own model family higher. Fatal when you are comparing models
      and using one of the contenders as the judge.
      
      **Mitigation:** judge with a different family from the system under test. If that is
      impossible, at minimum report the judge model alongside every score and never compare
      scores produced by different judges.
      
      ### Scale drift and clustering
      
      1-5 scores cluster (almost everything gets a 4), and the cluster shifts when the judge
      model version changes — silently re-baselining your whole history.
      
      **Mitigations:**
      
      - **Binary pass/fail against explicit criteria** wherever the decision is genuinely binary.
        Easier to calibrate, easier to act on, far more stable across model versions.
      - **Pin the judge model version** and treat a judge upgrade like a dataset version bump:
        re-baseline, do not compare across it.
      - If you need a scale, define each point with a concrete example, not an adjective.
      
      ### Other effects worth knowing
      
      | Effect | Note |
      |---|---|
      | **Sycophancy toward the prompt** | A judge asked "confirm this is correct" confirms. Phrase neutrally, or invert — see `adversarial-verification.md` |
      | **Format preference** | Markdown/bulleted answers score above equivalent prose. Normalise formatting before judging where you can |
      | **Anchoring on the reference** | Given a reference answer, judges penalise correct-but-different. Say explicitly that alternative correct answers are acceptable |
      
      ## Rubric design
      
      The rubric is where most judge quality lives — far more than the model choice.
      
      - **One criterion per question.** A rubric asking "is it accurate, helpful and well-written?"
        returns an unactionable blend. Three separate binary questions return three actionable
        answers.
      - **Concrete and checkable.** "Every factual claim appears in the provided source" beats
        "is accurate".
      - **Case-specific criteria where possible.** Criteria stored with the golden case
        (`criteria: [...]`) calibrate dramatically better than one global rubric stretched over a
        heterogeneous set.
      - **Require reasoning before the verdict.** Chain-of-thought judging (the G-Eval line of
        work) improves human agreement; and the reasoning is what lets you debug a disagreement
        instead of shrugging at it.
      - **Demand structured output** — `{"reason": "...", "verdict": "pass"}` — so scoring is
        parseable and the reason is stored, not discarded.
      - **Include the failure examples.** Two or three worked examples of what a `fail` looks like
        do more for agreement than a page of prose.
      
      ## Panels vs N-identical
      
      Running the same judge three times with the same rubric mostly buys the same bias three
      times. It measures the judge's *own* variance — worth knowing once, not worth paying for
      every run.
      
      A **panel with distinct lenses** is different in kind: each judge is asked a different
      question, so their failure modes do not overlap.
      
      | Shape | Use when |
      |---|---|
      | Single judge, absolute scoring | Default. Cheapest thing that works |
      | N-identical, same rubric | One-off: measure judge variance to set your CI noise floor |
      | **Panel, distinct lenses** | The thing can fail in several independent ways (correct? policy-compliant? reproducible?) |
      | **Panel, distinct model families** | High-stakes scoring; aggregate by majority to damp single-model bias |
      
      Majority voting across heterogeneous judges measurably improves correlation with human
      judgment. It also multiplies cost — reserve it for the scores you gate on.
      
      ## Calibration — the step that makes a judge trustworthy
      
      Calibration has gone from nice-to-have to table stakes. The method:
      
      1. **Sample 50-200 cases** from the golden set, stratified across buckets and across the
         judge's own verdicts (include cases it passes *and* fails, or you cannot measure both
         error directions).
      2. **Label them by hand.** Same rubric the judge gets. Two humans on a subset gives you a
         human-human ceiling — a judge cannot beat the agreement humans achieve with each other,
         so that number tells you what "good" even means here.
      3. **Compute Cohen kappa**, not raw agreement. On an imbalanced set (90% pass) a judge that
         says "pass" unconditionally scores 90% agreement and is worthless; kappa discounts the
         agreement you would get by chance and lands it near zero.
      
      ```bash
      python3 scripts/judge-calibration.py labels.jsonl --min-kappa 0.6
      ```
      
      | kappa | Reading |
      |---|---|
      | >= 0.8 | Strong. Production-ready; safe to gate on with a margin |
      | 0.6 - 0.8 | Substantial. Usable, keep it advisory or gate loosely |
      | < 0.6 | The rubric is the problem, not the model. Rewrite before spending more on judges |
      
      4. **Read the confusion matrix, not just the headline.** A judge with kappa 0.65 that is
         wrong only in the false-*negative* direction is safe for a gate (it under-passes, never
         over-passes). One with the same kappa that hallucinates passes is not. The two need
         opposite fixes.
      5. **Re-calibrate on a schedule** — roughly 50 fresh cases periodically, and *always* after
         a judge model version change or a rubric edit.
      
      ## Cross-reference
      
      - Refuting rather than confirming: `adversarial-verification.md`
      - Where the labelled cases come from: `golden-datasets.md`
      - Turning a calibrated judge into a CI gate: `regression-gating.md`
      
    • regression-gating.md 8.4 KB
      # Regression Gating — making evals block CI without killing CI
      
      An eval gate has exactly one job: stop a regression from merging. It fails at that job in
      two ways — by letting regressions through, and by going red so often that everyone learns
      to click merge anyway. The second failure is more common and much harder to reverse.
      
      ## The rule
      
      **A blocking check must never be flaky.** Once a team has seen three red-for-no-reason eval
      runs, the gate is socially dead even while it is still technically enforced. Design for that
      first and for coverage second.
      
      ## The tier ladder
      
      | Tier | Checks | Gate | Runs on |
      |---|---|---|---|
      | **0. Deterministic** | Schema validity, tool-call assertions, exact matches, forbidden strings | **Blocking**, zero tolerance | Every push |
      | **1. Judge, uncalibrated** | Any new rubric, first few weeks | **Advisory** — comment the delta, never fail | Every PR |
      | **2. Judge, calibrated** | kappa >= 0.6, variance measured | **Blocking with a margin** below rolling baseline | Every PR |
      | **3. Consistency** | pass^3 over the full set | **Blocking on the nightly**, advisory on PRs | Nightly |
      | **4. Cost / latency** | Tokens and p95 per case | **Blocking on an absolute ceiling**, advisory on trend | Every PR |
      
      Promotion from tier 1 to tier 2 is an explicit decision backed by a calibration run
      (`scripts/judge-calibration.py`), not something that happens because a rubric has been
      around a while.
      
      ## The noise floor
      
      You cannot set a threshold without knowing how much the suite moves when *nothing changes*.
      
      Measure it once, properly: run the unchanged system against the frozen set N times (5 is
      usually enough) and record the spread.
      
      ```
      run 1: 0.88   run 2: 0.85   run 3: 0.89   run 4: 0.86   run 5: 0.88
      baseline 0.872,  spread 0.04
      ```
      
      Then gate **below the baseline by more than the spread**: threshold 0.80, not 0.87. A gate
      inside the noise band fails on identical code, which is the fastest possible route to a
      dead gate.
      
      Keep it honest over time by committing a **rolling window of run results to git** — a small
      JSON file, appended per run on the main branch:
      
      ```json
      {"date": "2026-08-30", "dataset": "golden-v3", "judge": "<pinned-model-id>",
       "score": 0.871, "pass_at_1": 0.86, "pass_hat_3": 0.79,
       "cost_usd": 2.14, "p95_ms": 4180, "n": 287}
      ```
      
      That file is what turns "today looks bad" into "today is 2.6 spreads below a stable
      baseline" — and it costs nothing. Treating every run as standalone is what makes teams
      unable to distinguish noise from regression.
      
      Re-baseline (and say so in the commit) on any of: dataset version bump, judge model change,
      rubric edit, or a deliberate accepted trade-off.
      
      ```bash
      python3 scripts/eval-baseline.py evals/history.jsonl --candidate /tmp/run.jsonl
      # prints baseline, noise floor, and the threshold your gate should use
      ```
      
      ## Is the drop real? — the paired test
      
      The noise floor tells you whether an aggregate score moved further than it usually
      does. It does **not** tell you whether the same cases moved, and that is the question
      you actually care about.
      
      Two runs over the same frozen set produce *paired binary outcomes*, and the right tool
      for those is **McNemar's exact test**. It looks only at the discordant pairs:
      
      |  | candidate passes | candidate fails |
      |---|---|---|
      | **baseline passes** | ignored | **b** — regressions |
      | **baseline fails** | **c** — fixes | ignored |
      
      Cases that behaved identically in both runs carry no information about whether the
      change helped. Under the null hypothesis b is a coin flip over b+c trials, so the
      p-value is a tail probability. Use the **exact** test rather than chi-square: eval
      sets routinely produce b+c under 25, where the approximation misleads.
      
      Why this matters more than the aggregate: a change that breaks 8 cases and fixes 7 moves
      the headline score by 0.01 - invisible against any noise floor - while having silently
      swapped which 15 things work. The paired view names those 15 cases; a score comparison
      structurally cannot.
      
      Note carefully what the test does and does not say there. 8-vs-7 gives p = 1.0: genuinely
      indistinguishable from chance, and the tool will correctly call it noise. The value in that
      run is not the verdict, it is the enumerated `regressed` and `fixed` lists telling you a
      churn happened at all. Significance answers "did the system get worse"; the lists answer
      "what moved" - and on a flat score only the second question has an answer worth having.
      
      ```bash
      python3 scripts/eval-baseline.py evals/history.jsonl \
        --baseline-results base.jsonl --candidate-results new.jsonl --alpha 0.05
      # exit 10 = significant regression, and it names the cases that flipped
      ```
      
      Three cautions:
      
      - **Significance is not magnitude.** With a large set, a trivially small real drop
        reaches p < 0.05. Read the count of regressed cases, not only the p-value.
      - **Don't run the test repeatedly until it agrees with you.** Testing every PR against
        the same baseline is many comparisons; treat a single surprising red as a prompt to
        look at the named cases, not as proof on its own.
      - **k > 1 breaks the pairing** unless you collapse each case to one outcome first
        (pass^k is the usual choice). Feed the collapsed per-case result, not every run.
      
      ## CI shape
      
      A workable three-tier cadence:
      
      | Trigger | Scope | Budget | Gate |
      |---|---|---|---|
      | **Every push** | Deterministic assertions on the full set | seconds, $0 | Blocking |
      | **PR** | Judge metrics on a stratified ~30% subset, k=1 | a few minutes | Per tier ladder |
      | **Nightly on main** | Full set, k=3, cost and latency recorded, appended to the history file | whatever it costs | Blocking; page on a real drop |
      
      Notes that matter in practice:
      
      - **Pin everything the score depends on** — judge model version, dataset version, prompt
        version, temperature (0 for the judge). An unpinned judge model is a silent
        re-baselining that will be blamed on your code.
      - **Cache aggressively.** Eval runs re-send near-identical prompts; caching the static
        prefix cuts the bill substantially and does not change scores.
      - **Report the delta, not the absolute.** "-0.04 vs main (noise floor 0.03)" is actionable;
        "0.83" is not.
      - **Name the failing cases in the CI output.** A gate that says "score dropped" without
        listing which 6 cases flipped forces a local re-run and gets ignored.
      - **Never auto-retry a failing eval to green.** Retry-until-pass converts a real regression
        into a flake report. If you retry, report all attempts.
      
      ## Cost and latency attribution
      
      Record per case, from day one:
      
      | Field | Why |
      |---|---|
      | `tokens_in` / `tokens_out` | The unit you actually pay for; also the best proxy for context bloat |
      | `ms` (wall) and step count | p95 latency is a product requirement, and step count catches loops |
      | `cost_usd` | Roll up per run so a "small" prompt change that doubles spend is visible immediately |
      | `model` / `judge_model` | Attribution is meaningless if you cannot tell which model produced the number |
      
      Two reasons this is not optional. First, an eval suite is the only place you learn that the
      accuracy win cost 4x the tokens — production tells you eventually, and much more expensively.
      Second, retrofitting attribution once the harness exists means touching every runner, every
      stored result and every dashboard; adding four fields at the start costs nothing.
      
      Gate on an **absolute ceiling** (cost per case must not exceed $X, p95 must not exceed Y ms)
      rather than on the trend. Trend gates fire on noise; ceilings encode a product decision.
      
      ## Failure modes to design against
      
      | Failure | Symptom | Fix |
      |---|---|---|
      | Flaky blocking judge | Reds nobody investigates | Demote to advisory until calibrated; widen the margin past the noise floor |
      | Threshold inside the noise band | Identical code fails intermittently | Measure the spread; gate below baseline minus spread |
      | Saturated suite | Green for months, then a production incident | The set stopped measuring — harvest harder cases (`golden-datasets.md`) |
      | Confounded comparison | Score moved and nobody knows why | Version the dataset; never compare across versions |
      | Judge upgrade drift | Step change in scores with no code change | Pin the judge; re-baseline explicitly on upgrade |
      | Test-fitting | Cases edited in the same commit as the prompt | Enforce in review: dataset changes land separately |
      
      ## Cross-reference
      
      - Case supply and freezing: `golden-datasets.md`
      - Getting a judge to kappa >= 0.6 so it can be promoted to blocking: `llm-judge.md`
      - Which metric belongs in which tier: `eval-taxonomy.md`
      
    • retrieval-eval.md 6.4 KB
      # Retrieval Eval — scoring RAG without conflating two different bugs
      
      Retrieval is the most common thing people build evals for, and the most commonly
      mis-measured. The mistake is universal: score the final answer, watch it drop, and
      have no idea whether the retriever failed or the generator did.
      
      ## Split the pipeline before you score it
      
      A RAG answer passes through two stages that fail independently. Score them
      separately or you cannot act on either.
      
      |  | Retrieved the right context | Retrieved the wrong context |
      |---|---|---|
      | **Answer correct** | Working as intended | **Lucky** — the model knew it anyway, or guessed. Will fail when the question shifts |
      | **Answer wrong** | **Generation bug** — chunking, prompt, or model | **Retrieval bug** — embeddings, index, query rewriting |
      
      The two off-diagonal cells need opposite fixes, and end-to-end accuracy averages
      them into one uninterpretable number. Worse, the top-right cell — right answer from
      wrong context — scores as a *pass* end-to-end and is a latent failure exactly like
      the lucky pass in `eval-taxonomy.md`.
      
      **Minimum viable split:** for every case, record the retrieved chunk ids alongside
      the answer. That single field turns an opaque score into a 2x2 you can act on.
      
      ## Retrieval metrics
      
      Retrieval is the one place in the eval stack where **deterministic scoring
      genuinely dominates** — you have ground-truth chunk ids, so no judge is required.
      Take the free signal.
      
      | Metric | Definition | Use when |
      |---|---|---|
      | **Recall@k** | Fraction of relevant chunks that appear in the top k | The headline. If the right chunk is not in the context, nothing downstream can save you |
      | **Precision@k** | Fraction of the top k that are relevant | Context budget is tight; noise crowds out signal |
      | **MRR** | Mean of 1/rank of the first relevant chunk | One right answer per query; you care that it ranks high |
      | **nDCG@k** | Rank-discounted gain over graded relevance | Multiple chunks matter and some matter more |
      | **Context precision** | Of the context actually passed to the model, how much was used | Diagnosing bloated prompts and cost |
      
      **Recall@k is the one to gate on.** Precision failures degrade an answer; recall
      failures make a correct answer impossible. Measure recall at the k you actually
      retrieve *and* at a larger k — if recall@20 is high while recall@5 is poor, you
      have a ranking problem, not an embedding problem, and those are different fixes.
      
      ## Building the ground truth
      
      The dataset is the hard part, as always (`golden-datasets.md`). What retrieval
      adds:
      
      - **Annotate chunk ids, not passages.** Ids survive re-chunking; quoted text does
        not. Store the chunk's stable id plus a content hash so a silent re-index shows
        up as drift rather than as a mysterious recall drop.
      - **Relevance is graded, not binary,** for anything but the simplest corpus:
        `2` = answers the question, `1` = useful context, `0` = irrelevant. nDCG needs
        this; recall@k works fine treating >= 1 as relevant.
      - **Multiple relevant chunks are normal.** A question answerable only by combining
        two documents is a different (harder) test than a single-hop lookup — label the
        hop count and report the two classes separately.
      - **Harvest queries from real traffic.** Synthetic questions generated *from* a
        chunk are trivially retrievable from that chunk — they share its vocabulary. They
        measure your embedding model's ability to match paraphrases, not your retriever's
        ability to handle how people actually ask.
      
      That last point is the single most common way a retrieval eval flatters itself.
      A set of "generate a question from this passage" pairs will show recall@5 above
      0.95 on a system that fails constantly in production.
      
      ## Failure classes worth their own bucket
      
      Each of these fails differently and needs its own cases:
      
      | Class | Why it breaks |
      |---|---|
      | **Vocabulary mismatch** | User says "can't log in", docs say "authentication failure". Pure semantic search handles this; keyword search does not — and hybrid exists for the reverse case |
      | **Exact identifiers** | Order numbers, error codes, SKUs, function names. Embeddings are *bad* at these; this is what BM25/keyword hybrid is for |
      | **Multi-hop** | The answer needs two documents. Single-shot retrieval structurally cannot |
      | **Negation / absence** | "Which plans do NOT include support?" Similarity retrieves the plans that *do* |
      | **Temporal** | "The current policy" retrieves a superseded version that is textually similar. Needs metadata filtering, not better embeddings |
      | **Nothing relevant exists** | The corpus does not contain the answer. Correct behaviour is to say so — and it must be a scored case, or the system learns to always answer |
      
      The last one deserves emphasis: **a golden set with no unanswerable questions
      cannot detect hallucination under retrieval failure**, which is the exact scenario
      users hit most often.
      
      ## Generation-side metrics
      
      Once the right context is in hand:
      
      | Metric | Question | Evaluator |
      |---|---|---|
      | **Faithfulness / groundedness** | Is every claim supported by the retrieved context? | Judge (`llm-judge.md`) — decompose into claims, check each |
      | **Citation accuracy** | Do the cited chunk ids actually contain the cited content? | **Deterministic** — verify the id exists and the claim maps to it |
      | **Answer relevance** | Does it address the question asked? | Judge |
      | **Refusal correctness** | Does it decline when the context does not contain the answer? | Deterministic, on the unanswerable bucket |
      
      Citation accuracy is quietly one of the highest-value checks available: it is free,
      it needs no judge, and it catches the specific failure where a model produces a
      confident answer and attaches a plausible-but-unrelated source.
      
      ## What to gate
      
      | Tier | Check |
      |---|---|
      | **Blocking** | recall@k on the frozen query set; citation-id validity; refusal rate on the unanswerable bucket |
      | **Blocking (ceiling)** | Context tokens per query — retrieval regressions often show up as cost before they show up as accuracy |
      | **Advisory until calibrated** | Faithfulness and answer-relevance judges |
      | **Diagnostic** | The 2x2 above, per bucket — this is what tells you which team owns the drop |
      
      ## Cross-reference
      
      - The levels this sits inside: `eval-taxonomy.md`
      - Case construction and freeze discipline: `golden-datasets.md`
      - Grading faithfulness without fooling yourself: `llm-judge.md`
      - Where these land in CI: `regression-gating.md`
      
    • tooling-landscape.md 5 KB
      # Tooling Landscape — which eval platform, when
      
      > **Verified 2026-08.** This is the fastest-moving part of the skill. Features, pricing and
      > OSS/commercial boundaries here change on a scale of months. **Re-verify with a web search
      > before quoting any specific claim** — treat everything below as a shape to check against
      > reality, not as a current fact sheet. No version numbers are quoted deliberately.
      
      ## The one structural fact
      
      **Trace-level observability and eval scoring have converged into the same products.** As of
      2026 you are not choosing a tracer and then an eval library; you are choosing one system
      that captures the nested span tree (model calls, tool calls, arguments, cost) and attaches
      scores to those traces, in dev and in production. Any comparison that treats them as two
      categories is out of date.
      
      That convergence is what makes trajectory- and step-level eval practical at all
      (`eval-taxonomy.md`) — you cannot score a path you did not capture.
      
      ## The honest default
      
      **Start with a JSONL file and a 40-line runner.** Read cases, call the system, apply
      assertions, write results, print a delta. It is a morning's work, it has no vendor coupling,
      and it forces you to decide what you are actually measuring — which is the hard part, and
      the part no platform does for you.
      
      Adopt a platform when you hit a specific wall:
      
      | Wall | Then you want |
      |---|---|
      | Non-engineers need to curate the dataset | A dataset-management UI (Braintrust is the archetype) |
      | You need to search production traces and score live traffic | An observability-first platform (Langfuse, Arize, LangSmith, Opik) |
      | Evals should feel like the test suite | A pytest-native framework (DeepEval) |
      | You already run ML infra and want one system | MLflow |
      | Scheduled runs, shared dashboards, alerting | Any hosted platform; this is what you are paying for |
      
      Do not adopt one because the eval list is long. Metric catalogs are cheap; a rubric
      calibrated against your humans is not, and no platform ships that.
      
      ## The players
      
      Open-source cores (self-hostable, no vendor lock on the data):
      
      | Tool | Shape | Reach for it when |
      |---|---|---|
      | **DeepEval** | pytest-native LLM eval framework, large research-backed metric library | Your team's mental model is "tests"; you want evals in the existing test command |
      | **MLflow** | Tracing with replay, prompt versioning, automated eval — one OSS platform | You already run MLflow, or you want the whole stack under one OSS licence |
      | **Opik** (Comet) | Tracing with cost tracking, built-in metrics, prompt versioning, broad framework integrations | You want hosted-or-self-hosted flexibility with wide framework coverage |
      | **Langfuse** | Observability-first: traces, datasets, scores, self-hostable | Production trace search is the primary need |
      | **Arize Phoenix** | OSS tracing/eval, OpenTelemetry-native | You are standardising on OTel semantics |
      
      Commercial-first:
      
      | Tool | Shape | Reach for it when |
      |---|---|---|
      | **Braintrust** | Eval-focused, strong collaborative dataset curation, scheduled runs, score-regression tracking | PMs and domain experts must own the golden set |
      | **LangSmith** | Tracing + eval, tight LangChain/LangGraph integration | You are already deep in that ecosystem |
      | **Arize** | Production ML/LLM observability at scale | Enterprise monitoring is the driver |
      | **AgentOps** | Agent-run-centric session replay, cost and step tracking | Debugging long agent trajectories is the pain |
      
      ## Selection checklist
      
      Score candidates on the things that actually bite six months in:
      
      1. **Can you export your traces and datasets?** The dataset is the durable asset
         (`golden-datasets.md`). If it only lives in a vendor UI, you have rented your history.
      2. **Does it capture the full nested span tree**, including tool arguments? Without
         arguments you cannot do step-level scoring.
      3. **Can it run in CI and fail a build** with a machine-readable result? A platform you can
         only read in a browser cannot gate anything.
      4. **Can you pin the judge model version?** Unpinned judges silently re-baseline
         (`llm-judge.md`).
      5. **Does it record cost and latency per case** natively, or must you thread it yourself?
      6. **Self-host option?** Matters the moment eval inputs contain customer data.
      7. **What happens to your custom metrics** — are they plain functions you own, or a DSL you
         would have to rewrite to migrate?
      
      ## Benchmarks vs your evals
      
      Public benchmarks (tau-bench and the agentic-benchmark family, SWE-bench and its
      descendants) are for *model selection* — they tell you which model to start from. They are
      not your eval suite: they are contaminated over time, they measure someone else's task
      distribution, and a model that tops them can still fail your product's specific policy.
      
      Use them once, at model-choice time. Then measure your own thing, on your own frozen set.
      
      ## Cross-reference
      
      - What to measure before you shop for a tool: `eval-taxonomy.md`
      - The asset that outlives whatever you pick: `golden-datasets.md`
      - Making the chosen tool gate CI: `regression-gating.md`
      
  • scripts
    • eval-baseline.py 20.3 KB
      #!/usr/bin/env python3
      """Tell a real eval regression from noise, using run history and McNemar's test.
      
      Reads the rolling run-history file the runner appends to, derives the baseline
      and the NOISE FLOOR, and judges one candidate run against them. Where per-case
      results for both runs are supplied it also runs McNemar's exact test, which is
      the honest answer to "did my change make it worse" on paired binary outcomes --
      a score drop inside the noise band is not evidence of anything.
      
      Usage:   eval-baseline.py [OPTIONS] <HISTORY.jsonl>
      Input:   HISTORY.jsonl -- one summary object per run, appended over time.
               Recognised: `score` (required), `n`, `date`, `dataset`, `judge`,
               `cost_usd`, `p95_ms`. `-` reads stdin.
               --candidate FILE     a one-row summary for the run under test
                                    (default: the last row of HISTORY)
               --baseline-results / --candidate-results  per-case JSONL of
                                    {"id": ..., "passed": true|false} for McNemar
      Output:  stdout -- human-readable verdict, or a --json envelope
               {"data": {...}, "meta": {...}} per SKILL-RESOURCE-PROTOCOL.md §4.
      Stderr:  headers, warnings, errors.
      Exit:    0 no regression (noise, improvement, or not enough evidence),
               2 usage, 3 not-found, 4 validation,
               10 REGRESSION CONFIRMED or a cost/latency ceiling breached
      
               --accept INVERTS this deliberately, for use as a hillclimb keep/discard
               gate:  0 = KEEP (a real improvement),  10 = DISCARD (noise or worse).
               Under --accept, "noise" is a DISCARD -- banking a change that is inside
               the noise band is how an improvement loop turns into a random walk.
               See references/hillclimbing.md.
      
      Examples:
        eval-baseline.py evals/history.jsonl
        eval-baseline.py evals/history.jsonl --candidate /tmp/run.jsonl
        eval-baseline.py evals/history.jsonl --baseline-results base.jsonl \
            --candidate-results new.jsonl --alpha 0.05
        eval-baseline.py evals/history.jsonl --max-cost-usd 2.50 --max-p95-ms 6000
        eval-baseline.py evals/history.jsonl --candidate iter.jsonl --accept  # keep/discard
        eval-baseline.py evals/history.jsonl --json | jq '.data.recommended_threshold'
      
      Offline and stdlib-only: it reads files you already have, it never calls a model.
      """
      
      import argparse
      import json
      import math
      import sys
      
      SCHEMA = "claude-mods.evals-ops.eval-baseline/v1"
      
      EXIT_OK, EXIT_USAGE, EXIT_NOT_FOUND, EXIT_VALIDATION, EXIT_REGRESSION = 0, 2, 3, 4, 10
      
      
      def load_jsonl(path, label):
          """Read JSONL from a path or stdin. Raises with a line number on bad input."""
          if path == "-":
              text = sys.stdin.read()
          else:
              try:
                  with open(path, "r", encoding="utf-8") as fh:
                      text = fh.read()
              except FileNotFoundError:
                  raise FileNotFoundError(path)
              except IsADirectoryError:
                  raise ValueError(f"{path} is a directory, not a JSONL file")
      
          rows = []
          for lineno, line in enumerate(text.splitlines(), start=1):
              line = line.strip()
              if not line or line.startswith("#"):
                  continue
              try:
                  obj = json.loads(line)
              except json.JSONDecodeError as exc:
                  raise ValueError(f"{label} line {lineno}: not valid JSON ({exc.msg})")
              if not isinstance(obj, dict):
                  raise ValueError(f"{label} line {lineno}: expected a JSON object")
              rows.append(obj)
          return rows
      
      
      def stdev(values):
          """Sample standard deviation. None below two points -- one run is not a spread."""
          n = len(values)
          if n < 2:
              return None
          mean = sum(values) / n
          return math.sqrt(sum((v - mean) ** 2 for v in values) / (n - 1))
      
      
      # Above this many discordant pairs the exact test is switched for the normal
      # approximation. The exact sum is O(n) big-integer work over 2**n: measured at
      # 133 ms for n=2000 but ~100 SECONDS for n=20000, which is a CI hang, not a
      # slow test. The approximation is reliable well below this bound (the usual
      # rule of thumb is b+c >= 25), so nothing accurate is lost by switching here.
      MCNEMAR_EXACT_MAX = 1000
      
      
      def mcnemar(b, c):
          """Two-sided McNemar p-value. Returns (p_value, method).
      
          b = passed before, fails now (regressions);  c = failed before, passes now.
          Only the DISCORDANT pairs carry information -- cases that behaved the same in
          both runs tell you nothing about whether the change helped. Under the null
          b ~ Binomial(b+c, 0.5), so the p-value is a coin-flip tail probability.
      
          Exact below MCNEMAR_EXACT_MAX because eval sets routinely produce b+c < 25,
          where the chi-square/normal approximation is unreliable; normal (with a
          continuity correction) above it, because the exact form does not terminate
          in useful time. The method used is reported so a caller never has to guess.
          """
          n = b + c
          if n == 0:
              return 1.0, "none-discordant"
          if n <= MCNEMAR_EXACT_MAX:
              k = min(b, c)
              tail = sum(math.comb(n, i) for i in range(0, k + 1)) / (2 ** n)
              return min(1.0, 2 * tail), "exact"
          # Continuity-corrected normal approximation; erfc gives the two-sided tail.
          z = (abs(b - c) - 1) / math.sqrt(n)
          return min(1.0, math.erfc(abs(z) / math.sqrt(2))), "normal-approx"
      
      
      def paired_counts(baseline_rows, candidate_rows):
          """Pair per-case results by id. Returns (b, c, both_pass, both_fail, unpaired)."""
          # Test `"id" in r`, NOT the truthiness of r["id"] -- an integer id of 0 or an
          # empty-string id is falsy, and the truthiness form silently DROPPED those
          # cases from the paired test, hiding real regressions. Stringify so 1 and "1"
          # pair with each other rather than becoming two unpaired cases.
          def index(rows):
              return {str(r["id"]): bool(r.get("passed")) for r in rows if "id" in r}
      
          base, cand = index(baseline_rows), index(candidate_rows)
          shared = set(base) & set(cand)
      
          b = sorted(i for i in shared if base[i] and not cand[i])
          c = sorted(i for i in shared if not base[i] and cand[i])
          both_pass = sum(1 for i in shared if base[i] and cand[i])
          both_fail = sum(1 for i in shared if not base[i] and not cand[i])
          unpaired = len(set(base) ^ set(cand))
          return b, c, both_pass, both_fail, unpaired
      
      
      def build_report(history, candidate, window, sigma, paired, alpha, ceilings):
          prior = [r for r in history if r is not candidate][-window:]
          scores = [float(r["score"]) for r in prior if isinstance(r.get("score"), (int, float))]
      
          warnings = []
          # A row whose score is absent or non-numeric (a JSON string "0.9", a null) is
          # unusable. Dropping it in silence let a whole malformed history report
          # "insufficient data" and exit 0 -- a gate failing open with no explanation.
          unusable = len(prior) - len(scores)
          if unusable:
              warnings.append(
                  f"{unusable} history row(s) had a missing or non-numeric 'score' and "
                  "were ignored (scores must be JSON numbers, not strings)"
              )
      
          baseline = round(sum(scores) / len(scores), 4) if scores else None
          spread = stdev(scores)
          spread = round(spread, 4) if spread is not None else None
      
          # A threshold inside the noise band fails on identical code. Sit below the
          # baseline by more than the measured spread -- that is the whole point of
          # keeping a history rather than judging each run standalone.
          #
          # `spread is not None` is deliberate, NOT a truthiness test: a spread of
          # exactly 0.0 is a MEASURED result (a deterministic metric, or a genuinely
          # stable suite) and the most informative history you can have. Treating it as
          # "no data" made the gate report insufficient-data and exit 0 on an
          # unambiguous 0.90 -> 0.70 regression.
          threshold = (round(baseline - sigma * spread, 4)
                       if (baseline is not None and spread is not None) else None)
      
          cand_score = candidate.get("score") if candidate else None
          cand_score = float(cand_score) if isinstance(cand_score, (int, float)) else None
          delta = round(cand_score - baseline, 4) if (cand_score is not None and baseline is not None) else None
          # z is undefined when the spread is exactly 0 (division by zero), but that is
          # the EASY case, not the hard one: with no measured variance any movement is
          # real. Handled explicitly in the verdict below rather than left as None.
          z = None
          if delta is not None and spread:
              z = round(delta / spread, 2)
      
          if len(scores) < 3:
              warnings.append(
                  f"only {len(scores)} prior run(s) in the window; a noise floor needs "
                  "at least 3, and 5 reruns of unchanged code is the honest way to get it"
              )
      
          # Comparing across a dataset, judge or rubric change is confounded: you cannot
          # tell whether the system moved or the measurement did.
          for field, human in (("dataset", "dataset version"), ("judge", "judge model")):
              seen = {r.get(field) for r in prior + ([candidate] if candidate else []) if r.get(field)}
              if len(seen) > 1:
                  warnings.append(
                      f"{human} changed within the window ({', '.join(sorted(map(str, seen)))}) "
                      "-- re-baseline; scores across that boundary are not comparable"
                  )
      
          # --- significance -------------------------------------------------------
          significance = None
          if paired is not None:
              b_ids, c_ids, both_pass, both_fail, unpaired = paired
              p, method = mcnemar(len(b_ids), len(c_ids))
              significance = {
                  "test": f"mcnemar-{method}",
                  "regressed": b_ids,           # passed before, fails now
                  "fixed": c_ids,               # failed before, passes now
                  "n_regressed": len(b_ids),
                  "n_fixed": len(c_ids),
                  "both_pass": both_pass,
                  "both_fail": both_fail,
                  "unpaired_ids": unpaired,
                  "p_value": round(p, 6),
                  "alpha": alpha,
                  "significant": p < alpha,
                  "direction": "worse" if len(b_ids) > len(c_ids)
                               else ("better" if len(c_ids) > len(b_ids) else "unchanged"),
              }
              if unpaired:
                  warnings.append(
                      f"{unpaired} case id(s) appear in only one result set; they are "
                      "excluded from the paired test"
                  )
              if len(b_ids) + len(c_ids) == 0:
                  warnings.append("no discordant cases -- the two runs agree everywhere")
      
          # --- ceilings -----------------------------------------------------------
          # Ceilings, not trends: a trend gate fires on noise, a ceiling encodes a
          # product decision someone actually made.
          breaches = []
          if candidate:
              for key, limit, label in (
                  ("cost_usd", ceilings.get("cost"), "cost_usd"),
                  ("p95_ms", ceilings.get("p95"), "p95_ms"),
              ):
                  value = candidate.get(key)
                  if limit is not None and isinstance(value, (int, float)) and value > limit:
                      breaches.append({"metric": label, "value": value, "ceiling": limit})
      
          # --- verdict ------------------------------------------------------------
          # Paired significance is the strongest evidence available; fall back to the
          # sigma rule only when per-case results were not supplied.
          if significance is not None:
              if significance["significant"] and significance["direction"] == "worse":
                  verdict, reason = "regression", (
                      f"{significance['n_regressed']} case(s) regressed vs "
                      f"{significance['n_fixed']} fixed, p={significance['p_value']} "
                      f"< alpha={alpha}"
                  )
              elif significance["significant"] and significance["direction"] == "better":
                  verdict, reason = "improvement", (
                      f"{significance['n_fixed']} case(s) fixed vs "
                      f"{significance['n_regressed']} regressed, p={significance['p_value']}"
                  )
              else:
                  verdict, reason = "noise", (
                      f"p={significance['p_value']} >= alpha={alpha}; the difference is "
                      "not distinguishable from chance"
                  )
          elif cand_score is None or threshold is None:
              verdict, reason = "insufficient-data", (
                  "no candidate score, or fewer than 2 usable history rows -- rerun the "
                  "suite against unchanged code a few times to establish a noise floor "
                  "before gating on it"
              )
          elif spread == 0:
              # Zero measured variance: every rerun of unchanged code gave the same
              # number, so ANY movement is signal and there is no band to be inside.
              if cand_score < baseline:
                  verdict, reason = "regression", (
                      f"score {cand_score} is below a baseline of {baseline} measured "
                      "with zero variance across the window"
                  )
              elif cand_score > baseline:
                  verdict, reason = "improvement", (
                      f"score {cand_score} exceeds a zero-variance baseline of {baseline}"
                  )
              else:
                  verdict, reason = "noise", f"score is identical to the baseline {baseline}"
          elif cand_score < threshold:
              verdict, reason = "regression", (
                  f"score {cand_score} is below the {sigma}-sigma threshold {threshold}"
              )
          elif z is not None and z > sigma:
              verdict, reason = "improvement", f"score is {z} sigma above baseline"
          else:
              verdict, reason = "noise", (
                  f"score {cand_score} is within the noise band around {baseline}"
              )
      
          if breaches:
              reason += "; " + ", ".join(
                  f"{x['metric']} {x['value']} exceeds ceiling {x['ceiling']}" for x in breaches
              )
      
          # The hillclimb decision is NOT the same question as the CI-gate decision.
          # CI asks "did this get worse?" (noise is fine). A hillclimb asks "is this
          # improvement real?" (noise is not good enough to bank). Only an improvement
          # that clears the floor, or is significant on the paired test, is a KEEP.
          decision = "keep" if verdict == "improvement" else "discard"
      
          return {
              "verdict": verdict,
              "reason": reason,
              "hillclimb_decision": decision,
              "failing": verdict == "regression" or bool(breaches),
              "baseline": baseline,
              "noise_floor": spread,
              "sigma": sigma,
              "recommended_threshold": threshold,
              "candidate_score": cand_score,
              "delta": delta,
              "z": z,
              "window": len(scores),
              "significance": significance,
              "ceiling_breaches": breaches,
              "warnings": warnings,
          }
      
      
      def print_human(r, source):
          out = [f"eval baseline: {source}"]
          out.append(f"  window          {r['window']} prior run(s)")
          out.append(f"  baseline        {r['baseline']}")
          out.append(f"  noise floor     {r['noise_floor']}  (sample stdev)")
          out.append(f"  gate at         {r['recommended_threshold']}  ({r['sigma']} sigma below baseline)")
          out.append(f"  candidate       {r['candidate_score']}   delta {r['delta']}   z {r['z']}")
          sig = r["significance"]
          if sig:
              out.append("")
              out.append("  paired test (McNemar exact)")
              out.append(f"    regressed     {sig['n_regressed']}   fixed {sig['n_fixed']}")
              out.append(f"    unchanged     {sig['both_pass']} pass / {sig['both_fail']} fail")
              out.append(f"    p-value       {sig['p_value']}  (alpha {sig['alpha']})")
              # Naming the flipped cases is what stops people ignoring a red gate.
              if sig["regressed"]:
                  out.append(f"    now failing   {', '.join(sig['regressed'][:15])}")
              if sig["fixed"]:
                  out.append(f"    now passing   {', '.join(sig['fixed'][:15])}")
          for x in r["ceiling_breaches"]:
              out.append(f"  CEILING         {x['metric']} {x['value']} > {x['ceiling']}")
          out.append("")
          out.append(f"  verdict         {r['verdict'].upper()} - {r['reason']}")
          out.append(f"  hillclimb       {r['hillclimb_decision'].upper()}")
          print("\n".join(out))
      
      
      def main(argv=None):
          ap = argparse.ArgumentParser(
              prog="eval-baseline.py",
              description="Distinguish an eval regression from noise using run history "
                          "and McNemar's exact test.",
          )
          ap.add_argument("history", help="rolling run-history JSONL, or - for stdin")
          ap.add_argument("--candidate", metavar="FILE",
                          help="one-row summary for the run under test "
                               "(default: the last row of HISTORY)")
          ap.add_argument("--baseline-results", metavar="FILE",
                          help="per-case JSONL of the baseline run, for the paired test")
          ap.add_argument("--candidate-results", metavar="FILE",
                          help="per-case JSONL of the candidate run, for the paired test")
          ap.add_argument("--window", type=int, default=10, metavar="N",
                          help="prior runs used for baseline and noise floor (default: 10)")
          ap.add_argument("--sigma", type=float, default=2.0, metavar="K",
                          help="how many noise-floor widths below baseline the gate sits "
                               "(default: 2.0)")
          ap.add_argument("--alpha", type=float, default=0.05, metavar="A",
                          help="significance level for the paired test (default: 0.05)")
          ap.add_argument("--max-cost-usd", type=float, metavar="USD",
                          help="fail if the candidate run exceeds this cost")
          ap.add_argument("--max-p95-ms", type=float, metavar="MS",
                          help="fail if the candidate run exceeds this p95 latency")
          ap.add_argument("--accept", action="store_true",
                          help="hillclimb keep/discard gate: exit 0 = KEEP a real improvement, "
                               "exit 10 = DISCARD noise or a regression (INVERTS the default "
                               "meaning of noise -- see the header)")
          ap.add_argument("--json", action="store_true", help="emit the JSON envelope on stdout")
      
          try:
              args = ap.parse_args(argv)
          except SystemExit as exc:
              raise SystemExit(exc.code)
      
          def fail(code, kind, message):
              if args.json:
                  print(json.dumps({"error": {"code": kind, "message": message, "details": {}}}))
              print(f"eval-baseline: {message}", file=sys.stderr)
              return code
      
          if args.window < 1:
              return fail(EXIT_USAGE, "VALIDATION", "--window must be at least 1")
          if args.sigma <= 0:
              return fail(EXIT_USAGE, "VALIDATION", "--sigma must be positive")
          if not (0.0 < args.alpha < 1.0):
              return fail(EXIT_USAGE, "VALIDATION", "--alpha must be strictly between 0 and 1")
          if bool(args.baseline_results) != bool(args.candidate_results):
              return fail(EXIT_USAGE, "VALIDATION",
                          "--baseline-results and --candidate-results must be given together")
      
          try:
              history = load_jsonl(args.history, "history")
              candidate = None
              if args.candidate:
                  rows = load_jsonl(args.candidate, "candidate")
                  if not rows:
                      return fail(EXIT_VALIDATION, "VALIDATION", "candidate file has no rows")
                  candidate = rows[-1]
              elif history:
                  candidate = history[-1]
      
              paired = None
              if args.baseline_results:
                  base_rows = load_jsonl(args.baseline_results, "baseline-results")
                  cand_rows = load_jsonl(args.candidate_results, "candidate-results")
                  if not base_rows or not cand_rows:
                      return fail(EXIT_VALIDATION, "VALIDATION",
                                  "per-case result files must not be empty")
                  paired = paired_counts(base_rows, cand_rows)
          except FileNotFoundError as exc:
              return fail(EXIT_NOT_FOUND, "NOT_FOUND", f"no such file: {exc}")
          except ValueError as exc:
              return fail(EXIT_VALIDATION, "VALIDATION", str(exc))
      
          if not history and candidate is None:
              return fail(EXIT_VALIDATION, "VALIDATION", "history is empty and no --candidate given")
      
          report = build_report(
              history, candidate, args.window, args.sigma, paired, args.alpha,
              {"cost": args.max_cost_usd, "p95": args.max_p95_ms},
          )
      
          if args.json:
              print(json.dumps({
                  "data": report,
                  "meta": {"count": report["window"], "schema": SCHEMA, "source": args.history},
              }, indent=2))
          else:
              print_human(report, args.history)
      
          for w in report["warnings"]:
              print(f"eval-baseline: warning: {w}", file=sys.stderr)
      
          if args.accept:
              return EXIT_OK if report["hillclimb_decision"] == "keep" else EXIT_REGRESSION
          return EXIT_REGRESSION if report["failing"] else EXIT_OK
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • goldenset-audit.py 16.5 KB
      #!/usr/bin/env python3
      """Audit a golden eval set for the rot that silently kills a regression suite.
      
      Checks duplicates and near-duplicates, bucket balance, undated/unexplained cases,
      missing expectations, and drift from a frozen manifest.
      
      Usage:   goldenset-audit.py [OPTIONS] <GOLDEN.jsonl>
      Input:   JSONL, one case per line. Recognised fields: `id`, `bucket`, `added`
               (ISO date), `why`, `input`, `expected`, `criteria`. All optional --
               a missing field becomes a finding, never a crash. `-` reads stdin.
      Output:  stdout -- human-readable report, or a --json envelope
               {"data": {"findings": [...], ...}, "meta": {...}} per §4.
      Stderr:  headers, progress, warnings, errors.
      Exit:    0 clean, 2 usage, 3 not-found, 4 validation (unparseable/empty),
               10 FINDINGS at or above --fail-on severity
      
      Examples:
        goldenset-audit.py evals/golden.jsonl
        goldenset-audit.py evals/golden.jsonl --json | jq '.data.findings[]'
        goldenset-audit.py evals/golden.jsonl --write-freeze manifest.json
        goldenset-audit.py evals/golden.jsonl --freeze manifest.json --fail-on warn
      
      Offline and stdlib-only. --write-freeze is the only write, and it is atomic.
      """
      
      import argparse
      import hashlib
      import json
      import os
      import re
      import sys
      from collections import Counter
      
      SCHEMA = "claude-mods.evals-ops.goldenset-audit/v1"
      FREEZE_SCHEMA = "claude-mods.evals-ops.goldenset-freeze/v1"
      
      EXIT_OK, EXIT_USAGE, EXIT_NOT_FOUND, EXIT_VALIDATION, EXIT_FINDINGS = 0, 2, 3, 4, 10
      
      SEVERITY_ORDER = {"info": 0, "warn": 1, "error": 2}
      
      # TARGET share per bucket, from references/golden-datasets.md, and the WARN band
      # around it. These are deliberately different numbers: the target is what you aim
      # for when composing the set, the band is where a warning fires. A band equal to
      # the target would warn on almost every healthy set and get ignored within a week
      # -- the same "a gate that cries wolf is a dead gate" rule the skill applies to CI.
      # Both numbers are reported, so a warning names the target it is measured against.
      BUCKET_TARGETS = {
          "production": (0.40, 0.50),
          "replay": (0.20, 0.30),
          "adversarial": (0.15, 0.20),
          "edge": (0.10, 0.15),
      }
      BUCKET_BANDS = {
          "production": (0.30, 0.65),
          "replay": (0.10, 0.40),
          "adversarial": (0.08, 0.35),
          "edge": (0.05, 0.30),
      }
      
      ISO_DATE = re.compile(r"^\d{4}-\d{2}-\d{2}")
      WORD = re.compile(r"[a-z0-9]+")
      
      
      def canonical(obj):
          """Stable JSON text for hashing: key order must not change a case's identity."""
          return json.dumps(obj, sort_keys=True, separators=(",", ":"), ensure_ascii=False)
      
      
      # A case's hash is its CONTENT, never its identity or metadata. Excluded:
      #   _line   bookkeeping we inject -- leaked in and made two identical cases on
      #           different lines hash differently, so duplicate detection missed them
      #   id      the whole point of duplicate detection is "two ids, one case"; if id
      #           were hashed, DUPLICATE_CASE could never fire (ids are unique by rule)
      #   bucket  reclassifying a case does not change what it tests
      #   added / why   documentation about the case, not the case
      # Excluding id also keeps the freeze manifest honest: it stores id and hash as
      # separate fields, so a rename shows as removed+added rather than as an edit.
      HASH_EXCLUDE = ("_line", "id", "bucket", "added", "why")
      
      
      def case_hash(case):
          payload = {k: case.get(k) for k in ("input", "expected", "criteria") if k in case}
          if not payload:
              payload = {k: v for k, v in case.items() if k not in HASH_EXCLUDE}
          return hashlib.sha256(canonical(payload).encode("utf-8")).hexdigest()
      
      
      def tokens(case):
          return set(WORD.findall(canonical(case.get("input", "")).lower()))
      
      
      def jaccard(a, b):
          if not a or not b:
              return 0.0
          return len(a & b) / len(a | b)
      
      
      def load_cases(path):
          if path == "-":
              text = sys.stdin.read()
          else:
              try:
                  with open(path, "r", encoding="utf-8") as fh:
                      text = fh.read()
              except FileNotFoundError:
                  raise FileNotFoundError(path)
              except IsADirectoryError:
                  raise ValueError(f"{path} is a directory, not a JSONL file")
      
          cases = []
          for lineno, line in enumerate(text.splitlines(), start=1):
              line = line.strip()
              if not line or line.startswith("#"):
                  continue
              try:
                  obj = json.loads(line)
              except json.JSONDecodeError as exc:
                  raise ValueError(f"line {lineno}: not valid JSON ({exc.msg})")
              if not isinstance(obj, dict):
                  raise ValueError(f"line {lineno}: expected a JSON object")
              obj.setdefault("_line", lineno)
              cases.append(obj)
          return cases
      
      
      def audit(cases, near_threshold, max_pairs):
          findings = []
      
          def add(severity, code, message, **details):
              findings.append(
                  {"severity": severity, "code": code, "message": message, "details": details}
              )
      
          n = len(cases)
          ids = [c.get("id") or f"line-{c['_line']}" for c in cases]
      
          # --- identity -----------------------------------------------------------
          missing_id = [c["_line"] for c in cases if not c.get("id")]
          if missing_id:
              add("warn", "MISSING_ID", f"{len(missing_id)} case(s) have no 'id'",
                  lines=missing_id[:20])
      
          dup_ids = [i for i, c in Counter(ids).items() if c > 1]
          if dup_ids:
              add("error", "DUPLICATE_ID", f"{len(dup_ids)} id(s) used more than once",
                  ids=sorted(dup_ids)[:20])
      
          # --- exact duplicates ---------------------------------------------------
          by_hash = {}
          for cid, case in zip(ids, cases):
              by_hash.setdefault(case_hash(case), []).append(cid)
          exact = {h: members for h, members in by_hash.items() if len(members) > 1}
          if exact:
              add("error", "DUPLICATE_CASE",
                  f"{len(exact)} group(s) of identical cases -- they double-weight one behaviour",
                  groups=[sorted(m) for m in list(exact.values())[:10]])
      
          # --- near-duplicates ----------------------------------------------------
          # O(n^2) on token sets; bounded by --max-pairs so a huge set degrades to a
          # skipped check with a note, never a hang.
          near = []
          if n * (n - 1) // 2 <= max_pairs:
              toks = [tokens(c) for c in cases]
              for i in range(n):
                  for j in range(i + 1, n):
                      sim = jaccard(toks[i], toks[j])
                      if sim >= near_threshold:
                          near.append({"a": ids[i], "b": ids[j], "similarity": round(sim, 3)})
              if near:
                  add("warn", "NEAR_DUPLICATE",
                      f"{len(near)} case pair(s) above {near_threshold} input similarity",
                      pairs=sorted(near, key=lambda p: -p["similarity"])[:15])
          else:
              add("info", "NEAR_DUPLICATE_SKIPPED",
                  f"near-duplicate scan skipped: {n} cases exceeds --max-pairs budget",
                  cases=n, max_pairs=max_pairs)
      
          # --- completeness -------------------------------------------------------
          no_expect = [cid for cid, c in zip(ids, cases)
                       if "expected" not in c and not c.get("criteria")]
          if no_expect:
              add("error", "NO_EXPECTATION",
                  f"{len(no_expect)} case(s) have neither 'expected' nor 'criteria' -- ungradeable",
                  ids=no_expect[:20])
      
          no_why = [cid for cid, c in zip(ids, cases) if not c.get("why")]
          if no_why:
              add("warn", "NO_RATIONALE",
                  f"{len(no_why)} case(s) have no 'why' -- a future maintainer cannot tell "
                  "whether deleting them loses coverage",
                  ids=no_why[:20])
      
          # --- dates --------------------------------------------------------------
          dated = [c.get("added") for c in cases
                   if isinstance(c.get("added"), str) and ISO_DATE.match(c["added"])]
          undated = [cid for cid, c in zip(ids, cases)
                     if not (isinstance(c.get("added"), str) and ISO_DATE.match(c.get("added", "")))]
          if undated:
              add("warn", "UNDATED",
                  f"{len(undated)} case(s) have no ISO 'added' date -- staleness is undetectable",
                  ids=undated[:20])
          newest = max(dated)[:10] if dated else None
      
          # --- bucket balance -----------------------------------------------------
          buckets = Counter(c.get("bucket") or "_unset" for c in cases)
          shares = {b: round(c / n, 4) for b, c in buckets.items()}
          if buckets.get("_unset"):
              add("warn", "NO_BUCKET",
                  f"{buckets['_unset']} case(s) have no 'bucket' -- balance cannot be tracked",
                  count=buckets["_unset"])
          known = [b for b in buckets if b in BUCKET_BANDS]
          if known:
              for bucket, (lo, hi) in BUCKET_BANDS.items():
                  share = shares.get(bucket, 0.0)
                  t_lo, t_hi = BUCKET_TARGETS[bucket]
                  if share < lo:
                      add("warn", "BUCKET_THIN",
                          f"bucket '{bucket}' is {share:.0%} of the set "
                          f"(target {t_lo:.0%}-{t_hi:.0%}, warns below {lo:.0%})",
                          bucket=bucket, share=share, floor=lo, target=[t_lo, t_hi])
                  elif share > hi:
                      add("warn", "BUCKET_HEAVY",
                          f"bucket '{bucket}' is {share:.0%} of the set "
                          f"(target {t_lo:.0%}-{t_hi:.0%}, warns above {hi:.0%})",
                          bucket=bucket, share=share, ceiling=hi, target=[t_lo, t_hi])
      
          # --- size ---------------------------------------------------------------
          if n < 20:
              add("warn", "TOO_SMALL",
                  f"{n} cases; below ~20 a percentage moves too coarsely to interpret", cases=n)
          elif n < 100:
              add("info", "SMALL",
                  f"{n} cases; 100-300 is the working range for a regression set", cases=n)
      
          return findings, {
              "cases": n,
              "buckets": dict(buckets),
              "bucket_shares": shares,
              "newest_added": newest,
              "dated_cases": len(dated),
              "near_duplicate_pairs": len(near),
          }
      
      
      def freeze_manifest(cases):
          entries = sorted(
              ({"id": c.get("id") or f"line-{c['_line']}", "hash": case_hash(c)} for c in cases),
              key=lambda e: (e["id"], e["hash"]),
          )
          digest = hashlib.sha256(
              canonical([[e["id"], e["hash"]] for e in entries]).encode("utf-8")
          ).hexdigest()
          return {"schema": FREEZE_SCHEMA, "count": len(entries), "digest": digest, "cases": entries}
      
      
      def compare_freeze(cases, manifest):
          """Diff the live set against a frozen manifest. Returns findings."""
          if not isinstance(manifest, dict) or "cases" not in manifest:
              return [{"severity": "error", "code": "FREEZE_MALFORMED",
                       "message": "freeze manifest has no 'cases' list", "details": {}}]
      
          # An entry with no id cannot be matched to a live case; index it by its hash so it
          # still reports as removed rather than colliding on a None key.
          frozen = {
              str(e.get("id") or f"unnamed-{e.get('hash', '?')}"): e.get("hash")
              for e in manifest["cases"]
              if isinstance(e, dict)
          }
          live = {c.get("id") or f"line-{c['_line']}": case_hash(c) for c in cases}
      
          added = sorted(set(live) - set(frozen))
          removed = sorted(set(frozen) - set(live))
          changed = sorted(i for i in set(live) & set(frozen) if live[i] != frozen[i])
      
          findings = []
          if changed:
              findings.append({
                  "severity": "error", "code": "FREEZE_CASE_CHANGED",
                  "message": f"{len(changed)} frozen case(s) were edited in place -- this is "
                             "fitting the test to the code; delete and re-add instead",
                  "details": {"ids": changed[:20]},
              })
          if removed:
              findings.append({
                  "severity": "error", "code": "FREEZE_CASE_REMOVED",
                  "message": f"{len(removed)} frozen case(s) are gone -- scores are no longer "
                             "comparable to the frozen baseline",
                  "details": {"ids": removed[:20]},
              })
          if added:
              findings.append({
                  "severity": "warn", "code": "FREEZE_CASE_ADDED",
                  "message": f"{len(added)} case(s) added since freeze -- re-baseline before "
                             "comparing scores across this boundary",
                  "details": {"ids": added[:20]},
              })
          return findings
      
      
      def atomic_write_json(path, payload):
          tmp = f"{path}.tmp"
          with open(tmp, "w", encoding="utf-8") as fh:
              json.dump(payload, fh, indent=2, sort_keys=True)
              fh.write("\n")
          os.replace(tmp, path)
      
      
      def print_human(findings, stats, source):
          out = [f"golden-set audit: {source}", f"  cases           {stats['cases']}"]
          if stats["buckets"]:
              shares = ", ".join(
                  f"{b} {stats['bucket_shares'][b]:.0%}" for b in sorted(stats["buckets"])
              )
              out.append(f"  buckets         {shares}")
          out.append(f"  newest added    {stats['newest_added'] or '-'}")
          out.append("")
          if not findings:
              out.append("  no findings")
          else:
              for sev in ("error", "warn", "info"):
                  rows = [f for f in findings if f["severity"] == sev]
                  for f in rows:
                      out.append(f"  [{sev.upper():<5}] {f['code']}: {f['message']}")
          print("\n".join(out))
      
      
      def main(argv=None):
          parser = argparse.ArgumentParser(
              prog="goldenset-audit.py",
              description="Audit a golden eval set for duplicates, imbalance, staleness and drift.",
          )
          parser.add_argument("golden", help="JSONL golden set, or - for stdin")
          parser.add_argument("--freeze", metavar="MANIFEST",
                              help="compare against a freeze manifest written by --write-freeze")
          parser.add_argument("--write-freeze", metavar="MANIFEST",
                              help="write a freeze manifest for this set (atomic)")
          parser.add_argument("--fail-on", choices=("error", "warn", "info"), default="error",
                              help="lowest severity that exits 10 (default: error)")
          parser.add_argument("--near-threshold", type=float, default=0.9, metavar="R",
                              help="Jaccard similarity at which two inputs are near-duplicates "
                                   "(default: 0.9)")
          parser.add_argument("--max-pairs", type=int, default=200000, metavar="N",
                              help="skip the O(n^2) near-duplicate scan above this pair count "
                                   "(default: 200000)")
          parser.add_argument("--json", action="store_true", help="emit the JSON envelope on stdout")
      
          try:
              args = parser.parse_args(argv)
          except SystemExit as exc:
              raise SystemExit(exc.code)
      
          def fail(code, kind, message):
              if args.json:
                  print(json.dumps({"error": {"code": kind, "message": message, "details": {}}}))
              print(f"goldenset-audit: {message}", file=sys.stderr)
              return code
      
          if not (0.0 < args.near_threshold <= 1.0):
              return fail(EXIT_USAGE, "VALIDATION", "--near-threshold must be in (0, 1]")
          if args.max_pairs < 0:
              return fail(EXIT_USAGE, "VALIDATION", "--max-pairs must not be negative")
          if args.freeze and args.write_freeze:
              return fail(EXIT_USAGE, "VALIDATION", "--freeze and --write-freeze are mutually exclusive")
      
          try:
              cases = load_cases(args.golden)
          except FileNotFoundError as exc:
              return fail(EXIT_NOT_FOUND, "NOT_FOUND", f"no such file: {exc}")
          except ValueError as exc:
              return fail(EXIT_VALIDATION, "VALIDATION", str(exc))
      
          if not cases:
              return fail(EXIT_VALIDATION, "VALIDATION", "golden set is empty")
      
          findings, stats = audit(cases, args.near_threshold, args.max_pairs)
      
          if args.freeze:
              try:
                  with open(args.freeze, "r", encoding="utf-8") as fh:
                      manifest = json.load(fh)
              except FileNotFoundError:
                  return fail(EXIT_NOT_FOUND, "NOT_FOUND", f"no such freeze manifest: {args.freeze}")
              except json.JSONDecodeError as exc:
                  return fail(EXIT_VALIDATION, "VALIDATION", f"freeze manifest is not JSON: {exc.msg}")
              findings.extend(compare_freeze(cases, manifest))
      
          if args.write_freeze:
              try:
                  atomic_write_json(args.write_freeze, freeze_manifest(cases))
              except OSError as exc:
                  return fail(EXIT_VALIDATION, "VALIDATION", f"cannot write manifest: {exc}")
              print(f"goldenset-audit: wrote freeze manifest {args.write_freeze}", file=sys.stderr)
      
          threshold = SEVERITY_ORDER[args.fail_on]
          triggering = [f for f in findings if SEVERITY_ORDER[f["severity"]] >= threshold]
      
          if args.json:
              print(json.dumps({
                  "data": {"findings": findings, "stats": stats, "fail_on": args.fail_on},
                  "meta": {"count": len(findings), "schema": SCHEMA, "source": args.golden},
              }, indent=2))
          else:
              print_human(findings, stats, args.golden)
      
          return EXIT_FINDINGS if triggering else EXIT_OK
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • judge-calibration.py 13.1 KB
      #!/usr/bin/env python3
      """Measure an LLM judge against human labels: Cohen kappa, confusion, bias probes.
      
      Usage:   judge-calibration.py [OPTIONS] <LABELS.jsonl>
      Input:   JSONL, one record per case. Required fields: `human`, `judge` (any
               hashable label -- "pass"/"fail", true/false, 1-5). Optional: `id`,
               `length` (or --verbosity-field NAME) for the verbosity-bias probe,
               and `position` ("first"/"second") for the position-bias probe.
               Use `-` to read the JSONL from stdin.
      Output:  stdout -- human-readable report, or a --json envelope
               {"data": {...}, "meta": {...}} per SKILL-RESOURCE-PROTOCOL.md §4.
      Stderr:  headers, progress, warnings, errors.
      Exit:    0 calibrated (kappa >= --min-kappa), 2 usage, 3 not-found,
               4 validation (unparseable/empty/missing fields),
               10 UNDER-CALIBRATED (ran fine, kappa below threshold)
      
      Examples:
        judge-calibration.py labels.jsonl
        judge-calibration.py labels.jsonl --min-kappa 0.8
        judge-calibration.py labels.jsonl --json | jq '.data.kappa'
        judge-calibration.py labels.jsonl --verbosity-field output_chars
        cat labels.jsonl | judge-calibration.py - --json
      
      Offline and stdlib-only: this scores labels you already have, it never calls a model.
      """
      
      import argparse
      import json
      import sys
      from collections import Counter, defaultdict
      
      SCHEMA = "claude-mods.evals-ops.judge-calibration/v1"
      
      EXIT_OK, EXIT_USAGE, EXIT_NOT_FOUND, EXIT_VALIDATION, EXIT_UNDER = 0, 2, 3, 4, 10
      
      # Landis & Koch style bands, as used for judge calibration in practice.
      # < 0.6 means the RUBRIC needs work, not the judge model -- see references/llm-judge.md.
      BANDS = [
          (0.80, "strong", "production-ready; safe to gate on with a margin"),
          (0.60, "substantial", "usable; keep advisory or gate loosely"),
          (0.40, "moderate", "rubric needs work before this gates anything"),
          (0.20, "fair", "rubric is the problem, not the model"),
          (float("-inf"), "poor", "no better than chance; rewrite the rubric"),
      ]
      
      
      def norm(value):
          """Normalise a label to a comparable string. true/1/'PASS' must all agree."""
          if isinstance(value, bool):
              return "pass" if value else "fail"
          if isinstance(value, str):
              return value.strip().lower()
          return str(value)
      
      
      def cohens_kappa(pairs):
          """Cohen kappa for two raters over nominal labels. Returns None if undefined."""
          n = len(pairs)
          if n == 0:
              return None
          observed = sum(1 for a, b in pairs if a == b) / n
          a_counts, b_counts = Counter(a for a, _ in pairs), Counter(b for _, b in pairs)
          expected = sum((a_counts[k] / n) * (b_counts.get(k, 0) / n) for k in a_counts)
          if expected >= 1.0:
              # Both raters used a single identical label: agreement is total but chance-
              # corrected agreement is undefined (0/0). Report it rather than dividing.
              return None
          return (observed - expected) / (1.0 - expected)
      
      
      def pearson(xs, ys):
          """Pearson correlation. Returns None when a series has no variance."""
          n = len(xs)
          if n < 3:
              return None
          mx, my = sum(xs) / n, sum(ys) / n
          sxy = sum((x - mx) * (y - my) for x, y in zip(xs, ys))
          sxx = sum((x - mx) ** 2 for x in xs)
          syy = sum((y - my) ** 2 for y in ys)
          if sxx <= 0 or syy <= 0:
              return None
          return sxy / ((sxx**0.5) * (syy**0.5))
      
      
      def band_for(kappa):
          for floor, name, advice in BANDS:
              if kappa >= floor:
                  return name, advice
          return "poor", "rewrite the rubric"
      
      
      def load_records(path):
          """Read JSONL from a path or stdin. Raises ValueError with a line number."""
          if path == "-":
              text = sys.stdin.read()
          else:
              try:
                  with open(path, "r", encoding="utf-8") as fh:
                      text = fh.read()
              except FileNotFoundError:
                  raise FileNotFoundError(path)
              except IsADirectoryError:
                  raise ValueError(f"{path} is a directory, not a JSONL file")
      
          records = []
          for lineno, line in enumerate(text.splitlines(), start=1):
              line = line.strip()
              if not line or line.startswith("#"):
                  continue
              try:
                  obj = json.loads(line)
              except json.JSONDecodeError as exc:
                  raise ValueError(f"line {lineno}: not valid JSON ({exc.msg})")
              if not isinstance(obj, dict):
                  raise ValueError(f"line {lineno}: expected a JSON object")
              records.append((lineno, obj))
          return records
      
      
      def analyse(records, verbosity_field):
          pairs, ids, skipped = [], [], []
          lengths, judge_numeric = [], []
          position_rows = defaultdict(list)
      
          for lineno, obj in records:
              if "human" not in obj or "judge" not in obj:
                  skipped.append({"line": lineno, "reason": "missing 'human' or 'judge'"})
                  continue
              h, j = norm(obj["human"]), norm(obj["judge"])
              pairs.append((h, j))
              ids.append(obj.get("id", f"line-{lineno}"))
      
              # Verbosity probe: does judge score track output length among cases the
              # HUMAN scored identically? Correlation there is bias, not signal.
              length = obj.get(verbosity_field)
              jn = obj["judge"]
              # bool is a subclass of int in Python, so a `true` in the length field
              # would silently read as 1.0 and yield a confident, meaningless
              # correlation. A length is never a boolean -- reject it explicitly.
              if isinstance(length, bool):
                  length = None
              if isinstance(length, (int, float)) and isinstance(jn, (int, float, bool)):
                  lengths.append(float(length))
                  judge_numeric.append(float(jn))
      
              pos = obj.get("position")
              if isinstance(pos, str):
                  position_rows[pos.strip().lower()].append(1 if h == j else 0)
      
          return pairs, ids, skipped, lengths, judge_numeric, position_rows
      
      
      def build_report(pairs, ids, skipped, lengths, judge_numeric, position_rows, min_kappa):
          n = len(pairs)
          agreement = sum(1 for a, b in pairs if a == b) / n
          kappa = cohens_kappa(pairs)
      
          confusion = Counter(pairs)
          labels = sorted({lab for pair in pairs for lab in pair})
      
          disagreements = [
              {"id": cid, "human": h, "judge": j}
              for cid, (h, j) in zip(ids, pairs)
              if h != j
          ]
      
          # Per-class recall from the human's point of view: where does the judge go wrong?
          # A judge that only ever under-passes is safe to gate on; one that over-passes is not.
          per_class = {}
          for lab in labels:
              total = sum(c for (h, _), c in confusion.items() if h == lab)
              hit = confusion.get((lab, lab), 0)
              per_class[lab] = {
                  "human_count": total,
                  "judge_agreed": hit,
                  "recall": round(hit / total, 4) if total else None,
              }
      
          verbosity_r = pearson(lengths, judge_numeric)
          position = {
              pos: {"n": len(v), "agreement": round(sum(v) / len(v), 4)}
              for pos, v in position_rows.items()
              if v
          }
      
          if kappa is None:
              band, advice, calibrated = "undefined", (
                  "every case shares one label -- kappa is undefined; "
                  "stratify the calibration sample across the judge's own verdicts"
              ), False
          else:
              band, advice = band_for(kappa)
              calibrated = kappa >= min_kappa
      
          warnings = []
          if n < 50:
              warnings.append(
                  f"only {n} labelled cases; 50-200 is the recommended range for a "
                  "meaningful kappa"
              )
          if len(labels) < 2:
              warnings.append("only one distinct label present -- the sample is not stratified")
          if verbosity_r is not None and abs(verbosity_r) >= 0.4:
              warnings.append(
                  f"verbosity probe: judge score correlates {verbosity_r:+.2f} with output "
                  "length -- separate correctness from style in the rubric"
              )
          if len(position) >= 2:
              vals = [v["agreement"] for v in position.values()]
              if max(vals) - min(vals) >= 0.1:
                  warnings.append(
                      "position probe: agreement differs by "
                      f"{max(vals) - min(vals):.2f} across presentation positions -- "
                      "run both orders and average, or score absolutely"
                  )
          if skipped:
              warnings.append(f"{len(skipped)} record(s) skipped (missing required fields)")
      
          return {
              "n": n,
              "kappa": round(kappa, 4) if kappa is not None else None,
              "band": band,
              "advice": advice,
              "raw_agreement": round(agreement, 4),
              "min_kappa": min_kappa,
              "calibrated": calibrated,
              "labels": labels,
              "confusion": [
                  {"human": h, "judge": j, "count": c} for (h, j), c in sorted(confusion.items())
              ],
              "per_class": per_class,
              "disagreements": disagreements,
              "probes": {
                  "verbosity_correlation": round(verbosity_r, 4) if verbosity_r is not None else None,
                  "position_agreement": position,
              },
              "skipped": skipped,
              "warnings": warnings,
          }
      
      
      def print_human(report, source):
          out = []
          out.append(f"judge calibration: {source}")
          out.append(f"  cases           {report['n']}")
          kappa = report["kappa"]
          out.append(
              f"  cohen kappa     {kappa if kappa is not None else 'undefined'}  "
              f"({report['band']})"
          )
          out.append(f"  raw agreement   {report['raw_agreement']}  (kappa is the honest one)")
          out.append(f"  threshold       {report['min_kappa']}")
          out.append("")
          out.append("  confusion (human -> judge)")
          for row in report["confusion"]:
              mark = " " if row["human"] == row["judge"] else "!"
              out.append(f"    {mark} {row['human']:>12} -> {row['judge']:<12} {row['count']}")
          out.append("")
          out.append("  per human class")
          for lab, stats in report["per_class"].items():
              recall = stats["recall"]
              out.append(
                  f"    {lab:>12}  n={stats['human_count']:<4} "
                  f"agreed={stats['judge_agreed']:<4} recall={recall if recall is not None else '-'}"
              )
          probes = report["probes"]
          if probes["verbosity_correlation"] is not None or probes["position_agreement"]:
              out.append("")
              out.append("  bias probes")
              if probes["verbosity_correlation"] is not None:
                  out.append(f"    verbosity r     {probes['verbosity_correlation']:+.4f}")
              for pos, stats in probes["position_agreement"].items():
                  out.append(f"    position {pos:<8} n={stats['n']:<4} agreement={stats['agreement']}")
          if report["disagreements"]:
              out.append("")
              out.append(f"  disagreements ({len(report['disagreements'])})")
              for d in report["disagreements"][:20]:
                  out.append(f"    {d['id']}: human={d['human']} judge={d['judge']}")
              if len(report["disagreements"]) > 20:
                  out.append(f"    ... {len(report['disagreements']) - 20} more")
          out.append("")
          out.append(f"  verdict         {report['advice']}")
          print("\n".join(out))
      
      
      def main(argv=None):
          parser = argparse.ArgumentParser(
              prog="judge-calibration.py",
              description="Cohen kappa and bias probes for an LLM judge vs human labels.",
              add_help=True,
          )
          parser.add_argument("labels", help="JSONL of {human, judge, ...} records, or - for stdin")
          parser.add_argument(
              "--min-kappa", type=float, default=0.6,
              help="kappa at or above which the judge counts as calibrated (default: 0.6)",
          )
          parser.add_argument(
              "--verbosity-field", default="length", metavar="NAME",
              help="numeric field carrying output length for the verbosity probe (default: length)",
          )
          parser.add_argument("--json", action="store_true", help="emit the JSON envelope on stdout")
      
          try:
              args = parser.parse_args(argv)
          except SystemExit as exc:
              # argparse exits 2 on bad args and 0 on --help; both already match the protocol.
              raise SystemExit(exc.code)
      
          if not (0.0 <= args.min_kappa <= 1.0):
              print("judge-calibration: --min-kappa must be between 0 and 1", file=sys.stderr)
              return EXIT_USAGE
      
          def fail(code, kind, message):
              if args.json:
                  print(json.dumps({"error": {"code": kind, "message": message, "details": {}}}))
              print(f"judge-calibration: {message}", file=sys.stderr)
              return code
      
          try:
              records = load_records(args.labels)
          except FileNotFoundError as exc:
              return fail(EXIT_NOT_FOUND, "NOT_FOUND", f"no such file: {exc}")
          except ValueError as exc:
              return fail(EXIT_VALIDATION, "VALIDATION", str(exc))
      
          pairs, ids, skipped, lengths, judge_numeric, position_rows = analyse(
              records, args.verbosity_field
          )
          if not pairs:
              return fail(
                  EXIT_VALIDATION, "VALIDATION",
                  "no usable records (every line missing 'human' or 'judge')",
              )
      
          report = build_report(
              pairs, ids, skipped, lengths, judge_numeric, position_rows, args.min_kappa
          )
      
          if args.json:
              print(json.dumps({
                  "data": report,
                  "meta": {"count": report["n"], "schema": SCHEMA, "source": args.labels},
              }, indent=2))
          else:
              print_human(report, args.labels)
      
          for warning in report["warnings"]:
              print(f"judge-calibration: warning: {warning}", file=sys.stderr)
      
          return EXIT_OK if report["calibrated"] else EXIT_UNDER
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
  • tests
    • run.sh 35.1 KB
      #!/usr/bin/env bash
      # Self-test for evals-ops scripts.
      #
      # Offline and deterministic: builds throwaway JSONL fixtures with KNOWN correct
      # answers (kappa computed by hand below), asserts the documented exit codes and
      # the actual numbers, then cleans up. Resolves paths relative to itself so it
      # works in the repo and once installed to ~/.claude/skills/evals-ops/.
      #
      # Usage:   bash tests/run.sh
      # Exit:    0 all pass, 1 one or more failures
      set -uo pipefail
      
      HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
      SKILL="$(dirname "$HERE")"
      SCRIPTS="$SKILL/scripts"
      CAL="$SCRIPTS/judge-calibration.py"
      AUD="$SCRIPTS/goldenset-audit.py"
      BAS="$SCRIPTS/eval-baseline.py"
      
      # Probe python by EXECUTING it. `command -v python3` finds the Windows Store
      # app-execution stub, which exists on PATH but exits 49 non-interactively.
      PYTHON=""
      for c in python3 python py; do
          if "$c" -c 'import sys' >/dev/null 2>&1; then PYTHON="$c"; break; fi
      done
      if [ -z "$PYTHON" ]; then
          echo "evals-ops tests: no working python3 found - skipping" >&2
          exit 0
      fi
      
      SB="$(mktemp -d)"; trap 'rm -rf "$SB"' EXIT
      PASS=0; FAIL=0
      ok()  { echo "  ok   $*"; PASS=$((PASS + 1)); }
      bad() { echo "  FAIL $*"; FAIL=$((FAIL + 1)); }
      
      # exit_is <want> <label> -- <command...>
      exit_is() {
          local want="$1" label="$2"; shift 3
          "$@" >/dev/null 2>&1; local got=$?
          if [ "$got" -eq "$want" ]; then ok "$label (exit $got)"; else bad "$label (want $want, got $got)"; fi
      }
      
      # json_eq <jq-ish python path> <expected> <label> -- <command...>
      # Reads stdout as JSON and compares a dotted path. Uses python, not jq, so the
      # suite has no dependency beyond the interpreter it already requires.
      json_eq() {
          local path="$1" want="$2" label="$3"; shift 4
          local out got
          out="$("$@" 2>/dev/null)"
          got="$(printf '%s' "$out" | "$PYTHON" -c '
      import json, sys
      doc = json.load(sys.stdin)
      for part in sys.argv[1].split("."):
          doc = doc[int(part)] if part.isdigit() else doc[part]
      print(doc)
      ' "$path" 2>/dev/null)"
          if [ "$got" = "$want" ]; then ok "$label ($path = $got)"; else bad "$label ($path: want $want, got '$got')"; fi
      }
      
      # emits <pattern> <label> -- <command...>
      # Capture stdout into a variable BEFORE grepping. Piping the script straight into
      # grep would let `set -o pipefail` surface the script's own domain exit code (10)
      # as the `if` condition, failing the assertion for the wrong reason.
      emits() {
          local pattern="$1" label="$2"; shift 3
          local out
          out="$("$@" 2>/dev/null)"
          if printf '%s' "$out" | grep -q -- "$pattern"; then ok "$label"; else bad "$label (no '$pattern' in output)"; fi
      }
      
      echo "== evals-ops self-test ($PYTHON)"
      
      # --- protocol surface -------------------------------------------------------
      exit_is 0 "judge-calibration --help"  -- "$PYTHON" "$CAL" --help
      exit_is 0 "goldenset-audit --help"    -- "$PYTHON" "$AUD" --help
      exit_is 0 "eval-baseline --help"      -- "$PYTHON" "$BAS" --help
      exit_is 2 "judge-calibration no args" -- "$PYTHON" "$CAL"
      exit_is 2 "goldenset-audit no args"   -- "$PYTHON" "$AUD"
      exit_is 2 "eval-baseline no args"     -- "$PYTHON" "$BAS"
      exit_is 2 "judge-calibration rejects out-of-range --min-kappa" \
          -- "$PYTHON" "$CAL" /dev/null --min-kappa 5
      exit_is 3 "judge-calibration missing file"  -- "$PYTHON" "$CAL" "$SB/nope.jsonl"
      exit_is 3 "goldenset-audit missing file"    -- "$PYTHON" "$AUD" "$SB/nope.jsonl"
      
      printf 'not json at all\n' > "$SB/bad.jsonl"
      exit_is 4 "judge-calibration malformed JSONL" -- "$PYTHON" "$CAL" "$SB/bad.jsonl"
      exit_is 4 "goldenset-audit malformed JSONL"   -- "$PYTHON" "$AUD" "$SB/bad.jsonl"
      
      : > "$SB/empty.jsonl"
      exit_is 4 "goldenset-audit empty set" -- "$PYTHON" "$AUD" "$SB/empty.jsonl"
      
      # --- judge-calibration: arithmetic --------------------------------------------
      # Perfect agreement on a balanced set: observed 1.0, expected 0.5, kappa = 1.0.
      {
        for i in 1 2 3 4 5; do echo "{\"id\":\"p$i\",\"human\":\"pass\",\"judge\":\"pass\"}"; done
        for i in 1 2 3 4 5; do echo "{\"id\":\"f$i\",\"human\":\"fail\",\"judge\":\"fail\"}"; done
      } > "$SB/perfect.jsonl"
      json_eq "data.kappa" "1.0" "kappa = 1.0 on perfect balanced agreement" \
          -- "$PYTHON" "$CAL" "$SB/perfect.jsonl" --json
      exit_is 0 "perfect agreement is calibrated" -- "$PYTHON" "$CAL" "$SB/perfect.jsonl"
      
      # Chance-level agreement. human 5 pass / 5 fail; judge 4 pass / 6 fail; 5 agree.
      # observed 0.5; expected (0.5*0.4)+(0.5*0.6) = 0.5  ->  kappa exactly 0.0.
      {
        echo '{"id":"a1","human":"pass","judge":"pass"}'
        echo '{"id":"a2","human":"pass","judge":"pass"}'
        echo '{"id":"a3","human":"pass","judge":"fail"}'
        echo '{"id":"a4","human":"pass","judge":"fail"}'
        echo '{"id":"a5","human":"pass","judge":"fail"}'
        echo '{"id":"b1","human":"fail","judge":"fail"}'
        echo '{"id":"b2","human":"fail","judge":"fail"}'
        echo '{"id":"b3","human":"fail","judge":"fail"}'
        echo '{"id":"b4","human":"fail","judge":"pass"}'
        echo '{"id":"b5","human":"fail","judge":"pass"}'
      } > "$SB/chance.jsonl"
      json_eq "data.kappa" "0.0" "chance-level agreement scores kappa 0.0" \
          -- "$PYTHON" "$CAL" "$SB/chance.jsonl" --json
      json_eq "data.raw_agreement" "0.5" "chance set raw agreement is 0.5" \
          -- "$PYTHON" "$CAL" "$SB/chance.jsonl" --json
      exit_is 10 "under-calibrated set exits 10" -- "$PYTHON" "$CAL" "$SB/chance.jsonl"
      
      # A genuinely moderate judge: 5/5 balanced both sides, 8 of 10 agree.
      # observed 0.8; expected 0.5  ->  kappa exactly 0.6, the gate boundary.
      {
        for i in 1 2 3 4; do echo "{\"id\":\"m$i\",\"human\":\"pass\",\"judge\":\"pass\"}"; done
        echo '{"id":"m5","human":"pass","judge":"fail"}'
        for i in 6 7 8 9; do echo "{\"id\":\"m$i\",\"human\":\"fail\",\"judge\":\"fail\"}"; done
        echo '{"id":"m10","human":"fail","judge":"pass"}'
      } > "$SB/moderate.jsonl"
      json_eq "data.kappa" "0.6" "moderate judge scores kappa 0.6" \
          -- "$PYTHON" "$CAL" "$SB/moderate.jsonl" --json
      json_eq "data.band" "substantial" "kappa 0.6 lands in the substantial band" \
          -- "$PYTHON" "$CAL" "$SB/moderate.jsonl" --json
      
      # THE point of using kappa at all: on an imbalanced set a judge that answers
      # "pass" unconditionally gets 90% RAW agreement and kappa 0.0. If these two ever
      # report the same number, the chance correction has been broken.
      {
        for i in $(seq 1 9); do echo "{\"id\":\"y$i\",\"human\":\"pass\",\"judge\":\"pass\"}"; done
        echo '{"id":"n1","human":"fail","judge":"pass"}'
      } > "$SB/imbalanced.jsonl"
      json_eq "data.raw_agreement" "0.9" "imbalanced set: raw agreement flatters at 0.9" \
          -- "$PYTHON" "$CAL" "$SB/imbalanced.jsonl" --json
      json_eq "data.kappa" "0.0" "imbalanced set: kappa correctly reports 0.0" \
          -- "$PYTHON" "$CAL" "$SB/imbalanced.jsonl" --json
      exit_is 10 "constant judge is not calibrated" -- "$PYTHON" "$CAL" "$SB/imbalanced.jsonl"
      
      # Degenerate case: both raters used one identical label, so chance agreement is
      # 1.0 and kappa is 0/0. Must report null and say so, not divide by zero.
      {
        for i in 1 2 3; do echo "{\"id\":\"s$i\",\"human\":\"pass\",\"judge\":\"pass\"}"; done
      } > "$SB/single.jsonl"
      json_eq "data.kappa" "None" "single-label sample reports kappa as null" \
          -- "$PYTHON" "$CAL" "$SB/single.jsonl" --json
      json_eq "data.band" "undefined" "single-label sample is banded 'undefined'" \
          -- "$PYTHON" "$CAL" "$SB/single.jsonl" --json
      exit_is 10 "undefined kappa is never treated as calibrated" -- "$PYTHON" "$CAL" "$SB/single.jsonl"
      
      # Label normalisation: true/false, 1/0 and "PASS" must compare like their peers.
      {
        echo '{"id":"n1","human":true,"judge":"PASS"}'
        echo '{"id":"n2","human":true,"judge":"pass"}'
        echo '{"id":"n3","human":false,"judge":"Fail"}'
        echo '{"id":"n4","human":false,"judge":"fail"}'
      } > "$SB/norm.jsonl"
      json_eq "data.kappa" "1.0" "boolean and cased labels normalise to agreement" \
          -- "$PYTHON" "$CAL" "$SB/norm.jsonl" --json
      
      # --min-kappa is the gate. moderate.jsonl scores exactly 0.6.
      exit_is 0  "--min-kappa 0.6 accepts kappa 0.6 (inclusive)" -- "$PYTHON" "$CAL" "$SB/moderate.jsonl" --min-kappa 0.6
      exit_is 0  "--min-kappa 0.5 accepts kappa 0.6"             -- "$PYTHON" "$CAL" "$SB/moderate.jsonl" --min-kappa 0.5
      exit_is 10 "--min-kappa 0.8 rejects kappa 0.6"             -- "$PYTHON" "$CAL" "$SB/moderate.jsonl" --min-kappa 0.8
      
      # Disagreements are enumerated so a human can go look at them.
      json_eq "meta.count" "10" "report counts every case" \
          -- "$PYTHON" "$CAL" "$SB/chance.jsonl" --json
      json_eq "data.disagreements.0.id" "a3" "first disagreement is reported by id" \
          -- "$PYTHON" "$CAL" "$SB/chance.jsonl" --json
      
      # Confusion direction matters more than the headline: a judge that only ever
      # under-passes is safe to gate on. Recall must be per-human-class, not global.
      json_eq "data.per_class.fail.recall" "0.0" "per-class recall exposes a judge that never says fail" \
          -- "$PYTHON" "$CAL" "$SB/imbalanced.jsonl" --json
      json_eq "data.per_class.pass.recall" "1.0" "per-class recall is 1.0 for the class it always picks" \
          -- "$PYTHON" "$CAL" "$SB/imbalanced.jsonl" --json
      
      # Verbosity probe: judge score rises monotonically with length while the human
      # scored every case identically -- that is the bias, and it must be detected.
      {
        echo '{"id":"v1","human":1,"judge":1,"length":10}'
        echo '{"id":"v2","human":1,"judge":2,"length":100}'
        echo '{"id":"v3","human":1,"judge":3,"length":200}'
        echo '{"id":"v4","human":1,"judge":4,"length":300}'
        echo '{"id":"v5","human":1,"judge":5,"length":400}'
      } > "$SB/verbose.jsonl"
      json_eq "data.probes.verbosity_correlation" "0.9998" "verbosity probe detects length correlation" \
          -- "$PYTHON" "$CAL" "$SB/verbose.jsonl" --json
      
      # stdin path
      if [ "$("$PYTHON" "$CAL" - --json < "$SB/perfect.jsonl" 2>/dev/null | "$PYTHON" -c 'import json,sys; print(json.load(sys.stdin)["data"]["n"])' 2>/dev/null)" = "10" ]; then
          ok "reads JSONL from stdin"
      else
          bad "reads JSONL from stdin"
      fi
      
      # stdout must stay parseable JSON under --json even when the script exits 10 and
      # warnings are firing on stderr. Capture first -- pipefail would mask this.
      _out="$("$PYTHON" "$CAL" "$SB/verbose.jsonl" --json 2>/dev/null)"
      if printf '%s' "$_out" | "$PYTHON" -c 'import json,sys; json.load(sys.stdin)' >/dev/null 2>&1; then
          ok "stdout stays pure JSON on a nonzero exit"
      else
          bad "stdout stays pure JSON on a nonzero exit"
      fi
      
      # ...and the human report must NOT be JSON, so the two modes cannot be confused.
      _out="$("$PYTHON" "$CAL" "$SB/moderate.jsonl" 2>/dev/null)"
      case "$_out" in
          *"cohen kappa"*) ok "human mode prints a readable report" ;;
          *) bad "human mode prints a readable report" ;;
      esac
      
      # --- goldenset-audit --------------------------------------------------------
      mk_case() { # id bucket
          printf '{"id":"%s","bucket":"%s","added":"2026-01-0%s","why":"case %s",' "$1" "$2" "$((RANDOM % 9 + 1))" "$1"
          printf '"input":{"q":"unique question about %s here"},"expected":{"a":"%s"}}\n' "$1" "$1"
      }
      {
        for i in $(seq 1 12); do mk_case "prod-$i" production; done
        for i in $(seq 1 6);  do mk_case "replay-$i" replay; done
        for i in $(seq 1 5);  do mk_case "adv-$i" adversarial; done
        for i in $(seq 1 4);  do mk_case "edge-$i" edge; done
      } > "$SB/healthy.jsonl"
      exit_is 0 "healthy balanced set audits clean" -- "$PYTHON" "$AUD" "$SB/healthy.jsonl"
      json_eq "data.stats.cases" "27" "case count is reported" -- "$PYTHON" "$AUD" "$SB/healthy.jsonl" --json
      
      # Duplicate id + identical case content are both errors (exit 10).
      { cat "$SB/healthy.jsonl"; mk_case "prod-1" production; } > "$SB/dupid.jsonl"
      exit_is 10 "duplicate id is a finding" -- "$PYTHON" "$AUD" "$SB/dupid.jsonl"
      emits "DUPLICATE_ID" "duplicate id reports DUPLICATE_ID" -- "$PYTHON" "$AUD" "$SB/dupid.jsonl" --json
      
      # A case with neither expected nor criteria is ungradeable.
      { cat "$SB/healthy.jsonl"; echo '{"id":"x1","bucket":"edge","added":"2026-02-01","why":"w","input":{"q":"z"}}'; } > "$SB/noexp.jsonl"
      emits "NO_EXPECTATION" "ungradeable case reports NO_EXPECTATION" -- "$PYTHON" "$AUD" "$SB/noexp.jsonl" --json
      exit_is 10 "ungradeable case is an error" -- "$PYTHON" "$AUD" "$SB/noexp.jsonl"
      
      # Bucket skew: 25 production, 1 of everything else -> BUCKET_HEAVY + BUCKET_THIN.
      {
        for i in $(seq 1 25); do mk_case "p-$i" production; done
        mk_case "r-1" replay; mk_case "a-1" adversarial; mk_case "e-1" edge
      } > "$SB/skew.jsonl"
      emits "BUCKET_HEAVY" "production-heavy set reports BUCKET_HEAVY" -- "$PYTHON" "$AUD" "$SB/skew.jsonl" --json
      emits "BUCKET_THIN" "starved buckets report BUCKET_THIN"      -- "$PYTHON" "$AUD" "$SB/skew.jsonl" --json
      # Balance findings are warnings, so the default --fail-on error must NOT trip.
      exit_is 0  "bucket skew is advisory under --fail-on error" -- "$PYTHON" "$AUD" "$SB/skew.jsonl"
      exit_is 10 "bucket skew trips --fail-on warn"              -- "$PYTHON" "$AUD" "$SB/skew.jsonl" --fail-on warn
      
      # Near-duplicate detection, and the --max-pairs escape hatch.
      {
        mk_case "u-1" production
        echo '{"id":"d-1","bucket":"edge","added":"2026-01-01","why":"w","input":{"q":"the quick brown fox jumps over the lazy dog"},"expected":{"a":1}}'
        echo '{"id":"d-2","bucket":"edge","added":"2026-01-01","why":"w","input":{"q":"the quick brown fox jumps over the lazy dog"},"expected":{"a":2}}'
      } > "$SB/near.jsonl"
      emits '"NEAR_DUPLICATE"' "identical inputs report NEAR_DUPLICATE" -- "$PYTHON" "$AUD" "$SB/near.jsonl" --json
      emits "NEAR_DUPLICATE_SKIPPED" "--max-pairs 0 skips the O(n^2) scan and says so" -- "$PYTHON" "$AUD" "$SB/near.jsonl" --max-pairs 0 --json
      
      # --- freeze manifest --------------------------------------------------------
      exit_is 0 "--write-freeze on a clean set" -- "$PYTHON" "$AUD" "$SB/healthy.jsonl" --write-freeze "$SB/m.json"
      [ -f "$SB/m.json" ] && ok "freeze manifest written" || bad "freeze manifest written"
      [ -f "$SB/m.json.tmp" ] && bad "temp file left behind" || ok "atomic write leaves no .tmp"
      exit_is 0 "unchanged set matches its freeze" -- "$PYTHON" "$AUD" "$SB/healthy.jsonl" --freeze "$SB/m.json"
      
      # Editing a frozen case in place is the cardinal sin -- must be an error.
      "$PYTHON" - "$SB/healthy.jsonl" "$SB/edited.jsonl" <<'PYEOF'
      import json, sys
      src, dst = sys.argv[1], sys.argv[2]
      lines = [json.loads(l) for l in open(src, encoding="utf-8") if l.strip()]
      lines[0]["expected"] = {"a": "quietly changed to make it pass"}
      with open(dst, "w", encoding="utf-8") as fh:
          for c in lines:
              fh.write(json.dumps(c) + "\n")
      PYEOF
      emits "FREEZE_CASE_CHANGED" "in-place edit of a frozen case reports FREEZE_CASE_CHANGED" -- "$PYTHON" "$AUD" "$SB/edited.jsonl" --freeze "$SB/m.json" --json
      exit_is 10 "in-place edit fails the freeze check" -- "$PYTHON" "$AUD" "$SB/edited.jsonl" --freeze "$SB/m.json"
      
      # Adding a case is a warning (re-baseline), not an error.
      { cat "$SB/healthy.jsonl"; mk_case "new-1" adversarial; } > "$SB/grown.jsonl"
      emits "FREEZE_CASE_ADDED" "added case reports FREEZE_CASE_ADDED" -- "$PYTHON" "$AUD" "$SB/grown.jsonl" --freeze "$SB/m.json" --json
      exit_is 0 "added case is advisory under --fail-on error" -- "$PYTHON" "$AUD" "$SB/grown.jsonl" --freeze "$SB/m.json"
      
      # Field order must not change a case hash -- otherwise every reformat is "drift".
      "$PYTHON" - "$SB/healthy.jsonl" "$SB/reordered.jsonl" <<'PYEOF'
      import json, sys
      src, dst = sys.argv[1], sys.argv[2]
      with open(dst, "w", encoding="utf-8") as fh:
          for line in open(src, encoding="utf-8"):
              if line.strip():
                  c = json.loads(line)
                  fh.write(json.dumps(dict(reversed(list(c.items())))) + "\n")
      PYEOF
      exit_is 0 "key reordering does not count as drift" -- "$PYTHON" "$AUD" "$SB/reordered.jsonl" --freeze "$SB/m.json"
      
      exit_is 2 "--freeze and --write-freeze are mutually exclusive" \
          -- "$PYTHON" "$AUD" "$SB/healthy.jsonl" --freeze "$SB/m.json" --write-freeze "$SB/m2.json"
      exit_is 3 "missing freeze manifest" -- "$PYTHON" "$AUD" "$SB/healthy.jsonl" --freeze "$SB/absent.json"
      
      # --- eval-baseline: noise floor, significance, ceilings ---------------------
      # History with a KNOWN spread. scores .88 .85 .89 .86 .88
      #   mean         = 0.872
      #   sample stdev = sqrt(0.00108/4) = 0.016432 -> 0.0164
      #   2-sigma gate = 0.872 - 2(0.0164) = 0.8392
      for v in 0.88 0.85 0.89 0.86 0.88; do
        echo "{\"score\":$v,\"n\":30,\"dataset\":\"v3\",\"judge\":\"j1\"}"
      done > "$SB/hist.jsonl"
      echo '{"score":0.87,"n":30,"dataset":"v3","judge":"j1","cost_usd":1.20,"p95_ms":3000}' > "$SB/cand.jsonl"
      
      json_eq "data.baseline" "0.872" "baseline is the window mean" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --json
      json_eq "data.noise_floor" "0.0164" "noise floor is the sample stdev" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --json
      json_eq "data.recommended_threshold" "0.8392" "gate sits 2 sigma below baseline" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --json
      json_eq "data.verdict" "noise" "a 0.002 drop inside the band is noise" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --json
      exit_is 0 "noise does not fail the gate" -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl"
      
      echo '{"score":0.62,"n":30,"dataset":"v3","judge":"j1"}' > "$SB/drop.jsonl"
      json_eq "data.verdict" "regression" "a drop past the threshold is a regression" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/drop.jsonl" --json
      exit_is 10 "regression exits 10" -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/drop.jsonl"
      
      # --sigma widens the band: 0.83 is below the 2-sigma gate (0.8392) but above the
      # 4-sigma gate (0.8064), so the same score must classify differently.
      echo '{"score":0.83,"n":30,"dataset":"v3","judge":"j1"}' > "$SB/mid.jsonl"
      json_eq "data.verdict" "regression" "0.83 is a regression at the default 2 sigma" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/mid.jsonl" --json
      json_eq "data.verdict" "noise" "0.83 is noise at 4 sigma (wider band)" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/mid.jsonl" --sigma 4 --json
      
      # --- McNemar exact test -----------------------------------------------------
      # 30 cases all passing at baseline; 8 now fail. b=8, c=0.
      # p = 2 * C(8,0)/2^8 = 2/256 = 0.007812
      for i in $(seq 1 30); do echo "{\"id\":\"c$i\",\"passed\":true}"; done > "$SB/base-res.jsonl"
      { for i in $(seq 1 22); do echo "{\"id\":\"c$i\",\"passed\":true}"; done
        for i in $(seq 23 30); do echo "{\"id\":\"c$i\",\"passed\":false}"; done; } > "$SB/cand-res.jsonl"
      
      PAIR=(--baseline-results "$SB/base-res.jsonl" --candidate-results "$SB/cand-res.jsonl")
      json_eq "data.significance.p_value" "0.007812" "McNemar exact p for b=8,c=0" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${PAIR[@]}" --json
      json_eq "data.verdict" "regression" "paired test overrides a flat aggregate score" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${PAIR[@]}" --json
      json_eq "data.significance.regressed.0" "c23" "regressed cases are named, not just counted" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${PAIR[@]}" --json
      exit_is 10 "significant paired regression exits 10" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${PAIR[@]}"
      
      # CHURN, not regression. 8 break and 7 are fixed: b=8, c=7, p = 1.0 exactly.
      # This is the case references/regression-gating.md calls out: the honest verdict
      # is "noise", and the value of the run is the enumerated lists, not the verdict.
      # If this ever flips to "regression", the exact test has been replaced by
      # something that over-claims - which is the failure mode the whole file warns about.
      { for i in $(seq 1 15); do echo "{\"id\":\"c$i\",\"passed\":true}"; done
        for i in $(seq 16 22); do echo "{\"id\":\"c$i\",\"passed\":false}"; done
        for i in $(seq 23 30); do echo "{\"id\":\"c$i\",\"passed\":true}"; done; } > "$SB/churn-base.jsonl"
      { for i in $(seq 1 15); do echo "{\"id\":\"c$i\",\"passed\":true}"; done
        for i in $(seq 16 22); do echo "{\"id\":\"c$i\",\"passed\":true}"; done
        for i in $(seq 23 30); do echo "{\"id\":\"c$i\",\"passed\":false}"; done; } > "$SB/churn-cand.jsonl"
      
      CHURN=(--baseline-results "$SB/churn-base.jsonl" --candidate-results "$SB/churn-cand.jsonl")
      json_eq "data.significance.p_value" "1.0" "8 broken / 7 fixed is p=1.0, not a regression" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${CHURN[@]}" --json
      json_eq "data.verdict" "noise" "churn is reported as noise, honestly" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${CHURN[@]}" --json
      json_eq "data.significance.n_fixed" "7" "...but the 7 fixed cases are still enumerated" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${CHURN[@]}" --json
      exit_is 0 "churn does not fail the gate" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${CHURN[@]}"
      
      # Identical runs: zero discordant pairs, p = 1.0 by definition.
      json_eq "data.significance.p_value" "1.0" "identical runs give p=1.0" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" \
             --baseline-results "$SB/base-res.jsonl" --candidate-results "$SB/base-res.jsonl" --json
      
      # Direction matters: the same pair reversed is an improvement, not a regression.
      REV=(--baseline-results "$SB/cand-res.jsonl" --candidate-results "$SB/base-res.jsonl")
      json_eq "data.verdict" "improvement" "a significant improvement is labelled as such" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${REV[@]}" --json
      exit_is 0 "an improvement never fails the gate" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" "${REV[@]}"
      
      # Cases present in only one result set cannot be paired and must be flagged.
      head -20 "$SB/base-res.jsonl" > "$SB/short-res.jsonl"
      json_eq "data.significance.unpaired_ids" "10" "unpaired case ids are counted" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" \
             --baseline-results "$SB/short-res.jsonl" --candidate-results "$SB/cand-res.jsonl" --json
      
      # --- confounded comparisons must warn ---------------------------------------
      # A dataset or judge change inside the window means the MEASUREMENT moved, not
      # necessarily the system. Silence here would be the worst possible failure.
      for v in 0.88 0.85 0.89; do echo "{\"score\":$v,\"dataset\":\"v3\",\"judge\":\"j1\"}"; done > "$SB/mixed.jsonl"
      echo '{"score":0.86,"dataset":"v4","judge":"j1"}' >> "$SB/mixed.jsonl"
      _err="$("$PYTHON" "$BAS" "$SB/mixed.jsonl" 2>&1 >/dev/null)"
      case "$_err" in
          *"dataset version changed"*) ok "warns when the dataset version changes mid-window" ;;
          *) bad "warns when the dataset version changes mid-window" ;;
      esac
      _err="$("$PYTHON" "$BAS" "$SB/hist.jsonl" --window 2 --candidate "$SB/cand.jsonl" 2>&1 >/dev/null)"
      case "$_err" in
          *"noise floor needs"*) ok "warns when the window is too small for a noise floor" ;;
          *) bad "warns when the window is too small for a noise floor" ;;
      esac
      
      # --- ceilings: absolute, not trend ------------------------------------------
      exit_is 10 "cost ceiling breach fails"    -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --max-cost-usd 0.50
      exit_is 0  "cost under ceiling passes"    -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --max-cost-usd 5.00
      exit_is 10 "latency ceiling breach fails" -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --max-p95-ms 1000
      exit_is 0  "latency under ceiling passes" -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --max-p95-ms 9000
      
      # --- argument validation ----------------------------------------------------
      exit_is 2 "paired flags must be given together" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --baseline-results "$SB/base-res.jsonl"
      exit_is 2 "--alpha must be in (0,1)"  -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --alpha 1.5
      exit_is 2 "--sigma must be positive"  -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --sigma 0
      exit_is 2 "--window must be >= 1"     -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --window 0
      exit_is 3 "missing history file"      -- "$PYTHON" "$BAS" "$SB/absent.jsonl"
      exit_is 4 "malformed history"         -- "$PYTHON" "$BAS" "$SB/bad.jsonl"
      
      # --- shipped assets must survive the skill's own tools ----------------------
      # An asset the skill's own scripts reject is worse than no asset at all.
      ASSETS="$SKILL/assets"
      exit_is 0 "example golden set passes its own audit" -- "$PYTHON" "$AUD" "$ASSETS/golden-set.example.jsonl"
      exit_is 0 "runner template compiles"                -- "$PYTHON" -m py_compile "$ASSETS/eval-runner.template.py"
      [ -f "$ASSETS/judge-rubric.template.md" ] && ok "judge rubric template ships" || bad "judge rubric template ships"
      [ -f "$ASSETS/eval-gate.template.yml" ]   && ok "CI gate template ships"      || bad "CI gate template ships"
      
      
      # --- --accept: the hillclimb keep/discard gate ------------------------------
      # The whole point of this flag is that it asks a DIFFERENT question from CI.
      # CI: "did this get worse?" -> noise is fine, exit 0.
      # Hillclimb: "is this improvement real?" -> noise is NOT good enough to bank.
      # These four assertions pin that inversion. If a refactor ever makes the noise
      # case exit 0 under --accept, the loop silently goes back to banking noise.
      echo '{"score":0.97,"n":30,"dataset":"v3","judge":"j1"}' > "$SB/win.jsonl"
      
      exit_is 0  "noise passes CI (did it get worse? no)" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl"
      exit_is 10 "noise is a DISCARD under --accept (is it real? no)" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --accept
      exit_is 0  "a real improvement is a KEEP under --accept" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/win.jsonl" --accept
      exit_is 10 "a regression is a DISCARD under --accept" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/drop.jsonl" --accept
      
      json_eq "data.hillclimb_decision" "discard" "noise reports decision=discard" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" --json
      json_eq "data.hillclimb_decision" "keep" "a real improvement reports decision=keep" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/win.jsonl" --json
      
      # The decision is exposed in the report regardless of the flag, so a caller can
      # read it without adopting the inverted exit semantics.
      json_eq "data.hillclimb_decision" "discard" "decision is reported without --accept too" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/drop.jsonl" --json
      
      # A significant paired improvement is a KEEP even when the aggregate barely moved -
      # this is the case a raw score comparison gets wrong in the generous direction.
      json_eq "data.hillclimb_decision" "keep" "paired significance drives the keep decision" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" \
             --baseline-results "$SB/cand-res.jsonl" --candidate-results "$SB/base-res.jsonl" --json
      # ...and churn (8 broken / 7 fixed, p=1.0) is NOT a keep, despite 7 fixes.
      json_eq "data.hillclimb_decision" "discard" "churn is never banked as an improvement" \
          -- "$PYTHON" "$BAS" "$SB/hist.jsonl" --candidate "$SB/cand.jsonl" \
             --baseline-results "$SB/churn-base.jsonl" --candidate-results "$SB/churn-cand.jsonl" --json
      
      # --- the iterate <-> evals-ops seam is documented in BOTH directions --------
      # A one-way cross-reference rots: whichever side gets edited without the other
      # leaves a dangling claim. These assert the link exists from each end.
      ITERATE="$SKILL/../iterate/SKILL.md"
      if [ -f "$ITERATE" ]; then
          grep -q "eval-baseline.py" "$ITERATE" \
              && ok "iterate points at the noise-floor gate" \
              || bad "iterate points at the noise-floor gate"
          grep -q "hillclimbing.md" "$ITERATE" \
              && ok "iterate links the hillclimbing reference" \
              || bad "iterate links the hillclimbing reference"
      else
          ok "iterate skill not present (installed standalone) - seam check skipped"
      fi
      grep -q "iterate" "$SKILL/references/hillclimbing.md" \
          && ok "hillclimbing defers loop mechanics to iterate" \
          || bad "hillclimbing defers loop mechanics to iterate"
      
      
      # ===========================================================================
      # Regressions found by the 2026-08-31 adversarial review. Every one of these
      # failed OPEN (exit 0 / silently dropped data) before the fix, which is the
      # dangerous direction for a gate: it reports "fine" while measuring nothing.
      # Each assertion names the defect so a future edit cannot quietly restore it.
      # ===========================================================================
      
      # R1. A MEASURED spread of exactly 0.0 is the most informative history possible
      # (a deterministic metric, or a genuinely stable suite). `if spread:` treated it
      # as "no data" -- so an unambiguous 0.90 -> 0.70 collapse reported
      # INSUFFICIENT-DATA and exited 0.
      for v in 1 2 3 4 5; do echo '{"score":0.90,"dataset":"v3","judge":"j1"}'; done > "$SB/flat.jsonl"
      echo '{"score":0.70,"dataset":"v3","judge":"j1"}' > "$SB/flat-down.jsonl"
      echo '{"score":0.95,"dataset":"v3","judge":"j1"}' > "$SB/flat-up.jsonl"
      echo '{"score":0.90,"dataset":"v3","judge":"j1"}' > "$SB/flat-same.jsonl"
      
      json_eq "data.noise_floor" "0.0" "zero-variance history reports a floor of 0.0, not null" \
          -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --candidate "$SB/flat-same.jsonl" --json
      json_eq "data.verdict" "regression" "R1: a drop on a zero-variance baseline is a regression" \
          -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --candidate "$SB/flat-down.jsonl" --json
      exit_is 10 "R1: and it actually fails the gate" \
          -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --candidate "$SB/flat-down.jsonl"
      json_eq "data.verdict" "improvement" "R1: a rise on a zero-variance baseline is an improvement" \
          -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --candidate "$SB/flat-up.jsonl" --json
      exit_is 0 "R1: that improvement is a KEEP under --accept" \
          -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --candidate "$SB/flat-up.jsonl" --accept
      # An identical score is NOT an improvement -- a no-op must never be banked.
      json_eq "data.hillclimb_decision" "discard" "R1: an unchanged score is never a KEEP" \
          -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --candidate "$SB/flat-same.jsonl" --json
      
      # R2. `if r.get("id")` is falsy for an integer id of 0, so case 0 was silently
      # dropped from the paired test -- hiding a real regression.
      printf '{"id":0,"passed":true}\n{"id":1,"passed":true}\n{"id":2,"passed":true}\n' > "$SB/id0-base.jsonl"
      printf '{"id":0,"passed":false}\n{"id":1,"passed":true}\n{"id":2,"passed":true}\n' > "$SB/id0-cand.jsonl"
      json_eq "data.significance.n_regressed" "1" "R2: an integer id of 0 is not dropped" \
          -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --baseline-results "$SB/id0-base.jsonl" \
             --candidate-results "$SB/id0-cand.jsonl" --json
      json_eq "data.significance.unpaired_ids" "0" "R2: numeric and string ids pair with each other" \
          -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --baseline-results "$SB/id0-base.jsonl" \
             --candidate-results "$SB/id0-cand.jsonl" --json
      
      # R3. A history row with a null or STRING score was dropped in silence, so a
      # whole malformed file reported insufficient-data and exited 0 with no clue why.
      printf '{"score":null}\n{"score":"0.9"}\n{"score":0.88}\n' > "$SB/badscores.jsonl"
      _err="$("$PYTHON" "$BAS" "$SB/badscores.jsonl" 2>&1 >/dev/null)"
      case "$_err" in
          *"non-numeric 'score'"*) ok "R3: non-numeric scores are reported, not swallowed" ;;
          *) bad "R3: non-numeric scores are reported, not swallowed" ;;
      esac
      
      # R4. The exact McNemar test is O(n) big-int work over 2**n: 133ms at b+c=2000
      # but ~100 SECONDS at 20000, i.e. a CI hang. Large inputs must switch method and
      # say which one they used.
      json_eq "data.significance.test" "mcnemar-exact" "R4: small discordant sets use the exact test" \
          -- "$PYTHON" "$BAS" "$SB/flat.jsonl" --baseline-results "$SB/base-res.jsonl" \
             --candidate-results "$SB/cand-res.jsonl" --json
      "$PYTHON" - "$SCRIPTS/eval-baseline.py" <<'PYEOF'
      import importlib.util, sys, time
      spec = importlib.util.spec_from_file_location("eb", sys.argv[1])
      m = importlib.util.module_from_spec(spec); spec.loader.exec_module(m)
      started = time.time()
      p, method = m.mcnemar(12000, 8000)          # 20k discordant pairs
      elapsed = time.time() - started
      assert method == "normal-approx", f"want normal-approx above the cutoff, got {method}"
      assert elapsed < 5, f"20k discordant pairs took {elapsed:.1f}s -- the exact test is back"
      # The approximation must still agree with the exact test near the cutoff.
      pe, _ = m.mcnemar(600, 400); pa = 2 * 0  # exact at 1000 pairs
      assert m.mcnemar(8, 0) == (0.0078125, "exact"), "exact path changed"
      PYEOF
      [ $? -eq 0 ] && ok "R4: large discordant sets switch to the normal approximation, fast" \
                   || bad "R4: large discordant sets switch to the normal approximation, fast"
      
      # R5. `_line` and `id` leaked into the content hash, so DUPLICATE_CASE could
      # never fire on the very thing it exists to catch: one case, two ids.
      #
      # THE FIXTURE MUST OMIT input/expected/criteria. case_hash() has two branches and
      # only the FALLBACK branch -- taken when none of those fields are present -- ever
      # saw _line and id. The first draft of this test used a fixture WITH `expected`,
      # so it took the other branch and passed against deliberately re-broken code.
      # Mutation-testing the assertion is what exposed that; do not "simplify" this
      # fixture by giving the cases an `expected`, or the test goes inert again.
      printf '{"id":"a","bucket":"edge","added":"2026-01-01","why":"w","note":"same"}\n' > "$SB/dupmin.jsonl"
      printf '{"id":"b","bucket":"production","added":"2026-05-05","why":"other","note":"same"}\n' >> "$SB/dupmin.jsonl"
      emits "DUPLICATE_CASE" "R5: same content under two ids is a duplicate (fallback hash)" \
          -- "$PYTHON" "$AUD" "$SB/dupmin.jsonl" --json
      
      # And the same on the primary branch, where content fields exist and only the
      # metadata differs.
      printf '{"id":"c","bucket":"edge","added":"2026-01-01","why":"w","input":{"q":"same"},"expected":{"a":1}}\n' > "$SB/dupfull.jsonl"
      printf '{"id":"d","bucket":"production","added":"2026-05-05","why":"other","input":{"q":"same"},"expected":{"a":1}}\n' >> "$SB/dupfull.jsonl"
      emits "DUPLICATE_CASE" "R5: differing bucket/added/why do not mask a duplicate" \
          -- "$PYTHON" "$AUD" "$SB/dupfull.jsonl" --json
      
      # R6. bool is a subclass of int, so `"length": true` read as a length of 1.0 and
      # produced a confident, entirely meaningless verbosity correlation (-0.866).
      printf '{"id":"a","human":1,"judge":1,"length":true}\n{"id":"b","human":1,"judge":5,"length":false}\n{"id":"c","human":1,"judge":3,"length":true}\n' > "$SB/boollen.jsonl"
      json_eq "data.probes.verbosity_correlation" "None" "R6: a boolean length yields no correlation" \
          -- "$PYTHON" "$CAL" "$SB/boollen.jsonl" --json
      
      # R7/R8. The shipped CI template asserted unverified action majors and a script
      # path that contradicted the one iterate documents. Both are now explicit ADAPT
      # points; asserting a bare major here again would be the regression.
      TPL="$SKILL/assets/eval-gate.template.yml"
      if grep -qE 'uses: actions/[a-z-]+@v[0-9]' "$TPL"; then
          bad "R7: template asserts an unverified action major (use a flagged placeholder)"
      else
          ok "R7: template does not assert unverified action majors"
      fi
      grep -q 'EVALS_OPS' "$TPL" \
          && ok "R8: template routes script paths through one adaptable variable" \
          || bad "R8: template routes script paths through one adaptable variable"
      
      # R9. The reference stated bucket TARGETS while the script warned on a wider
      # BAND, with nothing saying they were different numbers on purpose.
      grep -q "targets to compose against" "$SKILL/references/golden-datasets.md" \
          && ok "R9: the reference distinguishes bucket targets from the warn band" \
          || bad "R9: the reference distinguishes bucket targets from the warn band"
      emits "target 40%-50%" "R9: a bucket warning names the target it is measured against" \
          -- "$PYTHON" "$AUD" "$SB/skew.jsonl" --json
      
      
      echo
      echo "evals-ops: $PASS passed, $FAIL failed"
      [ "$FAIL" -eq 0 ] || exit 1
      
  • SKILL.md 17.4 KB
    ---
    name: evals-ops
    description: "Build and run evals for LLM and agent systems: golden datasets, LLM-as-a-judge with bias control, trajectory/step/outcome scoring, adversarial refuters, and CI regression gates. Triggers on: eval, evals, eval harness, golden dataset, golden set, LLM-as-a-judge, judge rubric, judge bias, judge calibration, Cohen kappa, agreement with human labels, pass@k, pass^k, trajectory eval, step-level eval, tool-call accuracy, regression gate, eval CI, blocking vs advisory eval, faithfulness score, DeepEval, Braintrust, Opik, Langfuse, AgentOps, did my prompt change make it worse, is my agent getting better."
    license: MIT
    metadata:
      author: claude-mods
      related-skills: "testing-ops, claude-api-ops, iterate, loop-ops, fleet-ops"
    ---
    
    # Evals Ops
    
    **Evals are the prerequisite, not the polish.** You cannot tune a prompt, a retriever, a
    compaction strategy or a memory layer without a harness that says whether the change made
    things better. Teams that skip this ship vibes and learn about regressions from users.
    
    This skill is the operational layer: what to measure, how to build the dataset, how to
    make a judge trustworthy, and how to gate CI on it without teaching everyone to ignore red.
    
    ## Route first
    
    | The ask | Go to |
    |---|---|
    | "What should I even measure?" | [Three levels](#three-levels-of-agent-eval) → `references/eval-taxonomy.md` |
    | "Where do the test cases come from?" | [Golden set](#the-golden-set) → `references/golden-datasets.md` |
    | "My judge disagrees with me / is it any good?" | [Judges](#llm-as-a-judge) → `references/llm-judge.md` |
    | "Verify a finding is real, not plausible" | [Refuters](#adversarial-verification) → `references/adversarial-verification.md` |
    | "Is my RAG retrieval any good?" | [Retrieval](#retrieval) → `references/retrieval-eval.md` |
    | "Where do the human labels come from?" | `references/annotation-workflow.md` |
    | "Should this block the merge?" | [Gating](#regression-gating) → `references/regression-gating.md` |
    | "Did this change really make it worse?" | [Is the drop real](#is-the-drop-real) |
    | "Optimise against the eval / run it overnight" | [Hillclimbing](#hillclimbing) → `references/hillclimbing.md` |
    | "Which platform should we use?" | `references/tooling-landscape.md` |
    | "Just give me a starting file" | [Assets](#assets) — golden set, rubric, runner, CI gate |
    
    ## The 60-second version
    
    1. **Write 20 cases before you write a metric.** A dataset you can eyeball beats a metric
       you cannot interpret. Grow to 100-300, then *freeze* it.
    2. **Prefer a deterministic assertion to any judge.** The JSON parsed, the tool was called
       with the right argument, the query returned 3 rows - free, instant, zero variance.
       Reach for a judge only where correctness is genuinely a matter of language.
    3. **Score the trajectory, not just the answer.** Most teams only check the final artifact
       and are surprised when a right answer came from a wrong path that breaks tomorrow.
    4. **Calibrate the judge against humans before trusting it.** Cohen kappa, not raw
       agreement. `scripts/judge-calibration.py` does the arithmetic and the verdict.
    5. **Blocking gates must be deterministic. Judge metrics start advisory.** One flaky red
       permanently devalues the signal.
    
    ## Three levels of agent eval
    
    Most teams do only the third, then wonder why quality is unpredictable.
    
    | Level | Question | Signal | Typical evaluator |
    |---|---|---|---|
    | **Outcome** | Is the final artifact correct? | Binary or scored end state | Deterministic assertion, unit test, judge |
    | **Step** | Was *this* tool call right? | Per-span: tool choice, arg shape, arg values | Schema/argument assertion, span-level judge |
    | **Trajectory** | Was the *path* sensible? | Sequence, loops, redundancy, cost | Reference-trajectory match, rubric judge |
    
    The failure that motivates all three: an agent reaches the right end state by an accidental
    route - the *lucky pass*. Outcome-only scoring records that as a win, and the same case
    fails next week when the accident does not recur. Conversely a trajectory-only score
    punishes a legitimately novel-but-correct path. **Gate on outcome; keep step and trajectory
    as the diagnostics that tell you why the gate moved.**
    
    For multi-turn or stateful agents also report **pass^k** (all k independent runs of the same
    case succeed) alongside **pass@k** (any of k succeeded). pass@k flatters a non-deterministic
    agent; pass^k is the number that predicts production. Full treatment:
    `references/eval-taxonomy.md`.
    
    ## The golden set
    
    A golden set is a **reviewed, versioned, deliberately frozen** collection of inputs with
    trusted expected outputs. Frozen matters: a set that grows every sprint cannot tell you
    whether last week's number moved because the system changed or because the set did.
    
    **Composition - four buckets, not one:**
    
    | Bucket | Source | Why |
    |---|---|---|
    | Production sample | Real traffic, stratified | Keeps the score connected to what users actually do |
    | Failure replays | Every incident that reached a human | Regression protection; the easiest cases to justify |
    | Adversarial | Injections, contradictions, refusal-bait | The class both agents and judges fail silently on |
    | Edge cases | Empty, huge, ambiguous, multilingual | Where deterministic code breaks first |
    
    **Sizing:** 20 to start, 100-300 for a working regression set, 200-500 once you have
    production traffic to sample. Beyond that you are usually buying latency, not signal - add
    cases when a new *failure class* appears, and record in the case itself why it exists.
    
    `scripts/goldenset-audit.py` checks a set for the rot that accumulates: duplicates, a
    bucket that quietly became 90% of the set, undated cases, and drift from a frozen manifest.
    
    ```bash
    python3 scripts/goldenset-audit.py evals/golden.jsonl --json | jq '.data.findings[]'
    ```
    
    Depth: `references/golden-datasets.md`.
    
    ## LLM-as-a-judge
    
    A judge is a measurement instrument. Instruments need calibration, and this one has
    documented, reproducible biases:
    
    | Bias | What it does | Mitigation |
    |---|---|---|
    | **Position** | Prefers whichever candidate was shown first | Run both orders and average; or score absolutely, not pairwise |
    | **Verbosity** | Rates longer answers higher regardless of quality | Separate correctness from style in the rubric; penalise unsupported length |
    | **Self-preference** | Rates its own model family's output higher | Judge with a different family than the one under test |
    | **Scale drift** | 1-5 scores cluster and shift between model versions | Binary pass/fail against explicit criteria; pin the judge model version |
    
    **Panel vs N-identical.** Three calls to the same judge with the same rubric mostly buys the
    same bias three times. A **panel with distinct lenses** - one asks "is this supported by the
    source?", one "does it follow the stated policy?", one "would this reproduce?" - finds
    failure modes redundancy structurally cannot. Use N-identical only to measure the judge's
    own variance, which is a different question worth asking once.
    
    **When a judge is the wrong tool:** if you can express the criterion as code, do. A judge
    costs money, adds latency, drifts across model versions, and has variance a regex does not.
    Judges earn their place on faithfulness, tone, policy compliance, and "is this a reasonable
    answer to an open question" - nowhere else.
    
    **Calibrate before you trust.** Label 50-200 cases by hand, run the judge on the same cases,
    and compute Cohen kappa (raw agreement lies when classes are imbalanced):
    
    ```bash
    python3 scripts/judge-calibration.py evals/labels.jsonl --min-kappa 0.6
    # exit 0 = calibrated;  exit 10 = below threshold, fix the rubric before shipping it
    ```
    
    kappa >= 0.8 production-ready - 0.6-0.8 substantial, usable with care - below 0.6 the rubric
    is the problem, not the model. Re-sample ~50 fresh cases periodically; judges drift when the
    underlying model version moves. Depth, including bias-probe design: `references/llm-judge.md`.
    
    **Measure the human-human ceiling first.** A judge cannot beat the agreement two people
    achieve with each other. Two annotators on 30-50 shared cases gives you that number - and
    if it is below ~0.6, the rubric is ambiguous and every label you produce against it is
    wasted. How to run the sessions, stratify the sample, and adjudicate disagreements:
    `references/annotation-workflow.md`.
    
    ## Adversarial verification
    
    For findings rather than scores - bug reports, audit results, review comments - flip the
    prompt: **ask the verifier to REFUTE, not to confirm.** "Try to refute this finding; default
    to refuted if uncertain" kills plausible-but-wrong results that an "is this correct?" prompt
    waves through, because agreement is the path of least resistance for a model.
    
    Then take a majority: run 3 refuters, keep the finding only if at least 2 fail to refute it.
    Prefer **perspective-diverse** refuters (correctness / security / does-it-actually-reproduce)
    over three identical skeptics - same reasoning as judge panels.
    
    This composes with the parallel-work skills rather than duplicating them: `fleet-ops` and
    `parallel-ops` own the fan-out mechanics; this skill owns the scoring contract the refuters
    return. See `references/adversarial-verification.md`.
    
    ## Retrieval
    
    RAG is the most common eval target and the most commonly mis-measured. Scoring only the
    final answer averages two independent failures into one uninterpretable number:
    
    |  | Right context | Wrong context |
    |---|---|---|
    | **Answer correct** | Working | **Lucky** - the model knew it anyway; scores as a pass |
    | **Answer wrong** | **Generation bug** - chunking, prompt, model | **Retrieval bug** - embeddings, index, query rewriting |
    
    Record the retrieved chunk ids next to every answer and that opaque score becomes a 2x2
    you can assign to a team. Gate on **recall@k** (a precision failure degrades an answer; a
    recall failure makes a correct one impossible) and on citation-id validity, which is free
    and catches confident answers attached to unrelated sources. Retrieval is the one place
    deterministic scoring genuinely dominates - you have ground-truth ids, so skip the judge.
    
    The bucket almost everyone omits: **questions the corpus cannot answer.** Without them the
    suite cannot detect hallucination under retrieval failure, which is what users hit most.
    Metrics, the six failure classes, and ground-truth construction: `references/retrieval-eval.md`.
    
    ## Regression gating
    
    The rule that keeps a gate alive: **a blocking check must never be flaky.**
    
    | Check | Gate |
    |---|---|
    | Deterministic assertions (schema, tool-call, exact match) | **Blocking.** Any failure fails CI. |
    | Judge metrics, first few weeks | **Advisory.** Post the delta as a PR comment. |
    | Judge metrics, calibrated (kappa >= 0.6) and variance-measured | **Blocking with a margin** below the rolling baseline |
    | Cost and p95 latency per case | **Blocking on a ceiling**, advisory on the trend |
    
    Set the threshold *below* the baseline by more than the measured noise floor: if the suite
    scores 0.88 +/- 0.03 across reruns, gate at 0.80, not 0.87. You cannot know the noise floor
    from a single run - commit a rolling window of run results to git and read the variance off
    it. That committed history is also what distinguishes "today is noisy" from "today broke".
    
    ### Is the drop real?
    
    A noise floor tells you the aggregate moved unusually far. It does not tell you *the same
    cases* moved. Two runs over one frozen set are paired binary outcomes, and the tool for
    those is **McNemar's exact test** over the discordant pairs only.
    
    This matters because a change that breaks 8 cases and fixes 7 moves the headline score by
    0.01 - invisible against any noise floor - while silently swapping which 15 things work. The
    paired view names those 15 cases; a score comparison structurally cannot. Whether 8-vs-7 is
    *significant* is a separate question - it is not - but knowing which cases flipped is what
    lets you go and look.
    
    ```bash
    python3 scripts/eval-baseline.py evals/history.jsonl   --baseline-results base.jsonl --candidate-results new.jsonl
    # exit 10 = significant regression, and it names the cases that flipped
    ```
    
    Attribute **cost and latency per eval run** from the start. An eval suite is the only place
    you find out the accuracy win cost 4x the tokens, and retrofitting attribution after the
    harness exists is far more annoying than a `tokens_in` / `tokens_out` / `ms` field per case.
    Full CI shape and the noise-floor method: `references/regression-gating.md`.
    
    ## Hillclimbing
    
    Once the harness measures, the obvious move is to optimise against it. That works, and it
    is also the fastest way to make a good suite useless.
    
    **The loop belongs to [`iterate`](../iterate/) - this skill owns what goes wrong.** Every
    hillclimbing failure is a property of the metric, not of the loop:
    
    1. **Banking noise.** `iterate` keeps a change when the metric beats the previous best -
       correct for line coverage, a coin flip for an eval score. At 0.88 +/- 0.03, a measured
       0.90 is not evidence. Gate the keep decision on the noise floor instead:
    
       ```bash
       python3 scripts/eval-baseline.py history.jsonl --candidate iter.jsonl --accept
       # exit 0 = KEEP (a real improvement), 10 = DISCARD (noise or worse)
       ```
    
       `--accept` deliberately inverts the CI meaning of "noise": CI asks *did this get worse*,
       a hillclimb asks *is this improvement real*. Noise fails the second question.
    2. **Overfitting the frozen set.** Split train / validation / held-out before optimising,
       never show the validation set to whatever proposes changes, and treat held-out as a
       *budget* you spend at milestones - not a dashboard.
    3. **Keeping a champion instead of a frontier.** One aggregate best hides which cases a
       candidate won. Retaining candidates that are best on at least one case is what stops the
       loop walling itself into a local optimum.
    
    And the eval-design consequence: **a scalar score gives an optimizer nothing to reflect on.**
    A judge returning `{"reason": ..., "verdict": ...}` can be improved against; one returning
    `0.4` cannot. That field costs nothing today and is what makes automated optimization
    tractable later.
    
    Splits, the optimizer landscape (GEPA, MIPROv2, APE/ORPO/SPO), the pre-flight checklist,
    and the extraction trigger for a future `prompt-optimization-ops`: `references/hillclimbing.md`.
    
    ## Tooling
    
    Trace-level observability and eval scoring have converged into the same products - you are
    picking one system, not two. Open-source cores worth knowing: DeepEval (pytest-native),
    MLflow (tracing, eval and prompt versioning in one OSS platform), Opik, Langfuse, Arize
    Phoenix. Commercial-first: Braintrust (dataset curation for non-engineers), AgentOps,
    LangSmith, Arize.
    
    Honest default: **start with a JSONL file and a 40-line runner.** Adopt a platform when you
    need shared dataset curation, trace search across production traffic, or scheduled runs -
    not before. Which-one-when: `references/tooling-landscape.md`.
    
    > The landscape moves fast. Treat every version, price and feature claim in that reference
    > as needing re-verification; it carries a datestamp for exactly that reason.
    
    ## Scripts
    
    | Script | Use |
    |---|---|
    | `scripts/judge-calibration.py` | Judge-vs-human agreement: Cohen kappa, confusion matrix, per-class breakdown, verbosity/position bias probes. Exit 10 = below `--min-kappa`. |
    | `scripts/goldenset-audit.py` | Golden-set health: duplicates, bucket balance, staleness, freeze-manifest drift. Exit 10 = findings. |
    | `scripts/eval-baseline.py` | Noise floor from run history, the threshold your gate should use, and McNemar's exact test naming the cases that flipped. Exit 10 = confirmed regression or a cost/latency ceiling breach; `--accept` turns it into a hillclimb keep/discard gate. |
    
    All three accept `--help` and `--json`, and are offline and stdlib-only.
    
    ```bash
    python3 scripts/judge-calibration.py labels.jsonl --json | jq '.data.kappa'
    python3 scripts/goldenset-audit.py golden.jsonl --freeze manifest.json
    python3 scripts/eval-baseline.py history.jsonl --json | jq '.data.recommended_threshold'
    ```
    
    ## Assets
    
    Copy-and-adapt starting points, so the first hour goes on deciding what to measure rather
    than on scaffolding:
    
    | Asset | What it is |
    |---|---|
    | `assets/golden-set.example.jsonl` | 10 worked cases across all four buckets, with `why` and `criteria` filled in |
    | `assets/eval-runner.template.py` | The 40-line runner this skill tells you to start with - two ADAPT blocks, cost/latency/pass^k pre-wired |
    | `assets/judge-rubric.template.md` | One-criterion rubric with the bias-counter instructions and the calibration checklist |
    | `assets/eval-gate.template.yml` | GitHub Actions workflow encoding the tier ladder: deterministic blocks on push, judge advisory on PR, k=3 nightly |
    
    ## References
    
    - `references/eval-taxonomy.md` - outcome/step/trajectory, lucky pass, pass@k vs pass^k, metric selection
    - `references/golden-datasets.md` - building, four-bucket composition, sizing, freeze discipline, rot
    - `references/llm-judge.md` - bias catalog and mitigations, rubric design, panels, calibration method
    - `references/adversarial-verification.md` - refute-not-confirm, majority thresholds, lens diversity
    - `references/retrieval-eval.md` - RAG: recall@k, the retrieval-vs-generation split, ground truth, failure classes
    - `references/annotation-workflow.md` - where human labels come from: the human-human ceiling, sampling, adjudication, drift
    - `references/regression-gating.md` - blocking vs advisory, noise floor, McNemar, CI shape, cost/latency attribution
    - `references/hillclimbing.md` - optimising against an eval without destroying it: noise, overfitting, frontiers, optimizers
    - `references/tooling-landscape.md` - platform comparison with verification datestamps
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related