{"slug":"evals-ops","title":"evals-ops","summary":"Build and run evals for LLM and agent systems: golden datasets, LLM-as-a-judge with bias control, trajectory/step/outcome scoring, adversarial refuters, and CI regression gates. Triggers on: eval, evals, eval harness, golden dataset, golden set, LLM-as-a-judge, judge rubric, judg","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-30T19:36:44.082589Z","repo":{"url":"https://github.com/0xDarkMatter/claude-mods","stars":43,"forks":7,"license":"MIT","updatedAt":"2026-09-30T15:18:48Z"},"bodyHtml":"<hr>\n<h2>name: evals-ops\ndescription: \"Build and run evals for LLM and agent systems: golden datasets, LLM-as-a-judge with bias control, trajectory/step/outcome scoring, adversarial refuters, and CI regression gates. Triggers on: eval, evals, eval harness, golden dataset, golden set, LLM-as-a-judge, judge rubric, judge bias, judge calibration, Cohen kappa, agreement with human labels, pass@k, pass^k, trajectory eval, step-level eval, tool-call accuracy, regression gate, eval CI, blocking vs advisory eval, faithfulness score, DeepEval, Braintrust, Opik, Langfuse, AgentOps, did my prompt change make it worse, is my agent getting better.\"\nlicense: MIT\nmetadata:\nauthor: claude-mods\nrelated-skills: \"testing-ops, claude-api-ops, iterate, loop-ops, fleet-ops\"</h2>\n<h1>Evals Ops</h1>\n<p><strong>Evals are the prerequisite, not the polish.</strong> You cannot tune a prompt, a retriever, a\ncompaction strategy or a memory layer without a harness that says whether the change made\nthings better. Teams that skip this ship vibes and learn about regressions from users.</p>\n<p>This skill is the operational layer: what to measure, how to build the dataset, how to\nmake a judge trustworthy, and how to gate CI on it without teaching everyone to ignore red.</p>\n<h2>Route first</h2>\n<table>\n<thead>\n<tr>\n<th>The ask</th>\n<th>Go to</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>\"What should I even measure?\"</td>\n<td><a href=\"#three-levels-of-agent-eval\">Three levels</a> → <code>references/eval-taxonomy.md</code></td>\n</tr>\n<tr>\n<td>\"Where do the test cases come from?\"</td>\n<td><a href=\"#the-golden-set\">Golden set</a> → <code>references/golden-datasets.md</code></td>\n</tr>\n<tr>\n<td>\"My judge disagrees with me / is it any good?\"</td>\n<td><a href=\"#llm-as-a-judge\">Judges</a> → <code>references/llm-judge.md</code></td>\n</tr>\n<tr>\n<td>\"Verify a finding is real, not plausible\"</td>\n<td><a href=\"#adversarial-verification\">Refuters</a> → <code>references/adversarial-verification.md</code></td>\n</tr>\n<tr>\n<td>\"Is my RAG retrieval any good?\"</td>\n<td><a href=\"#retrieval\">Retrieval</a> → <code>references/retrieval-eval.md</code></td>\n</tr>\n<tr>\n<td>\"Where do the human labels come from?\"</td>\n<td><code>references/annotation-workflow.md</code></td>\n</tr>\n<tr>\n<td>\"Should this block the merge?\"</td>\n<td><a href=\"#regression-gating\">Gating</a> → <code>references/regression-gating.md</code></td>\n</tr>\n<tr>\n<td>\"Did this change really make it worse?\"</td>\n<td><a href=\"#is-the-drop-real\">Is the drop real</a></td>\n</tr>\n<tr>\n<td>\"Optimise against the eval / run it overnight\"</td>\n<td><a href=\"#hillclimbing\">Hillclimbing</a> → <code>references/hillclimbing.md</code></td>\n</tr>\n<tr>\n<td>\"Which platform should we use?\"</td>\n<td><code>references/tooling-landscape.md</code></td>\n</tr>\n<tr>\n<td>\"Just give me a starting file\"</td>\n<td><a href=\"#assets\">Assets</a> — golden set, rubric, runner, CI gate</td>\n</tr>\n</tbody>\n</table>\n<h2>The 60-second version</h2>\n<ol>\n<li><strong>Write 20 cases before you write a metric.</strong> A dataset you can eyeball beats a metric\nyou cannot interpret. Grow to 100-300, then <em>freeze</em> it.</li>\n<li><strong>Prefer a deterministic assertion to any judge.</strong> The JSON parsed, the tool was called\nwith the right argument, the query returned 3 rows - free, instant, zero variance.\nReach for a judge only where correctness is genuinely a matter of language.</li>\n<li><strong>Score the trajectory, not just the answer.</strong> Most teams only check the final artifact\nand are surprised when a right answer came from a wrong path that breaks tomorrow.</li>\n<li><strong>Calibrate the judge against humans before trusting it.</strong> Cohen kappa, not raw\nagreement. <code>scripts/judge-calibration.py</code> does the arithmetic and the verdict.</li>\n<li><strong>Blocking gates must be deterministic. Judge metrics start advisory.</strong> One flaky red\npermanently devalues the signal.</li>\n</ol>\n<h2>Three levels of agent eval</h2>\n<p>Most teams do only the third, then wonder why quality is unpredictable.</p>\n<table>\n<thead>\n<tr>\n<th>Level</th>\n<th>Question</th>\n<th>Signal</th>\n<th>Typical evaluator</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Outcome</strong></td>\n<td>Is the final artifact correct?</td>\n<td>Binary or scored end state</td>\n<td>Deterministic assertion, unit test, judge</td>\n</tr>\n<tr>\n<td><strong>Step</strong></td>\n<td>Was <em>this</em> tool call right?</td>\n<td>Per-span: tool choice, arg shape, arg values</td>\n<td>Schema/argument assertion, span-level judge</td>\n</tr>\n<tr>\n<td><strong>Trajectory</strong></td>\n<td>Was the <em>path</em> sensible?</td>\n<td>Sequence, loops, redundancy, cost</td>\n<td>Reference-trajectory match, rubric judge</td>\n</tr>\n</tbody>\n</table>\n<p>The failure that motivates all three: an agent reaches the right end state by an accidental\nroute - the <em>lucky pass</em>. Outcome-only scoring records that as a win, and the same case\nfails next week when the accident does not recur. Conversely a trajectory-only score\npunishes a legitimately novel-but-correct path. <strong>Gate on outcome; keep step and trajectory\nas the diagnostics that tell you why the gate moved.</strong></p>\n<p>For multi-turn or stateful agents also report <strong>pass^k</strong> (all k independent runs of the same\ncase succeed) alongside <strong>pass@k</strong> (any of k succeeded). pass@k flatters a non-deterministic\nagent; pass^k is the number that predicts production. Full treatment:\n<code>references/eval-taxonomy.md</code>.</p>\n<h2>The golden set</h2>\n<p>A golden set is a <strong>reviewed, versioned, deliberately frozen</strong> collection of inputs with\ntrusted expected outputs. Frozen matters: a set that grows every sprint cannot tell you\nwhether last week's number moved because the system changed or because the set did.</p>\n<p><strong>Composition - four buckets, not one:</strong></p>\n<table>\n<thead>\n<tr>\n<th>Bucket</th>\n<th>Source</th>\n<th>Why</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Production sample</td>\n<td>Real traffic, stratified</td>\n<td>Keeps the score connected to what users actually do</td>\n</tr>\n<tr>\n<td>Failure replays</td>\n<td>Every incident that reached a human</td>\n<td>Regression protection; the easiest cases to justify</td>\n</tr>\n<tr>\n<td>Adversarial</td>\n<td>Injections, contradictions, refusal-bait</td>\n<td>The class both agents and judges fail silently on</td>\n</tr>\n<tr>\n<td>Edge cases</td>\n<td>Empty, huge, ambiguous, multilingual</td>\n<td>Where deterministic code breaks first</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Sizing:</strong> 20 to start, 100-300 for a working regression set, 200-500 once you have\nproduction traffic to sample. Beyond that you are usually buying latency, not signal - add\ncases when a new <em>failure class</em> appears, and record in the case itself why it exists.</p>\n<p><code>scripts/goldenset-audit.py</code> checks a set for the rot that accumulates: duplicates, a\nbucket that quietly became 90% of the set, undated cases, and drift from a frozen manifest.</p>\n<pre><code>python3 scripts/goldenset-audit.py evals/golden.jsonl --json | jq '.data.findings[]'\n</code></pre>\n<p>Depth: <code>references/golden-datasets.md</code>.</p>\n<h2>LLM-as-a-judge</h2>\n<p>A judge is a measurement instrument. Instruments need calibration, and this one has\ndocumented, reproducible biases:</p>\n<table>\n<thead>\n<tr>\n<th>Bias</th>\n<th>What it does</th>\n<th>Mitigation</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Position</strong></td>\n<td>Prefers whichever candidate was shown first</td>\n<td>Run both orders and average; or score absolutely, not pairwise</td>\n</tr>\n<tr>\n<td><strong>Verbosity</strong></td>\n<td>Rates longer answers higher regardless of quality</td>\n<td>Separate correctness from style in the rubric; penalise unsupported length</td>\n</tr>\n<tr>\n<td><strong>Self-preference</strong></td>\n<td>Rates its own model family's output higher</td>\n<td>Judge with a different family than the one under test</td>\n</tr>\n<tr>\n<td><strong>Scale drift</strong></td>\n<td>1-5 scores cluster and shift between model versions</td>\n<td>Binary pass/fail against explicit criteria; pin the judge model version</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Panel vs N-identical.</strong> Three calls to the same judge with the same rubric mostly buys the\nsame bias three times. A <strong>panel with distinct lenses</strong> - one asks \"is this supported by the\nsource?\", one \"does it follow the stated policy?\", one \"would this reproduce?\" - finds\nfailure modes redundancy structurally cannot. Use N-identical only to measure the judge's\nown variance, which is a different question worth asking once.</p>\n<p><strong>When a judge is the wrong tool:</strong> if you can express the criterion as code, do. A judge\ncosts money, adds latency, drifts across model versions, and has variance a regex does not.\nJudges earn their place on faithfulness, tone, policy compliance, and \"is this a reasonable\nanswer to an open question\" - nowhere else.</p>\n<p><strong>Calibrate before you trust.</strong> Label 50-200 cases by hand, run the judge on the same cases,\nand compute Cohen kappa (raw agreement lies when classes are imbalanced):</p>\n<pre><code>python3 scripts/judge-calibration.py evals/labels.jsonl --min-kappa 0.6\n# exit 0 = calibrated;  exit 10 = below threshold, fix the rubric before shipping it\n</code></pre>\n<p>kappa &gt;= 0.8 production-ready - 0.6-0.8 substantial, usable with care - below 0.6 the rubric\nis the problem, not the model. Re-sample ~50 fresh cases periodically; judges drift when the\nunderlying model version moves. Depth, including bias-probe design: <code>references/llm-judge.md</code>.</p>\n<p><strong>Measure the human-human ceiling first.</strong> A judge cannot beat the agreement two people\nachieve with each other. Two annotators on 30-50 shared cases gives you that number - and\nif it is below ~0.6, the rubric is ambiguous and every label you produce against it is\nwasted. How to run the sessions, stratify the sample, and adjudicate disagreements:\n<code>references/annotation-workflow.md</code>.</p>\n<h2>Adversarial verification</h2>\n<p>For findings rather than scores - bug reports, audit results, review comments - flip the\nprompt: <strong>ask the verifier to REFUTE, not to confirm.</strong> \"Try to refute this finding; default\nto refuted if uncertain\" kills plausible-but-wrong results that an \"is this correct?\" prompt\nwaves through, because agreement is the path of least resistance for a model.</p>\n<p>Then take a majority: run 3 refuters, keep the finding only if at least 2 fail to refute it.\nPrefer <strong>perspective-diverse</strong> refuters (correctness / security / does-it-actually-reproduce)\nover three identical skeptics - same reasoning as judge panels.</p>\n<p>This composes with the parallel-work skills rather than duplicating them: <code>fleet-ops</code> and\n<code>parallel-ops</code> own the fan-out mechanics; this skill owns the scoring contract the refuters\nreturn. See <code>references/adversarial-verification.md</code>.</p>\n<h2>Retrieval</h2>\n<p>RAG is the most common eval target and the most commonly mis-measured. Scoring only the\nfinal answer averages two independent failures into one uninterpretable number:</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>Right context</th>\n<th>Wrong context</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Answer correct</strong></td>\n<td>Working</td>\n<td><strong>Lucky</strong> - the model knew it anyway; scores as a pass</td>\n</tr>\n<tr>\n<td><strong>Answer wrong</strong></td>\n<td><strong>Generation bug</strong> - chunking, prompt, model</td>\n<td><strong>Retrieval bug</strong> - embeddings, index, query rewriting</td>\n</tr>\n</tbody>\n</table>\n<p>Record the retrieved chunk ids next to every answer and that opaque score becomes a 2x2\nyou can assign to a team. Gate on <strong>recall@k</strong> (a precision failure degrades an answer; a\nrecall failure makes a correct one impossible) and on citation-id validity, which is free\nand catches confident answers attached to unrelated sources. Retrieval is the one place\ndeterministic scoring genuinely dominates - you have ground-truth ids, so skip the judge.</p>\n<p>The bucket almost everyone omits: <strong>questions the corpus cannot answer.</strong> Without them the\nsuite cannot detect hallucination under retrieval failure, which is what users hit most.\nMetrics, the six failure classes, and ground-truth construction: <code>references/retrieval-eval.md</code>.</p>\n<h2>Regression gating</h2>\n<p>The rule that keeps a gate alive: <strong>a blocking check must never be flaky.</strong></p>\n<table>\n<thead>\n<tr>\n<th>Check</th>\n<th>Gate</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Deterministic assertions (schema, tool-call, exact match)</td>\n<td><strong>Blocking.</strong> Any failure fails CI.</td>\n</tr>\n<tr>\n<td>Judge metrics, first few weeks</td>\n<td><strong>Advisory.</strong> Post the delta as a PR comment.</td>\n</tr>\n<tr>\n<td>Judge metrics, calibrated (kappa &gt;= 0.6) and variance-measured</td>\n<td><strong>Blocking with a margin</strong> below the rolling baseline</td>\n</tr>\n<tr>\n<td>Cost and p95 latency per case</td>\n<td><strong>Blocking on a ceiling</strong>, advisory on the trend</td>\n</tr>\n</tbody>\n</table>\n<p>Set the threshold <em>below</em> the baseline by more than the measured noise floor: if the suite\nscores 0.88 +/- 0.03 across reruns, gate at 0.80, not 0.87. You cannot know the noise floor\nfrom a single run - commit a rolling window of run results to git and read the variance off\nit. That committed history is also what distinguishes \"today is noisy\" from \"today broke\".</p>\n<h3>Is the drop real?</h3>\n<p>A noise floor tells you the aggregate moved unusually far. It does not tell you <em>the same\ncases</em> moved. Two runs over one frozen set are paired binary outcomes, and the tool for\nthose is <strong>McNemar's exact test</strong> over the discordant pairs only.</p>\n<p>This matters because a change that breaks 8 cases and fixes 7 moves the headline score by\n0.01 - invisible against any noise floor - while silently swapping which 15 things work. The\npaired view names those 15 cases; a score comparison structurally cannot. Whether 8-vs-7 is\n<em>significant</em> is a separate question - it is not - but knowing which cases flipped is what\nlets you go and look.</p>\n<pre><code>python3 scripts/eval-baseline.py evals/history.jsonl   --baseline-results base.jsonl --candidate-results new.jsonl\n# exit 10 = significant regression, and it names the cases that flipped\n</code></pre>\n<p>Attribute <strong>cost and latency per eval run</strong> from the start. An eval suite is the only place\nyou find out the accuracy win cost 4x the tokens, and retrofitting attribution after the\nharness exists is far more annoying than a <code>tokens_in</code> / <code>tokens_out</code> / <code>ms</code> field per case.\nFull CI shape and the noise-floor method: <code>references/regression-gating.md</code>.</p>\n<h2>Hillclimbing</h2>\n<p>Once the harness measures, the obvious move is to optimise against it. That works, and it\nis also the fastest way to make a good suite useless.</p>\n<p><strong>The loop belongs to <a href=\"../iterate/\"><code>iterate</code></a> - this skill owns what goes wrong.</strong> Every\nhillclimbing failure is a property of the metric, not of the loop:</p>\n<ol>\n<li><p><strong>Banking noise.</strong> <code>iterate</code> keeps a change when the metric beats the previous best -\ncorrect for line coverage, a coin flip for an eval score. At 0.88 +/- 0.03, a measured\n0.90 is not evidence. Gate the keep decision on the noise floor instead:</p>\n<pre><code>python3 scripts/eval-baseline.py history.jsonl --candidate iter.jsonl --accept\n# exit 0 = KEEP (a real improvement), 10 = DISCARD (noise or worse)\n</code></pre>\n<p><code>--accept</code> deliberately inverts the CI meaning of \"noise\": CI asks <em>did this get worse</em>,\na hillclimb asks <em>is this improvement real</em>. Noise fails the second question.</p>\n</li>\n<li><p><strong>Overfitting the frozen set.</strong> Split train / validation / held-out before optimising,\nnever show the validation set to whatever proposes changes, and treat held-out as a\n<em>budget</em> you spend at milestones - not a dashboard.</p>\n</li>\n<li><p><strong>Keeping a champion instead of a frontier.</strong> One aggregate best hides which cases a\ncandidate won. Retaining candidates that are best on at least one case is what stops the\nloop walling itself into a local optimum.</p>\n</li>\n</ol>\n<p>And the eval-design consequence: <strong>a scalar score gives an optimizer nothing to reflect on.</strong>\nA judge returning <code>{\"reason\": ..., \"verdict\": ...}</code> can be improved against; one returning\n<code>0.4</code> cannot. That field costs nothing today and is what makes automated optimization\ntractable later.</p>\n<p>Splits, the optimizer landscape (GEPA, MIPROv2, APE/ORPO/SPO), the pre-flight checklist,\nand the extraction trigger for a future <code>prompt-optimization-ops</code>: <code>references/hillclimbing.md</code>.</p>\n<h2>Tooling</h2>\n<p>Trace-level observability and eval scoring have converged into the same products - you are\npicking one system, not two. Open-source cores worth knowing: DeepEval (pytest-native),\nMLflow (tracing, eval and prompt versioning in one OSS platform), Opik, Langfuse, Arize\nPhoenix. Commercial-first: Braintrust (dataset curation for non-engineers), AgentOps,\nLangSmith, Arize.</p>\n<p>Honest default: <strong>start with a JSONL file and a 40-line runner.</strong> Adopt a platform when you\nneed shared dataset curation, trace search across production traffic, or scheduled runs -\nnot before. Which-one-when: <code>references/tooling-landscape.md</code>.</p>\n<blockquote>\n<p>The landscape moves fast. Treat every version, price and feature claim in that reference\nas needing re-verification; it carries a datestamp for exactly that reason.</p>\n</blockquote>\n<h2>Scripts</h2>\n<table>\n<thead>\n<tr>\n<th>Script</th>\n<th>Use</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>scripts/judge-calibration.py</code></td>\n<td>Judge-vs-human agreement: Cohen kappa, confusion matrix, per-class breakdown, verbosity/position bias probes. Exit 10 = below <code>--min-kappa</code>.</td>\n</tr>\n<tr>\n<td><code>scripts/goldenset-audit.py</code></td>\n<td>Golden-set health: duplicates, bucket balance, staleness, freeze-manifest drift. Exit 10 = findings.</td>\n</tr>\n<tr>\n<td><code>scripts/eval-baseline.py</code></td>\n<td>Noise floor from run history, the threshold your gate should use, and McNemar's exact test naming the cases that flipped. Exit 10 = confirmed regression or a cost/latency ceiling breach; <code>--accept</code> turns it into a hillclimb keep/discard gate.</td>\n</tr>\n</tbody>\n</table>\n<p>All three accept <code>--help</code> and <code>--json</code>, and are offline and stdlib-only.</p>\n<pre><code>python3 scripts/judge-calibration.py labels.jsonl --json | jq '.data.kappa'\npython3 scripts/goldenset-audit.py golden.jsonl --freeze manifest.json\npython3 scripts/eval-baseline.py history.jsonl --json | jq '.data.recommended_threshold'\n</code></pre>\n<h2>Assets</h2>\n<p>Copy-and-adapt starting points, so the first hour goes on deciding what to measure rather\nthan on scaffolding:</p>\n<table>\n<thead>\n<tr>\n<th>Asset</th>\n<th>What it is</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>assets/golden-set.example.jsonl</code></td>\n<td>10 worked cases across all four buckets, with <code>why</code> and <code>criteria</code> filled in</td>\n</tr>\n<tr>\n<td><code>assets/eval-runner.template.py</code></td>\n<td>The 40-line runner this skill tells you to start with - two ADAPT blocks, cost/latency/pass^k pre-wired</td>\n</tr>\n<tr>\n<td><code>assets/judge-rubric.template.md</code></td>\n<td>One-criterion rubric with the bias-counter instructions and the calibration checklist</td>\n</tr>\n<tr>\n<td><code>assets/eval-gate.template.yml</code></td>\n<td>GitHub Actions workflow encoding the tier ladder: deterministic blocks on push, judge advisory on PR, k=3 nightly</td>\n</tr>\n</tbody>\n</table>\n<h2>References</h2>\n<ul>\n<li><code>references/eval-taxonomy.md</code> - outcome/step/trajectory, lucky pass, pass@k vs pass^k, metric selection</li>\n<li><code>references/golden-datasets.md</code> - building, four-bucket composition, sizing, freeze discipline, rot</li>\n<li><code>references/llm-judge.md</code> - bias catalog and mitigations, rubric design, panels, calibration method</li>\n<li><code>references/adversarial-verification.md</code> - refute-not-confirm, majority thresholds, lens diversity</li>\n<li><code>references/retrieval-eval.md</code> - RAG: recall@k, the retrieval-vs-generation split, ground truth, failure classes</li>\n<li><code>references/annotation-workflow.md</code> - where human labels come from: the human-human ceiling, sampling, adjudication, drift</li>\n<li><code>references/regression-gating.md</code> - blocking vs advisory, noise floor, McNemar, CI shape, cost/latency attribution</li>\n<li><code>references/hillclimbing.md</code> - optimising against an eval without destroying it: noise, overfitting, frontiers, optimizers</li>\n<li><code>references/tooling-landscape.md</code> - platform comparison with verification datestamps</li>\n</ul>\n","files":[{"path":"assets/eval-gate.template.yml","sizeBytes":6439,"isText":true},{"path":"assets/eval-runner.template.py","sizeBytes":6448,"isText":true},{"path":"assets/golden-set.example.jsonl","sizeBytes":6065,"isText":false},{"path":"assets/judge-rubric.template.md","sizeBytes":3522,"isText":true},{"path":"references/adversarial-verification.md","sizeBytes":5370,"isText":true},{"path":"references/annotation-workflow.md","sizeBytes":6178,"isText":true},{"path":"references/eval-taxonomy.md","sizeBytes":5543,"isText":true},{"path":"references/golden-datasets.md","sizeBytes":6993,"isText":true},{"path":"references/hillclimbing.md","sizeBytes":10013,"isText":true},{"path":"references/llm-judge.md","sizeBytes":7902,"isText":true},{"path":"references/regression-gating.md","sizeBytes":8595,"isText":true},{"path":"references/retrieval-eval.md","sizeBytes":6505,"isText":true},{"path":"references/tooling-landscape.md","sizeBytes":5101,"isText":true},{"path":"scripts/eval-baseline.py","sizeBytes":20795,"isText":true},{"path":"scripts/goldenset-audit.py","sizeBytes":16914,"isText":true},{"path":"scripts/judge-calibration.py","sizeBytes":13461,"isText":true},{"path":"SKILL.md","sizeBytes":17838,"isText":true},{"path":"tests/run.sh","sizeBytes":35901,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-30T19:37:29.68123Z","sha256":"CD8E76C9325CD2075B1CFA81CFA868EA29276E8C89830CD68AA220C854E9412E","sizeBytes":76507},"review":null,"source":{"repositoryUrl":"https://github.com/0xDarkMatter/claude-mods","path":"skills/evals-ops","license":"MIT","commit":"3dfaf0ba5753026a99ee13f9d9ed56b9793bb6e8","subtreeSha":"635D372F827051E63FA70D34C494E89D7CB0E7F4DD5EF0ABC6BFB3C9C1FEDFAE","lastSyncedAt":"2026-09-30T19:37:28.226022Z"},"reviewedAt":"2026-09-30T19:38:38.132618Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/evals-ops"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart"},{"target":"git","command":"git clone https://github.com/0xDarkMatter/claude-mods.git"}]}