Claude Skill

gating

Build and audit deterministic verification gates — a check that blocks a pipeline and can be shown to go red. Use when writing a calibration gate, CI check, validation script or pre-publication check for a numeric or empirical result; when a plausible-but-wrong value would surviv

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download oaustegard-claude-skills-plugins_ai-and-reasoning_skills_gating-e39c726.zip · 21 KB
Part of oaustegard/claude-skills — 39 skills

Install

skills CLI npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/ai-and-reasoning/skills/gating
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
Git git clone https://github.com/oaustegard/claude-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.

README

gating

Build and audit verification gates — deterministic checks that block a pipeline and can be shown to go red.

The characteristic failure of a gate is not a wrong check. A wrong check gets noticed. It is a check that cannot fail, which reports PASS forever and is indistinguishable from a working one from the outside.

So the object under suspicion here is the check itself.

The three obligations

Anchor something your own code did not produce published constant, closed form, invariant, theoretical bound
Known-bad a case the gate demonstrably rejects break the subject plausibly, run the gate, confirm red
Coverage limit what it cannot catch, in writing holes are invisible from inside a green run

scripts/gate.py enforces the last two: a gate that registers no known-bad or no coverage limit reports INCONCLUSIVE and exits 2, not PASS.

Scripts

# mutation pass — which single-token changes does the gate NOT notice?
python3 scripts/mutate.py --target src/codec.py -- python3 calibrate.py

mutate.py is stdlib-only, uses tokenize (so it never corrupts strings or comments), restores the file even on interrupt, and refuses to run when the gate is already red. It works against any gate command; mutmut and cosmic-ray go deeper once it stops finding survivors.

gate.py is a ~150-line harness: anchor(), bracket(), known_bad(), coverage(), note(), and a report() that returns a process exit code.

Relationship to challenging

challenging asks would a careful reader object? — LLM judgement against a persona, verdicts of SHIP/REVISE/RETHINK. gating asks can this check go red? — deterministic, anchored, exit-coded.

Complements with different blind spots: an adversarial reviewer will not recompute your constants, and a gate will not notice that your framing is wrong. Note also that a same-model reviewer is an independent context, not an independent reviewer; where shared priors are the risk, an anchor beats a reviewer because an anchor is not negotiable.

See SKILL.md for the build and audit procedures and the anti-pattern table, references/anchors.md for choosing an oracle, and references/auditing.md for the six-pass "can this fail?" sweep over an existing suite.

Skill manifest

gating

A gate is a check that blocks. Its only job is to go red when it should.

The characteristic failure is not a wrong check — a wrong check gets noticed. It is a check that cannot fail, which reports PASS forever and is indistinguishable from a working one from the outside. That is what makes this different from ordinary testing: the object under suspicion is the check.

When NOT to use this skill

Scope is ONE check and whether it can be made to fail.

Situation Use
Sequence several steps with branches and retries flowing
Run the repo's existing suite run it
Decide what to test at all this skill has no opinion; that is design

A gate is a thing that goes red. If nothing here can go red, there is no gate to audit.

The three obligations

Every gate owes these. A gate missing any of them is not yet a gate.

1. An anchor outside your own code. Something the check compares against that your implementation did not produce: a published constant, a closed-form answer, a conservation law, a degenerate case with a known result, an independent implementation. A check that compares this run to the last run only ever tells you the code still does what it did. See references/anchors.md.

2. A known-bad it demonstrably rejects. Break the subject the way it would plausibly break, run the gate, confirm red. Until you have done this you have not shown the gate works — you have shown it runs. This is the obligation people skip, because a passing gate feels like evidence.

Two things about known-bads that are easy to get wrong:

  • Validate it at the configuration it will run in. A case tuned on a small or fast setting can stop being bad at full size. An "untrained" grid built from one Lloyd iteration was genuinely zero-gain at m=2/K=16 and earned a real +0.10 dB at m=8/K=65536, where one iteration relocates ~63,000 empty cells toward the mode. It passed the fast gate and certified nothing about the real one. This matters more than it sounds, because mutate.py needs a fast gate variant and it is tempting to validate everything there.
  • Measure its reach. One known-bad is the floor, not the goal. Name which checks it exercises (known_bad(..., covers=(...))); the harness prints the checks no known-bad reaches. An audited gate had a single known-bad covering 1 of 8 checks — and the check its whole result rested on accepted the same bad case.

3. A written statement of what it cannot catch. Coverage holes are invisible from inside a green run: the gate is silent about the thing it does not look at, in exactly the same tone it uses for the thing it looked at and approved. The author has to assert the hole; nothing else will.

scripts/gate.py enforces obligations 2 and 3 mechanically — it returns exit code 2 (INCONCLUSIVE, not PASS) when a gate registers no known-bad or no coverage limit.

Building a gate

Work in this order. The first step is the one that determines whether the rest is worth anything.

Name the wrong conclusion, not the component. Not "check the quantizer is correct" but "prevent shipping scalar wins at high bit rates when that would really be an optimizer artifact." A gate aimed at a conclusion knows what counts as a near-miss; a gate aimed at a component just exercises the code.

Find an anchor. references/anchors.md lists the kinds, in rough order of strength, with the questions that find each one.

Prefer brackets to point checks. Assert a value lies strictly between two things it cannot legitimately pass: better than a baseline, worse than a theoretical bound. A one-sided check passes for a result that collapsed as readily as for one that is right — which is how an implementation that silently does nothing gets certified.

Derive the tolerance from measured noise. Run the thing several times, see how much it moves, put the threshold outside that. A tolerance picked for comfort tends to land wider than the defect you are trying to catch, and then it swallows it.

Then check it is not too tight to mean anything. The opposite failure is real and less obvious: a margin can be statistically impeccable and practically empty. A paired estimator — scoring both arms on one shared sample so the common fluctuation cancels — is the right way to measure a difference, and its standard error shrinks as the two arms converge. So "beats the baseline by 3 se" degenerates: a codebook perturbed by N(0, 1e-3) gained +0.0001 dB against a 3-se margin of 1.2e-06 and was accepted, while real ones gained 0.35–1.41 dB. The check certified the effect is real, not the effect is worth having. Those are different assertions and need different thresholds — and the second one has to come from an anchor (there, a published lattice codebook), never from the estimator, which knows nothing about what magnitude would matter.

Build the known-bad and confirm red. Then run scripts/mutate.py for the failures you did not think of.

Wire it to a non-zero exit and run it before the thing it gates, not after. A gate that runs after the results are written is a report.

Auditing an existing check suite

Given tests, a linter config, a CI job, or a gate someone already wrote, the question is not "do these pass" but "can these fail". Full procedure in references/auditing.md; the fast version:

  • For each assertion, name a concrete input that makes it fail. If you cannot, it is decoration — delete it or fix it.
  • Check each oracle's range against the range you actually operate in. A published table that stops short of your regime is a hole with a green light on it.
  • Run scripts/mutate.py against the code the suite covers. Every survivor is a behaviour nothing checks.
  • Look for assertions whose truth does not depend on the subject at all.
  • Look for tolerances wider than the effect being measured.

Scripts

# harness: refuses to report PASS without a known-bad and a coverage limit
python3 scripts/gate.py           # importable; see the module docstring

# mutation pass: which single-token changes does the gate NOT notice?
python3 scripts/mutate.py --target src/codec.py -- python3 calibrate.py
python3 scripts/mutate.py --target grids.py --max 40 -- pytest -q

mutate.py requires the gate to pass on unmutated code first and refuses to run otherwise, because survivor counts against an already-red gate mean nothing. It restores the file even on interrupt, and uses tokenize so string literals and comments are never corrupted. It is the zero-dependency pass that works against any gate command; once it stops finding survivors, mutmut or cosmic-ray go deeper on Python test suites specifically.

Anti-patterns

Each of these has shipped a wrong result somewhere. They are ordered by how convincingly they impersonate a working gate.

Anti-pattern Why it survives review
Slack wider than the defect A tolerance chosen for comfort. The gate passes the real thing and the broken thing, and reports PASS for both. Derive the threshold from noise, then confirm the known-bad falls outside it.
An oracle with a coverage hole Published anchors end somewhere. If the defect is past the end of the table, the check is structurally incapable of catching it and looks fine. State the range the anchor covers.
An assertion whose truth doesn't depend on the subject "The output has at least N distinct colours" is equally true of an unchanged frame. Prefer differential checks: the state must change when it should, and a toggle applied twice must return to the byte-identical original.
Confirming the check ran, not that it can fail "Invoke it and confirm the step appears in the output" catches a check that was never wired up. It says nothing about a check that is wired up and toothless.
Comparing against your own previous output Regenerated goldens ratify drift. If the golden came from the code under test, it is a changelog, not an oracle.
A cache keyed on the problem rather than the method cache[(m, K)] cannot notice that the code producing the value changed. Version-stamp the artifact and delete on mismatch instead of trusting.
A self-matching predicate until ! pgrep -f trainer never exits, because the watching shell's own argv contains trainer. Worse, a malformed variant exits immediately and reports the job finished while it runs. Wait on a PID.
A margin that is significant but not meaningful A threshold derived purely from estimator noise certifies that an effect is real, not that it is worth having — and a paired estimator's noise shrinks as the arms converge, so the margin can approach zero. Pair every noise-derived floor with a magnitude an anchor says would matter.
A strict bracket at an attainable optimum A theoretical bound is often reachable, and reaching it is the best possible outcome. A strict edge then goes red on a perfect result and blocks real work. Ask of each edge whether the subject can legitimately sit exactly there.
A gate written by whatever produced the artifact Shared assumptions produce shared blind spots, and the convention both inherited is the one neither questions. Anchors are the defence, because an anchor is the one input the producer did not choose.

Division of labour

This skill is for results and pipelines where the failure mode is a plausible wrong number that would survive a careful read.

Use When the risk is
gating (this) A number or empirical result is about to be published or acted on, and a wrong-but-reasonable value would pass unnoticed. Output: a gate that blocks.
challenging An artifact would draw a specific objection from a skeptical reader — prose, analysis, a recommendation, a diff. LLM judgement against a persona. Output: findings and a SHIP/REVISE/RETHINK verdict.
verifying-claims Documentation says something about code that is no longer true. Output: prose-vs-code disagreements.
A test suite / TDD Code you wrote does not behave as specified. Output: red tests.

challenging asks would a careful reader object? gating asks can this check go red? They are complements and they miss different things: an adversarial reviewer will not recompute your constants, and a gate will not notice that your framing is wrong.

A gate checks correctness, and cannot check comparability. If two arms of a comparison are each individually correct but not comparably implemented, no anchor and no mutant will see it: nothing is broken, so nothing goes red. A published ablation reported one transform 11–24× slower than another and carried a caveat saying the number was implementation-bound; both arms passed every correctness check, and the real finding — one arm was a tuned BLAS call and the other an interpreted loop — was found by a human reviewer months of gate-work later. Related: performance claims have no anchor in this framework at all. There is no published constant for how fast your code should be. Wall-clock belongs to benchmarking discipline (matched implementation effort, min-of-trials, stated hardware), not to gating.

One caution about pairing them. A same-model reviewer is an independent context, not an independent reviewer — it shares your priors, so the convention you did not question is the one it will not question either. Where that matters, an anchor beats a reviewer, because an anchor is not negotiable.

References

  • references/anchors.md — kinds of oracle, strongest first, and how to find one when nothing published exists.
  • references/auditing.md — the full "can this fail?" pass over an existing suite, including how to read a mutation report.
Files (claude-skills)
  • references
    • anchors.md 7.1 KB
      # Anchors — where a gate gets its authority
      
      An anchor is the part of a check your implementation did not produce. It is
      what separates a gate from a changelog.
      
      Without one, the strongest thing a check can say is "this run agrees with the
      last run," which is true of a codebase that has been quietly wrong since the
      first commit.
      
      Ordered strongest first. Take the highest one available; they compose, and a
      gate with two anchors of different kinds is much harder to fool than one with
      two of the same kind.
      
      ---
      
      ## 1. A published constant
      
      Someone measured or proved it, in a paper or a standard, and printed the
      digits.
      
      **Questions that find one.** What field has been studying this for decades?
      What is the textbook version of my quantity? Does my problem have a named
      constant attached to it?
      
      **Watch the range.** Published tables stop somewhere. Max's 1960 quantizer
      table stops at 5 bits; a check built on it is structurally incapable of
      noticing an error at 8 bits, and will report PASS in exactly the tone it uses
      for the rates it covers. Record the covered range as a coverage limit.
      
      **Watch the precision.** A hand-computed 1960 table has a last digit that may
      be worse than your float64 fixed point. When your converged value disagrees
      with a published one by a fraction of a percent, the table is sometimes the
      one that is wrong — but say so with evidence (convergence residual, stability
      across iteration counts, a uniqueness argument), not by widening the tolerance
      until the test goes green.
      
      ---
      
      ## 2. A closed form
      
      An exact expression you can evaluate independently of the machinery under
      test — an integral, a recurrence, a combinatorial count.
      
      This is the strongest anchor available for anything with a tractable special
      case, because it does not merely constrain the answer, it *is* the answer.
      
      **Pattern that works well:** find a degenerate configuration where a complex
      implementation must reduce to a simple formula. Lift a scalar quantizer's
      levels into an m-dimensional product grid, and nearest-neighbour assignment
      decomposes per coordinate — so a KD-tree measurement and a closed-form
      integral must agree to sampling error. That agreement exonerates the
      *instrument*, which converts "one of these two things is broken" into "this
      one thing is broken."
      
      Always exonerate the instrument before blaming the subject.
      
      ---
      
      ## 3. A conservation law or invariant
      
      Something that must hold regardless of the answer: a total that is preserved,
      a norm left unchanged by an orthogonal map, a round-trip that must return the
      input, an operation applied twice returning to the byte-identical original.
      
      Cheap, and unusually good at catching the class of bug that produces
      plausible-looking output. An inverse-transform round-trip that agrees to 1e-7
      rules out a large family of indexing and sign errors in one line.
      
      **Also invariant-shaped:** monotonicity that theory requires (quality improves
      with more bits, error shrinks with more samples), and orderings that must hold
      between arms.
      
      ---
      
      ## 4. A theoretical bound
      
      A quantity that cannot be exceeded — an information-theoretic limit, a
      complexity lower bound, a physical constraint.
      
      Bounds are one-sided, so they are best used as the far side of a bracket:
      better than the baseline, *and* not better than the bound. A result that beats
      its own theoretical limit is not a triumph, it is a bug, and this is the check
      that catches it.
      
      Bonus signal: watch the ratio-to-bound as a parameter sweeps. If theory says
      it should approach a constant, a smooth monotone approach is corroboration
      that many independent parts are right at once.
      
      ---
      
      ## 5. An independent implementation
      
      A second implementation of the same computation — a reference library, a slow
      brute-force version, a different language.
      
      **Weaker than it looks.** Two implementations agree on a false result whenever
      they share a modelling assumption, and if the same author or the same model
      wrote both, they share more than they appear to. Vary the assumption and the
      author, not just the code. A deliberately naive O(n²) brute force is usually a
      better second implementation than a clever one, because it shares less.
      
      ---
      
      ## 6. A known-answer fixture
      
      An input whose output you know for reasons outside the code: a worked example
      from a textbook, a case computed by hand, a case with an answer fixed by
      symmetry.
      
      Small ones are fine. The value is that the answer's provenance is external.
      
      ---
      
      ## When nothing published exists
      
      Some quantities have no literature. Options, in order:
      
      **Construct a degenerate case.** Set a parameter to a value where the answer
      becomes obvious — zero noise, one dimension, an identity transform, a
      single-element input. Assert the obvious answer.
      
      **Use a differential.** Absolute correctness may be unavailable while a
      *difference* is provable: this input must score higher than that one; adding
      data must not make the fit worse; the treated arm must beat the control.
      Differential anchors survive a great deal of implementation drift.
      
      **Manufacture a ground truth.** Generate synthetic data with a known answer
      baked in, then check the pipeline recovers it. Guard against the pipeline
      recovering it *by construction* — plant a case where the answer is deliberately
      not what the pipeline would assume.
      
      **Bound it from both sides with two crude methods** that are wrong in opposite
      directions. Neither is the answer; together they are a bracket.
      
      ---
      
      ## An anchor is also where a *magnitude* comes from
      
      Anchors are usually discussed as sources of correctness — is this value right?
      They are equally the only source of **practical significance**: is a difference
      big enough to care about?
      
      A threshold derived from measurement noise answers "is this effect real". It
      cannot answer "is this effect worth having", and the two diverge badly when the
      things being compared are similar: a paired estimator's standard error shrinks
      as the arms converge, so a 3-sigma margin can approach zero and admit an effect
      three orders of magnitude below anything that matters.
      
      So when a check exists to certify that something *helps*, it needs two floors:
      
      - a **statistical** floor from measured noise — the effect is not sampling luck;
      - a **practical** floor from an anchor — the effect is at least as large as
        <published method / theoretical gain / the smallest difference that would
        change a decision>.
      
      State which anchor supplies the second one. If none does, say so as a coverage
      limit: "this check certifies the effect is real, not that it is useful."
      
      ## Anchor hygiene
      
      - **Write the source next to the number.** `0.0716821  # Conway & Sloane,
        normalised second moment of E8`. An unattributed constant becomes a golden
        value within a release, and nobody can tell whether it is authority or
        history.
      - **Never derive an anchor from the code under test**, including "I ran it
        once and it looked right."
      - **Version-stamp cached artifacts.** If an anchor or a fitted object is
        cached, key the cache on the *method* as well as the problem, and delete on
        mismatch. A cache keyed only on the problem silently serves values produced
        by code you have since changed.
      - **State the covered range** every time. It is the single most common way an
        anchored gate turns out to have been blind.
      
    • auditing.md 7.2 KB
      # Auditing — can this check fail?
      
      For a check suite that already exists: tests, a linter config, a CI job, a
      calibration script. The question is not whether it passes. It is whether it
      *could* fail.
      
      Run this when a suite passes on the first try, when it has never gone red, when
      it was written by the same process that produced the code, or before you rely
      on it to gate something that matters.
      
      ---
      
      ## Pass 1 — name the failing input
      
      For each assertion, name a concrete input that makes it fail.
      
      If you cannot, it is decoration. Two things usually explain it:
      
      - **The assertion is independent of the subject.** "The screenshot contains at
        least N distinct colours" is equally true of an unchanged frame, so it passed
        for all seven keyboard shortcuts while none of the keystrokes were being
        delivered. Two tells were visible and read as success: every state reported
        the *identical* count, and the suite passed first try.
      - **The assertion restates the implementation.** `assert total == sum(xs)`
        where `total` is computed as `sum(xs)`. Tautologies pass forever.
      
      Replace both with **differential** checks: the state must change when it should,
      and an operation applied twice must return to the byte-identical original.
      
      ---
      
      ## Pass 2 — check the oracle's range
      
      For every check anchored on an external value, ask what range that anchor
      covers, and compare it to the range you actually operate in.
      
      The failure is quiet and total: past the end of the anchor's range there is no
      check at all, and the suite reports the same green it does everywhere else.
      
      A published quantizer table stopping at 5 bits meant a 16% error at 8 bits was
      unreachable by the table comparison — and the error's direction *loosened* a
      downstream check that depended on the same quantity, so the defect made the
      suite more permissive rather than less.
      
      Ask also: does the error direction of a wrong anchor tighten or loosen the
      checks downstream of it? A wrong value that loosens is far more dangerous than
      one that tightens, because tightening announces itself.
      
      ---
      
      ## Pass 3 — mutate
      
      ```bash
      python3 scripts/mutate.py --target src/thing.py -- <your gate command>
      ```
      
      Every survivor is a behaviour nothing checks.
      
      **First, disable any on-disk cache the subject keeps.** If the code under test
      caches artifacts keyed on the *problem* (`grid_m8_K65536.npz`) rather than on
      the *method*, a mutation of the producing code changes nothing observable: the
      gate loads the pre-mutation artifact and reports green. Every mutation to the
      trainer then "survives" for a reason that has nothing to do with the gate, and
      the survivor count is noise. Point the cache at a temporary directory for the
      mutation run. This is the same anti-pattern the gate section warns about, and
      it is far more damaging here because it inflates rather than hides.
      
      **Budget for it.** `mutate.py` runs the gate command once per mutation, so a
      90-second gate over 100 sites is two and a half hours. That pressure is what
      produces a fast gate variant — which is fine, as long as you remember that a
      known-bad validated only in the fast variant may not be bad at full size, and
      that the fast variant's reduced dimensions may make whole classes of mutation
      unobservable. Record which ones in the gate's stated coverage limits.
      
      **Reading the report.** Survivors cluster, and the cluster is the diagnosis:
      
      - Survivors on one function → that function is untested. Usually the fix is
        one test, not one assertion.
      - Survivors on comparison operators only → boundaries are untested. Add cases
        at, just below, and just above each threshold.
      - Survivors on numeric literals → thresholds and tolerances are unpinned.
        This is where *slack wider than the defect* lives: if the constant can move
        and nothing notices, it was never doing work.
      - Survivors on `and`/`or` → compound conditions are exercised on one branch
        only.
      
      **Legitimate survivors exist.** Equivalent mutants (a change with no
      observable effect), and code genuinely outside the gate's remit. Do not chase
      them to zero — move them to the gate's stated coverage limits, which converts
      an invisible hole into a written one. Write down *why* each is equivalent: an
      unexplained survivor and an equivalent mutant look identical in a report, and
      six months later nobody can tell which they are looking at.
      
      **Survivor line numbers are ephemeral — pin permanent fixtures by content.**
      The report says `grids.py:166`, and that reference is correct for exactly as
      long as nobody edits the file above line 166. If you promote survivors into a
      standing fixture (recommended — see the meta-check below), locate each target
      by a unique snippet of the line, not by its number. A fixture pinned by line
      survived until an upstream refactor shifted one file by ~50 lines, then failed
      with "line 206 does not contain '*'" — an error about the fixture, not about
      the subject, which is the most confusing kind to receive.
      
      ---
      
      ## Pass 4 — tolerances
      
      For every numeric tolerance, ask where it came from.
      
      If the answer is "it seemed reasonable," measure instead: run the thing several
      times, observe the spread, set the threshold outside the spread — then check
      that the *known-bad* also falls outside it. A tolerance that admits both the
      good and the bad case is worse than none, because it reports PASS with
      authority.
      
      A 2% slack once let a deliberately under-trained model through a check its
      properly-trained counterpart cleared by 2.6%. The slack was not derived from
      anything; it was generosity. Tightening it to a margin the real subject clears
      comfortably and the bad one does not restored the check's whole point.
      
      ---
      
      ## Pass 5 — provenance
      
      - Did any expected value come from running the code under test? Then it is a
        changelog entry, not an oracle.
      - Are cached or fitted artifacts keyed on the *method* that produced them, or
        only on the problem? A cache keyed `(m, K)` cannot notice that the trainer
        changed, and will serve stale artifacts indefinitely — including across a
        deliberate cache wipe, if a background writer loses the race.
      - Can you tell a stale artifact from a fresh one by looking at it? If the only
        difference is schema drift you happened to notice, add a version stamp.
      
      ---
      
      ## Pass 6 — the meta-check
      
      Break something on purpose and confirm the suite goes red.
      
      Not "invoke it and confirm the step appears in the output" — that catches a
      check that was never wired up, and says nothing about one that is wired up and
      toothless. Actually break the subject, actually watch it fail, actually put it
      back.
      
      Then keep it: a permanent known-bad fixture is worth more than the transient
      experiment, because it re-runs forever and catches the day someone widens a
      tolerance for an unrelated reason.
      
      ---
      
      ## Reporting an audit
      
      Say which of these is true, per check:
      
      - **CONFIRMED** — a named input makes it fail, and it was observed failing.
      - **PLAUSIBLE** — a named input should make it fail, not yet observed.
      - **CANNOT FAIL** — no input makes it fail. Delete or repair.
      - **BLIND** — it can fail, but not on the regime that matters. State the gap.
      
      A suite of forty checks where thirty-eight are CONFIRMED and two are BLIND is
      in far better shape than forty green ones of unknown status, and the audit is
      only worth writing down at that resolution.
      
  • scripts
    • gate.py 9.5 KB
      #!/usr/bin/env python3
      """A tiny harness for gates that can actually fail.
      
      The whole point of a gate is to go red when it should.  The characteristic
      failure is not a wrong check — it is a check that *cannot* fail, which reports
      PASS forever and reads exactly like a working one.
      
      So this harness refuses to let a gate pass unless it registered at least one
      `known_bad` case that its own criteria rejected.  A gate with no known-bad is
      reported as INCONCLUSIVE and exits non-zero, because "nothing looked wrong" is
      not evidence when nothing could have looked wrong.
      
          from gate import Gate
      
          g = Gate("codebook calibration")
      
          g.anchor("scalar quantizer, 2 bits", measured=0.117482, published=0.1175,
                   rel_tol=2e-3, source="Max (1960) table 1")
      
          g.bracket("trained grid sits between scalar and the Shannon bound",
                    value=0.0887, lo=0.0625, hi=0.1175,
                    why="must beat scalar; cannot beat the rate-distortion bound")
      
          g.known_bad("an under-trained grid is rejected",
                      rejected=under_trained_mse > threshold,
                      detail=f"{under_trained_mse:.5f} > {threshold:.5f}")
      
          g.coverage("Max's table stops at 5 bits — rates above that are unanchored")
      
          raise SystemExit(g.report())
      
      Stdlib only.  Import it, or copy the class into your gate script; it is small
      on purpose.
      """
      from __future__ import annotations
      
      import sys
      from dataclasses import dataclass, field
      
      
      @dataclass
      class _Check:
          ok: bool
          label: str
          detail: str
          kind: str  # "anchor" | "bracket" | "check" | "known-bad"
          covers: tuple[str, ...] = ()  # known-bad only: checks it exercises
      
      
      @dataclass
      class Gate:
          """Collects checks, then reports and returns a process exit code."""
      
          name: str = "gate"
          checks: list[_Check] = field(default_factory=list)
          notes: list[str] = field(default_factory=list)
          limits: list[str] = field(default_factory=list)
          stream = sys.stdout
      
          # -- primitives --------------------------------------------------------
      
          def check(self, ok: bool, label: str, detail: str = "", *,
                    kind: str = "check", covers: tuple[str, ...] = ()) -> bool:
              """Record a plain boolean check."""
              self.checks.append(_Check(bool(ok), label, detail, kind, covers))
              return bool(ok)
      
          def anchor(self, label: str, *, measured: float, published: float,
                     rel_tol: float, source: str) -> bool:
              """Compare a measurement against a value from outside your own code.
      
              `source` is required and is printed: an anchor whose provenance is not
              written down decays into a golden value, and a golden value only ever
              tells you the code still does what it did.
              """
              rel = abs(measured - published) / abs(published) if published else float("inf")
              return self.check(
                  rel < rel_tol, f"{label} [anchor: {source}]",
                  f"measured={measured:.6g} published={published:.6g} "
                  f"rel={rel:.2e} tol={rel_tol:.0e}", kind="anchor")
      
          def bracket(self, label: str, *, value: float, lo: float, hi: float,
                      why: str = "", lo_inclusive: bool = False,
                      hi_inclusive: bool = False) -> bool:
              """Assert lo < value < hi, or <= at either edge.
      
              Two-sided by construction.  A one-sided check passes for a value that
              collapsed as readily as for one that is right, which is how an
              implementation that silently does nothing gets certified.
      
              **Use the inclusive flag when an edge is ATTAINABLE.**  A theoretical
              bound is frequently reachable, and reaching it is optimal, not a
              failure.  A strict bound then reports red on the best possible result:
              a randomized-Hadamard rotation maps a coordinate spike to exactly
              ±1/sqrt(d), the information-theoretic floor for max|coord|, so
              `lo=1/sqrt(d)` rejected a perfect rotation and blocked a run.  Ask of
              each edge: can the subject legitimately sit exactly here?  If yes, make
              it inclusive.
              """
              lo_ok = (lo <= value) if lo_inclusive else (lo < value)
              hi_ok = (value <= hi) if hi_inclusive else (value < hi)
              lo_op = "<=" if lo_inclusive else "<"
              hi_op = "<=" if hi_inclusive else "<"
              suffix = f" — {why}" if why else ""
              return self.check(lo_ok and hi_ok, label,
                                f"{lo:.6g} {lo_op} {value:.6g} {hi_op} {hi:.6g}{suffix}",
                                kind="bracket")
      
          def known_bad(self, label: str, *, rejected: bool, detail: str = "",
                        covers: tuple[str, ...] = ()) -> bool:
              """Register a case the gate MUST reject, and whether it did.
      
              Build the bad case out of the same machinery the real subject uses and
              break it the way it would plausibly break — an under-trained model, a
              dropped term, an off-by-one — not a nonsense input that anything would
              catch.  A known-bad that is too obviously bad certifies nothing.
      
              **Validate it at the configuration it will actually run in.**  A
              known-bad tuned on a small/fast setting can stop being bad at full
              size: an "untrained" grid built from one Lloyd iteration was genuinely
              zero-gain at m=2/K=16 and earned a real +0.10 dB at m=8/K=65536, where
              a single iteration relocates ~63,000 empty cells toward the mode.  It
              passed the fast gate and certified nothing about the real one.
      
              `covers` names the check labels (or substrings of them) this case
              exercises.  Supplying it turns "we have a known-bad" into "we know
              WHICH checks have been shown to fire", which is a different and much
              stronger claim — see `report()`.
              """
              return self.check(rejected, label, detail, kind="known-bad",
                                covers=tuple(covers))
      
          def note(self, text: str) -> None:
              """Record a number worth seeing that is not itself pass/fail."""
              self.notes.append(text)
      
          def coverage(self, text: str) -> None:
              """Record something this gate CANNOT catch.
      
              Coverage holes are invisible from inside a green run, so they have to
              be asserted by the author.  Anchors that stop short of the range you
              operate in belong here.
              """
              self.limits.append(text)
      
          # -- reporting ---------------------------------------------------------
      
          @property
          def failures(self) -> list[_Check]:
              return [c for c in self.checks if not c.ok]
      
          @property
          def known_bads(self) -> list[_Check]:
              return [c for c in self.checks if c.kind == "known-bad"]
      
          def report(self) -> int:
              """Print the result; return 0 to pass, 1 to fail, 2 if inconclusive."""
              w = self.stream
              line = "=" * 74
              print(f"{line}\nGATE — {self.name}\n{line}", file=w)
              for c in self.checks:
                  tag = "PASS" if c.ok else "FAIL"
                  print(f"  [{tag}] ({c.kind}) {c.label}"
                        + (f": {c.detail}" if c.detail else ""), file=w)
              for n in self.notes:
                  print(f"  [note] {n}", file=w)
              for lim in self.limits:
                  print(f"  [cannot catch] {lim}", file=w)
      
              if self.failures:
                  print(f"\nFAILED — {len(self.failures)} check(s):", file=w)
                  for c in self.failures:
                      print(f"  * {c.label}: {c.detail}", file=w)
                  return 1
              if not self.known_bads:
                  print("\nINCONCLUSIVE — every check passed and no known-bad case was "
                        "registered.\nA gate that was never shown to reject anything has "
                        "not been shown to work.\nAdd g.known_bad(...) with a case its "
                        "own criteria must catch.", file=w)
                  return 2
              if not self.limits:
                  print("\nINCONCLUSIVE — no coverage limit recorded. State at least one "
                        "thing\nthis gate cannot catch (g.coverage(...)); if you truly "
                        "believe there is\nnothing, say that explicitly.", file=w)
                  return 2
              # Known-bad REACH.  "We have a known-bad" and "we know which checks
              # have been shown to fire" are different claims, and only the second
              # is worth much.  An audit of a real gate found its single known-bad
              # exercised 1 of 8 checks -- and the check the whole result rested on
              # ACCEPTED the same bad case.  That gate reported PASS.
              substantive = [c for c in self.checks if c.kind != "known-bad"]
              if any(kb.covers for kb in self.known_bads):
                  claimed = [t for kb in self.known_bads for t in kb.covers]
                  reached = {c.label for c in substantive
                             if any(t in c.label for t in claimed)}
                  unreached = [c.label for c in substantive if c.label not in reached]
                  print(f"\n  known-bad reach: {len(reached)}/{len(substantive)} "
                        f"checks exercised by a known-bad", file=w)
                  for lab in unreached[:12]:
                      print(f"    [unreached] {lab}", file=w)
                  if len(unreached) > 12:
                      print(f"    ... and {len(unreached) - 12} more", file=w)
              else:
                  print(f"\n  known-bad reach: UNDECLARED — {len(self.known_bads)} "
                        f"known-bad(s), none naming the checks they exercise.\n"
                        f"  Pass covers=(...) to known_bad() to turn 'we have one' "
                        f"into 'we know which checks fire'.", file=w)
      
              print(f"\nPASSED — {len(self.checks)} checks, "
                    f"{len(self.known_bads)} known-bad rejected, "
                    f"{len(self.limits)} coverage limit(s) stated.", file=w)
              return 0
      
      
      __all__ = ["Gate"]
      
    • mutate.py 6.4 KB
      #!/usr/bin/env python3
      """Break the code on purpose and see whether the gate notices.
      
      Writing a known-bad case by hand tests the one failure you thought of.  This
      tests the ones you did not: it perturbs single tokens in the target file, runs
      your gate after each perturbation, and reports the mutations that **survived**
      — the ones where the code changed and the gate stayed green.  Each survivor is
      a behaviour nothing is currently checking.
      
          python3 mutate.py --target grids.py -- python3 calibrate.py
          python3 mutate.py --target src/codec.py --max 40 -- pytest -q tests/
      
      The gate command must pass on unmutated code first; if it does not, that is
      reported and nothing else runs, because survivor counts against an already-red
      gate mean nothing.
      
      Uses `tokenize`, so string literals and comments are never touched — a
      regex-based mutator will happily corrupt a docstring and report a "kill" that
      is really a SyntaxError.
      
      Stdlib only.  This is the zero-dependency pass that works against ANY gate
      command; for a Python test suite specifically, `mutmut` and `cosmic-ray` are
      more thorough and worth reaching for once this stops finding survivors.
      """
      from __future__ import annotations
      
      import argparse
      import io
      import subprocess
      import sys
      import tokenize
      from dataclasses import dataclass
      from pathlib import Path
      
      #: Operator swaps.  Each flips a decision boundary, which is where gates that
      #: only ever see well-formed input tend to be blind.
      OPS = {
          "==": "!=", "!=": "==",
          "<": "<=", "<=": "<", ">": ">=", ">=": ">",
          "+": "-", "-": "+", "*": "/", "//": "/",
          "and": "or", "or": "and",
          "True": "False", "False": "True",
          "is": "is not",
          "min": "max", "max": "min",
          "all": "any", "any": "all",
      }
      
      
      @dataclass
      class Mutation:
          line: int
          col: int
          end_col: int
          before: str
          after: str
      
          def label(self) -> str:
              return f"line {self.line}: {self.before} -> {self.after}"
      
      
      def find_mutations(src: str) -> list[Mutation]:
          """Token-level mutation sites, skipping strings, comments and numbers-in-strings."""
          out: list[Mutation] = []
          try:
              toks = list(tokenize.generate_tokens(io.StringIO(src).readline))
          except (tokenize.TokenError, IndentationError, SyntaxError) as e:
              raise SystemExit(f"cannot tokenize target: {e}") from e
          for tok in toks:
              if tok.type not in (tokenize.OP, tokenize.NAME, tokenize.NUMBER):
                  continue
              if tok.start[0] != tok.end[0]:
                  continue  # multi-line token; skip
              s = tok.string
              if tok.type == tokenize.NUMBER:
                  # perturb the literal: catches thresholds and tolerances that no
                  # test actually pins, which is where "slack larger than the defect"
                  # lives
                  try:
                      after = str(int(s) + 1) if s.isdigit() else None
                  except ValueError:
                      after = None
                  if after is None:
                      continue
                  out.append(Mutation(tok.start[0], tok.start[1], tok.end[1], s, after))
              elif s in OPS:
                  out.append(Mutation(tok.start[0], tok.start[1], tok.end[1], s, OPS[s]))
          return out
      
      
      def apply_mutation(src: str, m: Mutation) -> str:
          lines = src.splitlines(keepends=True)
          ln = lines[m.line - 1]
          lines[m.line - 1] = ln[: m.col] + m.after + ln[m.end_col :]
          return "".join(lines)
      
      
      def run(cmd: list[str], timeout: int) -> tuple[bool, str]:
          """True when the gate PASSES (exit 0)."""
          try:
              p = subprocess.run(cmd, capture_output=True, text=True,
                                 timeout=timeout, check=False)
              return p.returncode == 0, f"exit {p.returncode}"
          except subprocess.TimeoutExpired:
              return False, "timeout"
          except OSError as e:
              return False, f"oserror {e}"
      
      
      def main() -> int:
          ap = argparse.ArgumentParser(
              description="Mutate a file and report which mutations the gate misses.")
          ap.add_argument("--target", action="append", required=True,
                          help="source file to mutate (repeatable)")
          ap.add_argument("--max", type=int, default=60,
                          help="cap on mutations per target (default 60)")
          ap.add_argument("--timeout", type=int, default=600,
                          help="seconds per gate run (default 600)")
          ap.add_argument("--stride", type=int, default=1,
                          help="sample every Nth site instead of the first --max")
          ap.add_argument("cmd", nargs=argparse.REMAINDER,
                          help="-- followed by the gate command")
          a = ap.parse_args()
          cmd = [c for c in a.cmd if c != "--"]
          if not cmd:
              ap.error("provide the gate command after --")
      
          ok, why = run(cmd, a.timeout)
          if not ok:
              print(f"BASELINE RED ({why}) — the gate fails on unmutated code.\n"
                    f"Fix that first; survivor counts against a red gate mean nothing.")
              return 2
          print(f"baseline: gate passes on unmutated code ({' '.join(cmd)})\n")
      
          survivors: list[tuple[str, Mutation]] = []
          killed = 0
          for target in a.target:
              path = Path(target)
              original = path.read_text()
              sites = find_mutations(original)[:: a.stride][: a.max]
              print(f"{target}: {len(sites)} mutation sites")
              try:
                  for i, m in enumerate(sites, 1):
                      path.write_text(apply_mutation(original, m))
                      passed, why = run(cmd, a.timeout)
                      if passed:
                          survivors.append((target, m))
                          mark = "SURVIVED"
                      else:
                          killed += 1
                          mark = f"killed ({why})"
                      print(f"  [{i}/{len(sites)}] {m.label():<34} {mark}", flush=True)
              finally:
                  path.write_text(original)  # always restore, even on Ctrl-C
      
          total = killed + len(survivors)
          print(f"\n{'=' * 70}")
          if not total:
              print("no mutation sites found — is the target the file the gate exercises?")
              return 2
          print(f"killed {killed}/{total}   survived {len(survivors)}/{total}")
          if survivors:
              print("\nSURVIVORS — the code changed here and the gate stayed green:")
              for target, m in survivors:
                  print(f"  {target}:{m.line}  {m.before} -> {m.after}")
              print("\nEach line is a behaviour nothing currently checks. Either add a "
                    "check,\nor record it under the gate's stated coverage limits.")
              return 1
          print("\nevery mutation was caught")
          return 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
  • CHANGELOG.md 3 KB
    # gating - Changelog
    
    All notable changes to the `gating` skill are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/).
    
    ## [0.3.0] - 2026-08-25
    
    ### Other
    
    - top skills: separate by omission, and correct the guidance that said otherwise (#777)
    
    ## [0.2.0] - 2026-08-03
    
    ### Other
    
    - gating 0.2.0: lessons from the first substantial use of the skill (#752)
    
    ## [0.2.0] - 2026-08-02
    
    Lessons from the first substantial use of the skill — auditing and rebuilding
    the calibration gate in `oaustegard/experiments/remex-vs-higgs-ablation`
    (8 checks -> 166, 91 mutants, four false reds in the rebuilt gate). Every item
    below is something the skill led into or failed to warn about, not a
    hypothetical.
    
    ### Added
    
    - `bracket()` takes `lo_inclusive` / `hi_inclusive`. A strict edge at an
      *attainable* theoretical bound reports red on a perfect result: a randomized
      Hadamard maps a coordinate spike to exactly 1/sqrt(d), the information-
      theoretic floor, and a strict `lo` blocked a real run.
    - `known_bad(..., covers=(...))` names the checks a case exercises, and
      `report()` prints the checks no known-bad reaches. "We have a known-bad" and
      "we know which checks fire" are different claims; an audited gate had one
      known-bad covering 1 of 8 checks, and the check its result rested on
      *accepted* the same bad case.
    - SKILL.md: validate a known-bad at the configuration it will run in — one
      tuned at m=2/K=16 stopped being bad at m=8/K=65536.
    - SKILL.md + `anchors.md`: a statistical margin is not a practical floor. A
      paired estimator's se shrinks as the arms converge, so a 3-sigma threshold
      admitted a +0.0001 dB effect where real ones were 0.35-1.41 dB. The
      magnitude that matters has to come from an anchor.
    - SKILL.md: two anti-pattern rows (significant-but-not-meaningful; strict
      bracket at an attainable optimum) and an explicit statement that a gate
      checks *correctness* and cannot check *comparability* — performance claims
      have no anchor in this framework.
    - `auditing.md`: disable on-disk caches before mutation testing, or every
      mutation of the producing code survives for the wrong reason; and pin
      permanent mutation fixtures by code snippet, because survivor line numbers
      go stale the moment anyone edits above them.
    
    ## [0.1.0] - 2026-08-02
    
    ### Added
    
    - initial release: build and audit deterministic verification gates
    - three obligations — anchor, known-bad, stated coverage limit
    - `scripts/gate.py`: harness that reports INCONCLUSIVE (exit 2) rather than
      PASS when no known-bad or no coverage limit was registered
    - `scripts/mutate.py`: stdlib token-level mutation pass over any gate command;
      refuses to run against an already-red gate, restores targets on interrupt
    - `references/anchors.md`: six kinds of oracle strongest-first, what to do when
      nothing published exists, anchor hygiene
    - `references/auditing.md`: six-pass "can this check fail?" sweep, with
      CONFIRMED / PLAUSIBLE / CANNOT FAIL / BLIND reporting resolution
  • README.md 2.3 KB
    # gating
    
    Build and audit verification gates — deterministic checks that block a
    pipeline and can be **shown** to go red.
    
    The characteristic failure of a gate is not a wrong check. A wrong check gets
    noticed. It is a check that *cannot* fail, which reports PASS forever and is
    indistinguishable from a working one from the outside.
    
    So the object under suspicion here is the check itself.
    
    ## The three obligations
    
    | | | |
    |---|---|---|
    | **Anchor** | something your own code did not produce | published constant, closed form, invariant, theoretical bound |
    | **Known-bad** | a case the gate demonstrably rejects | break the subject plausibly, run the gate, confirm red |
    | **Coverage limit** | what it cannot catch, in writing | holes are invisible from inside a green run |
    
    `scripts/gate.py` enforces the last two: a gate that registers no known-bad or
    no coverage limit reports **INCONCLUSIVE** and exits 2, not PASS.
    
    ## Scripts
    
    ```bash
    # mutation pass — which single-token changes does the gate NOT notice?
    python3 scripts/mutate.py --target src/codec.py -- python3 calibrate.py
    ```
    
    `mutate.py` is stdlib-only, uses `tokenize` (so it never corrupts strings or
    comments), restores the file even on interrupt, and refuses to run when the
    gate is already red. It works against *any* gate command; `mutmut` and
    `cosmic-ray` go deeper once it stops finding survivors.
    
    `gate.py` is a ~150-line harness: `anchor()`, `bracket()`, `known_bad()`,
    `coverage()`, `note()`, and a `report()` that returns a process exit code.
    
    ## Relationship to `challenging`
    
    `challenging` asks *would a careful reader object?* — LLM judgement against a
    persona, verdicts of SHIP/REVISE/RETHINK. `gating` asks *can this check go
    red?* — deterministic, anchored, exit-coded.
    
    Complements with different blind spots: an adversarial reviewer will not
    recompute your constants, and a gate will not notice that your framing is
    wrong. Note also that a same-model reviewer is an independent *context*, not
    an independent *reviewer*; where shared priors are the risk, an anchor beats a
    reviewer because an anchor is not negotiable.
    
    See [`SKILL.md`](SKILL.md) for the build and audit procedures and the
    anti-pattern table, [`references/anchors.md`](references/anchors.md) for
    choosing an oracle, and [`references/auditing.md`](references/auditing.md) for
    the six-pass "can this fail?" sweep over an existing suite.
    
  • SKILL.md 12.3 KB
    ---
    name: gating
    description: >-
      Build and audit deterministic verification gates — a check that blocks a
      pipeline and can be shown to go red. Use when writing a calibration gate,
      CI check, validation script or pre-publication check for a numeric or
      empirical result; when a plausible-but-wrong value would survive review;
      when asking whether an existing test, linter rule or check could actually
      fail; and when a suite passes first try, passes suspiciously often, or was
      written by whatever produced the thing it checks. Triggers on "can this
      check fail", "known-bad", "negative control", "calibration gate", "sanity
      check my results", "is this test actually testing anything".
    metadata:
      version: 0.3.0
    ---
    
    # gating
    
    A gate is a check that blocks. Its only job is to go red when it should.
    
    The characteristic failure is not a wrong check — a wrong check gets noticed.
    It is a check that **cannot fail**, which reports PASS forever and is
    indistinguishable from a working one from the outside. That is what makes this
    different from ordinary testing: the object under suspicion is the check.
    
    ## When NOT to use this skill
    
    Scope is ONE check and whether it can be made to fail.
    
    | Situation | Use |
    |---|---|
    | Sequence several steps with branches and retries | flowing |
    | Run the repo's existing suite | run it |
    | Decide what to test at all | this skill has no opinion; that is design |
    
    A gate is a thing that goes red. If nothing here can go red, there is no gate
    to audit.
    
    ## The three obligations
    
    Every gate owes these. A gate missing any of them is not yet a gate.
    
    **1. An anchor outside your own code.** Something the check compares against
    that your implementation did not produce: a published constant, a closed-form
    answer, a conservation law, a degenerate case with a known result, an
    independent implementation. A check that compares this run to the last run only
    ever tells you the code still does what it did. See `references/anchors.md`.
    
    **2. A known-bad it demonstrably rejects.** Break the subject the way it would
    plausibly break, run the gate, confirm red. Until you have done this you have
    not shown the gate works — you have shown it runs. This is the obligation
    people skip, because a passing gate feels like evidence.
    
    Two things about known-bads that are easy to get wrong:
    
    - **Validate it at the configuration it will run in.** A case tuned on a small
      or fast setting can stop being bad at full size. An "untrained" grid built
      from one Lloyd iteration was genuinely zero-gain at m=2/K=16 and earned a
      real +0.10 dB at m=8/K=65536, where one iteration relocates ~63,000 empty
      cells toward the mode. It passed the fast gate and certified nothing about
      the real one. This matters more than it sounds, because `mutate.py` needs a
      fast gate variant and it is tempting to validate everything there.
    - **Measure its *reach*.** One known-bad is the floor, not the goal. Name which
      checks it exercises (`known_bad(..., covers=(...))`); the harness prints the
      checks no known-bad reaches. An audited gate had a single known-bad covering
      1 of 8 checks — and the check its whole result rested on *accepted* the same
      bad case.
    
    **3. A written statement of what it cannot catch.** Coverage holes are
    invisible from inside a green run: the gate is silent about the thing it does
    not look at, in exactly the same tone it uses for the thing it looked at and
    approved. The author has to assert the hole; nothing else will.
    
    `scripts/gate.py` enforces obligations 2 and 3 mechanically — it returns exit
    code 2 (INCONCLUSIVE, not PASS) when a gate registers no known-bad or no
    coverage limit.
    
    ## Building a gate
    
    Work in this order. The first step is the one that determines whether the rest
    is worth anything.
    
    **Name the wrong conclusion, not the component.** Not "check the quantizer is
    correct" but "prevent shipping *scalar wins at high bit rates* when that would
    really be an optimizer artifact." A gate aimed at a conclusion knows what
    counts as a near-miss; a gate aimed at a component just exercises the code.
    
    **Find an anchor.** `references/anchors.md` lists the kinds, in rough order of
    strength, with the questions that find each one.
    
    **Prefer brackets to point checks.** Assert a value lies strictly between two
    things it cannot legitimately pass: better than a baseline, worse than a
    theoretical bound. A one-sided check passes for a result that collapsed as
    readily as for one that is right — which is how an implementation that
    silently does nothing gets certified.
    
    **Derive the tolerance from measured noise.** Run the thing several times, see
    how much it moves, put the threshold outside that. A tolerance picked for
    comfort tends to land wider than the defect you are trying to catch, and then
    it swallows it.
    
    **Then check it is not too tight to mean anything.** The opposite failure is
    real and less obvious: a margin can be statistically impeccable and practically
    empty. A *paired* estimator — scoring both arms on one shared sample so the
    common fluctuation cancels — is the right way to measure a difference, and its
    standard error *shrinks as the two arms converge*. So "beats the baseline by
    3 se" degenerates: a codebook perturbed by N(0, 1e-3) gained +0.0001 dB against
    a 3-se margin of 1.2e-06 and was accepted, while real ones gained 0.35–1.41 dB.
    The check certified *the effect is real*, not *the effect is worth having*.
    Those are different assertions and need different thresholds — and the second
    one has to come from an anchor (there, a published lattice codebook), never
    from the estimator, which knows nothing about what magnitude would matter.
    
    **Build the known-bad and confirm red.** Then run `scripts/mutate.py` for the
    failures you did not think of.
    
    **Wire it to a non-zero exit** and run it before the thing it gates, not after.
    A gate that runs after the results are written is a report.
    
    ## Auditing an existing check suite
    
    Given tests, a linter config, a CI job, or a gate someone already wrote, the
    question is not "do these pass" but "can these fail". Full procedure in
    `references/auditing.md`; the fast version:
    
    - For each assertion, name a concrete input that makes it fail. If you cannot,
      it is decoration — delete it or fix it.
    - Check each oracle's range against the range you actually operate in. A
      published table that stops short of your regime is a hole with a green light
      on it.
    - Run `scripts/mutate.py` against the code the suite covers. Every survivor is
      a behaviour nothing checks.
    - Look for assertions whose truth does not depend on the subject at all.
    - Look for tolerances wider than the effect being measured.
    
    ## Scripts
    
    ```bash
    # harness: refuses to report PASS without a known-bad and a coverage limit
    python3 scripts/gate.py           # importable; see the module docstring
    
    # mutation pass: which single-token changes does the gate NOT notice?
    python3 scripts/mutate.py --target src/codec.py -- python3 calibrate.py
    python3 scripts/mutate.py --target grids.py --max 40 -- pytest -q
    ```
    
    `mutate.py` requires the gate to pass on unmutated code first and refuses to
    run otherwise, because survivor counts against an already-red gate mean
    nothing. It restores the file even on interrupt, and uses `tokenize` so string
    literals and comments are never corrupted. It is the zero-dependency pass that
    works against *any* gate command; once it stops finding survivors, `mutmut` or
    `cosmic-ray` go deeper on Python test suites specifically.
    
    ## Anti-patterns
    
    Each of these has shipped a wrong result somewhere. They are ordered by how
    convincingly they impersonate a working gate.
    
    | Anti-pattern | Why it survives review |
    |---|---|
    | **Slack wider than the defect** | A tolerance chosen for comfort. The gate passes the real thing *and* the broken thing, and reports PASS for both. Derive the threshold from noise, then confirm the known-bad falls outside it. |
    | **An oracle with a coverage hole** | Published anchors end somewhere. If the defect is past the end of the table, the check is structurally incapable of catching it and looks fine. State the range the anchor covers. |
    | **An assertion whose truth doesn't depend on the subject** | "The output has at least N distinct colours" is equally true of an unchanged frame. Prefer differential checks: the state must *change* when it should, and a toggle applied twice must return to the byte-identical original. |
    | **Confirming the check ran, not that it can fail** | "Invoke it and confirm the step appears in the output" catches a check that was never wired up. It says nothing about a check that is wired up and toothless. |
    | **Comparing against your own previous output** | Regenerated goldens ratify drift. If the golden came from the code under test, it is a changelog, not an oracle. |
    | **A cache keyed on the problem rather than the method** | `cache[(m, K)]` cannot notice that the code producing the value changed. Version-stamp the artifact and delete on mismatch instead of trusting. |
    | **A self-matching predicate** | `until ! pgrep -f trainer` never exits, because the watching shell's own argv contains `trainer`. Worse, a malformed variant exits immediately and reports the job finished while it runs. Wait on a PID. |
    | **A margin that is significant but not meaningful** | A threshold derived purely from estimator noise certifies that an effect is *real*, not that it is *worth having* — and a paired estimator's noise shrinks as the arms converge, so the margin can approach zero. Pair every noise-derived floor with a magnitude an anchor says would matter. |
    | **A strict bracket at an attainable optimum** | A theoretical bound is often reachable, and reaching it is the best possible outcome. A strict edge then goes red on a perfect result and blocks real work. Ask of each edge whether the subject can legitimately sit exactly there. |
    | **A gate written by whatever produced the artifact** | Shared assumptions produce shared blind spots, and the convention both inherited is the one neither questions. Anchors are the defence, because an anchor is the one input the producer did not choose. |
    
    ## Division of labour
    
    This skill is for results and pipelines where the failure mode is a
    **plausible wrong number** that would survive a careful read.
    
    | Use | When the risk is |
    |---|---|
    | **`gating`** (this) | A number or empirical result is about to be published or acted on, and a wrong-but-reasonable value would pass unnoticed. Output: a gate that blocks. |
    | **`challenging`** | An artifact would draw a specific objection from a skeptical reader — prose, analysis, a recommendation, a diff. LLM judgement against a persona. Output: findings and a SHIP/REVISE/RETHINK verdict. |
    | **`verifying-claims`** | Documentation says something about code that is no longer true. Output: prose-vs-code disagreements. |
    | **A test suite / TDD** | Code you wrote does not behave as specified. Output: red tests. |
    
    `challenging` asks *would a careful reader object?* `gating` asks *can this
    check go red?* They are complements and they miss different things: an
    adversarial reviewer will not recompute your constants, and a gate will not
    notice that your framing is wrong.
    
    **A gate checks correctness, and cannot check comparability.** If two arms of a
    comparison are each individually correct but not *comparably implemented*, no
    anchor and no mutant will see it: nothing is broken, so nothing goes red. A
    published ablation reported one transform 11–24× slower than another and
    carried a caveat saying the number was implementation-bound; both arms passed
    every correctness check, and the real finding — one arm was a tuned BLAS call
    and the other an interpreted loop — was found by a human reviewer months of
    gate-work later. Related: **performance claims have no anchor in this
    framework at all.** There is no published constant for how fast your code
    should be. Wall-clock belongs to benchmarking discipline (matched
    implementation effort, min-of-trials, stated hardware), not to gating.
    
    One caution about pairing them. A same-model reviewer is an independent
    *context*, not an independent *reviewer* — it shares your priors, so the
    convention you did not question is the one it will not question either. Where
    that matters, an anchor beats a reviewer, because an anchor is not
    negotiable.
    
    ## References
    
    - `references/anchors.md` — kinds of oracle, strongest first, and how to find
      one when nothing published exists.
    - `references/auditing.md` — the full "can this fail?" pass over an existing
      suite, including how to read a mutation report.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related