Claude Skill

test-audit

Audit whether tests detect regressions in the behavior they claim to protect. Find mocked-away subjects, weak or circular assertions, undiscriminating fixtures, swallowed failures, and tests missing from gates. Use for test-quality audits, suspected false confidence in generated

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download iliaal-ai-skills-skills_test-audit-7b0520a.zip · 56 KB
Part of iliaal/ai-skills — 26 skills

Install

skills CLI npx skills add https://github.com/iliaal/ai-skills/tree/master/skills/test-audit
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install iliaal-ai-skills@llmmart
Git git clone https://github.com/iliaal/ai-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole iliaal/ai-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Test audit

Find tests that remain green when their claimed behavior breaks. Assess behavior, not apparent authorship, mock count, assertion count, or a target finding percentage. Default to a read-only audit. Change tests or production code only when requested.

Adapted from the OpenClaw test-audit skill (MIT).

What counts as a finding

Evaluate each claimed behavior separately. A test can protect one obligation and miss another. Use these dispositions:

  • Ineffective for the claim: the named behavior is replaced, never reached, or disconnected from a verdict that can fail.
  • Partially effective: real behavior is tested, but a specific promised distinction is absent from the input or assertions. Preserve the protection that exists.
  • Not enforced: the test is uncollected, skipped, or absent from a blocking gate. Report local collection and CI enforcement separately from assertion quality.
  • Redundant: another test protects the same behavior at the same boundary under equivalent inputs and setup. This requires separate evidence; weakness alone is not redundancy.
  • Adequate for its scope: the test detects a credible failure of its actual contract.
  • Unresolved: missing context or execution prevents a defensible decision.

A missing additional edge case is not automatically a defective existing test. Tie each quality finding to an existing claim, documented contract, or misleading assertion. Distinguish an individual test's weakness from a suite-wide coverage gap.

1. Establish scope and execution

Read repository instructions, runner configuration, shared fixtures, test helpers, and CI commands. Record the exact focused command and applicable completion gates. Determine which tests are collected, execute, and block a merge. An advisory job is not a blocking gate. A local green result does not establish CI enforcement.

Build an inventory from the runner's collection output when practical. Otherwise label the source inventory approximate. Count files, test definitions, and parameter cases separately. Do not treat helpers as tests or silently omit unrecognized syntax.

For a bounded request, inspect every in-scope test. For a large suite, declare the review budget and selection method before discovery. Cover distinct layers, major fixtures, and mocking styles. Combine risk-directed selection with a reproducible sample independent of detector hits. Keep those two sample results separate. Record the seed and sample unit if claiming random sampling. Expand recurring patterns into their owner family when practical; report anything left unreviewed.

Do not infer suite prevalence from targeted examples. If a prevalence estimate is requested, use a representative sample with an explicit denominator and uncertainty. Do not tune findings to an expected rate such as 10–20%.

2. Discover through behavior

Read behavioral-review.md before auditing test bodies. For each reviewed test, trace:

  1. Claim: what observable behavior does the name, fixture, or requirement promise?
  2. Real subject: which production operation actually runs, including setup, imported helpers, autouse fixtures, module mocks, and dependency injection?
  3. Input: does the fixture create the distinction the claim depends on?
  4. Oracle: where does the expected answer come from, and which assertion checks it?
  5. Verdict: will an incorrect result reach the runner as a failure?
  6. Counterexample: what plausible wrong implementation would still pass this test?

Construct a concrete counterexample before declaring an assertion sufficient. Examples include ignoring one filter, selecting the first item instead of the minimum, dropping a forwarded prop, or returning an empty result where only the type is asserted. Prefer a changed observable result over a crash or syntax error.

Inspect both positive and negative tests. Trace assertion helpers to their actual checks. Read the relevant implementation and neighboring tests before reporting a gap. For mocks, identify whether the replaced operation owns the claimed behavior or is a collaborator of a real subject. Check arguments, call requirements, returned data, and observable effects at that boundary.

When delegation is available, split semantic discovery by owner or directory. Each reviewer discovers independently of detector output. Require exact locations, real subject, surviving regression, retained value, and evidence status. Independently verify the strongest findings before reporting them.

3. Use detectors as additional leads

Run each script by its path in this skill's directory, not the audited repository's scripts/, with the audited repository root as the working directory. Write output to a temporary directory. Use the repository's pinned tooling for runner commands. The scripts require Python 3.11+ and the standard library.

Script What its results mean
weak_negatives.py <test globs> Heuristic candidates for ambiguous refusal assertions. Does not assess positive behavior or assertion data flow.
duplicate_tests.py <test globs> --json Structural duplicate/consolidation candidates, not an effectiveness score.
junit_census.py <report.xml> --test-files <glob> Reported skips, possible collection gaps, and PHPUnit zero-assertion leads (adequate when not throwing is the documented contract).
guard_pins.py --src <glob> --tests <glob> Refusal strings without textual matches; not proof of guard coverage.
pytest_reach_probe.py Runs the target suite with a recording plugin; use a scratch copy. See stacks.md.

Before trusting any output, read detectors.md: for the four static detectors, status 2 means known incomplete input, and each detector has blind spots that require source review. Read stacks.md for runner/probe commands and duplicate parser limits. Confirm tool flags locally before use. No detector hits means only no matches to that detector's rules; continue the semantic sample.

4. Confirm regression sensitivity

Static evidence can establish a finding when the data flow is explicit, such as a test asserting a value it assigned itself. Label it source-confirmed. Otherwise state the counterexample as inferred until executed. Use execution-confirmed only when the exact test ran against a verified behavioral change.

For high-impact or uncertain findings, use a focused mutation in an isolated scratch copy. Do not mutate the user's checkout for an audit. Keep network calls, paid APIs, and shared databases outside the probe unless separately authorized and isolated. Before running a mutation, read mutation-and-repair.md for the procedure and how to read a surviving mutant.

5. Report actionable evidence

For each finding, give:

  • Exact test name and file:line, plus the relevant production location.
  • Claimed behavior, actual real subject, and input/assertion/verdict defect.
  • Concrete surviving regression and retained useful protection.
  • Evidence status, command/results when executed, and unresolved limitations.
  • Individual-test versus suite scope, neighboring coverage checked, and repair direction.

Report reviewed scope and selection method alongside findings. Separate confirmed findings, unresolved candidates, and retained counterexamples. Do not count every missing obligation as a separate bad test. State commands not run and reasons.

Before proposing deletion of a production seam, or when changes are authorized, or when authoring a new test, read mutation-and-repair.md.

Files (ai-skills)
  • references
    • behavioral-review.md 7.3 KB
      # Behavioral review
      
      Use this checklist on independently selected tests, including files with no detector
      hits. Follow fixture and assertion-helper definitions before judging a test body.
      
      ## Trace the source of the answer
      
      Identify the exact behavior under review. Compare the operation that should compute
      the answer with the operation the test actually invokes.
      
      | Suspect shape | Evidence needed | Useful counterexample or repair |
      |---|---|---|
      | Mock supplies the subject's answer | The replaced operation itself owns the claimed transformation, persistence, or decision. No real owner runs. | Change that owner to return an incorrect value. Call the real owner with a fake only at its dependency boundary. |
      | Fixture constructs the finished result | The test hand-builds the saved record, rendered payload, receipt, or callback order that production should produce. | Stop production from producing that result. Feed the test raw input through the real producer. |
      | Self-derived expected value | Expected and actual share the same implementation or the expected object is compared to itself. | Break the shared computation. Use an independently specified answer or property that constrains the result. |
      | Type/truthiness/presence only | The name promises particular values, selection, order, or effects, but any object/non-null/array succeeds. | Return a wrong value of the same type, an empty collection, or omit the claimed effect. Assert meaningful contents. |
      | Test reimplements the algorithm | Only the test's copy executes or the oracle repeats the same decision with the same failure modes. | Change the production algorithm and trace whether any assertion depends on it. |
      
      An independent oracle may compute its answer. A simple mathematical relation or a
      different reference algorithm is not automatically circular. Shared bugs require a
      concrete explanation, not merely a visual resemblance between calculations.
      
      ## Make competing behaviors produce different answers
      
      | Claim | Fixture that hides the regression | Discriminating input |
      |---|---|---|
      | Sort or select minimum | Input is already in the expected order. | Entries arrive out of order with distinct keys. |
      | Compute median | One/two values, or symmetric values where mean equals median. | Skewed values such as 0, 100, 900. |
      | Apply several filters | Every rejected row also fails another filter. | One excluded row per filter that satisfies all other filters. |
      | Preserve manual selection | Automatic selection would choose the same item. | Observer input favors a different item while the override is active. |
      | Reject just one invalid property | Fixture violates multiple earlier checks. | Valid control plus exactly one invalid dimension; inspect the intended refusal. |
      | Keep no writes/queries/events | Empty inputs, disabled path, fixed-empty fake, or unchanged coarse timestamp. | Reach the live path; observe the effect independently; prove a planted effect is visible. |
      | All results satisfy a property | Result collection can be empty. | Assert expected population or identities before the universal property. |
      | Coalesce repeated requests | Only one request or requests that never overlap. | Same-key overlapping requests with controlled scheduling; observe one shared operation. |
      | Keep independent requests separate | Only one key, so incorrect merging remains invisible. | Distinct keys with independently asserted outcomes. |
      | Process concurrently | One item or synchronous completion hides serialization. | Multiple items and controlled scheduling that distinguishes concurrent entry from sequential work. |
      | Handle unknown ID | No existing records, so ignoring the requested ID still returns empty. | Seed a competing record and assert that its values do not appear. |
      
      Do not invent unsupported input shapes. Trace producer constraints when reachability
      is uncertain. A test of one valid equivalence class need not cover every other class;
      report the missing distinction only within the behavior the test claims.
      
      ## Follow failure into the runner
      
      - In pytest/unittest, returning `False` does not fail the test. Printed errors and
        caught exceptions can also finish successfully. Check failure branches for an
        assertion, rethrow, or runner failure call.
      - Check whether a prerequisite branch returns normally, skips visibly, or fails.
        A skip is honest nonexecution but still leaves the behavior unverified.
      - Check that async work and rejection assertions are awaited or returned to the
        runner. Confirm assertions in callbacks actually execute before test completion.
      - Inspect `if`, loops, catch blocks, and helper return values around assertions.
        An assertion that is never reached provides no protection.
      - Acceptance of success and failure together, such as status in `(200, 500)`, cannot
        establish successful service behavior. Identify the exact contract before judging
        an allowed set of outcomes.
      - For negatives, inspect the actual error and the path producing it. A bare exception
        or status can be sufficient when that is the contract and the fixture reaches the
        relevant operation. Do not require fragile message text universally.
      - Read matcher semantics. An `any`/`assertSent` predicate returning true for unrelated
        items lets those items satisfy the check. A `not called` assertion on a detached
        spy never observes the live operation.
      
      ## Keep legitimate doubles and narrow tests
      
      Retain these when the real subject makes the asserted decision:
      
      - External transport mocked while real code builds and sends an exact payload.
      - A callback spy asserting arguments, call count, or ordering when delegation is the
        documented contract. It does not also prove downstream persistence.
      - A fake child component exposing props or lifecycle events while a real parent
        decides forwarding, mounting, or remounting.
      - A throwing collaborator that drives a real transaction rollback or fallback.
      - A fake clock, scheduler, or barrier that makes real retry/concurrency logic observable.
      - A no-throw test for a documented input that actually invokes the real operation and
        propagates exceptions. A dummy assertion adds no value but does not erase that check.
      - Source, snapshot, or configuration assertions that independently pin required bytes,
        schema, packaging, or release contracts. Exactness alone is not evidence of junk.
      
      Do not classify tests by mock volume, class-name matching, incident IDs, assertion
      counts, or whether they use a database. An incident reference explains intent; the
      assertion still needs to distinguish the regression.
      
      ## Sources and calibration
      
      Google's [test-double guidance](https://testing.googleblog.com/2013/07/testing-on-toilet-know-your-test-doubles.html)
      distinguishes stubs, interaction mocks, and functional fakes. Their roles depend on the
      boundary being tested.
      
      [Testing Library's principles](https://testing-library.com/docs/guiding-principles/)
      favor tests exercised through the interfaces users interact with. Apply that principle
      within the stated test layer rather than demanding every test be end-to-end.
      
      [Stryker's mutation states](https://stryker-mutator.io/docs/mutation-testing-elements/mutant-states-and-metrics/)
      distinguish survivors, uncovered code, and invalid mutants. A mutation result measures
      response to selected changes; it is not a percentage of bad tests. Check equivalent
      mutants and assertion relevance before turning survivors into quality findings.
      
    • detectors.md 2.3 KB
      # Detector diagnostics and blind spots
      
      Read before trusting any detector output. The detectors are discovery aids; their
      results never bound the semantic audit.
      
      ## Exit status and completeness
      
      All four static detectors (weak-negative, duplicate, guard, and junit census) return
      status 2 when an input glob matches no file. The weak-negative and duplicate CLIs also
      return 2 for other known incomplete input and can still print useful partial findings.
      The census also returns 2 on an unreadable report. The reach probe is a pytest plugin,
      not a detector CLI; its report mode does not signal completeness. Inspect diagnostics.
      Run each detector's `--help` for its contract instead of reading the source.
      Duplicate JSON includes `input_complete`, `unmatched_patterns`, `unsupported_files`,
      `unparsed_files`, and `unrecognized_files`.
      
      These describe known file-level omissions; `input_complete: true` does not prove every
      test or syntax form was recognized. A zero-test file may be a helper or unsupported
      syntax. Reconcile omissions with the inventory before proceeding.
      
      ## Known blind spots (require source review)
      
      - An unrelated assertion, a printed exception, or a comment can make the weak-negative
        detector consider an error pinned. Status membership and some unittest assertions
        are not detected. JS/PHP/Rust parsing is heuristic; supported extensions do not imply
        complete framework syntax support.
      - The guard scanner counts textual occurrences, including unused strings. Its unpinned
        total is not a coverage estimate or an estimate of defective tests.
      - JUnit names may lack file paths. The census matches a file's path, name, or stem on
        path-segment boundaries, so `test_user` is not hidden by `test_user_admin`, but the
        name and stem fallbacks still conflate same-name files in different directories.
        Confirm suspected collection gaps with runner output.
      - Duplicate detection does not resolve every fixture, provider, hook, or module mock.
        Same executed lines or body shape do not establish equivalent contracts.
      
      ## Consolidation leads
      
      Treat REDUNDANT/SUBSUMED/FOLD/PARAMETRIZE detector groups as separate consolidation
      leads, not findings. Inspect setup and contract differences before merging. Preserve
      distinct risks across unit, integration, and end-to-end layers. Report deletions and
      LOC only when deletion or consolidation is part of the requested work.
      
    • mutation-and-repair.md 3.1 KB
      # Mutation probes and repairs
      
      Read before running a focused mutation to confirm a finding, before changing any test
      or production code, and before proposing deletion of a production seam. The
      evidence-status labels and scratch-copy rule in `SKILL.md` step 4 still apply.
      
      ## Focused mutation procedure
      
      Work in an isolated scratch copy that preserves the current working-tree contents,
      including relevant uncommitted changes.
      
      1. Run the unchanged selected test and record executed/skipped counts.
      2. Apply one plausible behavior regression to production code in the copy. Before
         writing, assert that the edited token occurs exactly once in the target file.
      3. Inspect the exact diff and verify that the tested import or binary uses that copy.
      4. Run the same test. Confirm its actual assertion result, not only process exit status.
      5. Establish that the mutation changes an in-scope contract on reachable inputs.
      6. Restore the source and verify the baseline again when reusing the copy.
      
      ## Reading the result
      
      A surviving mutant proves a gap only for that mutation and selected tests. A compile
      error, collection error, unrelated failure, timeout, or skipped test is not evidence
      that the intended assertion detects the regression. Equivalent mutations are not gaps.
      
      If feasible, add a scratch-only discriminating assertion and show that baseline passes
      and mutant fails for the intended reason. Label that proof separately from the unchanged
      test's survival. Do not present scratch controls as shipped tests.
      
      Use targeted mutation as discovery as well as confirmation when useful. Do not require
      an expensive whole-suite mutation campaign. Search neighboring coverage before making
      a suite-wide claim; identify stronger existing tests when they already catch the defect.
      
      ## Repairs and authoring
      
      When changes are authorized, repair the claimed protection first. Keep valid boundary
      mocks, no-throw smoke contracts, static contract checks, and distinct test layers.
      Do not replace every mock with infrastructure or assert private details unnecessarily.
      
      For each repair, choose an input and assertion that distinguish the correct behavior
      from the named regression. Validate unchanged-test survival where feasible, then prove
      the repaired test passes on baseline and fails on that regression. Run focused tests
      and repository-required gates for the affected files. Never weaken an assertion or
      regenerate expected output merely to obtain green.
      
      History, absence of production callers, and replacement coverage are required when
      proposing deletion of a production seam. They are not prerequisites for reporting a
      weak assertion. Public APIs, plugins, reflection, and dynamic registration can be live
      without an obvious local caller. No-callers search alone does not prove dead code.
      
      Handle detector consolidation groups per [detectors.md](./detectors.md).
      
      When authoring a new test, answer the same six trace questions in `SKILL.md` step 2. A
      regression test must fail on the defective behavior for the intended reason. A test
      need not prove the entire system to provide independent, useful protection at its
      stated boundary.
      
    • stacks.md 6.5 KB
      # Per-stack commands
      
      Run instrumentation and mutations only in an isolated scratch copy for read-only
      audits. Preserve the user's current source and relevant local changes in that copy.
      The commands below support the semantic review in `SKILL.md`; they do not replace it.
      
      Tool names and flags below were current when written. Confirm each with `--help` or the
      tool's docs before running it, and prefer the repository's pinned wrapper over a global
      binary. Install audit-only tools ad hoc, for example with `uvx`, `npx`, or a scratch
      Composer project. Never add them as project dependencies without approval.
      
      Each section covers three things:
      - **junit**: how to get the report `junit_census.py` reads. Produce it with the gate's
        own command plus the reporter flag, so the census reflects gate conditions.
      - **reach**: how to see which guard actually refused.
      - **mutation**: a tool for confirming one candidate. Scope it to the candidate's
        file; whole-suite runs are too slow to be routine.
      
      ## Python
      
      - **junit:** `pytest --junitxml=/tmp/gate.xml`, appended to the gate's pytest command.
      - **reach, CLI-driven suites:** in the scratch copy, run
        `PYTHONDONTWRITEBYTECODE=1 PYTHONPATH=<this skill's directory>/scripts REACH_PROBE_OUT=/tmp/reach.jsonl
        pytest -p pytest_reach_probe <tests>`, then
        `python3 <this skill's directory>/scripts/pytest_reach_probe.py /tmp/reach.jsonl`.
        The probe runs the suite and patches only `subprocess.Popen.communicate`, so it sees
        `subprocess.run` and wrappers built on `run` or `communicate`. It does not observe
        `subprocess.call`, `check_call`, a bare `Popen(...).wait()`, `os.system`, or asyncio
        subprocesses; silence from the probe is not evidence that no call failed.
      - **reach, in-process code:** temporarily replace `pytest.raises(X)` with
        `pytest.raises(X) as e` and print `e.value`. For a web client, print
        `response.json()` next to the status assertion.
      - **mutation:** `mutmut`, scoped to one module.
      
      ## Rust
      
      - **junit:** `cargo nextest` writes junit when the selected profile sets
        `[profile.<name>.junit] path = "junit.xml"` in the nextest config; with
        `--profile ci` the report lands in `target/nextest/ci/junit.xml`. Plain
        `cargo test` has no junit output. Use `cargo test -- --list` against
        `cargo test -- --list --ignored` to find `#[ignore]` tests. Also check whether the
        gate enables the features that `#[cfg(feature = "...")]` test modules need.
      - **reach:** replace `assert!(r.is_err())` with a `println!` of `r.unwrap_err()` and run
        with `-- --nocapture`. For binaries driven by a Python or shell harness, use that
        harness's reach probe.
      - **mutation:** `cargo mutants`, restricted to the file under suspicion. Its test command
        can be the repository's process-test runner.
      
      ## PHP (PHPUnit, Pest, Laravel)
      
      - **junit:** `vendor/bin/phpunit --log-junit /tmp/gate.xml`; Pest accepts the same
        flag. PHPUnit's junit carries an `assertions` count per test, so `junit_census.py`
        also reports executed tests with zero assertions.
      - **uncollected files:** tests outside the `<testsuite>` directories in `phpunit.xml`,
        or classes whose file name lacks the configured suffix (default `Test.php`), never
        run. Pass `--test-files 'tests/**/*.php'` to the census.
      - **reach:** in a copy of the candidate test, add `$this->withoutExceptionHandling()`
        before the request, so the real exception class and message surface instead of a
        rendered 403, 404, 409, or 422. A 403 can come from a policy, a gate, middleware,
        or a form request's `authorize()`. A 422 can come from any rule. Assert the
        validation key (`assertJsonValidationErrors(['field'])`) or the message.
      - **mutation:** Infection, filtered to the candidate's source file.
      
      ## JavaScript and TypeScript (Vitest, Jest, Playwright)
      
      - **junit:** `vitest run --reporter=junit --outputFile=/tmp/gate.xml`. Jest needs the
        `jest-junit` reporter. Playwright's `--reporter=junit` prints to stdout
        unless `PLAYWRIGHT_JUNIT_OUTPUT_FILE=/tmp/gate.xml` (or the reporter's `outputFile`
        config option) names a file. In a pnpm or other
        workspace, run the census per package, or pass several reports.
      - **uncollected files:** Vitest and Jest `include`/`testMatch` globs, and workspace
        packages whose `test` script is missing or ends in `|| true`, silently drop tests.
      - **reach:** change a bare `.toThrow()` to `.toThrow(/expected message/)`, or
        `console.log` the rejected error. For request tests, log the response body beside
        the status assertion.
      - **skips:** `it.skip`, `describe.skip`, `test.todo`, and `it.skipIf(...)`/`it.runIf(...)`
        whose condition is constant in CI all appear as skipped in junit.
      - **mutation:** StrykerJS, with its mutate glob set to the candidate file.
      
      ## Guard-message extraction hints for `guard_pins.py`
      
      | Stack | Useful `--src` | Useful `--exclude` |
      |---|---|---|
      | Rust | `'src/**/*.rs'` (in-file `#[cfg(test)]` code is counted as tests) | `'^[a-z_]+$'` for identifier-like literals |
      | Python | `--src 'src/**/*.py' --src 'app/**/*.py'` | log-format strings |
      | PHP/Laravel | `'app/**/*.php'` | `'^[\w.]+$'` for translation keys and config paths |
      | JS/TS | `--src 'src/**/*.ts' --src 'apps/*/src/**/*.ts'` | i18n keys |
      
      In Laravel, validation messages mostly come from language files rather than
      `app/**`. For those rules, a key assertion such as `assertJsonValidationErrors` is the
      pin, not the message text.
      
      ## What `duplicate_tests.py` resolves per stack
      
      | Stack | Resolved | Not resolved |
      |---|---|---|
      | Python | fixture parameters, `parametrize` cases (a plain test can match one case), module constants, same-module helper defaults, observation bindings | conftest fixture bodies, imported helpers and constants, `setUp` state |
      | Rust | `let` locals; `#[rstest]`/`#[test_case]` attributes as the data source | helper bodies; each parametrized fn is one unit |
      | PHP | `$variables`; `->assert*` chains split from the request; data provider name and parameters as the data source | `setUp`, `$this->` properties (a stateful body pairs only within its own file), provider rows |
      | JS/TS | `const`/`let`/`var` locals; `it.each` tables as the data source | `beforeEach`, module-level mocks, reordered statements |
      
      The status, polarity, harness, and clique gates apply to every stack. Two refinements are
      Python only: literals that select what to read are not counted as signal, and a test whose
      only calls load fixtures counts as observation-only.
      
      Pairs stay within one file by default. `--cross-file` hunts copy-pasted tests, but a
      copied body usually tests a different copy of the code, so it is not redundant with the
      original.
      
  • scripts
    • duplicate_tests.py 104.5 KB
      #!/usr/bin/env python3
      """Flag likely redundant tests by structure: same action, same or nested assertions, same shape.
      
      Each test is reduced to an action signature (the statements that build input and exercise
      the code, in order) and an assertion set. Groups are candidates, not verdicts: two tests
      with one signature can still guard different contracts through state the detector cannot
      see (a helper's computed default, a fixture's state, an environment variable). Pairs are
      formed within one file unless --cross-file is given, because identical text in two files
      usually exercises two modules (a per-file import, constant, fixture, or setUp).
      
      Kinds:
        REDUNDANT    same action, identical assertion sets: one of the tests adds nothing.
        SUBSUMED     same action, one assertion set a strict subset of another. The subsumer
                     keeps; the first-listed (subsumed) test is the candidate.
        FOLD         same action, different assertion sets: one input, several observables.
                     Fold into one test that asserts all of them.
        PARAMETRIZE  one table-driven test could replace the group. Match `shape`: identical
                     structure once every literal is abstracted; the same action shape where one
                     test adds one assertion shape, or replaces one with another on the same
                     subject; or identical result assertions (at least one beyond an exit or
                     status code) over setups at least --ratio similar with the same final act,
                     whose differing statements call the same callee or act on the same object,
                     plus at most one statement only one test runs. Match `kwarg-superset`:
                     identical calls and identical or nested assertions, but one test also passes
                     keyword arguments; that is a separate table row, or a duplicate when those
                     values are the callee's defaults (the reason names them).
      
      When full assertion sets are unrelated but the assertions after the act are equal or
      nested, REDUNDANT/SUBSUMED is still reported and the reason names the precondition asserts
      (fixture checks before the act) that were set aside.
      
      Gates on FOLD and PARAMETRIZE (every pair in a group must pass; no transitive chaining):
        - The expected status is an observable: exit codes and success/refusal polarity (`== 0`,
          `!= 0`, `is_err`, `assertStatus(422)`, `pytest.raises`) must match. PARAMETRIZE may add
          a row with another status only to a parametrized family whose cases already differ in
          status. FOLD refuses two different expected statuses.
        - PARAMETRIZE refuses assertion shapes that differ only in polarity or strength (`in` /
          `not in`, `==` / `!=`, `contains` / `!contains` / `starts_with`, `is_some` / `is_none`),
          and never pairs cases of two different parametrized families.
        - Harness: a statement shape found in more than half of a file's tests is harness. With no
          shared distinguishing literal (IDF-weighted literal lines of eight characters or more,
          excluding literals that only select what to read, such as `text.index("## Phase 3")`),
          at least 70% of the pair's statements must be non-harness.
        - Observation-only tests (Python: no call that feeds a literal or acts on an object, such
          as a document read and sliced) pair only when their needles overlap: a distinguishing
          shared needle covering at least half of the smaller needle set.
      Groups are cliques of qualifying pairs, strongest first, at most six members; overflow is
      named in the reason. `score` ranks PARAMETRIZE groups within a file by shared literal
      weight: identical distinguishing inputs across differently named tests rank first.
      
      Python (ast):
        - Tests are module-level and class-level `test*` functions. Fixture parameters are part
          of the signature (scoped to the file that defines them, else to its directory), as are
          `usefixtures` marks.
        - `pytest.mark.parametrize` expands into cases (at most 64): argument names are replaced
          by each case's values and `if`/ternary tests that become constant are folded, so a
          plain test can match one case. Cases of one function are siblings and never paired.
          Non-literal argvalues stay symbolic parameters.
        - Module-level constants bound to literals are substituted, so `approvals=NAMED` and the
          same dict literal compare equal. A keyword equal to a same-module helper's default is
          dropped; defaults come from the signature, or from a literal dict the helper merges
          `**kwargs` over (keys the helper assigns again are computed, so unknown). Helpers and
          constants imported from other modules are not resolved.
        - Normalization: assert messages dropped; skip guards (`if ...: pytest.skip()`) dropped;
          single-assignment bindings with no call (`world = fixture`, `p = root / "x"`) inlined;
          observation bindings (values built only from readers such as `json.loads`,
          `read_text`, `exists`) inlined into the assertions that use them, unless an action
          runs between the read and its use (a snapshot stays an action); loops whose body is
          only assertions count as one assertion; remaining locals renamed v0, v1, ... in order
          of first appearance, action statements first.
        - Incidental literals: an integer of two or more digits, or a digit run in a string not
          preceded by a letter or digit, that occurs at least twice in the setup and act (not
          counting assertions) and never as the operand of an exit-code or status comparison, is
          a threaded identifier (a round number, a record id) and becomes ID0, ID1, ... by first
          appearance, everywhere in the test. String path segments joined onto a `tmp*` fixture become TMP0, ....
          All other literals are kept for REDUNDANT/SUBSUMED/FOLD and abstracted for PARAMETRIZE.
        - Assertions: `assert`, calls named `assert*`/`expect*`/`verify*`/`check*` (including
          `self.assertEqual` and mock `assert_called_*`), `pytest.fail`, and `pytest.raises` /
          `pytest.warns` / `assertRaises` context managers. An assertion that exercises code
          (`assert run(...).returncode == 0`, or a non-reader call given a string, bytes, or
          f-string literal input) is also recorded as an action. A call whose only literal
          input is a number (`assert f(10) == 1`) is not, so identical tests built only from
          such assertions are not grouped.
      
      Rust (#[test], #[tokio::test], #[rstest], #[test_case]), PHP (PHPUnit test methods,
      #[Test]/@test, Pest it()/test()), JS/TS (it/test blocks, including .each): token-based.
      Comments and whitespace are stripped; locals are canonicalized (Rust `let` bindings, PHP
      `$variables`, JS `const`/`let`/`var` bindings); statements are split at top-level `;` (and
      line breaks in JS). A statement is an assertion when it starts with `assert*!` (Rust),
      `$this->assert*`/`self::assert*`/`expect(` (PHP), or `expect(`/`assert` (JS); PHP
      `->assert*` and supertest `.expect(` chains are split so the call before them is the action.
      An assertion whose argument calls something other than a known query or reader
      (`assertTrue($this->policy->view($user))`) is also recorded as an action. Data-driven
      tests (PHPUnit data providers and parameters, Pest `->with()`, `it.each` tables, rstest
      cases) carry their data source in the signature. Literals are kept for exact kinds and
      abstracted for PARAMETRIZE. Limits: no fixture, `beforeEach`/`setUp`, helper-default, or
      constant resolution; no case expansion (a dataset or table is one unit); statements are
      compared as token strings, so reordered setup or an extracted helper hides a duplicate.
      """
      
      from __future__ import annotations
      
      import argparse
      import ast
      import difflib
      import glob
      import itertools
      import json
      import math
      import os
      import re
      import sys
      from collections import Counter, defaultdict
      from concurrent.futures import ProcessPoolExecutor
      from dataclasses import dataclass, field
      from pathlib import Path
      
      if sys.version_info < (3, 11):
          print("duplicate_tests: Python 3.11+ required", file=sys.stderr)
          sys.exit(2)
      
      SKIP_DIRS = {
          "node_modules",
          "vendor",
          ".git",
          "target",
          ".venv",
          "venv",
          "__pycache__",
          "dist",
          "build",
          "coverage",
          ".claude",
          ".next",
          ".turbo",
          ".worktrees",
      }
      
      
      def expand(patterns: list[str]) -> list[str]:
          """Glob, dropping dependency, build, and agent-worktree directories."""
          found = {f for p in patterns for f in glob.glob(p, recursive=True)}
          return sorted(f for f in found if Path(f).is_file() and not SKIP_DIRS.intersection(Path(f).parts))
      
      
      @dataclass
      class Unit:
          path: str
          line: int
          func: str
          name: str
          fixtures: tuple[str, ...]
          actions: tuple[str, ...]
          action_pos: tuple[str, ...]
          action_kwargs: tuple[tuple[tuple[str, str], ...], ...]
          pre: frozenset[str]
          post: frozenset[str]
          action_shape: tuple[str, ...]
          assert_shape: tuple[str, ...]
          act_shape: str
          setup_shape: tuple[str, ...] = ()
          params: bool = False
          status: frozenset[str] = frozenset()
          inputs: frozenset[str] = frozenset()
          needles: frozenset[str] = frozenset()
          real_act: bool = True
          notes: list[str] = field(default_factory=list)
      
          @property
          def shape(self) -> tuple[str, ...]:
              return self.action_shape + self.assert_shape
      
          @property
          def full(self) -> frozenset[str]:
              return self.pre | self.post
      
          def ref(self) -> dict[str, object]:
              return {"file": self.path, "line": self.line, "test": self.name}
      
      
      # ---------------------------------------------------------------- observables shared by every language
      
      STRING_LITERAL = re.compile(r"""(?:[bBrRfFuU]{1,2})?("""
                                  r'''"""|\'\'\'|"|'|`)(?:\\.|(?!\1).)*?\1''', re.DOTALL)  # fmt: skip
      STATUS_TEXT = re.compile(
          r"returncode|status_code|exit_code|exitcode|exitCode|statusCode|exit_status|\bstatus\b|\bcode\s*\(\s*\)"
          r"|\brc\b|\.\s*ok\b|\bsuccess\s*\(|\bis_ok\b|\bis_err\b|\bunwrap_err\b|\bOk\s*\(|\bErr\s*\(|\btoThrow"
          r"|\brejects\b|\bresolves\b|[Rr]aises\b|\bexpectException|\bassert(?:Ok|Successful|Created|Accepted|NoContent"
          r"|Status|ExitCode|Failed|Forbidden|NotFound|Unauthorized|Unprocessable\w*|ServerError|BadRequest|Conflict)\b"
      )
      STATUS_OK = re.compile(
          r"\.\s*ok\b|\bsuccess\s*\(|\bis_ok\b|\bOk\s*\(|\bresolves\b|\bassert(?:Ok|Successful|Created|Accepted|NoContent)\b"
      )
      STATUS_ERR = re.compile(
          r"\bis_err\b|\bunwrap_err\b|\bErr\s*\(|\btoThrow|\brejects\b|[Rr]aises\b|\bexpectException"
          r"|\bassert(?:Failed|Forbidden|NotFound|Unauthorized|Unprocessable\w*|ServerError|BadRequest|Conflict)\b"
      )
      STATUS_NUMERIC = re.compile(r"code|status|\brc\b|Status|Some")
      NEGATED = re.compile(r"!=|\bassert_ne\b|\.\s*not\s*\.|\bassertNot|\bnot\b|[(,]\s*!\s*(?=[\w$(])")
      STATUS_INT = re.compile(r"(?<![\w.#$])-?\d+(?![\w.])")
      
      
      CALL_ARGS = re.compile(
          r"(?<![\w$])(?!(?:in|not|and|or|is|if|else|return|await|match)\b)([A-Za-z_$][\w$]*)\s*(!?)\s*\(([^()]*)\)"
      )
      KEEP_ARGS = re.compile(r"^(?:Some|Ok|Err|assert\w*|expect\w*|to[A-Z]\w*|check\w*|verify\w*)$")
      
      
      def drop_call_args(text: str) -> str:
          """Erase the arguments of ordinary calls (`parse(14).returncode`), keeping those of assertion wrappers."""
          for _ in range(8):
              new = CALL_ARGS.sub(
                  lambda m: f"{m[1]}{m[2]}[{m[3]}]" if KEEP_ARGS.match(m[1]) else f"{m[1]}{m[2]}[]", text
              )
              if new == text:
                  break
              text = new
          return text
      
      
      def status_sig(texts) -> frozenset[str]:
          """Expected exit codes and success/refusal polarity: an observable, never an abstractable literal."""
          out = set()
          for text in texts:
              bare = STRING_LITERAL.sub("S", text)
              if not STATUS_TEXT.search(bare):
                  continue
              ints = STATUS_INT.findall(drop_call_args(bare)) if STATUS_NUMERIC.search(bare) else []
              flags = [f for f, rx in (("ok", STATUS_OK), ("err", STATUS_ERR)) if rx.search(bare)]
              if ints or flags:
                  out.add(("!" if NEGATED.search(bare) else "") + ",".join(sorted(ints) + flags))
          return frozenset(out)
      
      
      PLACEHOLDER = re.compile(r"^(?:#?ID\d+#?|TMP\d+)$")
      
      
      def atoms(text: str) -> list[str]:
          """Literal lines worth comparing across tests: one per physical or escaped line, eight characters or
          more; shorter pieces are identifiers and flags (`Bash`, `--attempt`) that name no scenario."""
          found = []
          for piece in re.split(r"\n|\\n", text)[:200]:
              piece = piece.strip()
              if len(piece) >= 8 and not PLACEHOLDER.match(piece):
                  found.append(piece)
          return found
      
      
      POLAR = [
          (re.compile(r"\bnot in\b"), "in"),
          (re.compile(r"\bis not\b"), "is"),
          (re.compile(r"!=="), "==="),
          (re.compile(r"!="), "=="),
          (re.compile(r"\.\s*not\s*(?=\.)"), ""),
          (re.compile(r"\bnot\s+"), ""),
          (re.compile(r"(?<=[(,\s])!\s*(?=[\w$(])"), ""),
          (re.compile(r"\bassert_ne\b"), "assert_eq"),
          (re.compile(r"(?<=[a-z])(?:DoesNot|Not)(?=[A-Z])|\b_?not_|_not\b"), ""),
          (re.compile(r"\bis_(?:none|err|empty|ok|some)\b"), "is_some"),
          (re.compile(r"\bassert(?:False|True)\b"), "assertTrue"),
          (re.compile(r"\btoBe(?:Falsy|Truthy|Null|Undefined|Defined)\b"), "toBeTruthy"),
          (re.compile(r"\b(?:True|False|None|true|false)\b"), "BOOL"),
          (re.compile(r"\b(?:startswith|endswith|starts_with|ends_with|startsWith|endsWith|contains|includes|"
                      r"toContain|toMatch|toBe|toEqual|toStrictEqual|assertStringContainsString|assertStringStartsWith|"
                      r"assertStringEndsWith|assertSame|assertEquals|assertEqual|assertIn)\b"), "REL"),
      ]  # fmt: skip
      REL_FORMS = [
          re.compile(r"^assert _ (?:in|==|is) (.+)$"),
          re.compile(r"^assert (.+?) (?:==|in|is) _$"),
          re.compile(r"^assert (.+?)\.REL\(_\)$"),
          re.compile(r"^assert_eq (?:! )?\( (.+?) , [SN] \)$"),
          re.compile(r"^assert_eq (?:! )?\( [SN] , (.+?) \)$"),
          re.compile(r"^assert (?:! )?\( (.+?) \. REL \( [SN] \) \)$"),
          re.compile(r"^expect \( (.+?) \) \. REL \( [SN] \)$"),
      ]
      
      
      def polar(shape: str) -> str:
          """An assertion shape with polarity and check strength erased: in/not in, ==/!=, contains/starts_with."""
          for rx, repl in POLAR:
              shape = rx.sub(repl, shape)
          for rx in REL_FORMS:
              m = rx.match(shape)
              if m:
                  return f"REL({m.group(1)})"
          return shape
      
      
      CALLEE = re.compile(r"([A-Za-z_$][\w$]*(?:\s*(?:\.|->|::|\?->)\s*[A-Za-z_$][\w$]*)*)\s*!?\s*\(")
      ASSIGNED = re.compile(
          r"^(?:(?:let|const|var|mut)\s+)*[\w$\s,()\[\]*.:>-]*?(?<![=!<>])=(?![=>])\s*(?:await\s+)?"
      )
      
      
      def head(shape: str) -> str:
          """The callee a setup statement runs (`world.write`), else the whole statement."""
          body = ASSIGNED.sub("", shape, count=1).strip()
          m = CALLEE.match(body)
          return re.sub(r"\s+", "", m.group(1)) if m else shape
      
      
      def receiver(shape: str) -> str | None:
          """The object a trailing method call acts on (`(root / _).chmod(_)` -> `(root / _)`), else None."""
          body = ASSIGNED.sub("", shape, count=1).strip().rstrip(";").rstrip()
          if not body.endswith(")"):
              return None
          depth = 0
          for i in range(len(body) - 1, -1, -1):
              depth += {")": 1, "(": -1}.get(body[i], 0)
              if depth == 0:
                  m = re.search(r"(?:\.|->|\?->)\s*[A-Za-z_$][\w$]*\s*$", body[:i])
                  return re.sub(r"\s+", "", body[: m.start()]) if m and m.start() > 0 else None
          return None
      
      
      SUBJECT = [
          re.compile(r"^assert .+? (?:not in|in) (.+)$"),
          re.compile(r"^assert (?:not )?(.+?) (?:==|!=|is not|is) .+$"),
          re.compile(r"^assert (?:not )?(.+?)\.(?:startswith|endswith)\(.*\)$"),
          re.compile(r"^assert (?:! )?\( (?:! )?(.+?) \. (?:contains|starts_with|ends_with) \( .*$"),
          re.compile(r"^assert_(?:eq|ne) ! \( (.+?) , .*$"),
          re.compile(r"^expect \( (.+?) \) \..*$"),
      ]
      
      
      def subject(shape: str) -> str:
          """What an assertion inspects (`v0.stderr` in `assert _ in v0.stderr`), else the whole shape."""
          for rx in SUBJECT:
              m = rx.match(shape)
              if m:
                  return m.group(1)
          return shape
      
      
      # ---------------------------------------------------------------- Python
      
      ASSERT_CALL = re.compile(r"^_*(?:assert|expect|verify|check)\w*$")
      FAIL_OWNERS = {"pytest", "self", "cls", "unittest"}
      RAISES_CTX = re.compile(r"(?:^|\.)(?:raises|warns|deprecated_call|assertRaises\w*|assertWarns\w*)$")
      SKIP_CALL = re.compile(r"(?:^|\.)(?:skip|skipTest|importorskip|xfail)$")
      READER_CALLS = {
          "loads", "load", "read_text", "read_bytes", "exists", "is_file", "is_dir", "is_symlink",
          "stat", "lstat", "json_lines", "readlines", "splitlines", "decode", "encode", "keys",
          "values", "items", "get", "len", "sorted", "list", "dict", "set", "tuple", "str", "bytes",
          "int", "float", "bool", "fspath", "glob", "rglob", "iterdir", "read", "strip", "split",
          "lower", "upper", "startswith", "endswith", "count", "index", "find", "hexdigest",
          "sha256", "any", "all", "sum", "min", "max", "isinstance", "type", "repr", "format",
          "join", "replace", "resolve", "readlink", "listdir", "with_name", "with_suffix", "joinpath",
          "relative_to", "as_posix", "read_json", "getvalue",
      }  # fmt: skip
      READER_PREFIX = re.compile(r"^_*(?:read|load|parse|show|lines|json|manifest|collect)")
      DIGIT_RUN = re.compile(r"(?<![A-Za-z0-9])\d{2,}(?!\d)")
      STATUS_WORDS = re.compile(r"returncode|[Ss]tatus|exit_code|\.code\b|\brc\b")
      
      
      def call_name(node: ast.Call) -> str:
          func = node.func
          if isinstance(func, ast.Name):
              return func.id
          if isinstance(func, ast.Attribute):
              return func.attr
          return ""
      
      
      def dotted(node: ast.AST) -> str:
          return text_of(node)
      
      
      def clone(node):
          """Copy an AST far faster than copy.deepcopy; scalars and context singletons are shared."""
          new = node.__class__.__new__(node.__class__)
          fields = new.__dict__
          for key, value in node.__dict__.items():
              if isinstance(value, list):
                  fields[key] = [clone(item) if isinstance(item, ast.AST) else item for item in value]
              elif isinstance(value, ast.AST) and not isinstance(value, ast.expr_context):
                  fields[key] = clone(value)
              else:
                  fields[key] = value
          return new
      
      
      def walk(root: ast.AST):
          """Pre-order walk in source order; much cheaper than ast.walk."""
          stack = [root]
          while stack:
              node = stack.pop()
              yield node
              for name in reversed(node._fields):
                  value = getattr(node, name, None)
                  if isinstance(value, list):
                      stack.extend(v for v in reversed(value) if isinstance(v, ast.AST))
                  elif isinstance(value, ast.AST) and not isinstance(value, ast.expr_context):
                      stack.append(value)
      
      
      def text_of(node: ast.AST) -> str:
          try:
              return ast.unparse(node)
          except (ValueError, RecursionError):  # 3.11 cannot unparse a backslash inside an f-string part
              return ast.dump(node, annotate_fields=False)
      
      
      class Subst(ast.NodeTransformer):
          def __init__(self, mapping: dict[str, ast.expr]) -> None:
              self.mapping = mapping
      
          def visit_Name(self, node: ast.Name) -> ast.AST:
              if isinstance(node.ctx, ast.Load) and node.id in self.mapping:
                  return clone(self.mapping[node.id])
              return node
      
      
      def literal(node: ast.AST) -> tuple[bool, object]:
          try:
              return True, ast.literal_eval(node)
          except Exception:  # noqa: BLE001 - literal_eval raises several types
              return False, None
      
      
      COMPARE_OPS = {
          ast.Eq: lambda a, b: a == b,
          ast.NotEq: lambda a, b: a != b,
          ast.In: lambda a, b: a in b,
          ast.NotIn: lambda a, b: a not in b,
          ast.Is: lambda a, b: a is b,
          ast.IsNot: lambda a, b: a is not b,
          ast.Lt: lambda a, b: a < b,
          ast.LtE: lambda a, b: a <= b,
          ast.Gt: lambda a, b: a > b,
          ast.GtE: lambda a, b: a >= b,
      }
      
      
      def const_eval(node: ast.AST) -> tuple[bool, object]:
          """Evaluate a test expression that parametrize substitution made constant."""
          ok, value = literal(node)
          if ok:
              return True, value
          if isinstance(node, ast.UnaryOp) and isinstance(node.op, ast.Not):
              ok, value = const_eval(node.operand)
              return ok, (not value) if ok else None
          if isinstance(node, ast.BoolOp):
              values = []
              for item in node.values:
                  ok, value = const_eval(item)
                  if not ok:
                      return False, None
                  values.append(value)
              return True, all(values) if isinstance(node.op, ast.And) else any(values)
          if isinstance(node, ast.Compare):
              ok, left = const_eval(node.left)
              if not ok:
                  return False, None
              for op, right_node in zip(node.ops, node.comparators, strict=True):
                  ok, right = const_eval(right_node)
                  fn = COMPARE_OPS.get(type(op))
                  if not ok or fn is None:
                      return False, None
                  try:
                      if not fn(left, right):
                          return True, False
                  except TypeError:
                      return False, None
                  left = right
              return True, True
          return False, None
      
      
      class FoldIfExp(ast.NodeTransformer):
          def visit_IfExp(self, node: ast.IfExp) -> ast.AST:
              self.generic_visit(node)
              ok, value = const_eval(node.test)
              if ok:
                  return node.body if value else node.orelse
              return node
      
      
      def fold(stmts: list[ast.stmt]) -> list[ast.stmt]:
          out: list[ast.stmt] = []
          for stmt in stmts:
              stmt = FoldIfExp().visit(stmt)
              if isinstance(stmt, ast.If):
                  ok, value = const_eval(stmt.test)
                  if ok:
                      out.extend(fold(stmt.body if value else stmt.orelse))
                      continue
              for name in ("body", "orelse", "finalbody"):
                  if isinstance(getattr(stmt, name, None), list):
                      setattr(stmt, name, fold(getattr(stmt, name)))
              if isinstance(stmt, ast.Try | ast.TryStar):
                  for handler in stmt.handlers:
                      handler.body = fold(handler.body)
              out.append(stmt)
          return out
      
      
      def marker(name: str, *args: ast.expr) -> ast.stmt:
          return ast.Expr(value=ast.Call(func=ast.Name(id=name, ctx=ast.Load()), args=list(args), keywords=[]))
      
      
      def is_assert_stmt(stmt: ast.stmt) -> bool:
          if isinstance(stmt, ast.Assert):
              return True
          if isinstance(stmt, ast.Expr):
              value = stmt.value.value if isinstance(stmt.value, ast.Await) else stmt.value
              return isinstance(value, ast.Call) and is_assert_call(value)
          return False
      
      
      def is_assert_call(node: ast.Call) -> bool:
          """`assert*`/`expect*`/... calls, and `fail()` only as pytest's or unittest's: `world.fail(x)` is setup."""
          name = call_name(node)
          if name == "fail":
              func = node.func
              return isinstance(func, ast.Name) or (
                  isinstance(func, ast.Attribute)
                  and isinstance(func.value, ast.Name)
                  and func.value.id in FAIL_OWNERS
              )
          return bool(ASSERT_CALL.match(name))
      
      
      RESULT_ATTRS = {"returncode", "status_code", "exit_code", "ok", "stdout", "stderr", "combined", "output"}
      
      
      PURE_CALLS = {"match", "fullmatch", "search", "findall", "compile", "sub", "dumps", "approx", "call", "ANY"}
      
      
      def acting(test: ast.AST, wrapper: ast.AST | None = None) -> bool:
          """True when the assertion itself exercises code: `assert run(...).returncode == 0`, or
          `assert f("input") == 1` (a non-reader call given a string literal inside the assertion)."""
          for n in walk(test):
              if (
                  isinstance(n, ast.Attribute)
                  and n.attr in RESULT_ATTRS
                  and isinstance(n.value, ast.Call)
                  and call_name(n.value) not in READER_CALLS
              ):
                  return True
              if isinstance(n, ast.Call) and n is not wrapper:
                  name = call_name(n)
                  if (
                      name in READER_CALLS
                      or name in PURE_CALLS
                      or READER_PREFIX.match(name)
                      or ASSERT_CALL.match(name)
                  ):
                      continue
                  if any(carries_literal(a) for a in [*n.args, *(k.value for k in n.keywords)]):
                      return True
          return False
      
      
      def blob_atoms(text: str) -> list[str]:
          """Atoms of the string literals inside a collapsed container, else of its text."""
          try:
              tree = ast.parse(text, mode="eval")
          except SyntaxError:
              return atoms(text)
          found: list[str] = []
          for n in walk(tree):
              if isinstance(n, ast.Constant) and isinstance(n.value, str | bytes):
                  found += atoms(n.value if isinstance(n.value, str) else n.value.decode("latin-1"))
          return found
      
      
      def carries_literal(node: ast.AST) -> bool:
          """A call argument that supplies input text: a literal or f-string, also nested (`[call("bash", command=...)]`)."""
          return any(
              isinstance(n, ast.JoinedStr)
              or isinstance(n, ast.Constant)
              and isinstance(n.value, str | bytes | Blob)
              for n in walk(node)
          )
      
      
      def selectors(stmt: ast.AST) -> set[int]:
          """Literals that pick what to read (`text.index("### Phase 3:")`, `read(root / "x.md")`): fixture, not signal."""
          chosen: set[int] = set()
          for node in walk(stmt):
              if isinstance(node, ast.Call) and observer(node):
                  picked = [*node.args, *(k.value for k in node.keywords)]
                  if isinstance(node.func, ast.Attribute):
                      picked.append(node.func.value)  # `(root / "x.md").read_text()`
                  for arg in picked:
                      if isinstance(arg, ast.Constant | ast.BinOp):
                          chosen.update(id(n) for n in walk(arg) if isinstance(n, ast.Constant))
          return chosen
      
      
      def observer(node: ast.Call) -> bool:
          """A call that only reads, formats, or asserts: a test made of these alone exercises no code."""
          name = call_name(node)
          return (
              name.startswith("__")
              or name in READER_CALLS
              or name in PURE_CALLS
              or bool(READER_PREFIX.match(name))
              or is_assert_call(node)
          )
      
      
      def is_skip_guard(stmt: ast.stmt) -> bool:
          if not isinstance(stmt, ast.If) or len(stmt.body) != 1:
              return False
          only = stmt.body[0]
          if isinstance(only, ast.Return):
              return True
          return (
              isinstance(only, ast.Expr)
              and isinstance(only.value, ast.Call)
              and bool(SKIP_CALL.search(dotted(only.value.func)))
          )
      
      
      def flatten(stmts: list[ast.stmt]) -> list[tuple[str, ast.stmt]]:
          out: list[tuple[str, ast.stmt]] = []
          for stmt in stmts:
              if isinstance(stmt, ast.Pass) or (
                  isinstance(stmt, ast.Expr) and isinstance(stmt.value, ast.Constant)
              ):
                  continue
              if is_skip_guard(stmt):
                  out.extend(flatten(stmt.orelse))  # type: ignore[attr-defined]
                  continue
              if isinstance(stmt, ast.If):
                  out.append(("action", ast.If(test=stmt.test, body=[ast.Pass()], orelse=[])))
                  out.extend(flatten(stmt.body))
                  if stmt.orelse:
                      out.append(("action", marker("__else__")))
                      out.extend(flatten(stmt.orelse))
                  out.append(("action", marker("__end__")))
              elif isinstance(stmt, ast.For | ast.AsyncFor | ast.While) and all(
                  is_assert_stmt(s) for s in stmt.body
              ):
                  loop = clone(stmt)
                  loop.body = [
                      ast.Assert(test=s.test, msg=None) if isinstance(s, ast.Assert) else s for s in loop.body
                  ]
                  out.append(("assert", loop))
              elif isinstance(stmt, ast.For | ast.AsyncFor):
                  out.append(("action", marker("__for__", stmt.target, stmt.iter)))
                  out.extend(flatten(stmt.body))
                  out.append(("action", marker("__end__")))
              elif isinstance(stmt, ast.While):
                  out.append(("action", marker("__while__", stmt.test)))
                  out.extend(flatten(stmt.body))
                  out.append(("action", marker("__end__")))
              elif isinstance(stmt, ast.With | ast.AsyncWith):
                  others = []
                  for item in stmt.items:
                      ctx = item.context_expr
                      if isinstance(ctx, ast.Call) and RAISES_CTX.search(dotted(ctx.func)):
                          out.append(("assert", ast.Expr(value=ctx)))
                      else:
                          others.append(item)
                  if others:
                      out.append(("action", ast.With(items=others, body=[ast.Pass()], lineno=0, type_comment=None)))
                  out.extend(flatten(stmt.body))
                  if others:
                      out.append(("action", marker("__end__")))
              elif isinstance(stmt, ast.Try | ast.TryStar):
                  out.extend(flatten(stmt.body))
                  for handler in stmt.handlers:
                      out.append(("action", marker("__except__", *([handler.type] if handler.type else []))))
                      out.extend(flatten(handler.body))
                  out.extend(flatten(stmt.orelse))
                  if stmt.finalbody:
                      out.append(("action", marker("__finally__")))
                      out.extend(flatten(stmt.finalbody))
              elif isinstance(stmt, ast.Assert):
                  if acting(stmt.test):
                      copy = ast.Expr(value=stmt.test)
                      copy.acting_copy = True  # type: ignore[attr-defined]
                      out.append(("action", copy))
                  out.append(("assert", ast.Assert(test=stmt.test, msg=None)))
              elif is_assert_stmt(stmt):
                  call = stmt.value.value if isinstance(stmt.value, ast.Await) else stmt.value  # type: ignore[attr-defined]
                  if acting(stmt, call):
                      copy = clone(stmt)
                      copy.acting_copy = True
                      out.append(("action", copy))
                  out.append(("assert", stmt))
              else:
                  out.append(("action", stmt))
          return out
      
      
      def body_names(stmts: list[ast.stmt]) -> tuple[Counter[str], set[str], bool]:
          """Store counts, loaded names, and whether any if/ternary exists, in one pass."""
          counts: Counter[str] = Counter()
          loads: set[str] = set()
          branches = False
          for root in stmts:
              for node in walk(root):
                  if isinstance(node, ast.Name):
                      if isinstance(node.ctx, ast.Load):
                          loads.add(node.id)
                      else:
                          counts[node.id] += 1
                  elif isinstance(node, ast.If | ast.IfExp):
                      branches = True
                  elif isinstance(node, ast.arg):
                      counts[node.arg] += 2
                  elif isinstance(node, ast.ExceptHandler) and node.name:
                      counts[node.name] += 2
                  elif isinstance(node, ast.AugAssign) and isinstance(node.target, ast.Name):
                      counts[node.target.id] += 1
                  elif isinstance(node, ast.FunctionDef | ast.AsyncFunctionDef | ast.ClassDef):
                      counts[node.name] += 2
          return counts, loads, branches
      
      
      def loaded_names(node: ast.AST) -> set[str]:
          return {n.id for n in walk(node) if isinstance(n, ast.Name) and isinstance(n.ctx, ast.Load)}
      
      
      INLINABLE = (
          ast.Name, ast.Attribute, ast.Subscript, ast.BinOp, ast.Constant, ast.JoinedStr,
          ast.FormattedValue, ast.Tuple, ast.UnaryOp, ast.Compare, ast.BoolOp, ast.Slice, ast.Starred,
          ast.operator, ast.unaryop, ast.cmpop, ast.boolop,
      )  # fmt: skip
      
      
      def call_free(node: ast.AST) -> bool:
          return all(isinstance(n, INLINABLE) for n in walk(node))
      
      
      def reader_only(node: ast.AST) -> bool:
          calls = [n for n in walk(node) if isinstance(n, ast.Call)]
          return bool(calls) and all(
              call_name(c) in READER_CALLS or READER_PREFIX.match(call_name(c)) for c in calls
          )
      
      
      def single_target(stmt: ast.stmt) -> str | None:
          if isinstance(stmt, ast.Assign) and len(stmt.targets) == 1 and isinstance(stmt.targets[0], ast.Name):
              return stmt.targets[0].id
          if isinstance(stmt, ast.AnnAssign) and stmt.value is not None and isinstance(stmt.target, ast.Name):
              return stmt.target.id
          return None
      
      
      def inline(flat: list[tuple[str, ast.stmt]], stores: Counter[str]) -> list[tuple[str, ast.stmt]]:
          """Inline call-free single bindings, then observation bindings used only by assertions."""
          mapping: dict[str, ast.expr] = {}
          first: list[tuple[str, ast.stmt, set[str]]] = []
          for kind, stmt in flat:
              loads = loaded_names(stmt)
              if mapping and loads & mapping.keys():
                  stmt = Subst(mapping).visit(stmt)
                  loads = loaded_names(stmt)
              name = single_target(stmt)
              value = getattr(stmt, "value", None)
              if kind == "action" and name and stores[name] == 1 and value is not None and call_free(value):
                  mapping[name] = value
                  continue
              first.append((kind, stmt, loads))
      
          candidates = {
              name
              for kind, stmt, _ in first
              if kind == "action"
              and (name := single_target(stmt))
              and stores[name] == 1
              and reader_only(stmt.value)  # type: ignore[union-attr]
          }
          changed = True
          while changed:
              changed = False
              for kind, stmt, loads in first:
                  if kind == "assert" or single_target(stmt) in candidates:
                      continue
                  used = loads & candidates
                  if used:
                      candidates -= used
                      changed = True
          # A read taken before a later action is a snapshot, not an observation of the result: keep it.
          for index, (kind, stmt, _) in enumerate(first):
              name = single_target(stmt)
              if kind != "action" or name not in candidates:
                  continue
              uses = [i for i, (_, _, loads) in enumerate(first) if i > index and name in loads]
              between = first[index + 1 : max(uses, default=index) + 1]
              if any(k == "action" and single_target(st) not in candidates for k, st, _ in between):
                  candidates.discard(name)
          mapping = {}
          out: list[tuple[str, ast.stmt]] = []
          for kind, stmt, loads in first:
              if mapping and loads & mapping.keys():
                  stmt = Subst(mapping).visit(stmt)
              name = single_target(stmt)
              if kind == "action" and name in candidates:
                  mapping[name] = stmt.value  # type: ignore[union-attr, assignment]
                  continue
              out.append((kind, stmt))
          return out
      
      
      class Blob:
          """A Constant payload that unparses verbatim: a collapsed literal container or a placeholder."""
      
          __slots__ = ("text",)
      
          def __init__(self, text: str) -> None:
              self.text = text
      
          def __repr__(self) -> str:
              return self.text
      
          def __eq__(self, other: object) -> bool:
              return isinstance(other, Blob) and other.text == self.text
      
          def __hash__(self) -> int:
              return hash(self.text)
      
      
      def all_literal(node: ast.AST) -> bool:
          if isinstance(node, ast.Constant):
              return True
          if isinstance(node, ast.List | ast.Tuple | ast.Set):
              return all(all_literal(e) for e in node.elts)
          if isinstance(node, ast.Dict):
              return all(k is not None and all_literal(k) for k in node.keys) and all(
                  all_literal(v) for v in node.values
              )
          if isinstance(node, ast.UnaryOp):
              return all_literal(node.operand)
          return False
      
      
      def replace_children(root: ast.AST, pick) -> ast.AST:
          """Replace, in place, every maximal subexpression for which pick() returns a node."""
          stack = [root]
          while stack:
              node = stack.pop()
              for name in node._fields:
                  value = getattr(node, name, None)
                  if isinstance(value, list):
                      for i, item in enumerate(value):
                          if isinstance(item, ast.expr) and (new := pick(item)) is not None:
                              value[i] = new
                          elif isinstance(item, ast.AST):
                              stack.append(item)
                  elif isinstance(value, ast.expr):
                      if (new := pick(value)) is not None:
                          setattr(node, name, new)
                      else:
                          stack.append(value)
                  elif isinstance(value, ast.AST) and not isinstance(value, ast.expr_context):
                      stack.append(value)
          return root
      
      
      def collapse(root: ast.AST) -> ast.AST:
          """Fold large literal containers into one Constant so later passes visit one node, not hundreds."""
      
          def pick(node: ast.expr) -> ast.expr | None:
              if isinstance(node, ast.List | ast.Tuple | ast.Set | ast.Dict) and all_literal(node):
                  size = sum(1 for _ in walk(node))
                  if size > 12:
                      return ast.Constant(value=Blob(text_of(node)))
              return None
      
          return replace_children(root, pick)
      
      
      UNDERSCORE = ast.Name(id="_", ctx=ast.Load())
      Unparser = getattr(ast, "_Unparser", None)
      
      
      if Unparser is not None:
      
          class ShapeUnparser(Unparser):  # type: ignore[misc, valid-type]
              """Unparse with every literal (and f-string) written as `_`, without copying the tree."""
      
              def traverse(self, node):
                  if isinstance(node, ast.expr) and abstracted(node):
                      self.write("_")
                      return None
                  return super().traverse(node)
      
      
      def literal_expr(node: ast.AST) -> bool:
          """A literal, or an expression built only from literals (`"row\\n" + "x" * 5200`): one table cell."""
          if isinstance(node, ast.BinOp):
              return literal_expr(node.left) and literal_expr(node.right)
          if isinstance(node, ast.UnaryOp):
              return literal_expr(node.operand)
          if isinstance(node, ast.List | ast.Tuple | ast.Set):
              return all(literal_expr(e) for e in node.elts)
          return all_literal(node)
      
      
      def abstracted(node: ast.expr) -> bool:
          """A literal a table row could vary; True/False/None stay, because they carry assertion polarity."""
          if isinstance(node, ast.Constant) and (node.value is None or isinstance(node.value, bool)):
              return False
          return isinstance(node, ast.JoinedStr) or literal_expr(node)
      
      
      def shape_of(stmt: ast.stmt) -> str:
          if Unparser is not None:
              try:
                  return ShapeUnparser().visit(stmt)
              except (ValueError, RecursionError):
                  pass
      
          def pick(node: ast.expr) -> ast.expr | None:
              return UNDERSCORE if abstracted(node) else None
      
          return text_of(replace_children(clone(stmt), pick))
      
      
      def placeholders(flat: list[tuple[str, ast.stmt]], tmp_roots: bool) -> None:
          """Replace tmp path segments and threaded identifiers in place (see module docstring)."""
          if tmp_roots:
              seen: dict[str, str] = {}
              for _, stmt in flat:
                  for node in walk(stmt):
                      if isinstance(node, ast.BinOp) and isinstance(node.op, ast.Div):
                          root = node.left
                          while isinstance(root, ast.BinOp) and isinstance(root.op, ast.Div):
                              root = root.left
                          right = node.right
                          if (
                              isinstance(root, ast.Name)
                              and "tmp" in root.id.lower()
                              and isinstance(right, ast.Constant)
                              and isinstance(right.value, str)
                          ):
                              right.value = Blob(seen.setdefault(right.value, f"TMP{len(seen)}"))
          counts: Counter[str] = Counter()
          compared: set[str] = set()
          for kind, stmt in flat:
              if getattr(stmt, "acting_copy", False):
                  continue
              top = getattr(stmt, "value", None)
              for node in walk(stmt):
                  if isinstance(node, ast.Constant):
                      v = node.value
                      if isinstance(v, int) and not isinstance(v, bool) and abs(v) >= 10:
                          values = [str(abs(v))]
                      elif isinstance(v, str):
                          values = DIGIT_RUN.findall(v)
                      elif isinstance(v, bytes):
                          values = DIGIT_RUN.findall(v.decode("latin-1"))
                      else:
                          continue
                      if kind == "action":
                          counts.update(values)
                  elif kind == "assert" and isinstance(node, ast.Compare | ast.Call):
                      operands = [node.left, *node.comparators] if isinstance(node, ast.Compare) else node.args
                      if isinstance(node, ast.Call) and node is not top:
                          continue
                      if not STATUS_WORDS.search(text_of(node)):
                          continue
                      for operand in operands:
                          if isinstance(operand, ast.Constant) and isinstance(operand.value, int):
                              compared.add(str(abs(operand.value)))
          ids: dict[str, str] = {}
          for value in counts:  # insertion order is first appearance
              if counts[value] >= 2 and value not in compared:
                  ids[value] = f"ID{len(ids)}"
          if not ids:
              return
      
          def sub(text: str) -> str:
              return DIGIT_RUN.sub(lambda m: f"#{ids[m.group()]}#" if m.group() in ids else m.group(), text)
      
          for _, stmt in flat:
              for node in walk(stmt):
                  if isinstance(node, ast.Constant):
                      v = node.value
                      if isinstance(v, int) and not isinstance(v, bool) and str(abs(v)) in ids:
                          node.value = Blob(ids[str(abs(v))])
                      elif isinstance(v, str):
                          node.value = sub(v)
                      elif isinstance(v, bytes):
                          node.value = sub(v.decode("latin-1")).encode("latin-1")
      
      
      def split_kwargs(
          stmt: ast.stmt, text: str, defaults: dict[str, dict[str, object]]
      ) -> tuple[str, tuple[tuple[tuple[str, str], ...], ...]]:
          """Positional text plus per-call keywords; `!key` marks a keyword whose callee default is known."""
          calls = [n for n in walk(stmt) if isinstance(n, ast.Call)]
          if not any(c.keywords for c in calls):
              return text, tuple(() for _ in calls)
          node = clone(stmt)
          found: list[tuple[tuple[str, str], ...]] = []
          stack: list[ast.AST] = [node]
          while stack:
              n = stack.pop()
              if isinstance(n, ast.Call):
                  known = defaults.get(n.func.id, {}) if isinstance(n.func, ast.Name) else {}
                  found.append(
                      tuple(
                          sorted(
                              (f"!{k.arg}" if k.arg in known else k.arg or f"**{i}", text_of(k.value))
                              for i, k in enumerate(n.keywords)
                          )
                      )
                  )
                  n.keywords = []
              stack.extend(reversed([c for c in ast.iter_child_nodes(n) if not isinstance(c, ast.expr_context)]))
          return text_of(node), tuple(found)
      
      
      def module_constants(tree: ast.Module) -> dict[str, ast.expr]:
          counts: Counter[str] = Counter()
          values: dict[str, ast.expr] = {}
          for stmt in tree.body:
              name = single_target(stmt)
              if name:
                  counts[name] += 1
                  value = stmt.value  # type: ignore[union-attr]
                  if value is not None and all_literal(value):
                      values[name] = value
          return {k: v for k, v in values.items() if counts[k] == 1}
      
      
      def module_defaults(tree: ast.Module) -> dict[str, dict[str, object]]:
          """Literal keyword defaults of module-level helpers, so passing a default explicitly is a no-op."""
          found: dict[str, dict[str, object]] = {}
          for stmt in tree.body:
              if isinstance(stmt, ast.FunctionDef | ast.AsyncFunctionDef):
                  args = stmt.args
                  positional = [*args.posonlyargs, *args.args]
                  pairs = list(zip(positional[len(positional) - len(args.defaults) :], args.defaults, strict=True))
                  pairs += [(a, d) for a, d in zip(args.kwonlyargs, args.kw_defaults, strict=True) if d is not None]
                  values = {}
                  for arg, default in pairs:
                      ok, value = literal(default)
                      if ok:
                          values[arg.arg] = value
                  if args.kwarg is not None:
                      values.update(dict_defaults(stmt, args.kwarg.arg))
                  if values:
                      found[stmt.name] = values
          # A helper that forwards **extra to another helper inherits that helper's keyword defaults.
          for _ in range(3):
              for stmt in tree.body:
                  if not isinstance(stmt, ast.FunctionDef | ast.AsyncFunctionDef) or stmt.args.kwarg is None:
                      continue
                  extra = stmt.args.kwarg.arg
                  for node in walk(stmt):
                      if (
                          isinstance(node, ast.Call)
                          and isinstance(node.func, ast.Name)
                          and node.func.id in found
                          and any(
                              k.arg is None and isinstance(k.value, ast.Name) and k.value.id == extra
                              for k in node.keywords
                          )
                      ):
                          inherited = {
                              k: v for k, v in found[node.func.id].items() if k not in found.get(stmt.name, {})
                          }
                          if inherited:
                              found.setdefault(stmt.name, {}).update(inherited)
          return found
      
      
      def dict_defaults(func: ast.FunctionDef | ast.AsyncFunctionDef, extra: str) -> dict[str, object]:
          """Keys of a literal dict that `**extra` is merged over (`body.update(extra)` or `{**d, **extra}`)."""
          merged = any(
              isinstance(n, ast.Call)
              and call_name(n) == "update"
              and any(isinstance(a, ast.Name) and a.id == extra for a in n.args)
              or isinstance(n, ast.Dict)
              and any(
                  k is None and isinstance(v, ast.Name) and v.id == extra
                  for k, v in zip(n.keys, n.values, strict=True)
              )
              for n in walk(func)
          )
          if not merged:
              return {}
          values: dict[str, object] = {}
          seen: Counter[str] = Counter()
          for node in walk(func):
              if isinstance(node, ast.Dict):
                  for key, value in zip(node.keys, node.values, strict=True):
                      if isinstance(key, ast.Constant) and isinstance(key.value, str):
                          seen[key.value] += 1
                          ok, literal_value = literal(value)
                          if ok:
                              values[key.value] = literal_value
              elif (
                  isinstance(node, ast.Subscript)
                  and isinstance(node.ctx, ast.Store)
                  and isinstance(node.slice, ast.Constant)
                  and isinstance(node.slice.value, str)
              ):
                  seen[node.slice.value] += 1
          # A key the helper assigns more than once has a computed default: treat it as unknown.
          return {k: v for k, v in values.items() if seen[k] == 1}
      
      
      def drop_default_kwargs(stmts: list[ast.stmt], defaults: dict[str, dict[str, object]]) -> None:
          for stmt in stmts:
              for node in walk(stmt):
                  if isinstance(node, ast.Call) and isinstance(node.func, ast.Name) and node.func.id in defaults:
                      known = defaults[node.func.id]
                      kept = []
                      for kw in node.keywords:
                          if kw.arg in known:
                              ok, value = literal(kw.value)
                              # Exact type, so True does not stand in for a default of 1.
                              if ok and type(value) is type(known[kw.arg]) and value == known[kw.arg]:  # noqa: E721
                                  continue
                          kept.append(kw)
                      node.keywords = kept
      
      
      def parametrize_cases(
          func: ast.FunctionDef | ast.AsyncFunctionDef, constants: dict[str, ast.expr]
      ) -> tuple[list[tuple[str, dict[str, ast.expr]]], set[str], set[str], bool]:
          """Return (cases, expanded names, symbolic names, parametrized?)."""
          cases: list[tuple[str, dict[str, ast.expr]]] = [("", {})]
          expanded: set[str] = set()
          symbolic: set[str] = set()
          parametrized = False
          for deco in func.decorator_list:
              if not isinstance(deco, ast.Call) or not dotted(deco.func).endswith("parametrize") or not deco.args:
                  continue
              parametrized = True
              names_node = deco.args[0]
              ok, names_value = literal(names_node)
              if not ok:
                  continue
              if isinstance(names_value, str):
                  names = [n.strip() for n in names_value.split(",") if n.strip()]
              else:
                  names = [str(n) for n in names_value]
              values_node = deco.args[1] if len(deco.args) > 1 else None
              for kw in deco.keywords:
                  if kw.arg == "argvalues":
                      values_node = kw.value
              if isinstance(values_node, ast.Name) and values_node.id in constants:
                  values_node = constants[values_node.id]
              ids_node = next((kw.value for kw in deco.keywords if kw.arg == "ids"), None)
              ids_ok, ids_value = literal(ids_node) if ids_node is not None else (False, None)
              if not isinstance(values_node, ast.List | ast.Tuple):
                  symbolic.update(names)
                  continue
              rows: list[tuple[str, dict[str, ast.expr]]] = []
              for index, element in enumerate(values_node.elts):
                  label = None
                  if isinstance(element, ast.Call) and call_name(element) == "param":
                      for kw in element.keywords:
                          if kw.arg == "id":
                              ok, label = literal(kw.value)
                      element = (
                          element.args[0] if len(names) == 1 and element.args else ast.Tuple(elts=element.args)
                      )
                  if len(names) == 1:
                      values = [element]
                  elif isinstance(element, ast.Tuple | ast.List) and len(element.elts) == len(names):
                      values = list(element.elts)
                  else:
                      symbolic.update(names)
                      rows = []
                      break
                  if label is None and ids_ok and isinstance(ids_value, list | tuple) and index < len(ids_value):
                      label = ids_value[index]
                  if label is None:
                      simple = [literal(v) for v in values]
                      label = (
                          "-".join(str(v) for _, v in simple)
                          if all(ok and isinstance(v, str | int | float | bool | None) for ok, v in simple)
                          else str(index)
                      )
                  rows.append((str(label), dict(zip(names, values, strict=True))))
              if not rows:
                  continue
              expanded.update(names)
              cases = [
                  ("-".join(p for p in (a, b) if p), {**ma, **mb})
                  for (a, ma), (b, mb) in itertools.product(cases, rows)
              ][:64]
          return cases, expanded, symbolic, parametrized
      
      
      def py_unit(
          path: str,
          qual: str,
          func: ast.FunctionDef | ast.AsyncFunctionDef,
          base: list[ast.stmt],
          label: str,
          subst: dict[str, ast.expr],
          constants: dict[str, ast.expr],
          defaults: dict[str, dict[str, object]],
          fixtures: tuple[str, ...],
          fixture_keys: tuple[str, ...],
          parametrized: bool,
      ) -> Unit:
          body = [Subst(subst).visit(clone(s)) if subst else clone(s) for s in base]
          stores, used, branches = body_names(body)
          cmap = {k: v for k, v in constants.items() if k in used and k not in stores and k not in fixtures}
          if cmap:
              body = [Subst(cmap).visit(s) for s in body]
          if branches and (subst or cmap):
              body = fold(body)
              stores = body_names(body)[0]
          if defaults:
              drop_default_kwargs(body, defaults)
          flat = inline(flatten(body), stores)
      
          fixture_set = set(fixtures)
          mapping: dict[str, str] = {}
          for _, stmt in sorted(flat, key=lambda item: item[0] == "assert"):
              for node in walk(stmt):
                  if isinstance(node, ast.Name):
                      if node.id in stores and node.id not in fixture_set:
                          node.id = mapping.setdefault(node.id, f"v{len(mapping)}")
                  elif isinstance(node, ast.arg) and node.arg in stores and node.arg not in fixture_set:
                      node.arg = mapping.setdefault(node.arg, f"v{len(mapping)}")
          placeholders(flat, any("tmp" in f.lower() for f in fixtures))
      
          texts = [text_of(stmt) for _, stmt in flat]
          shapes = [shape_of(stmt) for _, stmt in flat]
          inputs: set[str] = set()
          needles: set[str] = set()
          real_act = False
          for kind, stmt in flat:
              bucket = needles if kind == "assert" else inputs
              chosen = selectors(stmt)
              for node in walk(stmt):
                  if isinstance(node, ast.Constant) and id(node) in chosen:
                      continue
                  if isinstance(node, ast.Constant):
                      v = node.value
                      if isinstance(v, str):
                          bucket.update(atoms(v))
                      elif isinstance(v, Blob):
                          bucket.update(blob_atoms(v.text))
                      elif isinstance(v, bytes):
                          bucket.update(atoms(v.decode("latin-1")))
                  # A module helper called with only fixtures (`spine_text(repo_root)`) loads a fixture.
                  elif (
                      kind == "action"
                      and isinstance(node, ast.Call)
                      and not observer(node)
                      and (
                          isinstance(node.func, ast.Attribute)
                          or any(carries_literal(a) for a in [*node.args, *(k.value for k in node.keywords)])
                      )
                  ):
                      real_act = True
          last_assert = max((i for i, (k, _) in enumerate(flat) if k == "assert"), default=-1)
          act = -1
          for i, (kind, _) in enumerate(flat):
              if kind == "action" and (last_assert < 0 or i < last_assert) and not texts[i].startswith("__"):
                  act = i
          actions, pos, kwargs, action_shape = [], [], [], []
          pre, post, assert_shape = set(), set(), []
          for i, (kind, stmt) in enumerate(flat):
              if kind == "assert":
                  (pre if i < act else post).add(texts[i])
                  assert_shape.append(shapes[i])
              else:
                  actions.append(texts[i])
                  p, k = split_kwargs(stmt, texts[i], defaults)
                  pos.append(p)
                  kwargs.extend(k)
                  action_shape.append(shapes[i])
          return Unit(
              path=path,
              line=func.lineno,
              func=f"{path}::{qual}",
              name=f"{qual}[{label}]" if label else qual,
              fixtures=fixture_keys,
              actions=tuple(actions),
              action_pos=tuple(pos),
              action_kwargs=tuple(kwargs),
              pre=frozenset(pre),
              post=frozenset(post),
              action_shape=tuple(action_shape),
              assert_shape=tuple(sorted(assert_shape)),
              act_shape=shapes[act] if act >= 0 else "",
              setup_shape=tuple(shapes[i] for i in range(act + 1) if flat[i][0] == "action"),
              params=parametrized,
              status=status_sig(texts[i] for i, (kind, _) in enumerate(flat) if kind == "assert"),
              inputs=frozenset(inputs),
              needles=frozenset(needles),
              real_act=real_act,
          )
      
      
      def scan_python(path: str, text: str, part: int = 0, parts: int = 1) -> list[Unit]:
          tree = ast.parse(text)
          constants = {k: collapse(ast.Expr(value=clone(v))).value for k, v in module_constants(tree).items()}  # type: ignore[attr-defined]
          defaults = module_defaults(tree)
          local_fixtures = {
              stmt.name
              for stmt in tree.body
              if isinstance(stmt, ast.FunctionDef | ast.AsyncFunctionDef)
              and any("fixture" in dotted(d.func if isinstance(d, ast.Call) else d) for d in stmt.decorator_list)
          }
          funcs: list[tuple[str, ast.FunctionDef | ast.AsyncFunctionDef]] = []
          for stmt in tree.body:
              if isinstance(stmt, ast.FunctionDef | ast.AsyncFunctionDef) and stmt.name.startswith("test"):
                  funcs.append((stmt.name, stmt))
              elif isinstance(stmt, ast.ClassDef):
                  for item in stmt.body:
                      if isinstance(item, ast.FunctionDef | ast.AsyncFunctionDef) and item.name.startswith("test"):
                          funcs.append((f"{stmt.name}::{item.name}", item))
          units: list[Unit] = []
          for qual, func in funcs[part::parts]:
              cases, expanded, symbolic, parametrized = parametrize_cases(func, constants)
              params = [a.arg for a in [*func.args.posonlyargs, *func.args.args, *func.args.kwonlyargs]]
              fixtures = [p for p in params if p not in ("self", "cls") and p not in expanded]
              for deco in func.decorator_list:
                  if isinstance(deco, ast.Call) and dotted(deco.func).endswith("usefixtures"):
                      fixtures += [str(v) for ok, v in map(literal, deco.args) if ok]
              del symbolic  # symbolic names stay ordinary parameters in `fixtures`
              # Fixtures are scoped: one defined in this file, or else in the directory's conftest chain.
              scope = str(Path(path).parent)
              fixture_keys = tuple(sorted(f"{f}@{path if f in local_fixtures else scope}" for f in fixtures))
              base = [collapse(clone(stmt)) for stmt in func.body]
              for label, subst in cases:
                  try:
                      units.append(
                          py_unit(
                              path,
                              qual,
                              func,
                              base,
                              label,
                              subst,
                              constants,
                              defaults,
                              tuple(fixtures),
                              fixture_keys,
                              parametrized,
                          )
                      )
                  except RecursionError as error:
                      raise ValueError(f"{qual}: too deeply nested") from error
          return units
      
      
      # ---------------------------------------------------------------- token languages
      
      PUNCT3 = {"===", "!==", "...", "<=>", "**=", "<<=", ">>=", "?->", "??=", "::<"}
      PUNCT2 = {
          "->", "=>", "::", "==", "!=", "<=", ">=", "&&", "||", "??", "?.", "+=", "-=", "*=", "/=",
          "%=", "++", "--", "<<", ">>", "..", "|=", "&=", "^=", "**",
      }  # fmt: skip
      JS_REGEX_PREV = set("(,=:[!&|?{};+-*%<>~^") | {
          "return",
          "typeof",
          "=>",
          "&&",
          "||",
          "??",
          "==",
          "===",
          "!=",
          "!==",
      }
      
      
      RS_RAW = re.compile(r'b?r(#*)"')
      RS_CHAR = re.compile(r"'(?:\\.[^']*|[^'\\])'")
      RS_LIFETIME = re.compile(r"'\w+")
      PHP_HEREDOC = re.compile(r"<<<\s*['\"]?(\w+)['\"]?")
      PHP_VAR = re.compile(r"\$\w+")
      NUMBER = re.compile(r"0[xXbBoO]\w+|\d[\d_]*(?:\.\d[\d_]*)?(?:[eE][+-]?\d+)?\w*")
      IDENT = re.compile(r"[\w$\\]+")
      RS_IDENT = re.compile(r"\w+")
      
      
      @dataclass
      class Tok:
          kind: str  # id var num str punct comment
          text: str
          line: int
      
      
      def tokenize(text: str, lang: str) -> list[Tok]:
          toks: list[Tok] = []
          i, n, line = 0, len(text), 1
      
          def prev_sig() -> Tok | None:
              for t in reversed(toks):
                  if t.kind != "comment":
                      return t
              return None
      
          while i < n:
              c = text[i]
              if c == "\n":
                  line += 1
                  i += 1
                  continue
              if c.isspace():
                  i += 1
                  continue
              start, start_line = i, line
              two = text[i : i + 2]
              if two == "//" or (lang == "php" and c == "#" and text[i : i + 2] != "#["):
                  end = text.find("\n", i)
                  end = n if end < 0 else end
                  toks.append(Tok("comment", text[i:end], line))
                  i = end
                  continue
              if two == "/*":
                  end = text.find("*/", i + 2)
                  end = n if end < 0 else end + 2
                  toks.append(Tok("comment", text[i:end], line))
                  line += text.count("\n", i, end)
                  i = end
                  continue
              if lang == "rs" and c in "br" and (m := RS_RAW.match(text, i)):
                  closer = '"' + m.group(1)
                  end = text.find(closer, m.end())
                  end = n if end < 0 else end + len(closer)
                  toks.append(Tok("str", text[i:end], line))
                  line += text.count("\n", i, end)
                  i = end
                  continue
              if lang == "php" and text.startswith("<<<", i):
                  m = PHP_HEREDOC.match(text, i)
                  if m:
                      close = re.compile(r"^\s*" + re.escape(m.group(1)) + r"\b", re.MULTILINE)
                      found = close.search(text, m.end())
                      end = n if found is None else found.end()
                      toks.append(Tok("str", text[i:end], line))
                      line += text.count("\n", i, end)
         
    • guard_pins.py 5.9 KB
      #!/usr/bin/env python3
      """List refusal messages in source that no test asserts.
      
      A guard whose message appears in no test is either untested or tested only by a
      status-only negative, which a different guard's refusal would also satisfy. The count
      sizes the problem. Individual rows are candidates: the extraction also picks up warnings
      and format fragments (over-count), and a short generic test needle can look like a pin
      (under-count).
      
      Source literals are taken from lines that look like a refusal (throw, raise, abort,
      Err(, bail!, exit, reject, deny, ValidationException, ...) and the two lines after
      them, so multi-line calls are covered. Test needles are every string literal in the
      test files. In Rust files, everything after the first `#[cfg(test)]` counts as test
      code, not source.
      
      Exit status 2 means known incomplete input: a --src or --tests glob matched no file.
      Output is still printed.
      """
      
      from __future__ import annotations
      
      import argparse
      import glob
      import re
      import sys
      from collections import Counter
      from pathlib import Path
      
      SKIP_DIRS = {
          "node_modules",
          "vendor",
          ".git",
          "target",
          ".venv",
          "venv",
          "__pycache__",
          "dist",
          "build",
          "coverage",
          ".claude",
          ".next",
          ".turbo",
          ".worktrees",
      }
      
      
      def expand(patterns: list[str]) -> list[str]:
          """Glob, dropping dependency, build, and agent-worktree directories."""
          found = {f for p in patterns for f in glob.glob(p, recursive=True)}
          return sorted(f for f in found if Path(f).is_file() and not SKIP_DIRS.intersection(Path(f).parts))
      
      
      GUARD_LINE = re.compile(
          r"\b(?:throw|raise|abort(?:_if|_unless)?|bail!|panic!|Err\(|Failure|fail|reject|deny|refuse|exit|die|"
          r"withMessages|withErrors|ValidationException|HttpException|AuthorizationException|Error\(|report_and_exit)",
          re.IGNORECASE,
      )
      DOUBLE = re.compile(r'"((?:[^"\\\n]|\\.)*)"')
      SINGLE = re.compile(r"'((?:[^'\\\n]|\\.)*)'")
      BACKTICK = re.compile(r"`((?:[^`\\]|\\.)*)`")
      HOLE = re.compile(r"\{[^{}]*\}|%[sdifx]|\$\{[^}]*\}|\$\w+|\{\d*\}")
      
      
      def literals(line: str, suffix: str) -> list[str]:
          found = [m.group(1) for m in DOUBLE.finditer(line)]
          if suffix not in (".rs",):
              found += [m.group(1) for m in SINGLE.finditer(line)]
          if suffix in (".js", ".jsx", ".ts", ".tsx", ".mjs", ".cjs"):
              found += [m.group(1) for m in BACKTICK.finditer(line)]
          return found
      
      
      def split_rust(text: str) -> tuple[str, str]:
          m = re.search(r"^\s*#\[cfg\(test\)\]", text, re.MULTILINE)
          return (text, "") if m is None else (text[: m.start()], text[m.start() :])
      
      
      def main() -> int:
          parser = argparse.ArgumentParser(
              description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter
          )
          parser.add_argument(
              "--src", action="append", required=True, metavar="GLOB", help="production source glob"
          )
          parser.add_argument("--tests", action="append", default=[], metavar="GLOB", help="test file glob")
          parser.add_argument("--min-length", type=int, default=20, help="shortest source literal considered")
          parser.add_argument("--needle-min", type=int, default=12, help="shortest test literal counted as a pin")
          parser.add_argument(
              "--exclude", action="append", default=[], metavar="REGEX", help="skip literals matching"
          )
          parser.add_argument("--limit", type=int, default=60, help="unpinned rows printed")
          args = parser.parse_args()
      
          pattern_files = [(pattern, expand([pattern])) for pattern in [*args.src, *args.tests]]
          unmatched = [pattern for pattern, names in pattern_files if not names]
          for pattern in unmatched:
              print(f"guard_pins: unmatched pattern (no eligible files): {pattern}", file=sys.stderr)
          src_files = expand(args.src)
          test_files = expand(args.tests)
          excludes = [re.compile(x) for x in args.exclude]
      
          needles: set[str] = set()
          guards: list[tuple[str, int, str]] = []
      
          for name in test_files:
              suffix = Path(name).suffix
              for line in Path(name).read_text(errors="replace").splitlines():
                  needles.update(n for n in literals(line, suffix) if len(n) >= args.needle_min)
      
          for name in src_files:
              suffix = Path(name).suffix
              text = Path(name).read_text(errors="replace")
              if suffix == ".rs":
                  text, test_part = split_rust(text)
                  for line in test_part.splitlines():
                      needles.update(n for n in literals(line, suffix) if len(n) >= args.needle_min)
              lines = text.splitlines()
              window = 0
              for number, line in enumerate(lines, 1):
                  if GUARD_LINE.search(line):
                      window = 3
                  if window:
                      window -= 1
                      for literal in literals(line, suffix):
                          if len(literal) >= args.min_length and not any(x.search(literal) for x in excludes):
                              guards.append((name, number, literal))
      
          seen: set[str] = set()
          unpinned: list[tuple[str, int, str]] = []
          for name, number, literal in guards:
              if literal in seen:
                  continue
              seen.add(literal)
              fragments = [f.strip() for f in HOLE.split(literal) if len(f.strip()) >= args.needle_min]
              pinned = any(n in literal or any(n in f for f in fragments) for n in needles)
              if not pinned:
                  unpinned.append((name, number, literal))
      
          per_file = Counter(name for name, _, _ in unpinned)
          print(
              f"{len(seen)} distinct guard literals in {len(src_files)} source files; "
              f"{len(unpinned)} contain no test string literal (>= {args.needle_min} chars)"
          )
          print(f"{len(needles)} test needles from {len(test_files)} test files\n")
          print("== unpinned by file")
          for name, count in per_file.most_common(25):
              print(f"{count:5d}  {name}")
          print(f"\n== first {args.limit} unpinned")
          for name, number, literal in unpinned[: args.limit]:
              print(f"{name}:{number}  {literal[:120]}")
          return 2 if unmatched else 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • junit_census.py 5.7 KB
      #!/usr/bin/env python3
      """Census of a gate run's junit reports: skips by reason, unreported test files, zero-assertion tests.
      
      Reads junit XML written by pytest (--junitxml), PHPUnit (--log-junit), Vitest
      (--reporter=junit), jest-junit, or cargo-nextest. Pass --test-files to list test files
      on disk that no report mentions: those were never collected by the gate. A file counts
      as mentioned only when its path, name, or stem appears between token boundaries, so a
      report naming tests/test_user_admin.py does not also mark tests/test_user.py.
      
      Exit status 2 means known incomplete input: an unreadable report, or a --test-files
      glob that matched no file. Output is still printed when a glob is unmatched.
      """
      
      from __future__ import annotations
      
      import argparse
      import glob
      import re
      import sys
      import xml.etree.ElementTree as ET
      from collections import defaultdict
      from pathlib import Path
      
      SKIP_DIRS = {
          "node_modules",
          "vendor",
          ".git",
          "target",
          ".venv",
          "venv",
          "__pycache__",
          "dist",
          "build",
          "coverage",
          ".claude",
          ".next",
          ".turbo",
          ".worktrees",
      }
      
      
      def expand(patterns: list[str]) -> list[str]:
          """Glob, dropping dependency, build, and agent-worktree directories."""
          found = {f for p in patterns for f in glob.glob(p, recursive=True)}
          return sorted(f for f in found if Path(f).is_file() and not SKIP_DIRS.intersection(Path(f).parts))
      
      
      def normalise(text: str) -> str:
          return text.replace("\\", "/").replace("::", "/").lower()
      
      
      def bounded(candidate: str) -> re.Pattern[str]:
          """Match a path, name, or stem only where it is not part of a longer token."""
          return re.compile(rf"(?<![\w-]){re.escape(candidate)}(?![\w-])")
      
      
      def main() -> int:
          parser = argparse.ArgumentParser(
              description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter
          )
          parser.add_argument("reports", nargs="+", help="junit XML files from the gate run")
          parser.add_argument(
              "--test-files",
              action="append",
              default=[],
              metavar="GLOB",
              help="glob (repeatable, ** allowed) of test files expected to be collected",
          )
          parser.add_argument("--limit", type=int, default=20, help="rows shown per section")
          args = parser.parse_args()
      
          total = skipped = failed = 0
          skips: dict[str, list[str]] = defaultdict(list)
          zero_assertions: list[str] = []
          mentioned: set[str] = set()
          counts_assertions = False
      
          for report in args.reports:
              try:
                  root = ET.parse(report).getroot()
              except (OSError, ET.ParseError) as error:
                  print(f"junit_census: cannot read {report}: {error}; audit input is incomplete", file=sys.stderr)
                  return 2
              for suite in root.iter("testsuite"):
                  for key in ("name", "file"):
                      if suite.get(key):
                          mentioned.add(normalise(suite.get(key, "")))
              for case in root.iter("testcase"):
                  total += 1
                  parts = [case.get(k, "") for k in ("file", "classname", "name")]
                  for part in parts:
                      if part:
                          mentioned.add(normalise(part))
                          mentioned.add(normalise(part.replace(".", "/")))
                  ident = "::".join(p for p in (case.get("classname"), case.get("name")) if p)
                  skip = case.find("skipped")
                  if skip is not None:
                      skipped += 1
                      reason = (skip.get("message") or skip.text or "(no reason)").strip()
                      reason = re.sub(r"\s+", " ", reason)[:200]
                      skips[reason].append(ident)
                  elif case.find("failure") is not None or case.find("error") is not None:
                      failed += 1
                  if case.get("assertions") is not None:
                      counts_assertions = True
                  if case.get("assertions") == "0" and skip is None:
                      zero_assertions.append(ident)
      
          print(f"testcases {total}  skipped {skipped}  failed/errored {failed}")
      
          print(f"\n== skip reasons ({len(skips)} distinct)")
          print("A reason that is constant in the gate environment (build profile, unset env var,")
          print("tool the gate never installs) makes every test under it dead in the gate.")
          for reason, cases in sorted(skips.items(), key=lambda kv: -len(kv[1]))[: args.limit]:
              print(f"{len(cases):5d}  {reason}")
              for ident in cases[:3]:
                  print(f"         e.g. {ident}")
      
          if counts_assertions:
              print(f"\n== executed tests reporting zero assertions ({len(zero_assertions)})")
              for ident in zero_assertions[: args.limit]:
                  print(f"  {ident}")
      
          unmatched: list[str] = []
          if args.test_files:
              pattern_files = [(pattern, expand([pattern])) for pattern in args.test_files]
              files = sorted({name for _, names in pattern_files for name in names})
              unmatched = [pattern for pattern, names in pattern_files if not names]
              for pattern in unmatched:
                  print(f"junit_census: unmatched pattern (no eligible files): {pattern}", file=sys.stderr)
              unseen = []
              for name in files:
                  path = Path(name)
                  stem = normalise(str(path.with_suffix("")))
                  candidates = [bounded(c) for c in {stem, normalise(path.name), normalise(path.stem)}]
                  if not any(c.search(m) for c in candidates for m in mentioned):
                      unseen.append(name)
              print(f"\n== test files on disk never mentioned by any report ({len(unseen)} of {len(files)})")
              print("Check runner include/exclude config, naming conventions, and suite lists.")
              for name in unseen[: args.limit * 5]:
                  print(f"  {name}")
          return 2 if unmatched else 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • pytest_reach_probe.py 4.7 KB
      """Pytest plugin: record every nonzero subprocess result per test.
      
      Answers "which guard actually refused?" for suites that drive a CLI through
      `subprocess` (`run`, or `Popen(...).communicate()`). Each nonzero result becomes one JSON line: test node id, argv[0]
      basename, exit code, and the first and last non-empty stderr lines.
      
          PYTHONDONTWRITEBYTECODE=1 PYTHONPATH=<skill>/scripts REACH_PROBE_OUT=/tmp/reach.jsonl \
              pytest -p pytest_reach_probe <candidate tests>
      
          python3 <skill>/scripts/pytest_reach_probe.py /tmp/reach.jsonl [--grep TEXT]
      
      The second form prints, per test, the LAST nonzero call (usually the one the assertion
      checks) and its last stderr line (the refusal, after any warnings), so each can be compared with the guard the test is
      named for. The plugin only observes; it never changes a result. The output file is
      truncated once per session, so rows from an earlier run never mix in. Without
      REACH_PROBE_OUT it is reach-probe-<pid>.jsonl in the system temp directory, and the path
      is printed in the terminal summary. It works under xdist: the controller truncates and
      picks the path, and each worker appends its own lines. A test module that binds `from subprocess
      import Popen` before the plugin loads still goes through the patched method, because
      the patch is on the class.
      """
      
      from __future__ import annotations
      
      import json
      import os
      import subprocess
      import tempfile
      from pathlib import Path
      
      _current: dict[str, str | None] = {"nodeid": None}
      _original_communicate = subprocess.Popen.communicate
      
      
      def _refusal_lines(data: object) -> tuple[str, str]:
          """First and last non-empty stderr lines; a refusal usually ends the stream after warnings."""
          if isinstance(data, bytes):
              data = data.decode("utf-8", "replace")
          if not isinstance(data, str):
              return "", ""
          lines = [line.strip() for line in data.splitlines() if line.strip()]
          return (lines[0][:400], lines[-1][:400]) if lines else ("", "")
      
      
      def _recording_communicate(self, *args, **kwargs):
          result = _original_communicate(self, *args, **kwargs)
          if self.returncode not in (0, None) and _current["nodeid"] is not None:
              argv = self.args
              first, last = _refusal_lines(result[1] if isinstance(result, tuple) else None)
              head = argv[0] if isinstance(argv, list | tuple) and argv else argv
              row = {
                  "test": _current["nodeid"],
                  "program": os.path.basename(os.fsdecode(head))
                  if isinstance(head, str | bytes | os.PathLike)
                  else str(head),
                  "rc": self.returncode,
                  "stderr_first": first,
                  "stderr_last": last,
              }
              with _output_path().open("a", encoding="utf-8") as handle:
                  handle.write(json.dumps(row) + "\n")
          return result
      
      
      def _output_path() -> Path:
          return Path(os.environ.get("REACH_PROBE_OUT") or Path(tempfile.gettempdir()) / f"reach-probe-{os.getpid()}.jsonl")
      
      
      def pytest_configure(config):
          if not hasattr(config, "workerinput"):
              # xdist workers inherit the environment, so they append to the controller's file.
              os.environ["REACH_PROBE_OUT"] = str(_output_path())
              flags = os.O_WRONLY | os.O_CREAT | os.O_TRUNC | getattr(os, "O_NOFOLLOW", 0)
              os.close(os.open(os.environ["REACH_PROBE_OUT"], flags, 0o600))
          subprocess.Popen.communicate = _recording_communicate
      
      
      def pytest_terminal_summary(terminalreporter, exitstatus, config):
          if not hasattr(config, "workerinput"):
              terminalreporter.write_line(f"reach probe rows: {os.environ.get('REACH_PROBE_OUT')}")
      
      
      def pytest_unconfigure(config):
          subprocess.Popen.communicate = _original_communicate
      
      
      def pytest_runtest_setup(item):
          _current["nodeid"] = item.nodeid
      
      
      def pytest_runtest_teardown(item, nextitem):
          _current["nodeid"] = None
      
      
      def _report() -> int:
          import argparse
      
          parser = argparse.ArgumentParser(description="summarise a reach-probe JSONL file")
          parser.add_argument("jsonl")
          parser.add_argument("--grep", help="only tests whose id contains TEXT")
          args = parser.parse_args()
          last: dict[str, dict] = {}
          calls: dict[str, int] = {}
          for line in Path(args.jsonl).read_text(encoding="utf-8").splitlines():
              row = json.loads(line)
              last[row["test"]] = row
              calls[row["test"]] = calls.get(row["test"], 0) + 1
          for test, row in last.items():
              if args.grep and args.grep not in test:
                  continue
              print(
                  f"{test}\n    rc {row['rc']}  {row['program']}  ({calls[test]} nonzero calls)\n    last:  {row['stderr_last']}"
              )
              if row["stderr_first"] != row["stderr_last"]:
                  print(f"    first: {row['stderr_first']}")
          return 0
      
      
      if __name__ == "__main__":
          raise SystemExit(_report())
      
    • weak_negatives.py 14.5 KB
      #!/usr/bin/env python3
      """Flag negative tests whose assertion cannot tell which guard refused.
      
      Each hit is a candidate for Ineffective for the claim (a test that may pass without
      reaching the guard it names). Confirm hits with a reach probe or a mutation check before acting on them.
      
      Kinds:
        status-only      asserts a nonzero exit code or 4xx/5xx status and nothing that names
                         the refusal (message, error type plus message, validation key, body)
        or-alternative   accepts either of two messages, so the earlier guard satisfies it
        bare-throw       expects a generic exception type (or any error, in JS) with no message
                         or matcher, and nothing else in the test names the refusal
        should-panic     Rust #[should_panic] with no expected = "..."
      
      Languages by extension: .py (AST), .rs, .php (PHPUnit and Pest), .js/.jsx/.ts/.tsx/.mjs/.cjs/.mts/.cts.
      """
      
      from __future__ import annotations
      
      import argparse
      import ast
      import glob
      import re
      import sys
      from collections import Counter
      from pathlib import Path
      
      SKIP_DIRS = {
          "node_modules",
          "vendor",
          ".git",
          "target",
          ".venv",
          "venv",
          "__pycache__",
          "dist",
          "build",
          "coverage",
          ".claude",
          ".next",
          ".turbo",
          ".worktrees",
      }
      
      
      def expand(patterns: list[str]) -> list[str]:
          """Glob, dropping dependency, build, and agent-worktree directories."""
          found = {f for p in patterns for f in glob.glob(p, recursive=True)}
          return sorted(f for f in found if Path(f).is_file() and not SKIP_DIRS.intersection(Path(f).parts))
      
      
      Hit = tuple[str, int, str, str, str]
      
      GENERIC_ERRORS = {
          "Exception",
          "BaseException",
          "RuntimeException",
          "RuntimeError",
          "InvalidArgumentException",
          "LogicException",
          "ValueError",
          "TypeError",
          "KeyError",
          "AssertionError",
          "OSError",
          "SystemExit",
          "CalledProcessError",
          "ValidationException",
          "HttpException",
          "AuthorizationException",
          "AuthenticationException",
          "QueryException",
          "Throwable",
          "Error",
          "ErrorException",
          "UnexpectedValueException",
          "DomainException",
      }
      PY_STATUS_ATTRS = {"returncode", "status_code", "exit_code", "status", "rc", "code"}
      PY_DETAIL_WORDS = re.compile(
          r"stderr|stdout|message|output|match|excinfo|body|error|detail|reason", re.IGNORECASE
      )
      PY_CHECKER_CALL = re.compile(r"assert|expect|refus|reject|deny|denied|fail", re.IGNORECASE)
      
      
      def _is_nonzero_int(node: ast.AST) -> bool:
          return (
              isinstance(node, ast.Constant)
              and isinstance(node.value, int)
              and not isinstance(node.value, bool)
              and node.value != 0
          )
      
      
      def _status_name(node: ast.AST) -> str | None:
          if isinstance(node, ast.Attribute) and node.attr in PY_STATUS_ATTRS:
              return node.attr
          if isinstance(node, ast.Name) and node.id in PY_STATUS_ATTRS:
              return node.id
          return None
      
      
      HTTP_STATUS_ATTRS = {"status_code", "status"}
      SUCCESS_NAME = re.compile(r"(?:^|_)(?:OK|SUCCESS|PASS|CLEAN|ZERO)(?:_|$)")
      
      
      def _is_refusal_value(attr: str, node: ast.AST) -> bool:
          if attr in HTTP_STATUS_ATTRS:
              return isinstance(node, ast.Constant) and isinstance(node.value, int) and 400 <= node.value <= 599
          if isinstance(node, ast.Name | ast.Attribute):
              name = ast.unparse(node).split(".")[-1]
              return name.isupper() and not SUCCESS_NAME.search(name)
          return _is_nonzero_int(node)
      
      
      def _is_zero(node: ast.AST) -> bool:
          return isinstance(node, ast.Constant) and node.value == 0
      
      
      def _is_exit_name(node: ast.AST) -> bool:
          return _status_name(node) not in (None, *HTTP_STATUS_ATTRS)
      
      
      def _py_status_assert(test: ast.AST) -> bool:
          if not isinstance(test, ast.Compare) or len(test.ops) != 1:
              return False
          left, right, op = test.left, test.comparators[0], test.ops[0]
          if isinstance(op, ast.Eq):
              if (attr := _status_name(left)) is not None and _is_refusal_value(attr, right):
                  return True
              return (attr := _status_name(right)) is not None and _is_refusal_value(attr, left)
          if isinstance(op, ast.NotEq):
              return (_is_exit_name(left) and _is_zero(right)) or (_is_exit_name(right) and _is_zero(left))
          return False
      
      
      def _py_detail_checked(func: ast.AST) -> bool:
          for node in ast.walk(func):
              if isinstance(node, ast.Assert) and not _py_status_assert(node.test):
                  test = node.test
                  membership = isinstance(test, ast.Compare) and any(
                      isinstance(o, ast.In | ast.NotIn) for o in test.ops
                  )
                  if PY_DETAIL_WORDS.search(ast.unparse(test)) or (
                      membership and _status_name(test.left) is None and not _is_nonzero_int(test.left)
                  ):
                      return True
              if isinstance(node, ast.Call):
                  name = ast.unparse(node.func)
                  if "raises" in name and any(k.arg == "match" for k in node.keywords):
                      return True
                  if PY_CHECKER_CALL.search(name) and any(
                      isinstance(n, ast.Constant) and isinstance(n.value, str | bytes) and len(n.value) >= 6
                      for arg in [*node.args, *(k.value for k in node.keywords)]
                      for n in ast.walk(arg)
                  ):
                      return True
          return False
      
      
      def scan_python(path: str, text: str) -> list[Hit]:
          tree = ast.parse(text)
          hits: list[Hit] = []
          for func in ast.walk(tree):
              if not isinstance(func, ast.FunctionDef | ast.AsyncFunctionDef) or not func.name.startswith("test"):
                  continue
              status_line = None
              for node in ast.walk(func):
                  if isinstance(node, ast.Assert):
                      if _py_status_assert(node.test) and status_line is None:
                          status_line = node.lineno
                      t = node.test
                      if isinstance(t, ast.BoolOp) and isinstance(t.op, ast.Or):
                          ins = [
                              v
                              for v in t.values
                              if isinstance(v, ast.Compare) and any(isinstance(o, ast.In) for o in v.ops)
                          ]
                          if len(ins) >= 2:
                              hits.append((path, node.lineno, "or-alternative", func.name, ast.unparse(t)[:140]))
                  if isinstance(node, ast.With | ast.AsyncWith):
                      for item in node.items:
                          call = item.context_expr
                          if (
                              isinstance(call, ast.Call)
                              and ast.unparse(call.func).endswith("raises")
                              and not any(k.arg == "match" for k in call.keywords)
                              and call.args
                              and ast.unparse(call.args[0]).split(".")[-1] in GENERIC_ERRORS
                          ):
                              var = item.optional_vars
                              used = var is not None and any(
                                  isinstance(n, ast.Name) and n.id == ast.unparse(var) and n is not var
                                  for stmt in func.body
                                  for n in ast.walk(stmt)
                                  if stmt is not node
                              )
                              if not used:
                                  hits.append((path, node.lineno, "bare-throw", func.name, ast.unparse(call)[:140]))
              if status_line is not None and not _py_detail_checked(func):
                  hits.append(
                      (path, status_line, "status-only", func.name, "status/exit code is the only refusal evidence")
                  )
          return hits
      
      
      def _blocks(text: str, start_re: re.Pattern[str]) -> list[tuple[int, str, str]]:
          starts = [(m.start(), m.group("name")) for m in start_re.finditer(text)]
          out = []
          for i, (offset, name) in enumerate(starts):
              end = starts[i + 1][0] if i + 1 < len(starts) else len(text)
              out.append((text.count("\n", 0, offset) + 1, name, text[offset:end]))
          return out
      
      
      JS_START = re.compile(r"\b(?:it|test)(?:\.(?:only|each\([^)]*\)|concurrent))?\s*\(\s*(['\"`])(?P<name>.+?)\1")
      JS_BARE_THROW = re.compile(r"\.(?:toThrow|toThrowError)\(\s*\)")
      JS_STATUS = re.compile(
          r"expect\([^)]*\.(?:status|statusCode|exitCode|code)\)\s*\.(?:toBe|toEqual|toStrictEqual)\(\s*(?:[45]\d\d|[1-9]\d?)\s*\)"
          r"|toHaveProperty\(\s*['\"](?:status|statusCode)['\"]\s*,\s*[45]\d\d"
          r"|\.expect\(\s*[45]\d\d\s*\)(?!\s*\.\s*expect)"
          r"|toHaveStatus\(\s*[45]\d\d"
      )
      JS_DETAIL = re.compile(
          r"expect\([^)]*(?:body|message|error|text|json|data|detail|stderr|stdout|errors)\b"
          r"|\.(?:toThrow|toThrowError|rejects\.toThrow)\(\s*[^\s)]"
          r"|toMatchObject|toMatchInlineSnapshot|toMatchSnapshot|toContain\(|toMatch\("
          r"|toHaveBeenCalledWith\([^)]*['\"`]"
      )
      JS_OR = re.compile(r"expect\([^;]*\|\|[^;]*\)\s*\.(?:toBe\(\s*true|toBeTruthy)")
      
      
      def scan_js(path: str, text: str) -> list[Hit]:
          hits: list[Hit] = []
          for line, name, body in _blocks(text, JS_START):
              if (m := JS_BARE_THROW.search(body)) and not JS_DETAIL.search(body):
                  hits.append((path, line, "bare-throw", name, m.group(0)))
              if JS_STATUS.search(body) and not JS_DETAIL.search(body):
                  hits.append((path, line, "status-only", name, JS_STATUS.search(body).group(0)[:140]))
              if m := JS_OR.search(body):
                  hits.append((path, line, "or-alternative", name, m.group(0)[:140]))
          return hits
      
      
      PHP_START = re.compile(
          r"function\s+(?P<name>test\w*)\s*\("
          r"|(?:#\[Test\]|@test)[\s\S]{0,200}?function\s+(?P<name2>\w+)\s*\("
          r"|^\s*(?:it|test)\s*\(\s*(['\"])(?P<name3>.+?)\3",
          re.MULTILINE,
      )
      PHP_STATUS = re.compile(
          r"->assert(?:Status\(\s*[45]\d\d|Forbidden|Unauthorized|NotFound|Unprocessable|BadRequest|Conflict|ServerError|MethodNotAllowed)\b"
      )
      PHP_DETAIL = re.compile(
          r"assertJsonValidationErrors|assertInvalid|assertJsonPath|assertJsonFragment|assertJson\(|assertExactJson|assertSee"
          r"|assertSessionHasErrors|expectExceptionMessage|assertStringContainsString|assertJsonStructure|->json\(\s*['\"]"
      )
      PHP_EXPECT = re.compile(r"expectException\(\s*(?P<cls>[\w\\]+)::class")
      
      
      def scan_php(path: str, text: str) -> list[Hit]:
          hits: list[Hit] = []
          starts = []
          for m in PHP_START.finditer(text):
              starts.append((m.start(), m.group("name") or m.group("name2") or m.group("name3")))
          for i, (offset, name) in enumerate(starts):
              end = starts[i + 1][0] if i + 1 < len(starts) else len(text)
              body, line = text[offset:end], text.count("\n", 0, offset) + 1
              if (m := PHP_STATUS.search(body)) and not PHP_DETAIL.search(body):
                  hits.append((path, line, "status-only", name, m.group(0)))
              for m in PHP_EXPECT.finditer(body):
                  if "expectExceptionMessage" not in body and m.group("cls").split("\\")[-1] in GENERIC_ERRORS:
                      hits.append(
                          (path, line, "bare-throw", name, f"expectException({m.group('cls')}) without a message")
                      )
                      break
              for m in re.finditer(r"->(?:throws|toThrow)\(\s*(?P<cls>[\w\\]+)::class\s*\)", body):
                  if m.group("cls").split("\\")[-1] in GENERIC_ERRORS:
                      hits.append(
                          (path, line, "bare-throw", name, f"Pest throws({m.group('cls')}) without a message")
                      )
                      break
          return hits
      
      
      RS_START = re.compile(r"#\[test\](?P<attrs>(?:\s*#\[[^\]]*\])*)\s*(?:async\s+)?fn\s+(?P<name>\w+)")
      RS_IS_ERR = re.compile(r"assert!\(\s*[^;]*?\.is_err\(\)\s*\)|matches!\([^;]*Err\(\s*_\s*\)\s*\)")
      RS_DETAIL = re.compile(
          r"unwrap_err\(\)|\.err\(\)|to_string\(\)|expect_err|Err\(\s*[A-Z]\w*(?:::\w+)+|contains\("
      )
      
      
      def scan_rust(path: str, text: str) -> list[Hit]:
          hits: list[Hit] = []
          starts = [(m.start(), m.group("name"), m.group("attrs")) for m in RS_START.finditer(text)]
          for i, (offset, name, attrs) in enumerate(starts):
              end = starts[i + 1][0] if i + 1 < len(starts) else len(text)
              body, line = text[offset:end], text.count("\n", 0, offset) + 1
              leading = []
              for prior in reversed(text[:offset].splitlines()):
                  if not prior.strip().startswith(("#[", "///")):
                      break
                  leading.append(prior)
              if "#[should_panic]" in attrs or any("#[should_panic]" in p for p in leading):
                  hits.append((path, line, "should-panic", name, "#[should_panic] without expected ="))
              if (m := RS_IS_ERR.search(body)) and not RS_DETAIL.search(body):
                  hits.append((path, line, "status-only", name, m.group(0)[:140]))
          return hits
      
      
      SCANNERS = {
          ".py": scan_python,
          ".rs": scan_rust,
          ".php": scan_php,
          **{ext: scan_js for ext in (".js", ".jsx", ".ts", ".tsx", ".mjs", ".cjs", ".mts", ".cts")},
      }
      
      
      def main() -> int:
          parser = argparse.ArgumentParser(
              description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter
          )
          parser.add_argument("patterns", nargs="+", help="globs of test files (** allowed)")
          parser.add_argument("--kind", action="append", help="only report these kinds")
          args = parser.parse_args()
      
          pattern_files = [(pattern, expand([pattern])) for pattern in args.patterns]
          matched = sorted({name for _, names in pattern_files for name in names})
          unmatched = [pattern for pattern, names in pattern_files if not names]
          for pattern in unmatched:
              print(f"weak_negatives: unmatched pattern (no eligible files): {pattern}", file=sys.stderr)
          files = [name for name in matched if Path(name).suffix in SCANNERS]
          unsupported = [name for name in matched if Path(name).suffix not in SCANNERS]
          if not files:
              print("weak_negatives: no supported files matched; audit input is incomplete", file=sys.stderr)
              return 2
          for name in unsupported:
              print(f"weak_negatives: unsupported file: {name}", file=sys.stderr)
          hits: list[Hit] = []
          failures = 0
          for name in files:
              scanner = SCANNERS[Path(name).suffix]
              try:
                  hits.extend(scanner(name, Path(name).read_text(errors="replace")))
              except (OSError, SyntaxError) as error:
                  failures += 1
                  print(f"weak_negatives: unparsed {name}: {error}", file=sys.stderr)
          if args.kind:
              hits = [h for h in hits if h[2] in args.kind]
          for path, line, kind, test, detail in sorted(hits):
              print(f"{path}:{line}  [{kind}]  {test}  -- {detail}")
          counts = Counter(h[2] for h in hits)
          summary = ", ".join(f"{k} {v}" for k, v in sorted(counts.items())) or "none"
          print(
              f"\n{len(files) - failures} supported files scanned; {failures} unparsed; "
              f"{len(unsupported)} unsupported; {len(unmatched)} unmatched patterns; "
              f"{len(hits)} candidates ({summary})",
              file=sys.stderr,
          )
          return 2 if failures or unsupported or unmatched else 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
  • SKILL.md 8.1 KB
    ---
    name: test-audit
    class: workflow
    description: >-
      Audit whether tests detect regressions in the behavior they claim to protect.
      Find mocked-away subjects, weak or circular assertions, undiscriminating fixtures,
      swallowed failures, and tests missing from gates. Use for test-quality audits,
      suspected false confidence in generated tests, and reviewing new tests. Use
      ia-writing-tests for writing tests.
    ---
    
    # Test audit
    
    Find tests that remain green when their claimed behavior breaks. Assess behavior,
    not apparent authorship, mock count, assertion count, or a target finding percentage.
    Default to a read-only audit. Change tests or production code only when requested.
    
    Adapted from the OpenClaw test-audit skill (MIT).
    
    ## What counts as a finding
    
    Evaluate each claimed behavior separately. A test can protect one obligation and miss
    another. Use these dispositions:
    
    - **Ineffective for the claim:** the named behavior is replaced, never reached, or
      disconnected from a verdict that can fail.
    - **Partially effective:** real behavior is tested, but a specific promised distinction
      is absent from the input or assertions. Preserve the protection that exists.
    - **Not enforced:** the test is uncollected, skipped, or absent from a blocking gate.
      Report local collection and CI enforcement separately from assertion quality.
    - **Redundant:** another test protects the same behavior at the same boundary under
      equivalent inputs and setup. This requires separate evidence; weakness alone is
      not redundancy.
    - **Adequate for its scope:** the test detects a credible failure of its actual contract.
    - **Unresolved:** missing context or execution prevents a defensible decision.
    
    A missing additional edge case is not automatically a defective existing test. Tie
    each quality finding to an existing claim, documented contract, or misleading assertion.
    Distinguish an individual test's weakness from a suite-wide coverage gap.
    
    ## 1. Establish scope and execution
    
    Read repository instructions, runner configuration, shared fixtures, test helpers, and
    CI commands. Record the exact focused command and applicable completion gates.
    Determine which tests are collected, execute, and block a merge. An advisory job is
    not a blocking gate. A local green result does not establish CI enforcement.
    
    Build an inventory from the runner's collection output when practical. Otherwise label
    the source inventory approximate. Count files, test definitions, and parameter cases
    separately. Do not treat helpers as tests or silently omit unrecognized syntax.
    
    For a bounded request, inspect every in-scope test. For a large suite, declare the
    review budget and selection method before discovery. Cover distinct layers, major
    fixtures, and mocking styles. Combine risk-directed selection with a reproducible
    sample independent of detector hits. Keep those two sample results separate. Record
    the seed and sample unit if claiming random sampling. Expand recurring patterns into
    their owner family when practical; report anything left unreviewed.
    
    Do not infer suite prevalence from targeted examples. If a prevalence estimate is
    requested, use a representative sample with an explicit denominator and uncertainty.
    Do not tune findings to an expected rate such as 10–20%.
    
    ## 2. Discover through behavior
    
    Read [behavioral-review.md](./references/behavioral-review.md) before auditing test bodies.
    For each reviewed test, trace:
    
    1. **Claim:** what observable behavior does the name, fixture, or requirement promise?
    2. **Real subject:** which production operation actually runs, including setup, imported
       helpers, autouse fixtures, module mocks, and dependency injection?
    3. **Input:** does the fixture create the distinction the claim depends on?
    4. **Oracle:** where does the expected answer come from, and which assertion checks it?
    5. **Verdict:** will an incorrect result reach the runner as a failure?
    6. **Counterexample:** what plausible wrong implementation would still pass this test?
    
    Construct a concrete counterexample before declaring an assertion sufficient. Examples
    include ignoring one filter, selecting the first item instead of the minimum, dropping
    a forwarded prop, or returning an empty result where only the type is asserted.
    Prefer a changed observable result over a crash or syntax error.
    
    Inspect both positive and negative tests. Trace assertion helpers to their actual
    checks. Read the relevant implementation and neighboring tests before reporting a gap.
    For mocks, identify whether the replaced operation owns the claimed behavior or is a
    collaborator of a real subject. Check arguments, call requirements, returned data, and
    observable effects at that boundary.
    
    When delegation is available, split semantic discovery by owner or directory. Each
    reviewer discovers independently of detector output. Require exact locations, real
    subject, surviving regression, retained value, and evidence status. Independently
    verify the strongest findings before reporting them.
    
    ## 3. Use detectors as additional leads
    
    Run each script by its path in this skill's directory, not the audited repository's
    `scripts/`, with the audited repository root as the working directory. Write output
    to a temporary directory.
    Use the repository's pinned tooling for runner commands. The scripts require Python
    3.11+ and the standard library.
    
    | Script | What its results mean |
    |---|---|
    | [weak_negatives.py](./scripts/weak_negatives.py) `<test globs>` | Heuristic candidates for ambiguous refusal assertions. Does not assess positive behavior or assertion data flow. |
    | [duplicate_tests.py](./scripts/duplicate_tests.py) `<test globs> --json` | Structural duplicate/consolidation candidates, not an effectiveness score. |
    | [junit_census.py](./scripts/junit_census.py) `<report.xml> --test-files <glob>` | Reported skips, possible collection gaps, and PHPUnit zero-assertion leads (adequate when not throwing is the documented contract). |
    | [guard_pins.py](./scripts/guard_pins.py) `--src <glob> --tests <glob>` | Refusal strings without textual matches; not proof of guard coverage. |
    | [pytest_reach_probe.py](./scripts/pytest_reach_probe.py) | Runs the target suite with a recording plugin; use a scratch copy. See [stacks.md](./references/stacks.md). |
    
    Before trusting any output, read [detectors.md](./references/detectors.md): for the
    four static detectors, status 2 means known incomplete input, and each detector has
    blind spots that require source review. Read [stacks.md](./references/stacks.md) for
    runner/probe commands and duplicate parser limits. Confirm tool flags locally before
    use. No detector hits means only no matches to that detector's rules; continue the
    semantic sample.
    
    ## 4. Confirm regression sensitivity
    
    Static evidence can establish a finding when the data flow is explicit, such as a
    test asserting a value it assigned itself. Label it **source-confirmed**. Otherwise
    state the counterexample as **inferred** until executed. Use **execution-confirmed**
    only when the exact test ran against a verified behavioral change.
    
    For high-impact or uncertain findings, use a focused mutation in an isolated scratch
    copy. Do not mutate the user's checkout for an audit. Keep network calls, paid APIs, and
    shared databases outside the probe unless separately authorized and isolated. Before
    running a mutation, read [mutation-and-repair.md](./references/mutation-and-repair.md)
    for the procedure and how to read a surviving mutant.
    
    ## 5. Report actionable evidence
    
    For each finding, give:
    
    - Exact test name and `file:line`, plus the relevant production location.
    - Claimed behavior, actual real subject, and input/assertion/verdict defect.
    - Concrete surviving regression and retained useful protection.
    - Evidence status, command/results when executed, and unresolved limitations.
    - Individual-test versus suite scope, neighboring coverage checked, and repair direction.
    
    Report reviewed scope and selection method alongside findings. Separate confirmed
    findings, unresolved candidates, and retained counterexamples. Do not count every
    missing obligation as a separate bad test. State commands not run and reasons.
    
    Before proposing deletion of a production seam, or when changes are authorized, or when
    authoring a new test, read [mutation-and-repair.md](./references/mutation-and-repair.md).
    
  • SPEC.md 5.3 KB
    # ia-test-audit Specification
    
    ## Intent
    
    `ia-test-audit` is a `workflow`-class skill (a multi-step process producing concrete artifacts). Audit whether tests detect regressions in the behavior they claim to protect. Find mocked-away subjects, weak or circular assertions, undiscriminating fixtures, swallowed failures, and tests missing from gates. Use for test-quality audits, suspected false confidence in generated tests, and reviewing new tests. Use ia-writing-tests for writing tests.
    
    ## Scope
    
    In scope:
    - Behaviors described in `SKILL.md` and routed via the should_trigger phrasings in `distillery/tests/fixtures/triggers/ia-test-audit.jsonl`.
    - Updates to runtime behavior, structure, trigger precision, references, and validation.
    
    Out of scope:
    - Acting as the runtime instructions themselves (those live in `SKILL.md`).
    - Trigger phrasings already covered by adjacent `ia-*` skills (`validate-plugin` flags >70% description overlap as DUPLICATE_TRIGGER).
    - <!-- to fill in: domain-specific exclusions when the skill drifts -->
    
    ## Trigger Context
    
    - Class: `workflow`
    - Hook regex: `plugins/whetstone/hooks/skill-patterns.sh` -> `SKILL_PATTERNS[ia-test-audit]`
    - Common requests (from fixture should_trigger):
      - "audit the tests in the billing module for false confidence"
      - "do these tests actually catch a regression in the discount logic?"
      - "assess the test suite for mocked-away subjects"
    - Should not trigger for (from fixture should_not_trigger):
      - "write tests for the user service"
      - "add test coverage for the auth module"
      - "the test quality is poor, improve it"
    
    ## Source And Evidence Model
    
    Authoritative sources:
    
    - `SKILL.md` -- runtime instructions and reference routing.
    - `references/*.md` -- bundled supplementary content (4 file(s)).
    - `scripts/*.py` -- five bundled stdlib-only detectors and probes (`weak_negatives.py`, `duplicate_tests.py`, `guard_pins.py`, `junit_census.py`, `pytest_reach_probe.py`); require Python 3.11+.
    - `distillery/scripts/test_ia_test_audit_detectors.py` -- pytest CLI tests for the bundled scripts.
    - `distillery/tests/fixtures/triggers/ia-test-audit.jsonl` -- positive and negative trigger phrasings under regression test.
    - `plugins/whetstone/hooks/skill-patterns.sh` -- regex pattern that fires this skill.
    - `distillery/.eval-data/ia-test-audit/` -- harvested session examples (when present).
    
    Data that must not be stored in this skill or its references:
    
    - Secrets, credentials, tokens.
    - Machine-specific filesystem paths (`/home/...`, `/Users/...`, `~/ai/...`). The validator (`MACHINE_PATH_LEAK`) flags these as HIGH.
    - Private URLs, customer data, or unredacted personal information.
    
    ### Coverage matrix
    
    | Dimension | Status | Evidence |
    |---|---|---|
    | Trigger fixtures | complete | distillery/tests/fixtures/triggers/ia-test-audit.jsonl (>=5 should_trigger, >=5 should_not_trigger) |
    | Hook regex pattern | complete | plugins/whetstone/hooks/skill-patterns.sh (`SKILL_PATTERNS[ia-test-audit]`) |
    | Reference architecture | complete | 4 file(s) under references/ |
    | Bundled scripts | partial | 5 scripts under scripts/; CLI tests in distillery/scripts/test_ia_test_audit_detectors.py; next: add CLI tests for guard_pins.py, junit_census.py, and pytest_reach_probe.py |
    | Real-usage signal | <!-- populated by harvest-sessions when sessions exist --> | distillery/.eval-data/ia-test-audit/ (created by harvest-sessions) |
    
    ## Evaluation
    
    Lightweight (run on every change):
    
    ```bash
    python3 distillery/scripts/distiller.py validate-plugin --component ia-test-audit
    python3 distillery/scripts/distiller.py test-triggers --skill ia-test-audit
    python3 -m pytest -q distillery/scripts/test_ia_test_audit_detectors.py
    ```
    
    Deeper (when behavior risk warrants):
    
    ```bash
    python3 distillery/scripts/distiller.py dspy-eval ia-test-audit
    python3 distillery/scripts/distiller.py diagnose-negatives ia-test-audit
    ```
    
    Acceptance gates:
    - `validate-plugin --component ia-test-audit` returns 0 HIGH findings.
    - `test-triggers --skill ia-test-audit` returns F1 = 1.0 with floors of 5 should_trigger and 5 should_not_trigger.
    - For dspy-eval, the composite score does not regress against the most recent saved baseline (see `distillery/.eval-data/ia-test-audit/history.json`).
    
    ## Known Limitations
    
    - The scripts require Python 3.11+ (`duplicate_tests.py` uses `ast.TryStar`); under 3.10 it fails with a traceback, not a version message.
    - `pytest_reach_probe.py` patches only `subprocess.Popen.communicate`; `subprocess.call`, `check_call`, bare `Popen.wait()`, `os.system`, and asyncio subprocesses are not observed. It runs the target suite, so it belongs in a scratch copy.
    - The detectors are lead generators with documented blind spots (`references/detectors.md`); no hits never bounds the semantic audit.
    
    ## Maintenance Notes
    
    - Update `SKILL.md` when the runtime workflow, branch conditions, or output contract changes.
    - Update this `SPEC.md` when intent, scope, evidence model, evaluation gates, or maintenance expectations change.
    - Update the trigger fixture when adding new positive phrasings, removing stale ones, or expanding scope (the 5/5 floor is a hard validator gate).
    - Update the hook regex in `skill-patterns.sh` whenever fixture positives expose a missed phrasing; verify F1 = 1.0 with `eval-triggers` before committing.
    - Run the full release pipeline via `/release` -- never bump versions or update CHANGELOG.md from a per-skill edit.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related