test-audit
Audit whether tests detect regressions in the behavior they claim to protect. Find mocked-away subjects, weak or circular assertions, undiscriminating fixtures, swallowed failures, and tests missing from gates. Use for test-quality audits, suspected false confidence in generated
Install
npx skills add https://github.com/iliaal/ai-skills/tree/master/skills/test-audit
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install iliaal-ai-skills@llmmart
git clone https://github.com/iliaal/ai-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole iliaal/ai-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Test audit
Find tests that remain green when their claimed behavior breaks. Assess behavior, not apparent authorship, mock count, assertion count, or a target finding percentage. Default to a read-only audit. Change tests or production code only when requested.
Adapted from the OpenClaw test-audit skill (MIT).
What counts as a finding
Evaluate each claimed behavior separately. A test can protect one obligation and miss another. Use these dispositions:
- Ineffective for the claim: the named behavior is replaced, never reached, or disconnected from a verdict that can fail.
- Partially effective: real behavior is tested, but a specific promised distinction is absent from the input or assertions. Preserve the protection that exists.
- Not enforced: the test is uncollected, skipped, or absent from a blocking gate. Report local collection and CI enforcement separately from assertion quality.
- Redundant: another test protects the same behavior at the same boundary under equivalent inputs and setup. This requires separate evidence; weakness alone is not redundancy.
- Adequate for its scope: the test detects a credible failure of its actual contract.
- Unresolved: missing context or execution prevents a defensible decision.
A missing additional edge case is not automatically a defective existing test. Tie each quality finding to an existing claim, documented contract, or misleading assertion. Distinguish an individual test's weakness from a suite-wide coverage gap.
1. Establish scope and execution
Read repository instructions, runner configuration, shared fixtures, test helpers, and CI commands. Record the exact focused command and applicable completion gates. Determine which tests are collected, execute, and block a merge. An advisory job is not a blocking gate. A local green result does not establish CI enforcement.
Build an inventory from the runner's collection output when practical. Otherwise label the source inventory approximate. Count files, test definitions, and parameter cases separately. Do not treat helpers as tests or silently omit unrecognized syntax.
For a bounded request, inspect every in-scope test. For a large suite, declare the review budget and selection method before discovery. Cover distinct layers, major fixtures, and mocking styles. Combine risk-directed selection with a reproducible sample independent of detector hits. Keep those two sample results separate. Record the seed and sample unit if claiming random sampling. Expand recurring patterns into their owner family when practical; report anything left unreviewed.
Do not infer suite prevalence from targeted examples. If a prevalence estimate is requested, use a representative sample with an explicit denominator and uncertainty. Do not tune findings to an expected rate such as 10–20%.
2. Discover through behavior
Read behavioral-review.md before auditing test bodies. For each reviewed test, trace:
- Claim: what observable behavior does the name, fixture, or requirement promise?
- Real subject: which production operation actually runs, including setup, imported helpers, autouse fixtures, module mocks, and dependency injection?
- Input: does the fixture create the distinction the claim depends on?
- Oracle: where does the expected answer come from, and which assertion checks it?
- Verdict: will an incorrect result reach the runner as a failure?
- Counterexample: what plausible wrong implementation would still pass this test?
Construct a concrete counterexample before declaring an assertion sufficient. Examples include ignoring one filter, selecting the first item instead of the minimum, dropping a forwarded prop, or returning an empty result where only the type is asserted. Prefer a changed observable result over a crash or syntax error.
Inspect both positive and negative tests. Trace assertion helpers to their actual checks. Read the relevant implementation and neighboring tests before reporting a gap. For mocks, identify whether the replaced operation owns the claimed behavior or is a collaborator of a real subject. Check arguments, call requirements, returned data, and observable effects at that boundary.
When delegation is available, split semantic discovery by owner or directory. Each reviewer discovers independently of detector output. Require exact locations, real subject, surviving regression, retained value, and evidence status. Independently verify the strongest findings before reporting them.
3. Use detectors as additional leads
Run each script by its path in this skill's directory, not the audited repository's
scripts/, with the audited repository root as the working directory. Write output
to a temporary directory.
Use the repository's pinned tooling for runner commands. The scripts require Python
3.11+ and the standard library.
| Script | What its results mean |
|---|---|
weak_negatives.py <test globs> |
Heuristic candidates for ambiguous refusal assertions. Does not assess positive behavior or assertion data flow. |
duplicate_tests.py <test globs> --json |
Structural duplicate/consolidation candidates, not an effectiveness score. |
junit_census.py <report.xml> --test-files <glob> |
Reported skips, possible collection gaps, and PHPUnit zero-assertion leads (adequate when not throwing is the documented contract). |
guard_pins.py --src <glob> --tests <glob> |
Refusal strings without textual matches; not proof of guard coverage. |
| pytest_reach_probe.py | Runs the target suite with a recording plugin; use a scratch copy. See stacks.md. |
Before trusting any output, read detectors.md: for the four static detectors, status 2 means known incomplete input, and each detector has blind spots that require source review. Read stacks.md for runner/probe commands and duplicate parser limits. Confirm tool flags locally before use. No detector hits means only no matches to that detector's rules; continue the semantic sample.
4. Confirm regression sensitivity
Static evidence can establish a finding when the data flow is explicit, such as a test asserting a value it assigned itself. Label it source-confirmed. Otherwise state the counterexample as inferred until executed. Use execution-confirmed only when the exact test ran against a verified behavioral change.
For high-impact or uncertain findings, use a focused mutation in an isolated scratch copy. Do not mutate the user's checkout for an audit. Keep network calls, paid APIs, and shared databases outside the probe unless separately authorized and isolated. Before running a mutation, read mutation-and-repair.md for the procedure and how to read a surviving mutant.
5. Report actionable evidence
For each finding, give:
- Exact test name and
file:line, plus the relevant production location. - Claimed behavior, actual real subject, and input/assertion/verdict defect.
- Concrete surviving regression and retained useful protection.
- Evidence status, command/results when executed, and unresolved limitations.
- Individual-test versus suite scope, neighboring coverage checked, and repair direction.
Report reviewed scope and selection method alongside findings. Separate confirmed findings, unresolved candidates, and retained counterexamples. Do not count every missing obligation as a separate bad test. State commands not run and reasons.
Before proposing deletion of a production seam, or when changes are authorized, or when authoring a new test, read mutation-and-repair.md.
Files (ai-skills)
-
references
-
behavioral-review.md 7.3 KB
# Behavioral review Use this checklist on independently selected tests, including files with no detector hits. Follow fixture and assertion-helper definitions before judging a test body. ## Trace the source of the answer Identify the exact behavior under review. Compare the operation that should compute the answer with the operation the test actually invokes. | Suspect shape | Evidence needed | Useful counterexample or repair | |---|---|---| | Mock supplies the subject's answer | The replaced operation itself owns the claimed transformation, persistence, or decision. No real owner runs. | Change that owner to return an incorrect value. Call the real owner with a fake only at its dependency boundary. | | Fixture constructs the finished result | The test hand-builds the saved record, rendered payload, receipt, or callback order that production should produce. | Stop production from producing that result. Feed the test raw input through the real producer. | | Self-derived expected value | Expected and actual share the same implementation or the expected object is compared to itself. | Break the shared computation. Use an independently specified answer or property that constrains the result. | | Type/truthiness/presence only | The name promises particular values, selection, order, or effects, but any object/non-null/array succeeds. | Return a wrong value of the same type, an empty collection, or omit the claimed effect. Assert meaningful contents. | | Test reimplements the algorithm | Only the test's copy executes or the oracle repeats the same decision with the same failure modes. | Change the production algorithm and trace whether any assertion depends on it. | An independent oracle may compute its answer. A simple mathematical relation or a different reference algorithm is not automatically circular. Shared bugs require a concrete explanation, not merely a visual resemblance between calculations. ## Make competing behaviors produce different answers | Claim | Fixture that hides the regression | Discriminating input | |---|---|---| | Sort or select minimum | Input is already in the expected order. | Entries arrive out of order with distinct keys. | | Compute median | One/two values, or symmetric values where mean equals median. | Skewed values such as 0, 100, 900. | | Apply several filters | Every rejected row also fails another filter. | One excluded row per filter that satisfies all other filters. | | Preserve manual selection | Automatic selection would choose the same item. | Observer input favors a different item while the override is active. | | Reject just one invalid property | Fixture violates multiple earlier checks. | Valid control plus exactly one invalid dimension; inspect the intended refusal. | | Keep no writes/queries/events | Empty inputs, disabled path, fixed-empty fake, or unchanged coarse timestamp. | Reach the live path; observe the effect independently; prove a planted effect is visible. | | All results satisfy a property | Result collection can be empty. | Assert expected population or identities before the universal property. | | Coalesce repeated requests | Only one request or requests that never overlap. | Same-key overlapping requests with controlled scheduling; observe one shared operation. | | Keep independent requests separate | Only one key, so incorrect merging remains invisible. | Distinct keys with independently asserted outcomes. | | Process concurrently | One item or synchronous completion hides serialization. | Multiple items and controlled scheduling that distinguishes concurrent entry from sequential work. | | Handle unknown ID | No existing records, so ignoring the requested ID still returns empty. | Seed a competing record and assert that its values do not appear. | Do not invent unsupported input shapes. Trace producer constraints when reachability is uncertain. A test of one valid equivalence class need not cover every other class; report the missing distinction only within the behavior the test claims. ## Follow failure into the runner - In pytest/unittest, returning `False` does not fail the test. Printed errors and caught exceptions can also finish successfully. Check failure branches for an assertion, rethrow, or runner failure call. - Check whether a prerequisite branch returns normally, skips visibly, or fails. A skip is honest nonexecution but still leaves the behavior unverified. - Check that async work and rejection assertions are awaited or returned to the runner. Confirm assertions in callbacks actually execute before test completion. - Inspect `if`, loops, catch blocks, and helper return values around assertions. An assertion that is never reached provides no protection. - Acceptance of success and failure together, such as status in `(200, 500)`, cannot establish successful service behavior. Identify the exact contract before judging an allowed set of outcomes. - For negatives, inspect the actual error and the path producing it. A bare exception or status can be sufficient when that is the contract and the fixture reaches the relevant operation. Do not require fragile message text universally. - Read matcher semantics. An `any`/`assertSent` predicate returning true for unrelated items lets those items satisfy the check. A `not called` assertion on a detached spy never observes the live operation. ## Keep legitimate doubles and narrow tests Retain these when the real subject makes the asserted decision: - External transport mocked while real code builds and sends an exact payload. - A callback spy asserting arguments, call count, or ordering when delegation is the documented contract. It does not also prove downstream persistence. - A fake child component exposing props or lifecycle events while a real parent decides forwarding, mounting, or remounting. - A throwing collaborator that drives a real transaction rollback or fallback. - A fake clock, scheduler, or barrier that makes real retry/concurrency logic observable. - A no-throw test for a documented input that actually invokes the real operation and propagates exceptions. A dummy assertion adds no value but does not erase that check. - Source, snapshot, or configuration assertions that independently pin required bytes, schema, packaging, or release contracts. Exactness alone is not evidence of junk. Do not classify tests by mock volume, class-name matching, incident IDs, assertion counts, or whether they use a database. An incident reference explains intent; the assertion still needs to distinguish the regression. ## Sources and calibration Google's [test-double guidance](https://testing.googleblog.com/2013/07/testing-on-toilet-know-your-test-doubles.html) distinguishes stubs, interaction mocks, and functional fakes. Their roles depend on the boundary being tested. [Testing Library's principles](https://testing-library.com/docs/guiding-principles/) favor tests exercised through the interfaces users interact with. Apply that principle within the stated test layer rather than demanding every test be end-to-end. [Stryker's mutation states](https://stryker-mutator.io/docs/mutation-testing-elements/mutant-states-and-metrics/) distinguish survivors, uncovered code, and invalid mutants. A mutation result measures response to selected changes; it is not a percentage of bad tests. Check equivalent mutants and assertion relevance before turning survivors into quality findings. -
detectors.md 2.3 KB
# Detector diagnostics and blind spots Read before trusting any detector output. The detectors are discovery aids; their results never bound the semantic audit. ## Exit status and completeness All four static detectors (weak-negative, duplicate, guard, and junit census) return status 2 when an input glob matches no file. The weak-negative and duplicate CLIs also return 2 for other known incomplete input and can still print useful partial findings. The census also returns 2 on an unreadable report. The reach probe is a pytest plugin, not a detector CLI; its report mode does not signal completeness. Inspect diagnostics. Run each detector's `--help` for its contract instead of reading the source. Duplicate JSON includes `input_complete`, `unmatched_patterns`, `unsupported_files`, `unparsed_files`, and `unrecognized_files`. These describe known file-level omissions; `input_complete: true` does not prove every test or syntax form was recognized. A zero-test file may be a helper or unsupported syntax. Reconcile omissions with the inventory before proceeding. ## Known blind spots (require source review) - An unrelated assertion, a printed exception, or a comment can make the weak-negative detector consider an error pinned. Status membership and some unittest assertions are not detected. JS/PHP/Rust parsing is heuristic; supported extensions do not imply complete framework syntax support. - The guard scanner counts textual occurrences, including unused strings. Its unpinned total is not a coverage estimate or an estimate of defective tests. - JUnit names may lack file paths. The census matches a file's path, name, or stem on path-segment boundaries, so `test_user` is not hidden by `test_user_admin`, but the name and stem fallbacks still conflate same-name files in different directories. Confirm suspected collection gaps with runner output. - Duplicate detection does not resolve every fixture, provider, hook, or module mock. Same executed lines or body shape do not establish equivalent contracts. ## Consolidation leads Treat REDUNDANT/SUBSUMED/FOLD/PARAMETRIZE detector groups as separate consolidation leads, not findings. Inspect setup and contract differences before merging. Preserve distinct risks across unit, integration, and end-to-end layers. Report deletions and LOC only when deletion or consolidation is part of the requested work. -
mutation-and-repair.md 3.1 KB
# Mutation probes and repairs Read before running a focused mutation to confirm a finding, before changing any test or production code, and before proposing deletion of a production seam. The evidence-status labels and scratch-copy rule in `SKILL.md` step 4 still apply. ## Focused mutation procedure Work in an isolated scratch copy that preserves the current working-tree contents, including relevant uncommitted changes. 1. Run the unchanged selected test and record executed/skipped counts. 2. Apply one plausible behavior regression to production code in the copy. Before writing, assert that the edited token occurs exactly once in the target file. 3. Inspect the exact diff and verify that the tested import or binary uses that copy. 4. Run the same test. Confirm its actual assertion result, not only process exit status. 5. Establish that the mutation changes an in-scope contract on reachable inputs. 6. Restore the source and verify the baseline again when reusing the copy. ## Reading the result A surviving mutant proves a gap only for that mutation and selected tests. A compile error, collection error, unrelated failure, timeout, or skipped test is not evidence that the intended assertion detects the regression. Equivalent mutations are not gaps. If feasible, add a scratch-only discriminating assertion and show that baseline passes and mutant fails for the intended reason. Label that proof separately from the unchanged test's survival. Do not present scratch controls as shipped tests. Use targeted mutation as discovery as well as confirmation when useful. Do not require an expensive whole-suite mutation campaign. Search neighboring coverage before making a suite-wide claim; identify stronger existing tests when they already catch the defect. ## Repairs and authoring When changes are authorized, repair the claimed protection first. Keep valid boundary mocks, no-throw smoke contracts, static contract checks, and distinct test layers. Do not replace every mock with infrastructure or assert private details unnecessarily. For each repair, choose an input and assertion that distinguish the correct behavior from the named regression. Validate unchanged-test survival where feasible, then prove the repaired test passes on baseline and fails on that regression. Run focused tests and repository-required gates for the affected files. Never weaken an assertion or regenerate expected output merely to obtain green. History, absence of production callers, and replacement coverage are required when proposing deletion of a production seam. They are not prerequisites for reporting a weak assertion. Public APIs, plugins, reflection, and dynamic registration can be live without an obvious local caller. No-callers search alone does not prove dead code. Handle detector consolidation groups per [detectors.md](./detectors.md). When authoring a new test, answer the same six trace questions in `SKILL.md` step 2. A regression test must fail on the defective behavior for the intended reason. A test need not prove the entire system to provide independent, useful protection at its stated boundary. -
stacks.md 6.5 KB
# Per-stack commands Run instrumentation and mutations only in an isolated scratch copy for read-only audits. Preserve the user's current source and relevant local changes in that copy. The commands below support the semantic review in `SKILL.md`; they do not replace it. Tool names and flags below were current when written. Confirm each with `--help` or the tool's docs before running it, and prefer the repository's pinned wrapper over a global binary. Install audit-only tools ad hoc, for example with `uvx`, `npx`, or a scratch Composer project. Never add them as project dependencies without approval. Each section covers three things: - **junit**: how to get the report `junit_census.py` reads. Produce it with the gate's own command plus the reporter flag, so the census reflects gate conditions. - **reach**: how to see which guard actually refused. - **mutation**: a tool for confirming one candidate. Scope it to the candidate's file; whole-suite runs are too slow to be routine. ## Python - **junit:** `pytest --junitxml=/tmp/gate.xml`, appended to the gate's pytest command. - **reach, CLI-driven suites:** in the scratch copy, run `PYTHONDONTWRITEBYTECODE=1 PYTHONPATH=<this skill's directory>/scripts REACH_PROBE_OUT=/tmp/reach.jsonl pytest -p pytest_reach_probe <tests>`, then `python3 <this skill's directory>/scripts/pytest_reach_probe.py /tmp/reach.jsonl`. The probe runs the suite and patches only `subprocess.Popen.communicate`, so it sees `subprocess.run` and wrappers built on `run` or `communicate`. It does not observe `subprocess.call`, `check_call`, a bare `Popen(...).wait()`, `os.system`, or asyncio subprocesses; silence from the probe is not evidence that no call failed. - **reach, in-process code:** temporarily replace `pytest.raises(X)` with `pytest.raises(X) as e` and print `e.value`. For a web client, print `response.json()` next to the status assertion. - **mutation:** `mutmut`, scoped to one module. ## Rust - **junit:** `cargo nextest` writes junit when the selected profile sets `[profile.<name>.junit] path = "junit.xml"` in the nextest config; with `--profile ci` the report lands in `target/nextest/ci/junit.xml`. Plain `cargo test` has no junit output. Use `cargo test -- --list` against `cargo test -- --list --ignored` to find `#[ignore]` tests. Also check whether the gate enables the features that `#[cfg(feature = "...")]` test modules need. - **reach:** replace `assert!(r.is_err())` with a `println!` of `r.unwrap_err()` and run with `-- --nocapture`. For binaries driven by a Python or shell harness, use that harness's reach probe. - **mutation:** `cargo mutants`, restricted to the file under suspicion. Its test command can be the repository's process-test runner. ## PHP (PHPUnit, Pest, Laravel) - **junit:** `vendor/bin/phpunit --log-junit /tmp/gate.xml`; Pest accepts the same flag. PHPUnit's junit carries an `assertions` count per test, so `junit_census.py` also reports executed tests with zero assertions. - **uncollected files:** tests outside the `<testsuite>` directories in `phpunit.xml`, or classes whose file name lacks the configured suffix (default `Test.php`), never run. Pass `--test-files 'tests/**/*.php'` to the census. - **reach:** in a copy of the candidate test, add `$this->withoutExceptionHandling()` before the request, so the real exception class and message surface instead of a rendered 403, 404, 409, or 422. A 403 can come from a policy, a gate, middleware, or a form request's `authorize()`. A 422 can come from any rule. Assert the validation key (`assertJsonValidationErrors(['field'])`) or the message. - **mutation:** Infection, filtered to the candidate's source file. ## JavaScript and TypeScript (Vitest, Jest, Playwright) - **junit:** `vitest run --reporter=junit --outputFile=/tmp/gate.xml`. Jest needs the `jest-junit` reporter. Playwright's `--reporter=junit` prints to stdout unless `PLAYWRIGHT_JUNIT_OUTPUT_FILE=/tmp/gate.xml` (or the reporter's `outputFile` config option) names a file. In a pnpm or other workspace, run the census per package, or pass several reports. - **uncollected files:** Vitest and Jest `include`/`testMatch` globs, and workspace packages whose `test` script is missing or ends in `|| true`, silently drop tests. - **reach:** change a bare `.toThrow()` to `.toThrow(/expected message/)`, or `console.log` the rejected error. For request tests, log the response body beside the status assertion. - **skips:** `it.skip`, `describe.skip`, `test.todo`, and `it.skipIf(...)`/`it.runIf(...)` whose condition is constant in CI all appear as skipped in junit. - **mutation:** StrykerJS, with its mutate glob set to the candidate file. ## Guard-message extraction hints for `guard_pins.py` | Stack | Useful `--src` | Useful `--exclude` | |---|---|---| | Rust | `'src/**/*.rs'` (in-file `#[cfg(test)]` code is counted as tests) | `'^[a-z_]+$'` for identifier-like literals | | Python | `--src 'src/**/*.py' --src 'app/**/*.py'` | log-format strings | | PHP/Laravel | `'app/**/*.php'` | `'^[\w.]+$'` for translation keys and config paths | | JS/TS | `--src 'src/**/*.ts' --src 'apps/*/src/**/*.ts'` | i18n keys | In Laravel, validation messages mostly come from language files rather than `app/**`. For those rules, a key assertion such as `assertJsonValidationErrors` is the pin, not the message text. ## What `duplicate_tests.py` resolves per stack | Stack | Resolved | Not resolved | |---|---|---| | Python | fixture parameters, `parametrize` cases (a plain test can match one case), module constants, same-module helper defaults, observation bindings | conftest fixture bodies, imported helpers and constants, `setUp` state | | Rust | `let` locals; `#[rstest]`/`#[test_case]` attributes as the data source | helper bodies; each parametrized fn is one unit | | PHP | `$variables`; `->assert*` chains split from the request; data provider name and parameters as the data source | `setUp`, `$this->` properties (a stateful body pairs only within its own file), provider rows | | JS/TS | `const`/`let`/`var` locals; `it.each` tables as the data source | `beforeEach`, module-level mocks, reordered statements | The status, polarity, harness, and clique gates apply to every stack. Two refinements are Python only: literals that select what to read are not counted as signal, and a test whose only calls load fixtures counts as observation-only. Pairs stay within one file by default. `--cross-file` hunts copy-pasted tests, but a copied body usually tests a different copy of the code, so it is not redundant with the original.
-
-
scripts
-
duplicate_tests.py 104.5 KB
#!/usr/bin/env python3 """Flag likely redundant tests by structure: same action, same or nested assertions, same shape. Each test is reduced to an action signature (the statements that build input and exercise the code, in order) and an assertion set. Groups are candidates, not verdicts: two tests with one signature can still guard different contracts through state the detector cannot see (a helper's computed default, a fixture's state, an environment variable). Pairs are formed within one file unless --cross-file is given, because identical text in two files usually exercises two modules (a per-file import, constant, fixture, or setUp). Kinds: REDUNDANT same action, identical assertion sets: one of the tests adds nothing. SUBSUMED same action, one assertion set a strict subset of another. The subsumer keeps; the first-listed (subsumed) test is the candidate. FOLD same action, different assertion sets: one input, several observables. Fold into one test that asserts all of them. PARAMETRIZE one table-driven test could replace the group. Match `shape`: identical structure once every literal is abstracted; the same action shape where one test adds one assertion shape, or replaces one with another on the same subject; or identical result assertions (at least one beyond an exit or status code) over setups at least --ratio similar with the same final act, whose differing statements call the same callee or act on the same object, plus at most one statement only one test runs. Match `kwarg-superset`: identical calls and identical or nested assertions, but one test also passes keyword arguments; that is a separate table row, or a duplicate when those values are the callee's defaults (the reason names them). When full assertion sets are unrelated but the assertions after the act are equal or nested, REDUNDANT/SUBSUMED is still reported and the reason names the precondition asserts (fixture checks before the act) that were set aside. Gates on FOLD and PARAMETRIZE (every pair in a group must pass; no transitive chaining): - The expected status is an observable: exit codes and success/refusal polarity (`== 0`, `!= 0`, `is_err`, `assertStatus(422)`, `pytest.raises`) must match. PARAMETRIZE may add a row with another status only to a parametrized family whose cases already differ in status. FOLD refuses two different expected statuses. - PARAMETRIZE refuses assertion shapes that differ only in polarity or strength (`in` / `not in`, `==` / `!=`, `contains` / `!contains` / `starts_with`, `is_some` / `is_none`), and never pairs cases of two different parametrized families. - Harness: a statement shape found in more than half of a file's tests is harness. With no shared distinguishing literal (IDF-weighted literal lines of eight characters or more, excluding literals that only select what to read, such as `text.index("## Phase 3")`), at least 70% of the pair's statements must be non-harness. - Observation-only tests (Python: no call that feeds a literal or acts on an object, such as a document read and sliced) pair only when their needles overlap: a distinguishing shared needle covering at least half of the smaller needle set. Groups are cliques of qualifying pairs, strongest first, at most six members; overflow is named in the reason. `score` ranks PARAMETRIZE groups within a file by shared literal weight: identical distinguishing inputs across differently named tests rank first. Python (ast): - Tests are module-level and class-level `test*` functions. Fixture parameters are part of the signature (scoped to the file that defines them, else to its directory), as are `usefixtures` marks. - `pytest.mark.parametrize` expands into cases (at most 64): argument names are replaced by each case's values and `if`/ternary tests that become constant are folded, so a plain test can match one case. Cases of one function are siblings and never paired. Non-literal argvalues stay symbolic parameters. - Module-level constants bound to literals are substituted, so `approvals=NAMED` and the same dict literal compare equal. A keyword equal to a same-module helper's default is dropped; defaults come from the signature, or from a literal dict the helper merges `**kwargs` over (keys the helper assigns again are computed, so unknown). Helpers and constants imported from other modules are not resolved. - Normalization: assert messages dropped; skip guards (`if ...: pytest.skip()`) dropped; single-assignment bindings with no call (`world = fixture`, `p = root / "x"`) inlined; observation bindings (values built only from readers such as `json.loads`, `read_text`, `exists`) inlined into the assertions that use them, unless an action runs between the read and its use (a snapshot stays an action); loops whose body is only assertions count as one assertion; remaining locals renamed v0, v1, ... in order of first appearance, action statements first. - Incidental literals: an integer of two or more digits, or a digit run in a string not preceded by a letter or digit, that occurs at least twice in the setup and act (not counting assertions) and never as the operand of an exit-code or status comparison, is a threaded identifier (a round number, a record id) and becomes ID0, ID1, ... by first appearance, everywhere in the test. String path segments joined onto a `tmp*` fixture become TMP0, .... All other literals are kept for REDUNDANT/SUBSUMED/FOLD and abstracted for PARAMETRIZE. - Assertions: `assert`, calls named `assert*`/`expect*`/`verify*`/`check*` (including `self.assertEqual` and mock `assert_called_*`), `pytest.fail`, and `pytest.raises` / `pytest.warns` / `assertRaises` context managers. An assertion that exercises code (`assert run(...).returncode == 0`, or a non-reader call given a string, bytes, or f-string literal input) is also recorded as an action. A call whose only literal input is a number (`assert f(10) == 1`) is not, so identical tests built only from such assertions are not grouped. Rust (#[test], #[tokio::test], #[rstest], #[test_case]), PHP (PHPUnit test methods, #[Test]/@test, Pest it()/test()), JS/TS (it/test blocks, including .each): token-based. Comments and whitespace are stripped; locals are canonicalized (Rust `let` bindings, PHP `$variables`, JS `const`/`let`/`var` bindings); statements are split at top-level `;` (and line breaks in JS). A statement is an assertion when it starts with `assert*!` (Rust), `$this->assert*`/`self::assert*`/`expect(` (PHP), or `expect(`/`assert` (JS); PHP `->assert*` and supertest `.expect(` chains are split so the call before them is the action. An assertion whose argument calls something other than a known query or reader (`assertTrue($this->policy->view($user))`) is also recorded as an action. Data-driven tests (PHPUnit data providers and parameters, Pest `->with()`, `it.each` tables, rstest cases) carry their data source in the signature. Literals are kept for exact kinds and abstracted for PARAMETRIZE. Limits: no fixture, `beforeEach`/`setUp`, helper-default, or constant resolution; no case expansion (a dataset or table is one unit); statements are compared as token strings, so reordered setup or an extracted helper hides a duplicate. """ from __future__ import annotations import argparse import ast import difflib import glob import itertools import json import math import os import re import sys from collections import Counter, defaultdict from concurrent.futures import ProcessPoolExecutor from dataclasses import dataclass, field from pathlib import Path if sys.version_info < (3, 11): print("duplicate_tests: Python 3.11+ required", file=sys.stderr) sys.exit(2) SKIP_DIRS = { "node_modules", "vendor", ".git", "target", ".venv", "venv", "__pycache__", "dist", "build", "coverage", ".claude", ".next", ".turbo", ".worktrees", } def expand(patterns: list[str]) -> list[str]: """Glob, dropping dependency, build, and agent-worktree directories.""" found = {f for p in patterns for f in glob.glob(p, recursive=True)} return sorted(f for f in found if Path(f).is_file() and not SKIP_DIRS.intersection(Path(f).parts)) @dataclass class Unit: path: str line: int func: str name: str fixtures: tuple[str, ...] actions: tuple[str, ...] action_pos: tuple[str, ...] action_kwargs: tuple[tuple[tuple[str, str], ...], ...] pre: frozenset[str] post: frozenset[str] action_shape: tuple[str, ...] assert_shape: tuple[str, ...] act_shape: str setup_shape: tuple[str, ...] = () params: bool = False status: frozenset[str] = frozenset() inputs: frozenset[str] = frozenset() needles: frozenset[str] = frozenset() real_act: bool = True notes: list[str] = field(default_factory=list) @property def shape(self) -> tuple[str, ...]: return self.action_shape + self.assert_shape @property def full(self) -> frozenset[str]: return self.pre | self.post def ref(self) -> dict[str, object]: return {"file": self.path, "line": self.line, "test": self.name} # ---------------------------------------------------------------- observables shared by every language STRING_LITERAL = re.compile(r"""(?:[bBrRfFuU]{1,2})?(""" r'''"""|\'\'\'|"|'|`)(?:\\.|(?!\1).)*?\1''', re.DOTALL) # fmt: skip STATUS_TEXT = re.compile( r"returncode|status_code|exit_code|exitcode|exitCode|statusCode|exit_status|\bstatus\b|\bcode\s*\(\s*\)" r"|\brc\b|\.\s*ok\b|\bsuccess\s*\(|\bis_ok\b|\bis_err\b|\bunwrap_err\b|\bOk\s*\(|\bErr\s*\(|\btoThrow" r"|\brejects\b|\bresolves\b|[Rr]aises\b|\bexpectException|\bassert(?:Ok|Successful|Created|Accepted|NoContent" r"|Status|ExitCode|Failed|Forbidden|NotFound|Unauthorized|Unprocessable\w*|ServerError|BadRequest|Conflict)\b" ) STATUS_OK = re.compile( r"\.\s*ok\b|\bsuccess\s*\(|\bis_ok\b|\bOk\s*\(|\bresolves\b|\bassert(?:Ok|Successful|Created|Accepted|NoContent)\b" ) STATUS_ERR = re.compile( r"\bis_err\b|\bunwrap_err\b|\bErr\s*\(|\btoThrow|\brejects\b|[Rr]aises\b|\bexpectException" r"|\bassert(?:Failed|Forbidden|NotFound|Unauthorized|Unprocessable\w*|ServerError|BadRequest|Conflict)\b" ) STATUS_NUMERIC = re.compile(r"code|status|\brc\b|Status|Some") NEGATED = re.compile(r"!=|\bassert_ne\b|\.\s*not\s*\.|\bassertNot|\bnot\b|[(,]\s*!\s*(?=[\w$(])") STATUS_INT = re.compile(r"(?<![\w.#$])-?\d+(?![\w.])") CALL_ARGS = re.compile( r"(?<![\w$])(?!(?:in|not|and|or|is|if|else|return|await|match)\b)([A-Za-z_$][\w$]*)\s*(!?)\s*\(([^()]*)\)" ) KEEP_ARGS = re.compile(r"^(?:Some|Ok|Err|assert\w*|expect\w*|to[A-Z]\w*|check\w*|verify\w*)$") def drop_call_args(text: str) -> str: """Erase the arguments of ordinary calls (`parse(14).returncode`), keeping those of assertion wrappers.""" for _ in range(8): new = CALL_ARGS.sub( lambda m: f"{m[1]}{m[2]}[{m[3]}]" if KEEP_ARGS.match(m[1]) else f"{m[1]}{m[2]}[]", text ) if new == text: break text = new return text def status_sig(texts) -> frozenset[str]: """Expected exit codes and success/refusal polarity: an observable, never an abstractable literal.""" out = set() for text in texts: bare = STRING_LITERAL.sub("S", text) if not STATUS_TEXT.search(bare): continue ints = STATUS_INT.findall(drop_call_args(bare)) if STATUS_NUMERIC.search(bare) else [] flags = [f for f, rx in (("ok", STATUS_OK), ("err", STATUS_ERR)) if rx.search(bare)] if ints or flags: out.add(("!" if NEGATED.search(bare) else "") + ",".join(sorted(ints) + flags)) return frozenset(out) PLACEHOLDER = re.compile(r"^(?:#?ID\d+#?|TMP\d+)$") def atoms(text: str) -> list[str]: """Literal lines worth comparing across tests: one per physical or escaped line, eight characters or more; shorter pieces are identifiers and flags (`Bash`, `--attempt`) that name no scenario.""" found = [] for piece in re.split(r"\n|\\n", text)[:200]: piece = piece.strip() if len(piece) >= 8 and not PLACEHOLDER.match(piece): found.append(piece) return found POLAR = [ (re.compile(r"\bnot in\b"), "in"), (re.compile(r"\bis not\b"), "is"), (re.compile(r"!=="), "==="), (re.compile(r"!="), "=="), (re.compile(r"\.\s*not\s*(?=\.)"), ""), (re.compile(r"\bnot\s+"), ""), (re.compile(r"(?<=[(,\s])!\s*(?=[\w$(])"), ""), (re.compile(r"\bassert_ne\b"), "assert_eq"), (re.compile(r"(?<=[a-z])(?:DoesNot|Not)(?=[A-Z])|\b_?not_|_not\b"), ""), (re.compile(r"\bis_(?:none|err|empty|ok|some)\b"), "is_some"), (re.compile(r"\bassert(?:False|True)\b"), "assertTrue"), (re.compile(r"\btoBe(?:Falsy|Truthy|Null|Undefined|Defined)\b"), "toBeTruthy"), (re.compile(r"\b(?:True|False|None|true|false)\b"), "BOOL"), (re.compile(r"\b(?:startswith|endswith|starts_with|ends_with|startsWith|endsWith|contains|includes|" r"toContain|toMatch|toBe|toEqual|toStrictEqual|assertStringContainsString|assertStringStartsWith|" r"assertStringEndsWith|assertSame|assertEquals|assertEqual|assertIn)\b"), "REL"), ] # fmt: skip REL_FORMS = [ re.compile(r"^assert _ (?:in|==|is) (.+)$"), re.compile(r"^assert (.+?) (?:==|in|is) _$"), re.compile(r"^assert (.+?)\.REL\(_\)$"), re.compile(r"^assert_eq (?:! )?\( (.+?) , [SN] \)$"), re.compile(r"^assert_eq (?:! )?\( [SN] , (.+?) \)$"), re.compile(r"^assert (?:! )?\( (.+?) \. REL \( [SN] \) \)$"), re.compile(r"^expect \( (.+?) \) \. REL \( [SN] \)$"), ] def polar(shape: str) -> str: """An assertion shape with polarity and check strength erased: in/not in, ==/!=, contains/starts_with.""" for rx, repl in POLAR: shape = rx.sub(repl, shape) for rx in REL_FORMS: m = rx.match(shape) if m: return f"REL({m.group(1)})" return shape CALLEE = re.compile(r"([A-Za-z_$][\w$]*(?:\s*(?:\.|->|::|\?->)\s*[A-Za-z_$][\w$]*)*)\s*!?\s*\(") ASSIGNED = re.compile( r"^(?:(?:let|const|var|mut)\s+)*[\w$\s,()\[\]*.:>-]*?(?<![=!<>])=(?![=>])\s*(?:await\s+)?" ) def head(shape: str) -> str: """The callee a setup statement runs (`world.write`), else the whole statement.""" body = ASSIGNED.sub("", shape, count=1).strip() m = CALLEE.match(body) return re.sub(r"\s+", "", m.group(1)) if m else shape def receiver(shape: str) -> str | None: """The object a trailing method call acts on (`(root / _).chmod(_)` -> `(root / _)`), else None.""" body = ASSIGNED.sub("", shape, count=1).strip().rstrip(";").rstrip() if not body.endswith(")"): return None depth = 0 for i in range(len(body) - 1, -1, -1): depth += {")": 1, "(": -1}.get(body[i], 0) if depth == 0: m = re.search(r"(?:\.|->|\?->)\s*[A-Za-z_$][\w$]*\s*$", body[:i]) return re.sub(r"\s+", "", body[: m.start()]) if m and m.start() > 0 else None return None SUBJECT = [ re.compile(r"^assert .+? (?:not in|in) (.+)$"), re.compile(r"^assert (?:not )?(.+?) (?:==|!=|is not|is) .+$"), re.compile(r"^assert (?:not )?(.+?)\.(?:startswith|endswith)\(.*\)$"), re.compile(r"^assert (?:! )?\( (?:! )?(.+?) \. (?:contains|starts_with|ends_with) \( .*$"), re.compile(r"^assert_(?:eq|ne) ! \( (.+?) , .*$"), re.compile(r"^expect \( (.+?) \) \..*$"), ] def subject(shape: str) -> str: """What an assertion inspects (`v0.stderr` in `assert _ in v0.stderr`), else the whole shape.""" for rx in SUBJECT: m = rx.match(shape) if m: return m.group(1) return shape # ---------------------------------------------------------------- Python ASSERT_CALL = re.compile(r"^_*(?:assert|expect|verify|check)\w*$") FAIL_OWNERS = {"pytest", "self", "cls", "unittest"} RAISES_CTX = re.compile(r"(?:^|\.)(?:raises|warns|deprecated_call|assertRaises\w*|assertWarns\w*)$") SKIP_CALL = re.compile(r"(?:^|\.)(?:skip|skipTest|importorskip|xfail)$") READER_CALLS = { "loads", "load", "read_text", "read_bytes", "exists", "is_file", "is_dir", "is_symlink", "stat", "lstat", "json_lines", "readlines", "splitlines", "decode", "encode", "keys", "values", "items", "get", "len", "sorted", "list", "dict", "set", "tuple", "str", "bytes", "int", "float", "bool", "fspath", "glob", "rglob", "iterdir", "read", "strip", "split", "lower", "upper", "startswith", "endswith", "count", "index", "find", "hexdigest", "sha256", "any", "all", "sum", "min", "max", "isinstance", "type", "repr", "format", "join", "replace", "resolve", "readlink", "listdir", "with_name", "with_suffix", "joinpath", "relative_to", "as_posix", "read_json", "getvalue", } # fmt: skip READER_PREFIX = re.compile(r"^_*(?:read|load|parse|show|lines|json|manifest|collect)") DIGIT_RUN = re.compile(r"(?<![A-Za-z0-9])\d{2,}(?!\d)") STATUS_WORDS = re.compile(r"returncode|[Ss]tatus|exit_code|\.code\b|\brc\b") def call_name(node: ast.Call) -> str: func = node.func if isinstance(func, ast.Name): return func.id if isinstance(func, ast.Attribute): return func.attr return "" def dotted(node: ast.AST) -> str: return text_of(node) def clone(node): """Copy an AST far faster than copy.deepcopy; scalars and context singletons are shared.""" new = node.__class__.__new__(node.__class__) fields = new.__dict__ for key, value in node.__dict__.items(): if isinstance(value, list): fields[key] = [clone(item) if isinstance(item, ast.AST) else item for item in value] elif isinstance(value, ast.AST) and not isinstance(value, ast.expr_context): fields[key] = clone(value) else: fields[key] = value return new def walk(root: ast.AST): """Pre-order walk in source order; much cheaper than ast.walk.""" stack = [root] while stack: node = stack.pop() yield node for name in reversed(node._fields): value = getattr(node, name, None) if isinstance(value, list): stack.extend(v for v in reversed(value) if isinstance(v, ast.AST)) elif isinstance(value, ast.AST) and not isinstance(value, ast.expr_context): stack.append(value) def text_of(node: ast.AST) -> str: try: return ast.unparse(node) except (ValueError, RecursionError): # 3.11 cannot unparse a backslash inside an f-string part return ast.dump(node, annotate_fields=False) class Subst(ast.NodeTransformer): def __init__(self, mapping: dict[str, ast.expr]) -> None: self.mapping = mapping def visit_Name(self, node: ast.Name) -> ast.AST: if isinstance(node.ctx, ast.Load) and node.id in self.mapping: return clone(self.mapping[node.id]) return node def literal(node: ast.AST) -> tuple[bool, object]: try: return True, ast.literal_eval(node) except Exception: # noqa: BLE001 - literal_eval raises several types return False, None COMPARE_OPS = { ast.Eq: lambda a, b: a == b, ast.NotEq: lambda a, b: a != b, ast.In: lambda a, b: a in b, ast.NotIn: lambda a, b: a not in b, ast.Is: lambda a, b: a is b, ast.IsNot: lambda a, b: a is not b, ast.Lt: lambda a, b: a < b, ast.LtE: lambda a, b: a <= b, ast.Gt: lambda a, b: a > b, ast.GtE: lambda a, b: a >= b, } def const_eval(node: ast.AST) -> tuple[bool, object]: """Evaluate a test expression that parametrize substitution made constant.""" ok, value = literal(node) if ok: return True, value if isinstance(node, ast.UnaryOp) and isinstance(node.op, ast.Not): ok, value = const_eval(node.operand) return ok, (not value) if ok else None if isinstance(node, ast.BoolOp): values = [] for item in node.values: ok, value = const_eval(item) if not ok: return False, None values.append(value) return True, all(values) if isinstance(node.op, ast.And) else any(values) if isinstance(node, ast.Compare): ok, left = const_eval(node.left) if not ok: return False, None for op, right_node in zip(node.ops, node.comparators, strict=True): ok, right = const_eval(right_node) fn = COMPARE_OPS.get(type(op)) if not ok or fn is None: return False, None try: if not fn(left, right): return True, False except TypeError: return False, None left = right return True, True return False, None class FoldIfExp(ast.NodeTransformer): def visit_IfExp(self, node: ast.IfExp) -> ast.AST: self.generic_visit(node) ok, value = const_eval(node.test) if ok: return node.body if value else node.orelse return node def fold(stmts: list[ast.stmt]) -> list[ast.stmt]: out: list[ast.stmt] = [] for stmt in stmts: stmt = FoldIfExp().visit(stmt) if isinstance(stmt, ast.If): ok, value = const_eval(stmt.test) if ok: out.extend(fold(stmt.body if value else stmt.orelse)) continue for name in ("body", "orelse", "finalbody"): if isinstance(getattr(stmt, name, None), list): setattr(stmt, name, fold(getattr(stmt, name))) if isinstance(stmt, ast.Try | ast.TryStar): for handler in stmt.handlers: handler.body = fold(handler.body) out.append(stmt) return out def marker(name: str, *args: ast.expr) -> ast.stmt: return ast.Expr(value=ast.Call(func=ast.Name(id=name, ctx=ast.Load()), args=list(args), keywords=[])) def is_assert_stmt(stmt: ast.stmt) -> bool: if isinstance(stmt, ast.Assert): return True if isinstance(stmt, ast.Expr): value = stmt.value.value if isinstance(stmt.value, ast.Await) else stmt.value return isinstance(value, ast.Call) and is_assert_call(value) return False def is_assert_call(node: ast.Call) -> bool: """`assert*`/`expect*`/... calls, and `fail()` only as pytest's or unittest's: `world.fail(x)` is setup.""" name = call_name(node) if name == "fail": func = node.func return isinstance(func, ast.Name) or ( isinstance(func, ast.Attribute) and isinstance(func.value, ast.Name) and func.value.id in FAIL_OWNERS ) return bool(ASSERT_CALL.match(name)) RESULT_ATTRS = {"returncode", "status_code", "exit_code", "ok", "stdout", "stderr", "combined", "output"} PURE_CALLS = {"match", "fullmatch", "search", "findall", "compile", "sub", "dumps", "approx", "call", "ANY"} def acting(test: ast.AST, wrapper: ast.AST | None = None) -> bool: """True when the assertion itself exercises code: `assert run(...).returncode == 0`, or `assert f("input") == 1` (a non-reader call given a string literal inside the assertion).""" for n in walk(test): if ( isinstance(n, ast.Attribute) and n.attr in RESULT_ATTRS and isinstance(n.value, ast.Call) and call_name(n.value) not in READER_CALLS ): return True if isinstance(n, ast.Call) and n is not wrapper: name = call_name(n) if ( name in READER_CALLS or name in PURE_CALLS or READER_PREFIX.match(name) or ASSERT_CALL.match(name) ): continue if any(carries_literal(a) for a in [*n.args, *(k.value for k in n.keywords)]): return True return False def blob_atoms(text: str) -> list[str]: """Atoms of the string literals inside a collapsed container, else of its text.""" try: tree = ast.parse(text, mode="eval") except SyntaxError: return atoms(text) found: list[str] = [] for n in walk(tree): if isinstance(n, ast.Constant) and isinstance(n.value, str | bytes): found += atoms(n.value if isinstance(n.value, str) else n.value.decode("latin-1")) return found def carries_literal(node: ast.AST) -> bool: """A call argument that supplies input text: a literal or f-string, also nested (`[call("bash", command=...)]`).""" return any( isinstance(n, ast.JoinedStr) or isinstance(n, ast.Constant) and isinstance(n.value, str | bytes | Blob) for n in walk(node) ) def selectors(stmt: ast.AST) -> set[int]: """Literals that pick what to read (`text.index("### Phase 3:")`, `read(root / "x.md")`): fixture, not signal.""" chosen: set[int] = set() for node in walk(stmt): if isinstance(node, ast.Call) and observer(node): picked = [*node.args, *(k.value for k in node.keywords)] if isinstance(node.func, ast.Attribute): picked.append(node.func.value) # `(root / "x.md").read_text()` for arg in picked: if isinstance(arg, ast.Constant | ast.BinOp): chosen.update(id(n) for n in walk(arg) if isinstance(n, ast.Constant)) return chosen def observer(node: ast.Call) -> bool: """A call that only reads, formats, or asserts: a test made of these alone exercises no code.""" name = call_name(node) return ( name.startswith("__") or name in READER_CALLS or name in PURE_CALLS or bool(READER_PREFIX.match(name)) or is_assert_call(node) ) def is_skip_guard(stmt: ast.stmt) -> bool: if not isinstance(stmt, ast.If) or len(stmt.body) != 1: return False only = stmt.body[0] if isinstance(only, ast.Return): return True return ( isinstance(only, ast.Expr) and isinstance(only.value, ast.Call) and bool(SKIP_CALL.search(dotted(only.value.func))) ) def flatten(stmts: list[ast.stmt]) -> list[tuple[str, ast.stmt]]: out: list[tuple[str, ast.stmt]] = [] for stmt in stmts: if isinstance(stmt, ast.Pass) or ( isinstance(stmt, ast.Expr) and isinstance(stmt.value, ast.Constant) ): continue if is_skip_guard(stmt): out.extend(flatten(stmt.orelse)) # type: ignore[attr-defined] continue if isinstance(stmt, ast.If): out.append(("action", ast.If(test=stmt.test, body=[ast.Pass()], orelse=[]))) out.extend(flatten(stmt.body)) if stmt.orelse: out.append(("action", marker("__else__"))) out.extend(flatten(stmt.orelse)) out.append(("action", marker("__end__"))) elif isinstance(stmt, ast.For | ast.AsyncFor | ast.While) and all( is_assert_stmt(s) for s in stmt.body ): loop = clone(stmt) loop.body = [ ast.Assert(test=s.test, msg=None) if isinstance(s, ast.Assert) else s for s in loop.body ] out.append(("assert", loop)) elif isinstance(stmt, ast.For | ast.AsyncFor): out.append(("action", marker("__for__", stmt.target, stmt.iter))) out.extend(flatten(stmt.body)) out.append(("action", marker("__end__"))) elif isinstance(stmt, ast.While): out.append(("action", marker("__while__", stmt.test))) out.extend(flatten(stmt.body)) out.append(("action", marker("__end__"))) elif isinstance(stmt, ast.With | ast.AsyncWith): others = [] for item in stmt.items: ctx = item.context_expr if isinstance(ctx, ast.Call) and RAISES_CTX.search(dotted(ctx.func)): out.append(("assert", ast.Expr(value=ctx))) else: others.append(item) if others: out.append(("action", ast.With(items=others, body=[ast.Pass()], lineno=0, type_comment=None))) out.extend(flatten(stmt.body)) if others: out.append(("action", marker("__end__"))) elif isinstance(stmt, ast.Try | ast.TryStar): out.extend(flatten(stmt.body)) for handler in stmt.handlers: out.append(("action", marker("__except__", *([handler.type] if handler.type else [])))) out.extend(flatten(handler.body)) out.extend(flatten(stmt.orelse)) if stmt.finalbody: out.append(("action", marker("__finally__"))) out.extend(flatten(stmt.finalbody)) elif isinstance(stmt, ast.Assert): if acting(stmt.test): copy = ast.Expr(value=stmt.test) copy.acting_copy = True # type: ignore[attr-defined] out.append(("action", copy)) out.append(("assert", ast.Assert(test=stmt.test, msg=None))) elif is_assert_stmt(stmt): call = stmt.value.value if isinstance(stmt.value, ast.Await) else stmt.value # type: ignore[attr-defined] if acting(stmt, call): copy = clone(stmt) copy.acting_copy = True out.append(("action", copy)) out.append(("assert", stmt)) else: out.append(("action", stmt)) return out def body_names(stmts: list[ast.stmt]) -> tuple[Counter[str], set[str], bool]: """Store counts, loaded names, and whether any if/ternary exists, in one pass.""" counts: Counter[str] = Counter() loads: set[str] = set() branches = False for root in stmts: for node in walk(root): if isinstance(node, ast.Name): if isinstance(node.ctx, ast.Load): loads.add(node.id) else: counts[node.id] += 1 elif isinstance(node, ast.If | ast.IfExp): branches = True elif isinstance(node, ast.arg): counts[node.arg] += 2 elif isinstance(node, ast.ExceptHandler) and node.name: counts[node.name] += 2 elif isinstance(node, ast.AugAssign) and isinstance(node.target, ast.Name): counts[node.target.id] += 1 elif isinstance(node, ast.FunctionDef | ast.AsyncFunctionDef | ast.ClassDef): counts[node.name] += 2 return counts, loads, branches def loaded_names(node: ast.AST) -> set[str]: return {n.id for n in walk(node) if isinstance(n, ast.Name) and isinstance(n.ctx, ast.Load)} INLINABLE = ( ast.Name, ast.Attribute, ast.Subscript, ast.BinOp, ast.Constant, ast.JoinedStr, ast.FormattedValue, ast.Tuple, ast.UnaryOp, ast.Compare, ast.BoolOp, ast.Slice, ast.Starred, ast.operator, ast.unaryop, ast.cmpop, ast.boolop, ) # fmt: skip def call_free(node: ast.AST) -> bool: return all(isinstance(n, INLINABLE) for n in walk(node)) def reader_only(node: ast.AST) -> bool: calls = [n for n in walk(node) if isinstance(n, ast.Call)] return bool(calls) and all( call_name(c) in READER_CALLS or READER_PREFIX.match(call_name(c)) for c in calls ) def single_target(stmt: ast.stmt) -> str | None: if isinstance(stmt, ast.Assign) and len(stmt.targets) == 1 and isinstance(stmt.targets[0], ast.Name): return stmt.targets[0].id if isinstance(stmt, ast.AnnAssign) and stmt.value is not None and isinstance(stmt.target, ast.Name): return stmt.target.id return None def inline(flat: list[tuple[str, ast.stmt]], stores: Counter[str]) -> list[tuple[str, ast.stmt]]: """Inline call-free single bindings, then observation bindings used only by assertions.""" mapping: dict[str, ast.expr] = {} first: list[tuple[str, ast.stmt, set[str]]] = [] for kind, stmt in flat: loads = loaded_names(stmt) if mapping and loads & mapping.keys(): stmt = Subst(mapping).visit(stmt) loads = loaded_names(stmt) name = single_target(stmt) value = getattr(stmt, "value", None) if kind == "action" and name and stores[name] == 1 and value is not None and call_free(value): mapping[name] = value continue first.append((kind, stmt, loads)) candidates = { name for kind, stmt, _ in first if kind == "action" and (name := single_target(stmt)) and stores[name] == 1 and reader_only(stmt.value) # type: ignore[union-attr] } changed = True while changed: changed = False for kind, stmt, loads in first: if kind == "assert" or single_target(stmt) in candidates: continue used = loads & candidates if used: candidates -= used changed = True # A read taken before a later action is a snapshot, not an observation of the result: keep it. for index, (kind, stmt, _) in enumerate(first): name = single_target(stmt) if kind != "action" or name not in candidates: continue uses = [i for i, (_, _, loads) in enumerate(first) if i > index and name in loads] between = first[index + 1 : max(uses, default=index) + 1] if any(k == "action" and single_target(st) not in candidates for k, st, _ in between): candidates.discard(name) mapping = {} out: list[tuple[str, ast.stmt]] = [] for kind, stmt, loads in first: if mapping and loads & mapping.keys(): stmt = Subst(mapping).visit(stmt) name = single_target(stmt) if kind == "action" and name in candidates: mapping[name] = stmt.value # type: ignore[union-attr, assignment] continue out.append((kind, stmt)) return out class Blob: """A Constant payload that unparses verbatim: a collapsed literal container or a placeholder.""" __slots__ = ("text",) def __init__(self, text: str) -> None: self.text = text def __repr__(self) -> str: return self.text def __eq__(self, other: object) -> bool: return isinstance(other, Blob) and other.text == self.text def __hash__(self) -> int: return hash(self.text) def all_literal(node: ast.AST) -> bool: if isinstance(node, ast.Constant): return True if isinstance(node, ast.List | ast.Tuple | ast.Set): return all(all_literal(e) for e in node.elts) if isinstance(node, ast.Dict): return all(k is not None and all_literal(k) for k in node.keys) and all( all_literal(v) for v in node.values ) if isinstance(node, ast.UnaryOp): return all_literal(node.operand) return False def replace_children(root: ast.AST, pick) -> ast.AST: """Replace, in place, every maximal subexpression for which pick() returns a node.""" stack = [root] while stack: node = stack.pop() for name in node._fields: value = getattr(node, name, None) if isinstance(value, list): for i, item in enumerate(value): if isinstance(item, ast.expr) and (new := pick(item)) is not None: value[i] = new elif isinstance(item, ast.AST): stack.append(item) elif isinstance(value, ast.expr): if (new := pick(value)) is not None: setattr(node, name, new) else: stack.append(value) elif isinstance(value, ast.AST) and not isinstance(value, ast.expr_context): stack.append(value) return root def collapse(root: ast.AST) -> ast.AST: """Fold large literal containers into one Constant so later passes visit one node, not hundreds.""" def pick(node: ast.expr) -> ast.expr | None: if isinstance(node, ast.List | ast.Tuple | ast.Set | ast.Dict) and all_literal(node): size = sum(1 for _ in walk(node)) if size > 12: return ast.Constant(value=Blob(text_of(node))) return None return replace_children(root, pick) UNDERSCORE = ast.Name(id="_", ctx=ast.Load()) Unparser = getattr(ast, "_Unparser", None) if Unparser is not None: class ShapeUnparser(Unparser): # type: ignore[misc, valid-type] """Unparse with every literal (and f-string) written as `_`, without copying the tree.""" def traverse(self, node): if isinstance(node, ast.expr) and abstracted(node): self.write("_") return None return super().traverse(node) def literal_expr(node: ast.AST) -> bool: """A literal, or an expression built only from literals (`"row\\n" + "x" * 5200`): one table cell.""" if isinstance(node, ast.BinOp): return literal_expr(node.left) and literal_expr(node.right) if isinstance(node, ast.UnaryOp): return literal_expr(node.operand) if isinstance(node, ast.List | ast.Tuple | ast.Set): return all(literal_expr(e) for e in node.elts) return all_literal(node) def abstracted(node: ast.expr) -> bool: """A literal a table row could vary; True/False/None stay, because they carry assertion polarity.""" if isinstance(node, ast.Constant) and (node.value is None or isinstance(node.value, bool)): return False return isinstance(node, ast.JoinedStr) or literal_expr(node) def shape_of(stmt: ast.stmt) -> str: if Unparser is not None: try: return ShapeUnparser().visit(stmt) except (ValueError, RecursionError): pass def pick(node: ast.expr) -> ast.expr | None: return UNDERSCORE if abstracted(node) else None return text_of(replace_children(clone(stmt), pick)) def placeholders(flat: list[tuple[str, ast.stmt]], tmp_roots: bool) -> None: """Replace tmp path segments and threaded identifiers in place (see module docstring).""" if tmp_roots: seen: dict[str, str] = {} for _, stmt in flat: for node in walk(stmt): if isinstance(node, ast.BinOp) and isinstance(node.op, ast.Div): root = node.left while isinstance(root, ast.BinOp) and isinstance(root.op, ast.Div): root = root.left right = node.right if ( isinstance(root, ast.Name) and "tmp" in root.id.lower() and isinstance(right, ast.Constant) and isinstance(right.value, str) ): right.value = Blob(seen.setdefault(right.value, f"TMP{len(seen)}")) counts: Counter[str] = Counter() compared: set[str] = set() for kind, stmt in flat: if getattr(stmt, "acting_copy", False): continue top = getattr(stmt, "value", None) for node in walk(stmt): if isinstance(node, ast.Constant): v = node.value if isinstance(v, int) and not isinstance(v, bool) and abs(v) >= 10: values = [str(abs(v))] elif isinstance(v, str): values = DIGIT_RUN.findall(v) elif isinstance(v, bytes): values = DIGIT_RUN.findall(v.decode("latin-1")) else: continue if kind == "action": counts.update(values) elif kind == "assert" and isinstance(node, ast.Compare | ast.Call): operands = [node.left, *node.comparators] if isinstance(node, ast.Compare) else node.args if isinstance(node, ast.Call) and node is not top: continue if not STATUS_WORDS.search(text_of(node)): continue for operand in operands: if isinstance(operand, ast.Constant) and isinstance(operand.value, int): compared.add(str(abs(operand.value))) ids: dict[str, str] = {} for value in counts: # insertion order is first appearance if counts[value] >= 2 and value not in compared: ids[value] = f"ID{len(ids)}" if not ids: return def sub(text: str) -> str: return DIGIT_RUN.sub(lambda m: f"#{ids[m.group()]}#" if m.group() in ids else m.group(), text) for _, stmt in flat: for node in walk(stmt): if isinstance(node, ast.Constant): v = node.value if isinstance(v, int) and not isinstance(v, bool) and str(abs(v)) in ids: node.value = Blob(ids[str(abs(v))]) elif isinstance(v, str): node.value = sub(v) elif isinstance(v, bytes): node.value = sub(v.decode("latin-1")).encode("latin-1") def split_kwargs( stmt: ast.stmt, text: str, defaults: dict[str, dict[str, object]] ) -> tuple[str, tuple[tuple[tuple[str, str], ...], ...]]: """Positional text plus per-call keywords; `!key` marks a keyword whose callee default is known.""" calls = [n for n in walk(stmt) if isinstance(n, ast.Call)] if not any(c.keywords for c in calls): return text, tuple(() for _ in calls) node = clone(stmt) found: list[tuple[tuple[str, str], ...]] = [] stack: list[ast.AST] = [node] while stack: n = stack.pop() if isinstance(n, ast.Call): known = defaults.get(n.func.id, {}) if isinstance(n.func, ast.Name) else {} found.append( tuple( sorted( (f"!{k.arg}" if k.arg in known else k.arg or f"**{i}", text_of(k.value)) for i, k in enumerate(n.keywords) ) ) ) n.keywords = [] stack.extend(reversed([c for c in ast.iter_child_nodes(n) if not isinstance(c, ast.expr_context)])) return text_of(node), tuple(found) def module_constants(tree: ast.Module) -> dict[str, ast.expr]: counts: Counter[str] = Counter() values: dict[str, ast.expr] = {} for stmt in tree.body: name = single_target(stmt) if name: counts[name] += 1 value = stmt.value # type: ignore[union-attr] if value is not None and all_literal(value): values[name] = value return {k: v for k, v in values.items() if counts[k] == 1} def module_defaults(tree: ast.Module) -> dict[str, dict[str, object]]: """Literal keyword defaults of module-level helpers, so passing a default explicitly is a no-op.""" found: dict[str, dict[str, object]] = {} for stmt in tree.body: if isinstance(stmt, ast.FunctionDef | ast.AsyncFunctionDef): args = stmt.args positional = [*args.posonlyargs, *args.args] pairs = list(zip(positional[len(positional) - len(args.defaults) :], args.defaults, strict=True)) pairs += [(a, d) for a, d in zip(args.kwonlyargs, args.kw_defaults, strict=True) if d is not None] values = {} for arg, default in pairs: ok, value = literal(default) if ok: values[arg.arg] = value if args.kwarg is not None: values.update(dict_defaults(stmt, args.kwarg.arg)) if values: found[stmt.name] = values # A helper that forwards **extra to another helper inherits that helper's keyword defaults. for _ in range(3): for stmt in tree.body: if not isinstance(stmt, ast.FunctionDef | ast.AsyncFunctionDef) or stmt.args.kwarg is None: continue extra = stmt.args.kwarg.arg for node in walk(stmt): if ( isinstance(node, ast.Call) and isinstance(node.func, ast.Name) and node.func.id in found and any( k.arg is None and isinstance(k.value, ast.Name) and k.value.id == extra for k in node.keywords ) ): inherited = { k: v for k, v in found[node.func.id].items() if k not in found.get(stmt.name, {}) } if inherited: found.setdefault(stmt.name, {}).update(inherited) return found def dict_defaults(func: ast.FunctionDef | ast.AsyncFunctionDef, extra: str) -> dict[str, object]: """Keys of a literal dict that `**extra` is merged over (`body.update(extra)` or `{**d, **extra}`).""" merged = any( isinstance(n, ast.Call) and call_name(n) == "update" and any(isinstance(a, ast.Name) and a.id == extra for a in n.args) or isinstance(n, ast.Dict) and any( k is None and isinstance(v, ast.Name) and v.id == extra for k, v in zip(n.keys, n.values, strict=True) ) for n in walk(func) ) if not merged: return {} values: dict[str, object] = {} seen: Counter[str] = Counter() for node in walk(func): if isinstance(node, ast.Dict): for key, value in zip(node.keys, node.values, strict=True): if isinstance(key, ast.Constant) and isinstance(key.value, str): seen[key.value] += 1 ok, literal_value = literal(value) if ok: values[key.value] = literal_value elif ( isinstance(node, ast.Subscript) and isinstance(node.ctx, ast.Store) and isinstance(node.slice, ast.Constant) and isinstance(node.slice.value, str) ): seen[node.slice.value] += 1 # A key the helper assigns more than once has a computed default: treat it as unknown. return {k: v for k, v in values.items() if seen[k] == 1} def drop_default_kwargs(stmts: list[ast.stmt], defaults: dict[str, dict[str, object]]) -> None: for stmt in stmts: for node in walk(stmt): if isinstance(node, ast.Call) and isinstance(node.func, ast.Name) and node.func.id in defaults: known = defaults[node.func.id] kept = [] for kw in node.keywords: if kw.arg in known: ok, value = literal(kw.value) # Exact type, so True does not stand in for a default of 1. if ok and type(value) is type(known[kw.arg]) and value == known[kw.arg]: # noqa: E721 continue kept.append(kw) node.keywords = kept def parametrize_cases( func: ast.FunctionDef | ast.AsyncFunctionDef, constants: dict[str, ast.expr] ) -> tuple[list[tuple[str, dict[str, ast.expr]]], set[str], set[str], bool]: """Return (cases, expanded names, symbolic names, parametrized?).""" cases: list[tuple[str, dict[str, ast.expr]]] = [("", {})] expanded: set[str] = set() symbolic: set[str] = set() parametrized = False for deco in func.decorator_list: if not isinstance(deco, ast.Call) or not dotted(deco.func).endswith("parametrize") or not deco.args: continue parametrized = True names_node = deco.args[0] ok, names_value = literal(names_node) if not ok: continue if isinstance(names_value, str): names = [n.strip() for n in names_value.split(",") if n.strip()] else: names = [str(n) for n in names_value] values_node = deco.args[1] if len(deco.args) > 1 else None for kw in deco.keywords: if kw.arg == "argvalues": values_node = kw.value if isinstance(values_node, ast.Name) and values_node.id in constants: values_node = constants[values_node.id] ids_node = next((kw.value for kw in deco.keywords if kw.arg == "ids"), None) ids_ok, ids_value = literal(ids_node) if ids_node is not None else (False, None) if not isinstance(values_node, ast.List | ast.Tuple): symbolic.update(names) continue rows: list[tuple[str, dict[str, ast.expr]]] = [] for index, element in enumerate(values_node.elts): label = None if isinstance(element, ast.Call) and call_name(element) == "param": for kw in element.keywords: if kw.arg == "id": ok, label = literal(kw.value) element = ( element.args[0] if len(names) == 1 and element.args else ast.Tuple(elts=element.args) ) if len(names) == 1: values = [element] elif isinstance(element, ast.Tuple | ast.List) and len(element.elts) == len(names): values = list(element.elts) else: symbolic.update(names) rows = [] break if label is None and ids_ok and isinstance(ids_value, list | tuple) and index < len(ids_value): label = ids_value[index] if label is None: simple = [literal(v) for v in values] label = ( "-".join(str(v) for _, v in simple) if all(ok and isinstance(v, str | int | float | bool | None) for ok, v in simple) else str(index) ) rows.append((str(label), dict(zip(names, values, strict=True)))) if not rows: continue expanded.update(names) cases = [ ("-".join(p for p in (a, b) if p), {**ma, **mb}) for (a, ma), (b, mb) in itertools.product(cases, rows) ][:64] return cases, expanded, symbolic, parametrized def py_unit( path: str, qual: str, func: ast.FunctionDef | ast.AsyncFunctionDef, base: list[ast.stmt], label: str, subst: dict[str, ast.expr], constants: dict[str, ast.expr], defaults: dict[str, dict[str, object]], fixtures: tuple[str, ...], fixture_keys: tuple[str, ...], parametrized: bool, ) -> Unit: body = [Subst(subst).visit(clone(s)) if subst else clone(s) for s in base] stores, used, branches = body_names(body) cmap = {k: v for k, v in constants.items() if k in used and k not in stores and k not in fixtures} if cmap: body = [Subst(cmap).visit(s) for s in body] if branches and (subst or cmap): body = fold(body) stores = body_names(body)[0] if defaults: drop_default_kwargs(body, defaults) flat = inline(flatten(body), stores) fixture_set = set(fixtures) mapping: dict[str, str] = {} for _, stmt in sorted(flat, key=lambda item: item[0] == "assert"): for node in walk(stmt): if isinstance(node, ast.Name): if node.id in stores and node.id not in fixture_set: node.id = mapping.setdefault(node.id, f"v{len(mapping)}") elif isinstance(node, ast.arg) and node.arg in stores and node.arg not in fixture_set: node.arg = mapping.setdefault(node.arg, f"v{len(mapping)}") placeholders(flat, any("tmp" in f.lower() for f in fixtures)) texts = [text_of(stmt) for _, stmt in flat] shapes = [shape_of(stmt) for _, stmt in flat] inputs: set[str] = set() needles: set[str] = set() real_act = False for kind, stmt in flat: bucket = needles if kind == "assert" else inputs chosen = selectors(stmt) for node in walk(stmt): if isinstance(node, ast.Constant) and id(node) in chosen: continue if isinstance(node, ast.Constant): v = node.value if isinstance(v, str): bucket.update(atoms(v)) elif isinstance(v, Blob): bucket.update(blob_atoms(v.text)) elif isinstance(v, bytes): bucket.update(atoms(v.decode("latin-1"))) # A module helper called with only fixtures (`spine_text(repo_root)`) loads a fixture. elif ( kind == "action" and isinstance(node, ast.Call) and not observer(node) and ( isinstance(node.func, ast.Attribute) or any(carries_literal(a) for a in [*node.args, *(k.value for k in node.keywords)]) ) ): real_act = True last_assert = max((i for i, (k, _) in enumerate(flat) if k == "assert"), default=-1) act = -1 for i, (kind, _) in enumerate(flat): if kind == "action" and (last_assert < 0 or i < last_assert) and not texts[i].startswith("__"): act = i actions, pos, kwargs, action_shape = [], [], [], [] pre, post, assert_shape = set(), set(), [] for i, (kind, stmt) in enumerate(flat): if kind == "assert": (pre if i < act else post).add(texts[i]) assert_shape.append(shapes[i]) else: actions.append(texts[i]) p, k = split_kwargs(stmt, texts[i], defaults) pos.append(p) kwargs.extend(k) action_shape.append(shapes[i]) return Unit( path=path, line=func.lineno, func=f"{path}::{qual}", name=f"{qual}[{label}]" if label else qual, fixtures=fixture_keys, actions=tuple(actions), action_pos=tuple(pos), action_kwargs=tuple(kwargs), pre=frozenset(pre), post=frozenset(post), action_shape=tuple(action_shape), assert_shape=tuple(sorted(assert_shape)), act_shape=shapes[act] if act >= 0 else "", setup_shape=tuple(shapes[i] for i in range(act + 1) if flat[i][0] == "action"), params=parametrized, status=status_sig(texts[i] for i, (kind, _) in enumerate(flat) if kind == "assert"), inputs=frozenset(inputs), needles=frozenset(needles), real_act=real_act, ) def scan_python(path: str, text: str, part: int = 0, parts: int = 1) -> list[Unit]: tree = ast.parse(text) constants = {k: collapse(ast.Expr(value=clone(v))).value for k, v in module_constants(tree).items()} # type: ignore[attr-defined] defaults = module_defaults(tree) local_fixtures = { stmt.name for stmt in tree.body if isinstance(stmt, ast.FunctionDef | ast.AsyncFunctionDef) and any("fixture" in dotted(d.func if isinstance(d, ast.Call) else d) for d in stmt.decorator_list) } funcs: list[tuple[str, ast.FunctionDef | ast.AsyncFunctionDef]] = [] for stmt in tree.body: if isinstance(stmt, ast.FunctionDef | ast.AsyncFunctionDef) and stmt.name.startswith("test"): funcs.append((stmt.name, stmt)) elif isinstance(stmt, ast.ClassDef): for item in stmt.body: if isinstance(item, ast.FunctionDef | ast.AsyncFunctionDef) and item.name.startswith("test"): funcs.append((f"{stmt.name}::{item.name}", item)) units: list[Unit] = [] for qual, func in funcs[part::parts]: cases, expanded, symbolic, parametrized = parametrize_cases(func, constants) params = [a.arg for a in [*func.args.posonlyargs, *func.args.args, *func.args.kwonlyargs]] fixtures = [p for p in params if p not in ("self", "cls") and p not in expanded] for deco in func.decorator_list: if isinstance(deco, ast.Call) and dotted(deco.func).endswith("usefixtures"): fixtures += [str(v) for ok, v in map(literal, deco.args) if ok] del symbolic # symbolic names stay ordinary parameters in `fixtures` # Fixtures are scoped: one defined in this file, or else in the directory's conftest chain. scope = str(Path(path).parent) fixture_keys = tuple(sorted(f"{f}@{path if f in local_fixtures else scope}" for f in fixtures)) base = [collapse(clone(stmt)) for stmt in func.body] for label, subst in cases: try: units.append( py_unit( path, qual, func, base, label, subst, constants, defaults, tuple(fixtures), fixture_keys, parametrized, ) ) except RecursionError as error: raise ValueError(f"{qual}: too deeply nested") from error return units # ---------------------------------------------------------------- token languages PUNCT3 = {"===", "!==", "...", "<=>", "**=", "<<=", ">>=", "?->", "??=", "::<"} PUNCT2 = { "->", "=>", "::", "==", "!=", "<=", ">=", "&&", "||", "??", "?.", "+=", "-=", "*=", "/=", "%=", "++", "--", "<<", ">>", "..", "|=", "&=", "^=", "**", } # fmt: skip JS_REGEX_PREV = set("(,=:[!&|?{};+-*%<>~^") | { "return", "typeof", "=>", "&&", "||", "??", "==", "===", "!=", "!==", } RS_RAW = re.compile(r'b?r(#*)"') RS_CHAR = re.compile(r"'(?:\\.[^']*|[^'\\])'") RS_LIFETIME = re.compile(r"'\w+") PHP_HEREDOC = re.compile(r"<<<\s*['\"]?(\w+)['\"]?") PHP_VAR = re.compile(r"\$\w+") NUMBER = re.compile(r"0[xXbBoO]\w+|\d[\d_]*(?:\.\d[\d_]*)?(?:[eE][+-]?\d+)?\w*") IDENT = re.compile(r"[\w$\\]+") RS_IDENT = re.compile(r"\w+") @dataclass class Tok: kind: str # id var num str punct comment text: str line: int def tokenize(text: str, lang: str) -> list[Tok]: toks: list[Tok] = [] i, n, line = 0, len(text), 1 def prev_sig() -> Tok | None: for t in reversed(toks): if t.kind != "comment": return t return None while i < n: c = text[i] if c == "\n": line += 1 i += 1 continue if c.isspace(): i += 1 continue start, start_line = i, line two = text[i : i + 2] if two == "//" or (lang == "php" and c == "#" and text[i : i + 2] != "#["): end = text.find("\n", i) end = n if end < 0 else end toks.append(Tok("comment", text[i:end], line)) i = end continue if two == "/*": end = text.find("*/", i + 2) end = n if end < 0 else end + 2 toks.append(Tok("comment", text[i:end], line)) line += text.count("\n", i, end) i = end continue if lang == "rs" and c in "br" and (m := RS_RAW.match(text, i)): closer = '"' + m.group(1) end = text.find(closer, m.end()) end = n if end < 0 else end + len(closer) toks.append(Tok("str", text[i:end], line)) line += text.count("\n", i, end) i = end continue if lang == "php" and text.startswith("<<<", i): m = PHP_HEREDOC.match(text, i) if m: close = re.compile(r"^\s*" + re.escape(m.group(1)) + r"\b", re.MULTILINE) found = close.search(text, m.end()) end = n if found is None else found.end() toks.append(Tok("str", text[i:end], line)) line += text.count("\n", i, end) -
guard_pins.py 5.9 KB
#!/usr/bin/env python3 """List refusal messages in source that no test asserts. A guard whose message appears in no test is either untested or tested only by a status-only negative, which a different guard's refusal would also satisfy. The count sizes the problem. Individual rows are candidates: the extraction also picks up warnings and format fragments (over-count), and a short generic test needle can look like a pin (under-count). Source literals are taken from lines that look like a refusal (throw, raise, abort, Err(, bail!, exit, reject, deny, ValidationException, ...) and the two lines after them, so multi-line calls are covered. Test needles are every string literal in the test files. In Rust files, everything after the first `#[cfg(test)]` counts as test code, not source. Exit status 2 means known incomplete input: a --src or --tests glob matched no file. Output is still printed. """ from __future__ import annotations import argparse import glob import re import sys from collections import Counter from pathlib import Path SKIP_DIRS = { "node_modules", "vendor", ".git", "target", ".venv", "venv", "__pycache__", "dist", "build", "coverage", ".claude", ".next", ".turbo", ".worktrees", } def expand(patterns: list[str]) -> list[str]: """Glob, dropping dependency, build, and agent-worktree directories.""" found = {f for p in patterns for f in glob.glob(p, recursive=True)} return sorted(f for f in found if Path(f).is_file() and not SKIP_DIRS.intersection(Path(f).parts)) GUARD_LINE = re.compile( r"\b(?:throw|raise|abort(?:_if|_unless)?|bail!|panic!|Err\(|Failure|fail|reject|deny|refuse|exit|die|" r"withMessages|withErrors|ValidationException|HttpException|AuthorizationException|Error\(|report_and_exit)", re.IGNORECASE, ) DOUBLE = re.compile(r'"((?:[^"\\\n]|\\.)*)"') SINGLE = re.compile(r"'((?:[^'\\\n]|\\.)*)'") BACKTICK = re.compile(r"`((?:[^`\\]|\\.)*)`") HOLE = re.compile(r"\{[^{}]*\}|%[sdifx]|\$\{[^}]*\}|\$\w+|\{\d*\}") def literals(line: str, suffix: str) -> list[str]: found = [m.group(1) for m in DOUBLE.finditer(line)] if suffix not in (".rs",): found += [m.group(1) for m in SINGLE.finditer(line)] if suffix in (".js", ".jsx", ".ts", ".tsx", ".mjs", ".cjs"): found += [m.group(1) for m in BACKTICK.finditer(line)] return found def split_rust(text: str) -> tuple[str, str]: m = re.search(r"^\s*#\[cfg\(test\)\]", text, re.MULTILINE) return (text, "") if m is None else (text[: m.start()], text[m.start() :]) def main() -> int: parser = argparse.ArgumentParser( description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter ) parser.add_argument( "--src", action="append", required=True, metavar="GLOB", help="production source glob" ) parser.add_argument("--tests", action="append", default=[], metavar="GLOB", help="test file glob") parser.add_argument("--min-length", type=int, default=20, help="shortest source literal considered") parser.add_argument("--needle-min", type=int, default=12, help="shortest test literal counted as a pin") parser.add_argument( "--exclude", action="append", default=[], metavar="REGEX", help="skip literals matching" ) parser.add_argument("--limit", type=int, default=60, help="unpinned rows printed") args = parser.parse_args() pattern_files = [(pattern, expand([pattern])) for pattern in [*args.src, *args.tests]] unmatched = [pattern for pattern, names in pattern_files if not names] for pattern in unmatched: print(f"guard_pins: unmatched pattern (no eligible files): {pattern}", file=sys.stderr) src_files = expand(args.src) test_files = expand(args.tests) excludes = [re.compile(x) for x in args.exclude] needles: set[str] = set() guards: list[tuple[str, int, str]] = [] for name in test_files: suffix = Path(name).suffix for line in Path(name).read_text(errors="replace").splitlines(): needles.update(n for n in literals(line, suffix) if len(n) >= args.needle_min) for name in src_files: suffix = Path(name).suffix text = Path(name).read_text(errors="replace") if suffix == ".rs": text, test_part = split_rust(text) for line in test_part.splitlines(): needles.update(n for n in literals(line, suffix) if len(n) >= args.needle_min) lines = text.splitlines() window = 0 for number, line in enumerate(lines, 1): if GUARD_LINE.search(line): window = 3 if window: window -= 1 for literal in literals(line, suffix): if len(literal) >= args.min_length and not any(x.search(literal) for x in excludes): guards.append((name, number, literal)) seen: set[str] = set() unpinned: list[tuple[str, int, str]] = [] for name, number, literal in guards: if literal in seen: continue seen.add(literal) fragments = [f.strip() for f in HOLE.split(literal) if len(f.strip()) >= args.needle_min] pinned = any(n in literal or any(n in f for f in fragments) for n in needles) if not pinned: unpinned.append((name, number, literal)) per_file = Counter(name for name, _, _ in unpinned) print( f"{len(seen)} distinct guard literals in {len(src_files)} source files; " f"{len(unpinned)} contain no test string literal (>= {args.needle_min} chars)" ) print(f"{len(needles)} test needles from {len(test_files)} test files\n") print("== unpinned by file") for name, count in per_file.most_common(25): print(f"{count:5d} {name}") print(f"\n== first {args.limit} unpinned") for name, number, literal in unpinned[: args.limit]: print(f"{name}:{number} {literal[:120]}") return 2 if unmatched else 0 if __name__ == "__main__": sys.exit(main()) -
junit_census.py 5.7 KB
#!/usr/bin/env python3 """Census of a gate run's junit reports: skips by reason, unreported test files, zero-assertion tests. Reads junit XML written by pytest (--junitxml), PHPUnit (--log-junit), Vitest (--reporter=junit), jest-junit, or cargo-nextest. Pass --test-files to list test files on disk that no report mentions: those were never collected by the gate. A file counts as mentioned only when its path, name, or stem appears between token boundaries, so a report naming tests/test_user_admin.py does not also mark tests/test_user.py. Exit status 2 means known incomplete input: an unreadable report, or a --test-files glob that matched no file. Output is still printed when a glob is unmatched. """ from __future__ import annotations import argparse import glob import re import sys import xml.etree.ElementTree as ET from collections import defaultdict from pathlib import Path SKIP_DIRS = { "node_modules", "vendor", ".git", "target", ".venv", "venv", "__pycache__", "dist", "build", "coverage", ".claude", ".next", ".turbo", ".worktrees", } def expand(patterns: list[str]) -> list[str]: """Glob, dropping dependency, build, and agent-worktree directories.""" found = {f for p in patterns for f in glob.glob(p, recursive=True)} return sorted(f for f in found if Path(f).is_file() and not SKIP_DIRS.intersection(Path(f).parts)) def normalise(text: str) -> str: return text.replace("\\", "/").replace("::", "/").lower() def bounded(candidate: str) -> re.Pattern[str]: """Match a path, name, or stem only where it is not part of a longer token.""" return re.compile(rf"(?<![\w-]){re.escape(candidate)}(?![\w-])") def main() -> int: parser = argparse.ArgumentParser( description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter ) parser.add_argument("reports", nargs="+", help="junit XML files from the gate run") parser.add_argument( "--test-files", action="append", default=[], metavar="GLOB", help="glob (repeatable, ** allowed) of test files expected to be collected", ) parser.add_argument("--limit", type=int, default=20, help="rows shown per section") args = parser.parse_args() total = skipped = failed = 0 skips: dict[str, list[str]] = defaultdict(list) zero_assertions: list[str] = [] mentioned: set[str] = set() counts_assertions = False for report in args.reports: try: root = ET.parse(report).getroot() except (OSError, ET.ParseError) as error: print(f"junit_census: cannot read {report}: {error}; audit input is incomplete", file=sys.stderr) return 2 for suite in root.iter("testsuite"): for key in ("name", "file"): if suite.get(key): mentioned.add(normalise(suite.get(key, ""))) for case in root.iter("testcase"): total += 1 parts = [case.get(k, "") for k in ("file", "classname", "name")] for part in parts: if part: mentioned.add(normalise(part)) mentioned.add(normalise(part.replace(".", "/"))) ident = "::".join(p for p in (case.get("classname"), case.get("name")) if p) skip = case.find("skipped") if skip is not None: skipped += 1 reason = (skip.get("message") or skip.text or "(no reason)").strip() reason = re.sub(r"\s+", " ", reason)[:200] skips[reason].append(ident) elif case.find("failure") is not None or case.find("error") is not None: failed += 1 if case.get("assertions") is not None: counts_assertions = True if case.get("assertions") == "0" and skip is None: zero_assertions.append(ident) print(f"testcases {total} skipped {skipped} failed/errored {failed}") print(f"\n== skip reasons ({len(skips)} distinct)") print("A reason that is constant in the gate environment (build profile, unset env var,") print("tool the gate never installs) makes every test under it dead in the gate.") for reason, cases in sorted(skips.items(), key=lambda kv: -len(kv[1]))[: args.limit]: print(f"{len(cases):5d} {reason}") for ident in cases[:3]: print(f" e.g. {ident}") if counts_assertions: print(f"\n== executed tests reporting zero assertions ({len(zero_assertions)})") for ident in zero_assertions[: args.limit]: print(f" {ident}") unmatched: list[str] = [] if args.test_files: pattern_files = [(pattern, expand([pattern])) for pattern in args.test_files] files = sorted({name for _, names in pattern_files for name in names}) unmatched = [pattern for pattern, names in pattern_files if not names] for pattern in unmatched: print(f"junit_census: unmatched pattern (no eligible files): {pattern}", file=sys.stderr) unseen = [] for name in files: path = Path(name) stem = normalise(str(path.with_suffix(""))) candidates = [bounded(c) for c in {stem, normalise(path.name), normalise(path.stem)}] if not any(c.search(m) for c in candidates for m in mentioned): unseen.append(name) print(f"\n== test files on disk never mentioned by any report ({len(unseen)} of {len(files)})") print("Check runner include/exclude config, naming conventions, and suite lists.") for name in unseen[: args.limit * 5]: print(f" {name}") return 2 if unmatched else 0 if __name__ == "__main__": sys.exit(main()) -
pytest_reach_probe.py 4.7 KB
"""Pytest plugin: record every nonzero subprocess result per test. Answers "which guard actually refused?" for suites that drive a CLI through `subprocess` (`run`, or `Popen(...).communicate()`). Each nonzero result becomes one JSON line: test node id, argv[0] basename, exit code, and the first and last non-empty stderr lines. PYTHONDONTWRITEBYTECODE=1 PYTHONPATH=<skill>/scripts REACH_PROBE_OUT=/tmp/reach.jsonl \ pytest -p pytest_reach_probe <candidate tests> python3 <skill>/scripts/pytest_reach_probe.py /tmp/reach.jsonl [--grep TEXT] The second form prints, per test, the LAST nonzero call (usually the one the assertion checks) and its last stderr line (the refusal, after any warnings), so each can be compared with the guard the test is named for. The plugin only observes; it never changes a result. The output file is truncated once per session, so rows from an earlier run never mix in. Without REACH_PROBE_OUT it is reach-probe-<pid>.jsonl in the system temp directory, and the path is printed in the terminal summary. It works under xdist: the controller truncates and picks the path, and each worker appends its own lines. A test module that binds `from subprocess import Popen` before the plugin loads still goes through the patched method, because the patch is on the class. """ from __future__ import annotations import json import os import subprocess import tempfile from pathlib import Path _current: dict[str, str | None] = {"nodeid": None} _original_communicate = subprocess.Popen.communicate def _refusal_lines(data: object) -> tuple[str, str]: """First and last non-empty stderr lines; a refusal usually ends the stream after warnings.""" if isinstance(data, bytes): data = data.decode("utf-8", "replace") if not isinstance(data, str): return "", "" lines = [line.strip() for line in data.splitlines() if line.strip()] return (lines[0][:400], lines[-1][:400]) if lines else ("", "") def _recording_communicate(self, *args, **kwargs): result = _original_communicate(self, *args, **kwargs) if self.returncode not in (0, None) and _current["nodeid"] is not None: argv = self.args first, last = _refusal_lines(result[1] if isinstance(result, tuple) else None) head = argv[0] if isinstance(argv, list | tuple) and argv else argv row = { "test": _current["nodeid"], "program": os.path.basename(os.fsdecode(head)) if isinstance(head, str | bytes | os.PathLike) else str(head), "rc": self.returncode, "stderr_first": first, "stderr_last": last, } with _output_path().open("a", encoding="utf-8") as handle: handle.write(json.dumps(row) + "\n") return result def _output_path() -> Path: return Path(os.environ.get("REACH_PROBE_OUT") or Path(tempfile.gettempdir()) / f"reach-probe-{os.getpid()}.jsonl") def pytest_configure(config): if not hasattr(config, "workerinput"): # xdist workers inherit the environment, so they append to the controller's file. os.environ["REACH_PROBE_OUT"] = str(_output_path()) flags = os.O_WRONLY | os.O_CREAT | os.O_TRUNC | getattr(os, "O_NOFOLLOW", 0) os.close(os.open(os.environ["REACH_PROBE_OUT"], flags, 0o600)) subprocess.Popen.communicate = _recording_communicate def pytest_terminal_summary(terminalreporter, exitstatus, config): if not hasattr(config, "workerinput"): terminalreporter.write_line(f"reach probe rows: {os.environ.get('REACH_PROBE_OUT')}") def pytest_unconfigure(config): subprocess.Popen.communicate = _original_communicate def pytest_runtest_setup(item): _current["nodeid"] = item.nodeid def pytest_runtest_teardown(item, nextitem): _current["nodeid"] = None def _report() -> int: import argparse parser = argparse.ArgumentParser(description="summarise a reach-probe JSONL file") parser.add_argument("jsonl") parser.add_argument("--grep", help="only tests whose id contains TEXT") args = parser.parse_args() last: dict[str, dict] = {} calls: dict[str, int] = {} for line in Path(args.jsonl).read_text(encoding="utf-8").splitlines(): row = json.loads(line) last[row["test"]] = row calls[row["test"]] = calls.get(row["test"], 0) + 1 for test, row in last.items(): if args.grep and args.grep not in test: continue print( f"{test}\n rc {row['rc']} {row['program']} ({calls[test]} nonzero calls)\n last: {row['stderr_last']}" ) if row["stderr_first"] != row["stderr_last"]: print(f" first: {row['stderr_first']}") return 0 if __name__ == "__main__": raise SystemExit(_report()) -
weak_negatives.py 14.5 KB
#!/usr/bin/env python3 """Flag negative tests whose assertion cannot tell which guard refused. Each hit is a candidate for Ineffective for the claim (a test that may pass without reaching the guard it names). Confirm hits with a reach probe or a mutation check before acting on them. Kinds: status-only asserts a nonzero exit code or 4xx/5xx status and nothing that names the refusal (message, error type plus message, validation key, body) or-alternative accepts either of two messages, so the earlier guard satisfies it bare-throw expects a generic exception type (or any error, in JS) with no message or matcher, and nothing else in the test names the refusal should-panic Rust #[should_panic] with no expected = "..." Languages by extension: .py (AST), .rs, .php (PHPUnit and Pest), .js/.jsx/.ts/.tsx/.mjs/.cjs/.mts/.cts. """ from __future__ import annotations import argparse import ast import glob import re import sys from collections import Counter from pathlib import Path SKIP_DIRS = { "node_modules", "vendor", ".git", "target", ".venv", "venv", "__pycache__", "dist", "build", "coverage", ".claude", ".next", ".turbo", ".worktrees", } def expand(patterns: list[str]) -> list[str]: """Glob, dropping dependency, build, and agent-worktree directories.""" found = {f for p in patterns for f in glob.glob(p, recursive=True)} return sorted(f for f in found if Path(f).is_file() and not SKIP_DIRS.intersection(Path(f).parts)) Hit = tuple[str, int, str, str, str] GENERIC_ERRORS = { "Exception", "BaseException", "RuntimeException", "RuntimeError", "InvalidArgumentException", "LogicException", "ValueError", "TypeError", "KeyError", "AssertionError", "OSError", "SystemExit", "CalledProcessError", "ValidationException", "HttpException", "AuthorizationException", "AuthenticationException", "QueryException", "Throwable", "Error", "ErrorException", "UnexpectedValueException", "DomainException", } PY_STATUS_ATTRS = {"returncode", "status_code", "exit_code", "status", "rc", "code"} PY_DETAIL_WORDS = re.compile( r"stderr|stdout|message|output|match|excinfo|body|error|detail|reason", re.IGNORECASE ) PY_CHECKER_CALL = re.compile(r"assert|expect|refus|reject|deny|denied|fail", re.IGNORECASE) def _is_nonzero_int(node: ast.AST) -> bool: return ( isinstance(node, ast.Constant) and isinstance(node.value, int) and not isinstance(node.value, bool) and node.value != 0 ) def _status_name(node: ast.AST) -> str | None: if isinstance(node, ast.Attribute) and node.attr in PY_STATUS_ATTRS: return node.attr if isinstance(node, ast.Name) and node.id in PY_STATUS_ATTRS: return node.id return None HTTP_STATUS_ATTRS = {"status_code", "status"} SUCCESS_NAME = re.compile(r"(?:^|_)(?:OK|SUCCESS|PASS|CLEAN|ZERO)(?:_|$)") def _is_refusal_value(attr: str, node: ast.AST) -> bool: if attr in HTTP_STATUS_ATTRS: return isinstance(node, ast.Constant) and isinstance(node.value, int) and 400 <= node.value <= 599 if isinstance(node, ast.Name | ast.Attribute): name = ast.unparse(node).split(".")[-1] return name.isupper() and not SUCCESS_NAME.search(name) return _is_nonzero_int(node) def _is_zero(node: ast.AST) -> bool: return isinstance(node, ast.Constant) and node.value == 0 def _is_exit_name(node: ast.AST) -> bool: return _status_name(node) not in (None, *HTTP_STATUS_ATTRS) def _py_status_assert(test: ast.AST) -> bool: if not isinstance(test, ast.Compare) or len(test.ops) != 1: return False left, right, op = test.left, test.comparators[0], test.ops[0] if isinstance(op, ast.Eq): if (attr := _status_name(left)) is not None and _is_refusal_value(attr, right): return True return (attr := _status_name(right)) is not None and _is_refusal_value(attr, left) if isinstance(op, ast.NotEq): return (_is_exit_name(left) and _is_zero(right)) or (_is_exit_name(right) and _is_zero(left)) return False def _py_detail_checked(func: ast.AST) -> bool: for node in ast.walk(func): if isinstance(node, ast.Assert) and not _py_status_assert(node.test): test = node.test membership = isinstance(test, ast.Compare) and any( isinstance(o, ast.In | ast.NotIn) for o in test.ops ) if PY_DETAIL_WORDS.search(ast.unparse(test)) or ( membership and _status_name(test.left) is None and not _is_nonzero_int(test.left) ): return True if isinstance(node, ast.Call): name = ast.unparse(node.func) if "raises" in name and any(k.arg == "match" for k in node.keywords): return True if PY_CHECKER_CALL.search(name) and any( isinstance(n, ast.Constant) and isinstance(n.value, str | bytes) and len(n.value) >= 6 for arg in [*node.args, *(k.value for k in node.keywords)] for n in ast.walk(arg) ): return True return False def scan_python(path: str, text: str) -> list[Hit]: tree = ast.parse(text) hits: list[Hit] = [] for func in ast.walk(tree): if not isinstance(func, ast.FunctionDef | ast.AsyncFunctionDef) or not func.name.startswith("test"): continue status_line = None for node in ast.walk(func): if isinstance(node, ast.Assert): if _py_status_assert(node.test) and status_line is None: status_line = node.lineno t = node.test if isinstance(t, ast.BoolOp) and isinstance(t.op, ast.Or): ins = [ v for v in t.values if isinstance(v, ast.Compare) and any(isinstance(o, ast.In) for o in v.ops) ] if len(ins) >= 2: hits.append((path, node.lineno, "or-alternative", func.name, ast.unparse(t)[:140])) if isinstance(node, ast.With | ast.AsyncWith): for item in node.items: call = item.context_expr if ( isinstance(call, ast.Call) and ast.unparse(call.func).endswith("raises") and not any(k.arg == "match" for k in call.keywords) and call.args and ast.unparse(call.args[0]).split(".")[-1] in GENERIC_ERRORS ): var = item.optional_vars used = var is not None and any( isinstance(n, ast.Name) and n.id == ast.unparse(var) and n is not var for stmt in func.body for n in ast.walk(stmt) if stmt is not node ) if not used: hits.append((path, node.lineno, "bare-throw", func.name, ast.unparse(call)[:140])) if status_line is not None and not _py_detail_checked(func): hits.append( (path, status_line, "status-only", func.name, "status/exit code is the only refusal evidence") ) return hits def _blocks(text: str, start_re: re.Pattern[str]) -> list[tuple[int, str, str]]: starts = [(m.start(), m.group("name")) for m in start_re.finditer(text)] out = [] for i, (offset, name) in enumerate(starts): end = starts[i + 1][0] if i + 1 < len(starts) else len(text) out.append((text.count("\n", 0, offset) + 1, name, text[offset:end])) return out JS_START = re.compile(r"\b(?:it|test)(?:\.(?:only|each\([^)]*\)|concurrent))?\s*\(\s*(['\"`])(?P<name>.+?)\1") JS_BARE_THROW = re.compile(r"\.(?:toThrow|toThrowError)\(\s*\)") JS_STATUS = re.compile( r"expect\([^)]*\.(?:status|statusCode|exitCode|code)\)\s*\.(?:toBe|toEqual|toStrictEqual)\(\s*(?:[45]\d\d|[1-9]\d?)\s*\)" r"|toHaveProperty\(\s*['\"](?:status|statusCode)['\"]\s*,\s*[45]\d\d" r"|\.expect\(\s*[45]\d\d\s*\)(?!\s*\.\s*expect)" r"|toHaveStatus\(\s*[45]\d\d" ) JS_DETAIL = re.compile( r"expect\([^)]*(?:body|message|error|text|json|data|detail|stderr|stdout|errors)\b" r"|\.(?:toThrow|toThrowError|rejects\.toThrow)\(\s*[^\s)]" r"|toMatchObject|toMatchInlineSnapshot|toMatchSnapshot|toContain\(|toMatch\(" r"|toHaveBeenCalledWith\([^)]*['\"`]" ) JS_OR = re.compile(r"expect\([^;]*\|\|[^;]*\)\s*\.(?:toBe\(\s*true|toBeTruthy)") def scan_js(path: str, text: str) -> list[Hit]: hits: list[Hit] = [] for line, name, body in _blocks(text, JS_START): if (m := JS_BARE_THROW.search(body)) and not JS_DETAIL.search(body): hits.append((path, line, "bare-throw", name, m.group(0))) if JS_STATUS.search(body) and not JS_DETAIL.search(body): hits.append((path, line, "status-only", name, JS_STATUS.search(body).group(0)[:140])) if m := JS_OR.search(body): hits.append((path, line, "or-alternative", name, m.group(0)[:140])) return hits PHP_START = re.compile( r"function\s+(?P<name>test\w*)\s*\(" r"|(?:#\[Test\]|@test)[\s\S]{0,200}?function\s+(?P<name2>\w+)\s*\(" r"|^\s*(?:it|test)\s*\(\s*(['\"])(?P<name3>.+?)\3", re.MULTILINE, ) PHP_STATUS = re.compile( r"->assert(?:Status\(\s*[45]\d\d|Forbidden|Unauthorized|NotFound|Unprocessable|BadRequest|Conflict|ServerError|MethodNotAllowed)\b" ) PHP_DETAIL = re.compile( r"assertJsonValidationErrors|assertInvalid|assertJsonPath|assertJsonFragment|assertJson\(|assertExactJson|assertSee" r"|assertSessionHasErrors|expectExceptionMessage|assertStringContainsString|assertJsonStructure|->json\(\s*['\"]" ) PHP_EXPECT = re.compile(r"expectException\(\s*(?P<cls>[\w\\]+)::class") def scan_php(path: str, text: str) -> list[Hit]: hits: list[Hit] = [] starts = [] for m in PHP_START.finditer(text): starts.append((m.start(), m.group("name") or m.group("name2") or m.group("name3"))) for i, (offset, name) in enumerate(starts): end = starts[i + 1][0] if i + 1 < len(starts) else len(text) body, line = text[offset:end], text.count("\n", 0, offset) + 1 if (m := PHP_STATUS.search(body)) and not PHP_DETAIL.search(body): hits.append((path, line, "status-only", name, m.group(0))) for m in PHP_EXPECT.finditer(body): if "expectExceptionMessage" not in body and m.group("cls").split("\\")[-1] in GENERIC_ERRORS: hits.append( (path, line, "bare-throw", name, f"expectException({m.group('cls')}) without a message") ) break for m in re.finditer(r"->(?:throws|toThrow)\(\s*(?P<cls>[\w\\]+)::class\s*\)", body): if m.group("cls").split("\\")[-1] in GENERIC_ERRORS: hits.append( (path, line, "bare-throw", name, f"Pest throws({m.group('cls')}) without a message") ) break return hits RS_START = re.compile(r"#\[test\](?P<attrs>(?:\s*#\[[^\]]*\])*)\s*(?:async\s+)?fn\s+(?P<name>\w+)") RS_IS_ERR = re.compile(r"assert!\(\s*[^;]*?\.is_err\(\)\s*\)|matches!\([^;]*Err\(\s*_\s*\)\s*\)") RS_DETAIL = re.compile( r"unwrap_err\(\)|\.err\(\)|to_string\(\)|expect_err|Err\(\s*[A-Z]\w*(?:::\w+)+|contains\(" ) def scan_rust(path: str, text: str) -> list[Hit]: hits: list[Hit] = [] starts = [(m.start(), m.group("name"), m.group("attrs")) for m in RS_START.finditer(text)] for i, (offset, name, attrs) in enumerate(starts): end = starts[i + 1][0] if i + 1 < len(starts) else len(text) body, line = text[offset:end], text.count("\n", 0, offset) + 1 leading = [] for prior in reversed(text[:offset].splitlines()): if not prior.strip().startswith(("#[", "///")): break leading.append(prior) if "#[should_panic]" in attrs or any("#[should_panic]" in p for p in leading): hits.append((path, line, "should-panic", name, "#[should_panic] without expected =")) if (m := RS_IS_ERR.search(body)) and not RS_DETAIL.search(body): hits.append((path, line, "status-only", name, m.group(0)[:140])) return hits SCANNERS = { ".py": scan_python, ".rs": scan_rust, ".php": scan_php, **{ext: scan_js for ext in (".js", ".jsx", ".ts", ".tsx", ".mjs", ".cjs", ".mts", ".cts")}, } def main() -> int: parser = argparse.ArgumentParser( description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter ) parser.add_argument("patterns", nargs="+", help="globs of test files (** allowed)") parser.add_argument("--kind", action="append", help="only report these kinds") args = parser.parse_args() pattern_files = [(pattern, expand([pattern])) for pattern in args.patterns] matched = sorted({name for _, names in pattern_files for name in names}) unmatched = [pattern for pattern, names in pattern_files if not names] for pattern in unmatched: print(f"weak_negatives: unmatched pattern (no eligible files): {pattern}", file=sys.stderr) files = [name for name in matched if Path(name).suffix in SCANNERS] unsupported = [name for name in matched if Path(name).suffix not in SCANNERS] if not files: print("weak_negatives: no supported files matched; audit input is incomplete", file=sys.stderr) return 2 for name in unsupported: print(f"weak_negatives: unsupported file: {name}", file=sys.stderr) hits: list[Hit] = [] failures = 0 for name in files: scanner = SCANNERS[Path(name).suffix] try: hits.extend(scanner(name, Path(name).read_text(errors="replace"))) except (OSError, SyntaxError) as error: failures += 1 print(f"weak_negatives: unparsed {name}: {error}", file=sys.stderr) if args.kind: hits = [h for h in hits if h[2] in args.kind] for path, line, kind, test, detail in sorted(hits): print(f"{path}:{line} [{kind}] {test} -- {detail}") counts = Counter(h[2] for h in hits) summary = ", ".join(f"{k} {v}" for k, v in sorted(counts.items())) or "none" print( f"\n{len(files) - failures} supported files scanned; {failures} unparsed; " f"{len(unsupported)} unsupported; {len(unmatched)} unmatched patterns; " f"{len(hits)} candidates ({summary})", file=sys.stderr, ) return 2 if failures or unsupported or unmatched else 0 if __name__ == "__main__": sys.exit(main())
-
-
SKILL.md 8.1 KB
--- name: test-audit class: workflow description: >- Audit whether tests detect regressions in the behavior they claim to protect. Find mocked-away subjects, weak or circular assertions, undiscriminating fixtures, swallowed failures, and tests missing from gates. Use for test-quality audits, suspected false confidence in generated tests, and reviewing new tests. Use ia-writing-tests for writing tests. --- # Test audit Find tests that remain green when their claimed behavior breaks. Assess behavior, not apparent authorship, mock count, assertion count, or a target finding percentage. Default to a read-only audit. Change tests or production code only when requested. Adapted from the OpenClaw test-audit skill (MIT). ## What counts as a finding Evaluate each claimed behavior separately. A test can protect one obligation and miss another. Use these dispositions: - **Ineffective for the claim:** the named behavior is replaced, never reached, or disconnected from a verdict that can fail. - **Partially effective:** real behavior is tested, but a specific promised distinction is absent from the input or assertions. Preserve the protection that exists. - **Not enforced:** the test is uncollected, skipped, or absent from a blocking gate. Report local collection and CI enforcement separately from assertion quality. - **Redundant:** another test protects the same behavior at the same boundary under equivalent inputs and setup. This requires separate evidence; weakness alone is not redundancy. - **Adequate for its scope:** the test detects a credible failure of its actual contract. - **Unresolved:** missing context or execution prevents a defensible decision. A missing additional edge case is not automatically a defective existing test. Tie each quality finding to an existing claim, documented contract, or misleading assertion. Distinguish an individual test's weakness from a suite-wide coverage gap. ## 1. Establish scope and execution Read repository instructions, runner configuration, shared fixtures, test helpers, and CI commands. Record the exact focused command and applicable completion gates. Determine which tests are collected, execute, and block a merge. An advisory job is not a blocking gate. A local green result does not establish CI enforcement. Build an inventory from the runner's collection output when practical. Otherwise label the source inventory approximate. Count files, test definitions, and parameter cases separately. Do not treat helpers as tests or silently omit unrecognized syntax. For a bounded request, inspect every in-scope test. For a large suite, declare the review budget and selection method before discovery. Cover distinct layers, major fixtures, and mocking styles. Combine risk-directed selection with a reproducible sample independent of detector hits. Keep those two sample results separate. Record the seed and sample unit if claiming random sampling. Expand recurring patterns into their owner family when practical; report anything left unreviewed. Do not infer suite prevalence from targeted examples. If a prevalence estimate is requested, use a representative sample with an explicit denominator and uncertainty. Do not tune findings to an expected rate such as 10–20%. ## 2. Discover through behavior Read [behavioral-review.md](./references/behavioral-review.md) before auditing test bodies. For each reviewed test, trace: 1. **Claim:** what observable behavior does the name, fixture, or requirement promise? 2. **Real subject:** which production operation actually runs, including setup, imported helpers, autouse fixtures, module mocks, and dependency injection? 3. **Input:** does the fixture create the distinction the claim depends on? 4. **Oracle:** where does the expected answer come from, and which assertion checks it? 5. **Verdict:** will an incorrect result reach the runner as a failure? 6. **Counterexample:** what plausible wrong implementation would still pass this test? Construct a concrete counterexample before declaring an assertion sufficient. Examples include ignoring one filter, selecting the first item instead of the minimum, dropping a forwarded prop, or returning an empty result where only the type is asserted. Prefer a changed observable result over a crash or syntax error. Inspect both positive and negative tests. Trace assertion helpers to their actual checks. Read the relevant implementation and neighboring tests before reporting a gap. For mocks, identify whether the replaced operation owns the claimed behavior or is a collaborator of a real subject. Check arguments, call requirements, returned data, and observable effects at that boundary. When delegation is available, split semantic discovery by owner or directory. Each reviewer discovers independently of detector output. Require exact locations, real subject, surviving regression, retained value, and evidence status. Independently verify the strongest findings before reporting them. ## 3. Use detectors as additional leads Run each script by its path in this skill's directory, not the audited repository's `scripts/`, with the audited repository root as the working directory. Write output to a temporary directory. Use the repository's pinned tooling for runner commands. The scripts require Python 3.11+ and the standard library. | Script | What its results mean | |---|---| | [weak_negatives.py](./scripts/weak_negatives.py) `<test globs>` | Heuristic candidates for ambiguous refusal assertions. Does not assess positive behavior or assertion data flow. | | [duplicate_tests.py](./scripts/duplicate_tests.py) `<test globs> --json` | Structural duplicate/consolidation candidates, not an effectiveness score. | | [junit_census.py](./scripts/junit_census.py) `<report.xml> --test-files <glob>` | Reported skips, possible collection gaps, and PHPUnit zero-assertion leads (adequate when not throwing is the documented contract). | | [guard_pins.py](./scripts/guard_pins.py) `--src <glob> --tests <glob>` | Refusal strings without textual matches; not proof of guard coverage. | | [pytest_reach_probe.py](./scripts/pytest_reach_probe.py) | Runs the target suite with a recording plugin; use a scratch copy. See [stacks.md](./references/stacks.md). | Before trusting any output, read [detectors.md](./references/detectors.md): for the four static detectors, status 2 means known incomplete input, and each detector has blind spots that require source review. Read [stacks.md](./references/stacks.md) for runner/probe commands and duplicate parser limits. Confirm tool flags locally before use. No detector hits means only no matches to that detector's rules; continue the semantic sample. ## 4. Confirm regression sensitivity Static evidence can establish a finding when the data flow is explicit, such as a test asserting a value it assigned itself. Label it **source-confirmed**. Otherwise state the counterexample as **inferred** until executed. Use **execution-confirmed** only when the exact test ran against a verified behavioral change. For high-impact or uncertain findings, use a focused mutation in an isolated scratch copy. Do not mutate the user's checkout for an audit. Keep network calls, paid APIs, and shared databases outside the probe unless separately authorized and isolated. Before running a mutation, read [mutation-and-repair.md](./references/mutation-and-repair.md) for the procedure and how to read a surviving mutant. ## 5. Report actionable evidence For each finding, give: - Exact test name and `file:line`, plus the relevant production location. - Claimed behavior, actual real subject, and input/assertion/verdict defect. - Concrete surviving regression and retained useful protection. - Evidence status, command/results when executed, and unresolved limitations. - Individual-test versus suite scope, neighboring coverage checked, and repair direction. Report reviewed scope and selection method alongside findings. Separate confirmed findings, unresolved candidates, and retained counterexamples. Do not count every missing obligation as a separate bad test. State commands not run and reasons. Before proposing deletion of a production seam, or when changes are authorized, or when authoring a new test, read [mutation-and-repair.md](./references/mutation-and-repair.md). -
SPEC.md 5.3 KB
# ia-test-audit Specification ## Intent `ia-test-audit` is a `workflow`-class skill (a multi-step process producing concrete artifacts). Audit whether tests detect regressions in the behavior they claim to protect. Find mocked-away subjects, weak or circular assertions, undiscriminating fixtures, swallowed failures, and tests missing from gates. Use for test-quality audits, suspected false confidence in generated tests, and reviewing new tests. Use ia-writing-tests for writing tests. ## Scope In scope: - Behaviors described in `SKILL.md` and routed via the should_trigger phrasings in `distillery/tests/fixtures/triggers/ia-test-audit.jsonl`. - Updates to runtime behavior, structure, trigger precision, references, and validation. Out of scope: - Acting as the runtime instructions themselves (those live in `SKILL.md`). - Trigger phrasings already covered by adjacent `ia-*` skills (`validate-plugin` flags >70% description overlap as DUPLICATE_TRIGGER). - <!-- to fill in: domain-specific exclusions when the skill drifts --> ## Trigger Context - Class: `workflow` - Hook regex: `plugins/whetstone/hooks/skill-patterns.sh` -> `SKILL_PATTERNS[ia-test-audit]` - Common requests (from fixture should_trigger): - "audit the tests in the billing module for false confidence" - "do these tests actually catch a regression in the discount logic?" - "assess the test suite for mocked-away subjects" - Should not trigger for (from fixture should_not_trigger): - "write tests for the user service" - "add test coverage for the auth module" - "the test quality is poor, improve it" ## Source And Evidence Model Authoritative sources: - `SKILL.md` -- runtime instructions and reference routing. - `references/*.md` -- bundled supplementary content (4 file(s)). - `scripts/*.py` -- five bundled stdlib-only detectors and probes (`weak_negatives.py`, `duplicate_tests.py`, `guard_pins.py`, `junit_census.py`, `pytest_reach_probe.py`); require Python 3.11+. - `distillery/scripts/test_ia_test_audit_detectors.py` -- pytest CLI tests for the bundled scripts. - `distillery/tests/fixtures/triggers/ia-test-audit.jsonl` -- positive and negative trigger phrasings under regression test. - `plugins/whetstone/hooks/skill-patterns.sh` -- regex pattern that fires this skill. - `distillery/.eval-data/ia-test-audit/` -- harvested session examples (when present). Data that must not be stored in this skill or its references: - Secrets, credentials, tokens. - Machine-specific filesystem paths (`/home/...`, `/Users/...`, `~/ai/...`). The validator (`MACHINE_PATH_LEAK`) flags these as HIGH. - Private URLs, customer data, or unredacted personal information. ### Coverage matrix | Dimension | Status | Evidence | |---|---|---| | Trigger fixtures | complete | distillery/tests/fixtures/triggers/ia-test-audit.jsonl (>=5 should_trigger, >=5 should_not_trigger) | | Hook regex pattern | complete | plugins/whetstone/hooks/skill-patterns.sh (`SKILL_PATTERNS[ia-test-audit]`) | | Reference architecture | complete | 4 file(s) under references/ | | Bundled scripts | partial | 5 scripts under scripts/; CLI tests in distillery/scripts/test_ia_test_audit_detectors.py; next: add CLI tests for guard_pins.py, junit_census.py, and pytest_reach_probe.py | | Real-usage signal | <!-- populated by harvest-sessions when sessions exist --> | distillery/.eval-data/ia-test-audit/ (created by harvest-sessions) | ## Evaluation Lightweight (run on every change): ```bash python3 distillery/scripts/distiller.py validate-plugin --component ia-test-audit python3 distillery/scripts/distiller.py test-triggers --skill ia-test-audit python3 -m pytest -q distillery/scripts/test_ia_test_audit_detectors.py ``` Deeper (when behavior risk warrants): ```bash python3 distillery/scripts/distiller.py dspy-eval ia-test-audit python3 distillery/scripts/distiller.py diagnose-negatives ia-test-audit ``` Acceptance gates: - `validate-plugin --component ia-test-audit` returns 0 HIGH findings. - `test-triggers --skill ia-test-audit` returns F1 = 1.0 with floors of 5 should_trigger and 5 should_not_trigger. - For dspy-eval, the composite score does not regress against the most recent saved baseline (see `distillery/.eval-data/ia-test-audit/history.json`). ## Known Limitations - The scripts require Python 3.11+ (`duplicate_tests.py` uses `ast.TryStar`); under 3.10 it fails with a traceback, not a version message. - `pytest_reach_probe.py` patches only `subprocess.Popen.communicate`; `subprocess.call`, `check_call`, bare `Popen.wait()`, `os.system`, and asyncio subprocesses are not observed. It runs the target suite, so it belongs in a scratch copy. - The detectors are lead generators with documented blind spots (`references/detectors.md`); no hits never bounds the semantic audit. ## Maintenance Notes - Update `SKILL.md` when the runtime workflow, branch conditions, or output contract changes. - Update this `SPEC.md` when intent, scope, evidence model, evaluation gates, or maintenance expectations change. - Update the trigger fixture when adding new positive phrasings, removing stale ones, or expanding scope (the 5/5 floor is a hard validator gate). - Update the hook regex in `skill-patterns.sh` whenever fixture positives expose a missed phrasing; verify F1 = 1.0 with `eval-triggers` before committing. - Run the full release pipeline via `/release` -- never bump versions or update CHANGELOG.md from a per-skill edit.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.