Claude Cursor Skill

audit

Check thesis chapters for consistency before submission — contradictory numbers, terminology drift, and broken cross-references — and read every rewritten or proposed sentence against the one it replaces before anyone sees the rewrite.

LLM Mart · 0 points · 1 views 1 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download yha9806-academic-writing-toolkit-.claude_skills_audit-184e482.zip · 116 KB
Part of yha9806/academic-writing-toolkit — 21 skills

Install

skills CLI npx skills add https://github.com/yha9806/academic-writing-toolkit/tree/main/.claude/skills/audit
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install yha9806-academic-writing-toolkit@llmmart
Git git clone https://github.com/yha9806/academic-writing-toolkit.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole yha9806/academic-writing-toolkit collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

/audit — Thesis Consistency Audit Skill

Running Python helpers

Choose the interpreter before running the examples. For a globally installed copy, use the private runtime recorded by its installer. In a checkout or linked workspace, use AWT_PYTHON when set, otherwise the toolkit's .venv (follow the skill directory link back to the toolkit): Scripts/python.exe on Windows, bin/python on macOS/Linux. Without that environment, check that python (Windows) or python3 (macOS/Linux) actually runs and has the helper's dependencies. Replace the example's python3 with that executable. In PowerShell, prefix a quoted executable with &; keep commands on one line and quote file paths.

A manuscript registered with the writing loop

Run loop coverage <workspace> --run first (experimental/writing-loop/bin/loop). It derives every check's inputs from the workspace (the tracked draft files at HEAD, the bibliography, the ledgers, the target venue's corpus) instead of the thesis defaults below, runs the ones the latest edits made due, and lists the checks that cannot run on this manuscript and why. Its table covers the script checks only: categories A, B and C below are read by the model, leave no run record, and are listed under the table as such. Report any check it shows as not current, missing a prerequisite, not applicable or failed, and never report a check it did not run as clean.

Purpose

Scan all thesis chapters for internal data consistency issues: contradictory numbers, inconsistent terminology, broken cross-references, and arithmetic errors. This is a pre-submission quality check.

Trigger Words

This skill activates on: audit, consistency check, check numbers, /audit.

Workflow

  1. Scan all chapter files in the chapters/ directory using Glob. Read each file to extract quantitative claims, terminology, and cross-references.

  2. Check the following categories:

    A. Numerical consistency — read by the model, not checked by a script, except for what audit H binds. The three bullets below are what to look for; nothing verifies that you looked.

    • The same statistic (e.g., accuracy, sample size, p-value) cited in multiple chapters must have the same value.
    • Percentages in a distribution must sum to 100% (with tolerance of +/-1% for rounding).
    • Counts (e.g., "42 models") must match between chapters.

    For numbers that matter, use audit H instead: it binds the printed value to the artifact it came from, so the check survives the next edit.

    B. Terminological consistency

    • The same concept must use the same term throughout. Flag cases where synonyms are used inconsistently (e.g., "structured review" vs "systematic review" for the same concept).
    • Abbreviations must be defined on first use in each chapter.

    C. Cross-reference validity

    • References to other sections (e.g., "as discussed in Section 3.2") must point to sections that exist.
    • References to tables and figures must match actual table/figure numbers.
    • Forward references ("Chapter 6 will show...") must be fulfilled.

    D. Citation checks — disabled in this release (disclosed gap)

    The deterministic citation tiers previously run here are disabled: measured against realistic thesis text they produced false high-severity "phantom citation" findings on ordinary parentheticals, missed multi-word institutional authors, and flagged the comma form that Cite Them Right Harvard mandates. Until the checker meets a measured, disclosed false-positive rate, do not run it and do not present citation consistency as audited. Reference integrity is still covered by /verify-refs (BibTeX records) and by the notes-file contract lint.

    E. Claim positioning (deterministic, runs before F and G — positioning is not repairable after review; style is)

    python3 .claude/skills/audit/scripts/audit-claim-positioning.py --base-dir chapters --bib references.bib --json
    

    (omit --bib when the project has no bibliography file). Report every issue it returns: unsourced-keyword and bare-novelty as High — a field's vocabulary in use without its literature, or a novelty claim in a paragraph that shows no search — uncited-method and dangling-entry as Medium.

    Measured precision, one manuscript, 2026-09-20: the checker returned 11 issues and 9 were false, all from three mechanisms now fixed — a \keywords{} block wrapped across source lines so a term matched only its own declaration; the bare word "novelty" inside the manuscript's own method name, and inside two sentences that refuse the claim; and a natbib optional argument (\citep[p.~12]{key}) that made the citation invisible, so a method cited in its own sentence was reported as uncited. After the fixes it returns the 2 that reading had already confirmed. That is n=1 and not a rate — but it is the reason to read every issue against the source before writing it into a report at High, exactly as category D's disabled tiers required.

    For a requested claim-scope or contribution review, consult references/argument-licence/argument-level-lock.md. Separate the field gap, delivered contribution, observed finding and extrapolation; report the six-line claim licence and any unsupported transition with a text or evidence anchor. This is an Advisory reading task. Do not create standing CSV ledgers or infer scientific validity from a checker result. Existing legacy packets can be interpreted with the schema and checks linked in references/argument-licence/README.md.

    H. Number ledger — is this the artifact's number, and is it scoped?

python3 .claude/skills/audit/scripts/audit-number-ledger.py --base-dir . --ledger numbers.tsv --json

A row binds one printed value to the file it came from: printed, in_artifact, scope, artifact, locator. Two value columns because they differ in practice — a figure writing 0.635 against prose printing 63.5\% — and the relation is recorded, never inferred. Report locator-not-in-artifact, value-not-in-locator, printed-artifact-mismatch, number-not-in-manuscript and scope-missing as Critical; unledgered-number is a coverage list, not a finding. An empty ledger exits 2.

scope is matched literally: pick a token the correct sentences contain, or -. On a real manuscript the first token tried flagged two sentences that carry the scope in other words; technique passed both.

I. Method ledger — is each sentence about what was done still bound to where it was done?

python3 .claude/skills/audit/scripts/audit-method-ledger.py --base-dir . --ledger method-ledger.tsv \
    --repo e0=~/dev/experiments@<commit> --full sections/03_data.tex --git . --state <ids.json> --json

For sentences with no claim, number or citation: methods, data, design, what a figure shows. A row (id loc sentence claim_type pointer evidence verdict note checked) points at code, config or output in a repository pinned at a commit. Pointers are read out of the cell's text: path, path:12-20 (pinned code only), path#a.b[*].c (resolved in the JSON or YAML), path@"fragment", \label{x}, commit:<sha>. Report sentence-changed, pointer-missing, pointer-by-line, no-pointer, open-verdict, retired-but-present and row-dropped as High; note-contradicts-verdict and unledgered as Medium. partial rows are pending, neither pass nor fail. A deleted sentence's row is kept as retired <commit>, never removed. --migrate-headings prints a diff that moves run-in \paragraph headings out of sentence cells and writes nothing. Nothing to examine exits 2.

G. Claim ledger — does a LaTeX manuscript's claim match its archived source?

python3 .claude/skills/audit/scripts/audit-claim-ledger.py --base-dir . --ledger ledger.tsv --json

For manuscripts kept as .tex (which audit F does not read). The ledger is a TSV with the columns claim, cite_key, snippet, source_file, level: one row binds one manuscript sentence to one verbatim snippet in one archived source. Report snippet-not-in-source, claim-not-in-manuscript, key-not-in-claim-sentence, source-file-missing and negative-claim-without-fulltext as High; unledgered-assertion as Medium; qualifier-dropped and unledgered-credit are prompts, not findings. Exit 2 means no ledger row was checked at all — report that as not audited, never as clean. --pairs prints the claim/snippet pairs so a reader can judge what the audit does not: whether a claim says more than its snippet.

F. Citation fidelity — does the citing sentence match its source?

node .claude/skills/audit/scripts/audit-citation-fidelity.mjs --base-dir . --json

Exit 2 with nothing_checked: true means the audit found no citation under chapters/**/*.md — report it as not audited, never as clean. For a LaTeX manuscript, not_covered counts the \cite commands, distinct keys and files this audit does not read — under the whole --base-dir, so point it at the manuscript directory rather than at a repository that also holds archived copies, or the count will be a multiple of the truth. --allow-empty accepts an empty workspace on purpose. An empty findings list is only a result when corpus_files and sentences_checked say something was read. Report quote-not-in-source and page-mismatch as High — a quoted span that is not verbatim in the source's notes or PDF, or a page that the source contradicts — and notes-missing as Medium. low-overlap is experimental: list it under Measurements as a prompt to re-read, never as an issue; no false-positive rate has been measured for it yet. State the tool's own limit in the report verbatim: it does not detect a sentence that inverts its source in the source's own words — the failure that mattered most on a real manuscript — and that still requires reading. Every finding here is a proxy; a finding is a reason to open the source, not a verdict.

G. Prose fingerprint (measurement only)

With no baseline corpus, report G as not audited and say which corpus is missing. Do not leave it out of the report: an absent section reads as a pass. The bibliography baseline needs the project's own reference PDFs (literature/, twenty or more, the author's own papers excluded):

python3 .claude/skills/audit/scripts/audit-prose-fingerprint.py --target chapters --baseline literature --exclude '<author-surname>*'

The tool now checks that precondition itself rather than trusting the flag: it measures how much of the target's 4-gram vocabulary each baseline document contains, and a document that looks like a draft or a copy of the target is named in baseline_suspect. When one is found it withholds every percentile and exits 2 — not 1, which means "measured, something is out of range". Exclude the file, or pass --allow-overlap if the overlap is intended; the waiver restores the percentiles but still reports preconditions_checked: false, because a waiver is not a met precondition.

Two further fields have to be read before the percentiles mean anything. baseline_too_short names every document that loaded but fell under the 1500-word floor, so baseline_documents plus the skipped, too-short and excluded lists account for every candidate in the directory; a count that does not close means the corpus is not what you think it is. pipeline_mismatch is true when the target is read as stripped markup and the baseline as printed PDF, which is what the documented invocation above does: a percentile then compares two readings, not two documents. When the target's own build sits beside it the tool measures both readings and puts them in pipeline_cross_check. On one real manuscript the word count -- the denominator of every per-1k rate -- differed by 40% between them and sentence-length lag-1 changed sign, so quote the cross-check rather than treating the percentiles as exact.

Two baselines, and a percentile that does not name its own is not a result. The corpus above is the project's bibliography: it answers whether the prose sits inside the literature the manuscript argues with. Build the other one before drafting, not at polish time:

python3 .claude/skills/audit/scripts/build-venue-baseline.py --venue "<journal>" --from-year 2021 --out <corpus-dir> --manifest <repo>/venue_manifest.json

It admits a record only when its DOI resolves to the venue's registered container-title, writes every candidate's disposition so the manifest's arithmetic closes, and exits 2 with VENUE_CORPUS_TOO_SMALL rather than returning a short corpus that reads like a complete one. Run the fingerprint a second time against it, with --target pointing at the manuscript's built PDF — against a venue corpus both sides go through the same extraction, which is the one case where pipeline_mismatch can be false. Read it once, to answer whether the manuscript reads like the venue's genre. Do not edit to move a venue percentile; the stop rule is unchanged.

Report the distributions under Measurements, never as issues: this is Advisory by nature. Out-of-range is the hard signal, a percentile is a soft one, and clustering matters more than count. Always report preconditions_checked beside them: outliers: [] on an unverified baseline says nothing, and reading it as a pass is the failure this scan was added for. Method and stop rules: references/prose-polish-method.md.

G2. Rewritten sentences, one at a time (before anyone reads a proposal)

The fingerprint and structure audits measure whole documents, so a handful of rewrites cannot move them, and neither reads a proposal that has not been applied. Every proposed rewrite, and every sentence a correction adds, goes through this first, as an id, old, new TSV (leave old empty for an added sentence; the file is read without quoting):

python3 .claude/skills/audit/scripts/audit-sentence-changes.py --pairs <rewrites.tsv> --baseline <venue-corpus-dir>

For a rewrite it names what was added: three or more words, a comma (when punctuation as a whole grew), a colon, semicolon, dash or parenthesis, a subordinate clause (counted as gained, so trading "because" for "which" counts), an adverb or a modifier (counted net, so a term swap does not), two prepositional phrases, a subordinate opener, or a merge of sentences. A split is judged as the old sentence against all its pieces. An added sentence is flagged for any colon, semicolon, dash or worded parenthesis and for density above the venue's 75th percentile. Modifiers are a stand-in for a part-of-speech tagger; the report's limits line says what it cannot see. A correction that stays faithful to its source by piling qualifiers onto the old sentence is the usual cause; split it or restructure it, and re-run until nothing is flagged or each remaining flag is one you can defend. Show the author the rewrites only after that.

Keep the file after the author has read it, with two more columns: verdict (accepted, or rejected; revised counts as rejected; empty while not judged) and reason. Run the audit on it again and it sets the flags against the verdicts and names the script by its hash. The thresholds were fitted on one round; a later round's verdicts, judged by the version frozen before that round, are their only test.

Once applied, the writing loop runs the same audit on each commit against the last commit at which it flagged nothing, so a round of several commits is read as a whole; the run record and the summary name that base. A flagged sentence keeps the base where it was until the sentence is fixed, or until the author accepts it by moving draft.base_ref forward.

G3. Topic and contribution type against the venue (before choosing it)

The venue corpus of build-venue-baseline.py answers whether the manuscript reads like the venue. It is built from preprints, which skew to the subfields whose authors post them, so it cannot say whether the venue publishes this topic or this kind of contribution. That takes every article the venue published in a window, from the registrar:

python3 .claude/skills/audit/scripts/venue-topic-fit.py fetch --issn <ISSN> --from 2024-01-01 --out <corpus.json>
python3 .claude/skills/audit/scripts/venue-topic-fit.py neighbors --corpus <corpus.json> --title "<title>" --abstract <abstract.txt>
python3 .claude/skills/audit/scripts/venue-topic-fit.py sample --corpus <corpus.json> --n 100 --seed <int> --out <sheet.tsv>
python3 .claude/skills/audit/scripts/venue-topic-fit.py tally --sheet <sheet.tsv>

neighbors gives the nearest articles and where the manuscript's own nearest-neighbour similarity falls among the articles': a low percentile means its wording sits at the edge of what the venue publishes. Read the nearest articles before drawing anything from the number. sample draws titles to code by contribution type (M method, E evaluation or analysis of existing systems, D dataset or benchmark, S study of people or science, R review; mark a doubtful code with ?); tally gives shares with 95% intervals. The coder is whoever codes the sheet, and the report should name them. Three limits go into any report of it: coding from titles is coarse; TF-IDF measures shared words, not fit; and every article in the corpus was accepted, so nothing here is an acceptance rate.

  1. Output the audit report using the format below.

Output Format

## Audit Report -- {YYYY-MM-DD}

### Summary

- **Critical**: {N} issues (contradictory data)
- **High**: {N} issues (broken references, missing definitions)
- **Medium**: {N} issues (terminology inconsistency, minor arithmetic)

### Issues

| # | Severity | Category | Location | Issue | Current | Expected |
|---|----------|----------|----------|-------|---------|----------|
| 1 | Critical | Numerical | Ch3 s3.2, Ch5 s5.4 | Sample size differs | 120 (Ch3) vs 125 (Ch5) | Should be consistent |
| 2 | High | Cross-ref | Ch4 s4.1 | Ref to "Section 3.7" | Section 3.7 | Section does not exist |

### Measurements (category G when a baseline exists; category F's experimental low-overlap prompts)

{Per metric: rate, clustering (gap CV), longest gap — with the baseline's
range and where the manuscript sits. Numbers, not verdicts.}

### Recommendations

{Grouped by severity, brief notes on how to resolve each issue.}

Severity Levels

  • Critical: The same quantitative claim has different values in different chapters. This directly undermines thesis credibility.
  • High: Broken cross-references, undefined abbreviations on first use, missing table/figure numbers.
  • Medium: Inconsistent terminology that does not cause factual error, minor rounding discrepancies within tolerance.

Constraints

  1. Never auto-fix. List all issues for the user to review and decide. The user may choose to fix selectively.
  2. No emoji in output.
  3. Report all instances, not just the first occurrence. If a statistic appears in 4 chapters with 2 different values, list all 4 locations.
  4. Be specific about locations. Provide chapter number, section number, and surrounding context so the user can find the issue quickly.
  5. Do not flag stylistic issues. This skill checks data consistency, not prose quality.
Files (academic-writing-toolkit)
  • scripts
    • audit-citation-fidelity.mjs 13.7 KB · in bundle
    • audit-claim-ledger.py 19.8 KB
      #!/usr/bin/env python3
      """Claim ledger audit for LaTeX manuscripts.
      
          python3 audit-claim-ledger.py --base-dir <manuscript> --ledger <ledger.tsv> [--ledger <more.tsv> ...]
                                        [--also-file <supplement.tex> ...] [--json] [--pairs] [--allow-empty]
      
      The gap this closes: the citation fidelity audit reads `chapters/**/*.md`
      against reading notes. A LaTeX manuscript's claims about its sources were
      checked by nothing at all.
      
      A ledger row binds one manuscript sentence to one verbatim snippet in one
      archived source, and the audit checks that binding in both directions:
      
        snippet-not-in-source          the snippet is not verbatim in the archived source
        claim-not-in-manuscript        the sentence was edited; the binding is stale
        key-not-in-claim-sentence      the row names a key the sentence does not cite
        source-file-missing            the archived source is not on disk
        negative-claim-without-fulltext  "X did not do Y" recorded against an abstract
        unledgered-assertion           a citing sentence that asserts something about
                                       its source and has no ledger row
        qualifier-dropped              PROMPT: a hedge in the snippet that the claim
                                       does not carry ("expected", "some", "may")
      
      What it does NOT do: judge whether a claim says more than its snippet in
      words the snippet never used. Machine judgement of that was measured on a
      small red-check set: with the claim and its passage paired it caught 7 of 8,
      and it missed a dropped "expected" every time. So the audit binds and prompts;
      a reader still decides. `--pairs` prints claim/snippet pairs for that reading.
      
      Commit gate (--gate-since GITREF, --credits FILE)
      ------------------------------------------------
      Run over a whole manuscript the audit produces a long coverage list that nobody
      reads at the moment a citation is written. `--gate-since` narrows it to the
      citing sentences this change added, and makes those hard findings, so the check
      can sit on the commit instead of on the submission.
      
        new-assertion-unledgered     a new sentence asserts something about its source
        new-citation-unaccounted     a new sentence cites a source with no account
        credit-outside-its-procedure a credited key used for a procedure it was not
                                     credited for
      
      `--credits` holds the method credits the author accepts, one per line, as
      `key = the procedure it may be cited for`. The procedure matters: run against
      the six commits that introduced the wrong citations in a real manuscript, a
      key-only allowlist let an equivalence-testing paper through as the source of a
      permutation test, because the key was listed and the sentence read like a
      credit. A bare key (no `=`) restores that hole. A key may be listed on
      several lines, one procedure each; any procedure the sentence names covers it.
      
      The full scan reads the same file: a citing sentence with no ledger row whose
      every key is covered is reported as `credited` and leaves the
      unledgered-assertion count, so a sentence the author has already accepted as a
      method credit stops reading as missing evidence.
      
      Historical check, 2026-09-20: over the four commits that introduced them, the
      gate flags all six wrong citations plus the dropped qualifier, as part of 34,
      14, 6 and 4 flagged sentences respectively. It demands an account; it does not
      judge whether the account is right.
      
      Exit: 1 on a hard finding, 2 when no ledger row was checked at all (an empty
      ledger verifies nothing, whatever the coverage list says) unless --allow-empty,
      0 otherwise. In gate mode a change that added no citation is a pass.
      """
      import argparse
      import json
      import re
      import subprocess
      import sys
      from pathlib import Path
      
      COLUMNS = ["claim", "cite_key", "snippet", "source_file", "level"]
      LEVELS = {"fulltext", "abstract-only", "metadata"}
      # Verbs that make a citing sentence a claim about its source rather than a
      # credit line ("we use X~\cite{y}").
      REPORTING = re.compile(
          r"\b(show|shows|showed|shown|find|finds|found|report|reports|reported|document|documents|documented"
          r"|demonstrate\w*|establish\w*|catalogu\w*|argue\w*|observe[sd]?|note[sd]?|propose[sd]?|examine[sd]?"
          r"|conclude[sd]?|reveal\w*|suggest\w*|claim[sd]?|describe[sd]?|warn[sd]?)\b", re.I)
      NEGATIVE = re.compile(
          r"\b(did not|does not|do not|never|no prior|nobody|none of|neither|is not|are not|was not|were not"
          r"|first to|the only)\b", re.I)
      HEDGE = ["expected", "may", "might", "can", "could", "some", "many", "often", "typically", "likely",
               "possibly", "approximately", "about", "partly", "partially", "only", "largely", "mostly"]
      CITE = re.compile(r"\\[a-zA-Z]*cite[a-zA-Z]*\*?(?:\[[^\]]*\])*\{([^}]*)\}")
      
      
      def clean_tex(text):
          text = re.sub(r"(?m)(?<!\\)%.*$", "", text)
          text = re.sub(r"\\(?:label|ref|eqref|input|include)\{[^}]*\}", " ", text)
          text = re.sub(r"\\(?:section|subsection|subsubsection|paragraph)\*?\{[^}]*\}", " ", text)
          text = re.sub(r"\\(?:emph|textbf|textit|texttt|text)\{([^}]*)\}", r"\1", text)
          text = text.replace("~", " ").replace("--", "-").replace("\\%", "%")
          return re.sub(r"\s+", " ", text)
      
      
      def norm(text):
          """Whitespace, hyphenated line breaks and quotation marks normalised for matching."""
          text = re.sub(r"-\s*\n\s*", "", text)
          text = text.replace("\u2019", "'").replace("\u2018", "'").replace("\u201c", '"').replace("\u201d", '"')
          text = text.replace("--", "-").replace("\u2013", "-").replace("\u2014", "-")
          return re.sub(r"\s+", " ", text).strip().lower()
      
      
      def tex_files(base):
          return [p for p in sorted(base.rglob("*.tex"))
                  if not any(part.startswith(".") or part in {"node_modules", "build"} for part in p.relative_to(base).parts)]
      
      
      def split_cited(text):
          """[(sentence, keys)] for every sentence of `text` that carries a \\cite."""
          out = []
          for sentence in re.split(r"(?<=[.])\s+(?=[A-Z\\])", clean_tex(text)):
              keys = [k.strip() for group in CITE.findall(sentence) for k in group.split(",") if k.strip()]
              if keys:
                  out.append((sentence.strip(), keys))
          return out
      
      
      def manuscript_files(base, also=()):
          """[(path, label)]: every .tex under base, then each --also-file (a supplement outside base, say), each once.
          09-27: moving text into a supplement at the repository root took it out of the base directory, and 13 ledger
          rows read as edited away while the sentences were only elsewhere."""
          out = [(p, str(p.relative_to(base))) for p in tex_files(base)]
          seen = {p.resolve() for p, _ in out}
          for x in also:
              p = Path(x).expanduser().resolve()
              if p not in seen:
                  seen.add(p)
                  out.append((p, str(x)))
          return out
      
      
      def citing_sentences(base, also=()):
          """[(file, sentence, keys)] for every sentence that carries a \\cite."""
          out = []
          for path, label in manuscript_files(base, also):
              for sentence, keys in split_cited(path.read_text(encoding="utf-8", errors="replace")):
                  out.append((label, sentence, keys))
          return out
      
      
      def sentences_at_ref(base, ref, also=()):
          """Normalised citing sentences as they stood at `ref`.
      
          Files that did not exist there contribute nothing, so every sentence of a
          newly added file counts as new.
          """
          top = subprocess.run(["git", "-C", str(base), "rev-parse", "--show-toplevel"],
                               capture_output=True, text=True)
          if top.returncode:
              sys.exit(f"GATE_NOT_A_REPO: {base} is not inside a git repository")
          root = Path(top.stdout.strip())
          if subprocess.run(["git", "-C", str(root), "rev-parse", "--verify", "--quiet", f"{ref}^{{commit}}"],
                            capture_output=True).returncode:
              sys.exit(f"GATE_BAD_REF: {ref} is not a commit in {root}")
          before = set()
          for path, _ in manuscript_files(base, also):
              rel = path.resolve().relative_to(root)
              shown = subprocess.run(["git", "-C", str(root), "show", f"{ref}:{rel}"],
                                     capture_output=True, text=True)
              if shown.returncode:
                  continue
              before.update(norm(s) for s, _ in split_cited(shown.stdout))
          return before
      
      
      def read_credits(path):
          """key -> [procedure, ...] as the author accepted them; an empty procedure accepts any use."""
          credits = {}
          for line in Path(path).expanduser().read_text(encoding="utf-8").splitlines():
              line = line.split("#")[0].strip()
              if not line:
                  continue
              key, _, procedure = line.partition("=")
              credits.setdefault(key.strip(), []).append(norm(procedure))
          return credits
      
      
      def uncovered(keys, sentence, credits):
          """The keys of `sentence` that no accepted credit covers.
      
          Match the procedure against the prose only. Key names carry the procedure's
          own words (lakens2017equivalence), so leaving the \\cite in would let a key
          vouch for itself."""
          prose = norm(CITE.sub(" ", sentence))
          return [k for k in keys if not any(p == "" or p in prose for p in credits.get(k, []))]
      
      
      def read_ledger(path):
          rows = []
          lines = [l.rstrip("\n") for l in path.read_text(encoding="utf-8").splitlines() if l.strip()]
          if not lines:
              return rows
          header = lines[0].split("\t")
          if header[:len(COLUMNS)] != COLUMNS:
              sys.exit(f"LEDGER_COLUMNS: expected {COLUMNS}, found {header}")
          for n, line in enumerate(lines[1:], 2):
              parts = line.split("\t")
              if len(parts) < len(COLUMNS):
                  sys.exit(f"LEDGER_COLUMNS: line {n} has {len(parts)} columns, expected {len(COLUMNS)}")
              row = dict(zip(COLUMNS, parts))
              row["line"] = n
              if row["level"] not in LEVELS:
                  sys.exit(f"LEDGER_LEVEL: line {n} has level {row['level']!r}, expected one of {sorted(LEVELS)}")
              rows.append(row)
          return rows
      
      
      def main(argv=None):
          ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
          ap.add_argument("--base-dir", default=".")
          ap.add_argument("--ledger", required=True, action="append",
                          help="a ledger; give it more than once (the text's ledger, a supplement's), and the rows are read together")
          ap.add_argument("--also-file", action="append", default=[], metavar="FILE",
                          help="a manuscript file outside --base-dir to read too (a supplement at the repository root)")
          ap.add_argument("--json", action="store_true")
          ap.add_argument("--pairs", action="store_true", help="print claim/snippet pairs for reading")
          ap.add_argument("--allow-empty", action="store_true")
          ap.add_argument("--gate-since", metavar="GITREF",
                          help="only citing sentences that are new since GITREF are hard findings")
          ap.add_argument("--credits", metavar="FILE",
                          help="method credits the author has accepted, one per line, as "
                               "'key = the procedure it may be cited for'; a bare key accepts any use. "
                               "Read by the full scan and by the commit gate")
          a = ap.parse_args(argv)
          base = Path(a.base_dir).expanduser().resolve()
          ledger_paths = [Path(x).expanduser().resolve() for x in a.ledger]
          ledger_path = ledger_paths[0]
          if not base.is_dir():
              sys.exit(f"BASE_MISSING: {base}")
          for lp in ledger_paths:
              if not lp.is_file():
                  sys.exit(f"LEDGER_MISSING: {lp}")
          for x in a.also_file:
              if not Path(x).expanduser().is_file():
                  sys.exit(f"ALSO_FILE_MISSING: {x}")
      
          sentences = citing_sentences(base, a.also_file)
          rows = []
          for lp in ledger_paths:
              for row in read_ledger(lp):
                  row["_ledger"] = lp
                  rows.append(row)
          credits = read_credits(a.credits) if a.credits else {}
          findings, pairs = [], []
      
          for row in rows:
              loc = f"{row['_ledger'].name}:{row['line']}"
              claim_n = norm(row["claim"])
              bound = [(f, s, k) for f, s, k in sentences if claim_n and claim_n in norm(s)]
              if not bound:
                  findings.append({"kind": "claim-not-in-manuscript", "location": loc, "cite_key": row["cite_key"],
                                   "detail": f"the ledger's claim is not a sentence in the manuscript: {row['claim'][:90]}"})
              elif row["cite_key"] not in {k for _, _, keys in bound for k in keys}:
                  findings.append({"kind": "key-not-in-claim-sentence", "location": loc, "cite_key": row["cite_key"],
                                   "detail": f"the bound sentence does not cite {row['cite_key']}"})
              source = (base / row["source_file"]) if not Path(row["source_file"]).is_absolute() else Path(row["source_file"])
              if not source.is_file():
                  source = row["_ledger"].parent / row["source_file"]
              if not source.is_file():
                  findings.append({"kind": "source-file-missing", "location": loc, "cite_key": row["cite_key"],
                                   "detail": f"archived source not found: {row['source_file']}"})
              else:
                  text = norm(source.read_text(encoding="utf-8", errors="replace"))
                  if norm(row["snippet"]) not in text:
                      findings.append({"kind": "snippet-not-in-source", "location": loc, "cite_key": row["cite_key"],
                                       "detail": f'"{row["snippet"][:90]}" is not verbatim in {row["source_file"]}'})
              if NEGATIVE.search(row["claim"]) and row["level"] != "fulltext":
                  findings.append({"kind": "negative-claim-without-fulltext", "location": loc, "cite_key": row["cite_key"],
                                   "detail": f"a negative claim recorded against level={row['level']}; read the full text or scope the sentence to what was read"})
              dropped = [h for h in HEDGE
                         if re.search(rf"\b{re.escape(h)}\b", row["snippet"], re.I)
                         and not re.search(rf"\b{re.escape(h)}\b", row["claim"], re.I)]
              if dropped:
                  findings.append({"kind": "qualifier-dropped", "prompt": True, "location": loc, "cite_key": row["cite_key"],
                                   "detail": f"the snippet hedges with {dropped}; the claim does not — read the pair before trusting it"})
              pairs.append({"cite_key": row["cite_key"], "claim": row["claim"], "snippet": row["snippet"],
                            "source_file": row["source_file"], "level": row["level"]})
      
          ledgered = {norm(r["claim"]) for r in rows if norm(r["claim"])}
          for f, s, keys in sentences:
              if any(c and c in norm(s) for c in ledgered):
                  continue
              if credits and not uncovered(keys, s, credits):
                  findings.append({"kind": "credited", "prompt": False, "location": f, "cite_key": ",".join(keys),
                                   "sentence": s, "detail": f"accepted as a method credit for {','.join(keys)}: {s[:90]}"})
                  continue
              kind = "unledgered-assertion" if REPORTING.search(s) else "unledgered-credit"
              findings.append({"kind": kind, "prompt": kind == "unledgered-credit", "location": f,
                               "cite_key": ",".join(keys), "sentence": s,
                               "detail": f"{'asserts something about' if kind == 'unledgered-assertion' else 'credits'} {','.join(keys)} with no ledger row: {s[:90]}"})
      
          gate = None
          if a.gate_since:
              # A credit is accepted for a named procedure, not for a key outright.
              # Run against the commit that introduced them, a key-only allowlist let
              # an equivalence-testing paper through as the source of a permutation
              # test: the key was listed, and the sentence read like a credit.
              before = sentences_at_ref(base, a.gate_since, a.also_file)
              added = [(f, s, k) for f, s, k in sentences if norm(s) not in before]
              for f, s, keys in added:
                  if any(c and c in norm(s) for c in ledgered):
                      continue
                  off = uncovered(keys, s, credits)
                  if not off:
                      continue
                  miscredited = [k for k in off if k in credits]
                  if miscredited:
                      kind, why = "credit-outside-its-procedure", (
                          "; ".join(f"{k} is accepted for \"{' / '.join(credits[k])}\", which this sentence does not mention"
                                    for k in miscredited))
                  else:
                      kind = "new-assertion-unledgered" if REPORTING.search(s) else "new-citation-unaccounted"
                      why = "no ledger row and no credit entry"
                  findings.append({"kind": kind, "location": f, "cite_key": ",".join(off),
                                   "detail": f"new since {a.gate_since}, {why}: {s[:90]}"})
              gate = {"since": a.gate_since, "new_citing_sentences": len(added),
                      "credits_file": a.credits, "credits": {k: v for k, v in sorted(credits.items())}}
      
          hard_kinds = {"snippet-not-in-source", "claim-not-in-manuscript", "key-not-in-claim-sentence",
                        "source-file-missing", "negative-claim-without-fulltext",
                        "new-assertion-unledgered", "new-citation-unaccounted",
                        "credit-outside-its-procedure"}
          hard = [f for f in findings if f["kind"] in hard_kinds]
          # Nothing verified is not a pass: an empty ledger over a citing manuscript
          # produces a coverage list and no verification at all.
          # In gate mode the question is what this change added, so a change that
          # added no citation is a legitimate pass even with an empty ledger.
          nothing = not rows and not a.gate_since
          payload = {
              "schema_version": 1,
              "base": str(base),
              "ledger": str(ledger_path),
              "ledgers": [str(p) for p in ledger_paths],
              "also_files": list(a.also_file),
              "citing_sentences": len(sentences),
              "ledger_rows": len(rows),
              "gate": gate,
              "credits_file": a.credits,
              "credited_sentences": sum(1 for f in findings if f["kind"] == "credited"),
              "findings": findings,
              "hard_finding_count": len(hard),
              "nothing_checked": nothing,
              "limits": {
                  "semantic_overreach": "NOT decided here: a claim that says more than its snippet in words the snippet never used passes every check. Measured on a small red-check set, a paired machine judge caught 7 of 8 and missed a dropped 'expected' every time; read the pairs.",
                  "coverage": "unledgered-assertion uses a reporting-verb heuristic, so it both misses claims phrased without one and flags some credit lines",
              },
          }
          if a.json:
              print(json.dumps(payload, indent=2, ensure_ascii=False))
          else:
              unledgered = sum(1 for f in findings if f["kind"] == "unledgered-assertion")
              print(f"claim ledger: {len(rows)} row(s) against {len(sentences)} citing sentence(s) under {base}")
              print(f"coverage: {unledgered} asserting sentence(s) carry no row, and nothing here checks them")
              if gate:
                  print(f"gate: {gate['new_citing_sentences']} citing sentence(s) new since {gate['since']}"
                        f"; {len(gate['credits'])} key(s) accepted as credits")
              for kind in ["new-assertion-unledgered", "new-citation-unaccounted", "credit-outside-its-procedure",
                           "snippet-not-in-source", "claim-not-in-manuscript", "key-not-in-claim-sentence",
                           "source-file-missing", "negative-claim-without-fulltext", "unledgered-assertion",
                           "qualifier-dropped", "unledgered-credit", "credited"]:
                  group = [f for f in findings if f["kind"] == kind]
                  if not group:
                      continue
                  print(f"\n{kind}{' (prompt, not a finding)' if group[0].get('prompt') else ''} ({len(group)})")
                  for f in group:
                      print(f"  {f['location']}\n    {f['detail']}")
              if nothing:
                  print("\nNOTHING CHECKED: no citing sentence and no ledger row. This is not a pass.")
              print(f"\nNot decided here: {payload['limits']['semantic_overreach']}")
          if a.pairs:
              for p in pairs:
                  print(f"\n[{p['cite_key']} · {p['level']}]\n  claim:   {p['claim']}\n  snippet: {p['snippet']}\n  source:  {p['source_file']}")
          return 1 if hard else (2 if nothing and not a.allow_empty else 0)
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • audit-claim-positioning.py 14.3 KB
      #!/usr/bin/env python3
      """Deterministic check that claims and named methods carry sources.
      
      This script does not judge whether a citation is the right one. It finds three
      things a manuscript can be inconsistent about entirely from the inside, without
      any knowledge of the field:
      
        unsourced-keyword   a term the manuscript advertises -- in its keywords, its
                            title, or a contribution statement -- that appears in the
                            body with no citation anywhere near it. A paper whose
                            keywords name a concept while its bibliography contains
                            no source for that concept is using a field's vocabulary
                            without its literature.
      
        uncited-method      a named statistical or experimental procedure used in the
                            body with no citation in the same paragraph. Papers whose
                            contribution is methodological rigour are the ones that
                            most often forget these.
      
        dangling-entry      a bibliography entry nothing cites.
      
        bare-novelty        a novelty claim ("to our knowledge", "we are not aware of
                            prior work") in a paragraph with no citation. The strength
                            of such a claim is the strength of the search behind it,
                            and an unsourced paragraph shows no search.
      
      Python 3.8 stdlib only. Reads LaTeX or Markdown. Citations are recognised as
      LaTeX \\cite{...}, pandoc [@key], and Harvard author-year in prose — "Smith
      (2024)", "(Smith and Doe, 2024, p. 12)" — the two forms the AWT guards accept;
      Markdown advertises terms with a "Keywords:" line.
      """
      
      import argparse
      import json
      import re
      import sys
      from pathlib import Path
      from typing import Dict, Iterable, List, Set, Tuple
      
      SCHEMA_VERSION = 1
      
      # Procedures whose names are proper nouns or fixed terms: if the manuscript runs
      # one, a reader expects a source. Deliberately conservative -- generic words
      # like "mean" or "correlation" are not here.
      METHODS = [
          "Benjamini", "Hochberg", "Bonferroni", "Holm", "Clopper", "Pearson correlation",
          "McNemar", "Friedman", "Wilcoxon", "Kruskal", "Mann-Whitney", "Fisher's exact",
          "Cohen's kappa", "Krippendorff", "Fleiss", "bootstrap", "permutation test",
          "equivalence test", "TOST", "preregistration", "preregistered",
          "Shapiro", "Levene", "ANOVA", "mixed-effects", "Cronbach",
      ]
      # A claim frame asserts priority outright. The bare word does not: it appears
      # in method names, and -- twice on one manuscript -- inside a sentence that
      # REFUSES the claim. The frames
      # carry their own negation ("we are not aware of"), so the guard below applies
      # to the bare word only.
      NOVELTY_CLAIM = re.compile(
          r"\b(?:to our knowledge|to the best of our knowledge|we are not aware of|"
          r"no prior work|first to (?:show|demonstrate|report|propose))",
          re.I,
      )
      # A hyphenated compound names a procedure or an object -- "novelty detection",
      # "novelty-seeking" -- and asserts nothing about priority.
      NOVELTY_WORD = re.compile(r"(?<!-)\bnovel(?:ty)?\b(?!-)", re.I)
      _NEGATOR = re.compile(r"\b(?:not|no|never|nor|without)\b[^.;]{0,60}$", re.I)
      
      
      def novelty_claim(text: str):
          """The first unhedged novelty assertion in `text`, or None."""
          m = NOVELTY_CLAIM.search(text)
          if m:
              return m
          for m in NOVELTY_WORD.finditer(text):
              if not _NEGATOR.search(text[: m.start()]):
                  return m
          return None
      
      
      # The numeric form [12] belongs to rendered text. In LaTeX source it collides
      # with set and interval notation -- $K \in \{1,5,10\}$ matched it, and a
      # manuscript full of such notation looked fully cited when it cited nothing.
      # natbib takes the locator in an optional argument -- \citep[p. 12]{key},
      # \citep[pp.~12--14]{key} -- and requiring "{" straight after the command
      # made exactly those invisible. The more precisely a manuscript cites, the less
      # this saw: on one manuscript a handful of citations vanished and took a finding with them.
      # Written once: the same shape decides whether a paragraph is cited AND which
      # keys count as used, and it lived in two copies. Fixing one left the other,
      # which is how a cited key kept reading as a dangling bibliography entry.
      _CITE_OPT = r"(?:\[[^\]]*\])*"
      CITE_TEX = re.compile(r"\\cite\w*" + _CITE_OPT + r"\{[^}]*\}")
      CITE_KEYS = re.compile(r"\\cite\w*" + _CITE_OPT + r"\{([^}]*)\}")
      CITE_MD = re.compile(r"\[@[^\]]+\]|\[\d+(?:,\s*\d+)*\]")
      # Harvard author-year, the two shapes the AWT guards' extractor accepts (plus
      # an optional "et al."): parenthetical "(Smith, 2024, p. 12)" / "(Smith and
      # Doe, 2024)" and narrative "Smith (2024)". Set notation cannot match: both
      # need a capitalised surname beside a four-digit year.
      _SURNAME = r"[A-Z][A-Za-z'’-]+(?:\s+(?:and|&)\s+[A-Z][A-Za-z'’-]+)?(?:\s+et al\.?)?"
      HARVARD_PAREN = re.compile(r"\((" + _SURNAME + r"),?\s+(\d{4})[a-z]?(?:,\s*pp?\.?\s*[\d,\s–-]+)?\)")
      HARVARD_NARR = re.compile(r"\b(" + _SURNAME + r")\s+\((\d{4})[a-z]?\)")
      CITE = re.compile("|".join([CITE_TEX.pattern, r"\[@[^\]]+\]", HARVARD_PAREN.pattern, HARVARD_NARR.pattern]))
      
      
      def strip_comments(text: str, suffix: str) -> str:
          if suffix == ".tex":
              text = re.sub(r"(?m)(?<!\\)%.*$", "", text)
              # Abstracts do not carry citations by convention; flagging a method
              # named there produces a finding no author can act on.
              text = re.sub(r"\\begin\{abstract\}.*?\\end\{abstract\}", " ", text, flags=re.S)
          return text
      
      
      def paragraphs(text: str) -> List[Dict]:
          out, pos = [], 0
          for block in re.split(r"\n\s*\n", text):
              line = text.count("\n", 0, pos) + 1
              pos += len(block) + 2
              flat = re.sub(r"\s+", " ", block).strip()
              if len(flat) > 40:
                  out.append({"line": line, "text": flat})
          return out
      
      
      def bib_keys(path: Path) -> Set[str]:
          return {e["key"] for e in bib_entries(path)}
      
      
      def bib_entries(path: Path) -> List[Dict]:
          """Each entry's key plus its first author's surname and year, for
          matching Harvard citations that never name a key."""
          if not path.exists():
              return []
          text = path.read_text(encoding="utf-8", errors="replace")
          starts = [m for m in re.finditer(r"@\w+\s*\{\s*([^,\s]+)\s*,", text)]
          out: List[Dict] = []
          for i, m in enumerate(starts):
              body = text[m.end(): starts[i + 1].start() if i + 1 < len(starts) else len(text)]
              author = re.search(r"author\s*=\s*[{\"](.*?)[}\"]\s*,?\s*\n?", body, re.S | re.I)
              year = re.search(r"year\s*=\s*[{\"]?\s*(\d{4})", body, re.I)
              surname = ""
              if author:
                  first = re.split(r"\s+and\s+", author.group(1).strip(), maxsplit=1)[0].strip()
                  surname = first.split(",")[0].strip() if "," in first else first.split()[-1]
                  surname = surname.strip("{} ")
              out.append({"key": m.group(1), "surname": surname.lower(), "year": year.group(1) if year else ""})
          return out
      
      
      def cited_harvard(text: str) -> Set[Tuple[str, str]]:
          pairs: Set[Tuple[str, str]] = set()
          for rx in (HARVARD_PAREN, HARVARD_NARR):
              for m in rx.finditer(text):
                  first = re.split(r"\s+(?:and|&)\s+|\s+et al", m.group(1))[0]
                  pairs.add((first.strip().lower(), m.group(2)))
          return pairs
      
      
      def cited_keys(text: str) -> Set[str]:
          keys = set()
          for m in CITE_KEYS.finditer(text):
              keys |= {k.strip() for k in m.group(1).split(",") if k.strip()}
          return keys
      
      
      def declared_terms(text: str) -> List[str]:
          """Terms the manuscript advertises: keywords, title, and contribution lines.
      
          A real \\keywords{} block wraps across source lines, so a term arrives as
          "label\nvariation" and then matches nothing in the body except the
          declaration it came from -- which sits in the preamble, where there is no
          citation to find. That reported a term used dozens of times, often beside a
          citation, as unsourced."""
          def norm(t: str) -> str:
              return re.sub(r"\s+", " ", t).strip()
          terms: List[str] = []
          for m in re.finditer(r"\\keywords\{([^}]*)\}", text):
              terms += [norm(t) for t in re.split(r"[,;]", m.group(1)) if norm(t)]
          # Markdown advertises terms with a "Keywords:" line (a thesis chapter has
          # no \keywords). Titles are not mined from Markdown headings: "Chapter 1"
          # is not a technical phrase.
          for m in re.finditer(r"(?im)^\**keywords\**\s*[::]\s*(.+?)\s*$", text):
              terms += [norm(t) for t in re.split(r"[,;]", m.group(1)) if norm(t)]
          for m in re.finditer(r"\\title\{([^}]*)\}", text):
              # only multi-word technical phrases from the title, not every word
              terms += [norm(p).strip(" .,:") for p in re.split(r"[:,]", m.group(1)) if len(p.split()) >= 2]
          return [t for t in terms if 1 <= len(t.split()) <= 5]
      
      
      def audit(base: Path, tex_files: List[Path], bib: Path) -> List[dict]:
          whole = ""
          per_file = []
          for p in tex_files:
              t = strip_comments(p.read_text(encoding="utf-8", errors="replace"), p.suffix.lower())
              per_file.append((p, t))
              whole += "\n\n" + t
      
          issues: List[dict] = []
          # A named procedure needs one source in the manuscript, not one per mention.
          sourced_methods: Set[str] = set()
          for para in paragraphs(whole):
              if not CITE.search(para["text"]):
                  continue
              for meth in METHODS:
                  if re.search(r"\b" + re.escape(meth), para["text"], re.I):
                      sourced_methods.add(meth)
          first_use: Dict[str, str] = {}
      
          cited, harvard = cited_keys(whole), cited_harvard(whole)
          for e in bib_entries(bib):
              if e["key"] in cited or (e["surname"] and (e["surname"], e["year"]) in harvard):
                  continue
              issues.append({"kind": "dangling-entry", "location": str(bib.name), "detail": e["key"]})
      
          terms = declared_terms(whole)
          # A declared term always matches its own declaration, so the keyword block,
          # the title and a Markdown "Keywords:" line are not body text for this
          # check. Counting them made a keyword that appears nowhere else look like
          # one used without a source -- a different finding with a different fix,
          # and the declaration sits in the preamble where no citation ever is.
          body = re.sub(r"\\keywords\{[^}]*\}", " ", whole)
          body = re.sub(r"\\title\{[^}]*\}", " ", body)
          body = re.sub(r"(?im)^\**keywords\**\s*[:\uff1a].*$", " ", body)
          body_low = body.lower()
          for term in terms:
              tl = term.lower()
              if body_low.count(tl) == 0:
                  # A term the paper advertises and never uses is a stronger signal
                  # than one used without a source: check its head words instead.
                  heads = [w for w in re.findall(r"[a-z-]{5,}", tl) if w not in
                           {"based", "using", "towards", "における"}]
                  if heads and not any(body_low.count(h) for h in heads):
                      issues.append({"kind": "unsourced-keyword", "location": "keywords/title",
                                     "detail": term + "  (advertised, absent from body)"})
                      continue
                  tl = max(heads, key=len) if heads else tl
              # is there any citation within 400 characters of any mention?
              supported = False
              for m in re.finditer(re.escape(tl), body_low):
                  window = body[max(0, m.start() - 400): m.end() + 400]
                  if CITE.search(window):
                      supported = True
                      break
              if not supported:
                  issues.append({"kind": "unsourced-keyword", "location": "keywords/title",
                                 "detail": term})
      
          for p, t in per_file:
              for para in paragraphs(t):
                  has_cite = bool(CITE.search(para["text"]))
                  loc = "{}:{}".format(p.relative_to(base) if base in p.parents or base == p.parent else p.name,
                                       para["line"])
                  for meth in METHODS:
                      if meth in sourced_methods:
                          continue
                      if re.search(r"\b" + re.escape(meth), para["text"], re.I):
                          first_use.setdefault(meth, loc)
                  m = novelty_claim(para["text"])
                  if m and not has_cite:
                      issues.append({"kind": "bare-novelty", "location": loc,
                                     "detail": para["text"][max(0, m.start() - 30): m.end() + 60]})
          for meth, loc in sorted(first_use.items()):
              issues.append({"kind": "uncited-method", "location": loc, "detail": meth})
          return issues
      
      
      def main() -> int:
          ap = argparse.ArgumentParser(
              description="Audit whether advertised claims and named methods carry sources.")
          ap.add_argument("--base-dir", required=True)
          ap.add_argument("--bib", help="bibliography file (default: first .bib under --base-dir)")
          ap.add_argument("--json", action="store_true", dest="emit_json")
          args = ap.parse_args()
      
          base = Path(args.base_dir)
          if not base.is_dir():
              sys.stderr.write("error: --base-dir is not a directory\n")
              return 2
          SKIP = {"release", "audit", "submission", "ref_pdfs", "tmp", "scripts",
                  ".git", "node_modules", "chapters_backup"}
          files = sorted(p for p in base.rglob("*")
                         if p.suffix.lower() in {".tex", ".md"} and p.is_file()
                         and not (SKIP & set(p.relative_to(base).parts[:-1]))
                         and not p.name.endswith(("_audit.md", "_audit_full.md")))
          if not files:
              sys.stderr.write("error: no .tex or .md files under --base-dir\n")
              return 2
          bib = Path(args.bib) if args.bib else next(iter(sorted(base.rglob("*.bib"))), base / "none.bib")
      
          issues = audit(base, files, bib)
          payload = {"schema_version": SCHEMA_VERSION, "issues": issues, "issue_count": len(issues)}
          if args.emit_json:
              print(json.dumps(payload, indent=2))
          else:
              if not issues:
                  print("no unsourced claims, uncited methods, or dangling entries found")
              for kind in ("unsourced-keyword", "bare-novelty", "uncited-method", "dangling-entry"):
                  group = [i for i in issues if i["kind"] == kind]
                  if not group:
                      continue
                  print("\n{} ({})".format(kind, len(group)))
                  for i in group:
                      print("  {location}: {detail}".format(**i))
              print("\nNote: this checks that a source is present, never that it is the "
                    "right one.\nA term the paper advertises with no citation near it is "
                    "the signal that matters:\nit means the vocabulary of a field is in "
                    "use while its literature is not.")
          return 1 if issues else 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • audit-float-reviews.py 42.4 KB
      #!/usr/bin/env python3
      r"""Say which figures and tables no one has looked at since they last changed.
      
          python3 audit-float-reviews.py --base-dir <manuscript> --main main.tex [--main supplement.tex ...]
                                         --reviews <reviews.tsv> [--json]
          python3 audit-float-reviews.py ... --render --pdf <x.pdf> --aux <x.aux> [--pdf .. --aux ..]
                                         [--built-from <commit>] [--claims <claims ledger>] --out <dir>
      
      The gap this closes. Every check the loop runs reads text. An overflowing label,
      a caption that no longer describes its figure, a figure that contradicts the
      numbers beside it: these are seen only by someone who opens the rendered page,
      and nothing recorded whether anyone had. This cannot judge a figure. It lists
      every float the draft includes with a version fingerprint and says which have
      no review of their current version.
      
      A float is a figure, table or longtable environment (spacing inside \begin
      allowed, and any environment a \newenvironment defines around one), a
      \captionof{figure|table} with the center, minipage or flush environment
      around it (its paragraph when there is none), or a tabular set in the running
      text outside both (id <file>#tabular<k>, by order), in the document body of the
      files reachable from --main through \input, \include, \subfile and \import.
      Not the document: comment and filecontents environments, and a block opened
      by \iffalse at the start of a line up to its matching \fi (TeX's conditionals
      and those a \newif declares are counted; with no matching \fi nothing is
      dropped). Its id is its first \label, in the environment or in a file it
      inputs, or <file>#<n>.
      
      The text is read as TeX reads it: whole-line comments and the text of inline
      comments dropped, a line break a space, a line ending in % joined to the next
      with nothing, a blank line a paragraph break. So re-wrapping a caption keeps a
      review, and a removed % or a new blank line between two panels reopens it.
      
      A float's fingerprint is sha256 over its text and every file it pulls in, by
      path and content: \input, \import (whose directory then leads),
      \includegraphics and \includesvg (the directories of \graphicspath tried
      too; extensions in driver order, the bare name last), and the data files that
      \addplot table (a file, or a table \pgfplotstableread filled anywhere in the
      draft), \pgfplotstableread, \includestandalone, \lstinputlisting,
      \verbatiminput, \csv... and \DTLloaddb read; inline data is not a file. A .tex
      file counts by its text read as above, so a regenerated table whose only
      change is a comment keeps its review.
      
      Each main file's preamble, with every file it pulls in, is reviewed as one item
      of its own (label preamble:<main>): a macro, a length or a font there can
      change how every float prints, and one row to look again says so without
      reopening every float at once. A file named by a macro (\input{\dir/x}), or a
      missing one, in the preamble or the body fails the check rather than being
      skipped. Not followed: what a macro expands to (an image named inside a
      \newcommand), packages and classes.
      
      The review record is a TSV with a header:
      
          label<TAB>fingerprint<TAB>reviewer<TAB>date<TAB>verdict<TAB>note
      
      Only rows whose fingerprint is the item's current one count; the last such row
      decides. A verdict that starts with "fix" keeps the item open.
      
        unreviewed    no row for the float's (or preamble's) current fingerprint
        review-open   the latest current row says fix
        missing-file  the float pulls in a file that is not there
        unfollowed    an \input that names a macro or a missing file: what is behind
                      it is not seen
        stale-row     PROMPT: a row whose label names nothing in the draft
      
      --render writes the PDF page of each float (pdftoppm) and a REVIEW.md sheet
      listing each float's number, page, PNG, fingerprint and caption, with a row to
      fill in, and a row for each preamble. The page is the physical one, from the
      PDF's named destination for the label's anchor (pdfinfo -dests); without one,
      the printed page number is used and the sheet says so. The .aux files are read
      with the \@input files they name. With --built-from, each item's fingerprint
      at that commit is compared with the current one, and one whose rendered page
      shows an older version is marked; without it the PDF's modification time is
      compared with the float's files.
      
      The sheet opens with what to check in each float, and under each float lists
      what it says in words: the sentences of its caption, TikZ nodes, table notes
      and the .tex files it pulls in, with TeX dropped. A sentence that predicts,
      infers or claims a cause is marked, because the overreach scan only knows the
      wordings someone listed and a new wording in a figure goes past it. With
      --claims, the claims ledger's allowed wordings are set out at the top, so each
      marked sentence can be read against what the paper may say.
      
      Exit: 1 on unreviewed, review-open, missing-file or unfollowed; 2 when no float
      is found, the review record is unreadable, or a --pdf or --aux is missing; 0
      otherwise. --render exits 2 when it renders nothing.
      """
      import argparse
      import hashlib
      import json
      import posixpath
      import re
      import shutil
      import subprocess
      import sys
      from pathlib import Path
      
      BASE_ENVS = ["figure", "figure*", "table", "table*", "sidewaysfigure", "sidewaysfigure*", "sidewaystable",
                   "sidewaystable*", "wrapfigure", "wraptable", "SCfigure", "SCtable", "longtable", "longtable*"]
      INPUTS = re.compile(r"\\(?:input|include|subfile)\s*\{([^{}]*)\}|\\input\s+([^\s{}\\%]+)"
                          r"|\\(?:sub)?import\*?\s*\{([^{}]*)\}\s*\{([^{}]+)\}")
      OPT = r"\s*(?:\[[^\]]*\]\s*)*"
      GRAPHICS = re.compile(r"\\include(?:graphics|svg)\*?" + OPT + r"\{([^{}]+)\}")
      ADDPLOT = re.compile(r"\\addplot3?\+?" + OPT + r"table" + OPT + r"\{([^{}]+)\}")
      DATA = [ADDPLOT] + [re.compile(p) for p in (
          r"\\pgfplotstableread" + OPT + r"\{([^{}]+)\}",
          r"\\includestandalone" + OPT + r"\{([^{}]+)\}",
          r"\\lstinputlisting" + OPT + r"\{([^{}]+)\}",
          r"\\verbatiminput\*?\s*\{([^{}]+)\}",
          r"\\csv(?:autotabular|autolongtable|autobooktabular|autobooklongtable|reader)\*?" + OPT + r"\{([^{}]+)\}",
          r"\\DTLloaddb" + OPT + r"\{[^{}]*\}\s*\{([^{}]+)\}")]
      TABLEREAD = re.compile(r"\\pgfplotstableread" + OPT + r"\{([^{}]+)\}\s*\{?\s*\\([A-Za-z@]+)")
      GPATH = re.compile(r"\\graphicspath\s*\{((?:\s*\{[^{}]*\})+)\s*\}")
      LABEL = re.compile(r"\\label\s*(?:\[[^\]]*\])?\s*\{([^{}]+)\}")
      CAPTIONOF = re.compile(r"\\captionof\s*\{(figure|table)\}")
      TABULAR = re.compile(r"\\begin\s*\{(tabular\*?|tabularx|tabulary|NiceTabular\*?)\}")
      DOC = re.compile(r"\\begin\s*\{document\}")
      NEWENV = re.compile(r"\\(?:re)?newenvironment\s*\{([A-Za-z@]+\*?)\}(?:\s*\[[^\]]*\])*\s*\{[^{}]*?\\begin\s*\{([^{}]+)\}")
      NEWIF = re.compile(r"\\newif\s*\\(if[A-Za-z@]+)")
      PRIMITIVE_IFS = {"if", "ifcat", "ifnum", "ifdim", "ifodd", "ifvmode", "ifhmode", "ifmmode", "ifinner", "ifvoid",
                       "ifhbox", "ifvbox", "ifx", "ifeof", "iftrue", "iffalse", "ifcase", "ifdefined", "ifcsname",
                       "iffontchar", "ifincsname", "ifpdfprimitive", "ifpdfabsnum", "ifpdfabsdim", "ifprimitive"}
      WRAPPERS = r"center|minipage|flushleft|flushright"
      VERBATIM = re.compile(r"\\begin\s*\{(verbatim\*?|Verbatim\*?|BVerbatim|LVerbatim|lstlisting|minted|alltt)\}(.*?)"
                            r"\\end\s*\{\1\}", re.S)
      INLINE_DATA = re.compile(r"(\\addplot3?\+?" + OPT + r"table" + OPT + r"|\\pgfplotstableread" + OPT + r")\{([^{}]*\n[^{}]*)\}")
      FILECONTENTS = re.compile(r"\\begin\s*\{(filecontents\*?)\}(.*?)\\end\s*\{\1\}", re.S)
      BRACED = r"(?:[^{}]|\{(?:[^{}]|\{(?:[^{}]|\{[^{}]*\})*\})*\})*"
      NEWLABEL = re.compile(r"\\newlabel\{([^{}]+)\}\{\{(" + BRACED + r")\}\{([^{}]*)\}(?:\{(" + BRACED + r")\}\{([^{}]*)\})?")
      AUX_INPUT = re.compile(r"\\@input\{([^{}]+)\}")
      IMAGE_EXT = (".pdf", ".png", ".jpg", ".jpeg", ".eps", ".mps", ".jbig2", ".jb2", ".svg", "")
      KNOWN_EXT = {".pdf", ".png", ".jpg", ".jpeg", ".eps", ".mps", ".jbig2", ".jb2", ".svg", ".tex", ".tikz", ".pgf"}
      COLUMNS = ["label", "fingerprint", "reviewer", "date", "verdict", "note"]
      DEPTH = 8
      
      
      def strip_comment(line, keep_mark=False):
          """The line without its comment; with keep_mark the % stays (it still joins the line to the next)."""
          i = 0
          while True:
              j = line.find("%", i)
              if j < 0:
                  return line
              k = j
              while k > 0 and line[k - 1] == "\\":
                  k -= 1
              if (j - k) % 2 == 0:
                  return line[:j + 1] if keep_mark else line[:j]
              i = j + 1
      
      
      def strip_dead(text):
          """Drop what TeX never typesets: comment and filecontents environments, and each \\iffalse that starts a line
          through its matching \\fi. Conditionals counted are TeX's and the ones a \\newif declares; a \\iffalse with no
          matching \\fi is left alone (a dead float counted is noise, a live one dropped is a false pass)."""
          text = re.sub(r"\\begin\s*\{comment\}.*?\\end\s*\{comment\}", "", text, flags=re.S)
          ifs = PRIMITIVE_IFS | set(NEWIF.findall(text))
          opener = re.compile(r"(?m)^[ \t]*\\iffalse(?![A-Za-z@])")
          token = re.compile(r"(\\newif\s*)?\\(if[A-Za-z@]*|fi|else)(?![A-Za-z@])")
          out, i = [], 0
          while True:
              m = opener.search(text, i)
              if not m:
                  out.append(text[i:])
                  return "".join(out)
              depth, j, closed = 1, m.end(), False
              while True:
                  t = token.search(text, j)
                  if not t:
                      break
                  j = t.end()
                  if t.group(1):
                      continue  # \newif\ifname declares a conditional; it opens nothing
                  if t.group(2) == "fi":
                      depth -= 1
                      if depth == 0:
                          closed = True
                          break
                  elif t.group(2) == "else" and depth == 1:
                      break  # an \else makes part of the block live: keep all of it rather than guess which part
                  elif t.group(2) in ifs:
                      depth += 1
              if closed:
                  out.append(text[i:m.start()])
              else:
                  out.append(text[i:m.end()])
                  j = m.end()
              i = j
      
      
      def _sealed(kind, raw):
          return f"\\sealed{kind}{{{hashlib.sha256(raw.encode('utf-8')).hexdigest()[:16]}}}"
      
      
      def seal(text):
          """Text TeX does not read as prose is kept as written: verbatim-like environments and inline plot data (a line
          break separates rows there, a % is a character) and filecontents blocks (their content is a file) are each
          replaced by a token carrying the hash of their raw text, so any change to them counts and none of it is read as
          document text."""
          text = VERBATIM.sub(lambda m: _sealed("verbatim", m.group(0)), text)
          text = FILECONTENTS.sub(lambda m: _sealed("filecontents", m.group(0)), text)
          return INLINE_DATA.sub(lambda m: m.group(1) + "{" + _sealed("data", m.group(2)) + "}", text)
      
      
      def plain(text):
          """Text as TeX reads it: whole-line comments and inline comment text dropped, dead blocks dropped, a line break a
          space, a line ending in % joined to the next with nothing, spaces collapsed, one paragraph per line with a blank
          line between. Verbatim, inline plot data and filecontents are sealed first (see seal)."""
          text = seal(text)
          lines = [strip_comment(l, keep_mark=True) for l in text.splitlines() if not l.lstrip().startswith("%")]
          text = strip_dead("\n".join(lines))
          paras, cur = [], []
          for line in text.splitlines():
              s = re.sub(r"[ \t]+", " ", line).strip()
              if s:
                  cur.append(s)
              elif cur:
                  paras.append(cur)
                  cur = []
          if cur:
              paras.append(cur)
          out = []
          for p in paras:
              acc = ""
              for s in p:
                  if acc.endswith("%") and not acc.endswith("\\%"):
                      acc = acc[:-1] + s
                  else:
                      acc = f"{acc} {s}" if acc else s
              out.append(acc)
          return "\n\n".join(out)
      
      
      class Tree:
          """Files as the working tree has them, or as a commit has them (--built-from), relative to the base directory."""
      
          def __init__(self, base, commit=None):
              self.base, self.commit, self._cache = Path(base), commit, {}
      
          def read(self, rel):
              if rel not in self._cache:
                  if self.commit:
                      r = subprocess.run(["git", "-C", str(self.base), "show", f"{self.commit}:./{rel}"],
                                         capture_output=True)
                      self._cache[rel] = r.stdout if r.returncode == 0 else None
                  else:
                      p = self.base / rel
                      self._cache[rel] = p.read_bytes() if p.is_file() else None
              return self._cache[rel]
      
          def text(self, rel):
              b = self.read(rel)
              return None if b is None else b.decode("utf-8", "replace")
      
          def digest(self, rel):
              """A .tex-like file by its text read as TeX reads it, anything else by its bytes."""
              if posixpath.splitext(rel)[1] in (".tex", ".tikz", ".pgf"):
                  return hashlib.sha256(plain(self.text(rel) or "").encode("utf-8")).hexdigest()
              return hashlib.sha256(self.read(rel)).hexdigest()
      
      
      def resolve(tree, name, here, exts):
          for base in here:
              for ext in exts:
                  cand = posixpath.normpath(posixpath.join(base, name + ext) if base else name + ext)
                  if not cand.startswith("../") and tree.read(cand) is not None:
                      return cand
          return None
      
      
      def input_refs(text):
          """[(name, import_dir or None)] for \\input, \\include, \\subfile, \\import in the text."""
          out = []
          for a, b, d, f in INPUTS.findall(text):
              if f:
                  name, sub = f.strip(), d.strip()
              else:
                  name, sub = (a or b).strip(), None
              if name:
                  out.append((name if posixpath.splitext(name)[1] else name + ".tex", sub))
          return out
      
      
      _TEX_TREE = {}
      
      
      def in_tex_tree(name):
          """True when the TeX distribution provides the file (\\input{glyphtounicode}): not the project's to track."""
          if name not in _TEX_TREE:
              k = shutil.which("kpsewhich")
              _TEX_TREE[name] = bool(k) and subprocess.run([k, name], capture_output=True, text=True).returncode == 0
          return _TEX_TREE[name]
      
      
      def graphic_exts(name):
          return ("",) if posixpath.splitext(name)[1].lower() in KNOWN_EXT else IMAGE_EXT
      
      
      class Ctx:
          """What every float of one draft shares: \\graphicspath directories and the files \\pgfplotstableread put in a
          table macro."""
      
          def __init__(self, gpath=(), tables=None):
              self.gpath, self.tables = list(gpath), dict(tables or {})
      
      
      def pulled(tree, text, here, acc, missing, unfollowed, ctx, depth=0):
          """Files the text pulls in, recursively through .tex: acc [(path, digest)], and names that resolve to nothing or
          are macros."""
          seen = {p for p, _ in acc}
      
          def add(got):
              if got not in seen:
                  seen.add(got)
                  acc.append((got, tree.digest(got)))
                  return True
              return False
          for name, sub in input_refs(text):
              if "#" in name:
                  continue  # a macro's parameter: the file is named where the macro is used, which is not followed
              if "\\" in name:
                  unfollowed.append(name)
                  continue
              got = resolve(tree, posixpath.join(sub, name) if sub else name, here, ("",))
              if not got:
                  if not in_tex_tree(name):
                      missing.append(name)
              elif add(got) and depth < DEPTH:
                  # \input is found from the main file's directory first, as LaTeX finds it; \import puts its own first
                  child = ([posixpath.normpath(posixpath.join(h, sub)) if h else posixpath.normpath(sub) for h in here]
                           + here) if sub else here + [posixpath.dirname(got)]
                  pulled(tree, plain(tree.text(got) or ""), child, acc, missing, unfollowed, ctx, depth + 1)
          for name in GRAPHICS.findall(text):
              name = name.strip()
              if "#" in name:
                  continue
              if "\\" in name:
                  unfollowed.append(name)
                  continue
              got = resolve(tree, name, here + [posixpath.join(here[0], g) if here[0] else g for g in ctx.gpath],
                            graphic_exts(name))
              if not got:
                  missing.append(name)
              else:
                  add(got)
          for rx in DATA:
              for name in rx.findall(text):
                  name = name.strip()
                  if name.startswith("\\"):
                      # a table held in a macro: the file \pgfplotstableread filled it from, wherever that was
                      src = ctx.tables.get(name[1:]) if rx is ADDPLOT else None
                      if src:
                          got = resolve(tree, src, here, ("",))
                          if got:
                              add(got)
                          else:
                              missing.append(src)
                      continue
                  got = resolve(tree, name, here, ("", ".tex"))
                  if got:
                      add(got)
                  elif not re.search(r"\s", name):
                      missing.append(name)  # a name with spaces or line breaks in it is inline data, not a file
          return acc
      
      
      def fingerprint(body, files):
          h = hashlib.sha256(body.encode("utf-8"))
          for path, digest in files:
              h.update(("\0" + path + "\0" + digest).encode())
          return h.hexdigest()[:16]
      
      
      def reach(tree, main):
          """(body files {rel: (here, text)}, preamble item or None, custom float environments, ctx, problems) for one
          main file."""
          main = posixpath.normpath(main)
          root = posixpath.dirname(main)
          text = plain(tree.text(main) or "")
          m = DOC.search(text)
          pre_text, body_text = (text[:m.start()], text[m.end():]) if m else ("", text)
          pre_files, pre_missing, pre_unf, problems = [], [], [], []
          ctx = Ctx()
          pulled(tree, pre_text, [root, ""], pre_files, pre_missing, pre_unf, ctx)
          for name in pre_unf:
              problems.append({"kind": "unfollowed", "float": f"preamble:{main}",
                               "detail": f"{name} names a macro: what it pulls in is not seen"})
          for name in pre_missing:
              problems.append({"kind": "unfollowed", "float": f"preamble:{main}", "detail": f"{name} is not there"})
          pre_sources = [pre_text] + [plain(tree.text(p) or "") for p, _ in pre_files if p.endswith(".tex")]
          ctx.gpath = [d for src in pre_sources for group in GPATH.findall(src) for d in re.findall(r"\{([^{}]*)\}", group)]
          envs = {name for src in pre_sources for name, inner in NEWENV.findall(src) if inner in BASE_ENVS}
          preamble = None
          if m:
              preamble = {"id": f"preamble:{main}", "env": "preamble", "file": main, "labels": [], "caption": "", "statements": [],
                          "pulled": [p for p, _ in pre_files], "missing": [],
                          "fingerprint": fingerprint(pre_text, pre_files)}
          body = {main: ([root, ""], body_text)}
      
          def walk(rel, here, depth):
              for name, sub in input_refs(body[rel][1]):
                  if "\\" in name:
                      problems.append({"kind": "unfollowed", "float": rel, "detail": f"\\input{{{name}}} names a macro: "
                                       "floats behind it are not seen"})
                      continue
                  got = resolve(tree, posixpath.join(sub, name) if sub else name, here, ("",))
                  if not got:
                      problems.append({"kind": "unfollowed", "float": rel, "detail": f"{name} is not there"})
                      continue
                  if got in body or depth >= DEPTH:
                      continue
                  child = ([posixpath.normpath(posixpath.join(h, sub)) if h else posixpath.normpath(sub) for h in here]
                           + here) if sub else [root, posixpath.dirname(got), ""]
                  body[got] = (child, plain(tree.text(got) or ""))
                  walk(got, child, depth + 1)
          walk(main, [root, ""], 0)
          for src in pre_sources + [t for _, t in body.values()]:
              for f, macro in TABLEREAD.findall(src):
                  if not re.search(r"\s", f.strip()) and "sealed" not in f:
                      ctx.tables.setdefault(macro, f.strip())
          return body, preamble, envs, ctx, problems
      
      
      def spans(text, envs):
          """[(env, start, end)] of the float environments in the text, outermost only."""
          names = sorted(set(BASE_ENVS) | envs, key=len, reverse=True)
          begin = re.compile(r"\\begin\s*\{(" + "|".join(re.escape(n) for n in names) + r")\}")
          out, pos = [], 0
          while True:
              m = begin.search(text, pos)
              if not m:
                  return out
              end = re.compile(r"\\end\s*\{" + re.escape(m.group(1)) + r"\}").search(text, m.end())
              stop = end.end() if end else len(text)
              out.append((m.group(1), m.start(), stop))
              pos = stop
      
      
      def around(text, pos, found):
          """The span of a \\captionof: the outermost center, minipage or flush environment around it (an image often sits
          in a sibling minipage of the one holding the caption), or else its paragraph; never reaching into a float
          environment beside it."""
          for m in re.finditer(r"\\begin\s*\{(" + WRAPPERS + r")\}", text[:pos]):
              depth, j = 0, m.end()
              rx = re.compile(r"\\(begin|end)\s*\{" + m.group(1) + r"\}")
              while True:
                  t = rx.search(text, j)
                  if not t:
                      break
                  j = t.end()
                  if t.group(1) == "begin":
                      depth += 1
                  elif depth:
                      depth -= 1
                  else:
                      break
              if t and t.group(1) == "end" and t.start() >= pos:
                  return m.start(), t.end()
          a = text.rfind("\n\n", 0, pos)
          a = max([a + 2 if a >= 0 else 0] + [e for _, _, e in found if e <= pos])
          b = text.find("\n\n", pos)
          b = min([b if b >= 0 else len(text)] + [s for _, s, _ in found if s >= pos])
          return a, b
      
      
      # What a float says in words, for the review sheet: the text of the float and the .tex files it pulls in, with TeX
      # commands, options, coordinates and math dropped. A sentence that predicts or asserts is marked for the reviewer.
      # (09-27: a figure's annotation predicted an ordering the paper's own evidence ruled out. The overreach scan read the
      # figure text but only knows wordings someone listed; the first look at the rendered figure checked its fonts.)
      # Prediction, inference, cause and universal claims. Not "only", "all", "every": in a figure they mostly describe what
      # is drawn (on a real manuscript they made most of the marks, and none of those asserted anything).
      ASSERTS = re.compile(r"\b(should|would|will|must|expect\w*|predict\w*|if|therefore|thus|hence|shows?|showed|"
                           r"demonstrat\w*|proves?|confirms?|cannot|can't|always|never|because|causes?)\b", re.I)
      _DROP_ARGS = re.compile(r"\\(?:label|ref|cref|Cref|autoref|eqref|cite\w*|includegraphics|includesvg|input|include|"
                              r"usetikzlibrary|definecolor|tikzset|pgfplotsset|graphicspath|addplot|pgfplotstableread)\*?\s*"
                              r"(?:\[[^\]]*\]\s*)*(?:\{[^{}]*\}\s*)*")
      
      
      def detex(text):
          t = plain(text)
          t = re.sub(r"\}\s*;", "}. ", t)                    # a TikZ node ends where its text ends
          t = t.replace("\\caption", ". \\caption")           # and a caption starts a sentence of its own
          t = re.sub(r"\bat\s*\([^()]*\)", " ", t)           # node placement: at (x,y)
          t = re.sub(r"\$[^$]*\$", " ", t)
          t = _DROP_ARGS.sub(" ", t)
          t = re.sub(r"\\(?:begin|end)\s*\{[^{}]*\}", " ", t)
          t = re.sub(r"\[[^\[\]]*\]", " ", t)
          t = re.sub(r"\([^()]*\)", lambda m: m.group(0) if re.search(r"[A-Za-z]{3,}\s+[A-Za-z]{3,}", m.group(0)) else " ", t)
          t = re.sub(r"\\(?:[A-Za-z@]+\*?|.)", " ", t)              # commands, and \\, \ , \, and the like
          t = re.sub(r"[{}~;&]", " ", t)
          t = re.sub(r"\s+", " ", t)
          t = re.sub(r"(?:\s*\.){2,}", ".", t).strip()
          return re.sub(r"^\.\s*", "", t)
      
      
      def statements(texts):
          """[{text, asserts}] for the sentences a float shows: at least four words of letters, in source order, each once."""
          out, seen = [], set()
          for text in texts:
              for sent in re.split(r"(?<=[.!?])\s+", detex(text)):
                  sent = sent.strip()
                  # key=value runs are plot options that crossed a line, not words the page shows
                  words, numbers = len(re.findall(r"[A-Za-z]{2,}", sent)), len(re.findall(r"\d+(?:[.,]\d+)*", sent))
                  # a table's column spec or its rows of numbers are not sentences
                  if words < 4 or "=" in sent or "@" in sent or numbers > words or sent in seen:
                      continue
                  seen.add(sent)
                  out.append({"text": sent[:400], "asserts": bool(ASSERTS.search(sent))})
          return out
      
      
      def claims_allowed(path):
          """The claims ledger's allowed wordings, one per claim, to set a float's words against. [] when unreadable."""
          try:
              text = Path(path).read_text(encoding="utf-8")
          except OSError:
              return None
          out, head = [], ""
          for line in text.splitlines():
              h = re.match(r"^#{2,3}\s+(.+?)\s*$", line)
              if h:
                  head = h.group(1)
                  continue
              m = re.match(r"^\s*(?:[-*]\s*)?(?:\*\*)?允许的说法(?:\*\*)?\s*[::]\s*(?:\*\*)?\s*(.+?)\s*$", line)
              if m and head:
                  out.append(f"{head}: {m.group(1)}")
          return out
      
      
      def caption_of(body):
          m = re.search(r"\\caption(?:of\s*\{[^{}]*\})?\*?\s*(?:\[[^\]]*\])?\s*\{", body)
          if not m:
              return ""
          depth, j, i = 0, m.end() - 1, m.end() - 1
          while j < len(body):
              c = body[j]
              if c == "\\":
                  j += 2
                  continue
              depth += c == "{"
              depth -= c == "}"
              if depth == 0:
                  break
              j += 1
          return re.sub(r"\s+", " ", body[i + 1:j]).strip()
      
      
      def collect(tree, mains):
          """(floats, preambles, problems). Mains with a document environment go first, so a section file listed before
          its main still belongs to that main."""
          order = sorted(mains, key=lambda m: 0 if DOC.search(plain(tree.text(posixpath.normpath(m)) or "")) else 1)
          owner, problems, preambles = {}, [], []
          for m in order:
              body, pre, envs, ctx, probs = reach(tree, m)
              problems += probs
              if pre and pre["id"] not in [p["id"] for p in preambles]:
                  preambles.append(pre)
              for rel, (here, text) in body.items():
                  owner.setdefault(rel, (here, text, envs, ctx))
          floats = []
          for rel, (here, text, envs, ctx) in owner.items():
              found = spans(text, envs)
              pieces = [(env, text[a:b], 0) for env, a, b in found]
              taken = [(a, b) for _, a, b in found]
              for m in CAPTIONOF.finditer(text):
                  if not any(a <= m.start() < b for _, a, b in found):
                      a, b = around(text, m.start(), found)
                      pieces.append(("captionof " + m.group(1), text[a:b], m.start() - a))
                      taken.append((a, b))
              # A table set in the running text, with no float and no caption, still prints: listed as its own item, by
              # its order in the file (a new one inserted before it renumbers it, and its review is asked for again).
              k = 0
              for m in TABULAR.finditer(text):
                  if any(a <= m.start() < b for a, b in taken):
                      continue
                  k += 1
                  a, b = around(text, m.start(), found)
                  if not re.match(r"\\begin\s*\{(" + WRAPPERS + r")\}", text[a:]):
                      # no center or minipage around it: the table itself, not the paragraph it follows
                      end = re.compile(r"\\end\s*\{" + re.escape(m.group(1)) + r"\}").search(text, m.end())
                      a, b = m.start(), end.end() if end else len(text)
                  taken.append((a, b))
                  pieces.append(("inline tabular", text[a:b], 0, f"{rel}#tabular{k}"))
              for n, piece in enumerate(pieces, 1):
                  env, body, at = piece[:3]
                  acc, missing, unfollowed = [], [], []
                  pulled(tree, body, here, acc, missing, unfollowed, ctx)
                  # a \captionof names the label that follows it
                  labels = LABEL.findall(body[at:]) or LABEL.findall(body)
                  if not labels:
                      for p, _ in acc:
                          if p.endswith(".tex"):
                              labels = LABEL.findall(plain(tree.text(p) or ""))
                              if labels:
                                  break
                  fid = piece[3] if len(piece) > 3 else (labels[0] if labels else f"{rel}#{n}")
                  floats.append({"id": fid, "env": env, "file": rel, "labels": labels,
                                 "caption": caption_of(body), "pulled": [p for p, _ in acc],
                                 "statements": statements([body] + [tree.text(p) or "" for p, _ in acc if p.endswith(".tex")]),
                                 "missing": missing + [f"{u} (a macro)" for u in unfollowed],
                                 "fingerprint": fingerprint(body, acc)})
          # a tabular in a file that a float pulls in is that float's content, not a table of its own
          inside_floats = {p for f in floats if f["env"] != "inline tabular" for p in f["pulled"]}
          floats = [f for f in floats if not (f["env"] == "inline tabular" and f["file"] in inside_floats)]
          return floats, preambles, problems
      
      
      def read_reviews(path):
          """[row] in file order, or raises ValueError. A missing file is no review at all."""
          if not path.is_file():
              return []
          lines = [l for l in path.read_text(encoding="utf-8").splitlines() if l.strip()]
          if not lines:
              return []
          header = lines[0].split("\t")
          if header[:len(COLUMNS)] != COLUMNS:
              raise ValueError(f"expected columns {COLUMNS}, found {header}")
          rows = []
          for n, line in enumerate(lines[1:], 2):
              parts = line.split("\t")
              if len(parts) < 5 or not parts[0].strip() or not parts[1].strip() or not parts[4].strip():
                  raise ValueError(f"line {n}: label, fingerprint and verdict are required")
              row = dict(zip(COLUMNS, parts + [""] * (len(COLUMNS) - len(parts))))
              row["line"] = n
              rows.append(row)
          return rows
      
      
      def judge(item, rows, findings):
          current = [r for r in rows if r["label"].strip() == item["id"] and r["fingerprint"].strip() == item["fingerprint"]]
          item["review"] = ({k: current[-1][k] for k in ("reviewer", "date", "verdict", "note")} if current else None)
          older = [r for r in rows if r["label"].strip() == item["id"] and r["fingerprint"].strip() != item["fingerprint"]]
          if item["missing"]:
              findings.append({"kind": "missing-file", "float": item["id"], "detail": "pulls in " + ", ".join(item["missing"])})
          if not current:
              item["status"] = "unreviewed"
              what = ("the preamble changed since; look again at the pages its macros, lengths or fonts reach"
                      if item["env"] == "preamble" else "changed since")
              findings.append({"kind": "unreviewed", "float": item["id"],
                               "detail": (f"reviewed at an older version ({older[-1]['fingerprint'].strip()}); {what}"
                                          if older else "no review")})
          elif current[-1]["verdict"].strip().lower().startswith("fix"):
              item["status"] = "open"
              findings.append({"kind": "review-open", "float": item["id"],
                               "detail": f"{current[-1]['reviewer']}: {current[-1]['note'][:160]}"})
          else:
              item["status"] = "reviewed"
      
      
      def check(a):
          base = Path(a.base_dir)
          tree = Tree(base)
          floats, preambles, problems = collect(tree, a.main)
          if not floats:
              print("NOTHING CHECKED: no figure or table environment is reachable from --main. This is not a pass.",
                    file=sys.stderr)
              return 2, None
          rpath = Path(a.reviews) if Path(a.reviews).is_absolute() else base / a.reviews
          try:
              rows = read_reviews(rpath)
          except (OSError, ValueError) as e:
              print(f"REVIEWS: cannot read {rpath}: {e}", file=sys.stderr)
              return 2, None
          findings = list(problems)
          for item in floats + preambles:
              judge(item, rows, findings)
          ids = {f["id"] for f in floats + preambles}
          for r in rows:
              if r["label"].strip() not in ids:
                  findings.append({"kind": "stale-row", "prompt": True, "float": r["label"].strip(),
                                   "detail": f"line {r['line']} names nothing in the draft"})
          hard = [x for x in findings if not x.get("prompt")]
          n_rev = sum(1 for f in floats if f["status"] == "reviewed")
          n_open = sum(1 for f in floats if f["status"] == "open")
          parts = [f"图表 {len(floats)} 个:看过这一版 {n_rev}"]
          if n_open:
              parts.append(f"看过但待改 {n_open}")
          unrev = [f["id"] for f in floats if f["status"] == "unreviewed"]
          if unrev:
              parts.append(f"没人看过这一版 {len(unrev)}({'、'.join(unrev[:3])}{' 等' if len(unrev) > 3 else ''})")
          pre_unrev = sum(1 for p in preambles if p["status"] == "unreviewed")
          pre_open = sum(1 for p in preambles if p["status"] == "open")
          if pre_unrev:
              parts.append(f"导言区 {pre_unrev} 份没人看过这一版")
          if pre_open:
              parts.append(f"导言区 {pre_open} 份待改")
          miss = sum(1 for x in findings if x["kind"] == "missing-file")
          if miss:
              parts.append(f"引用的文件找不到 {miss}")
          unf = sum(1 for x in findings if x["kind"] == "unfollowed")
          if unf:
              parts.append(f"跟不进去的引入 {unf} 处")
          keys = ("id", "env", "file", "labels", "caption", "statements", "fingerprint", "pulled", "missing", "status", "review")
          reviewers = sorted({f["review"]["reviewer"] for f in floats + preambles if f.get("review")})
          payload = {"schema_version": 3, "base": str(base), "reviews": str(rpath),
                     "summary_zh": ";".join(parts),
                     "floats": [{k: f[k] for k in keys} for f in floats],
                     "preambles": [{k: p[k] for k in keys} for p in preambles],
                     "reviewers": reviewers, "findings": findings, "hard_finding_count": len(hard),
                     "limits": "whether a figure is right is the reviewer's call; this records who looked at which "
                               "version. Not followed: what a macro expands to, packages and classes."}
          return (1 if hard else 0), payload
      
      
      def aux_labels(pdf, aux, out, seen=None):
          """{label: [(pdf, number, printed page, anchor)]} from an .aux and the .aux files it \\@input's."""
          seen = set() if seen is None else seen
          p = Path(aux).resolve()
          if p in seen or not p.is_file():
              return out
          seen.add(p)
          text = p.read_text(encoding="utf-8", errors="replace")
          for label, number, page, _title, anchor in NEWLABEL.findall(text):
              out.setdefault(label, []).append((pdf, re.sub(r"[{}]|\\relax\s*", "", number).strip(), page.strip(),
                                                (anchor or "").strip()))
          for sub in AUX_INPUT.findall(text):
              aux_labels(pdf, p.parent / sub, out, seen)
          return out
      
      
      def dests(pdf):
          """{named destination: physical page} from pdfinfo, or {} when it is not available."""
          if not shutil.which("pdfinfo"):
              return {}
          r = subprocess.run(["pdfinfo", "-dests", pdf], capture_output=True, text=True, errors="replace")
          out = {}
          for line in (r.stdout or "").splitlines():
              m = re.match(r"\s*(\d+)\s+\[.*\]\s+\"(.*)\"\s*$", line)
              if m:
                  out.setdefault(m.group(2), int(m.group(1)))
          return out
      
      
      def render(a, payload):
          if len(a.pdf) != len(a.aux) or not a.pdf:
              print("RENDER: give --pdf and --aux in pairs", file=sys.stderr)
              return 2
          gone = [p for p in a.pdf + a.aux if not Path(p).is_file()]
          if gone:
              print("RENDER: not found: " + ", ".join(gone), file=sys.stderr)
              return 2
          if not shutil.which("pdftoppm"):
              print("RENDER: pdftoppm is not installed", file=sys.stderr)
              return 2
          out = Path(a.out)
          out.mkdir(parents=True, exist_ok=True)
          labels = {}
          for pdf, aux in zip(a.pdf, a.aux):
              aux_labels(pdf, aux, labels)
          named = {pdf: dests(pdf) for pdf in a.pdf}
          stems = [Path(p).stem for p in a.pdf]
          name_of = {p: (Path(p).stem if stems.count(Path(p).stem) == 1 else f"{k + 1}-{Path(p).stem}")
                     for k, p in enumerate(a.pdf)}
          then = None
          if a.built_from:
              fl, pre, _ = collect(Tree(a.base_dir, a.built_from), a.main)
              then = {f["id"]: f["fingerprint"] for f in fl + pre}
          rendered, rows = {}, []
          for f in payload["floats"]:
              hits = [h for l in f["labels"] for h in labels.get(l, [])]
              if not hits:
                  rows.append((f, None, None, None, ["not in any .aux given: not rendered"]))
                  continue
              pdf, number, printed, anchor = hits[0]
              notes = []
              if len({h[0] for h in hits}) > 1:
                  notes.append(f"the label is in {len({h[0] for h in hits})} of the .aux files given; rendered from {pdf}")
              page = named[pdf].get(anchor) if anchor else None
              if page is None:
                  page = int(printed) if printed.isdigit() else None
                  notes.append("no named destination for the label: the printed page number was used as the PDF page")
              if page is None:
                  rows.append((f, number, printed, None, notes + ["no usable page"]))
                  continue
              key = (pdf, page)
              if key not in rendered:
                  stem = f"{name_of[pdf]}-p{page:03d}"
                  r = subprocess.run(["pdftoppm", "-r", str(a.dpi), "-f", str(page), "-l", str(page), "-png", "-singlefile",
                                      pdf, str(out / stem)], capture_output=True, text=True)
                  rendered[key] = (out / (stem + ".png")) if r.returncode == 0 else None
              png = rendered[key]
              if png is None:
                  notes.append("pdftoppm failed")
              elif then is not None:
                  if then.get(f["id"]) != f["fingerprint"]:
                      notes.append(f"PDF built from {a.built_from}, where this float was a different version: "
                                   "rebuild before reviewing")
              else:
                  pre = [x for p in payload["preambles"] for x in [p["file"]] + p["pulled"]]
                  files = [Path(a.base_dir) / q for q in [f["file"]] + f["pulled"] + pre]
                  newest = max((p.stat().st_mtime for p in files if p.is_file()), default=0)
                  if newest > Path(pdf).stat().st_mtime:
                      notes.append("a file of this float is newer than the PDF: rebuild before reviewing")
              rows.append((f, number, page, png, notes))
          pre_rows = []
          for p in payload["preambles"]:
              notes = []
              if then is not None and then.get(p["id"]) != p["fingerprint"]:
                  notes.append(f"PDF built from {a.built_from}, where the preamble was a different version: rebuild first")
              pre_rows.append((p, notes))
          lines = ["# Figure and table review sheet", "",
                   f"Built from: {a.built_from or 'unknown (mtime compared)'}. PDFs: "
                   + ", ".join(f"{p} (sha256 {hashlib.sha256(Path(p).read_bytes()).hexdigest()[:12]})" for p in a.pdf), "",
                   "Look at each page; then add a row to the review record (verdict ok, or fix with a note).", "",
                   "For each figure and table, check:",
                   "1. Every sentence it shows, against the paper's conclusions: does the evidence support it? A sentence marked "
                   "⚑ predicts or asserts; a prediction the results did not bear out is a fix.",
                   "2. Every number it shows, against the text and the data it was drawn from.",
                   "3. Nothing overflows, overlaps or is too small to read; every glyph is in the intended font.",
                   "4. The caption describes what the figure shows now.", ""]
          if a.claims:
              allowed = claims_allowed(a.claims)
              if allowed is None:
                  lines += [f"**The claims ledger could not be read: {a.claims}. Check sentence 1 against the text instead.**", ""]
              else:
                  lines += ["What the paper may say (the claims ledger's allowed wordings):", ""] \
                      + [f"- {x}" for x in allowed] + (["- (the ledger names no allowed wording)"] if not allowed else []) + [""]
          for f, number, page, png, notes in rows:
              kind = "Figure" if "figure" in f["env"].lower() else "Table"
              cap = re.sub(r"(?<!\\)%", "", f["caption"])
              lines += [f"## {kind} {number or '?'} · `{f['id']}` · PDF page {page or '-'}", "",
                        f"- PNG: {png.name if png else '-'}", f"- fingerprint: `{f['fingerprint']}` · now: {f['status']}",
                        f"- caption: {cap[:400] or '(none)'}"]
              lines += [f"- **{n}**" for n in notes]
              says = f.get("statements") or []
              if says:
                  lines += ["- what it says in words (⚑ predicts or asserts: check it against the conclusions):"] \
                      + [f"  - {'⚑ ' if x['asserts'] else ''}{x['text']}" for x in says]
              lines += ["", "```", f"{f['id']}\t{f['fingerprint']}\t<reviewer>\t<date>\t<ok|fix>\t<note>", "```", ""]
          for p, notes in pre_rows:
              lines += [f"## Preamble · `{p['id']}`", "",
                        "- A preamble change can alter every float it reaches (a macro, a length, a font); a review of it "
                        "means the float pages above were looked at under this preamble.",
                        f"- fingerprint: `{p['fingerprint']}` · now: {p['status']}",
                        f"- pulls in: {', '.join(p['pulled']) or '(nothing)'}"]
              lines += [f"- **{n}**" for n in notes]
              lines += ["", "```", f"{p['id']}\t{p['fingerprint']}\t<reviewer>\t<date>\t<ok|fix>\t<note>", "```", ""]
          (out / "REVIEW.md").write_text("\n".join(lines), encoding="utf-8")
          n = sum(1 for *_, png, _ in rows if png)
          print(f"rendered {len(set(p for *_, p, _ in rows if p))} page(s) for {n} of {len(rows)} float(s) -> "
                f"{out / 'REVIEW.md'}")
          for f, number, page, png, notes in rows:
              for note in notes:
                  print(f"  {f['id']}: {note}")
          for p, notes in pre_rows:
              for note in notes:
                  print(f"  {p['id']}: {note}")
          return 0 if n else 2
      
      
      def main(argv=None):
          ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
          ap.add_argument("--base-dir", default=".")
          ap.add_argument("--main", action="append", required=True)
          ap.add_argument("--reviews", required=True)
          ap.add_argument("--json", action="store_true")
          ap.add_argument("--render", action="store_true")
          ap.add_argument("--claims", help="with --render: the claims ledger, whose allowed wordings the sheet sets out")
          ap.add_argument("--pdf", action="append", default=[])
          ap.add_argument("--aux", action="append", default=[])
          ap.add_argument("--built-from")
          ap.add_argument("--dpi", type=int, default=110)
          ap.add_argument("--out")
          a = ap.parse_args(argv)
          code, payload = check(a)
          if payload is None:
              return code
          if a.render:
              if not a.out:
                  print("RENDER: --out is required", file=sys.stderr)
                  return 2
              return render(a, payload)
          if a.json:
              print(json.dumps(payload, indent=2, ensure_ascii=False))
          else:
              print(payload["summary_zh"])
              for x in payload["findings"]:
                  print(f"  {x['kind']}{' (prompt)' if x.get('prompt') else ''}  {x['float']}: {x['detail']}")
              print(f"\n{payload['limits']}")
          return code
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • audit-generated-copies.py 24.3 KB
      #!/usr/bin/env python3
      r"""Check that each table or figure a manuscript copies from a generator is what the generator emits now.
      
          python3 audit-generated-copies.py --base-dir <manuscript> --manifest <generated.json> [--json] [--timeout 240]
      
      The gap this closes. A data repository holds the analysis artifacts and the
      scripts that turn them into LaTeX tables and TikZ figures; the manuscript holds
      copies of what those scripts wrote. The number ledger binds the prose to the
      copies, so a copy is treated as the truth. Nothing asked whether the copy is
      still what its generator emits: a cell edited by hand, a copy left behind when
      an artifact was rerun, or a fix made in the copy and never in the generator
      (the next regeneration undoes it) all passed.
      
      A presence test ("is each cell's number somewhere in the artifact") was the
      first idea and is not used: a two-decimal value in [0, 1] turns up by chance in
      any artifact holding a few hundred floats, so it passes almost everything.
      Rerunning the generator is exact.
      
      The manifest (JSON, in the manuscript repository):
      
          {"covers": ["tables/*.tex", "figures/*.tex"],
           "hand": {"tables/stats.tex": "written by hand; cells checked by the number ledger only"},
           "generators": [
             {"name": "tables", "repo": "~/data-repo",
              "export": ["scripts", "outputs/*.json"],
              "run": ["{repo}/.venv/bin/python", "scripts/make_tables.py", "{out}"],
              "copies": {"tables/main.tex": "{out}/main.tex"}}]}
      
      For each generator the repository's HEAD commit (never its working tree) is
      archived, `export` pathspecs only, into a temporary directory; `run` is run
      there with `{tree}` (the archive) and `{out}` (an empty directory)
      substituted, and `{repo}` for the interpreter: a virtual environment usually
      lives in the repository, so the first word may point there when it is a
      python in a bin/ directory; any other word that points into the working tree
      (through {repo}, an absolute path or ~) is refused. A relative `repo` is read from
      the manuscript directory, which the loop replaces with a copy: give an absolute
      or ~ path there. Links in the archive are not extracted. A produced
      path must resolve inside the archive or `{out}`; any file already there is
      removed before the run, so a generator that writes nothing is caught instead
      of the committed copy being read back.
      
      The working tree can still reach a run through the interpreter: PYTHONPATH,
      user site-packages and an editable install that points into the repository.
      The first two are removed from the environment; the third is looked for in
      the .pth and editable-finder files of each interpreter the run reaches (the
      python the command runs, through /usr/bin/env, nice or timeout; when the
      command runs a shell or a script instead, every python3 and python on its
      PATH), absolute and relative paths both, and a hit fails the generator. A
      python the generator itself starts from PATH is not searched when the command
      names its own interpreter. Uncommitted changes under the export paths are then reported as
      not used, which is only true after these steps.
      
      Copies are read before any generator runs and read again after; a copy that
      changed during the run (a generator that syncs into the manuscript) fails.
      
      Comparison is line by line, bytes kept (an invalid UTF-8 byte is not folded
      into another), trailing spaces and blank lines ignored, and only whole-line
      comments dropped: a provenance comment added to a copy is not a difference,
      while a trailing `%` (which joins lines in LaTeX, or sits in a URL) is
      compared as written. Output with no content line is not a match for anything.
      
        differs          the copy is not what the generator emits; the first
                         differing lines are shown
        generator-failed the archive, the run or its timeout failed
        output-missing   the generator ran and did not write a listed file
        copy-missing     a listed copy is not in the manuscript
        unlisted         a file under `covers` that no generator and no `hand`
                         entry names: nobody said where it comes from (matched
                         without regard to case: the disk may not care either)
        hand-missing     a `hand` entry names a file that is not there
        covers-empty     a `covers` pattern that matches no file (a typo turns the
                         unlisted check off)
        listed-twice     one copy named by two generators
        empty-output     the generator wrote a file with no content line
        copy-changed     a copy changed while the generators ran
      
      A `hand` file is listed, not checked. The manifest's commands run on this
      machine with the user's rights, as a project check does; they come from the
      manuscript repository, which the author controls.
      
      What this does NOT do: decide whether a generator computes the right thing, or
      whether the artifact it reads is the right one; nor sandbox the generator,
      which runs with the user's rights and could read or write anything it names by
      absolute path. It checks that the copy is what the committed generator and
      artifacts produce.
      
      Exit: 1 on any finding above, 2 when the manifest is missing or unreadable or
      nothing was compared, 0 otherwise. Timeouts kill the generator's whole process
      group.
      """
      import argparse
      import difflib
      import io
      import json
      import os
      import posixpath
      import re
      import shutil
      import signal
      import subprocess
      import sys
      import tarfile
      import tempfile
      import time
      from pathlib import Path
      
      HARD = ("differs", "generator-failed", "output-missing", "copy-missing", "unlisted", "hand-missing", "covers-empty",
              "listed-twice", "empty-output", "copy-changed")
      SCRUB = ("PYTHONPATH", "PYTHONHOME", "PYTHONSTARTUP", "PYTHONUSERBASE")
      
      
      def content_lines(data):
          """Lines to compare: bytes kept through surrogateescape, trailing space stripped, blank lines and whole-line
          comments dropped."""
          out = []
          for line in data.decode("utf-8", "surrogateescape").splitlines():
              line = line.rstrip()
              if line.strip() and not line.lstrip().startswith("%"):
                  out.append(line)
          return out
      
      
      def first_differences(copy_lines, gen_lines, limit=6):
          """The first changed lines, as the copy has them (-) and as the generator emits them (+)."""
          out = []
          for line in difflib.unified_diff(copy_lines, gen_lines, lineterm="", n=0):
              if line.startswith(("---", "+++", "@@")):
                  continue
              out.append(line.encode("utf-8", "backslashreplace").decode("utf-8")[:200])
              if len(out) >= limit:
                  break
          return out
      
      
      def git(repo, *args, binary=False):
          """(output, None) or (None, [last error line]). Text output keeps its leading spaces: porcelain status puts the
          first line's status in column one, and stripping it cut the first path short."""
          r = subprocess.run(["git", "-C", repo] + list(args), capture_output=True)
          if r.returncode:
              return None, r.stderr.decode("utf-8", "replace").strip().splitlines()[-1:] or ["git failed"]
          return (r.stdout if binary else r.stdout.decode("utf-8", "replace").rstrip("\n")), None
      
      
      def export(repo, specs, dest):
          """Archive HEAD's export paths into dest: regular files and directories only (a link could point anywhere).
          Returns (commit, error)."""
          head, err = git(repo, "rev-parse", "HEAD")
          head = head.strip() if head else head
          if err:
              return None, f"not a repository with a commit: {err[0]}"
          tar, err = git(repo, "archive", "--format=tar", "HEAD", "--", *specs, binary=True)
          if err:
              return head, f"git archive: {err[0]}"
          with tarfile.open(fileobj=io.BytesIO(tar)) as t:
              members = [m for m in t.getmembers() if (m.isfile() or m.isdir())
                         and not (m.name.startswith("/") or ".." in Path(m.name).parts)]
              t.extractall(dest, members=members)
          return head, None
      
      
      def dirty(repo, specs):
          out, err = git(repo, "status", "--porcelain", "--", *specs)
          return [] if err or not out else [l[3:] for l in out.splitlines()]
      
      
      def inside(path, dirs):
          real = os.path.realpath(path)
          return any(real == d or real.startswith(d + os.sep) for d in dirs)
      
      
      def editable_hits(python, repo, env):
          """Files through which this interpreter imports code from the repository's working tree: .pth lines and
          editable-install finders naming a path inside it."""
          probe = "import site,sys;print('\\n'.join(site.getsitepackages()+[site.getusersitepackages()]))"
          try:
              r = subprocess.run([python, "-c", probe], capture_output=True, text=True, timeout=30, env=env)
          except (OSError, subprocess.TimeoutExpired):
              return []
          hits, real = [], os.path.realpath(repo)
          for d in (r.stdout or "").splitlines():
              p = Path(d)
              if not p.is_dir():
                  continue
              for f in sorted(list(p.glob("*.pth")) + list(p.glob("__editable__*finder*.py"))):
                  try:
                      text = f.read_text(encoding="utf-8", errors="replace")
                  except OSError:
                      continue
                  # Paths are compared resolved: a temporary directory or a home path is often written through a link. A
                  # .pth line that is not an import is a path, relative to the .pth's own directory as site.py reads it.
                  paths = re.findall(r"/[^\s'\"\],;)]+", text)
                  if f.suffix == ".pth":
                      # site.py adds a relative line only when the directory exists
                      paths += [str(p / l.strip()) for l in text.splitlines()
                                if l.strip() and not l.startswith(("#", "import ", "import\t")) and (p / l.strip()).exists()]
                  if any(inside(m, [real]) for m in paths):
                      hits.append(str(f))
          return hits
      
      
      PYTHON_NAME = re.compile(r"python(\d+(\.\d+)*)?(\.exe)?$")
      
      
      def is_interpreter(word):
          """A python (python, python3, python3.11) in a bin/ directory: where a virtual environment keeps it. A script
          whose name merely starts with python is not one."""
          return bool(PYTHON_NAME.match(os.path.basename(word))) and os.path.basename(os.path.dirname(word)) == "bin"
      
      
      def program(argv):
          """The program a command line runs once /usr/bin/env, nice and timeout are looked through: env's own options,
          including -u NAME and -C DIR, and NAME=value assignments are skipped."""
          i = 0
          while i < len(argv):
              base = os.path.basename(argv[i])
              if base == "env":
                  i += 1
                  while i < len(argv) and (argv[i].startswith("-") or "=" in argv[i]):
                      if argv[i] in ("-S", "--split-string") and i + 1 < len(argv):
                          return program(argv[i + 1].split() + argv[i + 2:])  # env -S "python3 x.py": the string is the command
                      i += 2 if argv[i] in ("-u", "-C", "--unset", "--chdir") else 1
                  continue
              if base in ("nice", "timeout", "nohup", "time"):
                  i += 1
                  while i < len(argv) and argv[i].startswith("-"):
                      i += 2 if argv[i] in ("-n", "-s", "-k", "--signal", "--kill-after", "--adjustment") else 1
                  if base == "timeout" and i < len(argv):
                      i += 1  # the duration
                  continue
              return argv[i]
          return ""
      
      
      def interpreters(argv, env):
          """The Python interpreters a run reaches: the program its command line runs when that is a python; otherwise (a
          shell, a script) every python3 and python on the run's PATH, since what it starts may call either."""
          out = []
          first = program(argv)
          words = [first] if PYTHON_NAME.match(os.path.basename(first)) else ["python3", "python"]
          for w in words:
              found = w if os.sep in w else shutil.which(w, path=env.get("PATH"))
              if found and found not in out:
                  out.append(found)
          return out
      
      
      def names_working_tree(argv, repo, roots=()):
          """The words of a command line that point into the repository's working tree. An interpreter may live there (a
          virtual environment usually does); a script, a module path or an argument may not. Paths inside the run's own
          archive and output directories are the run's, wherever TMPDIR puts them."""
          real = os.path.realpath(repo)
          bad = []
          for a in argv:
              if is_interpreter(a):
                  continue
              for m in re.findall(r"(?:~|/)[^\s'\";:,]*", a):
                  path = os.path.expanduser(m)
                  if roots and inside(path, roots):
                      continue
                  if inside(path, [real]) or m.startswith(repo):
                      bad.append(a)
                      break
          return bad
      
      
      def run_argv(argv, cwd, timeout, env):
          """(returncode, stdout, stderr) or raises TimeoutError; on timeout the whole process group is killed."""
          p = subprocess.Popen(argv, cwd=cwd, stdout=subprocess.PIPE, stderr=subprocess.PIPE, env=env,
                               start_new_session=True)
          try:
              out, err = p.communicate(timeout=timeout)
          except subprocess.TimeoutExpired:
              try:
                  os.killpg(p.pid, signal.SIGKILL)
              except OSError:
                  pass
              p.communicate()
              raise TimeoutError
          return p.returncode, out.decode("utf-8", "replace"), err.decode("utf-8", "replace")
      
      
      def run_generator(g, base, timeout):
          """Run one generator. Returns (info, {copy: produced bytes or None}, error)."""
          raw = os.path.expanduser(g.get("repo") or "") if isinstance(g.get("repo"), str) else ""
          repo = str((base / raw).resolve()) if raw else ""
          info = {"name": g.get("name") or "?", "repo": repo, "commit": None, "dirty": [], "seconds": None}
          specs = [str(s) for s in g.get("export") or []]
          run = [str(a) for a in g.get("run") or []]
          if not repo or not specs or not run or not isinstance(g.get("copies"), dict):
              return info, {}, "generator needs repo, export, run and copies"
          env = {k: v for k, v in os.environ.items() if k not in SCRUB}
          env.update(PYTHONDONTWRITEBYTECODE="1", PYTHONNOUSERSITE="1")
          with tempfile.TemporaryDirectory(prefix="gen-tree-") as tree, tempfile.TemporaryDirectory(prefix="gen-out-") as out:
              commit, err = export(repo, specs, tree)
              info["commit"] = commit
              if err:
                  return info, {}, err
              info["dirty"] = dirty(repo, specs)
              subst = lambda s: os.path.expanduser(s).replace("{repo}", repo).replace("{tree}", tree).replace("{out}", out)
              argv = [subst(a) for a in run]
              bad = names_working_tree(argv, repo, [os.path.realpath(tree), os.path.realpath(out)])
              if bad:
                  return info, {}, ("the command names the working tree of the repository (" + ", ".join(bad[:2]) + "); "
                                    "only an interpreter in its bin/ may live there")
              hits = [h for py in interpreters(argv, env) for h in editable_hits(py, repo, env)]
              if hits:
                  return info, {}, "the interpreter imports the working tree of the repository through " + ", ".join(hits[:3])
              roots = [os.path.realpath(tree), os.path.realpath(out)]
              produced = {}
              for copy, target in g["copies"].items():
                  p = Path(subst(str(target)))
                  p = Path(os.path.normpath(str(p if p.is_absolute() else Path(tree) / p)))
                  if not inside(p, roots):
                      # A file outside the archive and the output directory would be read as it already is, and removing
                      # it would remove someone's file.
                      return info, {}, f"{copy}: the produced path {target} resolves outside {{tree}} and {{out}}"
                  if p.is_dir():
                      return info, {}, f"{copy}: the produced path {target} is a directory"
                  if p.exists() or p.is_symlink():
                      p.unlink()
                  produced[copy] = p
              t0 = time.time()
              try:
                  code, so, se = run_argv(argv, tree, timeout, env)
              except TimeoutError:
                  return info, {}, f"timed out after {timeout} s"
              except OSError as e:
                  return info, {}, f"could not start: {e}"
              info["seconds"] = round(time.time() - t0, 2)
              if code:
                  tail = (se or so or "").strip().splitlines()[-1:] or ["no output"]
                  return info, {}, f"exit {code}: {tail[0][:200]}"
              data = {c: (p.read_bytes() if p.is_file() and inside(p, roots) else None) for c, p in produced.items()}
          return info, data, None
      
      
      def covered_files(base, patterns):
          """{pattern: [paths]} under base: * and ? within a segment, ** across segments, case ignored."""
          def rx(pat):
              out, i = "", 0
              while i < len(pat):
                  if pat.startswith("**/", i):
                      out, i = out + "(?:.*/)?", i + 3
                  elif pat.startswith("**", i):
                      out, i = out + ".*", i + 2
                  elif pat[i] == "*":
                      out, i = out + "[^/]*", i + 1
                  elif pat[i] == "?":
                      out, i = out + "[^/]", i + 1
                  else:
                      out, i = out + re.escape(pat[i]), i + 1
              return re.compile(out, re.I)
          files = []
          for d, dirs, names in os.walk(base):
              dirs[:] = [x for x in dirs if not x.startswith(".")]
              rel = os.path.relpath(d, base)
              files += [posixpath.normpath(posixpath.join("" if rel == "." else rel.replace(os.sep, "/"), n)) for n in names]
          return {pat: sorted(f for f in files if rx(posixpath.normpath(str(pat))).fullmatch(f)) for pat in patterns}
      
      
      def main(argv=None):
          ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
          ap.add_argument("--base-dir", default=".")
          ap.add_argument("--manifest", required=True)
          ap.add_argument("--timeout", type=int, default=240, help="seconds per generator")
          ap.add_argument("--json", action="store_true")
          a = ap.parse_args(argv)
          base = Path(a.base_dir).resolve()
          mpath = Path(a.manifest) if Path(a.manifest).is_absolute() else base / a.manifest
          try:
              manifest = json.loads(mpath.read_text(encoding="utf-8"))
              if not isinstance(manifest, dict):
                  raise ValueError("not an object")
          except (OSError, ValueError) as e:
              print(f"MANIFEST: cannot read {mpath}: {e}", file=sys.stderr)
              return 2
          gens = manifest.get("generators") or []
          hand = manifest.get("hand") or {}
          covers = manifest.get("covers") or []
          if not isinstance(gens, list) or not isinstance(hand, dict) or not isinstance(covers, list):
              print("MANIFEST: generators must be a list, hand an object, covers a list", file=sys.stderr)
              return 2
          norm = lambda c: posixpath.normpath(str(c))
          hand = {norm(c): r for c, r in hand.items()}
          findings, same, generators, owner = [], [], [], {}
          for g in gens:
              for c in (g.get("copies") or {}) if isinstance(g, dict) and isinstance(g.get("copies"), dict) else {}:
                  if norm(c) in owner:
                      findings.append({"kind": "listed-twice", "copy": norm(c), "generator": g.get("name") or "?",
                                       "detail": f"also named by {owner[norm(c)]}"})
                  owner.setdefault(norm(c), g.get("name") or "?")
          if not owner and not covers:
              print("NOTHING CHECKED: the manifest lists no copy and covers no path. This is not a pass.", file=sys.stderr)
              return 2
          before = {c: ((base / c).read_bytes() if (base / c).is_file() else None) for c in owner}
          compared = set()
          for g in gens:
              if not isinstance(g, dict):
                  findings.append({"kind": "generator-failed", "copy": "-", "generator": "?", "detail": "not an object"})
                  continue
              info, data, err = run_generator(g, base, a.timeout)
              generators.append(info)
              copies = g.get("copies") if isinstance(g.get("copies"), dict) else {}
              for key in copies:
                  copy = norm(key)
                  if owner.get(copy) != (g.get("name") or "?") or copy in compared:
                      continue
                  compared.add(copy)
                  if err:
                      findings.append({"kind": "generator-failed", "copy": copy, "generator": info["name"], "detail": err})
                  elif before.get(copy) is None:
                      findings.append({"kind": "copy-missing", "copy": copy, "generator": info["name"],
                                       "detail": "listed in the manifest, not in the manuscript"})
                  elif data.get(key) is None:
                      findings.append({"kind": "output-missing", "copy": copy, "generator": info["name"],
                                       "detail": f"{info['name']} ran and did not write {copies[key]}"})
                  else:
                      cl, gl = content_lines(before[copy]), content_lines(data[key])
                      if not gl:
                          findings.append({"kind": "empty-output", "copy": copy, "generator": info["name"],
                                           "detail": f"{info['name']} wrote {copies[key]} with no content line"})
                      elif cl == gl:
                          same.append(copy)
                      else:
                          findings.append({"kind": "differs", "copy": copy, "generator": info["name"],
                                           "detail": f"{info['name']} at {(info['commit'] or '')[:7]} emits something else",
                                           "lines": first_differences(cl, gl)})
          for c, b in before.items():
              now = (base / c).read_bytes() if (base / c).is_file() else None
              if now != b:
                  findings.append({"kind": "copy-changed", "copy": c, "generator": owner[c],
                                   "detail": "the manuscript copy changed while the generators ran; compared as it was before"})
          hand_out = []
          for copy, reason in hand.items():
              if (base / copy).is_file():
                  hand_out.append({"copy": copy, "reason": str(reason)})
              else:
                  findings.append({"kind": "hand-missing", "copy": copy, "generator": "-",
                                   "detail": "the manifest says it is made by hand; it is not in the manuscript"})
          listed_lower = {c.lower() for c in list(owner) + list(hand)}
          for pat, files in covered_files(base, covers).items():
              if not files:
                  findings.append({"kind": "covers-empty", "copy": str(pat), "generator": "-",
                                   "detail": "this covers pattern matches no file, so nothing under it is checked"})
              for f in files:
                  if f.lower() not in listed_lower and not any(x["copy"] == f for x in findings if x["kind"] == "unlisted"):
                      findings.append({"kind": "unlisted", "copy": f, "generator": "-",
                                       "detail": "under covers, named by no generator and no hand entry"})
          checked = len(same) + sum(1 for f in findings if f["kind"] == "differs")
          hard = [f for f in findings if f["kind"] in HARD]
          parts = [f"{checked} 份副本重跑对照:一致 {len(same)}"]
          diff = [f["copy"] for f in findings if f["kind"] == "differs"]
          if diff:
              parts.append(f"不一致 {len(diff)}({'、'.join(diff[:3])}{' 等' if len(diff) > 3 else ''})")
          other = [f for f in hard if f["kind"] != "differs"]
          if other:
              parts.append("另有 " + "、".join(f"{k} {sum(1 for f in other if f['kind'] == k)}"
                                                for k in HARD if any(f["kind"] == k for f in other)))
          if hand_out:
              parts.append(f"手做 {len(hand_out)} 份不查")
          payload = {
              "schema_version": 1,
              "base": str(base),
              "manifest": str(mpath),
              "summary_zh": ";".join(parts),
              "copies_checked": checked,
              "same": same,
              "findings": findings,
              "hand": hand_out,
              "generators": generators,
              "hard_finding_count": len(hard),
              "limits": "not decided here: whether a generator computes the right thing or reads the right artifact; the "
                        "generator is not sandboxed. A copy is checked against what the committed generator and artifacts "
                        "produce",
          }
          if a.json:
              print(json.dumps(payload, indent=2, ensure_ascii=False))
          else:
              print(payload["summary_zh"])
              for g in generators:
                  print(f"  {g['name']}: {g['repo']} @ {(g['commit'] or '?')[:7]}"
                        + (f", {g['seconds']} s" if g["seconds"] is not None else "")
                        + (f"; uncommitted, not used: {', '.join(g['dirty'][:5])}" if g["dirty"] else ""))
              for f in findings:
                  print(f"\n{f['kind']}  {f['copy']}  [{f['generator']}]\n    {f['detail']}")
                  for line in f.get("lines") or []:
                      print(f"    {line}")
              for h in hand_out:
                  print(f"\nhand  {h['copy']}: {h['reason']}")
              print(f"\n{payload['limits'][0].upper()}{payload['limits'][1:]}")
          if hard:
              return 1
          if not checked:
              print("NOTHING CHECKED: no copy was compared with a generator's output. This is not a pass.", file=sys.stderr)
              return 2
          return 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • audit-method-ledger.py 28 KB
      #!/usr/bin/env python3
      r"""Bind the sentences that say what was done to where it was done.
      
          python3 audit-method-ledger.py --base-dir <manuscript tree> --ledger <method-ledger.tsv>
              [--repo NAME=PATH[@COMMIT]] [--repo-commit-file NAME=FILE] [--repo-prefix NAME=PREFIX]
              [--manuscript-prefix PREFIX] [--full FILE] [--note-flags REGEX] [--git <manuscript repository>]
              [--state <ids.json>] [--json] [--migrate-headings]
      
      The gap this closes (spec docs/specs/2026-09-29-method-ledger.md). A sentence about methods, data, design, what a
      figure shows or a result judged without a value carries no claim, number or citation, so the other ledgers never
      read it. On one manuscript every check was current while a figure drew two independent controls as nested, and the
      methods named the wrong scoring function. A method ledger binds each such sentence to the code, config or output
      that shows it. It is a TSV:
      
          id  loc  sentence  claim_type  pointer  evidence  verdict  note  checked
      
      This script keeps it true after it is written. It reports:
      
        sentence-changed          the row's sentence is no longer in the file `loc` names (headings are ignored both sides)
        pointer-missing           a pointer names a file, line range, JSON/YAML key, verbatim fragment, \label or commit
                                  that is not there
        pointer-by-line           a pointer gives a line number into the manuscript's own .tex (line numbers drift with
                                  every edit; point by \label, heading or a fragment instead)
        no-pointer                a row with a claim has nothing checkable in its pointer
        open-verdict              mismatch / unlocated, or partial without `checked`
        bad-verdict               a verdict the ledger does not define
        retired-but-present       a row retired for a deleted sentence whose sentence is back
        row-dropped               a row id seen on an earlier run is gone without having been retired (needs --state)
        duplicate-id              two rows share an id
        note-contradicts-verdict  a `match` row whose note says the sentence is unmarked or says too much
        unledgered                a sentence of a --full file with no row
      
      Pointers are read out of the cell's text, as the ledger writes them: `;` or `;` between them, prose around them.
      A path is looked up in the pinned repositories (at their commit) and then in the manuscript tree. Line numbers into a
      pinned repository are allowed, since a locked commit does not move.
      
      Exit 0 when nothing is reported, 1 when anything is, 2 when nothing could be examined: the ledger is missing, empty
      or has no row with a pointer, or a pinned repository or commit is not there.
      """
      import argparse
      import csv
      import difflib
      import fnmatch
      import json
      import re
      import subprocess
      import sys
      from pathlib import Path
      
      VERDICTS = ("match", "partial", "mismatch", "unlocated", "author", "none", "retired")
      OPEN = ("mismatch", "unlocated")
      NOTE_FLAGS = r"未标|没标|说过头|not marked|unlabell?ed|overclaim|post[- ]hoc"
      MANUSCRIPT_PREFIXES = ("overleaf:", "manuscript:", "ms:")
      EXT = r"jsonld|jsonl|json|ya?ml|py|mjs|js|md|tsv|csv|tex|txt|npy|npz|lock|sh|bib|zip|html|pdf|png|jpg|R|ipynb|toml|cfg"
      PATH_RX = re.compile(r"(?<![\w/.~-])((?:[A-Za-z][\w-]*:)?(?:~/)?(?:[\w\-.*]+/)*[\w\-.*]+\.(?:" + EXT + r")(?![\w.])"
                           r"|\bLICENSE\b)"
                           r"(?::(\d+(?:[-–]\d+)?(?:,\d+(?:[-–]\d+)?)*))?"
                           r"(?:#([^\s;;((、]+))?"
                           r"(?:@(?:\"([^\"]+)\"|「([^」]+)」))?")
      LABEL_RX = re.compile(r"\\label\{([^}]*)\}")
      COMMIT_RX = re.compile(r"\bcommit:([0-9a-f]{7,40})\b")
      HEADING = re.compile(r"\\(?:(?:sub)*section|(?:sub)?paragraph|chapter)\*?\s*")
      
      
      class Stop(Exception):
          """Nothing could be examined: exit 2 with the reason."""
      
      
      def git(repo, *args, binary=False):
          r = subprocess.run(["git", "-C", str(repo)] + list(args), capture_output=True)
          if r.returncode:
              return None
          return r.stdout if binary else r.stdout.decode("utf-8", "replace")
      
      
      def _group(text, i):
          """The end of the {...} group starting at text[i] == '{', or -1."""
          depth = 0
          for j in range(i, len(text)):
              c = text[j]
              if c == "\\":
                  continue
              if c == "{" and (j == 0 or text[j - 1] != "\\"):
                  depth += 1
              elif c == "}" and text[j - 1] != "\\":
                  depth -= 1
                  if depth == 0:
                      return j + 1
          return -1
      
      
      def strip_headings(text):
          """Sectioning commands and the \\label right after them, removed: a run-in \\paragraph{...} stored in front of a
          paragraph's first sentence is not part of that sentence (FOR-AWT: inserting a sentence after the heading broke the
          row)."""
          out, i = [], 0
          for m in HEADING.finditer(text):
              if m.start() < i:
                  continue
              j = m.end()
              if j < len(text) and text[j] == "[":
                  k = text.find("]", j)
                  j = k + 1 if k >= 0 else j
              if j >= len(text) or text[j] != "{":
                  continue
              end = _group(text, j)
              if end < 0:
                  continue
              lab = re.match(r"\s*\\label\{[^}]*\}", text[end:])
              out.append(text[i:m.start()])
              out.append(" ")
              i = end + (lab.end() if lab else 0)
          out.append(text[i:])
          return "".join(out)
      
      
      def norm(t):
          t = re.sub(r"(?m)(?<!\\)%.*$", "", t or "")
          return re.sub(r"\s+", " ", t).strip()
      
      
      def body(t):
          return norm(strip_headings(norm(t)))
      
      
      def split_sentences(t):
          t = re.sub(r"\\begin\{(figure|table)\*?\}.*?\\end\{\1\*?\}", " ", t, flags=re.S)
          t = body(t)
          return [s for s in re.split(r"(?<=[.?!])\s+(?=[A-Z\\])", t) if len(s.split()) >= 4]
      
      
      class GitSource:
          """A repository at a commit, read through one `git cat-file --batch` process."""
      
          def __init__(self, path, commit):
              self.path, self.commit = Path(path).expanduser(), commit
              self._files, self._proc, self._cache = None, None, {}
      
          def files(self):
              if self._files is None:
                  out = git(self.path, "ls-tree", "-r", "--name-only", self.commit)
                  self._files = out.split("\n") if out else []
              return self._files
      
          def read(self, rel):
              if rel in self._cache:
                  return self._cache[rel]
              if self._proc is None:
                  self._proc = subprocess.Popen(["git", "-C", str(self.path), "cat-file", "--batch"],
                                                stdin=subprocess.PIPE, stdout=subprocess.PIPE)
              self._proc.stdin.write(f"{self.commit}:{rel}\n".encode())
              self._proc.stdin.flush()
              head = self._proc.stdout.readline().split()
              out = None
              if len(head) == 3 and head[1] == b"blob":
                  out = self._proc.stdout.read(int(head[2]))
                  self._proc.stdout.read(1)
              elif len(head) == 3:
                  self._proc.stdout.read(int(head[2]) + 1)
              self._cache[rel] = out
              return out
      
      
      SKIP_DIRS = {".git", "node_modules", ".venv", "venv", "__pycache__", ".mypy_cache"}
      
      
      class DiskSource:
          """Files on disk: the manuscript tree the loop extracted, or a working copy (tracked or not)."""
      
          def __init__(self, base):
              self.base = Path(base).expanduser()
              self._files = None
      
          def files(self):
              if self._files is None:
                  import os
                  out = []
                  for root, dirs, names in os.walk(self.base):
                      dirs[:] = [d for d in dirs if d not in SKIP_DIRS]
                      rel = Path(root).relative_to(self.base).as_posix()
                      out += [n if rel == "." else f"{rel}/{n}" for n in names]
                  self._files = out
              return self._files
      
          def read(self, rel):
              p = self.base / rel
              try:
                  return p.read_bytes() if p.is_file() else None
              except OSError:
                  return None
      
      
      class Repo:
          """A named repository: pinned (its commit, with its working copy kept only to say a file is outside the pin) or
          not (its working copy)."""
      
          def __init__(self, name, path, commit):
              self.name, self.path, self.commit = name, Path(path).expanduser(), commit
              self.git = GitSource(self.path, commit) if commit else None
              self.disk = DiskSource(self.path)
              self.prefixes = []
      
      
      def _key_tokens(expr):
          out = []
          for name, br in re.findall(r"([^.\[\]]+)|\[([^\]]*)\]", expr):
              out.append(("name", name) if name else ("index", br))
          return out
      
      
      def _resolve(obj, toks):
          if not toks:
              return True
          (kind, v), rest = toks[0], toks[1:]
          if kind == "index":
              rng = re.match(r"^(\w+)\.\.(\w+)$", v)
              if rng:  # [c01..c18]: both ends are there, and the rest holds at each
                  ends = [rng.group(1), rng.group(2)]
                  if isinstance(obj, dict):
                      # an end names a key, or a whole word inside one (c01 in "pool/c01.name.jpg")
                      def at(e):
                          return [k for k in obj if k == e or re.search(r"(?<![A-Za-z0-9])" + re.escape(e) + r"(?![A-Za-z0-9])", str(k))]
                      return all(at(e) and any(_resolve(obj[k], rest) for k in at(e)) for e in ends)
                  if isinstance(obj, list) and all(e.isdigit() for e in ends):
                      return all(int(e) < len(obj) and _resolve(obj[int(e)], rest) for e in ends)
                  return False
              sel = re.match(r"^([^=]+)=(.*)$", v)
              if sel:  # [model=x]: an element whose field has that value
                  k, want = sel.group(1).strip(), sel.group(2).strip()
                  items = obj if isinstance(obj, list) else (list(obj.values()) if isinstance(obj, dict) else [])
                  return any(isinstance(x, dict) and str(x.get(k)) == want and _resolve(x, rest) for x in items)
              if v in ("*", ""):
                  items = list(obj.values()) if isinstance(obj, dict) else (obj if isinstance(obj, list) else [])
                  return any(_resolve(x, rest) for x in items)
              if isinstance(obj, list):
                  if v.lstrip("-").isdigit():
                      i = int(v)
                      return -len(obj) <= i < len(obj) and _resolve(obj[i], rest)
                  return any(_resolve(x, rest) for x in obj
                             if isinstance(x, dict) and any(str(y) == v for y in x.values()) or x == v)
              return isinstance(obj, dict) and v in obj and _resolve(obj[v], rest)
          if isinstance(obj, dict):
              return v in obj and _resolve(obj[v], rest)
          if isinstance(obj, list):
              return any(_resolve(x, toks) for x in obj if isinstance(x, (dict, list)))
          return False
      
      
      def key_found(obj, expr):
          """A key path resolves; in a JSON Schema, a field is found by name among its properties and definitions."""
          if _resolve(obj, _key_tokens(expr)):
              return True
          if isinstance(obj, dict) and "$schema" in obj:
              names = [v for k, v in _key_tokens(expr) if k == "name"]
      
              def props(o):
                  if isinstance(o, dict):
                      for k, v in o.items():
                          if k in ("properties", "$defs", "definitions") and isinstance(v, dict):
                              yield from v.keys()
                          yield from props(v)
                  elif isinstance(o, list):
                      for v in o:
                          yield from props(v)
              return bool(names) and names[-1] in set(props(obj))
          return False
      
      
      def key_alternatives(expr):
          """`a.b[*].c,d` is a.b[*].c and a.b[*].d."""
          expr = expr.rstrip(".,,。")
          parts = [p for p in re.split(r",(?![^\[]*\])", expr) if p]
          head = parts[0]
          stem = head.rsplit(".", 1)[0] + "." if "." in head else ""
          # each later part is a sibling of the first key's last segment, or a key of its own: either reading will do
          return [head] + [[stem + p, p] if stem else [p] for p in parts[1:]]
      
      
      def parse_structured(rel, raw):
          """(object, None) or (None, why it cannot be read)."""
          text = raw.decode("utf-8", "replace")
          if rel.endswith((".json", ".jsonld", ".ipynb")):
              try:
                  return json.loads(text), None
              except ValueError as e:
                  return None, f"不是合法 JSON({e})"
          if rel.endswith(".jsonl"):
              try:
                  return [json.loads(x) for x in text.splitlines() if x.strip()], None
              except ValueError as e:
                  return None, f"不是合法 JSONL({e})"
          if rel.endswith((".yaml", ".yml")):
              try:
                  import yaml
              except ImportError:
                  return None, "unverifiable"
              try:
                  return yaml.safe_load(text), None
              except Exception as e:  # yaml's errors share no public base worth naming here
                  return None, f"不是合法 YAML({e})"
          return None, "text"
      
      
      class Checker:
          def __init__(self, a):
              self.tree = DiskSource(a.base_dir)
              self.repos = []
              prefixes = {}
              for spec in a.repo_prefix:
                  name, _, pre = spec.partition("=")
                  prefixes.setdefault(name, []).append(pre)
              commit_files = dict(x.split("=", 1) for x in a.repo_commit_file)
              for spec in a.repo:
                  name, _, rest = spec.partition("=")
                  path, _, commit = rest.partition("@")
                  if name in commit_files:
                      raw = self.tree.read(commit_files[name])
                      if raw is None:
                          raise Stop(f"仓 {name} 的锁定提交文件读不到:{commit_files[name]}")
                      commit = raw.decode().strip()
                  if not Path(path).expanduser().is_dir():
                      raise Stop(f"仓 {name} 不在:{path}")
                  if commit and git(path, "cat-file", "-e", f"{commit}^{{commit}}") is None:
                      raise Stop(f"仓 {name} 里没有锁定的提交 {commit}")
                  r = Repo(name, path, commit or None)
                  r.prefixes = prefixes.get(name, [])
                  self.repos.append(r)
              self.tree = DiskSource(a.base_dir)
              self.working = DiskSource(a.git) if a.git else None
              self.ms_prefixes = tuple(a.manuscript_prefix or MANUSCRIPT_PREFIXES)
              self.git = a.git
              self.notes, self.outside, self.keys_checked = [], [], 0
              self._labels = None
              self._text = {}
      
          def labels(self):
              if self._labels is None:
                  found = set()
                  for f in self.tree.files():
                      if f.endswith(".tex"):
                          found |= set(LABEL_RX.findall(self.tree.read(f).decode("utf-8", "replace")))
                  self._labels = found
              return self._labels
      
          def sources(self, path):
              """(relative path, [(source, outside the pin?)]) to try, in order: pinned repositories at their commit, the
              manuscript at the index commit, unpinned working copies, the manuscript's working copy, and last the pinned
              repositories' working copies (found there is said: the file is not in the locked commit)."""
              for pre in self.ms_prefixes:
                  if path.startswith(pre):
                      return path[len(pre):], [(self.tree, False)] + ([(self.working, False)] if self.working else [])
              for r in self.repos:
                  for pre in [r.name + ":"] + r.prefixes:
                      if path.startswith(pre):
                          return path[len(pre):], ([(r.git, False), (r.disk, True)] if r.git else [(r.disk, False)])
              m = re.match(r"([0-9a-f]{7,40}):(.+)$", path)
              if m:
                  return m.group(2), [(GitSource(p, m.group(1)), False)
                                      for p in [r.path for r in self.repos] + ([self.working.base] if self.working else [])
                                      if git(p, "cat-file", "-e", f"{m.group(1)}^{{commit}}") is not None]
              if path.startswith("~/"):
                  return path, [(DiskSource("/"), False)]
              out = [(r.git, False) for r in self.repos if r.git] + [(self.tree, False)]
              out += [(r.disk, False) for r in self.repos if not r.git]
              out += [(self.working, False)] if self.working else []
              out += [(r.disk, True) for r in self.repos if r.git]
              return path, out
      
          @staticmethod
          def lookup(src, rel):
              if rel.startswith("~/"):
                  rel = str(Path(rel).expanduser()).lstrip("/")
                  return rel if src.read(rel) is not None else None  # an absolute path is read, never searched for
              if any(ch in rel for ch in "*?"):
                  pat = "*/" + rel
                  return next((f for f in src.files() if fnmatch.fnmatch(f, rel) or fnmatch.fnmatch(f, pat)), None)
              if src.read(rel) is not None:
                  return rel
              if "/" not in rel:
                  return next((f for f in src.files() if f.split("/")[-1] == rel), None)
              return next((f for f in src.files() if f.endswith("/" + rel)), None)  # a path given from a subdirectory
      
          def find(self, path, cell=""):
              """(source, relative path, bytes) or None. A bare file name is also looked for inside the zip archives the
              same cell names."""
              rel, srcs = self.sources(path)
              for src, outside in srcs:
                  if src is None:
                      continue
                  if isinstance(src, DiskSource) and src.base == Path("/") and not rel.startswith("~/"):
                      continue
                  hit = self.lookup(src, rel)
                  if hit is not None:
                      raw = src.read(hit)
                      if raw is not None:
                          if outside:
                              self.outside.append(path)
                          return src, hit, raw
              if "/" not in rel:
                  import io
                  import zipfile
                  for z in re.findall(r"((?:~/)?[\w\-./]+\.zip)", cell):
                      if z == path:
                          continue
                      got = self.find(z)
                      if got is None:
                          continue
                      try:
                          names = zipfile.ZipFile(io.BytesIO(got[2])).namelist()
                      except zipfile.BadZipFile:
                          continue
                      if any(fnmatch.fnmatch(n.split("/")[-1], rel) for n in names):
                          return got[0], f"{z}!{rel}", b""
              return None
      
          def check_pointer(self, rid, m, cell=""):
              """Errors for one path pointer."""
              path, lines, key, frag = m.group(1).rstrip("."), m.group(2), m.group(3), m.group(4) or m.group(5)
              got = self.find(path, cell)
              if got is None:
                  return [("pointer-missing", rid, path)]
              src, rel, raw = got
              errs = []
              if lines:
                  if src is self.tree and rel.endswith(".tex") and "!" not in rel:
                      errs.append(("pointer-by-line", rid, f"{rel}:{lines}"))
                  n = raw.count(b"\n") + (0 if raw.endswith(b"\n") else 1)
                  top = max(int(x) for x in re.findall(r"\d+", lines))
                  if top > n:
                      errs.append(("pointer-missing", rid, f"{rel}:{lines} 超出 {n} 行"))
              if key:
                  obj, why = parse_structured(rel, raw)
                  if why == "unverifiable":
                      self.notes.append(f"{rid} {rel}#{key}:没法核(缺 PyYAML)")
                  elif why == "text":
                      if key.rstrip(".,,。") not in raw.decode("utf-8", "replace"):
                          errs.append(("pointer-missing", rid, f"{rel}#{key}(文件里没有这个词)"))
                  elif why:
                      errs.append(("pointer-missing", rid, f"{rel}#{key}:{why}"))
                  else:
                      self.keys_checked += 1
                      for alt in key_alternatives(key):
                          readings = alt if isinstance(alt, list) else [alt]
                          if not any(key_found(obj, r) for r in readings):
                              errs.append(("pointer-missing", rid, f"{rel}#{readings[0]}(没有这个键)"))
              if frag and norm(frag) not in norm(raw.decode("utf-8", "replace")):
                  errs.append(("pointer-missing", rid, f"{rel}@「{frag[:40]}」(文件里没有这段)"))
              return errs
      
          def file_body(self, rel):
              if rel not in self._text:
                  raw = None
                  for pre in self.ms_prefixes:
                      if rel.startswith(pre):
                          rel = rel[len(pre):]
                  raw = self.tree.read(rel)
                  self._text[rel] = None if raw is None else (norm(raw.decode("utf-8", "replace")),
                                                               body(raw.decode("utf-8", "replace")))
              return self._text[rel]
      
          def sentence_present(self, loc, sentence):
              got = self.file_body(loc.split(":")[0])
              if got is None:
                  return False
              raw, stripped = got
              s = body(sentence)
              # a row can hold a heading's own words (a run-in \paragraph{...} that is a sentence): found as written too
              return bool(s) and s in stripped or norm(sentence) in raw
      
          def commit_exists(self, sha):
              repos = [r.path for r in self.repos] + ([self.git] if self.git else [])
              return any(git(p, "cat-file", "-e", f"{sha}^{{commit}}") is not None for p in repos)
      
      
      def read_ledger(path):
          try:
              with open(path, encoding="utf-8", newline="") as f:
                  rows = list(csv.DictReader(f, delimiter="\t"))
          except (OSError, UnicodeDecodeError) as e:
              raise Stop(f"台账读不了:{path}({e})")
          if not rows:
              raise Stop(f"台账为空:{path}")
          need = {"id", "loc", "sentence", "claim_type", "pointer", "verdict"}
          missing = need - set(rows[0].keys())
          if missing:
              raise Stop(f"台账缺列:{', '.join(sorted(missing))}")
          return rows
      
      
      def migrate(path, rows):
          """A diff that removes heading prefixes from sentence cells. Printed, never written."""
          old = Path(path).read_text(encoding="utf-8").splitlines(keepends=True)
          fields = list(rows[0].keys())
          new_rows = []
          for r in rows:
              s = r["sentence"] or ""
              stripped = norm(strip_headings(s))
              if stripped and stripped != norm(s) and HEADING.match(s.lstrip()):
                  r = dict(r, sentence=stripped)
              new_rows.append(r)
          import io
          buf = io.StringIO()
          w = csv.DictWriter(buf, fieldnames=fields, delimiter="\t", lineterminator="\n", quoting=csv.QUOTE_MINIMAL)
          w.writeheader()
          w.writerows(new_rows)
          new = buf.getvalue().splitlines(keepends=True)
          sys.stdout.writelines(difflib.unified_diff(old, new, fromfile=str(path), tofile=str(path) + "(去掉标题前缀)"))
      
      
      def audit(a):
          ledger = Path(a.ledger) if Path(a.ledger).is_absolute() else Path(a.base_dir) / a.ledger
          rows = read_ledger(ledger)
          if a.migrate_headings:
              migrate(ledger, rows)
              return None
          ck = Checker(a)
          flags = re.compile(a.note_flags or NOTE_FLAGS, re.I)
          errors, pending, seen, with_pointer = [], 0, {}, 0
          for r in rows:
              rid = (r.get("id") or "").strip()
              if rid in seen:
                  errors.append(("duplicate-id", rid, f"与第 {seen[rid]} 行同号"))
              seen.setdefault(rid, len(seen) + 2)
              verdict = (r.get("verdict") or "").strip()
              vkind = verdict.split()[0] if verdict else ""
              sentence, loc = r.get("sentence") or "", r.get("loc") or ""
              if vkind == "retired":
                  parts = verdict.split()
                  if len(parts) < 2 or not ck.commit_exists(parts[1]):
                      errors.append(("bad-verdict", rid, f"retired 要带删去它的提交号,且提交要在仓里:{verdict[:40]}"))
                  if sentence and ck.sentence_present(loc, sentence):
                      errors.append(("retired-but-present", rid, loc))
                  continue
              if not ck.sentence_present(loc, sentence):
                  errors.append(("sentence-changed", rid, loc))
              if (r.get("claim_type") or "").strip() == "none":
                  continue
              if vkind not in VERDICTS:
                  errors.append(("bad-verdict", rid, verdict[:30] or "(空)"))
              if vkind in OPEN:
                  errors.append(("open-verdict", rid, vkind))
              if vkind == "partial":
                  pending += 1
                  if not (r.get("checked") or "").strip():
                      errors.append(("open-verdict", rid, "partial 没写 checked"))
              if vkind == "match" and flags.search(r.get("note") or ""):
                  errors.append(("note-contradicts-verdict", rid, (r.get("note") or "")[:60]))
              pointer = r.get("pointer") or ""
              paths = list(PATH_RX.finditer(pointer))
              labels = LABEL_RX.findall(pointer)
              commits = COMMIT_RX.findall(pointer)
              if not (paths or labels or commits):
                  errors.append(("no-pointer", rid, pointer[:60] or "(空)"))
                  continue
              with_pointer += 1
              for m in paths:
                  errors += ck.check_pointer(rid, m, pointer)
              for lab in labels:
                  if lab not in ck.labels():
                      errors.append(("pointer-missing", rid, "\\label{" + lab + "}"))
              for sha in commits:
                  if not ck.commit_exists(sha):
                      errors.append(("pointer-missing", rid, "commit:" + sha))
          if not with_pointer:
              raise Stop("台账里没有一行带可检查的出处,不把这当成通过")
          ledgered = [body(r.get("sentence") or "") for r in rows]
          for f in a.full:
              raw = ck.tree.read(f)
              if raw is None:
                  raise Stop(f"要全覆盖的文件不在:{f}")
              for s in split_sentences(raw.decode("utf-8", "replace")):
                  if not any(x and (s in x or x in s) for x in ledgered):
                      errors.append(("unledgered", f, s[:70]))
          if a.state:
              sp = Path(a.state)
              try:
                  known = set(json.loads(sp.read_text(encoding="utf-8")).get("ids") or [])
              except (OSError, ValueError, AttributeError):
                  known = None
              now = set(seen)
              if known is not None:
                  for rid in sorted(known - now):
                      errors.append(("row-dropped", rid, "上次还在,现在没了;删句要写 retired <提交号>,不删行"))
              sp.parent.mkdir(parents=True, exist_ok=True)
              sp.write_text(json.dumps({"ids": sorted((known or set()) | now)}, ensure_ascii=False), encoding="utf-8")
          kinds = {}
          for k, _, _ in errors:
              kinds[k] = kinds.get(k, 0) + 1
          claims = sum(1 for r in rows if (r.get("claim_type") or "").strip() != "none")
          summary = (f"台账 {len(rows)} 行(有主张 {claims});报错 {len(errors)}"
                     + ("(" + "、".join(f"{k} {n}" for k, n in sorted(kinds.items())) + ")" if kinds else "")
                     + f";待定(partial){pending} 行")
          summary += f";核了 {ck.keys_checked} 个 JSON/YAML 键"
          if ck.notes:
              summary += f",{len(ck.notes)} 个没法核"
          if ck.outside:
              summary += f";{len(set(ck.outside))} 个出处只在磁盘上、不在锁定的提交里"
          return {"rows": len(rows), "claims": claims, "pending": pending, "kinds": kinds,
                  "errors": [{"kind": k, "id": i, "detail": d} for k, i, d in errors], "notes": ck.notes,
                  "outside_pin": sorted(set(ck.outside)), "keys_checked": ck.keys_checked,
                  "pinned": {r.name: r.commit for r in ck.repos}, "summary_zh": summary}
      
      
      def main(argv=None):
          ap = argparse.ArgumentParser(description=__doc__.split("\n\n")[0])
          ap.add_argument("--base-dir", required=True)
          ap.add_argument("--ledger", required=True)
          ap.add_argument("--repo", action="append", default=[], help="NAME=PATH[@COMMIT]")
          ap.add_argument("--repo-commit-file", action="append", default=[], help="NAME=FILE in the manuscript tree")
          ap.add_argument("--repo-prefix", action="append", default=[], help="NAME=PREFIX that routes a pointer to NAME")
          ap.add_argument("--manuscript-prefix", action="append", default=[])
          ap.add_argument("--full", action="append", default=[])
          ap.add_argument("--note-flags")
          ap.add_argument("--git", help="the manuscript repository, for retired and commit: pointers")
          ap.add_argument("--state", help="row ids seen on earlier runs (kept, never pruned)")
          ap.add_argument("--json", action="store_true")
          ap.add_argument("--migrate-headings", action="store_true")
          a = ap.parse_args(argv)
          try:
              out = audit(a)
          except Stop as e:
              print(f"audit-method-ledger: {e}", file=sys.stderr)
              return 2
          if out is None:
              return 0
          if a.json:
              print(json.dumps(out, ensure_ascii=False))
          else:
              print(out["summary_zh"])
              for e in out["errors"]:
                  print(f"  {e['kind']}\t{e['id']}\t{e['detail']}")
              for n in out["notes"]:
                  print(f"  注:{n}")
              for n in out["outside_pin"]:
                  print(f"  注:{n} 只在磁盘上找到,不在锁定的提交里")
          return 1 if out["errors"] else 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • audit-number-ledger.py 18.1 KB
      #!/usr/bin/env python3
      r"""Bind the numbers a manuscript reports to the artifacts they come from.
      
          python3 audit-number-ledger.py --base-dir <manuscript> --ledger <numbers.tsv> [--ledger <more.tsv> ...]
                                         [--ledger-files <numbers.tsv>=<files,...> ...] [--json] [--pairs] [--allow-empty]
      
      The gap this closes. A manuscript already has one guard on its numbers: a
      frequency table of numeric tokens taken before and after a prose pass, so an
      edit that moves a number shows up. That answers "did this pass change a
      number". It cannot answer either of the two questions that have actually gone
      wrong:
      
        - is this the artifact's number? Hand-copied figures drift, and three
          hand-counted totals in one day were each wrong by one or two.
        - is the sentence carrying the scope the number is only true within? A ratio
          true of 15 cells was written as if it held for all of them, and a pooled
          share was quoted alone against the manuscript's own warning not to.
      
      A row binds one reported number to one artifact:
      
          printed<TAB>in_artifact<TAB>scope<TAB>artifact<TAB>locator
          63.5<TAB>0.635<TAB>pooled<TAB>figures/fig.tex<TAB>explained variance 0.635
      
      `printed` is the value as the prose prints it; `in_artifact` is the value as the
      file writes it. They are two columns because on a real manuscript they differ:
      the figure carries `0.635` and the text prints `63.5\%`. A version with one
      column reported that as a broken binding. Converting silently would have been
      worse — the conversion is recorded and shown, never inferred. `scope` is the
      qualifier every sentence reporting the number must carry, or `-`. `locator` is
      the verbatim string in the artifact, so the binding can be re-read.
      
        locator-not-in-artifact    the locator is not verbatim in the artifact
        value-not-in-locator       the locator does not itself carry in_artifact, so
                                   the row looks bound and binds nothing
        printed-artifact-mismatch  printed and in_artifact are neither equal nor the
                                   same value as a proportion and a percentage
        number-not-in-manuscript   the printed value is no longer reported; the row
                                   is stale
        artifact-missing           the artifact is not on disk
        scope-missing              a sentence reports the number without its scope
        unledgered-number          PROMPT: a number reported with no row. Years,
                                   counts of sections and sample sizes all land here,
                                   so this is a coverage list, not a finding.
      
      An optional sixth column, `copies`, is how many prose sentences report the
      number. Without it the only test was "does the value appear somewhere", so a
      number printed in the abstract and the introduction could drift in one of them
      and still pass on the other copy:
      
        copies-changed             the number of sentences reporting it is not the
                                   count the ledger records
      
      `scope` is matched literally, and that is its limit. A scope ending in a digit
      must not continue into another digit (K=1 is not found inside K=10), and a
      locator ending in a digit must not continue into another digit or a decimal
      point (0.635 is not found inside 0.6357). On the real manuscript,
      the first scope token tried flagged two sentences that carry the scope in
      other words; a token those sentences actually contain passed both. Pick
      a token the correct sentences actually contain, or the column produces noise
      rather than a guard. A scope of `-` switches the check off for that row.
      
      What this does NOT do: decide whether the artifact is the right one, or whether
      a number is correctly derived from it. Nor does it see two numbers swap places
      inside one sentence ("from 2.1 to 30.5" for "from 30.5 to 2.1"): every value
      and every scope is still present. Only a reader catches that. It checks that the value printed in the
      prose is the value written in a file, under the scope the ledger records.
      
      Exit: 1 on a hard finding, 2 when no row was checked at all unless
      --allow-empty, 0 otherwise.
      """
      import argparse
      import fnmatch
      import json
      import re
      import sys
      from pathlib import Path
      
      COLUMNS = ["printed", "in_artifact", "scope", "artifact", "locator"]
      OPTIONAL = ["copies"]
      # A reported number: a decimal, a percentage or an integer of two digits or
      # more. Single digits are almost always prose ("the three requirements") and
      # would drown the coverage list.
      # A number written with thousands separators (4,207; LaTeX 4{,}207) is one number: read digit by digit it became
      # "207", a value the manuscript never reports.
      REPORTED = re.compile(r"(?<![\w.,])(\d{1,3}(?:,\d{3})+(?:\.\d+)?|\d+\.\d+|\d{2,})(?![\w.])")
      
      
      def clean_tex(text):
          text = re.sub(r"(?m)(?<!\\)%.*$", "", text)
          text = re.sub(r"\\(?:label|ref|eqref|cite|citep|citet|input|include)\*?\{[^}]*\}", " ", text)
          text = re.sub(r"\\(?:section|subsection|subsubsection|paragraph)\*?\{[^}]*\}", " ", text)
          text = re.sub(r"\\(?:emph|textbf|textit|texttt|text)\{([^}]*)\}", r"\1", text)
          text = text.replace("~", " ").replace("\\%", "%").replace("$", "").replace("{,}", ",")
          return re.sub(r"\s+", " ", text)
      
      
      def norm(text):
          return re.sub(r"\s+", " ", text).strip().lower()
      
      
      def source_files(base):
          out = []
          for suffix in ("*.tex", "*.md"):
              out += [p for p in sorted(base.rglob(suffix))
                      if not any(part.startswith(".") or part in {"node_modules", "build", "results", "audit"}
                                 for part in p.relative_to(base).parts)]
          return out
      
      
      def sentences(base):
          """[(file, sentence)] over the manuscript prose."""
          out = []
          for path in source_files(base):
              text = clean_tex(path.read_text(encoding="utf-8", errors="replace"))
              for sentence in re.split(r"(?<=[.])\s+(?=[A-Z\\])", text):
                  if sentence.strip():
                      out.append((str(path.relative_to(base)), sentence.strip()))
          return out
      
      
      def governs(rel, globs):
          """Whether a prose file is one a ledger answers for: a path, a directory prefix, or a glob (* crosses /)."""
          return any(rel == g or rel.startswith(g.rstrip("/") + "/") or fnmatch.fnmatch(rel, g) for g in globs)
      
      
      def read_ledger(path):
          rows = []
          lines = [l for l in path.read_text(encoding="utf-8").splitlines() if l.strip()]
          if not lines:
              return rows
          header = lines[0].split("\t")
          if header[:len(COLUMNS)] != COLUMNS or any(h not in OPTIONAL for h in header[len(COLUMNS):]):
              sys.exit(f"LEDGER_COLUMNS: expected {COLUMNS} (+ optional {OPTIONAL}), found {header}")
          for n, line in enumerate(lines[1:], 2):
              parts = line.split("\t")
              if len(parts) < len(COLUMNS):
                  sys.exit(f"LEDGER_COLUMNS: line {n} has {len(parts)} columns, expected {len(COLUMNS)}")
              row = dict(zip(header, parts))
              row["line"] = n
              rows.append(row)
          return rows
      
      
      def main(argv=None):
          ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
          ap.add_argument("--base-dir", default=".")
          ap.add_argument("--ledger", required=True, action="append",
                          help="a ledger; give it more than once (the text's numbers, a supplement's), and the rows are read together")
          ap.add_argument("--ledger-files", action="append", default=[], metavar="LEDGER=FILES",
                          help="the prose files a ledger answers for, comma-separated paths, directories or globs: its "
                               "numbers are looked for, counted and scope-checked there only. A ledger named in none reads "
                               "every file. Unledgered numbers are looked for everywhere")
          ap.add_argument("--json", action="store_true")
          ap.add_argument("--pairs", action="store_true", help="print number/locator pairs for reading")
          ap.add_argument("--allow-empty", action="store_true")
          a = ap.parse_args(argv)
          base = Path(a.base_dir).expanduser().resolve()
          ledger_paths = [Path(x).expanduser().resolve() for x in a.ledger]
          ledger_path = ledger_paths[0]
          if not base.is_dir():
              sys.exit(f"BASE_MISSING: {base}")
          for lp in ledger_paths:
              if not lp.is_file():
                  sys.exit(f"LEDGER_MISSING: {lp}")
          # 09-27: a supplement read beside the text put a second copy of the text's numbers in view, and every copies count
          # in the text's ledger read as changed. Each ledger answers for its own files.
          covers = {}
          for spec in a.ledger_files:
              name, _, files = spec.partition("=")
              lp = Path(name).expanduser().resolve()
              globs = [g.strip() for g in files.split(",") if g.strip()]
              if lp not in ledger_paths:
                  sys.exit(f"LEDGER_FILES_UNKNOWN: {name} is not one of the --ledger files")
              if not globs:
                  sys.exit(f"LEDGER_FILES_EMPTY: {name} names no files")
              covers[lp] = covers.get(lp, []) + globs
      
          prose = sentences(base)
          rows = []
          for lp in ledger_paths:
              for row in read_ledger(lp):
                  row["_ledger"] = lp
                  rows.append(row)
          # An artifact is where a number comes from, not where the prose reports it. A table file read as prose kept a
          # number "reported" after the sentence carrying it was cut, and put every cell on the coverage list.
          sources = {(base / r["artifact"]).resolve() for r in rows}
          # Each ledger's own artifacts are not its prose, and only its own: a table that the text's ledger draws numbers
          # from is prose that the supplement's ledger counts copies in (09-27: read together, the table vanished from the
          # supplement's count). With one ledger this is the rule it always was.
          own_sources = {lp: {(base / r["artifact"]).resolve() for r in rows if r["_ledger"] == lp} for lp in ledger_paths}
          every = prose
          prose = [(f, s) for f, s in every if (base / f).resolve() not in sources]
          read = {f for f, _ in every}
          for lp, globs in covers.items():
              if not any(governs(f, globs) for f in read):
                  # A declared scope that reads nothing would pass every row as absent-or-fine: fail instead.
                  sys.exit(f"LEDGER_FILES_MATCH_NOTHING: {lp.name} answers for {', '.join(globs)}, and no prose file matches")
          findings, pairs = [], []
          ledgered_numbers = set()
      
          for row in rows:
              where = f"{row['_ledger'].name}:{row['line']}"
              number = row["printed"].strip()
              in_artifact = row["in_artifact"].strip()
              scope = row["scope"].strip()
              ledgered_numbers.add(number)
      
              relation = None
              if number == in_artifact:
                  relation = "identical"
              else:
                  try:
                      bare = number.replace(",", "")
                      printed_value = float(bare)
                      artifact_value = float(in_artifact.replace("{,}", "").replace(",", ""))
                      # The prose prints the artifact's value to fewer decimals. Only exact rounding to the printed
                      # precision counts: 36.18 printed as 36.2 is a relation, printed as 36.1 is a finding.
                      places = len(bare.split(".")[1]) if "." in bare else 0
                      if printed_value == artifact_value:
                          relation = "the same value, written with thousands separators"
                      elif artifact_value and abs(printed_value - artifact_value * 100) < 1e-6:
                          relation = "printed as a percentage of the artifact's proportion"
                      elif printed_value and abs(artifact_value - printed_value * 100) < 1e-6:
                          relation = "printed as a proportion of the artifact's percentage"
                      elif abs(round(artifact_value, places) - printed_value) < 1e-9:
                          relation = f"the artifact's value rounded to {places} decimal place(s)"
                      elif abs(round(artifact_value * 100, places) - printed_value) < 1e-9:
                          relation = f"the artifact's proportion as a percentage rounded to {places} decimal place(s)"
                  except ValueError:
                      pass
              if relation is None:
                  findings.append({"kind": "printed-artifact-mismatch", "location": where, "number": number,
                                   "detail": f"the prose prints {number} and the artifact writes {in_artifact}; "
                                             "record them as equal, or as a proportion and its percentage, or look again"})
      
              artifact = base / row["artifact"]
              if not artifact.is_file():
                  artifact = row["_ledger"].parent / row["artifact"]
              if not artifact.is_file():
                  findings.append({"kind": "artifact-missing", "location": where, "number": number,
                                   "detail": f"artifact not found: {row['artifact']}"})
              else:
                  text = norm(artifact.read_text(encoding="utf-8", errors="replace"))
                  locator = norm(row["locator"])
                  if not re.search(re.escape(locator) + (r"(?![\d.])" if locator[-1:].isdigit() else ""), text):
                      findings.append({"kind": "locator-not-in-artifact", "location": where, "number": number,
                                       "detail": f'"{row["locator"][:70]}" is not verbatim in {row["artifact"]}'})
              if in_artifact not in row["locator"]:
                  findings.append({"kind": "value-not-in-locator", "location": where, "number": number,
                                   "detail": f'the locator "{row["locator"][:60]}" does not carry {in_artifact}'})
      
              mine, skip = covers.get(row["_ledger"]), own_sources[row["_ledger"]]
              reporting = [(f, s) for f, s in every if (base / f).resolve() not in skip and (mine is None or governs(f, mine))
                           and re.search(rf"(?<![\w.,]){re.escape(number)}(?![\w.])", s)]
              if not reporting:
                  findings.append({"kind": "number-not-in-manuscript", "location": where, "number": number,
                                   "detail": f"{number} is no longer reported anywhere in the manuscript"})
              if reporting and (row.get("copies") or "").strip() not in ("", "-"):
                  want = int(row["copies"])
                  if len(reporting) != want:
                      findings.append({"kind": "copies-changed", "location": where, "number": number,
                                       "detail": f"{number} is reported in {len(reporting)} sentence(s), the ledger records "
                                                 f"{want}: a copy was changed, added or cut"})
              if reporting and scope and scope != "-":
                  scope_rx = re.escape(norm(scope)) + (r"(?!\d)" if norm(scope)[-1:].isdigit() else "")
                  for f, s in reporting:
                      if not re.search(scope_rx, norm(s)):
                          findings.append({"kind": "scope-missing", "location": f, "number": number,
                                           "detail": f'reports {number} without "{scope}", which it is only true within: {s[:90]}'})
              pairs.append({"printed": number, "in_artifact": in_artifact, "relation": relation,
                            "scope": scope, "artifact": row["artifact"], "locator": row["locator"]})
      
          seen = set()
          for f, s in prose:
              for m in REPORTED.finditer(s):
                  value = m.group(1)
                  if value in ledgered_numbers or (f, value) in seen:
                      continue
                  seen.add((f, value))
                  findings.append({"kind": "unledgered-number", "prompt": True, "location": f, "number": value,
                                   "detail": f"reported with no ledger row: {s[:90]}"})
      
          hard_kinds = {"locator-not-in-artifact", "value-not-in-locator", "printed-artifact-mismatch",
                        "number-not-in-manuscript", "artifact-missing", "scope-missing", "copies-changed"}
          hard = [f for f in findings if f["kind"] in hard_kinds]
          nothing = not rows
          payload = {
              "schema_version": 1,
              "base": str(base),
              "ledger": str(ledger_path),
              "ledgers": [str(p) for p in ledger_paths],
              "ledger_files": {p.name: g for p, g in covers.items()},
              "prose_sentences": len(prose),
              "ledger_rows": len(rows),
              "findings": findings,
              "hard_finding_count": len(hard),
              "unledgered_count": sum(1 for f in findings if f["kind"] == "unledgered-number"),
              "nothing_checked": nothing,
              "limits": {
                  "derivation": "not decided here: whether the artifact is the right one, or whether the number is correctly derived from it. This checks that the printed value is a value written in a file, under the scope the ledger records.",
                  "coverage": "unledgered-number ignores single digits, so counts written as words or as one digit are invisible to it",
                  "scope": "matched literally: a sentence that carries the scope in other words is flagged, so the token has to be one the correct sentences contain",
              },
          }
          if a.json:
              print(json.dumps(payload, indent=2, ensure_ascii=False))
          else:
              unledgered = payload["unledgered_count"]
              covered = len(rows) + unledgered
              print(f"number ledger: {len(rows)} row(s) against {len(prose)} sentence(s) under {base}")
              print(f"coverage: {len(rows)} of {covered} reported number(s) carry a row; "
                    f"{unledgered} do not, and nothing here checks them")
              for kind in ["locator-not-in-artifact", "value-not-in-locator", "printed-artifact-mismatch",
                           "number-not-in-manuscript", "artifact-missing", "scope-missing", "copies-changed",
                           "unledgered-number"]:
                  group = [f for f in findings if f["kind"] == kind]
                  if not group:
                      continue
                  print(f"\n{kind}{' (prompt, not a finding)' if group[0].get('prompt') else ''} ({len(group)})")
                  for f in group[:40]:
                      print(f"  {f['location']}  [{f['number']}]\n    {f['detail']}")
                  if len(group) > 40:
                      print(f"  … {len(group) - 40} more")
              if nothing:
                  print("\nNOTHING CHECKED: no ledger row. This is not a pass.")
              print(f"\nNot decided here: {payload['limits']['derivation']}")
          if a.pairs:
              for p in pairs:
                  print(f"\n[{p['printed']} printed · {p['in_artifact']} in the artifact · {p['relation']}]"
                        f"\n  scope:    {p['scope']}\n  artifact: {p['artifact']}\n  locator:  {p['locator']}")
          return 1 if hard else (2 if nothing and not a.allow_empty else 0)
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • audit-prose-fingerprint.py 36.3 KB
      #!/usr/bin/env python3
      """Deterministic prose-fingerprint audit against a corpus of published papers.
      
      This script does not rewrite prose. It measures a small set of stylistic
      constructions and reports where the target sits relative to a baseline corpus.
      
      The point is the *distribution*, not the count. A construction used twelve
      times as often as published work is a problem; so is one used at exactly the
      right rate but spread perfectly evenly, because human writers cluster. Every
      metric that can be distributed is therefore reported three ways: how often, how
      bunched (coefficient of variation of the gaps between occurrences, which is
      rate-independent), and how long the longest stretch without one runs.
      
      Baseline percentiles are only shown when a baseline corpus is supplied. Without
      one the numbers are still reported, but nothing here says what a good value is:
      "published papers in your own bibliography" is the only reference this tool
      trusts, and pooling gaps across papers of different rates inflates CV, so
      dispersion baselines are computed per document and never pooled.
      
      Python 3.8 stdlib only. PDF baselines require `pdftotext` on PATH; without it,
      PDFs are skipped and named in the report.
      """
      
      import argparse
      import fnmatch
      import json
      import math
      import re
      import shutil
      import statistics
      import subprocess
      import sys
      from pathlib import Path
      from typing import Dict, List, Optional, Sequence, Tuple
      
      SCHEMA_VERSION = 3
      
      # --- constructions -----------------------------------------------------------
      # Each is a family of surface forms, not a single string. The corrective
      # diptych in particular has five common shapes and counting only "rather than"
      # understates it by roughly half.
      CONTRAST = (
          r"\brather than\b|\band not\b|\b, not\b|\bnot\b[^.]{0,40}\bbut\b|\binstead of\b"
      )
      EXPLANATORY_COLON = r"[a-z]:\s+[a-z]"
      SEMICOLON = r";"
      DISCOURSE = (
          r"\b(?:[Hh]owever|[Mm]oreover|[Ff]urthermore|[Tt]hus|[Tt]herefore"
          r"|[Nn]evertheless|[Cc]onsequently|[Nn]onetheless)\b"
      )
      # A sentence that opens with a linking adverbial says how it stands to the sentence before it: a contrast, a
      # consequence, an example, an addition. Published papers open a steady share of their sentences this way; a draft
      # that opens almost none of them leaves the reader to supply every relation. The count is of sentences, not of
      # words, and only of the opening: "however" in mid-sentence is the discourse-marker rate above.
      # Not counted: enumerators (First, Second, Finally), which order a list without saying how its items relate, and
      # subordinators (Although, Because, While), which relate two clauses inside one sentence. "Instead of" and
      # "In addition to" open a phrase, not a link. Words that are links only when a comma follows them need the comma.
      # The list is closed: a link it does not name is not counted. audit-sentence-changes.py reads it from here.
      LINKING_OPENER = re.compile(
          r"^[\"'\u201c\u2018]?(?:"
          r"(?:However|Thus|Therefore|Hence|Moreover|Furthermore|Consequently|Nevertheless|Nonetheless|Accordingly"
          r"|Conversely|Similarly|Likewise|Indeed|Yet|Additionally|Specifically|Notably|Importantly|Meanwhile"
          r"|As a result|As a consequence|In contrast|By contrast|For example|For instance|In particular"
          r"|On the other hand|In other words|In turn|In summary|In short|To this end|Even so|Instead(?! of)"
          r"|In addition(?! to))\b"
          r"|(?:That is|Overall|Still|Otherwise|In practice|Taken together|Put differently)\s*,)")
      NOMINALISATION = r"\b\w+(?:tion|ment|ness|ity)s?\b"
      HEDGE = r"\b(?:may|might|could|appears?|suggests?|seems?|likely|plausibl\w+)\b"
      
      DISTRIBUTED = ("contrast", "explanatory_colon", "semicolon", "discourse_marker")
      PATTERNS = {
          "contrast": CONTRAST,
          "explanatory_colon": EXPLANATORY_COLON,
          "semicolon": SEMICOLON,
          "discourse_marker": DISCOURSE,
          "nominalisation": NOMINALISATION,
          "hedge": HEDGE,
      }
      
      TEXT_SUFFIXES = {".md", ".tex", ".txt"}
      STOP_HEADINGS = re.compile(r"\n\s*(?:References|REFERENCES|Bibliography|Works Cited)\s*\n")
      
      
      # --- extraction --------------------------------------------------------------
      
      # Share of whitespace tokens that are words. Measured 2026-09-22 on 81 real papers: lowest 0.579 (a table-heavy
      # paper), median 0.886; a PDF whose fonts extracted as symbols scored 0.036.
      MIN_WORD_SHARE = 0.30
      _WORD = re.compile(r"^[(\[\"'\u201c\u2018]?[A-Za-z][A-Za-z'\u2019\-]*[.,;:!?)\]\"'\u201d\u2019]*$")
      
      
      def word_share(text: str) -> float:
          tokens = text.split()
          return sum(1 for x in tokens if _WORD.match(x)) / len(tokens) if tokens else 0.0
      
      
      def read_pdf(path: Path) -> Optional[str]:
          if not shutil.which("pdftotext"):
              return None
          try:
              out = subprocess.run(
                  ["pdftotext", "-q", str(path), "-"],
                  capture_output=True, text=True, timeout=120,
              )
          except (OSError, subprocess.SubprocessError):
              return None
          if not out.stdout:
              return None
          # pdftotext preserves the printed line breaks, so a word hyphenated across
          # a line arrives as two tokens and inflates the sentence it sits in.
          return re.sub(r"-\n", "", out.stdout)
      
      
      def strip_markup(text: str, suffix: str) -> str:
          """Remove markup that is not prose the reader sees."""
          if suffix == ".tex":
              # Everything before \begin{document} is package loading and macro
              # definitions. It is not prose, and counting it inflates the word
              # total that every per-1k rate is divided by.
              body = re.split(r"\\begin\{document\}", text, maxsplit=1)
              if len(body) == 2:
                  text = body[1]
              text = re.sub(r"(?m)^\s*%.*$", "", text)
              # Alt text and comments are not read aloud in the flow of the argument,
              # but they do drift, so they are measured separately by --include-alt.
              text = re.sub(r"\\Description\{(?:[^{}]|\{[^{}]*\})*\}", " ", text)
              text = re.sub(r"\\(?:label|ref|eqref|cite\w*)\{[^}]*\}", " ", text)
              text = re.sub(r"\\begin\{(?:tabular|table|figure|itemize|enumerate|description|equation|align)\*?\}"
                            r".*?\\end\{(?:tabular|table|figure|itemize|enumerate|description|equation|align)\*?\}",
                            " ", text, flags=re.S)
              # An environment's name is not prose: \begin{center} used to leave the word "center" behind, and a
              # colon before it ("character by character: center") was counted as an explanatory colon.
              text = re.sub(r"\\(?:begin|end)\{[^}]*\}", " ", text)
              text = re.sub(r"\\[a-zA-Z]+\*?", " ", text)
              text = re.sub(r"[{}$&~\\]", " ", text)
          elif suffix == ".md":
              text = re.sub(r"```.*?```", " ", text, flags=re.S)
              text = re.sub(r"(?m)^\s*\|.*$", " ", text)
              text = re.sub(r"(?m)^#+\s*", "", text)
          m = list(STOP_HEADINGS.finditer(text))
          if m:
              text = text[: m[-1].start()]
          return re.sub(r"\s+", " ", text).strip()
      
      
      def load(path: Path) -> Optional[str]:
          if path.suffix.lower() == ".pdf":
              raw = read_pdf(path)
              return strip_markup(raw, ".txt") if raw else None
          if path.suffix.lower() in TEXT_SUFFIXES:
              return strip_markup(path.read_text(encoding="utf-8", errors="replace"), path.suffix.lower())
          return None
      
      
      def pipeline_of(suffix: str) -> str:
          """Which reading a suffix gets.
      
          `.tex` and `.md` have their markup removed: float environments, tables,
          captions, list bodies and citation commands all go. A PDF cannot be read
          that way -- pdftotext returns what was printed, furniture included. On one
          real manuscript the two readings of the SAME document differed by 40% in
          word count, which is the denominator of every per-1k rate, and the
          sentence-length lag-1 changed sign between them. A percentile computed
          across the two is comparing two different readings, not two documents.
          """
          if suffix == ".pdf":
              return "pdf-as-printed"
          if suffix in {".tex", ".md"}:
              return "markup-stripped"
          return "plain-text"
      
      
      def collect(target: Path) -> List[Tuple[str, str, str]]:
          """(display name, prose, suffix). The suffix travels with the text because
          the pipeline that produced it is part of what the number means."""
          if target.is_file():
              t = load(target)
              return [(target.name, t, target.suffix.lower())] if t else []
          out = []
          for p in sorted(target.rglob("*")):
              if not p.is_file():
                  continue
              if p.suffix.lower() in TEXT_SUFFIXES or p.suffix.lower() == ".pdf":
                  t = load(p)
                  if t and len(t.split()) >= 400:
                      out.append((str(p.relative_to(target)), t, p.suffix.lower()))
          return out
      
      
      # --- statistics --------------------------------------------------------------
      def word_positions(text: str, pattern: str) -> List[int]:
          """Occurrence positions measured in words, so gaps are rate-comparable."""
          words = text.split()
          offsets, cursor = [], 0
          for w in words:
              cursor = text.find(w, cursor)
              offsets.append(cursor)
              cursor += len(w)
          positions, j = [], 0
          for m in re.finditer(pattern, text):
              while j + 1 < len(offsets) and offsets[j + 1] <= m.start():
                  j += 1
              positions.append(j)
          return positions
      
      
      def cv(values: Sequence[float]) -> Optional[float]:
          if len(values) < 3:
              return None
          mean = statistics.mean(values)
          return statistics.pstdev(values) / mean if mean else None
      
      
      def sentence_lengths(text: str) -> List[Optional[int]]:
          """Word counts in reading order, with None where a span is not a sentence.
      
          The positions matter, not just the values: lag-1 must not treat two spans
          as adjacent when something was dropped between them, so the gaps are kept.
      
          A sentence may not begin with "(" here. It is tempting to allow it, because
          a LaTeX source really does open sentences with a numbered label, but in a
          PDF-derived baseline the same rule splits every inline author-year citation,
          so a fragment such as "(2019); Smith et al." is counted as a sentence. It
          does that in proportion to how much author-year citation each paper happens
          to use -- several percent for some, nothing at all for numerically cited
          ones -- which corrupts the baseline unevenly, while costing a clean source
          well under one percent of its sentences.
      
          The 120-word ceiling exists for PDF extraction runs, not for prose. Raising
          it from 90 matters: at 90 this metric flips sign depending on whether the
          gaps above are honoured, and at 120 it does not."""
          out: List[Optional[int]] = []
          for s in re.split(r"(?<=[.!?])\s+(?=[A-Z])", text):
              n = len(s.split())
              out.append(n if 4 <= n <= 120 else None)
          return out
      
      
      def linking_opener_share(text: str) -> Optional[float]:
          """Share of sentences that open with a linking adverbial (LINKING_OPENER). Sentences are split and kept as in
          sentence_lengths, so the denominator is the same set of sentences the length metrics read."""
          kept = [s for s in re.split(r"(?<=[.!?])\s+(?=[A-Z])", text) if 4 <= len(s.split()) <= 120]
          if len(kept) < 30:
              return None
          return sum(1 for s in kept if LINKING_OPENER.match(s.strip())) / len(kept)
      
      
      def lag1(values: Sequence[Optional[int]]) -> Optional[float]:
          """Positive: long sentences cluster with long ones, as in human drafts.
          Near zero: each length drawn independently.
          Negative: systematic long-short alternation, a manufactured rhythm.
      
          Only pairs that were genuinely adjacent contribute. Filtering first and
          correlating afterwards silently joins the sentences on either side of a
          dropped span, which is how the same manuscript reads inside the published
          range under one implementation and outside it under another."""
          kept = [v for v in values if v is not None]
          if len(kept) < 30:
              return None
          mean = statistics.mean(kept)
          den = sum((v - mean) ** 2 for v in kept)
          if not den:
              return None
          num = sum((values[i] - mean) * (values[i + 1] - mean)
                    for i in range(len(values) - 1)
                    if values[i] is not None and values[i + 1] is not None)
          return num / den
      
      
      def repeat_ngram_rate(text: str, n: int = 4) -> float:
          words = re.findall(r"[a-z']+", text.lower())
          if len(words) <= n:
              return 0.0
          grams: Dict[str, int] = {}
          for i in range(len(words) - n + 1):
              g = " ".join(words[i:i + n])
              grams[g] = grams.get(g, 0) + 1
          repeated = sum(v - 1 for v in grams.values() if v > 1)
          return 1000.0 * repeated / (len(words) - n + 1)
      
      
      def ngram_hashes(text: str, n: int = 4) -> set:
          """Hashed word n-grams, for asking how much of one text sits inside another.
      
          Hashes rather than the tuples: a 131-document baseline carries a few million
          n-grams, and the only thing asked of them here is set intersection.
          """
          words = re.findall(r"[a-z']+", text.lower())
          return {hash(tuple(words[i:i + n])) for i in range(len(words) - n + 1)}
      
      
      # A share alone is not enough. A document whose prose is highly repetitive has
      # few distinct n-grams, so a single shared stock phrase can be a large fraction
      # of them -- a synthetic fixture of one sentence repeated 150 times shares 8% of
      # its 4-grams with any text that opens the same way. A real earlier draft shares
      # thousands of distinct n-grams with its successor. The absolute count is what
      # separates those two, so both conditions must hold.
      MIN_SHARED_NGRAMS = 100
      
      
      def containment(target: set, other: set) -> tuple:
          """Share of the target's n-grams also in the other document, and how many."""
          if not target:
              return 0.0, 0
          shared = len(target & other)
          return shared / len(target), shared
      
      
      def measure(text: str) -> Dict[str, Optional[float]]:
          total = len(text.split())
          out: Dict[str, Optional[float]] = {"words": total}
          for name, pat in PATTERNS.items():
              pos = word_positions(text, pat)
              out[name + "_per_1k"] = 1000.0 * len(pos) / total if total else 0.0
              if name in DISTRIBUTED:
                  gaps = [b - a for a, b in zip(pos, pos[1:]) if b > a]
                  out[name + "_gap_cv"] = cv(gaps)
                  out[name + "_longest_dry_fraction"] = (max(gaps) / total) if gaps else None
          lengths = sentence_lengths(text)
          out["sentence_length_cv"] = cv([n for n in lengths if n is not None])
          out["sentence_length_lag1"] = lag1(lengths)
          out["linking_opener_share"] = linking_opener_share(text)
          out["repeat_4gram_per_1k"] = repeat_ngram_rate(text)
          verbs = re.findall(r"\b[Ww]e ([a-z]+)\b", text)
          out["first_person_verb_diversity"] = (len(set(verbs)) / len(verbs)) if verbs else None
          return out
      
      
      def sections(text: str, suffix: str, raw: str) -> Dict[str, str]:
          """Per-section rates are the uniformity signal that needs no baseline:
          when a rhetorical device runs at the same rate in the dataset section and
          the discussion, that evenness is itself the finding."""
          if suffix == ".tex":
              parts = re.split(r"\\section\*?\{([^}]*)\}", raw)
          elif suffix == ".md":
              parts = re.split(r"(?m)^#{1,2}\s+(.+)$", raw)
          else:
              return {}
          if len(parts) < 3:
              return {}
          out = {}
          for i in range(1, len(parts) - 1, 2):
              body = strip_markup(parts[i + 1], suffix)
              if len(body.split()) >= 200:
                  out[parts[i].strip()[:60]] = body
          return out
      
      
      # The range alone calls a value at the baseline's first or last paper "inside". A draft can sit there on several
      # metrics at once, below all but one or two published papers, and read as clean. The band is the middle 90% of the
      # baseline (the 5th and 95th percentiles, interpolated); a value inside the range but outside the band is at the edge:
      # reported beside the outliers, never counted as one, and it does not change the exit code.
      EDGE_BAND = (5, 95)
      
      
      def band(population: List[float]) -> Tuple[float, float]:
          """The 5th and 95th percentiles of a sorted population, by linear interpolation between its values."""
          if len(population) < 2:
              return population[0], population[-1]
          cuts = statistics.quantiles(population, n=20, method="inclusive")
          return cuts[EDGE_BAND[0] // 5 - 1], cuts[EDGE_BAND[1] // 5 - 1]
      
      
      def percentile(value: Optional[float], population: List[float]) -> Optional[float]:
          if value is None or not population:
              return None
          return 100.0 * sum(1 for x in population if x < value) / len(population)
      
      
      # --- reporting ---------------------------------------------------------------
      KEY_ORDER = [
          ("contrast_per_1k", "corrective diptych  /1k"),
          ("contrast_gap_cv", "  gap CV (bunching)"),
          ("contrast_longest_dry_fraction", "  longest dry run"),
          ("explanatory_colon_per_1k", "explanatory colon  /1k"),
          ("semicolon_per_1k", "semicolon  /1k"),
          ("discourse_marker_per_1k", "discourse marker  /1k"),
          ("sentence_length_cv", "sentence length CV"),
          ("sentence_length_lag1", "sentence length lag-1"),
          ("linking_opener_share", "linking opener share"),
          ("repeat_4gram_per_1k", "repeated 4-gram  /1k"),
          ("nominalisation_per_1k", "nominalisation  /1k"),
          ("hedge_per_1k", "hedging  /1k"),
          ("first_person_verb_diversity", "we+verb diversity"),
      ]
      
      # --per-file: a directory target measured file by file. The whole-paper numbers above are unchanged; this adds where
      # each rhetorical device peaks. A section's rate is NOT compared with the baseline's range: that range is built from
      # whole-paper averages, and a paper's average is never above its highest section, so every lively discussion section
      # would read as out of range and every flat method section as fine. What needs no baseline is the shape (the global
      # prose rules): plain sections (methods, limitations) near zero, the voice in the discussion and conclusion. A peak in
      # a plain section is backwards; in related work or the dataset section it is worth a look.
      DEVICE_KEYS = [("contrast_per_1k", "corrective diptych"),
                     ("explanatory_colon_per_1k", "explanatory colon"),
                     ("semicolon_per_1k", "semicolon")]
      MIN_SECTION_WORDS = 800
      ROLES = [("abstract", r"abstract"), ("introduction", r"intro"), ("related", r"related|background|prior"),
               ("limitations", r"limit"), ("method", r"method|pipeline|approach|procedure|protocol"),
               ("dataset", r"dataset|data|corpus|benchmark|collection"),
               ("results", r"result|finding|experiment|evaluat"), ("discussion", r"discuss"), ("conclusion", r"conclu")]
      PLAIN_ROLES = {"method", "limitations"}
      LOOK_ROLES = {"related", "dataset"}
      
      
      def role_of(name: str) -> Optional[str]:
          """The section a file holds, guessed from its name only; None when the name says nothing."""
          stem = Path(name).stem.lower()
          for role, pat in ROLES:
              if re.search(pat, stem):
                  return role
          return None
      
      
      def per_file(target: Path) -> Dict[str, dict]:
          """Every .tex/.md/.txt file under a directory target, short ones included (and marked), measured alone."""
          out = {}
          for p in sorted(target.rglob("*")):
              if not p.is_file() or p.suffix.lower() not in TEXT_SUFFIXES:
                  continue
              t = load(p)
              if not t:
                  # A file with no prose after markup is removed (a stub kept so the main file need not change) is listed,
                  # not dropped: a reader cannot otherwise tell an empty file from one that was never read.
                  out[str(p.relative_to(target))] = {"words": 0, "short": True, "role": role_of(p.name),
                                                     "metrics": {k: None for k, _ in DEVICE_KEYS}}
                  continue
              m = measure(t)
              out[str(p.relative_to(target))] = {
                  "words": int(m["words"]), "short": m["words"] < MIN_SECTION_WORDS, "role": role_of(p.name),
                  "metrics": {k: m.get(k) for k, _ in DEVICE_KEYS}}
          return out
      
      
      def peaks(files: Dict[str, dict]) -> Dict[str, dict]:
          """Where each device peaks among the files long enough to judge."""
          judged = {n: f for n, f in files.items() if not f["short"]}
          out = {}
          for key, _ in DEVICE_KEYS:
              vals = {n: f["metrics"][key] for n, f in judged.items() if f["metrics"].get(key) is not None}
              if len(vals) < 2 or max(vals.values()) <= 0:
                  continue
              top = max(vals, key=vals.get)
              role = judged[top]["role"]
              out[key] = {"file": top, "value": vals[top], "role": role,
                          "verdict": ("backwards" if role in PLAIN_ROLES else "look" if role in LOOK_ROLES
                                      else "ok" if role else "unknown")}
          return out
      
      
      def main() -> int:
          ap = argparse.ArgumentParser(
              description="Audit prose fingerprint against a baseline corpus of published work.")
          ap.add_argument("--target", required=True,
                          help="manuscript file, or a directory of .md/.tex/.txt/.pdf")
          ap.add_argument("--baseline",
                          help="directory of published papers to compare against "
                               "(a project's own literature/ is the intended corpus)")
          ap.add_argument("--exclude", action="append", default=[], metavar="GLOB",
                          help="skip baseline files whose name matches (repeatable). "
                               "Use it for the authors' own papers: a baseline that "
                               "contains them is partly the thing being measured, and "
                               "it is usually they who set the extreme")
          ap.add_argument("--min-baseline", type=int, default=5,
                          help="refuse to report percentiles below this many baseline documents")
          ap.add_argument("--overlap-threshold", type=float, default=1.0, metavar="PCT",
                          help="a baseline document sharing more than this percentage of "
                               "the target's 4-grams is reported as a suspected draft or "
                               "copy of the target itself (default: 1.0)")
          ap.add_argument("--allow-overlap", action="store_true",
                          help="report percentiles even when that scan finds one. Use only "
                               "when the overlap is intended and you can say why")
          ap.add_argument("--json", action="store_true", dest="emit_json")
          ap.add_argument("--per-file", action="store_true", dest="per_file",
                          help="with a directory target, also measure each file alone and report where each "
                               "rhetorical device peaks (see DEVICE_KEYS)")
          args = ap.parse_args()
      
          target = Path(args.target)
          if not target.exists():
              sys.stderr.write("error: --target does not exist\n")
              return 2
          docs = collect(target)
          if not docs:
              sys.stderr.write("error: no readable prose found in --target\n")
              return 2
          text = " ".join(t for _, t, _ in docs)
          target_pipelines = sorted({pipeline_of(sfx) for _, _, sfx in docs})
          mine = measure(text)
      
          base_rows: List[Dict[str, Optional[float]]] = []
          skipped: List[str] = []
          excluded: List[str] = []
          suspect: List[Dict[str, object]] = []
          target_grams = ngram_hashes(text)
          too_short: List[Dict] = []
          garbled: List[Dict] = []
          base_pipelines: Dict[str, int] = {}
          if args.baseline:
              bdir = Path(args.baseline)
              if not bdir.is_dir():
                  sys.stderr.write("error: --baseline is not a directory\n")
                  return 2
              for p in sorted(bdir.rglob("*")):
                  if not p.is_file() or p.suffix.lower() not in TEXT_SUFFIXES | {".pdf"}:
                      continue
                  if any(fnmatch.fnmatch(p.name, g) for g in args.exclude):
                      excluded.append(p.name)
                      continue
                  t = load(p)
                  if t is None:
                      skipped.append(p.name)
                      continue
                  words = len(t.split())
                  share = word_share(t)
                  if share < MIN_WORD_SHARE:
                      # A PDF whose fonts extract as symbols still yields thousands of
                      # "words". Counted as a baseline document, one such file set the
                      # lower bound of three ranges while the structure audit, which
                      # needs sentences, skipped it.
                      garbled.append({"file": p.name, "word_share": round(share, 3)})
                      continue
                  if words < 1500:
                      # A document also leaves the baseline by being too short to
                      # measure, and that exit had no name. `baseline_skipped` held
                      # only the unreadable ones, so a corpus of 179 files reported
                      # 129 documents and nothing said where the rest went. The same
                      # silence cost a test run its diagnosis: a fixture at 1470
                      # words dropped out and the failure looked unrelated.
                      too_short.append({"file": p.name, "words": words})
                      continue
                  base_rows.append(measure(t))
                  base_pipelines[pipeline_of(p.suffix.lower())] = (
                      base_pipelines.get(pipeline_of(p.suffix.lower()), 0) + 1)
                  frac, shared = containment(target_grams, ngram_hashes(t))
                  share = 100.0 * frac
                  if share > args.overlap_threshold and shared >= MIN_SHARED_NGRAMS:
                      suspect.append({"file": p.name,
                                      "target_4gram_share_pct": round(share, 3),
                                      "shared_4grams": shared})
      
          # The method's precondition is that the baseline holds none of the author's
          # own work. It used to live in prose, and `baseline_excluded: []` meant both
          # "checked, nothing to exclude" and "never checked". The scan above tells
          # them apart, and a percentile is an assertion about a baseline, so an
          # unverified baseline yields no percentile.
          # Two different questions, and the first draft of this conflated them:
          # whether the precondition HOLDS, and whether percentiles are withheld.
          # --allow-overlap answers the second, never the first. Reporting
          # preconditions_checked: true because someone passed a waiver would rebuild
          # the ambiguous green light this scan exists to remove.
          preconditions_ok = not suspect
          waived = bool(suspect) and args.allow_overlap
          contaminated = bool(suspect) and not args.allow_overlap
          enough = len(base_rows) >= args.min_baseline and not contaminated
          report = {"schema_version": SCHEMA_VERSION, "target": str(target),
                    "target_words": mine["words"], "baseline_documents": len(base_rows),
                    "baseline_sufficient": len(base_rows) >= args.min_baseline,
                    "baseline_skipped": skipped,
                    "baseline_too_short": too_short,
                    "baseline_garbled": garbled,
                    "target_pipeline": target_pipelines,
                    "baseline_pipeline_mix": base_pipelines,
                    "pipeline_mismatch": bool(base_pipelines)
                                         and set(target_pipelines) != set(base_pipelines),
                    "exclude_patterns": list(args.exclude),
                    "baseline_excluded": excluded,
                    "baseline_suspect": suspect,
                    "overlap_threshold_pct": args.overlap_threshold,
                    "preconditions_checked": preconditions_ok,
                    "overlap_waived": waived,
                    "metrics": {}}
      
          outliers, edges = [], []
          for key, _ in KEY_ORDER:
              pop = [r[key] for r in base_rows if r.get(key) is not None]
              entry = {"value": mine.get(key)}
              if enough and pop and mine.get(key) is not None:
                  pop_sorted = sorted(pop)
                  entry.update({
                      "baseline_median": statistics.median(pop_sorted),
                      "baseline_min": pop_sorted[0],
                      "baseline_max": pop_sorted[-1],
                      "percentile": percentile(mine[key], pop_sorted),
                      "outside_range": not (pop_sorted[0] <= mine[key] <= pop_sorted[-1]),
                  })
                  low, high = band(pop_sorted)
                  entry.update({"band_low": low, "band_high": high,
                                "edge": not entry["outside_range"] and not (low <= mine[key] <= high)})
                  if entry["outside_range"]:
                      outliers.append(key)
                  elif entry["edge"]:
                      edges.append(key)
              report["metrics"][key] = entry
      
          # Per-section evenness needs no baseline at all, which makes it the one
          # measurement here that is never blocked -- and the one whose absence used
          # to look like nothing. The key was simply missing when the target was a
          # directory, so a run that never computed it and a run with nothing to say
          # produced the same report. It now always appears, with its reason.
          if not (target.is_file() and target.suffix.lower() in {".tex", ".md"}):
              report["per_section_cv"] = None
              report["per_section_note"] = (
                  "NOT COMPUTED: per-section rates need a single .tex or .md target and "
                  "--target is {}. This metric needs no baseline and is the one that shows "
                  "whether the device runs evenly across sections, so its absence is a hole "
                  "in the reading, not a clean result.".format(
                      "a directory" if target.is_dir() else "a " + (target.suffix or "file")))
          else:
              raw = target.read_text(encoding="utf-8", errors="replace")
              secs = sections(text, target.suffix.lower(), raw)
              if len(secs) < 3:
                  report["per_section_cv"] = None
                  report["per_section_note"] = (
                      "NOT COMPUTED: {} section(s) were found in {}; three are needed for a "
                      "cross-section spread.".format(len(secs), target.name))
              else:
                  rates = {}
                  for name, body in secs.items():
                      w = len(body.split())
                      rates[name] = 1000.0 * len(re.findall(CONTRAST, body)) / w
                  spread = cv(list(rates.values()))
                  report["per_section_contrast_per_1k"] = rates
                  report["per_section_cv"] = spread
                  report["per_section_note"] = (
                      "A low cross-section CV means the device runs at the same rate in the "
                      "dutiful sections as in the discussion. That evenness is the signature; "
                      "the fix is to redistribute, not merely to reduce.")
          if args.per_file and target.is_dir():
              files = per_file(target)
              judged = [f for f in files.values() if not f["short"]]
              report["per_file"] = files
              report["peaks"] = peaks(files)
              if len(judged) >= 3:
                  report["per_section_cv"] = {k: cv([f["metrics"][k] for f in judged if f["metrics"].get(k) is not None])
                                              for k, _ in DEVICE_KEYS}
                  report["per_section_note"] = None
              else:
                  report["per_section_cv"] = None
                  report["per_section_note"] = (
                      "NOT COMPUTED: {} file(s) have at least {} words; three are needed for a "
                      "cross-section spread.".format(len(judged), MIN_SECTION_WORDS))
          # When the target is read one way and the baseline another, every
          # percentile above compares two readings, not two documents. The tool
          # cannot remove the asymmetry -- a published PDF has no source to strip --
          # but it does not have to warn about it in prose either. When the target's
          # own build sits beside it, the same metrics are measured through the
          # baseline's pipeline as well, so the cost of the mismatch is a number in
          # this report rather than a caution the reader is asked to remember.
          report["pipeline_cross_check"] = None
          if report["pipeline_mismatch"]:
              note = ("target and baseline are read through different pipelines "
                      "({} vs {}); percentiles compare two readings of a document, not "
                      "two documents".format(", ".join(target_pipelines),
                                             ", ".join(sorted(base_pipelines))))
              twin = target.with_suffix(".pdf") if target.is_file() else None
              if twin is not None and twin != target and twin.exists():
                  twin_text = load(twin)
                  if twin_text:
                      twin_m = measure(twin_text)
                      report["pipeline_cross_check"] = {
                          "as_read": target.name,
                          "as_baseline_would_read": twin.name,
                          "metrics": {k: {"as_read": mine.get(k),
                                          "as_baseline_would_read": twin_m.get(k)}
                                      for k, _ in KEY_ORDER},
                          "words": {"as_read": mine["words"],
                                    "as_baseline_would_read": twin_m["words"]},
                      }
                      note += ("; measured on this document, the two readings are in "
                               "pipeline_cross_check")
              else:
                  note += ("; no built .pdf was found next to --target, so the size of "
                           "the difference on THIS document is not measured here")
              report["pipeline_note"] = note
      
          report["outliers"] = outliers
          report["edge"] = edges
      
          if contaminated:
              sys.stderr.write(
                  "error: the baseline contains {} document(s) that share more than "
                  "{:.2f}% of the target's 4-grams, which is what a draft or a copy of "
                  "the target looks like:\n".format(len(suspect), args.overlap_threshold))
              for s in suspect:
                  sys.stderr.write("  {}  {}%\n".format(s["file"], s["target_4gram_share_pct"]))
              sys.stderr.write(
                  "  Percentiles are withheld. Exclude them with --exclude, or pass "
                  "--allow-overlap if the overlap is intended.\n")
      
          if args.emit_json:
              print(json.dumps(report, indent=2))
          else:
              print("target: {}  ({:,} words)".format(target, int(mine["words"])))
              if args.baseline:
                  print("baseline: {} documents{}".format(
                      len(base_rows), "" if enough else "  [too few for percentiles]"))
                  if skipped:
                      print("  unreadable (install pdftotext?): {}".format(", ".join(skipped[:6])))
              print()
              head = "{:<26}{:>10}".format("metric", "target")
              if enough:
                  head += "{:>10}{:>18}{:>8}".format("median", "range", "pct")
              if suspect:
                  print("  suspected drafts/copies of the target in the baseline: {}".format(
                      ", ".join("{} ({}%)".format(s["file"], s["target_4gram_share_pct"])
                                for s in suspect)))
                  print()
              print(head)
              for key, label in KEY_ORDER:
                  e = report["metrics"][key]
                  if e["value"] is None:
                      continue
                  line = "{:<26}{:>10.3f}".format(label, e["value"])
                  if "baseline_median" in e:
                      line += "{:>10.3f}{:>18}{:>8.0f}".format(
                          e["baseline_median"],
                          "{:.2f}-{:.2f}".format(e["baseline_min"], e["baseline_max"]),
                          e["percentile"])
                      if e["outside_range"]:
                          line += "  *"
                      elif e.get("edge"):
                          line += "  ~"
                  print(line)
              if report.get("per_file"):
                  print("\nper file, /1k: " + ", ".join(label for _, label in DEVICE_KEYS) + "  (* too short to judge)")
                  for name, f in report["per_file"].items():
                      vals = "".join("{:>9.2f}".format(f["metrics"][k] or 0.0) for k, _ in DEVICE_KEYS)
                      print("  {:<44}{}{}".format(name[:44], vals, "  *" if f["short"] else ""))
                  for key, pk in report.get("peaks", {}).items():
                      if pk["verdict"] in ("backwards", "look"):
                          print("  {} peaks in {} ({}): {}".format(
                              key, pk["file"], pk["role"],
                              "backwards, this section should be plain" if pk["verdict"] == "backwards" else "worth a look"))
              if isinstance(report.get("per_section_cv"), float):
                  print("\nper-section corrective-diptych rate  (cross-section CV {:.2f})".format(
                      report["per_section_cv"]))
                  for name, r in report["per_section_contrast_per_1k"].items():
                      print("  {:<44}{:>7.1f}".format(name, r))
              if outliers:
                  print("\noutside the baseline range: {}".format(", ".join(outliers)))
              if edges:
                  print("at the edge (inside the range, outside the baseline's {}th-{}th percentile band): {}".format(
                      EDGE_BAND[0], EDGE_BAND[1], ", ".join(edges)))
              print("\nNote: none of these values is a target to hit. Editing to move a "
                    "number\nrather than to fix a sentence produces a different artefact, "
                    "not a better one.")
      
          # 2, not 1: 1 means "measured, and something is outside the range"; 2 means
          # "not measured, because the baseline could not be trusted". Without the
          # distinction a contaminated run returns 0, since withheld percentiles
          # produce no outliers -- a failed precondition reading as success.
          if contaminated:
              return 2
          return 1 if outliers else 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • audit-prose-structure.py 10.9 KB
      #!/usr/bin/env python3
      """Audit sentence structure against a baseline corpus: what makes prose hard to read that punctuation counts miss.
      
          python3 audit-prose-structure.py --target <file or dir> --baseline <dir> [--exclude GLOB ...] [--json]
      
      audit-prose-fingerprint.py counts marks. This counts structure. On a real manuscript every mark-based rate sat in the
      published range while readers still called the prose hard: the sentences were shorter than most published papers',
      and carried more subordinate clauses per comma than any of them. A sentence that makes the reader hold a condition
      before the claim, or a clause between subject and verb, is hard at any length.
      
      Measured per document:
        median_len       median sentence length in words
        long_share       share of words in sentences of 35 words or more
        sub_per_100w     subordinate clauses per 100 words
        sub_per_comma    subordinate clauses per comma: the same structure with fewer places to breathe
        comma_per_100w   commas per 100 words
        dash_per_100w    em dashes per 100 words
        opens_with_sub   share of sentences that open with a subordinator (Where / When / If / Whether / Although ...)
        length_lag1      lag-1 autocorrelation of sentence length (published prose is positive: long sentences cluster)
      
      "that" is not counted: it is a determiner, a complementiser and a relativiser, and telling them apart needs a
      parser; counting it inflates whichever side one set out to find. Every subordinator below is unambiguous.
      
      Text is read by audit-prose-fingerprint.py's own `load` (markup stripped from .tex/.md, pdftotext for PDFs), so the
      two audits read the same prose. A target that is a directory is read whole, every file joined; each baseline
      document is measured on its own, never pooled, and is kept only with 25 sentences or more.
      
      Output: each metric with the baseline's min / median / max and the target's percentile; `outliers` names the ones
      outside the baseline's range. These are measurements, not targets: edit a sentence because it is hard to read,
      not to move a number.
      
      Exit: 0 inside the range on every metric; 1 at least one outside; 2 nothing to measure (no target prose, fewer
      baseline documents than --min-baseline, the fingerprint reader missing) or an argument it does not recognise.
      """
      import argparse
      import fnmatch
      import importlib.util
      import json
      import re
      import statistics
      import sys
      from pathlib import Path
      
      HERE = Path(__file__).resolve().parent
      SUB = re.compile(r"\b(?:which|where|when|while|whereas|although|though|because|since|unless|whether|if|after|"
                       r"before|until|who|whom|whose)\b", re.I)
      OPENER = re.compile(r"^(?:Where|When|If|Whether|Although|Though|While|Because|Since|Unless|Once|Whereas)\b")
      DASH = re.compile(r"---|—")
      SPLIT = re.compile(r"(?<=[.!?])\s+(?=[A-Z])")
      MIN_SENTENCES = 25
      KEYS = ["median_len", "long_share", "sub_per_100w", "sub_per_comma", "comma_per_100w", "dash_per_100w",
              "opens_with_sub", "length_lag1"]
      
      
      def die(msg):
          sys.stderr.write(f"audit-prose-structure: {msg}\n")
          sys.exit(2)
      
      
      def reader():
          path = HERE / "audit-prose-fingerprint.py"
          if not path.is_file():
              die(f"the fingerprint audit is not beside this script ({path}); both must read prose the same way")
          spec = importlib.util.spec_from_file_location("fingerprint", path)
          mod = importlib.util.module_from_spec(spec)
          spec.loader.exec_module(mod)
          return mod
      
      
      def spans(text):
          return [s.strip() for s in SPLIT.split(text)]
      
      
      def lag1(lengths):
          """Only adjacent sentences both kept count: a dropped span breaks the chain rather than joining its neighbours."""
          kept = [v for v in lengths if v is not None]
          if len(kept) < 30:
              return None
          mean = statistics.mean(kept)
          den = sum((v - mean) ** 2 for v in kept)
          if not den:
              return None
          num = sum((lengths[i] - mean) * (lengths[i + 1] - mean) for i in range(len(lengths) - 1)
                    if lengths[i] is not None and lengths[i + 1] is not None)
          return num / den
      
      
      def measure(text):
          raw = spans(text)
          lengths = [(len(s.split()) if 4 <= len(s.split()) <= 120 else None) for s in raw]
          sents = [s for s, n in zip(raw, lengths) if n is not None]
          if len(sents) < MIN_SENTENCES:
              return None
          words = sum(len(s.split()) for s in sents)
          sub = sum(len(SUB.findall(s)) for s in sents)
          commas = sum(s.count(",") for s in sents)
          return {"sentences": len(sents), "words": words,
                  "median_len": statistics.median(len(s.split()) for s in sents),
                  "long_share": sum(len(s.split()) for s in sents if len(s.split()) >= 35) / words,
                  "sub_per_100w": 100 * sub / words,
                  "sub_per_comma": sub / max(1, commas),
                  "comma_per_100w": 100 * commas / words,
                  "dash_per_100w": 100 * sum(len(DASH.findall(s)) for s in sents) / words,
                  "opens_with_sub": sum(1 for s in sents if OPENER.match(s)) / len(sents),
                  "length_lag1": lag1(lengths)}
      
      
      def read_target(load, target):
          """(whole text, {file: text}). A pattern concentrated in one section averages away over the whole paper, so the
          per-file figures are reported beside it; they are descriptive, and never outliers on their own."""
          files = [target] if target.is_file() else sorted(p for p in target.rglob("*") if p.is_file())
          parts = {}
          for p in files:
              t = load(p)
              # None: not a text file this reader handles. "" : a text file with no prose (a stub kept so the main file
              # need not change), kept so that per_file lists it instead of losing it.
              if t is not None:
                  parts[str(p.relative_to(target)) if target.is_dir() else p.name] = t
          return " ".join(parts.values()), parts
      
      
      def main(argv=None):
          ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
          ap.add_argument("--target", required=True)
          ap.add_argument("--baseline", required=True)
          ap.add_argument("--exclude", action="append", default=[], help="skip baseline files whose name matches")
          ap.add_argument("--min-baseline", type=int, default=20)
          ap.add_argument("--json", action="store_true")
          try:
              a = ap.parse_args(argv)
          except SystemExit as e:
              sys.exit(2 if e.code else 0)
          fp = reader()
          load = fp.load
          target, base = Path(a.target), Path(a.baseline)
          if not target.exists():
              die(f"target not found: {target}")
          if not base.is_dir():
              die(f"baseline is not a directory: {base}")
          whole, parts = read_target(load, target)
          tm = measure(whole)
          if tm is None:
              die(f"fewer than {MIN_SENTENCES} measurable sentences in the target: nothing measured is not a pass")
          rows, skipped, excluded = [], [], []
          for p in sorted(base.rglob("*")):
              if not p.is_file() or p.suffix.lower() not in {".pdf", ".tex", ".md", ".txt"}:
                  continue
              if any(fnmatch.fnmatch(p.name, g) for g in a.exclude):
                  excluded.append(p.name)
                  continue
              m = measure(load(p) or "")
              (rows.append({"doc": p.name, **m}) if m else skipped.append(p.name))
          if len(rows) < a.min_baseline:
              die(f"{len(rows)} baseline document(s) measurable (skipped {len(skipped)}, excluded {len(excluded)}); "
                  f"no percentile below {a.min_baseline}")
          metrics, outliers, edges = {}, [], []
          for k in KEYS:
              vals = sorted(r[k] for r in rows if r[k] is not None)
              v = tm[k]
              if v is None or not vals:
                  metrics[k] = {"value": v, "note": "not measurable"}
                  continue
              pct = round(100 * sum(x < v for x in vals) / len(vals))
              out = v < vals[0] or v > vals[-1]
              low, high = fp.band(vals)
              edge = not out and not (low <= v <= high)
              metrics[k] = {"value": v, "min": vals[0], "median": statistics.median(vals), "max": vals[-1],
                            "percentile": pct, "outside": out, "band_low": low, "band_high": high, "edge": edge}
              if out:
                  outliers.append(k)
              elif edge:
                  edges.append(k)
          per_file = {}
          for name, t in parts.items():
              m = measure(t)
              if m:
                  per_file[name] = {"short": False, **{k: m[k] for k in ("sentences", "median_len", "sub_per_comma",
                                                                          "opens_with_sub")}}
              else:
                  # Below the sentence floor: listed and marked, not dropped, so an unmeasured file is not mistaken for
                  # one that was never read.
                  per_file[name] = {"short": True,
                                    "sentences": sum(1 for s in spans(t) if 4 <= len(s.split()) <= 120)}
          # Where the structure is densest. Descriptive, like the rest of per_file: the baseline's range is built from
          # whole papers and a paper's value is an average of its sections, so no single section is an outlier on its own.
          judged = {n: f for n, f in per_file.items() if not f["short"] and f.get("sub_per_comma") is not None}
          densest = None
          if len(judged) >= 2:
              top = max(judged, key=lambda n: judged[n]["sub_per_comma"])
              densest = {"metric": "sub_per_comma", "file": top, "value": judged[top]["sub_per_comma"]}
          report = {"target": str(target), "target_sentences": tm["sentences"], "target_words": tm["words"],
                    "per_file": per_file, "densest": densest,
                    "baseline_documents": len(rows), "baseline_skipped": skipped, "baseline_excluded": excluded,
                    "metrics": metrics, "outliers": outliers, "edge": edges}
          if a.json:
              print(json.dumps(report, ensure_ascii=False, indent=1))
          else:
              print(f"target: {target} ({tm['sentences']} sentences, {tm['words']} words)   baseline: {len(rows)} documents")
              for k in KEYS:
                  m = metrics[k]
                  if "min" in m:
                      flag = "  *" if m["outside"] else ("  ~" if m["edge"] else "")
                      print(f"  {k:<16}{m['value']:>9.3f}   median {m['median']:.3f}   range {m['min']:.3f}–{m['max']:.3f}"
                            f"   pct {m['percentile']:>3}{flag}")
                  else:
                      print(f"  {k:<16}not measurable")
              print("outside the baseline range: " + (", ".join(outliers) or "none"))
              if edges:
                  print("at the edge (inside the range, outside the baseline's 5th-95th percentile band): " + ", ".join(edges))
              if len(per_file) > 1:
                  print("per file (descriptive): sentences / median length / clauses per comma / opens with a subordinator")
                  for name, m in per_file.items():
                      if m["short"]:
                          print(f"  {name:<40}{m['sentences']:>5}   too few sentences to measure")
                          continue
                      print(f"  {name:<40}{m['sentences']:>5}{m['median_len']:>6.0f}{m['sub_per_comma']:>8.3f}"
                            f"{m['opens_with_sub']:>8.3f}")
                  if densest:
                      print(f"  densest (clauses per comma): {densest['file']} {densest['value']:.3f} (descriptive)")
          return 1 if outliers else 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • audit-sentence-changes.py 57.5 KB
      #!/usr/bin/env python3
      """Check each rewritten sentence against the sentence it replaced, and against the venue's own sentences.
      
          python3 audit-sentence-changes.py --target <file or dir> --base <file or dir> [--baseline <dir>] [--carriers <file>] [--json]
          python3 audit-sentence-changes.py --pairs <tsv with columns id, old, new> [--baseline <dir>] [--json]
          (a pairs file may add the columns verdict and reason once the author has read the rewrites)
          (either form: --venue-cache <file> keeps the venue's measured sentences between runs)
      
      audit-prose-fingerprint.py and audit-prose-structure.py measure a whole document: rates per 1,000 words and
      distributions across sections. A dozen rewritten sentences barely move those rates, and neither audit reads a
      proposal that has not yet been applied. On one real round of proposed corrections most rewrites came back longer
      than the sentences they replaced, several added an explanatory colon or a semicolon (the construction the venue
      audit had already flagged in that manuscript), and others added a relative clause, a participial phrase or an
      adjective. Neither document audit could see it: the rewrites sat in a separate file, and once applied they were a
      few sentences in fifteen thousand words.
      
      So this reads sentences one at a time, and only the ones that changed. Every changed sentence is judged; none is
      waved through for lack of a predecessor. A sentence of the target that does not occur in the base is paired with
      the most similar base sentence that no longer occurs (a revision). If it has none, it may be a piece of a sentence
      that was split (most of its words come from one old sentence) or the result of a merge (it holds most of the words
      of several old sentences); a split is judged as the old sentence against the sum of its pieces, a merge as the sum
      of the old sentences against the new one. What is left is an addition, judged against the venue.
      
      For a revision, a split or a merge it reports what the rewrite added: words (three or more), subordinate clauses
      counted as occurrences gained (dropping one "because" and adding one "which" is still an added clause), adverbs and
      modifiers counted net (swapping one term for another adds none), commas (when punctuation as a whole grew), colons, semicolons,
      dashes, parentheses, subordinate clauses, adverbs, modifiers, prepositional phrases (two or more), a subordinate
      opener. A merge is flagged as such: it is how short sentences become long ones. An addition is flagged for any
      colon, semicolon, dash or worded parenthesis, and for length or density above the venue's 75th percentile of
      sentences (pooled across its documents: one sentence is compared with published sentences, not with documents); a
      revision that grows past the 90th is flagged for that too. Without --baseline additions are held to default ceilings
      measured on one journal, and the report says so.
      
      Adverbs are words ending in -ly and a closed list of others (also, still, only, even, very, often, instead ...);
      connectives such as however, thus and consequently are not counted. Modifiers stand in for a part-of-speech
      tagger, which this stdlib script does not have: participles after a noun ("the images collected from", "applied to
      characters"), a participle or adjective-suffixed word after a determiner ("a controlled test"), a comma followed by
      a participle, words with adjective suffixes (-al, -ive, -ous, -ic, -able, -ful, -less, -ary; a noun stoplist
      applies), a closed list of common adjectives, and hyphenated prenominal compounds. It misses adjectives outside
      those forms and relative clauses opened by "that", which is not counted as a clause for the reason
      audit-prose-structure.py gives. Because gains are counted, a noun such as "manual" that the suffix rule mistakes
      for an adjective only matters when the rewrite introduces it.
      
      Before splitting, LaTeX list items and figure or table captions are kept as prose (the fingerprint audit drops
      them), blank lines and Markdown headings and list items end a sentence, reference and citation commands and inline comments are removed, inline math becomes one placeholder word,
      footnotes become parentheses, and headings are dropped. A pairs file is read with no quoting: a stray quotation
      mark stays text.
      
      The --pairs form is for proposals: run it on the rewrites before anyone reads them. The --target/--base form is for
      a draft after the edit; the writing loop runs it against the last version at which it flagged nothing.
      
      The thresholds (three words, two prepositions, the 75th and 90th percentiles) were set against one real round: a
      handful of rewrites an author rejected for adding length, clauses, modifiers and punctuation, and a few dozen changes
      the author accepted. That is the data they were fitted on, not a test of them. The test is a later round: give the
      pairs file a verdict column (accepted or rejected, as the author decided; revised counts as rejected, empty as not
      yet judged) and a reason column, and the report sets the flags against the verdicts, naming the script by its hash
      so that the thresholds that judged are the ones frozen before the round. The flags are prompts to re-read a sentence, not targets: a revision
      may need a clause to stay faithful to its source, and then the flag is the reason to check that it earns it.
      
      A sentence removed without a successor (nothing in the target revises, splits or merges it) is judged too. Every check
      of the draft rewards deletion: word and rate ceilings fall, a sentence nobody wrote cannot be flagged. On one
      manuscript a research question, three qualifiers and a denominator were deleted and every check passed. A removal is
      flagged when it carried something: it matches a pattern in --carriers (one regular expression per line; the writing
      loop passes the claims ledger's required wordings), or it holds a number, or a limiting qualifier (only, at most, may,
      exploratory, of N, in this study ...). A flagged removal needs a reason like a flagged rewrite. The report counts
      removals and flagged removals apart, so "nothing flagged" is not read as "nothing removed".
      
      A removal is also flagged when it took an antecedent: a later sentence of the same paragraph, kept or revised, still
      says "the X" (or this, these, those; at most one word between) where X is a noun of the removed sentence and no earlier
      sentence of the file names X. On one manuscript an abstract lost the only sentence naming its subject, and its last
      sentence, ten sentences on, still spoke of it. The whole paragraph is read because a nearer window missed that case; the
      whole file triples the hits and catches nothing more. Ordinals, modals and a word followed by an article (a verb) are
      not taken for nouns. A changed or added sentence that gains a share ("one of three later sensors") is flagged when the
      draft states another share of the same total of the same noun ("two of the three later sensors"); the other sentences
      are listed beside it. A changed, added or copied sentence of six words or more that matches another sentence of the
      draft on its letters and digits (a cross-reference or a parenthesis the reader cannot see is not a difference) is
      flagged as duplicating it, with the other place named; a verbatim copy is found by count, since pairing takes a
      sentence the base already had as unchanged, and a move keeps the count. These need the whole draft, so the --pairs
      form does not run them. The report counts each apart.
      
      Exit: 0 no changed sentence is flagged (including no change at all, reported as such); 1 at least one flagged;
      2 nothing to compare (no prose in the target or the base, an empty or malformed pairs file, a baseline too small
      to give percentiles) or an argument it does not recognise.
      """
      import argparse
      import csv
      import difflib
      import importlib.util
      import json
      import re
      import sys
      from collections import Counter
      from pathlib import Path
      
      HERE = Path(__file__).resolve().parent
      
      SUB_WORDS = {"which", "where", "when", "while", "whereas", "although", "though", "because", "since", "unless",
                   "whether", "if", "after", "before", "until", "who", "whom", "whose", "once", "whereby", "wherein"}
      SUB = re.compile(r"\b(" + "|".join(sorted(SUB_WORDS)) + r")\b", re.I)
      # after / before / since / until followed by a number or an -ing word are prepositions ("since 2019", "after training");
      # "once a year" and "once more" are adverbs.
      TEMPORAL_PREP = re.compile(r"\b(?:after|before|since|until)\s+(?:\d|\w+ing\b)|\bonce\s+(?:a|an|per|every|more|again|or|and)\b"
                                 r"|\bonce[.,;]", re.I)
      AS_CLAUSE = re.compile(r"\bas\s+(?:(?:the|a|an|this|these|those|it|they|we|he|she|its|their|our)\s+)?\w+\s+"
                             r"(?:is|are|was|were|has|have|had|did|does|do|can|could|will|would|may|might|must|should|"
                             r"required|requires|expected|noted|shown|described|reported|suggested)\b", re.I)
      SO_THAT = re.compile(r"\b(?:so|such)\s+that\b", re.I)
      OPENER = re.compile(r"^(?:Where|When|If|Whether|Although|Though|While|Because|Since|Unless|Once|Whereas|As)\b")
      PREP = re.compile(r"\b(?:of|in|for|with|by|between|from|on|at|into|across|against|within|without|under|over|"
                        r"through|about|among|beyond|via)\b", re.I)
      LY = re.compile(r"\b([A-Za-z]{3,}ly)\b")
      NOT_ADVERB = {"early", "family", "apply", "supply", "reply", "rely", "assembly", "anomaly", "italy", "july",
                    "daily", "weekly", "monthly", "yearly", "ply", "multiply", "comply", "imply", "poly", "holy", "belly",
                    "rally", "tally", "ally", "fly", "jelly", "bully", "lily", "folly",
                    # adjectives more often than adverbs in this register ("is likely to"); counted as modifiers instead
                    "likely", "unlikely", "costly", "timely", "friendly", "scholarly", "orderly",
                    # connectives: they order an argument; the prose audits count them separately
                    "consequently", "accordingly", "similarly", "conversely", "finally", "firstly", "secondly", "lastly",
                    "thirdly"}
      OTHER_ADVERBS = {"also", "still", "only", "even", "very", "often", "instead", "just", "rather", "quite", "already",
                       "almost", "never", "always", "merely", "perhaps", "indeed", "too", "again", "sometimes", "somewhat",
                       "seldom", "soon", "hardly", "nearly", "far", "much", "well"}
      # "far", "much" and "well" are adverbs when they modify ("far fewer", "much larger", "well known"); as other parts of
      # speech they rarely appear in a rewrite that did not have them before, and only gains are counted.
      DET = r"(?:a|an|the|this|these|those|its|their|our|his|her|each|every|no|any|some|one|two|three|four|such|" \
            r"several|all|both|many)"
      # "that" is left out: as a relative pronoun it would read the verb after it ("two that reported") as a modifier.
      BE_HAVE = {"is", "are", "was", "were", "be", "been", "being", "has", "have", "had", "get", "gets", "got", "seem",
                 "seems", "seemed", "appear", "appears", "appeared", "become", "becomes", "became", "remain", "remains",
                 "remained", "not", "it", "they", "we", "he", "she", "i", "you", "can", "could", "will", "would", "may",
                 "might", "must", "should"}
      IRREGULAR_PARTICIPLES = {"given", "taken", "written", "shown", "known", "chosen", "drawn", "seen", "grown", "thrown",
                               "spoken", "broken", "hidden", "made", "built", "found", "held", "kept", "left", "set", "sent",
                               "told", "taught", "brought", "bought", "done", "cut", "put", "spent", "lost", "paid", "laid"}
      PARTICIPLE = r"(?:\w{3,}ed|" + "|".join(sorted(IRREGULAR_PARTICIPLES)) + r")"
      POST_PREP = r"(?:by|from|in|on|to|with|for|at|into|across|under|over|through|as|during|within|between|against|via)"
      PARTICIPLE_AFTER_NOUN = re.compile(r"\b(\w+)\s+(" + PARTICIPLE + r")\s+" + POST_PREP + r"\b", re.I)
      PARTICIPLE_AFTER_COMMA = re.compile(r",\s+(\w{3,}(?:ing|ed))\b")
      PRENOMINAL = re.compile(r"\b" + DET + r"\s+(\w{3,}(?:ed|ing))\s+[a-z]", re.I)
      ADJ_SUFFIX = re.compile(r"\b([a-z]{3,}(?:al|ive|ous|ic|able|ible|ful|less|ary))\b")
      NOUN_STOP = {"manual", "manuals", "retrieval", "proposal", "signal", "journal", "material", "potential", "individual",
                   "interval", "trial", "portal", "approval", "removal", "arrival", "rival", "animal", "capital", "festival",
                   "terminal", "principal", "criminal", "survival", "renewal", "denial", "referral", "tutorial", "rival",
                   "archive", "objective", "perspective", "initiative", "alternative", "incentive", "motive", "narrative",
                   "representative", "directive", "detective", "native", "topic", "music", "logic", "public", "mechanic",
                   "graphic", "fabric", "clinic", "critic", "rhetoric", "arithmetic", "epic", "traffic", "table", "variable",
                   "cable", "summary", "library", "dictionary", "boundary", "vocabulary", "anniversary", "salary",
                   "glossary", "commentary", "diary", "itinerary", "secretary", "sanctuary", "century", "several",
                   "general", "usual", "total", "normal", "original", "local", "global", "final", "initial", "central",
                   "legal", "equal", "natural", "social", "special", "digital", "visual", "textual", "cultural",
                   "historical", "technical", "typical", "critical", "classical", "musical", "physical", "practical",
                   "logical", "chemical", "medical", "political", "statistical", "empirical", "numerical", "vertical",
                   "identical", "lexical", "theoretical", "biblical", "ethical", "optical", "radical", "tropical",
                   "heuristic", "metric", "statistic", "rubric", "characteristic", "academic", "arithmetic"}
      # The -al adjectives in the stoplist are real adjectives; they are removed because a manuscript's own subject words
      # ("historical", "cultural", "lexical") recur in every rewrite and are not modifiers the rewrite chose to add.
      COMMON_ADJ = {"new", "old", "large", "small", "big", "open", "same", "different", "other", "own", "single", "whole",
                    "full", "main", "major", "minor", "key", "clear", "strong", "weak", "good", "bad", "high", "low", "long",
                    "short", "wide", "broad", "narrow", "simple", "specific", "real", "true", "false", "various", "certain",
                    "possible", "similar", "recent", "modern", "late", "best", "better", "worse", "worst", "rich", "poor",
                    "deep", "fine", "hard", "easy", "free", "fair", "fast", "slow", "great", "entire", "overall", "prior",
                    "particular", "multiple", "unique", "common", "rare", "unknown", "robust", "reliable", "relevant",
                    "important", "significant", "substantial", "considerable", "notable", "strict", "careful", "novel",
                    "complex", "diverse", "distinct", "explicit", "implicit", "direct", "indirect", "sharp", "trained",
                    "independent", "comprehensive", "extensive", "systematic", "rigorous", "subtle", "crucial", "genuine",
                    "likely", "unlikely", "costly", "timely", "friendly", "scholarly", "orderly"}
      HYPHEN_PRENOMINAL = re.compile(r"\b([a-z]+(?:-[a-z]+)+)\s+(?!(?:and|or|but|nor|the|for|with|from|into|that|than|are|was|were|has|have)\b)[a-z]{3,}")
      DASH = re.compile(r"---|—|―|\s--\s|\s–\s|\s-\s")
      PAREN_WORDED = re.compile(r"\(([^()]*)\)")
      ABBREV = re.compile(r"\b(e\.g|i\.e|et al|cf|vs|Fig|Figs|Eq|Eqs|Sec|Secs|No|approx|resp|ca)\.", re.I)
      SPLIT = re.compile(r"(?<=[.!?])\s+(?=[A-Z(])")
      # pre_tex puts this where LaTeX ends a unit without a full stop (a list item, a caption, a heading); the fingerprint's
      # markup stripping collapses the line breaks that would otherwise have marked it.
      BREAK = " \u00b6 "
      FEATURES = ["words", "clauses", "commas", "colons", "semicolons", "dashes", "parentheses", "adverbs", "modifiers",
                  "prepositions", "opener"]
      # A revision is "longer" when it gains at least this many words. In the round the thresholds were set on, the
      # rejected rewrites that grew gained well over this, and no accepted revision gained more than two.
      LONGER_WORDS = 3
      # Prepositional phrases gained before a revision is flagged: one is often the fact the correction adds.
      MORE_PREPOSITIONS = 2
      MATCH_MIN = 0.40
      # What a removed sentence can carry besides a pattern the caller names. Numbers: a count, a denominator, a result.
      # Qualifiers: the words that keep a claim no larger than its evidence. Both lists catch only what is written here.
      REMOVED_NUMBER = re.compile(r"\d")
      REMOVED_QUALIFIER = re.compile(
          r"\b(only|at most|at least|no more than|may|might|could|approximately|roughly|estimated|exploratory|"
          r"preliminary|tentative|suggests?|limited to|except|unless|in (?:this|our) (?:study|benchmark|sample|pool|setting|data)|"
          r"(?:out )?of \d+)\b|仅|只有|至多|至少|可能|大约|估计|探索性|初步|除非|在本(?:研究|基准|文)", re.I)
      PIECE_MIN = 0.60   # share of a sentence's content words found in another before one is read as part of the other
      VENUE_PCT = 90      # a revision that grows past this is long or dense for the venue
      ADDITION_PCT = 75   # an added sentence has nothing to be compared with, so it is held to a typical published one
      FP = None
      MIN_VENUE_SENTENCES = 1000
      MIN_VENUE_DOCUMENTS = 5
      STOPWORDS = {"the", "a", "an", "of", "in", "for", "with", "by", "and", "or", "to", "is", "are", "was", "were", "be",
                   "that", "this", "it", "its", "on", "at", "as", "from", "not", "we", "our", "their", "they", "than"}
      LIMITS = ("not measured: adjectives outside the suffix and word lists, nouns used as modifiers, relative clauses "
                "opened by 'that', a comma splice beyond the comma it adds, a pronoun (it, they) a removal left without its "
                "antecedent, a noun a revision dropped, a repeat reworded rather than copied; the writing loop reads committed "
                "versions only")
      
      
      def die(msg):
          sys.stderr.write(f"audit-sentence-changes: {msg}\n")
          sys.exit(2)
      
      
      def fingerprint():
          path = HERE / "audit-prose-fingerprint.py"
          if not path.is_file():
              die(f"the fingerprint audit is not beside this script ({path}); both must read prose the same way")
          spec = importlib.util.spec_from_file_location("fingerprint", path)
          mod = importlib.util.module_from_spec(spec)
          spec.loader.exec_module(mod)
          return mod
      
      
      # ---------------------------------------------------------------- reading LaTeX
      
      def _braced(text, start):
          """(content, end) of the brace group opening at text[start] == '{', nested groups included."""
          depth, i = 0, start
          while i < len(text):
              c = text[i]
              if c == "\\":
                  i += 2
                  continue
              if c == "{":
                  depth += 1
              elif c == "}":
                  depth -= 1
                  if depth == 0:
                      return text[start + 1:i], i + 1
              i += 1
          return text[start + 1:], len(text)
      
      
      def _replace_command(text, name, fn):
          out, i = [], 0
          pat = re.compile(r"\\" + name + r"\*?\s*(?:\[[^\]]*\])?\s*\{")
          while True:
              m = pat.search(text, i)
              if not m:
                  out.append(text[i:])
                  return "".join(out)
              body, end = _braced(text, m.end() - 1)
              out.append(text[i:m.start()])
              out.append(fn(body))
              i = end
      
      
      def pre_tex(text):
          """LaTeX to the prose a reader sees, before the fingerprint's own stripping: lists and captions are prose here."""
          body = re.split(r"\\begin\{document\}", text, maxsplit=1)
          text = body[1] if len(body) == 2 else text
          text = re.sub(r"(?<!\\)%.*", "", text)
          text = re.sub(r"\n[ \t]*\n", BREAK, text)   # a sentence never runs across a blank line
          text = re.sub(r"\\(?:part|chapter|section|subsection|subsubsection|paragraph|subparagraph)\*?\s*(?:\[[^\]]*\])?"
                        r"\s*\{(?:[^{}]|\{[^{}]*\})*\}", BREAK, text)
          # A float becomes its caption; the table body and graphics are not sentences.
          def float_to_caption(m):
              caps = []
              _replace_command(m.group(0), "caption", lambda b: caps.append(b) or "")
              return BREAK + BREAK.join(caps) + BREAK
          text = re.sub(r"\\begin\{(figure|table)\*?\}.*?\\end\{\1\*?\}", float_to_caption, text, flags=re.S)
          text = re.sub(r"\\(?:begin|end)\{(?:itemize|enumerate|description)\}", BREAK, text)
          text = re.sub(r"\\item(?:\[[^\]]*\])?", BREAK, text)
          text = re.sub(r"\\S\s*~?\s*(?=\\(?:c|C|auto|name|page|eq|v)?ref)", " ", text)
          text = re.sub(r"\\(?:c|C|auto|name|page|eq|v)?ref\*?\{[^}]*\}", " ", text)
          text = re.sub(r"\\(?:cite\w*|parencite|textcite|autocite|footcite)\*?(?:\s*\[[^\]]*\]){0,2}\s*\{[^}]*\}", " ", text)
          text = re.sub(r"\\label\{[^}]*\}", " ", text)
          text = _replace_command(text, "footnote", lambda b: " (" + b + ") ")
          text = text.replace("\\textemdash", " --- ").replace("\\textendash", " -- ")
          text = re.sub(r"(?<!\\)\$[^$]+\$|\\\(.*?\\\)", " MATH ", text)
          return text
      
      
      def pre_md(text):
          """Markdown headings, list items and blank lines end a unit of prose; without this a heading is read as the start
          of the sentence after it."""
          text = re.sub(r"(?m)^\s{0,3}#+.*$", BREAK, text)
          text = re.sub(r"(?m)^\s*(?:[-*+]|\d+[.)])\s+", BREAK, text)
          return re.sub(r"\n[ \t]*\n", BREAK, text)
      
      
      def read_text(fp, path):
          suffix = path.suffix.lower()
          if suffix == ".tex":
              raw = path.read_text(encoding="utf-8", errors="replace")
              return fp.strip_markup(pre_tex(raw), ".tex")
          if suffix == ".md":
              raw = path.read_text(encoding="utf-8", errors="replace")
              return fp.strip_markup(pre_md(raw), ".md")
          return fp.load(path)
      
      
      def sentence_units(text):
          """[(unit, sentence)]: a unit is what a blank line, a heading or a list item closes, a paragraph in prose."""
          text = ABBREV.sub(lambda m: m.group(1), re.sub(r"\s+", " ", text or ""))
          return [(n, s.strip()) for n, unit in enumerate(text.split(BREAK.strip())) for s in SPLIT.split(unit.strip())
                  if len(s.split()) >= 3]
      
      
      def sentences(text):
          return [s for _, s in sentence_units(text)]
      
      
      # ---------------------------------------------------------------- measuring
      
      def _words(s):
          return re.findall(r"[A-Za-z][A-Za-z'-]*", s)
      
      
      def clause_tokens(s):
          toks = [m.group(1).lower() for m in SUB.finditer(s)]
          for m in TEMPORAL_PREP.finditer(s):
              w = m.group(0).split()[0].lower()
              if w in toks:
                  toks.remove(w)
          toks += ["as"] * len(AS_CLAUSE.findall(s)) + ["so that"] * len(SO_THAT.findall(s))
          return toks
      
      
      def adverb_tokens(s):
          out = []
          words = _words(s)
          for i, w in enumerate(words):
              lw = w.lower()
              if lw in OTHER_ADVERBS:
                  out.append(lw)
              elif LY.fullmatch(w) and lw not in NOT_ADVERB and not (w[0].isupper() and i > 0):
                  out.append(lw)
          return out
      
      
      def modifier_tokens(s):
          out = []
          for m in PARTICIPLE_AFTER_NOUN.finditer(s):
              if m.group(1).lower() not in BE_HAVE:
                  out.append(m.group(2).lower())
          out += [m.group(1).lower() for m in PARTICIPLE_AFTER_COMMA.finditer(s)]
          out += [m.group(1).lower() for m in PRENOMINAL.finditer(s)]
          low = s[0].lower() + s[1:] if s else s
          for w in ADJ_SUFFIX.findall(low):
              if w not in NOUN_STOP:
                  out.append(w)
          out += [w.lower() for w in _words(s) if w.lower() in COMMON_ADJ]
          out += [m.group(1) for m in HYPHEN_PRENOMINAL.finditer(low)]
          return out
      
      
      def worded_parentheses(s):
          """Parentheses holding a lowercase word: an aside ("(roughly)", "(both historians)"), not a number, a symbol or an
          abbreviation being defined ("(CC)")."""
          return sum(1 for inner in PAREN_WORDED.findall(s) if re.search(r"\b[a-z]{3,}\b", inner))
      
      
      # A link to the sentence before ("However,", "For example,", ", in turn,") is not what makes a rewrite long or dense,
      # and a draft that states too few of them is its own problem (the fingerprint's linking opener share). Before a
      # rewrite is set against the sentence it replaced, a sentence-initial linking adverbial (the fingerprint's
      # LINKING_OPENER) and a linking adverbial set off by commas inside the sentence are taken out of both, so adding one
      # is not a comma, an adverb, an opener or three more words. What was added is reported as links_added, unflagged.
      # Held against the venue, a sentence is measured whole, as the venue's own sentences are.
      MID_LINK = re.compile(r",\s+(?:however|therefore|thus|hence|moreover|furthermore|consequently|nevertheless|nonetheless|"
                            r"accordingly|conversely|similarly|likewise|indeed|instead|specifically|notably|in turn|"
                            r"for example|for instance|in particular|in contrast|by contrast|as a result|in other words|"
                            r"that is|in addition),(?=\s)", re.I)
      
      
      def links(s):
          """The linking adverbials unlinked() takes out, lower-cased, in order."""
          m = FP.LINKING_OPENER.match(s) if FP else None
          found = [m.group(0).strip(" ,\"'\u201c\u2018").lower()] if m else []
          return found + [x.group(0).strip(" ,").lower() for x in MID_LINK.finditer(s)]
      
      
      def unlinked(s):
          m = FP.LINKING_OPENER.match(s) if FP else None
          if m:
              rest = s[m.end():].lstrip(" ,")
              s = rest[:1].upper() + rest[1:] if rest else s
          return MID_LINK.sub("", s)
      
      
      def features(s):
          return {
              "words": len(s.split()),
              "clauses": len(clause_tokens(s)),
              "commas": s.count(","),
              "colons": len(re.findall(r":(?!\d)", s)),
              "semicolons": s.count(";"),
              "dashes": len(DASH.findall(s)),
              "parentheses": s.count("("),
              "adverbs": len(adverb_tokens(s)),
              "modifiers": len(modifier_tokens(s)),
              "prepositions": len(PREP.findall(s)),
              "opener": 1 if OPENER.match(s) else 0,
          }
      
      
      def summed(sents):
          total = {k: 0 for k in FEATURES}
          for s in sents:
              for k, v in features(s).items():
                  total[k] += v
          total["opener"] = min(total["opener"], 1)
          return total
      
      
      def gained(old_sents, new_sents, fn):
          """Occurrences the new text has beyond the old, as a sorted list (a multiset difference, not a net count)."""
          extra = Counter(t for s in new_sents for t in fn(s)) - Counter(t for s in old_sents for t in fn(s))
          return sorted(extra.elements())
      
      
      # ---------------------------------------------------------------- reading files and pairs
      
      def read_prose(fp, path, units=None):
          """{file: [sentences]} for a file, or for every prose file under a directory. Directories whose names start with
          a dot are skipped: the loop puts the previous version beside the draft in one. units, when given, is filled with
          {file: [paragraph number of each sentence]}."""
          path = Path(path)
          files = [path] if path.is_file() else sorted(
              p for p in path.rglob("*") if p.is_file() and not any(part.startswith(".") for part in p.relative_to(path).parts))
          out = {}
          for p in files:
              if p.suffix.lower() not in fp.TEXT_SUFFIXES:
                  continue
              t = read_text(fp, p)
              if t and t.strip():
                  key = p.name if path.is_file() else str(p.relative_to(path))
                  pairs = sentence_units(t)
                  out[key] = [s for _, s in pairs]
                  if units is not None:
                      units[key] = [u for u, _ in pairs]
          return out
      
      
      def norm(s):
          return re.sub(r"\s+", " ", s).strip().lower()
      
      
      def content(s):
          return {w.lower() for w in _words(s) if len(w) >= 3 and w.lower() not in STOPWORDS | SUB_WORDS}
      
      
      def share(part, whole):
          a, b = content(part), content(whole)
          return (len(a & b) / len(a)) if a else 0.0, len(a & b)
      
      
      def pair_changes(target, base):
          """[(file, [old sentences], [new sentences])] for every target sentence that does not occur in the base, grouped:
          a revision is one old and one new, a split one old and several new, a merge several old and one new, an addition
          no old. Second value: the removed base sentences matched to nothing, as (file, sentence)."""
          base_all = {norm(s) for ss in base.values() for s in ss}
          target_all = {norm(s) for ss in target.values() for s in ss}
          removed = {f: [s for s in ss if norm(s) not in target_all] for f, ss in base.items()}
          pool = [(f, s) for f, ss in removed.items() for s in ss]
          used = set()
          matched, unmatched = [], []
          for f, ss in target.items():
              for s in ss:
                  if norm(s) in base_all:
                      continue
                  best, score = None, 0.0
                  words = s.split()
                  for pf, ps in sorted(pool, key=lambda x: x[0] != f):   # prefer the same file, then anywhere
                      if (pf, ps) in used:
                          continue
                      r = difflib.SequenceMatcher(None, words, ps.split(), autojunk=False).ratio()
                      if r > score:
                          best, score = (pf, ps), r
                  if best and score >= MATCH_MIN:
                      used.add(best)
                      matched.append([f, [best], [s]])
                  else:
                      unmatched.append((f, s))
          groups = {id(g): g for g in matched}
          by_old = {g[1][0]: g for g in matched}
          rest = []
          for f, s in unmatched:
              # a piece of a split: most of its content words come from one old sentence (matched or not)
              best, best_share = None, 0.0
              for old in pool:
                  sh, n = share(s, old[1])
                  if n >= 3 and sh >= PIECE_MIN and sh > best_share:
                      best, best_share = old, sh
              if best is not None:
                  g = by_old.get(best)
                  if g is None:
                      used.add(best)
                      g = [f, [best], []]
                      by_old[best] = g
                      groups[id(g)] = g
                  g[2].append(s)
              else:
                  rest.append((f, s))
          def holds(new, old):
              # most of the old sentence's content words are in the new one; a short old sentence needs all of them
              sh, n = share(old, new)
              return sh >= PIECE_MIN and n >= min(3, len(content(old))) and n >= 1
      
          for f, s in rest:
              olds = [old for old in pool if old not in used and holds(s, old[1])]
              if olds:
                  used.update(olds)
                  g = [f, olds, [s]]
              else:
                  g = [f, [], [s]]
              groups[id(g)] = g
          out = []
          for g in groups.values():
              f, olds, news = g
              out.append((f, [o[1] for o in olds], news))
          # a merged sentence matched by similarity to one old sentence may hold others that nothing else claimed
          for g in out:
              if len(g[1]) == 1 and len(g[2]) == 1:
                  extra = [old for old in pool if old not in used and holds(g[2][0], old[1])]
                  if extra:
                      used.update(extra)
                      g[1].extend(o[1] for o in extra)
          return out, [old for old in pool if old not in used]
      
      
      def carried(sentence, carriers):
          """What a removed sentence carried: the caller's patterns it matches, then a number, then a qualifier."""
          out = [f"pattern {rx.pattern}" for rx in carriers if rx.search(sentence)]
          if REMOVED_NUMBER.search(sentence):
              out.append("number")
          q = REMOVED_QUALIFIER.search(sentence)
          if q:
              out.append(f"qualifier '{q.group(0)}'")
          return out
      
      
      # A removal can take away what a later "the X" refers to (spec 2026-09-29-deletion-side-effects D1). On one manuscript
      # an abstract lost the only sentence that named its subject, and its last sentence, ten sentences on in the same
      # paragraph, still said "the <subject>"; a reader panel caught it, no check did. The rest of the paragraph is read with
      # no cap (five sentences would have missed that case); the whole file triples the hits and catches nothing more.
      # "that" is left out: in this prose it opens a clause far more often than it points back ("a baseline that knows").
      # At most one word between the determiner and the noun: two reach past the noun to a verb ("the collection already
      # records"). Every real case so far had none or one ("the same failure").
      REFERS = r"\b(?:the|this|these|those)\s+(?:[A-Za-z-]+\s+)?"
      NOT_NOUNS = {"first", "second", "third", "fourth", "fifth", "other", "others", "former", "latter", "same",
                   "following", "last", "next", "above", "below",
                   # modals and auxiliaries: "those scores cannot" is not a reference to "cannot"
                   "cannot", "could", "would", "should", "might", "must", "will", "does", "have", "been", "being"}
      
      
      PARTICIPLE = re.compile(r"[^e]ed$", re.I)   # added, used; need, speed and seed stay nouns
      
      
      def noun_key(w):
          """A word as a noun to look for: lower case, a possessive dropped, a plural folded onto its singular."""
          w = re.sub(r"['’]s?$", "", w.lower())
          return w[:-1] if w.endswith("s") and not w.endswith("ss") and len(w) > 4 else w
      
      
      def names(key):
          return re.compile(r"\b" + re.escape(key) + r"(?:s|es)?(?![a-z])", re.I)
      
      
      def lost_antecedents(removed, changes, base, units, target):
          """{(file, removed sentence): [{phrase, sentence}]}: a later sentence of the same paragraph, kept or revised, that
          says the/this/these + a noun of the removed sentence, when no earlier sentence of the file still names that noun."""
          news_of = {}
          for _, olds, news in changes:
              for o in olds:
                  news_of.setdefault(o, []).extend(news)
          out = {}
          for f, r in removed:
              ss, us, ts = base.get(f) or [], units.get(f) or [], target.get(f) or []
              if r not in ss or not ts or len(us) != len(ss):
                  continue
              i = ss.index(r)
              # a past participle is not a noun: "the sensors added later" does not point back to "was added"
              keys = sorted({noun_key(w) for w in content(r) if len(w) >= 4 and not PARTICIPLE.search(w)} - NOT_NOUNS)
              tn = [norm(t) for t in ts]
              hits = []
              for k in range(i + 1, len(ss)):
                  if us[k] != us[i]:
                      break
                  for t in ([ss[k]] if norm(ss[k]) in tn else news_of.get(ss[k], [])):
                      if norm(t) not in tn:
                          continue
                      before = " ".join(ts[:tn.index(norm(t))])
                      for key in keys:
                          # followed by an article, the word is a verb ("this manuscript addresses a question")
                          m = re.search(REFERS + re.escape(key) + r"(?:s|es)?(?![a-z])(?!\s+(?:a|an|the)\b)", t, re.I)
                          # one entry per noun phrase: "the candidate" and "the candidate pool" are one place to read
                          if m and not names(key).search(before) and not any(
                                  h["sentence"] == t and (h["phrase"] in m.group(0) or m.group(0) in h["phrase"])
                                  for h in hits):
                              hits.append({"phrase": m.group(0), "sentence": t})
              if hits:
                  out[(f, r)] = hits
          return out
      
      
      # A share that disagrees with the draft's other shares of the same total (D2). On one manuscript a denominator added for
      # clarity ("one of three later ...") sat against "two of the three later ..." elsewhere; the gate flagged the sentence
      # for a comma only. Listing every count of the same noun set about twenty unrelated sentences beside it; keeping the
      # total fixed and the share different left the ones that disagreed.
      COUNT_WORDS = {"one": 1, "two": 2, "three": 3, "four": 4, "five": 5, "six": 6, "seven": 7, "eight": 8, "nine": 9,
                     "ten": 10, "eleven": 11, "twelve": 12}
      _N = r"(" + "|".join(COUNT_WORDS) + r"|\d+)"
      SHARE = re.compile(r"\b" + _N + r"\s+of\s+(?:the\s+)?" + _N + r"\s+(?:[A-Za-z-]+\s+){0,2}?([A-Za-z-]+s)\b", re.I)
      
      
      def _count(w):
          return COUNT_WORDS.get(w.lower(), int(w) if w.isdigit() else None)
      
      
      def shares(s):
          """[(share, total, noun, text)] for each "K of (the) N <noun>s" in a sentence."""
          out = []
          for m in SHARE.finditer(s):
              k, n = _count(m.group(1)), _count(m.group(2))
              if k is not None and n is not None and k <= n:
                  out.append((k, n, noun_key(m.group(3)), m.group(0)))
          return out
      
      
      def shares_elsewhere(olds, news, draft):
          """[{share, sentences}]: each share a new sentence gained whose noun and total the draft counts with another share."""
          had = {(k, n, key) for o in olds for k, n, key, _ in shares(o)}
          found = []
          for s in news:
              for k, n, key, text in shares(s):
                  if (k, n, key) in had:
                      continue
                  others = [t for t in draft if t not in news and
                            any(n2 == n and key2 == key and k2 != k for k2, n2, key2, _ in shares(t))]
                  if others:
                      found.append({"share": text, "sentences": others})
          return found
      
      
      # A sentence written the same as another in the draft. On one manuscript a round that removed repeated statements
      # rewrote a results sentence into one that matched a dataset sentence letter for letter; only a cross-reference told
      # them apart, and the gate flagged a parenthesis. Sentences are compared on letters and digits alone. Pairing takes a
      # sentence already in the base as unchanged (so a move is free), which hides a copy; a copy is found by count instead:
      # the draft holds the sentence more often than the base did.
      DUP_MIN_WORDS = 6
      
      
      # What is left of a parenthesis once its \ref is gone: "(Section )", "(see Table )", "( )".
      REF_LEFTOVER = re.compile(r"\(\s*(?:see\s+)?(?:(?:supplementary\s+)?(?:sections?|tables?|figures?|figs?\.?|appendix|"
                                r"appendices|eqs?\.?|equations?)\s*(?:and\s+)?)*\)", re.I)
      
      
      def same_key(s):
          return re.sub(r"[^a-z0-9]", "", REF_LEFTOVER.sub("", s).lower())
      
      
      def duplicated(changes, target, base):
          """({(file, new sentence): [{where, sentence}]} for changed or added sentences that match another sentence of the
          draft, [(file, sentence, [{where, sentence}])] for sentences the draft holds more often than the base did)."""
          tpos = [(f, s) for f, ss in target.items() for s in ss if len(s.split()) >= DUP_MIN_WORDS]
          tkeys = Counter(same_key(s) for _, s in tpos)
          bkeys = Counter(same_key(s) for ss in base.values() for s in ss if len(s.split()) >= DUP_MIN_WORDS)
      
          def others(f, s):
              k, out, skipped = same_key(s), [], False
              for f2, t in tpos:
                  if same_key(t) != k:
                      continue
                  if not skipped and f2 == f and t == s:
                      skipped = True
                      continue
                  out.append({"where": f2, "sentence": t})
              return out
      
          in_changes, changed_keys = {}, set()
          for f, _, news in changes:
              for s in news:
                  changed_keys.add(same_key(s))
                  if len(s.split()) >= DUP_MIN_WORDS and tkeys[same_key(s)] >= 2:
                      in_changes[(f, s)] = others(f, s)
          copies = []
          # the copy is reported where the file gained the sentence, not where it already stood
          tfile = Counter((f, same_key(s)) for f, s in tpos)
          bfile = Counter((f, same_key(s)) for f, ss in base.items() for s in ss if len(s.split()) >= DUP_MIN_WORDS)
          for k, n in tkeys.items():
              if n >= 2 and n > bkeys.get(k, 0) and k not in changed_keys:
                  at = [(f, s) for f, s in tpos if same_key(s) == k]
                  f, s = next(((f, s) for f, s in at if tfile[(f, k)] > bfile.get((f, k), 0)), at[0])
                  copies.append((f, s, others(f, s)))
          return in_changes, copies
      
      
      def read_carriers(path):
          if not Path(path).is_file():
              die(f"no carriers file at {path}")
          out = []
          for n, line in enumerate(Path(path).read_text(encoding="utf-8").splitlines(), 1):
              line = line.strip()
              if not line or line.startswith("#"):
                  continue
              try:
                  out.append(re.compile(line, re.I))
              except re.error as e:
                  die(f"{path}:{n}: not a regular expression ({e}): {line}")
          return out
      
      
      def prose(cell):
          """A proposal is usually written in the draft's markup; read it the way the draft is read."""
          cell = (cell or "").strip()
          if "\\" in cell or "~" in cell or "$" in cell or "%" in cell:
              cell = FP.strip_markup(pre_tex(cell), ".tex")
          return re.sub(r"\s+", " ", cell).strip()
      
      
      VERDICTS = {"accepted": "accepted", "accept": "accepted", "kept": "accepted", "接受": "accepted", "留": "accepted",
                  "保留": "accepted", "rejected": "rejected", "reject": "rejected", "revised": "rejected",
                  "退回": "rejected", "拒": "rejected", "不要": "rejected", "改掉": "rejected", "作者改": "rejected"}
      
      
      def read_pairs(path):
          """[(id, old sentences, new sentences, verdict, reason)]. verdict is accepted, rejected or None (not judged)."""
          rows = []
          with open(path, newline="", encoding="utf-8") as fh:
              reader = csv.reader(fh, delimiter="\t", quoting=csv.QUOTE_NONE)
              header = next(reader, None)
              if not header or not {"id", "old", "new"} <= {h.strip() for h in header}:
                  die(f"{path}: the header must name the columns id, old, new (tab-separated)")
              col = {h.strip(): i for i, h in enumerate(header)}
              for alias, name in (("作者裁定", "verdict"), ("原因", "reason")):
                  if alias in col and name not in col:
                      col[name] = col[alias]
              for n, r in enumerate(reader, start=2):
                  if not any(c.strip() for c in r):
                      continue
                  if len(r) > len(header):
                      die(f"{path}:{n}: {len(r)} cells where the header has {len(header)} (a tab inside a cell?)")
                  get = lambda k: r[col[k]] if k in col and col[k] < len(r) else ""  # noqa: E731
                  raw = get("verdict").strip().lower()
                  if raw and raw not in VERDICTS:
                      die(f"{path}:{n}: verdict {get('verdict')!r} is not one of accepted, rejected, revised (or empty)")
                  new = prose(get("new"))
                  if new:
                      old = prose(get("old"))
                      rows.append((get("id"), sentences(old) if old else [], sentences(new) or [new],
                                   VERDICTS.get(raw), get("reason").strip() or None))
          return rows
      
      
      def _venue_key(baseline):
          """What the venue's distribution depends on: the corpus files (name, size, modification time) and the code that
          measures them (this script and the fingerprint audit it reads prose with)."""
          import hashlib
          h = hashlib.sha256()
          for f in (Path(__file__).resolve(), HERE / "audit-prose-fingerprint.py"):
              h.update(f.read_bytes())
          for f in sorted(p for p in Path(baseline).rglob("*") if p.is_file()):
              st = f.stat()
              h.update(f"{f.relative_to(baseline)}\0{st.st_size}\0{st.st_mtime_ns}\0".encode("utf-8"))
          return h.hexdigest()
      
      
      def venue_distribution(fp, baseline, cache=None):
          """The venue's sentences, measured. Reading a corpus of PDFs takes seconds; with --venue-cache the columns are kept
          and reused while neither the corpus nor this code has changed."""
          key = _venue_key(baseline) if cache else None
          if cache:
              try:
                  got = json.loads(Path(cache).read_text(encoding="utf-8"))
                  if got.get("key") == key:
                      return {"documents": got["documents"], "sentences": got["sentences"], "columns": got["columns"],
                              "cached": True}
              except (OSError, ValueError, KeyError):
                  pass
          docs = fp.collect(Path(baseline))
          kept, sents = 0, []
          for _, text, _ in docs:
              if hasattr(fp, "word_share") and fp.word_share(text) < getattr(fp, "MIN_WORD_SHARE", 0.3):
                  continue
              ss = [s for s in sentences(text) if 5 <= len(s.split()) <= 120]
              if ss:
                  kept += 1
                  sents.extend(ss)
          if kept < MIN_VENUE_DOCUMENTS or len(sents) < MIN_VENUE_SENTENCES:
              die(f"the baseline {baseline} gave {len(sents)} sentences from {kept} documents; percentiles need at least "
                  f"{MIN_VENUE_SENTENCES} sentences from {MIN_VENUE_DOCUMENTS} documents")
          cols = {k: sorted(features(s)[k] for s in sents) for k in FEATURES}
          if cache:
              Path(cache).parent.mkdir(parents=True, exist_ok=True)
              Path(cache).write_text(json.dumps({"key": key, "documents": kept, "sentences": len(sents), "columns": cols}),
                                     encoding="utf-8")
          return {"documents": kept, "sentences": len(sents), "columns": cols}
      
      
      def percentile(col, x):
          below = sum(1 for v in col if v < x)
          equal = sum(1 for v in col if v == x)
          return round(100 * (below + equal / 2) / len(col))
      
      
      def at(col, pct):
          return col[int(pct / 100 * (len(col) - 1))]
      
      
      DENSITY = ("clauses", "commas", "prepositions", "adverbs", "modifiers")
      # Ceilings for an added sentence when no venue corpus is given: the 75th percentiles of one journal's published papers
      # as this script measures them. A venue corpus replaces them, and the report says which was used. Why the 75th: on the
      # sentences added in one real round, the 90th flagged none, the 50th flagged about half (most for a single "also" or
      # "often"), and the 75th flagged only the densest. That is the round the value was chosen on.
      DEFAULT_CEILING = {"words": 31, "clauses": 1, "commas": 2, "prepositions": 3, "adverbs": 1, "modifiers": 3}
      
      
      # ---------------------------------------------------------------- judging
      
      def judge(olds, news, venue):
          """One group: the old sentences (none for an addition) and the new ones."""
          kind = ("added" if not olds else "split" if len(news) > 1 and len(olds) == 1
                  else "merged" if len(olds) > 1 else "revised")
          raw_news, old_text = news, olds
          links_added = sorted((Counter(t for x in news for t in links(x)) - Counter(t for x in olds for t in links(x))).elements())
          olds, news = [unlinked(x) for x in olds], [unlinked(x) for x in news]
          fn = summed(news)
          fo = summed(olds) if olds else None
          flags, added = [], {}
          if fo:
              if fn["words"] - fo["words"] >= LONGER_WORDS:
                  flags.append("longer")
              punct = ("commas", "colons", "semicolons", "dashes")
              # A comma that replaces a semicolon or a colon is not added punctuation.
              if fn["commas"] > fo["commas"] and sum(fn[k] for k in punct) > sum(fo[k] for k in punct):
                  flags.append("comma")
              for k, name in (("colons", "colon"), ("semicolons", "semicolon"), ("dashes", "dash"),
                              ("parentheses", "parenthesis")):
                  if fn[k] > fo[k]:
                      flags.append(name)
              # Clauses are counted as occurrences gained: trading "because" for "which" turns a reason into a relative
              # clause. Adverbs and modifiers are counted net: a correction that swaps one term for another adds none.
              g = gained(olds, news, clause_tokens)
              if g:
                  flags.append("clause")
                  added["clauses"] = g
              for k, name, fnc in (("adverbs", "adverb", adverb_tokens), ("modifiers", "modifier", modifier_tokens)):
                  g = gained(olds, news, fnc)
                  if fn[k] > fo[k] and g:
                      flags.append(name)
                      added[k] = g
              if fn["prepositions"] - fo["prepositions"] >= MORE_PREPOSITIONS:
                  flags.append("prepositions")
              if fn["opener"] and not fo["opener"]:
                  flags.append("opener")
              if kind == "merged":
                  flags.append("merged")
          else:
              for k, name in (("colons", "colon"), ("semicolons", "semicolon"), ("dashes", "dash")):
                  if fn[k]:
                      flags.append(name)
              if any(worded_parentheses(s) for s in news):
                  flags.append("parenthesis")
              added = {k: [t for s in news for t in f(s)] for k, f in
                       (("clauses", clause_tokens), ("adverbs", adverb_tokens), ("modifiers", modifier_tokens))}
              added = {k: v for k, v in added.items() if v}
          pct, dense_hits = {}, set()
          pct_used = ADDITION_PCT if fo is None else VENUE_PCT
          ceiling = ({k: at(venue["columns"][k], pct_used) for k in ("words",) + DENSITY} if venue else DEFAULT_CEILING)
          # Without a venue, only additions are held to the default ceilings: a revision is already judged against the
          # sentence it replaced.
          if venue or fo is None:
              for s in raw_news:
                  f1 = features(s)
                  if venue:
                      pct = {k: percentile(venue["columns"][k], f1[k]) for k in ("words",) + DENSITY}
                  longest_old = max((features(o)["words"] for o in olds), default=0)
                  grew = fo is None or features(unlinked(s))["words"] > (fo["words"] if kind != "merged" else longest_old)
                  if f1["words"] > ceiling["words"] and grew:
                      flags.append("long_for_venue")
                  for k in DENSITY:
                      if f1[k] > ceiling[k] and (fo is None or fn[k] > fo[k]):
                          dense_hits.add(k)
              if dense_hits:
                  flags.append("dense_for_venue")
          return {"old": " ".join(old_text) or None, "new": " ".join(raw_news), "kind": kind, "features_old": fo,
                  "features_new": fn, "added": added, "links_added": links_added, "dense": sorted(dense_hits),
                  "venue_percentile": pct,
                  "flags": sorted(set(flags), key=flags.index), "pieces": len(news), "olds": len(olds)}
      
      
      def verdict_table(results):
          """The flags set against the author's verdicts, and which version of this script (its thresholds) set them."""
          import hashlib
          t = {"judged": 0, "flagged_rejected": 0, "flagged_accepted": 0, "unflagged_rejected": 0, "unflagged_accepted": 0,
               "script": hashlib.sha256(Path(__file__).resolve().read_bytes()).hexdigest()}
          for r in results:
              v = r.get("verdict")
              if v:
                  t["judged"] += 1
                  t[("flagged_" if r["flags"] else "unflagged_") + v] += 1
          return t
      
      
      def main():
          ap = argparse.ArgumentParser(description="Check rewritten sentences one by one.")
          ap.add_argument("--target")
          ap.add_argument("--base")
          ap.add_argument("--pairs")
          ap.add_argument("--baseline")
          ap.add_argument("--venue-cache", help="keep the venue's measured sentences here and reuse them while unchanged")
          ap.add_argument("--carriers", help="patterns (one per line) that a removed sentence must not take with it unflagged")
          ap.add_argument("--json", action="store_true")
          a = ap.parse_args()
          global FP
          fp = FP = fingerprint()
          removed = []
          carriers = read_carriers(a.carriers) if a.carriers else []
          verdicts = {}
          if a.pairs:
              if a.target or a.base:
                  die("give --pairs, or --target with --base, not both")
              if not Path(a.pairs).is_file():
                  die(f"no pairs file at {a.pairs}")
              rows = read_pairs(a.pairs)
              if not rows:
                  die(f"{a.pairs} holds no rewritten sentence")
              verdicts = {rid: (v, why) for rid, _, _, v, why in rows}
              changes = [(rid, old, new) for rid, old, new, _, _ in rows if norm(" ".join(old)) != norm(" ".join(new))]
              compared = {"pairs": len(rows)}
          else:
              if not (a.target and a.base):
                  die("give --target and --base (the version before the edit), or --pairs")
              target = read_prose(fp, a.target) if Path(a.target).exists() else {}
              units = {}
              base = read_prose(fp, a.base, units) if Path(a.base).exists() else {}
              if not any(target.values()):
                  die(f"no prose in the target {a.target}")
              if not any(base.values()):
                  die(f"no prose in the base {a.base}: nothing to compare the draft with")
              changes, removed = pair_changes(target, base)
              compared = {"target_sentences": sum(map(len, target.values())), "base_sentences": sum(map(len, base.values()))}
          # What a removal took with it and a share that disagrees need the whole draft; a pairs file has none.
          lost = lost_antecedents(removed, changes, base, units, target) if not a.pairs else {}
          draft = [s for ss in target.values() for s in ss] if not a.pairs else []
          dups, copies = duplicated(changes, target, base) if not a.pairs else ({}, [])
          venue = venue_distribution(fp, a.baseline, a.venue_cache) if a.baseline else None
          results = []
          for where, olds, news in changes:
              r = judge(olds, news, venue)
              r["where"] = where
              found = shares_elsewhere(olds, news, draft) if draft else []
              if found:
                  r["flags"].append("count_elsewhere")
                  r["shares_elsewhere"] = found
              same = [d for s in news for d in dups.get((where, s), [])]
              if same:
                  r["flags"].append("duplicates_elsewhere")
                  r["duplicates"] = same
              if where in verdicts:
                  r["verdict"], r["reason"] = verdicts[where]
              results.append(r)
          removed_flagged = 0
          for where, old in removed:
              what = carried(old, carriers)
              took = lost.get((where, old)) or []
              if not what and not took:
                  continue
              removed_flagged += 1
              results.append({"old": old, "new": "", "kind": "removed", "where": where, "features_old": features(old),
                              "features_new": None, "added": {}, "dense": [], "venue_percentile": {}, "carried": what,
                              "antecedent": took,
                              "flags": (["removed_carrier"] if what else []) + (["took_antecedent"] if took else []),
                              "pieces": 0, "olds": 1})
              if where in verdicts:
                  results[-1]["verdict"], results[-1]["reason"] = verdicts[where]
          for where, s, same in copies:
              # a copy of a sentence the draft already had: pairing took it as unchanged, the count shows it
              results.append({"old": None, "new": s, "kind": "copied", "where": where, "features_old": None,
                              "features_new": features(s), "added": {}, "dense": [], "venue_percentile": {},
                              "duplicates": same, "flags": ["duplicates_elsewhere"], "pieces": 1, "olds": 0})
          flagged = [r for r in results if r["flags"]]
          kinds = Counter(r["kind"] for r in results)
          compared.update({"changed": len(results) - removed_flagged, "revised": kinds["revised"], "split": kinds["split"],
                           "merged": kinds["merged"], "added": kinds["added"], "removed": len(removed),
                           "removed_flagged": removed_flagged,
                           # counted apart (D3), so their volume on a real draft can be read
                           "took_antecedent": sum(1 for r in results if "took_antecedent" in r["flags"]),
                           "count_elsewhere": sum(1 for r in results if "count_elsewhere" in r["flags"]),
                           "copied": kinds["copied"],
                           "duplicates_elsewhere": sum(1 for r in results if "duplicates_elsewhere" in r["flags"])})
          unjudged = kinds["added"] if venue is None else 0   # judged against DEFAULT_CEILING, not a venue
          linked = Counter(t for r in results for t in r.get("links_added", []))
          out = {"schema_version": 2, "compared": compared, "changed": len(results) - removed_flagged, "flagged": len(flagged),
                 "added_without_venue": unjudged,
                 "links_added": dict(sorted(linked.items())),
                 "venue": ({"documents": venue["documents"], "sentences": venue["sentences"],
                            "p90": {k: at(venue["columns"][k], VENUE_PCT) for k in ("words",) + DENSITY},
                            "p75": {k: at(venue["columns"][k], ADDITION_PCT) for k in ("words",) + DENSITY}}
                           if venue else None),
                 "limits": LIMITS,
                 "verdicts": verdict_table(results) if verdicts else None,
                 "sentences": results,
                 "issues": [{"where": r["where"], "flags": r["flags"], "added": r["added"], "new": r["new"],
                             **({"old": r["old"], "carried": r["carried"], "antecedent": r["antecedent"]}
                                if r["kind"] == "removed" else {}),
                             **({"shares_elsewhere": r["shares_elsewhere"]} if r.get("shares_elsewhere") else {}),
                             **({"duplicates": r["duplicates"]} if r.get("duplicates") else {})}
                            for r in flagged]}
          if a.json:
              print(json.dumps(out, ensure_ascii=False, indent=1))
          else:
              print(f"changed sentences: {len(results) - removed_flagged} ({kinds['revised']} revised, {kinds['split']} split, "
                    f"{kinds['merged']} merged, {kinds['added']} added); removed without a successor: {len(removed)}, "
                    f"{removed_flagged} of them carrying something; flagged: {len(flagged)}")
              if venue:
                  print(f"venue: {venue['sentences']} sentences from {venue['documents']} documents; p90 "
                        + ", ".join(f"{k} {v}" for k, v in out["venue"]["p90"].items())
                        + "; added sentences held to p75 " + ", ".join(f"{k} {v}" for k, v in out["venue"]["p75"].items()))
              elif unjudged:
                  print(f"{unjudged} added sentence(s) held to default ceilings, not to a venue: give --baseline "
                        + "(" + ", ".join(f"{k} {v}" for k, v in DEFAULT_CEILING.items()) + ")")
              for r in flagged:
                  if r["kind"] == "removed":
                      print(f"\n[{r['where']}] removed: {', '.join(r['carried'] + ['took an antecedent'] * bool(r['antecedent']))}")
                      print(f"  was: {r['old']}")
                      for h in r["antecedent"]:
                          print(f"  a later sentence still says '{h['phrase']}': {h['sentence']}")
                      continue
                  fo, fn = r["features_old"], r["features_new"]
                  size = f"{fo['words']}->{fn['words']} words" if fo else f"{fn['words']} words, {r['kind']}"
                  extra = "; ".join(f"{k}: {', '.join(v)}" for k, v in r["added"].items())
                  print(f"\n[{r['where']}] {r['kind']}: {', '.join(r['flags'])}  ({size})" + (f"  [{extra}]" if extra else ""))
                  if r["old"]:
                      print(f"  was: {r['old']}")
                  print(f"  now: {r['new']}")
                  for d in r.get("duplicates") or []:
                      print(f"  the same, letter for letter, as [{d['where']}]: {d['sentence']}")
                  for x in r.get("shares_elsewhere") or []:
                      print(f"  '{x['share']}' beside the draft's other shares of the same total:")
                      for t in x["sentences"]:
                          print(f"    {t}")
              if not results and not removed:
                  print("no sentence changed")
              vt = out["verdicts"]
              if vt:
                  print(f"\nauthor verdicts on {vt['judged']} of {len(results)} changed sentences (script {vt['script'][:12]}): "
                        f"flagged and rejected {vt['flagged_rejected']}, flagged but kept {vt['flagged_accepted']}, "
                        f"not flagged but rejected {vt['unflagged_rejected']}, not flagged and kept {vt['unflagged_accepted']}")
              if linked:
                  n = sum(1 for r in results if r.get("links_added"))
                  print(f"\nlinking adverbials added, not flagged: {sum(linked.values())} in {n} sentence(s) ("
                        + ", ".join(f"{k} x{v}" for k, v in sorted(linked.items())) + ")")
              print(f"\n{LIMITS}")
          sys.exit(1 if flagged else 0)
      
      
      if __name__ == "__main__":
          main()
      
    • build-venue-baseline.py 13.3 KB
      #!/usr/bin/env python3
      """Build a baseline corpus of papers published in a target venue.
      
      The prose fingerprint compares a manuscript against a corpus. Until now that
      corpus was always an input someone had assembled by hand, and nothing checked
      what it represented. On one project the directory was named for the target
      journal, held the manuscript's own bibliography, and contained no article that
      journal had ever published -- so "inside the published range" meant "inside the
      range of the papers we cite" for seven rounds without anyone saying so.
      
      This builds the other corpus: papers the venue actually published. It is a
      different question from the bibliography baseline, not a better answer to the
      same one, and the two are never merged.
      
      Membership is verified against the registrar, never the author. arXiv's
      journal_ref is free text -- one real record reads "Just accpeted by ACM
      Computing Surveys 2026", typo included -- so a record is admitted only when it
      carries a DOI whose registered container-title equals the venue. That is the
      same DOI content negotiation the writing rules already require for
      bibliography entries.
      
      Every candidate leaves a trace. The manifest's dispositions sum to the number
      of candidates, and a run that cannot reach the corpus floor fails with a named
      code rather than returning a short corpus that reads like a complete one.
      
      Python 3.8 stdlib only. Network: export.arxiv.org, doi.org, arxiv.org.
      """
      
      import argparse
      import html
      import json
      import re
      import sys
      import time
      import urllib.error
      import urllib.parse
      import urllib.request
      import xml.etree.ElementTree as ET
      from pathlib import Path
      from typing import Dict, List, Optional
      
      ATOM = "{http://www.w3.org/2005/Atom}"
      ARXIV = "{http://arxiv.org/schemas/atom}"
      API = "http://export.arxiv.org/api/query"
      # arxiv.org answers 406 to a VERSIONED pdf path from a non-browser client --
      # /pdf/2407.17215v1 is refused while /pdf/2407.17215 is served. Dropping the
      # version would fix the fetch and lose the record of which version was
      # measured. export.arxiv.org is the host arXiv designates for automated
      # access, and it serves the versioned path. Measured 2026-09-20: 24 of 75
      # downloads failed this way against the main site and none against export.
      PDF_HOST = "https://export.arxiv.org/pdf/"
      DOWNLOAD_ATTEMPTS = 3
      UA = "awt-venue-baseline/0.1 (academic prose baseline construction)"
      PAGE = 100
      # The prose method refuses to report percentiles against fewer than twenty
      # published papers. A corpus under that is not a small corpus, it is not a
      # corpus, and returning it quietly is how a short baseline gets quoted.
      CORPUS_FLOOR = 20
      
      
      def die(code: str, message: str, remedy: str = "") -> None:
          sys.stderr.write("%s: %s\n" % (code, message))
          if remedy:
              sys.stderr.write("  remedy: %s\n" % remedy)
          sys.exit(2)
      
      
      def norm(s: Optional[str]) -> str:
          return " ".join((s or "").split()).strip()
      
      
      def fetch(url: str, accept: Optional[str] = None, timeout: int = 90) -> bytes:
          headers = {"User-Agent": UA}
          if accept:
              headers["Accept"] = accept
          return urllib.request.urlopen(
              urllib.request.Request(url, headers=headers), timeout=timeout).read()
      
      
      def query_arxiv(venue: str, delay: float) -> List[Dict]:
          """Every arXiv record whose journal_ref names the venue. Author-supplied."""
          out, start = [], 0
          while True:
              url = API + "?" + urllib.parse.urlencode(
                  {"search_query": 'jr:"%s"' % venue, "start": start,
                   "max_results": PAGE, "sortBy": "submittedDate",
                   "sortOrder": "descending"})
              try:
                  root = ET.fromstring(fetch(url))
              except (urllib.error.URLError, ET.ParseError) as exc:
                  die("VENUE_ARXIV_UNREACHABLE",
                      "the arXiv API did not answer: %s" % exc,
                      "check the network and re-run; nothing has been written")
              entries = root.findall(ATOM + "entry")
              for e in entries:
                  doi = e.findtext(ARXIV + "doi")
                  pub = e.findtext(ATOM + "published") or ""
                  out.append({
                      "arxiv_id": (e.findtext(ATOM + "id") or "").rsplit("/", 1)[-1],
                      "title": norm(e.findtext(ATOM + "title")),
                      "year": int(pub[:4]) if pub[:4].isdigit() else None,
                      "published": pub or None,
                      "primary_category": (e.find(ARXIV + "primary_category").get("term")
                                           if e.find(ARXIV + "primary_category") is not None else None),
                      "journal_ref": norm(e.findtext(ARXIV + "journal_ref")),
                      "doi": norm(doi) if doi else None,
                  })
              if len(entries) < PAGE:
                  return out
              start += PAGE
              time.sleep(delay)
      
      
      def venue_phrasings(venue: str) -> List[str]:
          """The venue as given, then with every free-standing "&" and "and" swapped.
      
          journal_ref is typed by authors, and they spell a journal whose name holds
          an ampersand both ways. The phrase search matches only the spelling it is
          given, so one query returns part of the frame. "&" inside a word ("R&D") is
          not a conjunction and is left alone.
          """
          out = [venue]
          for alt in (re.sub(r"(?<=\s)&(?=\s)", "and", venue),
                      re.sub(r"(?<=\s)and(?=\s)", "&", venue, flags=re.IGNORECASE)):
              if alt not in out:
                  out.append(alt)
          return out
      
      
      def container_title(doi: str, delay: float) -> Optional[str]:
          """What the registrar says the DOI was published in, not what the author typed."""
          try:
              d = json.loads(fetch("https://doi.org/" + urllib.parse.quote(doi),
                                   accept="application/vnd.citationstyles.csl+json",
                                   timeout=45).decode("utf-8", "replace"))
          except Exception:
              return None
          finally:
              time.sleep(delay)
          ct = d.get("container-title")
          if isinstance(ct, list):
              ct = ct[0] if ct else None
          # The registrar can send the title HTML-escaped: a journal named "X & Y"
          # arrives as "X &amp; Y", and compared verbatim it is some other journal.
          return norm(html.unescape(ct)) if ct else None
      
      
      def main() -> int:
          ap = argparse.ArgumentParser(
              description="Build a prose baseline from papers a venue published.")
          ap.add_argument("--venue", required=True,
                          help='journal name, matched against the registrar\'s '
                               'container-title (e.g. "ACM Computing Surveys")')
          ap.add_argument("--from-year", type=int, default=2021,
                          help="earliest arXiv submission year to admit (default 2021). "
                               "House genre drifts; the window is part of the frame")
          ap.add_argument("--out", help="directory the PDFs are written to")
          ap.add_argument("--manifest", required=True,
                          help="where the record of the sampling frame is written")
          ap.add_argument("--max", type=int, default=200, help="cap on admitted records")
          ap.add_argument("--allow-unverified", action="store_true",
                          help="admit records whose DOI does not resolve. They are "
                               "recorded as unverified and the manifest says so")
          ap.add_argument("--dry-run", action="store_true",
                          help="build the manifest and download nothing")
          ap.add_argument("--delay", type=float, default=3.0,
                          help="seconds between network calls (default 3.0)")
          args = ap.parse_args()
      
          if not args.dry_run and not args.out:
              die("VENUE_NO_OUTPUT_DIR", "--out is required unless --dry-run is passed",
                  "name a directory for the PDFs, outside the repository")
      
          venue_key = args.venue.lower()
          phrasings = venue_phrasings(args.venue)
          candidates: List[Dict] = []
          seen = set()
          for i, phrase in enumerate(phrasings):
              if i:
                  time.sleep(args.delay)
              for c in query_arxiv(phrase, args.delay):
                  # A journal_ref found under both spellings is one candidate.
                  if c["arxiv_id"] not in seen:
                      seen.add(c["arxiv_id"])
                      candidates.append(c)
          # Each query comes back newest first. Merged, keep that order, so --max
          # still cuts the oldest rather than whichever spelling was queried second.
          candidates.sort(key=lambda c: c.get("published") or "", reverse=True)
          if not candidates:
              die("VENUE_NO_CANDIDATES",
                  "no arXiv record carries a journal_ref naming %s"
                  % " or ".join('"%s"' % p for p in phrasings),
                  "check the venue string; the API matches it as a phrase")
      
          admitted: List[Dict] = []
          rejected: Dict[str, List[Dict]] = {
              "out_of_window": [], "no_doi": [], "doi_unresolvable": [],
              "container_title_mismatch": [], "over_max": [], "download_failed": []}
      
          for c in candidates:
              if c["year"] is None or c["year"] < args.from_year:
                  rejected["out_of_window"].append(c); continue
              if not c["doi"]:
                  # journal_ref alone is the author's own claim of where it appeared.
                  rejected["no_doi"].append(c); continue
              ct = container_title(c["doi"], args.delay)
              if ct is None:
                  if args.allow_unverified:
                      c["container_title"] = None
                      c["venue_verified"] = False
                      admitted.append(c)
                  else:
                      rejected["doi_unresolvable"].append(c)
                  continue
              c["container_title"] = ct
              if ct.lower() != venue_key:
                  rejected["container_title_mismatch"].append(c); continue
              c["venue_verified"] = True
              admitted.append(c)
      
          if len(admitted) > args.max:
              rejected["over_max"] = admitted[args.max:]
              admitted = admitted[: args.max]
      
          downloaded = 0
          if not args.dry_run:
              outdir = Path(args.out)
              outdir.mkdir(parents=True, exist_ok=True)
              kept = []
              for c in admitted:
                  dest = outdir / ("%s.pdf" % c["arxiv_id"].replace("/", "_"))
                  if dest.exists() and dest.stat().st_size > 0:
                      c["file"] = dest.name; kept.append(c); continue
                  last = None
                  for attempt in range(DOWNLOAD_ATTEMPTS):
                      try:
                          blob = fetch(PDF_HOST + c["arxiv_id"], accept="application/pdf",
                                       timeout=180)
                          if not blob.startswith(b"%PDF"):
                              raise ValueError("not a PDF")
                          dest.write_bytes(blob)
                          c["file"] = dest.name
                          kept.append(c); downloaded += 1
                          last = None
                          break
                      except Exception as exc:
                          # A truncated read is transient and was 4 of 28 failures in
                          # one real run; retrying it is the difference between a
                          # corpus and a corpus with holes in it.
                          last = exc
                          time.sleep(args.delay * (attempt + 1))
                  if last is not None:
                      c["error"] = "%s (after %d attempts)" % (str(last)[:100], DOWNLOAD_ATTEMPTS)
                      rejected["download_failed"].append(c)
                  time.sleep(args.delay)
              admitted = kept
      
          manifest = {
              # 2: "query" is a list, one entry per phrasing queried (1 held a string).
              "schema_version": 2,
              "venue": args.venue,
              "retrieved": time.strftime("%Y-%m-%dT%H:%M:%S%z"),
              "api": API,
              "query": ['jr:"%s"' % p for p in phrasings],
              "from_year": args.from_year,
              "verification": "arXiv journal_ref is author-supplied free text; a record "
                              "is admitted only when its DOI resolves to a registered "
                              "container-title equal to the venue",
              "stage": "authors' accepted manuscripts on arXiv, not the publisher's "
                       "copyedited pages",
              "corpus_dir": str(Path(args.out).resolve()) if args.out else None,
              "dry_run": bool(args.dry_run),
              "allow_unverified": bool(args.allow_unverified),
              "candidates": len(candidates),
              "admitted": len(admitted),
              "rejected": {k: len(v) for k, v in rejected.items()},
              "accounting_closes": len(admitted) + sum(len(v) for v in rejected.values())
                                   == len(candidates),
              "downloaded_now": downloaded,
              "records": admitted,
              "rejected_records": {k: v for k, v in rejected.items() if v},
          }
          Path(args.manifest).parent.mkdir(parents=True, exist_ok=True)
          Path(args.manifest).write_text(json.dumps(manifest, indent=1), encoding="utf-8")
      
          sys.stderr.write(
              "venue baseline: %d candidate(s) -> %d admitted (%s)\n" % (
                  len(candidates), len(admitted),
                  ", ".join("%s %d" % (k, len(v)) for k, v in rejected.items() if v) or "none rejected"))
          sys.stderr.write("manifest: %s\n" % args.manifest)
      
          if not manifest["accounting_closes"]:
              die("VENUE_ACCOUNTING_OPEN",
                  "dispositions do not sum to the candidate count; the manifest is "
                  "not a complete record of the frame")
          if len(admitted) < CORPUS_FLOOR:
              die("VENUE_CORPUS_TOO_SMALL",
                  "%d document(s) admitted; the prose method reports no percentile "
                  "below %d" % (len(admitted), CORPUS_FLOOR),
                  "widen --from-year, or accept that this venue has too little on "
                  "arXiv to measure against and say so rather than measuring anyway")
          return 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • citations.mjs 1.5 KB · in bundle
    • pdf-pages.mjs 530 B · in bundle
    • prose-view.py 6.2 KB
      #!/usr/bin/env python3
      """Write the prose a reader sees as Markdown chapters, so that checks written for chapters/*.md can read any draft.
      
          python3 prose-view.py --out <dir> [--format latex|markdown] FILE...
      
      Writes <dir>/chapters/NN-<name>.md, one per FILE, in the order given. A Markdown file is copied as it is. A LaTeX
      file is read the way the changed-sentence audit reads it (pre_tex in audit-sentence-changes.py): lists, citations,
      references, labels, footnotes and inline math are handled there; what is left of the markup is then stripped by the
      fingerprint audit's rules. On top of that, for the paragraph structure those checks need:
      
      - section headings become Markdown headings (`## Title`), and the abstract environment becomes `## Abstract`;
      - floats (figure, table) leave the running text; their captions are collected into one fenced block at the end of
        the chapter. A check that reads paragraphs skips fenced blocks; a check that reads every word still sees them;
      - display math, tabulars outside floats and verbatim environments are dropped.
      
      It is an approximation of the text, not a LaTeX parser. Macros the author defined are not expanded, so a word that
      lives only inside a custom macro is lost. Exit 0 with the files written; 2 when a file cannot be read or holds no
      prose at all.
      """
      import argparse
      import importlib.util
      import re
      import sys
      from pathlib import Path
      
      HERE = Path(__file__).resolve().parent
      DROP_ENVS = ("equation", "align", "gather", "multline", "eqnarray", "displaymath", "tabular", "tikzpicture",
                   "verbatim", "lstlisting", "algorithm", "algorithmic", "thebibliography", "CCSXML")
      # Commands whose arguments are file names, addresses or definitions, not prose: dropped with their arguments.
      NOT_PROSE = re.compile(r"\\(?:input|include|subfile|bibliography|bibliographystyle|addbibresource|includegraphics|"
                             r"url|graphicspath|usepackage|newcommand|renewcommand|providecommand|setcounter|setlength|"
                             r"hypersetup|pagestyle|thispagestyle|vspace|hspace)(?![A-Za-z])\*?(?:\s*\[[^\]]*\])?"
                             r"(?:\s*\{(?:[^{}]|\{[^{}]*\})*\})*")
      HEADING = re.compile(r"\\(?:part|chapter|section|subsection|subsubsection)\*?\s*(?:\[[^\]]*\])?\s*"
                           r"\{((?:[^{}]|\{[^{}]*\})*)\}(?:\s*\\label\{[^}]*\})?")
      
      
      def _load(name, filename):
          spec = importlib.util.spec_from_file_location(name, HERE / filename)
          mod = importlib.util.module_from_spec(spec)
          saved = sys.argv
          sys.argv = [filename]
          try:
              spec.loader.exec_module(mod)
          finally:
              sys.argv = saved
          return mod
      
      
      SC = _load("audit_sentence_changes", "audit-sentence-changes.py")
      FP = SC.fingerprint()
      MARK = "\u2063HEADING\u2063"
      
      
      def _plain(tex):
          """A fragment of LaTeX (a heading, a caption) as plain words."""
          return FP.strip_markup(SC.pre_tex(tex).replace(SC.BREAK, " "), ".tex")
      
      
      def _tidy(s):
          """What removing a citation or a reference leaves behind: empty brackets and a space before punctuation."""
          s = re.sub(r"\(\s*\)|\[\s*\]", "", s)
          s = re.sub(r"\s+([,.;:!?)])", r"\1", s)
          return re.sub(r"\s{2,}", " ", s).strip()
      
      
      def latex_view(text):
          body = re.split(r"\\begin\{document\}", text, maxsplit=1)
          text = body[1] if len(body) == 2 else text
          text = re.sub(r"(?<!\\)%.*", "", text)
          text = re.sub(r"\\end\{document\}.*", "", text, flags=re.S)
          captions = []
      
          def take_float(m):
              SC._replace_command(m.group(0), "caption", lambda b: captions.append(_plain(b)) or "")
              return "\n\n"
          text = re.sub(r"\\begin\{(figure|table)\*?\}.*?\\end\{\1\*?\}", take_float, text, flags=re.S)
          for env in DROP_ENVS:
              text = re.sub(r"\\begin\{" + env + r"\*?\}.*?\\end\{" + env + r"\*?\}", "\n\n", text, flags=re.S)
          text = re.sub(r"\\\[.*?\\\]|\$\$.*?\$\$", "\n\n", text, flags=re.S)
          text = re.sub(r"\\begin\{abstract\}", "\n\n" + MARK + "Abstract" + MARK + "\n\n", text)
          text = re.sub(r"\\end\{abstract\}", "\n\n", text)
          text = HEADING.sub(lambda m: "\n\n" + MARK + _plain(m.group(1)) + MARK + "\n\n", text)
          text = re.sub(r"\\makeatletter.*?\\makeatother", " ", text, flags=re.S)   # internal macros, never prose
          text = NOT_PROSE.sub(" ", text)
          text = re.sub(r"\\(?:maketitle|tableofcontents|printbibliography|appendix|clearpage|newpage|noindent)\b", " ", text)
          # An environment's name is not prose: \begin{center} would otherwise leave the word "center" behind.
          text = re.sub(r"\\(?:begin|end)\{[^}]*\}(?:\[[^\]]*\])?", " ", text)
          out = []
          for chunk in SC.pre_tex(text).split(SC.BREAK):
              chunk = chunk.strip()
              if not chunk:
                  continue
              if chunk.startswith(MARK) and chunk.endswith(MARK) and len(chunk) > 2 * len(MARK):
                  out.append("## " + chunk[len(MARK):-len(MARK)].strip())
                  continue
              prose = _tidy(FP.strip_markup(chunk, ".tex"))
              if prose:
                  out.append(prose)
          if captions:
              out.append("```captions\n" + "\n".join(c for c in captions if c) + "\n```")
          return "\n\n".join(out) + "\n"
      
      
      def main():
          ap = argparse.ArgumentParser(description="Write a draft's prose as Markdown chapters.")
          ap.add_argument("--out", required=True)
          ap.add_argument("--format", choices=("latex", "markdown"), default=None,
                          help="default: by file suffix (.tex is LaTeX, anything else Markdown)")
          ap.add_argument("files", nargs="+")
          a = ap.parse_args()
          dest = Path(a.out) / "chapters"
          dest.mkdir(parents=True, exist_ok=True)
          written = 0
          for n, name in enumerate(a.files, 1):
              p = Path(name)
              try:
                  raw = p.read_text(encoding="utf-8")
              except (OSError, UnicodeDecodeError) as e:
                  sys.stderr.write(f"prose-view: cannot read {name}: {e}\n")
                  return 2
              fmt = a.format or ("latex" if p.suffix.lower() == ".tex" else "markdown")
              view = latex_view(raw) if fmt == "latex" else raw
              if not view.strip():
                  continue
              (dest / f"{n:02d}-{p.stem}.md").write_text(view, encoding="utf-8")
              written += 1
          if not written:
              sys.stderr.write("prose-view: no prose in any of the files given\n")
              return 2
          print(f"prose-view: {written} chapter(s) in {dest}")
          return 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • quote-fidelity.mjs 1.9 KB · in bundle
    • venue-topic-fit.py 13.4 KB
      #!/usr/bin/env python3
      """Where a manuscript's topic and contribution type sit among the papers a journal publishes.
      
          python3 venue-topic-fit.py fetch --issn <ISSN> [--from 2024-01-01] --out <corpus.json> [--dry-run]
          python3 venue-topic-fit.py neighbors --corpus <corpus.json> --title <text> [--abstract <file>] [--top 15] [--json]
          python3 venue-topic-fit.py sample --corpus <corpus.json> --n 100 --seed <int> --out <sheet.tsv>
          python3 venue-topic-fit.py tally --sheet <sheet.tsv> [--json]
      
      build-venue-baseline.py and the prose audits answer whether a manuscript reads like the venue: sentence length,
      punctuation, structure. They were built from the authors' preprints of accepted papers, which skew towards the
      subfields whose authors post preprints, so they cannot say whether the manuscript is about what the venue publishes,
      or whether it makes the kind of contribution the venue usually takes. This measures those two things on every
      article the venue published in a window, from the registrar's records rather than from preprints.
      
      fetch    Every article with the ISSN from OpenAlex: id, DOI, title, year and, where OpenAlex has it, the abstract.
               Publishers withhold many abstracts; the corpus says how many it has. The request carries the tool's name
               and nothing else: no e-mail address, no mailto parameter. --dry-run prints the first request and stops.
      neighbors  TF-IDF cosine between the manuscript and each article, on titles (every article) and on title plus
               abstract (the articles that have one). It reports the nearest articles and where the manuscript's own
               nearest-neighbour similarity falls among the articles' nearest-neighbour similarities: a low percentile
               means its wording sits at the edge of what the venue publishes. Word overlap, not a judgement of fit.
      sample   A seeded random sample of titles with an empty code column, for a person (or a model, recorded as such)
               to code by contribution type: M method, model or system; E evaluation or analysis of existing systems or
               benchmarks; D dataset or benchmark as the main contribution; S study of people, organisations or science
               that does not evaluate a system; R review. Coding from titles is quick and coarse; say so when citing it.
      tally    Counts per code with Wilson 95% intervals. An uncoded row or an unknown code is an error, not a skip.
      
      What this cannot tell you: whether the manuscript would be accepted. Every article in the corpus was accepted; the
      rejected ones are not in it, so no share computed here is an acceptance rate.
      
      Exit: 0 done; 2 nothing to work on (an empty corpus, a sheet with no rows, an uncoded or unknown code, an argument it
      does not recognise). fetch exits 2 when the venue returns no article.
      """
      import argparse
      import csv
      import datetime as dt
      import json
      import math
      import random
      import re
      import sys
      import time
      import urllib.parse
      import urllib.request
      from collections import Counter, defaultdict
      from pathlib import Path
      
      API = "https://api.openalex.org/works"
      USER_AGENT = "academic-writing-toolkit venue-topic-fit"
      CODES = {"M": "method, model or system", "E": "evaluation or analysis of existing systems",
               "D": "dataset or benchmark", "S": "study of people, organisations or science", "R": "review"}
      STOP = set("""a an the of in on for to with by from and or but as at is are was were be been being this that these
      those it its their our we they which who whom whose what when where how why than then there here into onto over
      under via between among across through about not no can could may might will would should must do does did has have
      had using based use new towards toward approach method study analysis""".split())
      
      
      def die(msg):
          sys.stderr.write(f"venue-topic-fit: {msg}\n")
          sys.exit(2)
      
      
      # ---------------------------------------------------------------- fetch
      
      def abstract_from_index(inv):
          """OpenAlex stores an abstract as {word: [positions]}; rebuild the text in order."""
          words = sorted((p, w) for w, ps in (inv or {}).items() for p in ps)
          return " ".join(w for _, w in words)
      
      
      def fetch(issn, since, out, dry_run=False):
          params = {"filter": f"primary_location.source.issn:{issn},from_publication_date:{since},type:article",
                    "select": "id,doi,title,publication_year,abstract_inverted_index", "per-page": "200", "cursor": "*"}
          works = []
          while True:
              url = API + "?" + urllib.parse.urlencode(params)
              req = urllib.request.Request(url, headers={"User-Agent": USER_AGENT})
              if dry_run:
                  print(json.dumps({"url": url, "headers": dict(req.header_items())}))
                  return 0
              with urllib.request.urlopen(req, timeout=60) as r:
                  d = json.loads(r.read())
              for w in d.get("results") or []:
                  works.append({"id": w.get("id"), "doi": w.get("doi"), "title": w.get("title") or "",
                                "year": w.get("publication_year"),
                                "abstract": abstract_from_index(w.get("abstract_inverted_index"))})
              nxt = (d.get("meta") or {}).get("next_cursor")
              if not nxt or not d.get("results"):
                  break
              params["cursor"] = nxt
              time.sleep(1)
          if not works:
              die(f"OpenAlex returned no article for ISSN {issn} since {since}")
          rec = {"retrieved": dt.datetime.now().astimezone().isoformat(timespec="seconds"), "api": API,
                 "filter": params["filter"], "n": len(works), "with_abstract": sum(1 for w in works if w["abstract"]),
                 "works": works}
          Path(out).write_text(json.dumps(rec, ensure_ascii=False), encoding="utf-8")
          print(f"{len(works)} articles, {rec['with_abstract']} with an abstract -> {out}")
          return 0
      
      
      def load_corpus(path):
          try:
              d = json.loads(Path(path).read_text(encoding="utf-8"))
          except (OSError, ValueError) as e:
              die(f"cannot read the corpus {path}: {e}")
          works = [w for w in (d.get("works") or []) if (w.get("title") or "").strip()]
          if not works:
              die(f"the corpus {path} holds no article with a title")
          return d, works
      
      
      # ---------------------------------------------------------------- TF-IDF
      
      def tokens(text):
          return [w for w in re.findall(r"[a-z][a-z0-9-]+", (text or "").lower()) if w not in STOP and len(w) > 2]
      
      
      def vectors(docs):
          """Sublinear TF-IDF with smoothed IDF, L2-normalised, as {term: weight} per document."""
          tfs = [Counter(tokens(d)) for d in docs]
          df = Counter(t for tf in tfs for t in tf)
          n = len(docs)
          idf = {t: math.log((1 + n) / (1 + c)) + 1 for t, c in df.items()}
          out = []
          for tf in tfs:
              v = {t: (1 + math.log(c)) * idf[t] for t, c in tf.items()}
              norm = math.sqrt(sum(x * x for x in v.values())) or 1.0
              out.append({t: x / norm for t, x in v.items()})
          return out
      
      
      def similarities(vecs, q):
          """Cosine of q with every vector, and each vector's highest cosine with any other (through an inverted index)."""
          index = defaultdict(list)
          for i, v in enumerate(vecs):
              for t, x in v.items():
                  index[t].append((i, x))
          to_q = [0.0] * len(vecs)
          for t, x in q.items():
              for i, y in index.get(t, ()):
                  to_q[i] += x * y
          best = []
          for i, v in enumerate(vecs):
              acc = defaultdict(float)
              for t, x in v.items():
                  for j, y in index[t]:
                      if j != i:
                          acc[j] += x * y
              best.append(max(acc.values(), default=0.0))
          return to_q, best
      
      
      def neighbors(corpus, title, abstract, top):
          _, works = load_corpus(corpus)
          out = {"title": title}
          sets = [("titles", works, lambda w: w["title"], title)]
          with_abs = [w for w in works if (w.get("abstract") or "").strip()]
          if abstract is not None and with_abs:
              sets.append(("title_abstract", with_abs, lambda w: w["title"] + " " + w["abstract"], title + " " + abstract))
          for name, ws, text, query in sets:
              vecs = vectors([text(w) for w in ws] + [query])
              q = vecs.pop()
              to_q, best = similarities(vecs, q)
              ours = max(to_q, default=0.0)
              below = sum(1 for b in best if b < ours)
              order = sorted(range(len(ws)), key=lambda i: -to_q[i])[:top]
              out[name] = {"n": len(ws), "ours_nearest": round(ours, 3),
                           "articles_nearest_median": round(sorted(best)[len(best) // 2], 3),
                           "ours_percentile": round(100 * below / len(best), 1),
                           "top": [{"similarity": round(to_q[i], 3), "year": ws[i].get("year"), "doi": ws[i].get("doi"),
                                    "title": ws[i]["title"]} for i in order]}
          return out
      
      
      # ---------------------------------------------------------------- sample and tally
      
      def sample(corpus, n, seed, out):
          _, works = load_corpus(corpus)
          if n <= 0:
              die("--n must be positive")
          rng = random.Random(seed)
          picks = rng.sample(range(len(works)), min(n, len(works)))
          with open(out, "w", newline="", encoding="utf-8") as fh:
              w = csv.writer(fh, delimiter="\t", quoting=csv.QUOTE_NONE, escapechar="\\")
              w.writerow(["i", "year", "doi", "title", "code"])
              for k, i in enumerate(picks, 1):
                  w.writerow([k, works[i].get("year"), works[i].get("doi") or "", re.sub(r"\s+", " ", works[i]["title"]), ""])
          print(f"{len(picks)} titles (seed {seed}) -> {out}; code each as one of " + ", ".join(CODES))
          return 0
      
      
      def wilson(k, n, z=1.96):
          if n == 0:
              return (0.0, 0.0)
          p = k / n
          d = 1 + z * z / n
          c = p + z * z / (2 * n)
          r = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n))
          return (round((c - r) / d, 3), round((c + r) / d, 3))
      
      
      def tally(sheet):
          try:
              with open(sheet, newline="", encoding="utf-8") as fh:
                  rows = list(csv.DictReader(fh, delimiter="\t", quoting=csv.QUOTE_NONE))
          except OSError as e:
              die(f"cannot read the sheet {sheet}: {e}")
          if not rows:
              die(f"{sheet} has no rows")
          codes = [(r.get("code") or "").strip().rstrip("?") for r in rows]
          bad = [(r.get("i"), c) for r, c in zip(rows, codes) if c not in CODES]
          if bad:
              die(f"{len(bad)} row(s) uncoded or with an unknown code (first: row {bad[0][0]} {bad[0][1]!r}); "
                  f"codes are {', '.join(CODES)}")
          n = len(codes)
          c = Counter(codes)
          doubtful = sum(1 for r in rows if (r.get("code") or "").strip().endswith("?"))
          return {"n": n, "doubtful": doubtful,
                  "codes": {k: {"meaning": CODES[k], "count": c[k], "share": round(c[k] / n, 3), "wilson95": wilson(c[k], n)}
                            for k in CODES}}
      
      
      def main():
          ap = argparse.ArgumentParser(description="Topic and contribution type against a venue's published articles.")
          sub = ap.add_subparsers(dest="cmd")
          f = sub.add_parser("fetch")
          f.add_argument("--issn", required=True)
          f.add_argument("--from", dest="since", default="2024-01-01")
          f.add_argument("--out", required=True)
          f.add_argument("--dry-run", action="store_true")
          nb = sub.add_parser("neighbors")
          nb.add_argument("--corpus", required=True)
          nb.add_argument("--title", required=True)
          nb.add_argument("--abstract")
          nb.add_argument("--top", type=int, default=15)
          nb.add_argument("--json", action="store_true")
          sp = sub.add_parser("sample")
          sp.add_argument("--corpus", required=True)
          sp.add_argument("--n", type=int, default=100)
          sp.add_argument("--seed", type=int, required=True)
          sp.add_argument("--out", required=True)
          tl = sub.add_parser("tally")
          tl.add_argument("--sheet", required=True)
          tl.add_argument("--json", action="store_true")
          a = ap.parse_args()
          if a.cmd == "fetch":
              if not re.fullmatch(r"\d{4}-\d{3}[\dXx]", a.issn):
                  die(f"not an ISSN: {a.issn}")
              return fetch(a.issn, a.since, a.out, a.dry_run)
          if a.cmd == "neighbors":
              abstract = None
              if a.abstract:
                  try:
                      abstract = Path(a.abstract).read_text(encoding="utf-8")
                  except OSError as e:
                      die(f"cannot read the abstract {a.abstract}: {e}")
              if not tokens(a.title):
                  die("the title has no content word to compare")
              out = neighbors(a.corpus, a.title, abstract, a.top)
              if a.json:
                  print(json.dumps(out, ensure_ascii=False, indent=1))
              else:
                  for name in ("titles", "title_abstract"):
                      if name not in out:
                          continue
                      o = out[name]
                      print(f"{name}: {o['n']} articles; nearest {o['ours_nearest']} "
                            f"(articles' own nearest, median {o['articles_nearest_median']}; "
                            f"the manuscript is above {o['ours_percentile']}% of them)")
                      for r in o["top"]:
                          print(f"  {r['similarity']:.3f} {r['year']} {r['title'][:120]}")
                  print("Word overlap only; every article here was accepted, so nothing here is an acceptance rate.")
              return 0
          if a.cmd == "sample":
              return sample(a.corpus, a.n, a.seed, a.out)
          if a.cmd == "tally":
              out = tally(a.sheet)
              if a.json:
                  print(json.dumps(out, ensure_ascii=False, indent=1))
              else:
                  print(f"{out['n']} coded titles ({out['doubtful']} marked doubtful with '?')")
                  for k, v in out["codes"].items():
                      print(f"  {k} {v['meaning']}: {v['count']} ({v['share']:.0%}, 95% {v['wilson95'][0]:.0%}-{v['wilson95'][1]:.0%})")
              return 0
          die("give a command: fetch, neighbors, sample or tally")
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
  • SKILL.md 19.8 KB
    ---
    name: audit
    description: Check thesis chapters for consistency before submission — contradictory numbers, terminology drift, and broken cross-references — and read every rewritten or proposed sentence against the one it replaces before anyone sees the rewrite.
    allowed-tools: Read, Glob, Grep, Bash
    ---
    
    # /audit — Thesis Consistency Audit Skill
    
    ## Running Python helpers
    
    Choose the interpreter before running the examples. For a globally installed
    copy, use the private runtime recorded by its installer. In a checkout or linked
    workspace, use `AWT_PYTHON` when set, otherwise the toolkit's `.venv` (follow the
    skill directory link back to the toolkit): `Scripts/python.exe` on Windows,
    `bin/python` on macOS/Linux. Without that environment, check that `python`
    (Windows) or `python3` (macOS/Linux) actually runs and has the helper's dependencies.
    Replace the example's `python3` with that executable. In PowerShell, prefix a
    quoted executable with `&`; keep commands on one line and quote file paths.
    
    ## A manuscript registered with the writing loop
    
    Run `loop coverage <workspace> --run` first (`experimental/writing-loop/bin/loop`). It
    derives every check's inputs from the workspace (the tracked draft files at HEAD,
    the bibliography, the ledgers, the target venue's corpus) instead of the thesis
    defaults below, runs the ones the latest edits made due, and lists the checks
    that cannot run on this manuscript and why. Its table covers the script checks
    only: categories A, B and C below are read by the model, leave no run record, and
    are listed under the table as such. Report any check it shows as not current,
    missing a prerequisite, not applicable or failed, and never report a check it did
    not run as clean.
    
    ## Purpose
    
    Scan all thesis chapters for internal data consistency issues: contradictory numbers, inconsistent terminology, broken cross-references, and arithmetic errors. This is a pre-submission quality check.
    
    ## Trigger Words
    
    This skill activates on: `audit`, `consistency check`, `check numbers`, `/audit`.
    
    ## Workflow
    
    1. **Scan all chapter files** in the `chapters/` directory using Glob. Read each file to extract quantitative claims, terminology, and cross-references.
    
    2. **Check the following categories:**
    
       **A. Numerical consistency** — read by the model, not checked by a script,
       except for what audit H binds. The three bullets below are what to look for;
       nothing verifies that you looked.
       - The same statistic (e.g., accuracy, sample size, p-value) cited in multiple chapters must have the same value.
       - Percentages in a distribution must sum to 100% (with tolerance of +/-1% for rounding).
       - Counts (e.g., "42 models") must match between chapters.
    
       For numbers that matter, use audit H instead: it binds the printed value to
       the artifact it came from, so the check survives the next edit.
    
       **B. Terminological consistency**
       - The same concept must use the same term throughout. Flag cases where synonyms are used inconsistently (e.g., "structured review" vs "systematic review" for the same concept).
       - Abbreviations must be defined on first use in each chapter.
    
       **C. Cross-reference validity**
       - References to other sections (e.g., "as discussed in Section 3.2") must point to sections that exist.
       - References to tables and figures must match actual table/figure numbers.
       - Forward references ("Chapter 6 will show...") must be fulfilled.
    
       **D. Citation checks — disabled in this release (disclosed gap)**
    
       The deterministic citation tiers previously run here are disabled: measured
       against realistic thesis text they produced false high-severity "phantom
       citation" findings on ordinary parentheticals, missed multi-word
       institutional authors, and flagged the comma form that Cite Them Right
       Harvard mandates. Until the checker meets a measured, disclosed
       false-positive rate, do not run it and do not present citation
       consistency as audited. Reference integrity is still covered by
       `/verify-refs` (BibTeX records) and by the notes-file contract lint.
    
       **E. Claim positioning (deterministic, runs before F and G — positioning
       is not repairable after review; style is)**
    
       ```
       python3 .claude/skills/audit/scripts/audit-claim-positioning.py --base-dir chapters --bib references.bib --json
       ```
    
       (omit `--bib` when the project has no bibliography file). Report every
       issue it returns: `unsourced-keyword` and `bare-novelty` as **High** — a
       field's vocabulary in use without its literature, or a novelty claim in a
       paragraph that shows no search — `uncited-method` and `dangling-entry` as
       **Medium**.
    
       Measured precision, one manuscript, 2026-09-20: the checker returned 11
       issues and 9 were false, all from three mechanisms now fixed — a
       `\keywords{}` block wrapped across source lines so a term matched only its
       own declaration; the bare word "novelty" inside the manuscript's own method
       name, and inside two sentences that *refuse* the claim; and a natbib
       optional argument (`\citep[p.~12]{key}`) that made the citation invisible,
       so a method cited in its own sentence was reported as uncited. After the
       fixes it returns the 2 that reading had already confirmed. That is n=1 and
       not a rate — but it is the reason to read every issue against the source
       before writing it into a report at High, exactly as category D's disabled
       tiers required.
    
       For a requested claim-scope or contribution review, consult
       `references/argument-licence/argument-level-lock.md`. Separate the field
       gap, delivered contribution, observed finding and extrapolation; report
       the six-line claim licence and any unsupported transition with a text or
       evidence anchor. This is an Advisory reading task. Do not create standing
       CSV ledgers or infer scientific validity from a checker result. Existing
       legacy packets can be interpreted with the schema and checks linked in
       `references/argument-licence/README.md`.
    
       **H. Number ledger — is this the artifact's number, and is it scoped?**
    
    ```
    python3 .claude/skills/audit/scripts/audit-number-ledger.py --base-dir . --ledger numbers.tsv --json
    ```
    
    A row binds one printed value to the file it came from:
    `printed`, `in_artifact`, `scope`, `artifact`, `locator`. Two value columns
    because they differ in practice — a figure writing `0.635` against prose
    printing `63.5\%` — and the relation is recorded, never inferred. Report
    `locator-not-in-artifact`, `value-not-in-locator`, `printed-artifact-mismatch`,
    `number-not-in-manuscript` and `scope-missing` as **Critical**;
    `unledgered-number` is a coverage list, not a finding. An empty ledger exits 2.
    
    `scope` is matched literally: pick a token the correct sentences contain, or
    `-`. On a real manuscript the first token tried flagged two sentences that carry
    the scope in other words; `technique` passed both.
    
    **I. Method ledger — is each sentence about what was done still bound to where it was done?**
    
    ```
    python3 .claude/skills/audit/scripts/audit-method-ledger.py --base-dir . --ledger method-ledger.tsv \
        --repo e0=~/dev/experiments@<commit> --full sections/03_data.tex --git . --state <ids.json> --json
    ```
    
    For sentences with no claim, number or citation: methods, data, design, what a
    figure shows. A row (`id loc sentence claim_type pointer evidence verdict note
    checked`) points at code, config or output in a repository pinned at a commit.
    Pointers are read out of the cell's text: `path`, `path:12-20` (pinned code
    only), `path#a.b[*].c` (resolved in the JSON or YAML), `path@"fragment"`,
    `\label{x}`, `commit:<sha>`. Report `sentence-changed`, `pointer-missing`,
    `pointer-by-line`, `no-pointer`, `open-verdict`, `retired-but-present` and
    `row-dropped` as **High**; `note-contradicts-verdict` and `unledgered` as
    **Medium**. `partial` rows are pending, neither pass nor fail. A deleted
    sentence's row is kept as `retired <commit>`, never removed. `--migrate-headings`
    prints a diff that moves run-in `\paragraph` headings out of sentence cells and
    writes nothing. Nothing to examine exits 2.
    
    **G. Claim ledger — does a LaTeX manuscript's claim match its archived source?**
    
       ```
       python3 .claude/skills/audit/scripts/audit-claim-ledger.py --base-dir . --ledger ledger.tsv --json
       ```
    
       For manuscripts kept as `.tex` (which audit F does not read). The ledger is a
       TSV with the columns `claim, cite_key, snippet, source_file, level`: one row
       binds one manuscript sentence to one verbatim snippet in one archived source.
       Report `snippet-not-in-source`, `claim-not-in-manuscript`,
       `key-not-in-claim-sentence`, `source-file-missing` and
       `negative-claim-without-fulltext` as **High**; `unledgered-assertion` as
       **Medium**; `qualifier-dropped` and `unledgered-credit` are prompts, not
       findings. Exit 2 means no ledger row was checked at all — report that as
       **not audited**, never as clean. `--pairs` prints the claim/snippet pairs so
       a reader can judge what the audit does not: whether a claim says more than
       its snippet.
    
       **F. Citation fidelity — does the citing sentence match its source?**
    
       ```
       node .claude/skills/audit/scripts/audit-citation-fidelity.mjs --base-dir . --json
       ```
    
       Exit 2 with `nothing_checked: true` means the audit found no citation under
       `chapters/**/*.md` — report it as **not audited**, never as clean. For a LaTeX
       manuscript, `not_covered` counts the `\cite` commands, distinct keys and
       files this audit does not read — under the whole `--base-dir`, so point it
       at the manuscript directory rather than at a repository that also holds
       archived copies, or the count will be a multiple of the truth.
       `--allow-empty` accepts an empty workspace on purpose. An empty `findings`
       list is only a result when `corpus_files` and `sentences_checked` say
       something was read.
       Report `quote-not-in-source` and `page-mismatch` as **High** — a quoted
       span that is not verbatim in the source's notes or PDF, or a page that the
       source contradicts — and `notes-missing` as **Medium**. `low-overlap` is
       **experimental**: list it under Measurements as a prompt to re-read, never
       as an issue; no false-positive rate has been measured for it yet. State
       the tool's own limit in the report verbatim: it does **not** detect a
       sentence that inverts its source in the source's own words — the failure
       that mattered most on a real manuscript — and that still requires
       reading. Every finding here is a proxy; a finding is a reason to open the
       source, not a verdict.
    
       **G. Prose fingerprint (measurement only)**
    
       With no baseline corpus, report G as **not audited** and say which corpus is
       missing. Do not leave it out of the report: an absent section reads as a
       pass. The bibliography baseline needs the project's *own* reference PDFs
       (`literature/`, twenty or more, the author's own papers excluded):
    
       ```
       python3 .claude/skills/audit/scripts/audit-prose-fingerprint.py --target chapters --baseline literature --exclude '<author-surname>*'
       ```
    
       The tool now checks that precondition itself rather than trusting the
       flag: it measures how much of the target's 4-gram vocabulary each baseline
       document contains, and a document that looks like a draft or a copy of the
       target is named in `baseline_suspect`. When one is found it **withholds
       every percentile and exits 2** — not 1, which means "measured, something is
       out of range". Exclude the file, or pass `--allow-overlap` if the overlap
       is intended; the waiver restores the percentiles but still reports
       `preconditions_checked: false`, because a waiver is not a met precondition.
    
       Two further fields have to be read before the percentiles mean anything.
       `baseline_too_short` names every document that loaded but fell under the
       1500-word floor, so `baseline_documents` plus the skipped, too-short and
       excluded lists account for every candidate in the directory; a count that
       does not close means the corpus is not what you think it is.
       `pipeline_mismatch` is true when the target is read as stripped markup and
       the baseline as printed PDF, which is what the documented invocation above
       does: a percentile then compares two readings, not two documents. When the
       target's own build sits beside it the tool measures both readings and puts
       them in `pipeline_cross_check`. On one real manuscript the word count --
       the denominator of every per-1k rate -- differed by 40% between them and
       sentence-length lag-1 changed sign, so quote the cross-check rather than
       treating the percentiles as exact.
    
       **Two baselines, and a percentile that does not name its own is not a
       result.** The corpus above is the project's bibliography: it answers whether
       the prose sits inside the literature the manuscript argues with. Build the
       other one before drafting, not at polish time:
    
       ```
       python3 .claude/skills/audit/scripts/build-venue-baseline.py --venue "<journal>" --from-year 2021 --out <corpus-dir> --manifest <repo>/venue_manifest.json
       ```
    
       It admits a record only when its DOI resolves to the venue's registered
       `container-title`, writes every candidate's disposition so the manifest's
       arithmetic closes, and exits 2 with `VENUE_CORPUS_TOO_SMALL` rather than
       returning a short corpus that reads like a complete one. Run the fingerprint
       a second time against it, with `--target` pointing at the manuscript's built
       **PDF** — against a venue corpus both sides go through the same extraction,
       which is the one case where `pipeline_mismatch` can be false. Read it once,
       to answer whether the manuscript reads like the venue's genre. Do not edit
       to move a venue percentile; the stop rule is unchanged.
    
       Report the distributions under **Measurements**, never as issues: this is
       Advisory by nature. Out-of-range is the hard signal, a percentile is a
       soft one, and clustering matters more than count. Always report
       `preconditions_checked` beside them: `outliers: []` on an unverified
       baseline says nothing, and reading it as a pass is the failure this scan
       was added for. Method and stop rules: `references/prose-polish-method.md`.
    
       **G2. Rewritten sentences, one at a time (before anyone reads a proposal)**
    
       The fingerprint and structure audits measure whole documents, so a handful
       of rewrites cannot move them, and neither reads a proposal that has not been
       applied. Every proposed rewrite, and every sentence a correction adds, goes
       through this first, as an `id, old, new` TSV (leave `old` empty for an
       added sentence; the file is read without quoting):
    
       ```
       python3 .claude/skills/audit/scripts/audit-sentence-changes.py --pairs <rewrites.tsv> --baseline <venue-corpus-dir>
       ```
    
       For a rewrite it names what was added: three or more words, a comma (when
       punctuation as a whole grew), a colon, semicolon, dash or parenthesis, a
       subordinate clause (counted as gained, so trading "because" for "which"
       counts), an adverb or a modifier (counted net, so a term swap does not),
       two prepositional phrases, a subordinate opener, or a merge of sentences.
       A split is judged as the old sentence against all its pieces. An added
       sentence is flagged for any colon, semicolon, dash or worded parenthesis
       and for density above the venue's 75th percentile. Modifiers are a
       stand-in for a part-of-speech tagger; the report's `limits` line says what
       it cannot see. A correction that stays faithful to its source by piling
       qualifiers onto the old sentence is the usual cause; split it or restructure
       it, and re-run until nothing is flagged or each remaining flag is one you
       can defend. Show the author the rewrites only after that.
    
       Keep the file after the author has read it, with two more columns:
       `verdict` (accepted, or rejected; revised counts as rejected; empty while
       not judged) and `reason`. Run the audit on it again and it sets the flags
       against the verdicts and names the script by its hash. The thresholds were
       fitted on one round; a later round's verdicts, judged by the version frozen
       before that round, are their only test.
    
       Once applied, the writing loop runs the same audit on each commit against
       the last commit at which it flagged nothing, so a round of several commits
       is read as a whole; the run record and the summary name that base. A
       flagged sentence keeps the base where it was until the sentence is fixed,
       or until the author accepts it by moving `draft.base_ref` forward.
    
       **G3. Topic and contribution type against the venue (before choosing it)**
    
       The venue corpus of `build-venue-baseline.py` answers whether the
       manuscript reads like the venue. It is built from preprints, which skew to
       the subfields whose authors post them, so it cannot say whether the venue
       publishes this topic or this kind of contribution. That takes every
       article the venue published in a window, from the registrar:
    
       ```
       python3 .claude/skills/audit/scripts/venue-topic-fit.py fetch --issn <ISSN> --from 2024-01-01 --out <corpus.json>
       python3 .claude/skills/audit/scripts/venue-topic-fit.py neighbors --corpus <corpus.json> --title "<title>" --abstract <abstract.txt>
       python3 .claude/skills/audit/scripts/venue-topic-fit.py sample --corpus <corpus.json> --n 100 --seed <int> --out <sheet.tsv>
       python3 .claude/skills/audit/scripts/venue-topic-fit.py tally --sheet <sheet.tsv>
       ```
    
       `neighbors` gives the nearest articles and where the manuscript's own
       nearest-neighbour similarity falls among the articles': a low percentile
       means its wording sits at the edge of what the venue publishes. Read the
       nearest articles before drawing anything from the number. `sample` draws
       titles to code by contribution type (M method, E evaluation or analysis of
       existing systems, D dataset or benchmark, S study of people or science,
       R review; mark a doubtful code with `?`); `tally` gives shares with 95%
       intervals. The coder is whoever codes the sheet, and the report should
       name them. Three limits go into any report of it: coding from titles is
       coarse; TF-IDF measures shared words, not fit; and every article in the
       corpus was accepted, so nothing here is an acceptance rate.
    
    3. **Output the audit report** using the format below.
    
    ## Output Format
    
    ```
    ## Audit Report -- {YYYY-MM-DD}
    
    ### Summary
    
    - **Critical**: {N} issues (contradictory data)
    - **High**: {N} issues (broken references, missing definitions)
    - **Medium**: {N} issues (terminology inconsistency, minor arithmetic)
    
    ### Issues
    
    | # | Severity | Category | Location | Issue | Current | Expected |
    |---|----------|----------|----------|-------|---------|----------|
    | 1 | Critical | Numerical | Ch3 s3.2, Ch5 s5.4 | Sample size differs | 120 (Ch3) vs 125 (Ch5) | Should be consistent |
    | 2 | High | Cross-ref | Ch4 s4.1 | Ref to "Section 3.7" | Section 3.7 | Section does not exist |
    
    ### Measurements (category G when a baseline exists; category F's experimental low-overlap prompts)
    
    {Per metric: rate, clustering (gap CV), longest gap — with the baseline's
    range and where the manuscript sits. Numbers, not verdicts.}
    
    ### Recommendations
    
    {Grouped by severity, brief notes on how to resolve each issue.}
    ```
    
    ## Severity Levels
    
    - **Critical**: The same quantitative claim has different values in different chapters. This directly undermines thesis credibility.
    - **High**: Broken cross-references, undefined abbreviations on first use, missing table/figure numbers.
    - **Medium**: Inconsistent terminology that does not cause factual error, minor rounding discrepancies within tolerance.
    
    ## Constraints
    
    1. **Never auto-fix.** List all issues for the user to review and decide. The user may choose to fix selectively.
    2. **No emoji** in output.
    3. **Report all instances**, not just the first occurrence. If a statistic appears in 4 chapters with 2 different values, list all 4 locations.
    4. **Be specific** about locations. Provide chapter number, section number, and surrounding context so the user can find the issue quickly.
    5. **Do not flag stylistic issues.** This skill checks data consistency, not prose quality.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related