readers
Open a reader panel. Instruction-bound amnesiac sub-agents read the abstract or introduction and report what they carried away, set against the author's intended points. Use when those sections change or the writing loop marks the panel stale.
Install
npx skills add https://github.com/yha9806/academic-writing-toolkit/tree/main/.claude/skills/readers
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install yha9806-academic-writing-toolkit@llmmart
git clone https://github.com/yha9806/academic-writing-toolkit.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole yha9806/academic-writing-toolkit collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
/readers — Reader Panel
What it answers, and what it does not
It answers what a first-time reader believes, remembers and could reuse after reading a part of the manuscript, set against the points the author wants carried (the intent card). It is how "does the contribution come across" becomes something counted rather than asserted.
It does not answer whether the paper is good, whether a sentence changed its meaning (the author judges that), or which single word is hard (word-level agreement with the author was too low to use). An eight-reader panel separates only large differences; do not report a one-round rise or fall as a result. The script prints these limits in every report.
Before running
- Intent card. The author's statement of who the reader is and the two to four points (M1, M2, ...) the reader
should carry away. The writing loop names it in
target.intent_card. If it lives outside the workspace'shuman/folder it is a draft, and the report says so. Do not write it for the author; draft it only when asked, and label it a draft. The template atexperimental/writing-loop/templates/intent-card.mdhas further sections (advantage sentence, narrative order, experiment roles, the scan and its acceptance checks); this skill reads only the reader and the points. - Directed questions (optional, recommended): one per suspected misreading or per point you need confirmed,
as
id<TAB>questionlines. Free-text summaries overstate misreadings; a directed question confirms one. Always consider one about reuse and one about which field the work belongs to. When the draft states few links between sentences, pass--ask-relations: it asks each reader for the two sentences between which they most had to guess how one follows from the other, quoted. Readers asked only what got in their way seldom name a missing link.
Steps
Build the packet. From a loop workspace (records which sentences, at which commit, the panel reads):
python3 .claude/skills/readers/scripts/build-reader-packet.py --workspace <workspace> --out <dir> [--questions q.tsv] --aux <main.aux>Or from any file:
--text <file> [--bib refs.bib]. A.texfile is read as LaTeX; any other file as plain text, where%is a percent sign (--format text|latexoverrides). Citations stay in author-year form; they are never replaced by a placeholder. Give--aux(the compiled draft's.aux) so cross-references show the numbers the page shows; without it they read "(number omitted)" and the prompt tells readers so. A placeholder the page does not have draws readers' complaints: in one panel most of "what got in the way" was about it. Math shows its symbols; a command with no symbol keeps its backslash and the build names it.While the loop is still indexing a commit, the build refuses (exit 2): wait for
loop updateto finish. A--sectionsother than the workspace'starget.readers.sectionsbuilds a targeted comparison: the build says how many of the configured sentences it reads, and the tally does not record it as the readers check's run. To change what the check reads, change the workspace's configuration.Open the readers as sub-agents: two personas (
prompt_R1.txt,prompt_R2.txt) × two models (a small and a larger one) × samples per cell. Eight (two samples) is the floor and shows only large differences; use sixteen (four samples) whenever the question is how many readers carried a given point, since a cell repeated on the same text agrees only moderately with itself. Personas and directed questions can live in the loop workspace (target.readers.personas,target.readers.questions) so every run asks the same thing. Give each sub-agent the prompt file's content verbatim and nothing else. Save each reply unedited as<dir>/outputs/<persona>_<model>_<n>.json, e.g.R1_haiku_1.json. Each reply carries the packet id the prompt names; a reply for another packet, or a copy of another reply, is rejected. Rebuild the packet into a new directory for a new version rather than over an old one.Check the outputs; an incomplete output is not a reading and is named, not repaired:
python3 .claude/skills/readers/scripts/check-reader-output.py --packet <dir>/packet.json --outputs <dir>/outputsJudge the intent points. Two judges, independently: you, and a separate sub-agent given the readers'
rememberand directed answers with the reader names shuffled and the version not named. Each writesreader<TAB>point<TAB>judge<TAB>✓|△|✗|≠rows to<dir>/judgments.tsv. A reader carries a point only when at least two judges wrote ✓; judges who disagree count as not carried, and a pair with one judge is not judged. Give each judge the facts of every point (who or what did it, the number), not only its name:≠is a point said but credited to the wrong thing, and a judge told only the point's name grades it ✓. Mix a set of answers whose grade you know into the judging (one correct, one misattributed, one reversed, one bare number per point is enough), write their grades to<dir>/injected.tsvasreader<TAB>point<TAB>truth, and pass--injected: more than two misses records the panel as a failure. Give each directed question its answer's key phrases in a third column of the questions file (id<TAB>question<TAB>key ‖ key): the build flags any question whose key the first paragraph prints, since a reader answers it by copying. A count you read off the outputs yourself (a misreading, a complaint) goes to--derivedwith its coder; one only you coded is reported as uncoded, because the side that revised the text is not a blind coder of the revision's effect. The tally cannot tell two judge names written by one hand: the second judge must be a separate sub-agent that has not seen the first judge's rows.Tally, and record the run in the workspace:
python3 .claude/skills/readers/scripts/tally-readers.py --packet <dir>/packet.json --outputs <dir>/outputs --judgments <dir>/judgments.tsvThe report lists what each reader said got in the way, verbatim, with a keyword sort into kinds (density, sentences, links, terms, repetition, numbers, placeholders). The sort is for reading, not a count: to compare versions, have a blind coder write
writing:<kind>rows for--derived. A kind, a paragraph re-read, or a paragraph named by--ask-relationsthat three readers share is marked ⚑.To compare two versions, run a full panel on each and pass
--compare-packet/--compare-outputs/--compare-judgments; the report puts a two-sided Fisher p beside each point. Run the current version's panel twice and pass the second as--repeat-outputs/--repeat-judgments: a change no larger than the two runs' spread is reported as inside the noise. Judgeblank_reader.json(written beside the packet, reader idBLANK) like any reader: a point it carries is scored by copying the first paragraph, and the report says so. To check that a rewrite fixed a misreading, ask the same directed question in both versions (the same--questionsline in both builds) and judge it as a point: that is the only comparison of a misreading the tally reports as paired. A question asked in one version only is marked not comparable and gets no p; a free-recall point is compared with a note that not mentioning a misreading is not avoiding it. The report gives qualified readers per model for each panel and says when the two panels differ in make-up.Report to the author in three lines per scale (whole text, then paragraphs): what you want the reader to carry (the intent card), what the readers carried (the tally, counts with their denominators), and what the text added that the author did not intend (misreadings, points the readers took that are not on the card). Mark every machine-produced reading as a draft. The author decides what is a gap.
Counts in the report come from the scripts, not from your own reading of the outputs: how many readers qualified is the
qualified N of Mline ofcheck-reader-output.py, and a panel is a reading of this version only whentally-readers.pysays it recorded it (已记为这一版的读者组). A panel judged with sheets or scripts of your own still ends with both; if the tally did not record it, say the panel was not run on this version. One panel skipped both, reported sixteen qualified readers where there were fifteen, and the loop kept calling the last recorded panel stale.
Where the output goes
report.md beside the packet. Real runs on unpublished work stay in the private workspace; never commit reader
outputs or packets to a public repository.
Fail closed
Each script exits 2 when there is nothing to read, check or tally. A panel with fewer than eight qualified readers, or fewer than two personas or two models, is recorded as a failure rather than as a reading of the version.
Files (academic-writing-toolkit)
-
scripts
-
build-reader-packet.py 32 KB
#!/usr/bin/env python3 """Build what a panel of readers reads: numbered paragraphs, one prompt per persona, and a record of the version. python3 build-reader-packet.py --workspace <loop workspace> --out <dir> [--sections A,I] [--questions q.tsv] [--aux main.aux] python3 build-reader-packet.py --text <file> --out <dir> [--format text|latex] [--bib refs.bib] [--questions q.tsv] [--aux main.aux] With --workspace the paragraphs come from the writing loop's index (the tracked draft at its head, the sections named in the workspace's target.readers.sections unless --sections is given), and packet.json records which sentences, at which commit, against which intent card, so `loop coverage` can tell when a later edit has made the panel's reading stale. An index behind the workspace's branch (an update still running) is refused: the packet would be of the older version. A packet whose --sections differ from the workspace's is a targeted comparison, not the readers check: the build says how many of the configured sentences it reads, and packet.json has no snapshot, so the tally does not record it as the check's run (one such packet read only the abstract and introduction of a panel configured for the whole text, and the loop then called every other sentence "new" and asked for the panel again). With --text any file is split on blank lines. It is read as plain text unless it ends in .tex or --format latex is given: in plain text a % is a percent sign, and read as a LaTeX comment it cut one packet off mid-sentence. LaTeX is made readable, not summarised: a citation becomes the author-year form a reader of the published paper would see (from --bib, or the workspace's inputs.bib), never "[cite]"; a cross-reference shows the number the page shows, read from the compiled --aux ("§3.2", "Figure 2", "Table 4"), and one the .aux does not have becomes "(number omitted)", which the prompt tells readers is the packet's limit, not the manuscript's. It used to become "§x" everywhere, and most readers of one panel spent "what got in the way" on that placeholder. Math shows its symbols (≤, ≥, ·, α, a/b): with only the backslash removed, $x \\le y$ reached the readers as "x le y", and readers reported it as unrendered markup. A math command the packet has no symbol for keeps its backslash and is counted in packet.json (residual_commands). Figures and their descriptions are dropped. Directed questions (--questions: one `id<TAB>question` per line) are asked of every reader after the free questions. A third column may give the answer's key phrases (`key ‖ key`), for the judges: they never reach a reader and do not change the packet id. A question whose key phrase the first paragraph prints verbatim is flagged in packet.json (copyable_questions): a reader can answer it by copying, so it cannot tell who understood. Output in --out: manuscript.txt, prompt_<persona>.txt per persona, packet.json, blank_reader.json. blank_reader.json is a reader who read nothing but the first paragraph and copied it into every answer. Judge it like the others (reader id BLANK): a point it carries can be scored by copying, so readers carrying it is no evidence the text got it across. With a workspace packet, packet.json also measures how much of the introduction's first paragraph repeats the abstract (shared four-word sequences, the longest verbatim run): readers' complaints are only a sign. Exit: 0 written; 2 nothing to read (no paragraph, unreadable input), or an argument it does not recognise. """ import argparse import datetime as dt import difflib import hashlib import json import os import re import sys from collections import Counter from pathlib import Path ROOT = Path(__file__).resolve().parents[4] # AWT_LOOP_ENGINE: the engine copy a mutation run is testing; otherwise the one in this checkout. ENGINE = Path(os.environ.get("AWT_LOOP_ENGINE") or ROOT / "experimental" / "writing-loop" / "engine") if not ENGINE.is_dir(): # A user-scope install has no engine beside it; the installer records the checkout's (references/loop-engine.txt). _rec = Path(__file__).resolve().parent.parent / "references" / "loop-engine.txt" if _rec.is_file(): ENGINE = Path(_rec.read_text(encoding="utf-8").strip()) PERSONAS = { "R1": "a researcher in the manuscript's field whose first language is not English; you read English papers daily", "R2": "a senior reviewer for the target venue whose first language is English", } INSTRUCTIONS = """READER INSTRUCTIONS You are a reader, not an editor. Persona: {persona}. You have NOT read this manuscript before and know nothing about its authors or its history. Ignore any background notes, memory, project files or earlier conversation you may be able to see: read only the text below, as this reader would. Below is part of a manuscript submitted to {venue}. Paragraphs are numbered [P1], [P2], ... Read them in order, once, carrying forward what you have read. For EACH paragraph report: - "believe": one sentence: what you now believe the manuscript is claiming or doing, given everything so far - "expect": one sentence: what you expect to read next - "reread": the first 4-6 words (verbatim) of any sentence in this paragraph you had to go back and re-read - "guessed": words or phrases (verbatim) whose meaning you had to guess After the last paragraph report: - "remember": the three things you will remember from this manuscript, most important first - "why_accept": one sentence: the best reason, if any, to accept this paper - "closest_prior_work": what kind of existing work it most resembles - "reuse": what, if anything, you could apply to your own work after reading this - "writing_got_in_way": anything about how it is written that got in your way (or "nothing") {directed_block}- "outside_knowledge": any knowledge you used that is not in the text (or "none"). Report this last. {refs_note} Output ONLY one JSON object with the keys packet, paragraphs, remember, why_accept, closest_prior_work, reuse, writing_got_in_way{directed_keys}, outside_knowledge, where packet is exactly "{packet_id}" and paragraphs is a list of {{"p": 1, "believe": "...", "expect": "...", "reread": [], "guessed": []}}. No other text. MANUSCRIPT {manuscript} """ # Readers asked only what got in their way named density and qualifiers, rarely a missing link between two sentences, # which is what a draft with almost no linking words leaves them to supply. Asked directly, with a quote, it can be # placed: tally-readers.py finds each quote in the paragraphs and counts readers per paragraph. RELATION_ID = "relation_guessed" RELATION_QUESTION = ('the two sentences, one right after the other, between which you most had to guess how the ' 'second follows from the first: quote the first 4-6 words of each in double quotes (or "none")') def die(msg, code=2): sys.stderr.write(f"build-reader-packet: {msg}\n") sys.exit(code) def sha(b): return hashlib.sha1(b if isinstance(b, bytes) else b.encode("utf-8")).hexdigest() # ------------------------------------------------------------------ bibliography def bib_entries(text): """{key: "Surname et al., 2020"} from a BibTeX file; only author and year are read.""" out = {} for m in re.finditer(r"@\w+\s*\{\s*([^,\s]+)\s*,(.*?)(?=\n@|\Z)", text or "", re.S): key, body = m.group(1), m.group(2) au = re.search(r"\bauthor\s*=\s*[{\"](.*?)[}\"]\s*,?\s*\n", body, re.S | re.I) yr = re.search(r"\byear\s*=\s*[{\"]?(\d{4})", body, re.I) names = [n.strip() for n in re.split(r"\s+and\s+", au.group(1))] if au else [] def surname(n): n = re.sub(r"[{}]", "", n) return n.split(",")[0].strip() if "," in n else (n.split()[-1] if n.split() else n) if not names: who = key elif len(names) == 1: who = surname(names[0]) elif len(names) == 2: who = f"{surname(names[0])} and {surname(names[1])}" else: who = f"{surname(names[0])} et al." out[key] = f"{who}, {yr.group(1)}" if yr else who return out # ------------------------------------------------------------------ LaTeX to reading text CITE = re.compile(r"\\(citet|citep|cite|citeauthor|citeyear)\*?(?:\[[^\]]*\])*\{([^}]*)\}") def balanced_end(t, k): """Index just past the group that opens at t[k] == "{". An escaped brace (\\{ or \\}) is text, not structure: counted, one unpaired \\{ in alt text swallowed the rest of the paragraph.""" depth, k = 1, k + 1 while k < len(t) and depth: if t[k] == "\\": k += 2 continue depth += {"{": 1, "}": -1}.get(t[k], 0) k += 1 return k def drop_command(t, name): """Remove \\name[...]{...} with its optional argument and its whole balanced argument; a figure's alt text (\\Description, whose acmart form takes an optional short text) is for screen readers, and a reader of the page never sees it.""" out, i = [], 0 rx = re.compile(r"\\" + name + r"(?![A-Za-z])\s*(?:\[[^\]]*\])?\s*\{") while True: m = rx.search(t, i) if not m: return "".join(out) + t[i:] out.append(t[i:m.start()]) i = balanced_end(t, m.end() - 1) def drop_env_args(t): """A tabular's column specification (and a tabular*/tabularx width) is typesetting, not text.""" rx = re.compile(r"\\begin\{(tabular\*?|tabularx|array)\}\s*(?:\[[^\]]*\])?") out, i = [], 0 while True: m = rx.search(t, i) if not m: return "".join(out) + t[i:] out.append(t[i:m.start()] + " ") k = m.end() for _ in range(1 if m.group(1) in ("tabular", "array") else 2): while k < len(t) and t[k].isspace(): k += 1 if k < len(t) and t[k] == "{": k = balanced_end(t, k) i = k # \S\ref{..}, \S~\ref{..} (the ~ is a space by then), \ref, \eqref, \autoref, \cref, \Cref. REF = re.compile(r"(\\S\s*)?\\(ref|eqref|autoref|cref|Cref)\{([^}]*)\}") # What \autoref and \cref print before the number, by the label's conventional prefix. REF_KIND = {"sec": "Section", "subsec": "Section", "ssec": "Section", "fig": "Figure", "tab": "Table", "eq": "Equation", "app": "Appendix", "alg": "Algorithm", "lst": "Listing"} OMITTED = "(number omitted)" def aux_labels(path): """{label: number as printed} from a compiled .aux (`\newlabel{key}{{number}{page}...}`).""" try: raw = Path(path).read_text(encoding="utf-8", errors="replace") except OSError as e: die(f"cannot read --aux {path}: {e}") out = {} for m in re.finditer(r"\\newlabel\{([^}]*)\}\{\{((?:[^{}]|\{[^{}]*\})*)\}", raw): num = re.sub(r"\\[a-zA-Z@]+\s*", "", m.group(2)).replace("{", "").replace("}", "").strip() if num: out[m.group(1)] = num return out def reference(m, refs): """One cross-reference as the page shows it; counted in refs (resolved or omitted).""" section, cmd, keys = m.group(1), m.group(2), [k.strip() for k in m.group(3).split(",") if k.strip()] labels = refs.get("labels") or {} nums = [labels.get(k) for k in keys] if not keys or any(n is None for n in nums): refs["omitted"] = refs.get("omitted", 0) + 1 return ("§" if section else "") + OMITTED refs["resolved"] = refs.get("resolved", 0) + 1 shown = ", ".join(nums) if section: return "§" + shown if cmd == "eqref": return f"({shown})" if cmd in ("autoref", "cref", "Cref"): kind = REF_KIND.get(keys[0].split(":")[0].lower()) return f"{kind} {shown}" if kind else shown return shown # What the page prints for a symbol command. Without it a symbol only lost its backslash and reached the readers as a # word ("x le y"). Sizing commands (\left, \big) print nothing of their own. SYMBOLS = { "le": "≤", "leq": "≤", "ge": "≥", "geq": "≥", "ne": "≠", "neq": "≠", "approx": "≈", "sim": "~", "simeq": "≃", "equiv": "≡", "pm": "±", "mp": "∓", "cdot": "·", "times": "×", "div": "÷", "ll": "≪", "gg": "≫", "propto": "∝", "infty": "∞", "to": "→", "rightarrow": "→", "leftarrow": "←", "Rightarrow": "⇒", "Leftrightarrow": "⇔", "in": "∈", "notin": "∉", "subset": "⊂", "subseteq": "⊆", "cup": "∪", "cap": "∩", "emptyset": "∅", "mid": "|", "ldots": "…", "dots": "…", "cdots": "⋯", "sum": "Σ", "prod": "Π", "partial": "∂", "nabla": "∇", "circ": "∘", "prime": "′", "degree": "°", "S": "§", "alpha": "α", "beta": "β", "gamma": "γ", "delta": "δ", "epsilon": "ε", "varepsilon": "ε", "zeta": "ζ", "eta": "η", "theta": "θ", "iota": "ι", "kappa": "κ", "lambda": "λ", "mu": "μ", "nu": "ν", "xi": "ξ", "pi": "π", "rho": "ρ", "sigma": "σ", "tau": "τ", "upsilon": "υ", "phi": "φ", "varphi": "φ", "chi": "χ", "psi": "ψ", "omega": "ω", "Gamma": "Γ", "Delta": "Δ", "Theta": "Θ", "Lambda": "Λ", "Xi": "Ξ", "Pi": "Π", "Sigma": "Σ", "Phi": "Φ", "Psi": "Ψ", "Omega": "Ω", "left": "", "right": "", "big": "", "Big": "", "bigg": "", "Bigg": "", "bigl": "", "bigr": "", "Bigl": "", "Bigr": "", } KEEP = "\x00" # a backslash that must reach the reader: a math command with no symbol here DOLLAR = "\x01" # an escaped \$, which is text and must not open a formula LB, RB = "\x02", "\x03" # the braces of such a command's argument, which stay: "\foo{z}", not "\fooz" def readable(text, bib, unknown, refs=None, residual=None): """What a reader of the typeset page sees, as plain text. Applied to a whole paragraph: an environment or a figure's alt text often spans several indexed sentences, and cleaning each alone leaves its markup behind. residual (a Counter), when given, counts the math commands left with their backslash.""" def math(m): def keep(c): if residual is not None: residual["\\" + c.group(1)] += 1 return KEEP + c.group(1) + (LB + c.group(2)[1:-1] + RB if c.group(2) else "") body = re.sub(r"\\([A-Za-z]+)(\{[^{}]*\})?", keep, m.group(1).strip()) return re.sub(r"[{}\\]", "", body) def cite(m): cmd, keys = m.group(1), [k.strip() for k in m.group(2).split(",") if k.strip()] parts = [] for k in keys: if k not in bib: unknown.add(k) parts.append(bib.get(k, k)) if cmd == "citet": return "; ".join(re.sub(r", (\d{4})$", r" (\1)", p) for p in parts) return "(" + "; ".join(parts) + ")" t = CITE.sub(cite, text.replace("\\$", DOLLAR)) t = drop_command(t, "Description") t = drop_env_args(t) t = re.sub(r"\\(?:input|include|includegraphics)\*?(?:\[[^\]]*\])?\{[^}]*\}", "", t) t = re.sub(r"\\caption\s*(?:\[[^\]]*\])?\s*\{", " Caption: {", t) t = re.sub(r"\\begin\{[^}]*\}(?:\[[^\]]*\])*", " ", t) t = re.sub(r"\\end\{[^}]*\}", " ", t) t = re.sub(r"\\item\[([^\]]*)\]", r"\1", t) t = re.sub(r"~", " ", t) t = REF.sub(lambda m: reference(m, refs if refs is not None else {}), t) t = re.sub(r"\\label\{[^}]*\}", "", t) t = re.sub(r"\\(?:emph|textit|textbf|texttt|textsc|textsf|text|mathrm|mathbf|mathit|mathsf|mathtt|mathcal|mathbb|" r"mathfrak|boldsymbol|operatorname|mbox)\{([^{}]*)\}", r"\1", t) t = re.sub(r"\\[dt]?frac\{([^{}]*)\}\{([^{}]*)\}", r"\1/\2", t) t = re.sub(r"\\sqrt\{([^{}]*)\}", lambda m: "√" + (m.group(1) if len(m.group(1)) == 1 else f"({m.group(1)})"), t) t = re.sub(r"\\(hat|bar|tilde|vec|dot)\{([^{}]*)\}", lambda m: m.group(2) + {"hat": "\u0302", "bar": "\u0304", "tilde": "\u0303", "vec": "\u20d7", "dot": "\u0307"}[m.group(1)], t) t = re.sub(r"\\([A-Za-z]+)", lambda m: SYMBOLS.get(m.group(1), m.group(0)), t) t = re.sub(r"\\%", "%", t) t = re.sub(r"\\,|\\;|\\!", " ", t) t = re.sub(r"\{=\}", "=", t) t = re.sub(r"\\\((.*?)\\\)|\\\[(.*?)\\\]", lambda m: "$" + (m.group(1) or m.group(2) or "") + "$", t, flags=re.S) t = re.sub(r"\$([^$]*)\$", math, t) t = re.sub(r"\\[a-zA-Z]+\*?", "", t) t = re.sub(r"[{}]", "", t) t = re.sub(r"---", "—", t).replace("--", "–") t = re.sub(r"\s+", " ", t).strip() return t.replace(KEEP, "\\").replace(LB, "{").replace(RB, "}").replace(DOLLAR, "$") # ------------------------------------------------------------------ sources of paragraphs def ref_commit(cfg): """The commit the workspace's branch is at now, or None when it cannot be read.""" if not cfg.get("repo") or not cfg.get("ref"): return None import subprocess r = subprocess.run(["git", "-C", str(cfg["repo"]), "rev-parse", "--verify", "--quiet", f"{cfg['ref']}^{{commit}}"], capture_output=True, text=True) return r.stdout.strip() or None def from_workspace(ws, sections_arg): if not ENGINE.is_dir(): die(f"--workspace needs the writing loop engine at {ENGINE}; this copy of the skill does not have it") sys.path.insert(0, str(ENGINE)) from loop import catalogue as K # noqa: E402 from loop import config as C # noqa: E402 from loop import coverage as V # noqa: E402 from loop import targets as TG # noqa: E402 try: cfg = C.load(ws) except (OSError, ValueError) as e: die(f"workspace config unreadable: {e}") check = dict(K.by_id("readers")) configured = K.get(cfg, "target.readers.sections") or check["scope"]["default"] sentences, head = V.current_sentences(ws) if sentences is None: die("the workspace has no index yet (run `loop update` first)") now = ref_commit(cfg) if head and now and now != head: # 09-28: a packet built while an update was running read the version before the one just committed. die(f"the index was built at {head[:7]} but {cfg['ref']} is at {now[:7]}: an update is still running or has " "not run; wait for `loop update` to finish, then build the packet") override = bool(sections_arg) and set(sections_arg) != set(configured) if sections_arg: cfg.setdefault("target", {}).setdefault("readers", {})["sections"] = sections_arg prefixes = K.get(cfg, "target.readers.sections") or check["scope"]["default"] kept = [s for s in sentences if V.in_sections(s.get("section"), prefixes)] paras, order = {}, [] for s in kept: key = (s.get("section"), s.get("par")) if key not in paras: paras[key] = [] order.append(key) paras[key].append(s) bibtext = "" b = K.get(cfg, "inputs.bib") if b: import subprocess r = subprocess.run(["git", "-C", str(cfg["repo"]), "show", f"{head}:{b}"], capture_output=True) bibtext = r.stdout.decode("utf-8", "replace") if r.returncode == 0 else "" # A packet of other sections is a targeted comparison: without a snapshot the tally does not record it as the # readers check's run, so the loop keeps judging the check by its last panel of the configured sections. snap = None if override else V.snapshot(check, cfg, sentences, head) state, card = TG.intent_card_state(cfg) personas = K.get(cfg, "target.readers.personas") if personas: PERSONAS.clear() PERSONAS.update(personas) questions = K.get(cfg, "target.readers.questions") source = {"workspace": str(Path(ws).resolve()), "commit": head, "sections": prefixes, "questions_file": questions, "format": V.draft_format(cfg), "venue": K.get(cfg, "target.venue"), "intent_card": {"path": card, "state": state, "sha1": sha(Path(card).read_bytes()) if card and Path(card).is_file() else None}} source["paragraph_sections"] = [k[0] for k in order] if override: source["scope_override"] = { "configured": list(configured), "read_sentences": len(kept), "configured_sentences": sum(1 for s in sentences if V.in_sections(s.get("section"), configured))} # When the draft on disk last changed: an .aux compiled before that may number cross-references the draft has moved. # Not the commit time: compiling and then committing is the usual order, and the .aux would always read as a few # seconds older than a commit of the same text (09-28, found on a real workspace). times = [] for pat in (cfg.get("draft") or {}).get("glob") or []: for f in Path(cfg["repo"]).glob(pat): try: times.append(f.stat().st_mtime) except OSError: pass source["draft_changed"] = max(times) if times else None return [[(s["text"], s["sid"], s["hash"]) for s in paras[k]] for k in order], bibtext, source, snap def aux_staleness(aux, source, text_path): """{aux, source} (local times) when the .aux is older than the last change to the draft it numbers, else None: the draft files on disk (workspace) or the file read (--text). 09-27: an .aux compiled at 18:09 numbered a draft changed again before its 18:34 commit, and nothing said so.""" if not aux: return None try: aux_t = Path(aux).stat().st_mtime except OSError: return None src_t = source.get("draft_changed") if src_t is None and text_path: try: src_t = Path(text_path).stat().st_mtime except OSError: src_t = None if src_t is None or aux_t >= src_t: return None fmt = lambda t: dt.datetime.fromtimestamp(t).strftime("%Y-%m-%d %H:%M:%S") # noqa: E731 return {"aux": fmt(aux_t), "source": fmt(src_t)} def text_format(path, fmt): """latex for a .tex file or when asked; plain text otherwise, where a % is a percent sign.""" return fmt or ("latex" if str(path).lower().endswith(".tex") else "text") def from_text(path, fmt): try: raw = Path(path).read_text(encoding="utf-8") except OSError as e: die(f"cannot read {path}: {e}") source = {"text": str(Path(path).resolve()), "sha1": sha(raw), "format": fmt} if fmt == "latex": # LaTeX drops the rest of a line after an unescaped %, so the packet does too; after a figure it is almost # always a percentage that was meant, and in a plain-text file read as LaTeX it cut the packet mid-sentence. source["percent_after_digit"] = [n for n, line in enumerate(raw.splitlines(), 1) if re.search(r"\d%", line)] raw = re.sub(r"(?m)(?<!\\)%.*$", "", raw) raw = re.sub(r"\\begin\{figure\*?\}.*?\\end\{figure\*?\}", "", raw, flags=re.S) blocks = [b.strip() for b in re.split(r"\n\s*\n", raw) if b.strip()] if fmt == "latex": blocks = [b for b in blocks if not re.fullmatch(r"\\[a-zA-Z]+\*?(\{[^}]*\})*", b)] return [[(b, None, sha(b)[:10])] for b in blocks], "", source, None def repetition(rendered, sections): """How much of the introduction's first paragraph repeats the abstract: the share of its four-word sequences that occur in the abstract, and its longest verbatim run of words. None without sections (a --text packet).""" if not sections or len(sections) != len(rendered): return None ab = [r for r, s in zip(rendered, sections) if str(s).upper().startswith("A")] intro = next((r for r, s in zip(rendered, sections) if str(s).upper().startswith("I")), None) if not ab or intro is None: return None a = re.findall(r"[\w'-]+", " ".join(r["text"] for r in ab).lower()) i = re.findall(r"[\w'-]+", intro["text"].lower()) grams = lambda w: {tuple(w[k:k + 4]) for k in range(len(w) - 3)} ig = grams(i) m = difflib.SequenceMatcher(None, a, i, autojunk=False).find_longest_match(0, len(a), 0, len(i)) return {"abstract": [r["p"] for r in ab], "introduction_first": intro["p"], "shared_four_word_share": round(len(ig & grams(a)) / len(ig), 3) if ig else 0.0, "longest_verbatim_words": m.size, "longest_verbatim": " ".join(i[m.b:m.b + m.size])} def blank_reader(rendered, questions, packet_id): """A reader who copies the first paragraph into every answer (reader id BLANK).""" first = rendered[0]["text"] sentences = [s for s in re.split(r"(?<=[.!?])\s+", first) if s.strip()] out = {"reader": "BLANK", "packet": packet_id, "note": "Not a reader: every answer is copied from the first paragraph. Judge it like the others; a point " "it carries can be scored by copying.", "remember": sentences[:3], "why_accept": first, "closest_prior_work": first, "reuse": first} for q in questions: out[q["id"]] = first return out def read_questions(path): out = [] if not path: return out for i, line in enumerate(Path(path).read_text(encoding="utf-8").splitlines(), 1): if not line.strip() or line.startswith("#"): continue qid, _, rest = line.partition("\t") q, _, keys = rest.partition("\t") if not q.strip(): die(f"{path}:{i}: a directed question is `id<TAB>question[<TAB>key ‖ key]`") item = {"id": qid.strip(), "question": q.strip()} if keys.strip(): item["keys"] = [k.strip() for k in keys.split("‖") if k.strip()] out.append(item) return out def main(argv=None): ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) src = ap.add_mutually_exclusive_group(required=True) src.add_argument("--workspace") src.add_argument("--text") ap.add_argument("--out", required=True) ap.add_argument("--sections", help="comma-separated section prefixes (workspace mode)") ap.add_argument("--bib", help="BibTeX file for author-year citations") ap.add_argument("--questions", help="directed questions: id<TAB>question per line") ap.add_argument("--ask-relations", action="store_true", help=f"add the directed question {RELATION_ID}: where the reader had to guess how one sentence " "follows from the one before (tally-readers.py places the quotes in paragraphs)") ap.add_argument("--venue", help="how the prompt names the venue (default: the workspace's target.venue)") ap.add_argument("--aux", help="the compiled .aux, for the numbers cross-references show on the page") ap.add_argument("--format", choices=("text", "latex"), help="how --text is read (default: latex for a .tex file, plain text otherwise)") try: a = ap.parse_args(argv) except SystemExit as e: sys.exit(2 if e.code else 0) if a.workspace: paras, bibtext, source, snap = from_workspace(a.workspace, a.sections.split(",") if a.sections else None) else: paras, bibtext, source, snap = from_text(a.text, text_format(a.text, a.format)) if a.bib: try: bibtext = Path(a.bib).read_text(encoding="utf-8") except OSError as e: die(f"cannot read --bib {a.bib}: {e}") if not paras: die("no paragraph to give the readers: nothing was built") bib, unknown = bib_entries(bibtext), set() refs = {"labels": aux_labels(a.aux) if a.aux else {}, "resolved": 0, "omitted": 0} stale_aux = aux_staleness(a.aux, source, a.text) plain = bool(a.text) and source.get("format") == "text" residual = Counter() rendered = [] for i, para in enumerate(paras, 1): joined = " ".join(t for t, _, _ in para) text = re.sub(r"\s+", " ", joined).strip() if plain else readable(joined, bib, unknown, refs, residual) rendered.append({"p": i, "text": text, "sids": [s for _, s, _ in para if s], "hashes": [h for _, _, h in para]}) manuscript = "\n\n".join(f"[P{r['p']}] {r['text']}" for r in rendered) keyed = read_questions(a.questions or (source.get("questions_file") and str(Path(source["questions_file"]).expanduser()))) # Keys are for the judges: what the readers read, and the packet id, hold the questions without them. questions = [{"id": q["id"], "question": q["question"]} for q in keyed] if a.ask_relations and all(q["id"] != RELATION_ID for q in questions): questions.append({"id": RELATION_ID, "question": RELATION_QUESTION}) first = rendered[0]["text"].lower() copyable = [q["id"] for q in keyed if any(k.lower() in first for k in q.get("keys") or [])] venue = a.venue or source.get("venue") or "a journal" directed_block = "".join(f'- "{q["id"]}": {q["question"]}\n' for q in questions) directed_keys = "".join(f", {q['id']}" for q in questions) out = Path(a.out) out.mkdir(parents=True, exist_ok=True) (out / "manuscript.txt").write_text(manuscript + "\n", encoding="utf-8") # The packet id travels through every reader's output, so an output written for another version of the text # cannot be tallied against this one. packet_id = sha(json.dumps([manuscript, questions, sorted(PERSONAS.items()), venue]))[:12] refs_note = (f'Cross-references this packet has no number for read "{OMITTED}". That is a limit of the packet, not of ' f'the manuscript: the published page shows the number. Do not report it under writing_got_in_way.\n' if refs["omitted"] else "") prompts = {} for pid, persona in PERSONAS.items(): text = INSTRUCTIONS.format(persona=persona, venue=venue, directed_block=directed_block, refs_note=refs_note, directed_keys=directed_keys, manuscript=manuscript, packet_id=packet_id) (out / f"prompt_{pid}.txt").write_text(text, encoding="utf-8") prompts[pid] = {"file": f"prompt_{pid}.txt", "sha1": sha(text)} packet = {"schema": 1, "packet_id": packet_id, "created_at": dt.datetime.now(dt.timezone.utc).isoformat(), "source": source, "paragraphs": rendered, "questions": questions, "personas": PERSONAS, "prompts": prompts, "unknown_citation_keys": sorted(unknown), "snapshot": snap, "residual_commands": dict(residual.most_common()), "references": {"resolved": refs["resolved"], "omitted": refs["omitted"], "aux": str(Path(a.aux).resolve()) if a.aux else None, "aux_older_than_source": stale_aux}, "repetition": repetition(rendered, source.get("paragraph_sections")), "question_keys": {q["id"]: q["keys"] for q in keyed if q.get("keys")}, "copyable_questions": copyable} (out / "packet.json").write_text(json.dumps(packet, ensure_ascii=False, indent=1), encoding="utf-8") (out / "blank_reader.json").write_text(json.dumps(blank_reader(rendered, questions, packet_id), ensure_ascii=False, indent=1), encoding="utf-8") words = sum(len(r["text"].split()) for r in rendered) print(f"packet: {len(rendered)} paragraphs, {words} words, {len(questions)} directed question(s), " f"{len(PERSONAS)} personas -> {out}") over = source.get("scope_override") if over: print(f" --sections {','.join(source['sections'])} is not the workspace's readers scope " f"({','.join(over['configured'])}): this packet reads {over['read_sentences']} of the " f"{over['configured_sentences']} sentences in that scope and will not be recorded as the readers check's " "run; to change what the check reads, change target.readers.sections") if source.get("percent_after_digit"): print(f" read as LaTeX, a % after a number drops the rest of the line (lines " f"{', '.join(map(str, source['percent_after_digit'][:8]))}); if the file is plain text, give --format text") if residual: print(f" math commands with no symbol in the packet, left with their backslash: " f"{', '.join(f'{c} ×{n}' for c, n in residual.most_common(8))}") if unknown: print(f" citation keys not in the bibliography, left as keys: {', '.join(sorted(unknown))}") if copyable: print(f" directed questions the first paragraph answers verbatim (a reader can copy the answer): {', '.join(copyable)}") if stale_aux: print(f" the .aux was compiled {stale_aux['aux']}, before the draft last changed ({stale_aux['source']}): cross-reference " "numbers may be out of date; recompile, then build the packet again") if refs["omitted"]: print(f" cross-references shown as {OMITTED}: {refs['omitted']} of {refs['omitted'] + refs['resolved']}" + ("" if a.aux else " (give --aux, the compiled .aux, for the numbers)")) return 0 if __name__ == "__main__": sys.exit(main()) -
check-reader-output.py 5.6 KB
#!/usr/bin/env python3 """Check every reader's output against the packet it read. A reader whose output is incomplete did not read. python3 check-reader-output.py --packet <dir>/packet.json <reader output .json> [...] python3 check-reader-output.py --packet <dir>/packet.json --outputs <dir of .json> A qualified output is one JSON object with: one entry per paragraph of the packet (p, believe, expect, reread, guessed); remember (a non-empty list); why_accept, closest_prior_work, reuse, writing_got_in_way; an answer to every directed question in the packet; and outside_knowledge, the reader's own report of what it knew beyond the text. The reader is a sub-agent told to forget, not a reader who never knew, so a missing self-report disqualifies. Prints one line per file, then "qualified N of M". --json prints the same as an object. Exit: 0 every output qualified; 1 some did not; 2 none did, or there was nothing to check. """ import argparse import json import sys from pathlib import Path PARA_KEYS = ("believe", "expect", "reread", "guessed") TOP_TEXT = ("why_accept", "closest_prior_work", "reuse", "writing_got_in_way", "outside_knowledge") def die(msg): sys.stderr.write(f"check-reader-output: {msg}\n") sys.exit(2) def parse(text): """The JSON object in a reader's reply; a reply wrapped in a code fence is accepted, prose around it is not.""" t = text.strip() if t.startswith("```"): t = t.strip("`") t = t[t.find("{"):] if "{" in t else t return json.loads(t) def problems(data, packet): out = [] if not isinstance(data, dict): return ["not a JSON object"] if data.get("packet") != packet.get("packet_id"): out.append(f"written for packet {data.get('packet')!r}, not this one ({packet.get('packet_id')!r})") paras = data.get("paragraphs") want = [p["p"] for p in packet["paragraphs"]] if not isinstance(paras, list): out.append("paragraphs missing") else: seen = {} for e in paras: if isinstance(e, dict) and isinstance(e.get("p"), int): seen[e["p"]] = e missing = [p for p in want if p not in seen] if missing: out.append(f"paragraphs {missing[:6]} missing") for p, e in seen.items(): bad = [k for k in PARA_KEYS if k not in e] if bad: out.append(f"P{p} lacks {', '.join(bad)}") for k in ("reread", "guessed"): if k in e and not isinstance(e[k], list): out.append(f"P{p} {k} is not a list") rem = data.get("remember") if isinstance(rem, str) and rem.strip(): # Said as it is, not as "missing": a reader that numbered its points inside one string did answer, in the # wrong shape, and a panel counted by hand around a vague message disagreed with the script's count. out.append("remember is one string, not a list") elif not isinstance(rem, list) or not [x for x in rem if isinstance(x, str) and x.strip()]: out.append("remember missing or empty") for k in TOP_TEXT: if not isinstance(data.get(k), str) or not data[k].strip(): out.append(f"{k} missing") for q in packet.get("questions") or []: if not isinstance(data.get(q["id"]), str) or not data[q["id"]].strip(): out.append(f"directed question {q['id']} unanswered") return out def _normal(x): if isinstance(x, str): return " ".join(x.split()).casefold() if isinstance(x, list): return [_normal(v) for v in x] if isinstance(x, dict): return {k: _normal(v) for k, v in x.items()} return x def duplicate_of(data, name, seen): """An output identical to one already read, up to spacing and case, is a copy, not another reader.""" key = json.dumps(_normal(data), sort_keys=True, ensure_ascii=False) if key in seen: return [f"identical to {seen[key]}"] seen[key] = name return [] def main(argv=None): ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("--packet", required=True) ap.add_argument("--outputs", help="a directory of reader outputs (*.json)") ap.add_argument("files", nargs="*") ap.add_argument("--json", action="store_true") try: a = ap.parse_args(argv) except SystemExit as e: sys.exit(2 if e.code else 0) try: packet = json.loads(Path(a.packet).read_text(encoding="utf-8")) except (OSError, ValueError) as e: die(f"packet unreadable: {e}") if not packet.get("paragraphs"): die("the packet has no paragraphs") files = [Path(f) for f in a.files] if a.outputs: files += sorted(Path(a.outputs).glob("*.json")) if not files: die("no reader output to check: nothing examined is not a pass") rows, seen = [], {} for f in files: try: data = parse(f.read_text(encoding="utf-8")) why = problems(data, packet) + duplicate_of(data, f.name, seen) except (OSError, ValueError) as e: why = [f"unreadable: {e}"] rows.append({"file": f.name, "qualified": not why, "problems": why}) ok = sum(r["qualified"] for r in rows) if a.json: print(json.dumps({"qualified": ok, "checked": len(rows), "outputs": rows}, ensure_ascii=False, indent=1)) else: for r in rows: print(f"{'ok ' if r['qualified'] else 'NO '} {r['file']}" + ("" if r["qualified"] else ": " + "; ".join(r["problems"]))) print(f"qualified {ok} of {len(rows)}") if ok == 0: return 2 return 0 if ok == len(rows) else 1 if __name__ == "__main__": sys.exit(main()) -
tally-readers.py 31.5 KB
#!/usr/bin/env python3 """Tally a reader panel: only what can be counted. Whether a reader carried away an intended point is judged by people and sub-agents (--judgments), never inferred here. python3 tally-readers.py --packet <dir>/packet.json --outputs <dir> [--judgments j.tsv] [--min-readers 8] [--injected truth.tsv] [--derived coded.tsv] [--repeat-outputs D --repeat-judgments J] [--compare-packet P --compare-outputs D --compare-judgments J] [--json] Reader output files are named <persona>_<model>_<n>.json (e.g. R1_haiku_1.json); the name is how the panel's cells are counted. Only outputs that pass check-reader-output.py's rules are tallied; the others are named. --judgments: TSV, one row per judge per reader per intended point: reader<TAB>point<TAB>judge<TAB>verdict, verdict one of ✓ △ ✗ ≠ (or hit / partial / miss / misattributed). ≠ is a point said but credited to the wrong thing (one model's result told as another's): it is not carried, and it is counted apart, because a judge asked only whether a point was mentioned graded it ✓. --injected: TSV reader<TAB>point<TAB>truth for answers whose grade is known (correct, misattributed, reversed, a bare number), judged with the panel under the same reader ids. More than two judge-by-answer misses on it records the panel as a failure: the judges cannot yet be trusted with the real answers. A reader counts as carrying a point only when every judge wrote ✓; judges who disagree count as not carried, which is the conservative reading. Agreement between judges is reported. --compare-*: a second panel on another version. Per point, a two-sided Fisher exact p is reported beside the counts. A single round's rise or fall is not a result: an eight-reader panel separates only large differences. --derived: TSV metric<TAB>reader<TAB>coder<TAB>0|1 for a count read off the outputs (a misreading, a complaint). A metric coded only by `main`, the side that revised the text, is reported as uncoded, not as a count: in one panel the reviser's coding showed a large drop after its own rewrite and a blind coder's showed none. With a blind coder the count is the blind coder's. --repeat-*: a second panel on the same packet. Its spread per point is the panel's own noise: a change between versions no larger than it is reported as inside the noise, whatever its p. Without a repeat the report says the comparison has no noise floor (one panel run twice on one text moved a point by three readers of sixteen). Counts are also given per model (the reader model mattered more than the persona in calibration), and the blank reader's judgments (reader id BLANK, see build-reader-packet.py) mark the points that copying the first paragraph already scores. With a packet built from a loop workspace, the tally is recorded as the readers check's last run, so the loop knows which version of which sections the panel read. A panel smaller than --min-readers, or with fewer than two personas or two models, is recorded as a failure: it is not the panel the method was calibrated on. Writes report.md beside the packet. Exit: 0 tallied; 2 nothing qualified to tally, or an unreadable packet. """ import argparse import datetime as dt import importlib.util import json import os import math import re import sys from collections import Counter from pathlib import Path HERE = Path(__file__).resolve().parent ROOT = HERE.parents[3] # AWT_LOOP_ENGINE: the engine copy a mutation run is testing; otherwise the one in this checkout. ENGINE = Path(os.environ.get("AWT_LOOP_ENGINE") or ROOT / "experimental" / "writing-loop" / "engine") if not ENGINE.is_dir(): # A user-scope install has no engine beside it; the installer records the checkout's (references/loop-engine.txt). _rec = Path(__file__).resolve().parent.parent / "references" / "loop-engine.txt" if _rec.is_file(): ENGINE = Path(_rec.read_text(encoding="utf-8").strip()) # The panel the method was calibrated on: two personas x two models x two samples. --min-readers may raise it. MIN_PANEL = 8 MIN_JUDGES = 2 BLANK = "BLANK" NAME = re.compile(r"^(?P<persona>[A-Za-z0-9]+)_(?P<model>[A-Za-z0-9.\-]+)_(?P<n>\d+)\.json$") HIT = {"✓": "hit", "hit": "hit", "△": "partial", "partial": "partial", "✗": "miss", "miss": "miss", "≠": "misattributed", "misattributed": "misattributed", "归属错": "misattributed"} INJECT_TOLERANCE = 2 REVISER = "main" # The same place, the same kind of complaint: from this many readers it is named with ⚑. Below it a single reader's # taste cannot be told from a property of the text. SAME_PLACE = 3 NOTHING = re.compile(r"^\s*(?:nothing|none|n/?a|no|无|没有)?\s*[.。]?\s*$", re.I) # writing_got_in_way is free text. A closed keyword list sorts it for reading, and the sort is descriptive: a reader # who writes "the logic jumps" and one who writes "hard to see why this follows" may land in different kinds, or # none. The count to compare versions by is a blind coder's, read with --derived as writing:<kind>. Checked against one # real panel: "reading flow" is not a complaint about links, "section numbers" is one about placeholders, "repeated use # of placeholders" is not repetition; the patterns below leave those out. WRITING_KINDS = [ ("density", r"dense|density|packed|qualif|hedg|caveat|stack|nested|parenthe|too much (?:in|per)|overload"), ("sentences", r"long sentence|sentence length|convoluted|run-on|clause|syntax|complex sentence|wordy|verbose"), ("links", r"transition|signpost|abrupt|logical (?:gap|jump|leap|link)|(?:jumps?|leaps?) (?:from|between|to)|" r"how .{0,40}(?:relates?|connects?|follows)|(?:relation|connection|link)s? between|hard to follow the " r"(?:argument|logic|reasoning)"), ("terms", r"jargon|acronym|abbreviat|undefined|terminolog|\bterms?\b|notation|coined|label"), ("repetition", r"repetitive|repetition|redundan|restat|repeats? (?:itself|the same|what)|said (?:twice|again)"), ("numbers", r"numbers?-heavy|(?:many|multiple|several|too many) (?:numbers|statistics|percentages)|statistic|" r"p-values?|q-values?|percentages|hit.counts|decimal"), ("placeholders", r"§x|placeholder|cross-ref|section numbers?|see section"), ] # The directed question build-reader-packet.py --ask-relations adds. Its quotes are located in the packet's paragraphs. RELATION_ID = "relation_guessed" QUOTE = re.compile(r"[\"\u201c]([^\"\u201d]{12,}?)[\"\u201d]") LIMITS = [ "Readers are sub-agents told to ignore what they can see beyond the text; they are not readers who never knew. " "Each reports the outside knowledge it used.", "Eight readers separate only large differences. When this method was calibrated, repeated panels on the same " "text agreed only moderately and the reader model mattered more than the persona; a one-round rise or fall " "after a wording change did not reproduce.", "Whether a sentence changed its meaning is the author's call: in calibration a machine checker missed most " "of the meaning changes planted for it.", "Guessed words are descriptive only: at the level of single words the readers rarely matched where the author " "got stuck. Punctuation is not seen.", "Free-text summaries overstate misreadings; a directed question is the way to confirm one.", "Free recall is zero-sum: three remember lines hold every point a reader takes, so one point rising pushes another " "out. The rank at which a point was recalled is not recorded; a drop in free recall alone, with the directed " "question holding, is not evidence the text got worse.", ] def die(msg): sys.stderr.write(f"tally-readers: {msg}\n") sys.exit(2) def _checker(): spec = importlib.util.spec_from_file_location("check_reader_output", HERE / "check-reader-output.py") mod = importlib.util.module_from_spec(spec) spec.loader.exec_module(mod) return mod def load_panel(packet_path, outputs_dir): chk = _checker() try: packet = json.loads(Path(packet_path).read_text(encoding="utf-8")) except (OSError, ValueError) as e: die(f"packet unreadable: {e}") readers, rejected, seen = [], [], {} for f in sorted(Path(outputs_dir).glob("*.json")): m = NAME.match(f.name) try: data = chk.parse(f.read_text(encoding="utf-8")) why = chk.problems(data, packet) + chk.duplicate_of(data, f.name, seen) except (OSError, ValueError) as e: data, why = None, [f"unreadable: {e}"] if not m: why = why + ["file name is not <persona>_<model>_<n>.json"] if why: rejected.append({"file": f.name, "problems": why}) else: readers.append({"file": f.name, "reader": f.stem, "persona": m["persona"], "model": m["model"], "data": data}) return packet, readers, rejected def load_judgments(path): rows = {} if not path: return None for i, line in enumerate(Path(path).read_text(encoding="utf-8").splitlines(), 1): if not line.strip() or line.startswith("#"): continue parts = line.split("\t") if len(parts) != 4 or parts[3].strip() not in HIT: die(f"{path}:{i}: expected reader<TAB>point<TAB>judge<TAB>verdict (✓ △ ✗ ≠)") reader, point, judge, verdict = (p.strip() for p in parts) rows.setdefault((reader, point), {})[judge] = HIT[verdict] return rows def carried(judgments, readers): """{point: {"carried": n, "judged": n}} and the judges' agreement.""" if judgments is None: return None, None names = {r["reader"] for r in readers} points = sorted({p for (_, p) in judgments}) out, agree, pairs = {}, 0, 0 for p in points: c = j = 0 for r in names: v = judgments.get((r, p)) if not v: continue vals = list(v.values()) if len(vals) < MIN_JUDGES: continue # one judge's reading is not a judgment here; the pair counts as not judged j += 1 pairs += 1 agree += len(set(vals)) == 1 c += all(x == "hit" for x in vals) out[p] = {"carried": c, "judged": j} return out, (agree / pairs if pairs else None) def misattributed(judgments, readers): """{point: readers every judge graded ≠}: said, but credited to the wrong thing.""" if judgments is None: return None names, out = {r["reader"] for r in readers}, {} for (reader, p), v in judgments.items(): vals = list(v.values()) if reader in names and len(vals) >= MIN_JUDGES and all(x == "misattributed" for x in vals): out[p] = out.get(p, 0) + 1 return out def injected_misses(path, judgments): """(misses, cells): each judge's grade of each injected answer against its known grade; a judge that left one ungraded missed it. None when no --injected was given.""" if not path: return None truth = {} for i, line in enumerate(Path(path).read_text(encoding="utf-8").splitlines(), 1): if not line.strip() or line.startswith("#"): continue parts = [x.strip() for x in line.split("\t")] if len(parts) != 3 or parts[2] not in HIT: die(f"{path}:{i}: expected reader<TAB>point<TAB>truth (✓ △ ✗ ≠)") truth[(parts[0], parts[1])] = HIT[parts[2]] if not truth: die(f"{path}: no injected answer: an empty set checks nothing") judges = sorted({j for v in (judgments or {}).values() for j in v}) if not judges: return len(truth), len(truth) misses = sum((judgments or {}).get(pair, {}).get(j) != want for pair, want in truth.items() for j in judges) return misses, len(truth) * len(judges) def derived_metrics(path, readers): """{metric: {"coded_by", "blind", "count"}}: count is the blind coders' readers with a 1, None when only the reviser coded it. A blind coder is anyone but REVISER; with several, a reader counts when every blind coder wrote 1.""" if not path: return None rows = {} names = {r["reader"] for r in readers} for i, line in enumerate(Path(path).read_text(encoding="utf-8").splitlines(), 1): if not line.strip() or line.startswith("#"): continue parts = [x.strip() for x in line.split("\t")] if len(parts) != 4 or parts[3] not in ("0", "1"): die(f"{path}:{i}: expected metric<TAB>reader<TAB>coder<TAB>0|1") metric, reader, coder, v = parts if reader in names: rows.setdefault(metric, {}).setdefault(reader, {})[coder] = v == "1" out = {} for metric, by_reader in rows.items(): coders = sorted({c for v in by_reader.values() for c in v}) blind = [c for c in coders if c != REVISER] count = sum(all(v.get(c, False) for c in blind) for v in by_reader.values() if any(c in v for c in blind)) \ if blind else None out[metric] = {"coded_by": coders, "blind": bool(blind), "count": count} return out def carried_by_model(judgments, readers): """{model: {point: {"carried", "judged"}}}: the same rule as carried(), one model's readers at a time.""" if judgments is None: return None return {m: carried(judgments, [r for r in readers if r["model"] == m])[0] for m in sorted({r["model"] for r in readers})} def blank_carried(judgments): """{point: carried} for the blank reader (reader id BLANK), under the rule readers are held to.""" if judgments is None: return None out = {} for (reader, p), v in judgments.items(): vals = list(v.values()) if reader == BLANK and len(vals) >= MIN_JUDGES: out[p] = all(x == "hit" for x in vals) return out def noise_floor(hits, rhits): """{point: {"carried", "judged", "spread"}}: a repeat panel on the same packet, and the gap between the two runs as a share of readers judged.""" out = {} for p, v in (hits or {}).items(): r = (rhits or {}).get(p) if r and v["judged"] and r["judged"]: out[p] = {**r, "spread": abs(v["carried"] / v["judged"] - r["carried"] / r["judged"])} return out def fisher_two_sided(a, n1, b, n2): """Two-sided Fisher exact p for a of n1 against b of n2.""" k, n = a + b, n1 + n2 def pmf(x): return math.comb(n1, x) * math.comb(n2, k - x) / math.comb(n, k) obs = pmf(a) lo, hi = max(0, k - n2), min(k, n1) return min(1.0, sum(pmf(x) for x in range(lo, hi + 1) if pmf(x) <= obs * (1 + 1e-9))) def tally(packet, readers): paras = [] for para in packet["paragraphs"]: rr = [r for r in readers if any(e.get("p") == para["p"] and e.get("reread") for e in r["data"]["paragraphs"])] guessed = Counter(w.strip() for r in readers for e in r["data"]["paragraphs"] if e.get("p") == para["p"] for w in e.get("guessed") or [] if isinstance(w, str) and w.strip()) paras.append({"p": para["p"], "reread_by": len(rr), "guessed": guessed.most_common(8)}) for p in paras: p["flag"] = p["reread_by"] >= SAME_PLACE directed = {q["id"]: [(r["reader"], r["data"][q["id"]]) for r in readers] for q in packet.get("questions") or []} writing = [(r["reader"], r["data"]["writing_got_in_way"]) for r in readers if not NOTHING.match(str(r["data"]["writing_got_in_way"]))] kinds = {k: sorted({rd for rd, txt in writing if re.search(pat, str(txt), re.I)}) for k, pat in WRITING_KINDS} return {"paragraphs": paras, "directed": directed, "writing": writing, "writing_kinds": {k: {"readers": v, "flag": len(v) >= SAME_PLACE} for k, v in kinds.items() if v}, "relations": relations(packet, directed.get(RELATION_ID)), "remember": [(r["reader"], r["data"]["remember"]) for r in readers], "closest_prior_work": [(r["reader"], r["data"]["closest_prior_work"]) for r in readers], "reuse": [(r["reader"], r["data"]["reuse"]) for r in readers], "outside_knowledge": [(r["reader"], r["data"]["outside_knowledge"]) for r in readers if r["data"]["outside_knowledge"].strip().lower() not in ("none", "none.", "无")]} def _norm(text): return re.sub(r"\s+", " ", re.sub(r"[^\w\s]", " ", str(text).lower())).strip() def relations(packet, answers): """Where readers had to guess how one sentence follows from another: each quote placed in the paragraph that holds it, readers counted once per paragraph. None when the packet did not ask.""" if answers is None: return None paras = [(p["p"], _norm(p["text"])) for p in packet["paragraphs"]] named, where, unplaced = [], {}, [] for reader, answer in answers: if NOTHING.match(str(answer)): continue named.append(reader) hit = set() for q in QUOTE.findall(str(answer)): q = _norm(q) hit |= {p for p, text in paras if q and q in text} if not hit: unplaced.append(reader) for p in hit: where.setdefault(p, []).append(reader) return {"named": named, "unplaced": unplaced, "paragraphs": {p: {"readers": sorted(v), "flag": len(v) >= SAME_PLACE} for p, v in sorted(where.items())}} def models(readers, rejected): """{model: {"qualified", "rejected"}}: a panel's make-up. One round qualified one reader of eight from one model and seven of eight from the other; two panels built that way are not the same instrument.""" out = {} for r in readers: out.setdefault(r["model"], {"qualified": 0, "rejected": 0})["qualified"] += 1 for r in rejected: m = NAME.match(r["file"]) out.setdefault(m["model"] if m else "?", {"qualified": 0, "rejected": 0})["rejected"] += 1 return dict(sorted(out.items())) def _makeup(ms): return "、".join(f"{m} {v['qualified']}/{v['qualified'] + v['rejected']}" for m, v in ms.items()) def compare_kind(point, packet, cpacket): """directed: both versions asked the same question; mismatch: one did not, or asked it differently, so the counts are not of one thing; recall: a free-recall point, where not mentioning a misreading is not avoiding it.""" qa = {q["id"]: q["question"] for q in packet.get("questions") or []} qb = {q["id"]: q["question"] for q in cpacket.get("questions") or []} if point in qa and point in qb and qa[point] == qb[point]: return "directed" return "mismatch" if point in qa or point in qb else "recall" def panel_shape(readers, min_readers): personas, models = {r["persona"] for r in readers}, {r["model"] for r in readers} problems = [] if len(readers) < min_readers: problems.append(f"只有 {len(readers)} 位合格读者,少于 {min_readers}") if len(personas) < 2: problems.append("画像少于 2 种") if len(models) < 2: problems.append("模型少于 2 种") return problems, sorted(personas), sorted(models) def report(packet, readers, rejected, t, hits, agreement, shape, compare, extra=None): extra = extra or {} L = [f"# 读者组 · {len(readers)} 位读者 × {len(packet['paragraphs'])} 段 · 机器草稿", "", f"读的是:{json.dumps(packet.get('source') or {}, ensure_ascii=False)[:300]}", ""] if shape[0]: L += ["**面板不全**:" + ";".join(shape[0]) + "。结果只作描述,不记为这一版已读过。", ""] stale_aux = (packet.get("references") or {}).get("aux_older_than_source") if stale_aux: L += [f"**交叉引用编号可能过期**:出题用的 .aux 编译于 {stale_aux['aux']},早于稿子最后一次改动({stale_aux['source']})。" "读者对编号的抱怨可能来自这里,不是稿子。", ""] if rejected: L += ["不合格、没计入的输出:", *[f"- {r['file']}:{'; '.join(r['problems'])}" for r in rejected], ""] ms = extra.get("models") if ms: L += [f"按模型的合格读者(合格/交回):{_makeup(ms)}", ""] cms = extra.get("compare_models") if ms and cms and {m: v["qualified"] for m, v in ms.items()} != {m: v["qualified"] for m, v in cms.items()}: L += [f"**两组合格读者的模型构成不同**:本版 {_makeup(ms)};对照版 {_makeup(cms)}。读者模型对结果的影响比画像大," "这次对照不同质,版本之间的差可能来自读者组成。", ""] L += ["## 记忆点(判定者判,不是机器判)"] if hits is not None and not any(v["judged"] for v in hits.values()): hits = None L.append(f"未判:每对「读者 × 记忆点」要 {MIN_JUDGES} 位判定者,给的判定不够。这里不报命中。") elif hits is None: L.append("未判:没有给 --judgments。这里不报命中。") else: for p, v in hits.items(): line = f"- {p}:{v['carried']} / {v['judged']} 位读者带走了" floor = (extra.get("floor") or {}).get(p) if floor: line += f";同包重跑 {floor['carried']} / {floor['judged']}(面板自身波动 {floor['spread']:.2f})" if compare and p in compare and compare[p].get("kind") == "mismatch": c = compare[p] line += f";对照版 {c['carried']} / {c['judged']},**不可比**:这道定向题只有一版问了,或两版问法不同" elif compare and p in compare: c = compare[p] line += f";对照版 {c['carried']} / {c['judged']},双侧 Fisher p = {c['p']:.3f}" if c.get("kind") == "directed": line += "(两版同一道定向题)" if c.get("inside_noise") is True: line += ",**在噪声内**(不大于同包重跑的波动)" elif c.get("inside_noise") is False: line += ",超过同包重跑的波动" wrong = (extra.get("misattributed") or {}).get(p) if wrong: line += f";{wrong} 位说到了但归属错(不算带走)" if (extra.get("blank") or {}).get(p): line += ";**空白读者也带走了**:照抄第一段就能得分,不能当作读懂的证据" L.append(line) inj = extra.get("injected") if inj is None: L.append("没有注入集(--injected):判定者没有先在答案已知的答卷上查过。") else: L.append(f"注入集:判定者判错 {inj[0]} / {inj[1]} 格(允许 {INJECT_TOLERANCE})。") if compare and any(c.get("kind") == "recall" for c in compare.values()): L.append("自由回忆的点在两版之间比的是「提到没有」:一处误读在新版没人提,不等于没人误读。要确认一处误读改好了," "在两版上问同一道定向题(配对设计),再比那道题。") if compare and not extra.get("floor"): L.append("没有同包重跑(--repeat-*):看不出版本之间的变化是否大于面板自身的波动。") by_model = extra.get("by_model") or {} if by_model: L.append("按模型:" + ";".join(f"{m} " + "、".join(f"{p} {v['carried']}/{v['judged']}" for p, v in hm.items()) for m, hm in by_model.items() if hm)) L.append(f"判定者一致率:{agreement:.2f}" if agreement is not None else "只有一位判定者:没有一致率") der = extra.get("derived") if der: L += ["", "## 派生指标(从输出里数的)"] for m, v in der.items(): L.append(f"- {m}:{v['count']} 位(盲编:{'、'.join(c for c in v['coded_by'] if c != REVISER)})" if v["blind"] else f"- {m}:未盲编(只有改稿的一方 {REVISER} 编过),不报数") rep = packet.get("repetition") if rep: L += ["", "## 引言第一段与摘要的重复(量的,不是问的)", f"P{rep['introduction_first']} 的四词组有 {rep['shared_four_word_share']:.0%} 也在摘要里;最长逐字重合 " f"{rep['longest_verbatim_words']} 词:「{rep['longest_verbatim']}」"] wr, n = t.get("writing") or [], len(readers) L += ["", f"## 写法挡路(原话):{len(wr)} / {n} 位读者说有"] L += [f"- {r}:{a}" for r, a in wr] if t.get("writing_kinds"): L.append("关键词预分(描述,不是编码;" + f"同一类 {SAME_PLACE} 位以上标 ⚑):" + ";".join( f"{k} {len(v['readers'])} 位" + (" ⚑" if v["flag"] else "") for k, v in t["writing_kinds"].items())) if wr: L.append("要跨版本比较,请盲编:每位读者每一类一行 `writing:<类>\t<读者>\t<编者>\t0|1`,用 --derived 读;类:" + "、".join(k for k, _ in WRITING_KINDS) + "。") rel = t.get("relations") if rel is not None: L += ["", f"## 句间关系要猜(定向问题 {RELATION_ID}):{len(rel['named'])} / {n} 位读者指出"] for p, v in rel["paragraphs"].items(): L.append(f"- P{p}:{len(v['readers'])} 位" + (" ⚑" if v["flag"] else "") + f"({'、'.join(v['readers'])})") if rel["unplaced"]: L.append(f"- 引文在稿里找不到、没能定位:{'、'.join(rel['unplaced'])}") L += ["", "## 定向问题"] for qid, answers in t["directed"].items(): L += [f"- {qid}"] + [f" - {r}:{a}" for r, a in answers] L += ["", "## 最像哪类已有工作(原话)"] + [f"- {r}:{a}" for r, a in t["closest_prior_work"]] L += ["", "## 能带走什么(原话)"] + [f"- {r}:{a}" for r, a in t["reuse"]] hot = [f"P{p['p']}" for p in t["paragraphs"] if p.get("flag")] L += ["", "## 逐段:要重读的读者数与猜着读的词" + (f"({SAME_PLACE} 位以上重读:{'、'.join(hot)})" if hot else "")] for p in t["paragraphs"]: g = ",".join(f"{w}×{n}" for w, n in p["guessed"]) L.append(f"- P{p['p']}:重读 {p['reread_by']} 位" + (" ⚑" if p.get("flag") else "") + (f";猜:{g}" if g else "")) if t["outside_knowledge"]: L += ["", "## 读者自报的文外知识"] + [f"- {r}:{a}" for r, a in t["outside_knowledge"]] L += ["", "## 这个方法的局限", *[f"- {x}" for x in LIMITS]] return "\n".join(L) + "\n" def record(packet, readers, shape, hits): """The readers check's last run in the loop workspace the packet was built from.""" src = packet.get("source") or {} ws = src.get("workspace") if not ws or not packet.get("snapshot"): return None if not ENGINE.is_dir(): die(f"cannot record the run: the writing loop engine is not at {ENGINE}") sys.path.insert(0, str(ENGINE)) from loop import coverage as V # noqa: E402 if shape[0]: verdict, summary = "failed", "面板不全:" + ";".join(shape[0]) elif hits is None or not any(v["judged"] for v in hits.values()): verdict, summary = "findings", f"{len(readers)} 位读者;记忆点未判" else: verdict = "findings" summary = f"{len(readers)} 位读者;" + ",".join(f"{p} {v['carried']}/{v['judged']}" for p, v in hits.items()) card = (src.get("intent_card") or {}).get("state") if card == "draft": summary += "(意图卡是草稿)" elif card == "delegated": summary += "(意图卡:作者授权 Claude 定稿)" rec = {"id": "readers", "commit": src.get("commit"), "at": dt.datetime.now(dt.timezone.utc).isoformat(), "snapshot": packet["snapshot"], "verdict": verdict, "summary": summary, "exit": None, "panel": {"readers": len(readers), "personas": shape[1], "models": shape[2]}} V.save_run(ws, rec) return rec def main(argv=None): ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("--packet", required=True) ap.add_argument("--outputs", required=True) ap.add_argument("--judgments") ap.add_argument("--min-readers", type=int, default=8) ap.add_argument("--injected", help="answers of known grade: reader<TAB>point<TAB>truth") ap.add_argument("--derived", help="counts read off the outputs: metric<TAB>reader<TAB>coder<TAB>0|1") ap.add_argument("--repeat-outputs", help="a second panel's outputs on the same packet") ap.add_argument("--repeat-judgments") ap.add_argument("--compare-packet") ap.add_argument("--compare-outputs") ap.add_argument("--compare-judgments") ap.add_argument("--json", action="store_true") try: a = ap.parse_args(argv) except SystemExit as e: sys.exit(2 if e.code else 0) packet, readers, rejected = load_panel(a.packet, a.outputs) if not readers: die(f"no qualified reader output in {a.outputs} ({len(rejected)} rejected): nothing tallied is not a result") t = tally(packet, readers) judgments = load_judgments(a.judgments) hits, agreement = carried(judgments, readers) floor = None if a.repeat_outputs and hits is not None: _, rreaders, _ = load_panel(a.packet, a.repeat_outputs) floor = noise_floor(hits, carried(load_judgments(a.repeat_judgments), rreaders)[0]) extra = {"floor": floor, "by_model": carried_by_model(judgments, readers), "blank": blank_carried(judgments), "misattributed": misattributed(judgments, readers), "injected": injected_misses(a.injected, judgments), "derived": derived_metrics(a.derived, readers), "models": models(readers, rejected)} compare = None if a.compare_packet and a.compare_outputs and hits is not None: cpacket, creaders, crejected = load_panel(a.compare_packet, a.compare_outputs) extra["compare_models"] = models(creaders, crejected) chits, _ = carried(load_judgments(a.compare_judgments), creaders) compare = {} for p, v in hits.items(): c = (chits or {}).get(p) if c and v["judged"] and c["judged"]: kind = compare_kind(p, packet, cpacket) compare[p] = {**c, "kind": kind, "p": None if kind == "mismatch" else fisher_two_sided(v["carried"], v["judged"], c["carried"], c["judged"])} f = (floor or {}).get(p) delta = abs(v["carried"] / v["judged"] - c["carried"] / c["judged"]) compare[p]["inside_noise"] = (delta <= f["spread"] + 1e-9) if f else None shape = panel_shape(readers, max(a.min_readers, MIN_PANEL)) if extra["injected"] and extra["injected"][0] > INJECT_TOLERANCE: shape[0].append(f"判定者在注入集上判错 {extra['injected'][0]} 格(允许 {INJECT_TOLERANCE})") text = report(packet, readers, rejected, t, hits, agreement, shape, compare, extra) out = Path(a.packet).parent / "report.md" out.write_text(text, encoding="utf-8") rec = record(packet, readers, shape, hits) if a.json: print(json.dumps({"readers": len(readers), "rejected": rejected, "carried": hits, "agreement": agreement, "panel_problems": shape[0], "compare": compare, "tally": t, "noise_floor": floor, "by_model": extra["by_model"], "blank": extra["blank"], "repetition": packet.get("repetition"), "misattributed": extra["misattributed"], "injected": extra["injected"], "derived": extra["derived"], "models": extra["models"], "compare_models": extra.get("compare_models"), "recorded": bool(rec)}, ensure_ascii=False, indent=1)) else: print(text.splitlines()[0]) print(f"report: {out}" + (";已记为这一版的读者组" if rec and rec["verdict"] != "failed" else ";面板不全,已记为失败" if rec else "")) return 0 if __name__ == "__main__": sys.exit(main())
-
-
SKILL.md 9.2 KB
--- name: readers description: Open a reader panel. Instruction-bound amnesiac sub-agents read the abstract or introduction and report what they carried away, set against the author's intended points. Use when those sections change or the writing loop marks the panel stale. allowed-tools: Read, Glob, Grep, Bash, Agent --- # /readers — Reader Panel ## What it answers, and what it does not It answers **what a first-time reader believes, remembers and could reuse** after reading a part of the manuscript, set against the points the author wants carried (the intent card). It is how "does the contribution come across" becomes something counted rather than asserted. It does not answer whether the paper is good, whether a sentence changed its meaning (the author judges that), or which single word is hard (word-level agreement with the author was too low to use). An eight-reader panel separates only large differences; do not report a one-round rise or fall as a result. The script prints these limits in every report. ## Before running 1. **Intent card.** The author's statement of who the reader is and the two to four points (M1, M2, ...) the reader should carry away. The writing loop names it in `target.intent_card`. If it lives outside the workspace's `human/` folder it is a draft, and the report says so. Do not write it for the author; draft it only when asked, and label it a draft. The template at `experimental/writing-loop/templates/intent-card.md` has further sections (advantage sentence, narrative order, experiment roles, the scan and its acceptance checks); this skill reads only the reader and the points. 2. **Directed questions** (optional, recommended): one per suspected misreading or per point you need confirmed, as `id<TAB>question` lines. Free-text summaries overstate misreadings; a directed question confirms one. Always consider one about reuse and one about which field the work belongs to. When the draft states few links between sentences, pass `--ask-relations`: it asks each reader for the two sentences between which they most had to guess how one follows from the other, quoted. Readers asked only what got in their way seldom name a missing link. ## Steps 1. Build the packet. From a loop workspace (records which sentences, at which commit, the panel reads): ``` python3 .claude/skills/readers/scripts/build-reader-packet.py --workspace <workspace> --out <dir> [--questions q.tsv] --aux <main.aux> ``` Or from any file: `--text <file> [--bib refs.bib]`. A `.tex` file is read as LaTeX; any other file as plain text, where `%` is a percent sign (`--format text|latex` overrides). Citations stay in author-year form; they are never replaced by a placeholder. Give `--aux` (the compiled draft's `.aux`) so cross-references show the numbers the page shows; without it they read "(number omitted)" and the prompt tells readers so. A placeholder the page does not have draws readers' complaints: in one panel most of "what got in the way" was about it. Math shows its symbols; a command with no symbol keeps its backslash and the build names it. While the loop is still indexing a commit, the build refuses (exit 2): wait for `loop update` to finish. A `--sections` other than the workspace's `target.readers.sections` builds a targeted comparison: the build says how many of the configured sentences it reads, and the tally does not record it as the readers check's run. To change what the check reads, change the workspace's configuration. 2. Open the readers as sub-agents: two personas (`prompt_R1.txt`, `prompt_R2.txt`) × two models (a small and a larger one) × samples per cell. Eight (two samples) is the floor and shows only large differences; use sixteen (four samples) whenever the question is how many readers carried a given point, since a cell repeated on the same text agrees only moderately with itself. Personas and directed questions can live in the loop workspace (`target.readers.personas`, `target.readers.questions`) so every run asks the same thing. Give each sub-agent the prompt file's content verbatim and nothing else. Save each reply unedited as `<dir>/outputs/<persona>_<model>_<n>.json`, e.g. `R1_haiku_1.json`. Each reply carries the packet id the prompt names; a reply for another packet, or a copy of another reply, is rejected. Rebuild the packet into a new directory for a new version rather than over an old one. 3. Check the outputs; an incomplete output is not a reading and is named, not repaired: ``` python3 .claude/skills/readers/scripts/check-reader-output.py --packet <dir>/packet.json --outputs <dir>/outputs ``` 4. Judge the intent points. Two judges, independently: you, and a separate sub-agent given the readers' `remember` and directed answers with the reader names shuffled and the version not named. Each writes `reader<TAB>point<TAB>judge<TAB>✓|△|✗|≠` rows to `<dir>/judgments.tsv`. A reader carries a point only when at least two judges wrote ✓; judges who disagree count as not carried, and a pair with one judge is not judged. Give each judge the facts of every point (who or what did it, the number), not only its name: `≠` is a point said but credited to the wrong thing, and a judge told only the point's name grades it ✓. Mix a set of answers whose grade you know into the judging (one correct, one misattributed, one reversed, one bare number per point is enough), write their grades to `<dir>/injected.tsv` as `reader<TAB>point<TAB>truth`, and pass `--injected`: more than two misses records the panel as a failure. Give each directed question its answer's key phrases in a third column of the questions file (`id<TAB>question<TAB>key ‖ key`): the build flags any question whose key the first paragraph prints, since a reader answers it by copying. A count you read off the outputs yourself (a misreading, a complaint) goes to `--derived` with its coder; one only you coded is reported as uncoded, because the side that revised the text is not a blind coder of the revision's effect. The tally cannot tell two judge names written by one hand: the second judge must be a separate sub-agent that has not seen the first judge's rows. 5. Tally, and record the run in the workspace: ``` python3 .claude/skills/readers/scripts/tally-readers.py --packet <dir>/packet.json --outputs <dir>/outputs --judgments <dir>/judgments.tsv ``` The report lists what each reader said got in the way, verbatim, with a keyword sort into kinds (density, sentences, links, terms, repetition, numbers, placeholders). The sort is for reading, not a count: to compare versions, have a blind coder write `writing:<kind>` rows for `--derived`. A kind, a paragraph re-read, or a paragraph named by `--ask-relations` that three readers share is marked ⚑. To compare two versions, run a full panel on each and pass `--compare-packet/--compare-outputs/--compare-judgments`; the report puts a two-sided Fisher p beside each point. Run the current version's panel twice and pass the second as `--repeat-outputs/--repeat-judgments`: a change no larger than the two runs' spread is reported as inside the noise. Judge `blank_reader.json` (written beside the packet, reader id `BLANK`) like any reader: a point it carries is scored by copying the first paragraph, and the report says so. To check that a rewrite fixed a misreading, ask the same directed question in both versions (the same `--questions` line in both builds) and judge it as a point: that is the only comparison of a misreading the tally reports as paired. A question asked in one version only is marked not comparable and gets no p; a free-recall point is compared with a note that not mentioning a misreading is not avoiding it. The report gives qualified readers per model for each panel and says when the two panels differ in make-up. 6. Report to the author in three lines per scale (whole text, then paragraphs): what you want the reader to carry (the intent card), what the readers carried (the tally, counts with their denominators), and what the text added that the author did not intend (misreadings, points the readers took that are not on the card). Mark every machine-produced reading as a draft. The author decides what is a gap. Counts in the report come from the scripts, not from your own reading of the outputs: how many readers qualified is the `qualified N of M` line of `check-reader-output.py`, and a panel is a reading of this version only when `tally-readers.py` says it recorded it (`已记为这一版的读者组`). A panel judged with sheets or scripts of your own still ends with both; if the tally did not record it, say the panel was not run on this version. One panel skipped both, reported sixteen qualified readers where there were fifteen, and the loop kept calling the last recorded panel stale. ## Where the output goes `report.md` beside the packet. Real runs on unpublished work stay in the private workspace; never commit reader outputs or packets to a public repository. ## Fail closed Each script exits 2 when there is nothing to read, check or tally. A panel with fewer than eight qualified readers, or fewer than two personas or two models, is recorded as a failure rather than as a reading of the version.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.