verifying-claims
Check that a document's claims about code are actually true by reading the prose, the code, and the tests and reporting (or fixing) where they disagree. Use whenever the user wants to verify a README, guide, spec, or docstring still matches the code; whenever they mention documen
Install
npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/ai-and-reasoning/skills/verifying-claims
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
git clone https://github.com/oaustegard/claude-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.
README
verifying-claims
Check that what a document says about code is true — by reading the document, the code, and the tests together and reporting where they disagree.
The reviewer is the agent, not a parser. There is no claim-comment DSL: the prose's meaning is read directly and compared to what the code does and what the tests assert. No shadow copy, because the thing checked is the thing the human reads.
See SKILL.md for the procedure, the verdicts (PASS / FAIL / UNSUPPORTED /
STALE), and the division of labor with TDD.
Quick start
python3 scripts/gather_context.py --doc README.md --src pkg/ --tests tests/
That bundles the document text, the public API surface (ast-parsed, never imported), and the test inventory into one report. The agent then reads the bundle and judges each prose claim against it.
What this is and isn't
- Is: a triggered, semantic review of whether documentation matches reality — run before docs ship, after a refactor, or as a sweep.
- Isn't: a CI merge gate or a test framework. The deterministic behavioral gate is your test suite (TDD). This is the prose layer tests can't reach.
Files
scripts/gather_context.py— deterministic input bundler.references/drift-report-example.md— what a review report looks like.
Skill manifest
verifying-claims
Check that what a document says about code is true, by reading the document, the code, and the tests together and reporting where they disagree.
What changed (v0.1 → v0.2)
v0.1 was a comment-DSL: you hand-wrote <!-- claim: ... --> next to prose and
a script checked the comment against the code. That had a fatal gap — the
comment and the prose were two artifacts stapled together, and only the comment
was checked, while humans read the prose. The prose could lie with a green run.
v0.2 drops the DSL. The reviewer is the agent: it reads the prose's meaning directly and compares it to what the code does and what the tests assert. No shadow copy, because the thing being checked is the thing the human reads. (Existing tools already own the alternatives — Gherkin binds executable scenarios, Lean's Verso transcludes facts into prose, TDD couples code to tests. This fills the remaining slot: free-prose documentation, judged.)
Division of labor — read this first
This skill does NOT gate merges and is NOT a test framework.
- The test suite (TDD/CI) owns the behavioral contract: deterministic, cheap, auditable, gated. A green check is something you can hold CI to.
- This skill owns the prose layer: does the documentation match reality? That needs semantic judgment across artifacts, which is non-deterministic and fallible — so it runs as a triggered review (before docs ship, on request, as a sweep), not as a per-commit gate. "The agent said the docs match" is not a guarantee you gate a merge on; it's a review you act on.
Tests are the anchor. The docs are correct when they agree with what the tests assert about the code. So write/keep good tests first; this skill keeps the prose pinned to them.
Procedure
- Identify the document(s) to check and the code + tests they describe.
- Gather consistent input: run
scripts/gather_context.py --doc DOC --src SRC --tests TESTS. It ast-parses source (no imports, no execution) and bundles the document text, the public API surface, and the test inventory. - Extract the claims the prose makes — every checkable assertion about the code (signatures, behavior, return shapes, defaults, guarantees, examples). Do this by reading; there are no claim markers.
- Judge each claim against the API surface and the tests:
- Does the code actually do what the prose says?
- Is the claim backed by a test, or merely asserted?
- Does it reference something that no longer exists?
- Report drift, ranked by severity, each finding citing the prose claim and the contradicting reality (file/function). Use the verdicts below.
- Optionally fix: rewrite the prose to match reality, and/or flag claims that need a test (an UNSUPPORTED claim is a missing test, not just a doc bug).
Verdicts
- PASS — the prose claim matches the code and is exercised by a test.
- FAIL — the code contradicts the claim (the doc is wrong, or the code regressed and the doc caught it).
- UNSUPPORTED — the claim matches the current code but no test backs it, so nothing protects it from future drift. Surface as a missing test.
- STALE — the claim refers to something removed or renamed.
Invoking
- "Check the README against the code before I publish it."
- "Does
docs/api.mdstill matchpkg/?" - "Sweep the docs for drift after this refactor."
Run it at moments that matter — pre-publish, post-refactor, on a docs PR — not on every commit. The deterministic gate is the test suite; this is the layer tests can't reach.
Honest limits
- Non-deterministic and fallible: a review can miss drift or misjudge. Treat output as a careful review, not a proof.
- Cost/latency: reading three artifacts and reasoning is expensive next to a test run. Don't wire it where a cheap deterministic check belongs.
- It checks prose against code+tests; it does not verify the tests themselves are correct. Garbage tests → confident-but-wrong PASS. TDD discipline upstream still matters.
When NOT to use
- As a CI merge gate (use the test suite).
- To verify behavior (write a test).
- On prose with no factual claims about code (nothing to check).
Files
scripts/gather_context.py— deterministic input bundler (doc + API surface + test inventory), ast-only, no imports.references/drift-report-example.md— what a review report looks like.
Files (claude-skills)
-
references
-
drift-report-example.md 1.8 KB
# Drift report — example What a review produces. This is the report for a small parser package whose `README.md` claims `parse(text)` turns text into records and `Reader(path).read()` streams them, checked against the source and tests via `gather_context.py`. --- **Document:** `README.md` · **Sources:** `pkg/parser.py` · **Tests:** `tests/` | Verdict | Claim (prose) | Reality | |---|---|---| | PASS | `parse(text)` turns text into records | `parse(text, strict=False)` exists; `test_parse_empty` exercises it | | UNSUPPORTED | `Reader(path).read()` streams records | `Reader.read(self, n=10)` exists and matches, but `test_reader_reads` asserts `True` — it never calls `read()`. No test protects this claim. | **Summary:** 1 PASS, 1 UNSUPPORTED, 0 FAIL, 0 STALE. **Recommended actions:** - The `read()` claim is accurate today but unprotected. Add a test that calls `Reader(path).read()` and asserts on its output, so a future change to `read` fails loudly instead of silently invalidating the README. --- Notes on reading this report: - **UNSUPPORTED is the interesting verdict.** A dumb signature check would have marked `read()` green — the signature matches. Reading the *test* shows the claim rests on nothing. That gap is what an agent review adds over a declarative check, and it points at a missing test rather than a doc edit. - **FAIL would mean the doc is wrong now** (e.g., README says `parse` returns a dict but the code returns a list). Those get a prose fix. - **STALE would mean the doc references something gone** (e.g., a removed `parse_strict` function). Those get a prose fix or removal. - The report never claims the *tests* are correct — only that the prose agrees with code+tests as they stand. Bad tests upstream still produce a confident PASS.
-
-
scripts
-
gather_context.py 6.7 KB
#!/usr/bin/env python3 """gather_context.py — assemble the inputs for an agent-driven doc/code/test consistency review. This does the *deterministic* half of verifying-claims: it extracts, from source, the things a reviewer needs to judge whether a document's prose still matches reality — the public API surface and the test inventory — and bundles them with the document text. It makes no judgments. The semantic comparison (does the prose agree with the code and the tests?) is the agent's job. Sources are parsed with `ast`, never imported, so no module top-level code runs. Usage: python3 gather_context.py --doc README.md --src pkg/ --tests tests/ python3 gather_context.py --doc docs/api.md --src a.py --src b.py python3 gather_context.py --doc README.md --src pkg/ --json """ import argparse import ast import json import os import sys def _iter_py(paths): for p in paths: if os.path.isdir(p): for dirpath, dirnames, filenames in os.walk(p): dirnames[:] = [d for d in dirnames if d not in ("__pycache__", ".git") and not d.startswith(".")] for fn in sorted(filenames): if fn.endswith(".py"): yield os.path.join(dirpath, fn) elif p.endswith(".py") and os.path.isfile(p): yield p def _sig(node: ast.FunctionDef) -> str: a = node.args parts = [] posonly = getattr(a, "posonlyargs", []) for arg in posonly: parts.append(arg.arg) if posonly: parts.append("/") defaults = list(a.defaults) pos = list(a.args) n_no_default = len(pos) - len(defaults) for i, arg in enumerate(pos): parts.append(arg.arg if i < n_no_default else f"{arg.arg}=...") if a.vararg: parts.append("*" + a.vararg.arg) elif a.kwonlyargs: parts.append("*") for arg, d in zip(a.kwonlyargs, a.kw_defaults): parts.append(arg.arg if d is None else f"{arg.arg}=...") if a.kwarg: parts.append("**" + a.kwarg.arg) return ", ".join(parts) def _doc1(node) -> str: d = ast.get_docstring(node) return d.strip().splitlines()[0].strip() if d else "" def api_surface(src_paths): out = [] for path in _iter_py(src_paths): try: tree = ast.parse(open(path, encoding="utf-8").read()) except (SyntaxError, UnicodeDecodeError) as e: out.append({"file": path, "error": str(e), "defs": []}) continue defs = [] for node in tree.body: if isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)) and not node.name.startswith("_"): defs.append({"name": node.name, "sig": _sig(node), "doc": _doc1(node)}) elif isinstance(node, ast.ClassDef) and not node.name.startswith("_"): methods = [] for m in node.body: if isinstance(m, (ast.FunctionDef, ast.AsyncFunctionDef)) and ( not m.name.startswith("_") or m.name == "__init__" ): methods.append({"name": f"{node.name}.{m.name}", "sig": _sig(m), "doc": _doc1(m)}) defs.append({"name": node.name, "sig": None, "doc": _doc1(node), "methods": methods}) if defs: out.append({"file": path, "defs": defs}) return out def test_inventory(test_paths): out = [] for path in _iter_py(test_paths): try: tree = ast.parse(open(path, encoding="utf-8").read()) except (SyntaxError, UnicodeDecodeError): continue tests = [] def scan(body, tests=tests): # bind per-iteration list (B023) for node in body: if isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)) and node.name.startswith("test"): n_assert = sum(isinstance(x, ast.Assert) for x in ast.walk(node)) tests.append({"name": node.name, "doc": _doc1(node), "asserts": n_assert}) elif isinstance(node, ast.ClassDef): scan(node.body) scan(tree.body) if tests: out.append({"file": path, "tests": tests}) return out def render_md(doc_path, doc_text, api, tests): L = [f"# Context bundle for `{doc_path}`", ""] L += ["## Document", "", "```markdown", doc_text.rstrip(), "```", ""] L += ["## API surface (from source, not imported)", ""] if not api: L.append("_(no source provided)_") for f in api: L.append(f"### `{f['file']}`") if f.get("error"): L.append(f"- parse error: {f['error']}") continue for d in f["defs"]: if d.get("sig") is None and "methods" in d: L.append(f"- class `{d['name']}`" + (f" — {d['doc']}" if d["doc"] else "")) for m in d["methods"]: L.append(f" - `{m['name']}({m['sig']})`" + (f" — {m['doc']}" if m["doc"] else "")) else: L.append(f"- `{d['name']}({d['sig']})`" + (f" — {d['doc']}" if d["doc"] else "")) L.append("") L += ["## Tests (names + assertion counts, from source)", ""] if not tests: L.append("_(no tests provided)_") for f in tests: L.append(f"### `{f['file']}`") for t in f["tests"]: note = t["doc"] or f"{t['asserts']} assert(s)" L.append(f"- `{t['name']}` — {note}") L.append("") L += ["---", "Reviewer: compare each claim the Document makes against the API surface and the", "Tests above. Report drift as PASS / FAIL / UNSUPPORTED (no test backs the claim) /", "STALE (claim refers to something that no longer exists). The behavioral gate is the", "test suite; this review covers only whether the prose matches reality."] return "\n".join(L) def main(): ap = argparse.ArgumentParser(description="Bundle doc + API surface + test inventory for review.") ap.add_argument("--doc", required=True) ap.add_argument("--src", action="append", default=[], help="source file or dir (repeatable)") ap.add_argument("--tests", action="append", default=[], help="test file or dir (repeatable)") ap.add_argument("--json", action="store_true") args = ap.parse_args() if not os.path.isfile(args.doc): print(f"doc not found: {args.doc}", file=sys.stderr) return 2 doc_text = open(args.doc, encoding="utf-8").read() api = api_surface(args.src) tests = test_inventory(args.tests) if args.json: print(json.dumps({"doc": args.doc, "doc_text": doc_text, "api": api, "tests": tests}, indent=2)) else: print(render_md(args.doc, doc_text, api, tests)) return 0 if __name__ == "__main__": sys.exit(main())
-
-
CHANGELOG.md 1.7 KB
# Changelog ## 0.2.0 — 2026-06-07 Pivot: from a claim-comment DSL to an agent-driven review. The v0.1 model stapled a `<!-- claim: ... -->` comment beside prose and checked the comment against the code. That left the prose — the half humans actually read — unverified, and was out-engineered on both flanks (Gherkin binds executable scenarios; Lean's Verso transcludes facts; TDD couples code to tests). v0.2 fills the remaining slot — free-prose documentation — by having the agent read the prose's meaning and compare it to the code and the tests directly. No DSL, no shadow copy. - Removed: `scripts/verify_claims.py` (the comment parser + signature/ command-output resolvers), `assets/verify-claims.yml`, `assets/ test_verify_claims.py`, `references/example-spec.md`. The CI/pytest forcing functions belonged to the DSL approach; the deterministic gate now lives where it should — the project's own test suite (TDD). - Added: `scripts/gather_context.py` — deterministic input bundler (document + ast-parsed API surface + test inventory), no imports, no execution. - Added verdict `UNSUPPORTED` — claim matches code but no test backs it (a missing test, surfaced as such). - Reframed as a triggered review, not a merge gate; the division of labor with TDD is now explicit in SKILL.md. ## 0.1.0 — 2026-06-06 Initial skill (working title "verso"). Claim-comment DSL with `signature` and `command-output` resolvers, `--watch`, import allowlist, and CI/pytest integration templates. Superseded by 0.2.0. ## [0.2.0] - 2026-06-13 ### Fixed - valid YAML frontmatter — colon-space in description broke parsing (#692) ### Other - verifying-claims v0.2: pivot from claim-DSL to agent-driven doc/code/test review (#691) -
README.md 1.3 KB
# verifying-claims Check that what a document *says* about code is true — by reading the document, the code, and the tests together and reporting where they disagree. The reviewer is the agent, not a parser. There is no claim-comment DSL: the prose's meaning is read directly and compared to what the code does and what the tests assert. No shadow copy, because the thing checked is the thing the human reads. See `SKILL.md` for the procedure, the verdicts (PASS / FAIL / UNSUPPORTED / STALE), and the division of labor with TDD. ## Quick start ``` python3 scripts/gather_context.py --doc README.md --src pkg/ --tests tests/ ``` That bundles the document text, the public API surface (ast-parsed, never imported), and the test inventory into one report. The agent then reads the bundle and judges each prose claim against it. ## What this is and isn't - **Is:** a triggered, semantic review of whether documentation matches reality — run before docs ship, after a refactor, or as a sweep. - **Isn't:** a CI merge gate or a test framework. The deterministic behavioral gate is your test suite (TDD). This is the prose layer tests can't reach. ## Files - `scripts/gather_context.py` — deterministic input bundler. - `references/drift-report-example.md` — what a review report looks like. -
SKILL.md 5.1 KB
--- name: verifying-claims description: Check that a document's claims about code are actually true by reading the prose, the code, and the tests and reporting (or fixing) where they disagree. Use whenever the user wants to verify a README, guide, spec, or docstring still matches the code; whenever they mention documentation drift, doc-code sync, "is this still accurate", stale docs, or keeping docs/tests/code consistent; before publishing or merging a docs change; or as a periodic doc-accuracy sweep. The agent reads the prose's meaning directly — there is no claim-comment DSL to maintain. Pairs with TDD — the test suite is the deterministic behavioral gate, this skill is the semantic prose-vs-reality review. metadata: version: 0.2.0 --- # verifying-claims Check that what a document *says* about code is true, by reading the document, the code, and the tests together and reporting where they disagree. ## What changed (v0.1 → v0.2) v0.1 was a comment-DSL: you hand-wrote `<!-- claim: ... -->` next to prose and a script checked the *comment* against the code. That had a fatal gap — the comment and the prose were two artifacts stapled together, and only the comment was checked, while humans read the prose. The prose could lie with a green run. v0.2 drops the DSL. The reviewer is the agent: it reads the prose's *meaning* directly and compares it to what the code does and what the tests assert. No shadow copy, because the thing being checked is the thing the human reads. (Existing tools already own the alternatives — Gherkin binds executable scenarios, Lean's Verso transcludes facts into prose, TDD couples code to tests. This fills the remaining slot: free-prose documentation, judged.) ## Division of labor — read this first This skill does NOT gate merges and is NOT a test framework. - **The test suite (TDD/CI)** owns the behavioral contract: deterministic, cheap, auditable, gated. A green check is something you can hold CI to. - **This skill** owns the prose layer: does the documentation match reality? That needs semantic judgment across artifacts, which is non-deterministic and fallible — so it runs as a *triggered review* (before docs ship, on request, as a sweep), not as a per-commit gate. "The agent said the docs match" is not a guarantee you gate a merge on; it's a review you act on. Tests are the anchor. The docs are correct when they agree with what the tests assert about the code. So write/keep good tests first; this skill keeps the prose pinned to them. ## Procedure 1. **Identify** the document(s) to check and the code + tests they describe. 2. **Gather** consistent input: run `scripts/gather_context.py --doc DOC --src SRC --tests TESTS`. It ast-parses source (no imports, no execution) and bundles the document text, the public API surface, and the test inventory. 3. **Extract the claims** the prose makes — every checkable assertion about the code (signatures, behavior, return shapes, defaults, guarantees, examples). Do this by reading; there are no claim markers. 4. **Judge each claim** against the API surface and the tests: - Does the code actually do what the prose says? - Is the claim backed by a test, or merely asserted? - Does it reference something that no longer exists? 5. **Report** drift, ranked by severity, each finding citing the prose claim and the contradicting reality (file/function). Use the verdicts below. 6. **Optionally fix**: rewrite the prose to match reality, and/or flag claims that need a test (an UNSUPPORTED claim is a missing test, not just a doc bug). ## Verdicts - **PASS** — the prose claim matches the code and is exercised by a test. - **FAIL** — the code contradicts the claim (the doc is wrong, or the code regressed and the doc caught it). - **UNSUPPORTED** — the claim matches the current code but no test backs it, so nothing protects it from future drift. Surface as a missing test. - **STALE** — the claim refers to something removed or renamed. ## Invoking - "Check the README against the code before I publish it." - "Does `docs/api.md` still match `pkg/`?" - "Sweep the docs for drift after this refactor." Run it at moments that matter — pre-publish, post-refactor, on a docs PR — not on every commit. The deterministic gate is the test suite; this is the layer tests can't reach. ## Honest limits - Non-deterministic and fallible: a review can miss drift or misjudge. Treat output as a careful review, not a proof. - Cost/latency: reading three artifacts and reasoning is expensive next to a test run. Don't wire it where a cheap deterministic check belongs. - It checks prose against code+tests; it does not verify the tests themselves are correct. Garbage tests → confident-but-wrong PASS. TDD discipline upstream still matters. ## When NOT to use - As a CI merge gate (use the test suite). - To verify behavior (write a test). - On prose with no factual claims about code (nothing to check). ## Files - `scripts/gather_context.py` — deterministic input bundler (doc + API surface + test inventory), ast-only, no imports. - `references/drift-report-example.md` — what a review report looks like.
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.