Claude Skill

verifying-claims

Check that a document's claims about code are actually true by reading the prose, the code, and the tests and reporting (or fixing) where they disagree. Use whenever the user wants to verify a README, guide, spec, or docstring still matches the code; whenever they mention documen

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download oaustegard-claude-skills-plugins_ai-and-reasoning_skills_verifying-claims-e39c726.zip · 7 KB
Part of oaustegard/claude-skills — 39 skills

Install

skills CLI npx skills add https://github.com/oaustegard/claude-skills/tree/main/plugins/ai-and-reasoning/skills/verifying-claims
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install oaustegard-claude-skills@llmmart
Git git clone https://github.com/oaustegard/claude-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole oaustegard/claude-skills collection as a plugin from our marketplace. Git is the plain clone.

README

verifying-claims

Check that what a document says about code is true — by reading the document, the code, and the tests together and reporting where they disagree.

The reviewer is the agent, not a parser. There is no claim-comment DSL: the prose's meaning is read directly and compared to what the code does and what the tests assert. No shadow copy, because the thing checked is the thing the human reads.

See SKILL.md for the procedure, the verdicts (PASS / FAIL / UNSUPPORTED / STALE), and the division of labor with TDD.

Quick start

python3 scripts/gather_context.py --doc README.md --src pkg/ --tests tests/

That bundles the document text, the public API surface (ast-parsed, never imported), and the test inventory into one report. The agent then reads the bundle and judges each prose claim against it.

What this is and isn't

  • Is: a triggered, semantic review of whether documentation matches reality — run before docs ship, after a refactor, or as a sweep.
  • Isn't: a CI merge gate or a test framework. The deterministic behavioral gate is your test suite (TDD). This is the prose layer tests can't reach.

Files

  • scripts/gather_context.py — deterministic input bundler.
  • references/drift-report-example.md — what a review report looks like.

Skill manifest

verifying-claims

Check that what a document says about code is true, by reading the document, the code, and the tests together and reporting where they disagree.

What changed (v0.1 → v0.2)

v0.1 was a comment-DSL: you hand-wrote <!-- claim: ... --> next to prose and a script checked the comment against the code. That had a fatal gap — the comment and the prose were two artifacts stapled together, and only the comment was checked, while humans read the prose. The prose could lie with a green run.

v0.2 drops the DSL. The reviewer is the agent: it reads the prose's meaning directly and compares it to what the code does and what the tests assert. No shadow copy, because the thing being checked is the thing the human reads. (Existing tools already own the alternatives — Gherkin binds executable scenarios, Lean's Verso transcludes facts into prose, TDD couples code to tests. This fills the remaining slot: free-prose documentation, judged.)

Division of labor — read this first

This skill does NOT gate merges and is NOT a test framework.

  • The test suite (TDD/CI) owns the behavioral contract: deterministic, cheap, auditable, gated. A green check is something you can hold CI to.
  • This skill owns the prose layer: does the documentation match reality? That needs semantic judgment across artifacts, which is non-deterministic and fallible — so it runs as a triggered review (before docs ship, on request, as a sweep), not as a per-commit gate. "The agent said the docs match" is not a guarantee you gate a merge on; it's a review you act on.

Tests are the anchor. The docs are correct when they agree with what the tests assert about the code. So write/keep good tests first; this skill keeps the prose pinned to them.

Procedure

  1. Identify the document(s) to check and the code + tests they describe.
  2. Gather consistent input: run scripts/gather_context.py --doc DOC --src SRC --tests TESTS. It ast-parses source (no imports, no execution) and bundles the document text, the public API surface, and the test inventory.
  3. Extract the claims the prose makes — every checkable assertion about the code (signatures, behavior, return shapes, defaults, guarantees, examples). Do this by reading; there are no claim markers.
  4. Judge each claim against the API surface and the tests:
    • Does the code actually do what the prose says?
    • Is the claim backed by a test, or merely asserted?
    • Does it reference something that no longer exists?
  5. Report drift, ranked by severity, each finding citing the prose claim and the contradicting reality (file/function). Use the verdicts below.
  6. Optionally fix: rewrite the prose to match reality, and/or flag claims that need a test (an UNSUPPORTED claim is a missing test, not just a doc bug).

Verdicts

  • PASS — the prose claim matches the code and is exercised by a test.
  • FAIL — the code contradicts the claim (the doc is wrong, or the code regressed and the doc caught it).
  • UNSUPPORTED — the claim matches the current code but no test backs it, so nothing protects it from future drift. Surface as a missing test.
  • STALE — the claim refers to something removed or renamed.

Invoking

  • "Check the README against the code before I publish it."
  • "Does docs/api.md still match pkg/?"
  • "Sweep the docs for drift after this refactor."

Run it at moments that matter — pre-publish, post-refactor, on a docs PR — not on every commit. The deterministic gate is the test suite; this is the layer tests can't reach.

Honest limits

  • Non-deterministic and fallible: a review can miss drift or misjudge. Treat output as a careful review, not a proof.
  • Cost/latency: reading three artifacts and reasoning is expensive next to a test run. Don't wire it where a cheap deterministic check belongs.
  • It checks prose against code+tests; it does not verify the tests themselves are correct. Garbage tests → confident-but-wrong PASS. TDD discipline upstream still matters.

When NOT to use

  • As a CI merge gate (use the test suite).
  • To verify behavior (write a test).
  • On prose with no factual claims about code (nothing to check).

Files

  • scripts/gather_context.py — deterministic input bundler (doc + API surface + test inventory), ast-only, no imports.
  • references/drift-report-example.md — what a review report looks like.
Files (claude-skills)
  • references
    • drift-report-example.md 1.8 KB
      # Drift report — example
      
      What a review produces. This is the report for a small parser package whose
      `README.md` claims `parse(text)` turns text into records and `Reader(path).read()`
      streams them, checked against the source and tests via `gather_context.py`.
      
      ---
      
      **Document:** `README.md`  ·  **Sources:** `pkg/parser.py`  ·  **Tests:** `tests/`
      
      | Verdict | Claim (prose) | Reality |
      |---|---|---|
      | PASS | `parse(text)` turns text into records | `parse(text, strict=False)` exists; `test_parse_empty` exercises it |
      | UNSUPPORTED | `Reader(path).read()` streams records | `Reader.read(self, n=10)` exists and matches, but `test_reader_reads` asserts `True` — it never calls `read()`. No test protects this claim. |
      
      **Summary:** 1 PASS, 1 UNSUPPORTED, 0 FAIL, 0 STALE.
      
      **Recommended actions:**
      - The `read()` claim is accurate today but unprotected. Add a test that calls
        `Reader(path).read()` and asserts on its output, so a future change to `read`
        fails loudly instead of silently invalidating the README.
      
      ---
      
      Notes on reading this report:
      
      - **UNSUPPORTED is the interesting verdict.** A dumb signature check would have
        marked `read()` green — the signature matches. Reading the *test* shows the
        claim rests on nothing. That gap is what an agent review adds over a
        declarative check, and it points at a missing test rather than a doc edit.
      - **FAIL would mean the doc is wrong now** (e.g., README says `parse` returns a
        dict but the code returns a list). Those get a prose fix.
      - **STALE would mean the doc references something gone** (e.g., a removed
        `parse_strict` function). Those get a prose fix or removal.
      - The report never claims the *tests* are correct — only that the prose agrees
        with code+tests as they stand. Bad tests upstream still produce a confident
        PASS.
      
  • scripts
    • gather_context.py 6.7 KB
      #!/usr/bin/env python3
      """gather_context.py — assemble the inputs for an agent-driven doc/code/test
      consistency review.
      
      This does the *deterministic* half of verifying-claims: it extracts, from
      source, the things a reviewer needs to judge whether a document's prose still
      matches reality — the public API surface and the test inventory — and bundles
      them with the document text. It makes no judgments. The semantic comparison
      (does the prose agree with the code and the tests?) is the agent's job.
      
      Sources are parsed with `ast`, never imported, so no module top-level code runs.
      
      Usage:
          python3 gather_context.py --doc README.md --src pkg/ --tests tests/
          python3 gather_context.py --doc docs/api.md --src a.py --src b.py
          python3 gather_context.py --doc README.md --src pkg/ --json
      """
      
      import argparse
      import ast
      import json
      import os
      import sys
      
      
      def _iter_py(paths):
          for p in paths:
              if os.path.isdir(p):
                  for dirpath, dirnames, filenames in os.walk(p):
                      dirnames[:] = [d for d in dirnames
                                     if d not in ("__pycache__", ".git") and not d.startswith(".")]
                      for fn in sorted(filenames):
                          if fn.endswith(".py"):
                              yield os.path.join(dirpath, fn)
              elif p.endswith(".py") and os.path.isfile(p):
                  yield p
      
      
      def _sig(node: ast.FunctionDef) -> str:
          a = node.args
          parts = []
          posonly = getattr(a, "posonlyargs", [])
          for arg in posonly:
              parts.append(arg.arg)
          if posonly:
              parts.append("/")
          defaults = list(a.defaults)
          pos = list(a.args)
          n_no_default = len(pos) - len(defaults)
          for i, arg in enumerate(pos):
              parts.append(arg.arg if i < n_no_default else f"{arg.arg}=...")
          if a.vararg:
              parts.append("*" + a.vararg.arg)
          elif a.kwonlyargs:
              parts.append("*")
          for arg, d in zip(a.kwonlyargs, a.kw_defaults):
              parts.append(arg.arg if d is None else f"{arg.arg}=...")
          if a.kwarg:
              parts.append("**" + a.kwarg.arg)
          return ", ".join(parts)
      
      
      def _doc1(node) -> str:
          d = ast.get_docstring(node)
          return d.strip().splitlines()[0].strip() if d else ""
      
      
      def api_surface(src_paths):
          out = []
          for path in _iter_py(src_paths):
              try:
                  tree = ast.parse(open(path, encoding="utf-8").read())
              except (SyntaxError, UnicodeDecodeError) as e:
                  out.append({"file": path, "error": str(e), "defs": []})
                  continue
              defs = []
              for node in tree.body:
                  if isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)) and not node.name.startswith("_"):
                      defs.append({"name": node.name, "sig": _sig(node), "doc": _doc1(node)})
                  elif isinstance(node, ast.ClassDef) and not node.name.startswith("_"):
                      methods = []
                      for m in node.body:
                          if isinstance(m, (ast.FunctionDef, ast.AsyncFunctionDef)) and (
                              not m.name.startswith("_") or m.name == "__init__"
                          ):
                              methods.append({"name": f"{node.name}.{m.name}", "sig": _sig(m), "doc": _doc1(m)})
                      defs.append({"name": node.name, "sig": None, "doc": _doc1(node), "methods": methods})
              if defs:
                  out.append({"file": path, "defs": defs})
          return out
      
      
      def test_inventory(test_paths):
          out = []
          for path in _iter_py(test_paths):
              try:
                  tree = ast.parse(open(path, encoding="utf-8").read())
              except (SyntaxError, UnicodeDecodeError):
                  continue
              tests = []
      
              def scan(body, tests=tests):  # bind per-iteration list (B023)
                  for node in body:
                      if isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)) and node.name.startswith("test"):
                          n_assert = sum(isinstance(x, ast.Assert) for x in ast.walk(node))
                          tests.append({"name": node.name, "doc": _doc1(node), "asserts": n_assert})
                      elif isinstance(node, ast.ClassDef):
                          scan(node.body)
      
              scan(tree.body)
              if tests:
                  out.append({"file": path, "tests": tests})
          return out
      
      
      def render_md(doc_path, doc_text, api, tests):
          L = [f"# Context bundle for `{doc_path}`", ""]
          L += ["## Document", "", "```markdown", doc_text.rstrip(), "```", ""]
          L += ["## API surface (from source, not imported)", ""]
          if not api:
              L.append("_(no source provided)_")
          for f in api:
              L.append(f"### `{f['file']}`")
              if f.get("error"):
                  L.append(f"- parse error: {f['error']}")
                  continue
              for d in f["defs"]:
                  if d.get("sig") is None and "methods" in d:
                      L.append(f"- class `{d['name']}`" + (f" — {d['doc']}" if d["doc"] else ""))
                      for m in d["methods"]:
                          L.append(f"    - `{m['name']}({m['sig']})`" + (f" — {m['doc']}" if m["doc"] else ""))
                  else:
                      L.append(f"- `{d['name']}({d['sig']})`" + (f" — {d['doc']}" if d["doc"] else ""))
              L.append("")
          L += ["## Tests (names + assertion counts, from source)", ""]
          if not tests:
              L.append("_(no tests provided)_")
          for f in tests:
              L.append(f"### `{f['file']}`")
              for t in f["tests"]:
                  note = t["doc"] or f"{t['asserts']} assert(s)"
                  L.append(f"- `{t['name']}` — {note}")
              L.append("")
          L += ["---", "Reviewer: compare each claim the Document makes against the API surface and the",
                "Tests above. Report drift as PASS / FAIL / UNSUPPORTED (no test backs the claim) /",
                "STALE (claim refers to something that no longer exists). The behavioral gate is the",
                "test suite; this review covers only whether the prose matches reality."]
          return "\n".join(L)
      
      
      def main():
          ap = argparse.ArgumentParser(description="Bundle doc + API surface + test inventory for review.")
          ap.add_argument("--doc", required=True)
          ap.add_argument("--src", action="append", default=[], help="source file or dir (repeatable)")
          ap.add_argument("--tests", action="append", default=[], help="test file or dir (repeatable)")
          ap.add_argument("--json", action="store_true")
          args = ap.parse_args()
      
          if not os.path.isfile(args.doc):
              print(f"doc not found: {args.doc}", file=sys.stderr)
              return 2
          doc_text = open(args.doc, encoding="utf-8").read()
          api = api_surface(args.src)
          tests = test_inventory(args.tests)
      
          if args.json:
              print(json.dumps({"doc": args.doc, "doc_text": doc_text, "api": api, "tests": tests}, indent=2))
          else:
              print(render_md(args.doc, doc_text, api, tests))
          return 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
  • CHANGELOG.md 1.7 KB
    # Changelog
    
    ## 0.2.0 — 2026-06-07
    
    Pivot: from a claim-comment DSL to an agent-driven review.
    
    The v0.1 model stapled a `<!-- claim: ... -->` comment beside prose and checked
    the comment against the code. That left the prose — the half humans actually
    read — unverified, and was out-engineered on both flanks (Gherkin binds
    executable scenarios; Lean's Verso transcludes facts; TDD couples code to
    tests). v0.2 fills the remaining slot — free-prose documentation — by having the
    agent read the prose's meaning and compare it to the code and the tests
    directly. No DSL, no shadow copy.
    
    - Removed: `scripts/verify_claims.py` (the comment parser + signature/
      command-output resolvers), `assets/verify-claims.yml`, `assets/
      test_verify_claims.py`, `references/example-spec.md`. The CI/pytest forcing
      functions belonged to the DSL approach; the deterministic gate now lives where
      it should — the project's own test suite (TDD).
    - Added: `scripts/gather_context.py` — deterministic input bundler (document +
      ast-parsed API surface + test inventory), no imports, no execution.
    - Added verdict `UNSUPPORTED` — claim matches code but no test backs it (a
      missing test, surfaced as such).
    - Reframed as a triggered review, not a merge gate; the division of labor with
      TDD is now explicit in SKILL.md.
    
    ## 0.1.0 — 2026-06-06
    
    Initial skill (working title "verso"). Claim-comment DSL with `signature` and
    `command-output` resolvers, `--watch`, import allowlist, and CI/pytest
    integration templates. Superseded by 0.2.0.
    
    ## [0.2.0] - 2026-06-13
    
    ### Fixed
    
    - valid YAML frontmatter — colon-space in description broke parsing (#692)
    
    ### Other
    
    - verifying-claims v0.2: pivot from claim-DSL to agent-driven doc/code/test review (#691)
    
  • README.md 1.3 KB
    # verifying-claims
    
    Check that what a document *says* about code is true — by reading the document,
    the code, and the tests together and reporting where they disagree.
    
    The reviewer is the agent, not a parser. There is no claim-comment DSL: the
    prose's meaning is read directly and compared to what the code does and what the
    tests assert. No shadow copy, because the thing checked is the thing the human
    reads.
    
    See `SKILL.md` for the procedure, the verdicts (PASS / FAIL / UNSUPPORTED /
    STALE), and the division of labor with TDD.
    
    ## Quick start
    
    ```
    python3 scripts/gather_context.py --doc README.md --src pkg/ --tests tests/
    ```
    
    That bundles the document text, the public API surface (ast-parsed, never
    imported), and the test inventory into one report. The agent then reads the
    bundle and judges each prose claim against it.
    
    ## What this is and isn't
    
    - **Is:** a triggered, semantic review of whether documentation matches reality
      — run before docs ship, after a refactor, or as a sweep.
    - **Isn't:** a CI merge gate or a test framework. The deterministic behavioral
      gate is your test suite (TDD). This is the prose layer tests can't reach.
    
    ## Files
    
    - `scripts/gather_context.py` — deterministic input bundler.
    - `references/drift-report-example.md` — what a review report looks like.
    
  • SKILL.md 5.1 KB
    ---
    name: verifying-claims
    description: Check that a document's claims about code are actually true by reading the prose, the code, and the tests and reporting (or fixing) where they disagree. Use whenever the user wants to verify a README, guide, spec, or docstring still matches the code; whenever they mention documentation drift, doc-code sync, "is this still accurate", stale docs, or keeping docs/tests/code consistent; before publishing or merging a docs change; or as a periodic doc-accuracy sweep. The agent reads the prose's meaning directly — there is no claim-comment DSL to maintain. Pairs with TDD — the test suite is the deterministic behavioral gate, this skill is the semantic prose-vs-reality review.
    metadata:
      version: 0.2.0
    ---
    
    # verifying-claims
    
    Check that what a document *says* about code is true, by reading the document,
    the code, and the tests together and reporting where they disagree.
    
    ## What changed (v0.1 → v0.2)
    
    v0.1 was a comment-DSL: you hand-wrote `<!-- claim: ... -->` next to prose and
    a script checked the *comment* against the code. That had a fatal gap — the
    comment and the prose were two artifacts stapled together, and only the comment
    was checked, while humans read the prose. The prose could lie with a green run.
    
    v0.2 drops the DSL. The reviewer is the agent: it reads the prose's *meaning*
    directly and compares it to what the code does and what the tests assert. No
    shadow copy, because the thing being checked is the thing the human reads.
    (Existing tools already own the alternatives — Gherkin binds executable
    scenarios, Lean's Verso transcludes facts into prose, TDD couples code to
    tests. This fills the remaining slot: free-prose documentation, judged.)
    
    ## Division of labor — read this first
    
    This skill does NOT gate merges and is NOT a test framework.
    
    - **The test suite (TDD/CI)** owns the behavioral contract: deterministic,
      cheap, auditable, gated. A green check is something you can hold CI to.
    - **This skill** owns the prose layer: does the documentation match reality?
      That needs semantic judgment across artifacts, which is non-deterministic and
      fallible — so it runs as a *triggered review* (before docs ship, on request,
      as a sweep), not as a per-commit gate. "The agent said the docs match" is not
      a guarantee you gate a merge on; it's a review you act on.
    
    Tests are the anchor. The docs are correct when they agree with what the tests
    assert about the code. So write/keep good tests first; this skill keeps the
    prose pinned to them.
    
    ## Procedure
    
    1. **Identify** the document(s) to check and the code + tests they describe.
    2. **Gather** consistent input: run `scripts/gather_context.py --doc DOC --src
       SRC --tests TESTS`. It ast-parses source (no imports, no execution) and
       bundles the document text, the public API surface, and the test inventory.
    3. **Extract the claims** the prose makes — every checkable assertion about the
       code (signatures, behavior, return shapes, defaults, guarantees, examples).
       Do this by reading; there are no claim markers.
    4. **Judge each claim** against the API surface and the tests:
       - Does the code actually do what the prose says?
       - Is the claim backed by a test, or merely asserted?
       - Does it reference something that no longer exists?
    5. **Report** drift, ranked by severity, each finding citing the prose claim and
       the contradicting reality (file/function). Use the verdicts below.
    6. **Optionally fix**: rewrite the prose to match reality, and/or flag claims
       that need a test (an UNSUPPORTED claim is a missing test, not just a doc bug).
    
    ## Verdicts
    
    - **PASS** — the prose claim matches the code and is exercised by a test.
    - **FAIL** — the code contradicts the claim (the doc is wrong, or the code
      regressed and the doc caught it).
    - **UNSUPPORTED** — the claim matches the current code but no test backs it, so
      nothing protects it from future drift. Surface as a missing test.
    - **STALE** — the claim refers to something removed or renamed.
    
    ## Invoking
    
    - "Check the README against the code before I publish it."
    - "Does `docs/api.md` still match `pkg/`?"
    - "Sweep the docs for drift after this refactor."
    
    Run it at moments that matter — pre-publish, post-refactor, on a docs PR — not
    on every commit. The deterministic gate is the test suite; this is the layer
    tests can't reach.
    
    ## Honest limits
    
    - Non-deterministic and fallible: a review can miss drift or misjudge. Treat
      output as a careful review, not a proof.
    - Cost/latency: reading three artifacts and reasoning is expensive next to a
      test run. Don't wire it where a cheap deterministic check belongs.
    - It checks prose against code+tests; it does not verify the tests themselves
      are correct. Garbage tests → confident-but-wrong PASS. TDD discipline upstream
      still matters.
    
    ## When NOT to use
    
    - As a CI merge gate (use the test suite).
    - To verify behavior (write a test).
    - On prose with no factual claims about code (nothing to check).
    
    ## Files
    
    - `scripts/gather_context.py` — deterministic input bundler (doc + API surface +
      test inventory), ast-only, no imports.
    - `references/drift-report-example.md` — what a review report looks like.
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related