{"slug":"verify-refs","title":"verify-refs","summary":"Audit-only verification of manuscript references against PubMed and CrossRef. Detects fabricated or mismatched citations and writes qc/reference_audit.json. Does not modify references/ or refs.bib.","platform":"Claude","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-14T20:48:09.599854Z","repo":{"url":"https://github.com/Aperivue/medsci-skills","stars":318,"forks":75,"license":"MIT","updatedAt":"2026-09-27T05:05:17Z"},"bodyHtml":"<hr>\n<h2>name: verify-refs\ndescription: Audit-only verification of manuscript references against PubMed and CrossRef. Detects fabricated or mismatched citations and writes qc/reference_audit.json. Does not modify references/ or refs.bib.\ntriggers: verify refs, verify references, citation audit, reference hallucination, fabricated references, bibliography check, PMID check, DOI check\ntools: Read, Write, Edit, Bash, Grep, Glob\nmodel: inherit</h2>\n<h1>Verify References (Audit-Only)</h1>\n<p>You help a medical researcher prevent reference hallucinations before submission.\nThis skill audits an existing manuscript or bibliography. It <strong>does not write</strong>\nto <code>references/</code> or <code>manuscript/_src/refs.bib</code>. It does not discover new\nliterature; use <code>/search-lit</code> for discovery and <code>/lit-sync</code> for bib management.</p>\n<h2>When to Use</h2>\n<ul>\n<li>Before journal submission, especially for <code>.docx</code> manuscripts inherited from\ncoauthors or external editors.</li>\n<li>After AI-assisted drafting or revision introduced or modified references.</li>\n<li>When a reviewer or collaborator flags a possibly fabricated citation.</li>\n<li>Before <code>/sync-submission</code> freezes a journal package.</li>\n</ul>\n<h2>Inputs</h2>\n<ol>\n<li>Manuscript or bibliography path: <code>.md</code>, <code>.docx</code>, <code>.bib</code>, <code>.txt</code>, or <code>.tsv</code>.</li>\n<li>Optional project root. Default: current working directory.</li>\n<li>Optional flags passed to the script:\n<ul>\n<li><code>--offline</code>: extract and classify references without API verification.</li>\n<li><code>--timeout N</code>: HTTP timeout seconds.</li>\n</ul>\n</li>\n</ol>\n<h2>Companion: pandoc citation key check</h2>\n<p>For markdown manuscripts using pandoc <code>[@bibkey]</code> citations, validate citation\nkeys first to catch undefined/unused keys before this audit. If you also use the\ncompanion <code>manage-refs</code> skill, run its <code>check_citation_keys.py</code> for this;\notherwise use your reference manager's citation-key check.</p>\n<p>Then run <code>verify_refs.py</code> against the .bib to validate each entry against\nPubMed/CrossRef. The two checks are complementary: a citation-key check catches\nmis-keyed cites; <code>verify_refs.py</code> catches fabricated metadata.</p>\n<h2>Deterministic Script</h2>\n<p>Run the bundled script rather than verifying citations by memory:</p>\n<pre><code>python \"${CLAUDE_SKILL_DIR}/scripts/verify_refs.py\" manuscript/manuscript.md --project-root .\n</code></pre>\n<p>For hooks or quick manual runs, use the wrapper:</p>\n<pre><code>\"${CLAUDE_SKILL_DIR}/scripts/verify_cli.sh\" manuscript/manuscript.md --offline\n</code></pre>\n<p><strong>Manual pre-submission strict run</strong> (Phase 1A.5):</p>\n<pre><code>\"${CLAUDE_SKILL_DIR}/scripts/verify_cli.sh\" manuscript/index.qmd --strict\n</code></pre>\n<p><code>--strict</code> forbids <code>--offline</code> and exits non-zero on any UNVERIFIED row.\nFull checkpoint protocol: <code>references/manual_checkpoint_guide.md</code>.</p>\n<p>The script uses DOI, PMID, CrossRef, PubMed E-utilities, and OpenAlex where\navailable. If network verification fails, it records <code>UNVERIFIED</code> rather than\nsilently passing.</p>\n<p><strong>OpenAlex tertiary index (existence recovery).</strong> PubMed covers only biomedical\nliterature and CrossRef's conference-proceedings coverage is uneven, so\nNeurIPS / ICLR / ACL-style citations — common in medical-AI manuscripts — fall\nthrough both and would be marked <code>UNVERIFIED</code>. After the PubMed and CrossRef tiers,\nthe script consults OpenAlex (<code>https://api.openalex.org</code>, free, no API key) <strong>only\nwhen no authoritative author list was obtained yet</strong> (so a reference already\nresolved by PubMed/CrossRef incurs no extra call). It resolves by DOI when present,\notherwise by a title search guarded by a token-similarity threshold so a fabricated\ntitle cannot earn a spurious <code>OK</code>. This is the free analogue of the second index\n(e.g. Scopus) that journal submission portals run alongside CrossRef. OpenAlex\ndisplay names carry no structured family/given split and mix <code>First Last</code> with\n<code>Last, First</code> forms, so OpenAlex-sourced authors support an existence check plus a\ntolerant first-author <em>membership</em> check, but never drive the strict positional or\nauthor-count MISMATCH (those stay reserved for PubMed efetch / CrossRef). An\nOpenAlex miss is recorded as <code>UNVERIFIED</code>, never <code>FABRICATED</code>. Pass <code>--no-openalex</code>\nto restrict verification to PubMed + CrossRef.</p>\n<h2>Output Contract (v1.3.0)</h2>\n<table>\n<thead>\n<tr>\n<th>Artifact</th>\n<th>Path</th>\n<th>Purpose</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Audit JSON</td>\n<td><code>qc/reference_audit.json</code></td>\n<td>Metadata audit output — row-level status (OK/MISMATCH/UNVERIFIED/FABRICATED), counts, <code>cited_authors[]</code>/<code>actual_authors[]</code>, <code>duplicate_findings[]</code>, submission-safe flag, full records</td>\n</tr>\n</tbody>\n</table>\n<p><strong>v1.2.0 (2026-05)</strong> adds <code>duplicate_findings[]</code> to the audit JSON. Verbatim PMID or DOI duplicates within the reference list are flagged as MAJOR findings (resolves <code>/peer-review</code> Phase 2A P7). DOI normalization strips <code>https://doi.org/</code>, <code>http://dx.doi.org/</code>, <code>doi:</code> prefixes plus trailing slashes before comparison so <code>https://doi.org/10.x/abc/</code> and <code>10.x/abc</code> collapse to one key. Both <code>submission_safe</code> and <code>fully_verified</code> now require <code>duplicate_findings</code> to be empty.</p>\n<p><strong>v1.3.0 (2026-05)</strong> extends the author cross-check from first-author-only to the <strong>full author list</strong> and bumps <code>schema_version</code> to 4. For BibTeX inputs, every cited author family name is compared index-by-index against the authoritative source, and the cited-vs-source author counts are compared. PubMed <code>efetch.fcgi</code> (XML full record) is the truth source when a PMID is present — it is authoritative for given/family names where CrossRef is not (a documented case where CrossRef returned a wrong given name that PubMed efetch corrected). Records now carry <code>cited_authors[]</code>, <code>actual_authors[]</code>, <code>cited_author_count</code>, and <code>actual_author_count</code>. A correct first author does not establish that the remaining author names are authentic. Plain-text / TSV inputs, which cannot be parsed into a confident full list, degrade gracefully to the first-author check.</p>\n<p><strong>Removed in Phase 1A.2</strong> (per <code>docs/artifact_contract.md</code>):</p>\n<ul>\n<li><code>references/verified_references.tsv</code> — record-level details now live inside <code>reference_audit.json</code> under <code>records[]</code>.</li>\n<li><code>references/library.bib</code> — never this skill's concern. <code>/search-lit</code> produces candidates; <code>/lit-sync</code> (via Better BibTeX) writes <code>manuscript/_src/refs.bib</code>.</li>\n</ul>\n<p>Sole-writer enforcement: <code>scripts/validate_project_contract.py</code> will flag any <code>references/*</code> file written by this skill as drift.</p>\n<h2>Workflow</h2>\n<ol>\n<li>Identify the input file and project root.</li>\n<li>Run <code>scripts/verify_refs.py</code>.</li>\n<li>Read <code>qc/reference_audit.json</code>.</li>\n<li>Report all <code>FABRICATED</code> and <code>MISMATCH</code> rows first (from <code>records[]</code>).</li>\n<li>Report all <code>duplicate_findings[]</code> entries (verbatim PMID/DOI duplicates — cite renumbering required).</li>\n<li>If <code>UNVERIFIED</code> rows remain, list them as manual checks and do not call the\nmanuscript fully submission-safe. Rows with <code>note = \"pagination_placeholder\"</code>\n(<code>e000–e000</code> / <code>in press</code> / <code>TBD</code> / <code>forthcoming</code>) need the citation resolved\nbefore submission; <code>/self-review</code> Phase 2.5c decides whether any is a P0 blocker.</li>\n<li>If the user needs a human-readable table, summarize from <code>records[]</code> in chat — do not write a TSV.</li>\n</ol>\n<h2>Quality Gates</h2>\n<ul>\n<li>Gate 1: stop submission if any row is <code>FABRICATED</code>.</li>\n<li>Gate 2: require user confirmation before accepting <code>UNVERIFIED</code> references.</li>\n<li>Gate 3: rerun after any reference edits.</li>\n<li>Gate 4 (added 2026-04-26; extended to full-author in v1.3.0): the cited\nauthor list is cross-checked against the authoritative source (PubMed efetch\npreferred, then CrossRef, then PubMed esummary). A row whose DOI/PMID resolves\nbut whose cited authors do not match — at any index, or in total count — is\ndowngraded to <code>MISMATCH</code>. First-author mismatches get\n<code>note = \"first-author hallucination suspected\"</code>; #2..#N family or count\nmismatches get <code>note = \"non-first-author hallucination or count mismatch\"</code>.\nThis catches the LLM failure mode where a real DOI is paired with invented\nauthor names anywhere in the list, not just the lead author. Intentional CSL\net-al truncation (cited fewer than source) can be silenced per-entry with a\nBibTeX <code>_audit_truncated = &lt;N&gt;</code> field.</li>\n<li>Gate 5 (added 2026-05, v1.2.0): PMID/DOI duplicate detection within the\nreference list. Verbatim duplicates (same PMID or normalized DOI) — a common\nLLM citation-compilation artifact — are flagged as MAJOR findings in\n<code>duplicate_findings[]</code>. <code>submission_safe == true</code> requires the list to be\nempty. Resolves <code>/peer-review</code> Phase 2A P7.</li>\n<li>Gate 6 (added 2026-06): pagination / publication-stage placeholders. A reference\nwhose raw entry still carries <code>e000–e000</code>, <code>in press</code>, <code>TBD</code>, or <code>forthcoming</code>\nis not yet a fully citable record. Each is marked <code>UNVERIFIED</code> with\n<code>note = \"pagination_placeholder\"</code> (a would-be <code>VERIFIED</code> record is downgraded; a\nworse status is left unchanged). <strong>verify-refs is manuscript-agnostic and does not\njudge centrality</strong> — it only flags. The escalation call (is this a method- or\nheadline-load-bearing citation, hence a P0 submission blocker?) is made by\n<code>/self-review</code> Phase 2.5c, which has the manuscript in hand.</li>\n</ul>\n<p><strong>Classification note — citation-metadata confusion is not fabrication.</strong> Digits\nin a DOI suffix sometimes look like a journal article number but differ from the\nreal one (e.g., a DOI tail \"77196\" against article number 26068, or a \"60466-1\"\nsuffix against article 6274). This is cosmetic metadata confusion, not a\nfabricated reference: do not record such rows as <code>FABRICATED</code> when the DOI/PMID\nresolves and the authors match. A genuine <code>FABRICATED</code> verdict requires a\nnon-resolving identifier or an author cross-check failure (Gate 4), not a\nmismatch between a DOI suffix and an article number.</p>\n<h2>Author Cross-Check (Detail)</h2>\n<p>Two failure patterns motivate the author checks: a real DOI can be paired with\nthe wrong first author, and a correct first author can be followed by fabricated\nco-author names. DOI resolution and first-author agreement alone cannot verify\nthe full author list.</p>\n<ul>\n<li>The authoritative author list is taken from PubMed <code>efetch.fcgi</code> (XML) when a\nPMID is present, falling back to CrossRef (DOI) and then PubMed esummary.\nefetch is preferred because CrossRef is unreliable for given names.</li>\n<li>For BibTeX inputs, the full cited list is parsed (<code>cited_authors[]</code>,\nbalanced-brace aware, LaTeX-accent tolerant) and compared family-by-family and\nby total count against <code>actual_authors[]</code>.</li>\n<li>Comparison is tolerant: case, diacritics (NFKD plus Turkish/Polish/Czech/\nGerman/Nordic special letters), hyphen vs space, and name particles\n(\"von\", \"van\", \"de\", ...) are normalized before matching.</li>\n<li>If the cited authors cannot be parsed confidently, the check degrades to the\nfirst-author surname comparison, and if even that is empty it is skipped\nsilently — no false MISMATCH from formatting ambiguity.</li>\n<li>Title-only PubMed search does not return an authoritative author and is\ntherefore excluded from this check.</li>\n<li>Intentional truncation (a bib that cites only the first author, or first five\n<ul>\n<li>et al., by design) would otherwise trip the count check; mark such entries\nwith <code>_audit_truncated = &lt;N&gt;</code> to downgrade the count mismatch to a note.</li>\n</ul>\n</li>\n</ul>\n<h2>Claim Fidelity — does the source say what you say it says?</h2>\n<p><code>verify_refs.py</code> answers whether a reference is real and whose it is. It cannot answer\nwhether the <em>sentence citing it</em> is true of it, and a citation can be perfectly real while\nthe claim attached to it is not. That gap is where the failure lives: the DOI resolves, the\nauthors match, the reference list renders, and the sentence is still wrong.</p>\n<p><code>scripts/check_claim_fidelity.py</code> checks the claims that have a checkable answer, against\nfull texts you have already downloaded and converted (<code>/fulltext-retrieval</code> produces exactly\nthat layout — it never fetches anything itself):</p>\n<pre><code>python3 \"${CLAUDE_SKILL_DIR}/scripts/check_claim_fidelity.py\" \\\n  --manuscript manuscript/manuscript.md \\\n  --fulltext-dir fulltext/ --bib manuscript/_src/refs.bib \\\n  --out qc/claim_fidelity.json --strict\n</code></pre>\n<table>\n<thead>\n<tr>\n<th>Verdict</th>\n<th>Severity</th>\n<th>Fires when</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>CITED_QUOTE_ABSENT</code></td>\n<td>major</td>\n<td>Quoted text attributed to a source is not in it in any reading order.</td>\n</tr>\n<tr>\n<td><code>CITED_QUOTE_UNRESOLVED</code></td>\n<td>prompt</td>\n<td>The quote matched only with foreign tokens wedged in, or a word or two missing — the signature of a dirty extraction, not of a fabrication. Look; do not assume.</td>\n</tr>\n<tr>\n<td><code>ATTRIBUTION_UNSUPPORTED</code></td>\n<td>prompt</td>\n<td>Not one content word of the attributed claim appears in the source, in any form. Paraphrase normally keeps at least one of the source's own terms.</td>\n</tr>\n<tr>\n<td><code>ORDINAL_CLAIM_UNSUPPORTED</code></td>\n<td>prompt</td>\n<td>\"reports three strategies [12]\" where the source discusses that noun but never that count near it.</td>\n</tr>\n</tbody>\n</table>\n<p>Only the quote verdict can fail <code>--strict</code>. Everything else is a prompt to go read the\nsource, because paraphrase is legitimate and a gate that blocks on it would be turned off.</p>\n<p><strong>Read the \"not checked\" lines.</strong> A citation with no full text on disk is reported as\nunresolved and never guessed at, and a source whose extracted text is an abstract is reported\nas too short to judge — absence proves nothing against an abstract. Silence from this\ndetector means \"nothing checkable was wrong\", which is not the same as \"everything is right\".</p>\n<h3>Sentence-level source evidence table</h3>\n<p>The same <code>qc/claim_fidelity.json</code> now includes <code>evidence_rows</code>: recognized prose\nsentence/citation pairs, manuscript coordinates, source-text and PDF hashes,\nadvisory retrieval identity, and a separate assessor-entered comparison. Initial\nrows are <code>not_assessed</code>, even when bibliographic status is OK and no probe fires.</p>\n<pre><code>python3 \"${CLAUDE_SKILL_DIR}/scripts/check_claim_fidelity.py\" \\\n  --manuscript manuscript/manuscript.md --bib manuscript/_src/refs.bib \\\n  --fulltext-dir fulltext/ --retrieval-report pdfs/retrieval_report.json \\\n  --reference-audit qc/reference_audit.json \\\n  --out qc/claim_fidelity.json --evidence-table qc/claim_fidelity.md\n</code></pre>\n<p>Inspect the actual source before entering pages, excerpts, metric/unit/denominator,\npopulation, direction, and a named assessment. Neither equal numbers nor matching\nwords establish support. Record whether the assessor used AI assistance; do not\ndescribe an AI-generated assessment as human approval. Rerun with\n<code>--reviewed-report qc/claim_fidelity.json</code> to retain annotations. Changed inputs\nleave old assessments unresolved; unmatched rows remain in the JSON for review.\nThe Markdown table is a derived view, not a second editable evidence store.</p>\n<p>See <code>references/claim_evidence_workflow.md</code> for field meanings, re-review steps,\nsource-identity limitations, and the difference between recorded and verified.</p>\n<h2>What This Skill Does NOT Do</h2>\n<ul>\n<li>Does not fetch full texts (use <code>/fulltext-retrieval</code>); claim fidelity reads converted text\noff disk so it stays deterministic and CI-runnable.</li>\n<li>Does not automatically judge topical fit or semantic support. The probes check limited\nwording patterns; the evidence table records attributed assessments, not verified facts.</li>\n<li>Does not generate new references from memory.</li>\n<li>Does not replace missing citations with plausible alternatives without\n<code>/search-lit</code> or user approval.</li>\n<li>Does not sync Zotero collections; use <code>/lit-sync</code> after this audit.</li>\n</ul>\n<h2>Anti-Hallucination</h2>\n<ul>\n<li>Never fabricate titles, DOIs, PMIDs, author lists, journal names, years,\nvolumes, or pages.</li>\n<li>Every OK row must be backed by DOI, PMID, CrossRef, or PubMed title evidence.</li>\n<li>If evidence is unavailable, mark <code>UNVERIFIED</code> and keep it visible.</li>\n</ul>\n","files":[{"path":"references/claim_evidence_workflow.md","sizeBytes":6466,"isText":true},{"path":"references/manual_checkpoint_guide.md","sizeBytes":4031,"isText":true},{"path":"scripts/check_claim_fidelity.py","sizeBytes":33348,"isText":true},{"path":"scripts/_claim_evidence.py","sizeBytes":13992,"isText":true},{"path":"scripts/claim_fidelity_challenge/fixture/fulltext/10.1000_synthetic.oversight.md","sizeBytes":3493,"isText":true},{"path":"scripts/claim_fidelity_challenge/fixture/fulltext/10.1000_synthetic.stub.md","sizeBytes":910,"isText":true},{"path":"scripts/claim_fidelity_challenge/fixture/manuscript_numbered.md","sizeBytes":769,"isText":true},{"path":"scripts/claim_fidelity_challenge/fixture/manuscript_shortsource.md","sizeBytes":506,"isText":true},{"path":"scripts/claim_fidelity_challenge/fixture/manuscript_supported.md","sizeBytes":732,"isText":true},{"path":"scripts/claim_fidelity_challenge/fixture/manuscript_unsupported.md","sizeBytes":692,"isText":true},{"path":"scripts/claim_fidelity_challenge/fixture/refs.bib","sizeBytes":588,"isText":false},{"path":"scripts/claim_fidelity_challenge/verify.sh","sizeBytes":8835,"isText":true},{"path":"scripts/_quote_match.py","sizeBytes":7904,"isText":true},{"path":"scripts/verify_cli.sh","sizeBytes":1923,"isText":true},{"path":"scripts/verify_refs.py","sizeBytes":49027,"isText":true},{"path":"SKILL.md","sizeBytes":15318,"isText":true},{"path":"skill.yml","sizeBytes":3422,"isText":true},{"path":"tests/fixtures/corporate_author.bib","sizeBytes":821,"isText":false},{"path":"tests/fixtures/pagination_placeholder.bib","sizeBytes":410,"isText":false},{"path":"tests/test_audit_source_path.sh","sizeBytes":4372,"isText":true},{"path":"tests/test_author_normalization.sh","sizeBytes":3948,"isText":true},{"path":"tests/test_bib_last_field.sh","sizeBytes":4746,"isText":true},{"path":"tests/test_bibtex_et_al.sh","sizeBytes":4686,"isText":true},{"path":"tests/test_claim_evidence.py","sizeBytes":16345,"isText":true},{"path":"tests/test_corporate_author.sh","sizeBytes":2215,"isText":true},{"path":"tests/test_fabricated_author.sh","sizeBytes":3147,"isText":true},{"path":"tests/test_openalex_tier.sh","sizeBytes":6153,"isText":true},{"path":"tests/test_pagination_placeholder.sh","sizeBytes":1701,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-14T20:48:29.310684Z","sha256":"663CABAECE30BDD9F77C1D00AF81CA592616FC118A5F502B19917A48F66F99CB","sizeBytes":77992},"review":null,"source":{"repositoryUrl":"https://github.com/Aperivue/medsci-skills","path":"skills/verify-refs","license":"MIT","commit":"5599b724675a1d788e03cd58dabd3db7c68ca86b","subtreeSha":"F3E1B3BE4D008657C78BE54784C250353A6E3C1E81B997BF7915B8D33FCB4789","lastSyncedAt":"2026-09-27T19:46:33.449845Z"},"reviewedAt":"2026-09-14T20:53:14.856423Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/Aperivue/medsci-skills/tree/main/skills/verify-refs"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install aperivue-medsci-skills@llmmart"},{"target":"git","command":"git clone https://github.com/Aperivue/medsci-skills.git"}]}