search-lit
Literature search and citation management for medical research. Searches PubMed, Semantic Scholar, and bioRxiv/medRxiv with verified citations. Anti-hallucination — every reference verified via API before inclusion. Generates BibTeX entries.
Install
npx skills add https://github.com/Aperivue/medsci-skills/tree/main/skills/search-lit
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install aperivue-medsci-skills@llmmart
git clone https://github.com/Aperivue/medsci-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole aperivue/medsci-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Literature Search Skill
Every reference you produce must come from a live database result — never generate a citation from memory alone, because a recalled citation can look real and not exist.
Search Tools: MCP (Primary) + E-utilities (Fallback)
Primary: MCP Tools (Claude.ai Remote)
| Database | MCP Tool | Purpose |
|---|---|---|
| PubMed | mcp__claude_ai_PubMed__search_articles |
Search by query, MeSH terms |
| PubMed | mcp__claude_ai_PubMed__get_article_metadata |
Full metadata for a PMID |
| PubMed | mcp__claude_ai_PubMed__find_related_articles |
Related articles for a PMID |
| PubMed | mcp__claude_ai_PubMed__lookup_article_by_citation |
Verify a citation |
| PubMed | mcp__claude_ai_PubMed__convert_article_ids |
Convert between PMID/DOI/PMCID |
| Semantic Scholar | mcp__claude_ai_Scholar_Gateway__semanticSearch |
Semantic search across all fields |
| bioRxiv/medRxiv | mcp__claude_ai_bioRxiv__search_preprints |
Search preprint servers |
| bioRxiv/medRxiv | mcp__claude_ai_bioRxiv__get_preprint |
Full preprint metadata |
| CrossRef | WebFetch with https://api.crossref.org/works/{DOI} |
DOI verification |
Fallback: NCBI E-utilities (Direct API via Bash)
If any mcp__claude_ai_PubMed__* call returns an error containing "terminated", "not found",
"not available", or "not connected", switch ALL subsequent PubMed calls in this session to the
bundled E-utilities scripts. Do not retry MCP after a disconnect — it will not recover within the
same conversation.
EUTILS="${CLAUDE_SKILL_DIR}/references/pubmed_eutils.sh"
PARSER="${CLAUDE_SKILL_DIR}/references/parse_pubmed.py"
# Search PubMed (returns PMIDs)
bash "$EUTILS" search "diagnostic test accuracy meta-analysis radiology" 20 \
| python3 "$PARSER" esearch
# Get article summaries as markdown table
bash "$EUTILS" fetch_json "16168343,16085191,31462531" \
| python3 "$PARSER" esummary
# Get detailed metadata
bash "$EUTILS" fetch "16168343" \
| python3 "$PARSER" efetch
# Generate BibTeX entries
bash "$EUTILS" fetch "16168343,16085191" \
| python3 "$PARSER" bibtex
# Verify a citation by exact title
bash "$EUTILS" cite_lookup "Bivariate analysis of sensitivity and specificity" \
| python3 "$PARSER" esearch
# Find related articles for a PMID
bash "$EUTILS" related "16168343" 10 \
| python3 "$PARSER" esummary
The script sleeps 350 ms between calls (NCBI allows 3 requests/s without an API key, 10/s with
NCBI_API_KEY); keep batch calls sequential.
| MCP Tool | E-utilities Command | Parser Mode |
|---|---|---|
search_articles |
search <query> [retmax] |
esearch |
get_article_metadata |
fetch <pmids> |
efetch or bibtex |
find_related_articles |
related <pmid> [retmax] |
esummary |
lookup_article_by_citation |
cite_lookup <title> |
esearch → fetch |
convert_article_ids |
Not available (use CrossRef DOI lookup) | — |
Workflow
Phase 1: Search Strategy
- Get the research topic, question, or manuscript section that needs references.
- Build the query from the key concepts (Population, Intervention/Exposure, Comparison, Outcome),
with MeSH terms for PubMed:
(concept1 OR synonym1) AND (concept2 OR synonym2). - Set scope: date range (default: last 10 years unless the user specifies), article types, and language (default: English).
- Present the Boolean query, databases, and filters to the user.
Gate: Wait for user approval before running searches.
Phase 2: Execute Search
- Search PubMed (
search_articles, Boolean query), Semantic Scholar (semanticSearch, natural language query), and bioRxiv/medRxiv (search_preprints) when preprints are relevant. - Deduplicate across databases by DOI or title similarity.
- Present the results in this table, and write the same rows to
references/search_results.tsv:
| # | Title | Authors (first + last) | Year | Journal | PMID/DOI | Relevance |
|---|-------|----------------------|------|---------|----------|-----------|
| 1 | ... | Kim J, ... Lee S | 2024 | Radiology | 12345678 | High |
- Ask the user to select which papers to include.
Record what the source said existed, not only what you downloaded
A wrong haul raises no error, and a PRISMA flow built on a wrong number is fiction that nothing downstream contradicts. Check both of these:
- A count that equals a page cap exactly. Every source reports a total:
esearchresult.count,opensearch:totalResults,meta.count. Recordapi_totalbesidedownloaded, and fail loudly whendownloaded < api_total, or whendownloadedequals a page or loop cap exactly (e.g. a loop's ownif start >= 2000: breakreporting 2,000 records). PrintTRUNCATEDand refuse to write the search record. - A boolean that was never applied. OpenAlex's
search=is a relevance-ranked free-text parameter that silently ignores AND/OR;filter=title_and_abstract.search:honours them. Run the query once more with one mandatory clause negated. If the hit count does not drop, the boolean is being ignored — the engine is ranking, not filtering.
PubMed via E-utilities is the one place where the naive pattern is safe. Everywhere else, do both.
A DOI in a screening row is not necessarily that row's DOI
When the pipeline filled a doi column (matched against Crossref by title similarity) rather
than receiving it with the record, a wrong match is a valid, resolvable DOI for a different paper.
Resolve it and read the title back before any decision rests on it:
python3 "${CLAUDE_SKILL_DIR}/scripts/check_doi_record_match.py" --table 2_Screening/round3.tsv \
--email <contact> --json qc/doi_record_match.json
DOI_NOT_THIS_RECORD is a DOI that resolves to another paper; DOI_IS_CONTAINER resolves to an
issue, supplement or proceedings rather than an article; DOI_IS_UPDATE_NOTICE resolves to a
correction, erratum or retraction notice (Crossref update-to) rather than to the article it
updates; DOI_UNRESOLVED is reported rather than dropped. This runs at screening, where a wrong DOI is still cheap; /verify-refs audits a finished
reference list.
Phase 2.5: Citation Searching (Snowballing)
Optional; recommended for systematic reviews and thorough background work (PRISMA item 7, "records identified through citation searching"). Expand a seed set along the citation graph with the Semantic Scholar Graph API helper:
python3 "${CLAUDE_SKILL_DIR}/references/snowball.py" \
--seed DOI:10.1000/synthetic.example,PMID:00000000 \
--direction all \
--pool references/library.bib \
--out references/library.bib
- Directions:
backward(references the seeds cite),forward(papers citing the seeds),similar(S2 recommendations), orall(default). Dedup is against the--pooland within the harvested set, by DOI and normalized title. - Trust flag: candidates are written
verified=false+verified_by=semantic_scholar. Run/verify-refs(or Phase 4 verification) on each before citing it. - Output contract: appends to
references/library.bibonly; the script hard-refusesmanuscript/_src/refs.bib. - PRISMA line: the script prints
Records identified through citation searching (snowballing): N raw (backward=…, forward=…, similar=…); after dedup against existing pool: M new candidates.Record M in the PRISMA flow's citation-searching box. - Incomplete runs: when any seed/direction fetch fails, or the source says more records exist
past
--limit(anextpage), the script prints aFAILEDorTRUNCATEDline per seed/direction, ends the PRISMA line withINCOMPLETE, and exits 1. Those counts are a lower bound. Do not record them; re-run (or raise--limit) until the run exits 0.
Phase 3: Deep Read
For each selected paper, retrieve full metadata (get_article_metadata for PubMed, get_preprint
for bioRxiv) and extract the study design, sample size/dataset, key methods, primary findings (with
specific numbers), and the authors' stated limitations. For several papers, present a literature
matrix for review:
| Paper | Design | N | Key Finding | Limitation | Relevance to Our Study |
|-------|--------|---|-------------|------------|----------------------|
Phase 4: Citation Management
Verification
- NEVER fabricate a DOI or PMID. If you cannot find one, mark the reference
[UNVERIFIED - NEEDS MANUAL CHECK]. - Cross-check every reference against the API result: first and last author, publication year, journal, article title (exact, not paraphrased), and volume/pages when available. Flag each field that does not match.
- Confirm each DOI resolves with WebFetch
https://api.crossref.org/works/{DOI}(on CrossRef errors, follow Error Handling).
BibTeX Generation
Generate an entry for every reference, verified or not, with an explicit verified flag:
@article{FirstAuthorLastName_Year_ShortKey,
author = {Last1, First1 and Last2, First2 and Last3, First3},
title = {Full Title As Retrieved From Database},
journal = {Journal Name},
year = {2024},
volume = {310},
number = {2},
pages = {e234567},
doi = {10.1001/jama.2024.12345},
pmid = {12345678},
verified = {true},
verified_by = {pubmed+crossref},
verified_on = {2026-04-24},
}
verified (required on every entry) |
Meaning |
|---|---|
true |
DOI or PMID confirmed via PubMed/CrossRef; title, authors, year all match |
false |
Parsed from text, but the API lookup failed or returned a mismatch; the manuscript MUST show [UNVERIFIED - NEEDS MANUAL CHECK] |
manual |
User explicitly added it despite the lookup failure; still unverified |
verified_by lists the confirming sources (pubmed, crossref, semantic_scholar, or a
combination); verified_on is the ISO date of the most recent successful verification. No
downstream script reads these fields — /verify-refs re-checks every entry — so they record trust
rather than grant it.
BibTeX key convention: FirstAuthorLastName_Year_OneWord (e.g., Kim_2024_Validation). The key
is provisional: Better BibTeX assigns the citable key when /lit-sync imports the entry into Zotero.
Output
- Append the entries to
references/library.bib(do not overwrite) — the candidate pool/lit-syncimports into Zotero. NEVER write tomanuscript/_src/refs.bib, because/lit-sync(via Better BibTeX) is its sole writer. - Print a summary with verification status:
Verified: 12 references (verified=true)
Unverified: 1 reference (verified=false) [NEEDS MANUAL CHECK]
Total: 13 references
Phase 4b: Zotero Library Integration
Importing into Zotero belongs to /lit-sync, which owns Zotero writes and
references/zotero_collection.json: hand it references/library.bib. If a Zotero MCP server is
connected, you may read from it here — zotero_search_items (by DOI) to mark candidates already in
the library, zotero_get_annotations to reference the user's prior reading notes.
Phase 5: Full-Text Retrieval
Delegate to /fulltext-retrieval, the single home of the open-access cascade; do not
re-implement OA fetching here. Pass the verified candidate DOIs from references/library.bib:
ENGINE="${CLAUDE_SKILL_DIR}/../fulltext-retrieval/fetch_oa.py"
# extract DOI + Title (and PMID/FirstAuthor when available) → worklist.tsv
python3 "$ENGINE" worklist.tsv -o pdfs/ -e <contact-email> --report pdfs/retrieval_report.json
Put the verified bibliographic title in the worklist rather than a DOI-only list. Keep
source_identity and file_sha256 from the retrieval report with the record. Download success and
title agreement alone do not verify the PDF: inspect conflicts, unresolved/unavailable evidence, and
files whose hashes have changed before citing them. Missing identity fields in older reports mean
unassessed; even consistent is advisory front-matter corroboration, not verification of the
paper's claims.
For Zotero-resident PDFs and proxy-aware retrieval, use /lit-sync Phase 2.7. For DOIs in
pdfs/manual_needed.txt, use only institutional access (your library's own subscriptions, proxy or
VPN), interlibrary loan, or the corresponding author. Never bypass paywalls or publisher access
controls, and do not configure unauthorized PDF mirrors.
Phase 6: Gap Analysis
When called during manuscript writing, extract the manuscript's inline citations, compare them with the search results, and report specific gaps: key papers not cited, outdated references with newer versions, and missing methodological references (statistical methods, reporting guidelines).
Specialized Search Modes
Mode: Manuscript Paper Reference Pool
Supplies a manuscript's reference pool — typically invoked by /write-paper Step 7.3c (or
/self-review Phase 2.5c-2) when the reference-adequacy gate finds the draft under target or a
named method uncited; usable directly for an original-research bibliography.
For an original-research article, return 25–40 verified candidates, not the ~10 a quick search settles on. If the field is genuinely sparse, say so explicitly rather than returning a thin list silently. Respect a narrower journal reference cap or user scope when one is given.
Cover six candidate categories:
- Background / disease burden / clinical context — why the question matters.
- Gap-defining prior studies — the work the manuscript extends or contradicts.
- Comparator / comparable-design cohorts — studies the Results will be measured against.
- Methods / statistical canonical sources — the originating reference for every named method, model, score, equation, or diagnostic criterion (e.g. competing-risk model, multiple imputation, E-value, eGFR equation, concordance statistic). This category clears Methods named-method gaps.
- Reporting-guideline sources — STROBE, TRIPOD(+AI), CONSORT, PRISMA(-DTA), STARD, etc.
- Interpretation / mechanism / limitation support — grounds Discussion claims.
For each candidate, report PMID/DOI, verification status, candidate category, the
target manuscript section, and a one-line why it is needed. Entries go through Phase 4 into
references/library.bib only. This mode produces candidates: the user decides inclusion, and it
does not insert references into the manuscript bib.
Mode: Crowding Check
Run before a study is designed. A background search ("what has been written about this topic") leaves the trap open; ask four narrower questions instead:
| Ask of | Verdict |
|---|---|
| the research question | taken / partly taken / open |
| the sampling frame (what population, which records, which years) | taken / partly taken / open |
| the measurement axis (what is being coded or measured, and at what granularity) | taken / partly taken / open |
| the target journal | already published there / adjacent / open |
Give each its own verdict: a design can be original on one axis and fully occupied on another, and collapsing the four into one answer hides that. The journal row is not vanity: a design once matched an existing paper on frame, coding axis and journal — its own first choice. Most of the time this mode narrows a claim rather than ending a project (e.g. from "nobody has looked at this" to "nobody has decomposed it by provenance"), and the narrowed claim survives review.
Search the way a competitor would: the exact frame, the exact measure, and the journal's own site, not only the topic. Report the four verdicts and the papers behind each, then let the user decide.
Mode: Systematic Search
For systematic reviews or comprehensive literature sections:
- Document the full search strategy (PRISMA-compliant).
- Record: database, date of search, query string, number of results.
- Track inclusion/exclusion at each screening step.
- Output a PRISMA flow diagram data summary.
Mode: Quick Cite
For a single reference the user describes ("that 2023 paper by Smith about AI in chest X-ray"): search PubMed and Semantic Scholar with the details, present the top 3 candidates, and generate the BibTeX entry for the one the user confirms.
Mode: Related Papers
From a PMID or DOI, get related papers with find_related_articles plus Semantic Scholar
citation-based recommendations, ranked by relevance. For a structured, dedup-aware,
PRISMA-countable expansion (backward + forward + similar), use Phase 2.5: Citation Searching
with references/snowball.py instead.
Mode: Embase Browser Automation
Embase has no public API. Read ${CLAUDE_SKILL_DIR}/references/embase_browser.md when the search
must include Embase — it has the Chrome-automation export steps, the CSV row format, and the
PubMed → Embase query translation.
Error Handling
- If a search returns 0 results, broaden the query (remove one concept or use broader MeSH terms) and retry.
- CrossRef HTTP errors (token-saving rules):
- 403 (rate-limited): Do NOT retry. Skip CrossRef → verify via PubMed title search instead.
- 303 (redirect): Follow the redirect if possible. If not, skip CrossRef → PubMed fallback.
- After the first CrossRef 403/303 in a session, skip CrossRef for ALL remaining references and go directly to PubMed title verification, to avoid N×retry token waste.
- Do not print raw error messages ("Request failed with status code 403."). Report one summary
line at the end:
CrossRef unavailable for {N} references (rate-limited). Verified via PubMed instead.
- If a DOI does not resolve via CrossRef (after the rules above), search PubMed by title to confirm the reference exists.
- If a reference cannot be verified by any method, state: "This reference could not be verified. Please check manually before submission." Never silently include an unverified reference.
Known limits
check_doi_record_match.pycompares titles at a similarity threshold and reads structured Crossref fields (type,update-to). A study protocol, or part 2 of a multi-part paper, whose title differs from the row's by a word or a number is not separated from the paper the row describes; no structured field marks it. Treat a silent run as "no mismatch detected", not as proof that every DOI is right.
Files (medsci-skills)
-
references
-
snowball_challenge
-
expected
-
snowball.bib 1.5 KB · in bundle
-
-
fixture
-
DOI_10_0_seed1.backward.json 600 B
{ "data": [ { "citedPaper": { "title": "Synthetic Backward Paper One", "year": 2019, "venue": "Journal of Synthetic Radiology", "externalIds": {"DOI": "10.0/back-one", "PubMed": "30000001"}, "authors": [{"name": "Ada Researcher"}, {"name": "Ben Coauthor"}] } }, { "citedPaper": { "title": "Synthetic Duplicate In Pool", "year": 2018, "venue": "Synthetic Imaging", "externalIds": {"DOI": "10.0/dup-backward", "PubMed": "30000002"}, "authors": [{"name": "Cara Author"}] } } ] } -
DOI_10_0_seed1.forward.json 571 B
{ "data": [ { "citingPaper": { "title": "Synthetic Forward Paper Alpha", "year": 2022, "venue": "AI in Synthetic Medicine", "externalIds": {"DOI": "10.0/fwd-alpha", "PubMed": "30000003"}, "authors": [{"name": "Dan Writer"}] } }, { "citingPaper": { "title": "Synthetic Forward Paper Beta", "year": 2023, "venue": "Synthetic Cardiology", "externalIds": {"DOI": "10.0/fwd-beta"}, "authors": [{"name": "Eve Scholar"}, {"name": "Finn Helper"}] } } ] } -
DOI_10_0_seed1.similar.json 263 B
{ "recommendedPapers": [ { "title": "Synthetic Similar Recommendation", "year": 2021, "venue": "Synthetic Methods", "externalIds": {"DOI": "10.0/sim-one", "PubMed": "30000005"}, "authors": [{"name": "Gail Expert"}] } ] } -
DOI_10_0_seed_err.backward.json 59 B
{ "error": "Paper with id DOI:10.0/seed_err not found" } -
DOI_10_0_seed_trunc.forward.json 678 B
{ "offset": 0, "next": 2, "data": [ { "citingPaper": { "title": "Synthetic Forward Page Paper One", "year": 2024, "venue": "Synthetic Imaging", "externalIds": { "DOI": "10.0/page-one" }, "authors": [ { "name": "Eve Scholar" } ] } }, { "citingPaper": { "title": "Synthetic Forward Page Paper Two", "year": 2025, "venue": "Synthetic Imaging", "externalIds": { "DOI": "10.0/page-two" }, "authors": [ { "name": "Finn Scholar" } ] } } ] } -
library.bib 277 B · in bundle
-
-
problem.md 2.6 KB
# Challenge card — Citation snowballing (search-lit Phase 2.5) ## Problem Keyword/Boolean search alone misses relevant studies that are only reachable through the **citation graph**. Recognized systematic-review practice ("snowballing" / "citation searching", PRISMA item 7) requires expanding a seed set **backward** (references the seeds cite), **forward** (papers citing the seeds), and **laterally** (algorithmically similar papers), then reporting how many records that step contributed. Before this gate, `search-lit` had only a lightweight "Related Papers" mode and **no structured citation-searching workflow, no dedup against the existing candidate pool, and no PRISMA citation-search count**. ## What the new gate does `references/snowball.py` expands seed DOIs/PMIDs via the Semantic Scholar Graph API (`/references`, `/citations`, `/recommendations`), deduplicates against the existing `references/library.bib` pool (by DOI and normalized title), and emits API-verified BibTeX candidates carrying `verified=false` + `verified_by=semantic_scholar` (downstream `/verify-refs` confirms each). It prints a PRISMA "records identified through citation searching" line. Output is **appended** to the candidate pool; it never writes `manuscript/_src/refs.bib` (that is `/lit-sync`'s sole path). ## Fixture (synthetic only — no real papers/PII) - `fixture/DOI_10_0_seed1.{backward,forward,similar}.json` — recorded Semantic-Scholar-shaped responses for one synthetic seed. - `fixture/library.bib` — an existing candidate pool containing one paper (`10.0/dup-backward`) that also appears in the backward results. ## Expected - `expected/snowball.bib` — 4 new candidates (1 backward, 2 forward, 1 similar). - The 2nd backward paper is **dropped** because its DOI already exists in the pool. PRISMA line: `5 raw (backward=1, forward=2, similar=1) ... 4 new`. ## Baseline vs new gate | | Baseline (Related Papers mode) | New snowballing gate | |---|---|---| | Directions | forward-ish recommendations only | backward + forward + similar | | Dedup vs pool | none | DOI + normalized title | | PRISMA citation-search count | none | printed | | Output | ad hoc | verified BibTeX appended to library.bib | ## Verifier (deterministic, no network) ```bash bash verify.sh ``` Reads the recorded fixtures, runs `snowball.py --offline-fixture`, and diffs against `expected/snowball.bib`. Exit 0 = match. ## Acknowledgement The reproducible "fixture + expected + deterministic verifier" packaging is inspired by public reproducible-audit layouts such as [EinsteinArena](https://einsteinarena.com/) (design inspiration only; no code, solutions, or data were copied). -
verify.sh 3.1 KB
#!/usr/bin/env bash # Deterministic verifier for the citation-snowballing challenge card. # No network: reads recorded Semantic Scholar responses from fixture/. # Exit 0 = output matches expected/snowball.bib ; non-zero = regression. set -euo pipefail HERE="$(cd "$(dirname "$0")" && pwd)" SNOWBALL="$HERE/../snowball.py" actual="$(python3 "$SNOWBALL" \ --seed DOI:10.0/seed1 \ --direction all \ --offline-fixture "$HERE/fixture" \ --pool "$HERE/fixture/library.bib" \ --as-of 2026-06-14 \ --stdout 2>/dev/null)" if diff -u "$HERE/expected/snowball.bib" <(printf '%s\n' "$actual"); then echo "PASS: snowball output matches expected (4 new candidates; 1 backward dup removed)." else echo "FAIL: snowball output drifted from expected/snowball.bib" >&2 exit 1 fi # A count is only a PRISMA count if the source answered in full. A fetch that failed, or a page # the source says continues (`next`), used to print "N raw ... M new candidates" and exit 0, so a # lower bound went into the flow diagram as the number that exists. fail=0 ERR="$(mktemp)" trap 'rm -f "$ERR"' EXIT run() { # run <seed> <limit> -> exit code; stderr lands in $ERR set +e python3 "$SNOWBALL" --seed "$1" --direction all --offline-fixture "$HERE/fixture" \ --limit "$2" --as-of 2026-06-14 --stdout >/dev/null 2>"$ERR" local rc=$? set -e echo "$rc" } rc="$(run DOI:10.0/seed_err 50)" if [ "$rc" -ne 0 ] && grep -q "FAILED backward" "$ERR" && grep -q "INCOMPLETE" "$ERR"; then echo "PASS: a Semantic Scholar error body exits non-zero and marks the PRISMA line INCOMPLETE." else echo "FAIL: an error body was read as 0 records (exit $rc)" >&2; cat "$ERR" >&2; fail=1 fi rc="$(run DOI:10.0/seed_trunc 2)" if [ "$rc" -ne 0 ] && grep -q "TRUNCATED forward" "$ERR" && grep -q "INCOMPLETE" "$ERR"; then echo "PASS: a page with 'next' (more records past --limit) prints TRUNCATED and exits non-zero." else echo "FAIL: a truncated page was reported as complete (exit $rc)" >&2; cat "$ERR" >&2; fail=1 fi # Negative control: a page exactly at --limit with no `next` is complete. rc="$(run DOI:10.0/seed1 2)" if [ "$rc" -eq 0 ] && ! grep -qE "TRUNCATED|FAILED|INCOMPLETE" "$ERR"; then echo "PASS: a full page with no 'next' is complete (exit 0, no TRUNCATED)." else echo "FAIL: a complete page was called incomplete (exit $rc)" >&2; cat "$ERR" >&2; fail=1 fi # A live fetch that raises (network down, 403, 429) is a failure, not zero results. set +e python3 - "$HERE/.." >"$ERR" 2>&1 <<'PY' import sys sys.path.insert(0, sys.argv[1]) import snowball def boom(url): raise OSError("HTTP Error 403: Forbidden") snowball._http_get_json = boom rc = snowball.main(["--seed", "10.1000/xyz", "--direction", "backward", "--stdout"]) sys.exit(10 + rc) PY rc=$? set -e if [ "$rc" -eq 11 ] && grep -q "FAILED backward for DOI:10.1000/xyz" "$ERR" && grep -q "INCOMPLETE" "$ERR"; then echo "PASS: a live fetch that raises exits non-zero and marks the PRISMA line INCOMPLETE." else echo "FAIL: a failed live fetch was read as 0 records (main returned $((rc-10)))" >&2; cat "$ERR" >&2; fail=1 fi [ "$fail" -eq 0 ] || exit 1
-
-
embase_browser.md 1.3 KB
# Embase search via browser automation Embase has no public API. Search and export it through Chrome browser automation (MCP): 1. Navigate to `embase.com` — institutional SSO authenticates automatically. If cookie error (`login?error#`), clear Elsevier/Embase cookies and retry. 2. Go to the **Advanced Search** tab. 3. Enter an Embase-syntax query (Emtree `/exp` + `:ab,ti` field tags). Uncheck "Map to preferred term in Emtree" when using explicit `/exp` terms. 4. After results appear, use the "Select number of items" dropdown → select the total count. 5. Click **Export** (in the Results section) → choose **CSV** format → check fields: Title, Author names, Source, Publication year, Publication type, DOI, Abstract, Language of article, Medline PMID. 6. Click Export → the Download tab opens → click Download. 7. The CSV is in **row format** (records separated by blank rows). Parse it as: ```python # Each record = consecutive rows until blank row # Row format: [FIELD_NAME, value1, value2, ...] # AUTHOR NAMES row has multiple values (one per author) ``` ## PubMed → Embase query translation - MeSH `[Mesh]` → Emtree `/exp` - `[tiab]` → `:ab,ti` - `[Title/Abstract]` → `:ab,ti` - Boolean operators stay the same (AND, OR) - Phrase search: use single quotes in Embase (`'artificial ascites'`) -
parse_pubmed.py 13.8 KB
#!/usr/bin/env python3 """ Parse PubMed E-utilities responses into structured data. Usage: # Parse esearch JSON → list of PMIDs echo '<json>' | python3 parse_pubmed.py esearch # Parse esummary JSON → markdown table echo '<json>' | python3 parse_pubmed.py esummary # Parse efetch XML → detailed metadata (for BibTeX generation) echo '<xml>' | python3 parse_pubmed.py efetch # Parse efetch XML → BibTeX entries echo '<xml>' | python3 parse_pubmed.py bibtex """ import sys import json import re import xml.etree.ElementTree as ET from datetime import date from textwrap import shorten # Heuristic for East Asian name reverse-encoding in PubMed XML. # Cases observed: <LastName>Qiaoling</LastName><ForeName>Fu</ForeName> where # Fu is the actual family name. Pattern: LastName looks like a long given # name (≥3 alpha chars, no spaces) AND ForeName looks like a short surname # fragment (1-2 chars, no period). The naive test catches the common reverse # encoding without flagging legitimate short-surname authors. _EAST_ASIAN_REVERSE_THRESHOLD = 3 # LastName length lower bound for suspicion def _full_text(parent, path: str, default: str = "") -> str: """All text inside the element at `path`, including text inside inline children. PubMed marks up titles and abstracts inline (`<i>Helicobacter pylori</i>`, `CO<sub>2</sub>`). `findtext()` / `.text` return only the text before the first child element, which cut "Eradication of <i>Helicobacter pylori</i> infection." down to "Eradication of " while the BibTeX entry was still stamped verified. Absent element -> `default`. """ el = parent.find(path) if parent is not None else None if el is None: return default return "".join(el.itertext()) def _looks_east_asian_reversed(last: str, fore: str) -> bool: """Return True if (LastName, ForeName) look swapped per PubMed encoding bug.""" if not last or not fore: return False # ForeName should look like a surname (1-2 chars, no spaces, no period) # AND LastName should look like a multi-char given name. return ( 1 <= len(fore) <= 2 and fore.isalpha() and "." not in fore and len(last) >= _EAST_ASIAN_REVERSE_THRESHOLD and last.isalpha() ) def _extract_authors(author_list_el): """Walk an <AuthorList> element. Return (bib_authors, display_authors, first_author_last, suspicions, has_collective_only). - bib_authors: list of "Family, Given" strings for BibTeX `author = {...}`. Corporate (<CollectiveName>) authors are double-braced. - display_authors: list of "Last First" strings for human-readable output. - first_author_last: surname of the first listed author (used for cite key). - suspicions: list of human-readable warning strings (East Asian reverse encoding, missing LastName, etc.). - has_collective_only: True if AuthorList contains only <CollectiveName> entries (no individual <LastName>). Caller should consider emitting `@misc` instead of `@article` (guideline / consortium pattern). """ bib_authors: list[str] = [] display_authors: list[str] = [] first_author_last = "" suspicions: list[str] = [] individual_count = 0 collective_count = 0 if author_list_el is None: return bib_authors, display_authors, "", suspicions, False for au in author_list_el.findall("Author"): last = au.findtext("LastName", "") or "" fore = au.findtext("ForeName", "") or "" collective = au.findtext("CollectiveName", "") or "" if collective: collective_count += 1 # Double-brace to prevent BibTeX from splitting on the comma / # spaces inside the corporate name. bib_authors.append("{" + collective + "}") display_authors.append(collective) if not first_author_last: first_author_last = re.sub(r"[^A-Za-z]+", "", collective.split()[0]) or "Group" continue if last: individual_count += 1 if _looks_east_asian_reversed(last, fore): suspicions.append( f"East Asian name order suspected for '{last} {fore}' — " "PubMed XML may have LastName/ForeName swapped" ) bib_authors.append(f"{last}, {fore}") display_authors.append(f"{last} {fore}".strip()) if not first_author_last: first_author_last = last continue # Author element with neither <LastName> nor <CollectiveName>: rare # but possible. Record as suspicion, otherwise skip. suspicions.append("Author element with no LastName and no CollectiveName") has_collective_only = collective_count > 0 and individual_count == 0 return bib_authors, display_authors, first_author_last, suspicions, has_collective_only def parse_esearch(data: str) -> None: """Parse esearch JSON response, print PMIDs and count.""" # A body without `esearchresult.count`, or with an `ERROR` field, is not a search that # found nothing: it is a search that did not run. Defaulting the count to "0" printed # "Total results: 0" with exit 0 for an error body, which a caller reads as "no papers" # (ma-scout's "MA = 0" verdict). Name the input and exit 2 instead. result = json.loads(data) esearch = result.get("esearchresult") if isinstance(result, dict) else None problem = None if not isinstance(esearch, dict): problem = "no 'esearchresult' object" if isinstance(result, dict) and result.get("error"): problem += f" (error: {result.get('error')})" elif esearch.get("ERROR"): problem = f"esearchresult.ERROR: {esearch.get('ERROR')}" elif "count" not in esearch: problem = "esearchresult has no 'count'" if problem is not None: print( f"ERROR: esearch response is not a search result ({problem}); " "the count is unknown, not 0.", file=sys.stderr, ) sys.exit(2) count = esearch["count"] ids = esearch.get("idlist", []) print(f"Total results: {count}") print(f"Returned: {len(ids)}") print(f"PMIDs: {','.join(ids)}") def parse_esummary(data: str) -> None: """Parse esummary JSON response into a markdown table.""" result = json.loads(data) docs = result.get("result", {}) uids = docs.get("uids", []) if not uids: print("No results found.") return print("| # | PMID | Year | Journal | Title | Authors |") print("|---|------|------|---------|-------|---------|") for i, uid in enumerate(uids, 1): doc = docs.get(uid, {}) title = shorten(doc.get("title", "N/A"), width=80, placeholder="...") authors_raw = doc.get("authors", []) if authors_raw: first = authors_raw[0].get("name", "") last = authors_raw[-1].get("name", "") if len(authors_raw) > 1 else "" authors = f"{first}, ... {last}" if last and last != first else first else: authors = "N/A" journal = shorten(doc.get("fulljournalname", doc.get("source", "N/A")), width=40, placeholder="...") pubdate = doc.get("pubdate", "N/A") year = pubdate[:4] if pubdate else "N/A" doi_list = doc.get("articleids", []) doi = next((d["value"] for d in doi_list if d.get("idtype") == "doi"), "") print(f"| {i} | {uid} | {year} | {journal} | {title} | {authors} |") print(f"\n*{len(uids)} articles retrieved*") def parse_efetch(data: str) -> None: """Parse efetch XML response into structured metadata.""" root = ET.fromstring(data) articles = root.findall(".//PubmedArticle") for article in articles: medline = article.find("MedlineCitation") if medline is None: continue pmid = medline.findtext("PMID", "N/A") art = medline.find("Article") if art is None: continue title = _full_text(art, "ArticleTitle", "N/A") journal_el = art.find("Journal") journal = journal_el.findtext("Title", "N/A") if journal_el is not None else "N/A" journal_abbrev = journal_el.findtext("ISOAbbreviation", "") if journal_el is not None else "" # Year ji = journal_el.find("JournalIssue") if journal_el is not None else None pd = ji.find("PubDate") if ji is not None else None year = pd.findtext("Year", "") if pd is not None else "" if not year: medline_date = pd.findtext("MedlineDate", "") if pd is not None else "" year = medline_date[:4] if medline_date else "N/A" volume = ji.findtext("Volume", "") if ji is not None else "" issue = ji.findtext("Issue", "") if ji is not None else "" # Pages pages = art.findtext("Pagination/MedlinePgn", "") # Authors (handles East Asian reverse encoding + CollectiveName) author_list = art.find("AuthorList") _, authors, _, suspicions, _ = _extract_authors(author_list) # DOI doi = "" for aid in art.findall("ELocationID"): if aid.get("EIdType") == "doi": doi = aid.text or "" # Abstract abstract_el = art.find("Abstract") abstract = "" if abstract_el is not None: parts = abstract_el.findall("AbstractText") abstract = " ".join( (p.get("Label", "") + ": " if p.get("Label") else "") + "".join(p.itertext()) for p in parts ) print(f"## PMID: {pmid}") print(f"**Title**: {title}") print(f"**Authors**: {'; '.join(authors)}") print(f"**Journal**: {journal} ({journal_abbrev})") print(f"**Year**: {year} **Volume**: {volume} **Issue**: {issue} **Pages**: {pages}") print(f"**DOI**: {doi}") if abstract: print(f"**Abstract**: {shorten(abstract, width=500, placeholder='...')}") for note in suspicions: print(f"> ⚠ {note}") print() def generate_bibtex(data: str) -> None: """Parse efetch XML and generate BibTeX entries.""" root = ET.fromstring(data) articles = root.findall(".//PubmedArticle") for article in articles: medline = article.find("MedlineCitation") if medline is None: continue pmid = medline.findtext("PMID", "") art = medline.find("Article") if art is None: continue title = _full_text(art, "ArticleTitle", "") journal_el = art.find("Journal") journal_abbrev = journal_el.findtext("ISOAbbreviation", "") if journal_el is not None else "" journal_full = journal_el.findtext("Title", "") if journal_el is not None else "" ji = journal_el.find("JournalIssue") if journal_el is not None else None pd = ji.find("PubDate") if ji is not None else None year = pd.findtext("Year", "") if pd is not None else "" if not year: md = pd.findtext("MedlineDate", "") if pd is not None else "" year = md[:4] if md else "" volume = ji.findtext("Volume", "") if ji is not None else "" issue = ji.findtext("Issue", "") if ji is not None else "" pages = art.findtext("Pagination/MedlinePgn", "") doi = "" for aid in art.findall("ELocationID"): if aid.get("EIdType") == "doi": doi = aid.text or "" author_list = art.find("AuthorList") bib_authors, _, first_author_last, suspicions, has_collective_only = \ _extract_authors(author_list) # Generate citation key key = f"{first_author_last}_{year}_{pmid}" if first_author_last else f"PMID_{pmid}" # Corporate / consortium guideline (e.g., KDIGO, AHA/ACC) — emit as # @misc so BibTeX styles render the body of the entry without trying # to format a personal author. Vancouver / AMA CSL handle both # @article and @misc with author = {{Organization Name}}. entry_type = "misc" if has_collective_only else "article" # Prepend suspicion comments so they survive .bib copy/paste audits. for note in suspicions: print(f"% [VERIFY] {note}") print(f"@{entry_type}{{{key},") print(f" author = {{{' and '.join(bib_authors)}}},") print(f" title = {{{title}}},") print(f" journal = {{{journal_full}}},") print(f" year = {{{year}}},") if volume: print(f" volume = {{{volume}}},") if issue: print(f" number = {{{issue}}},") if pages: print(f" pages = {{{pages}}},") if doi: print(f" doi = {{{doi}}},") print(f" pmid = {{{pmid}}},") # Anti-hallucination verification flag. Entries emitted by this script # originate from PubMed efetch XML, so a non-empty PMID is proof of # API provenance (verified=true). Missing PMID → verified=false and # downstream tooling (/verify-refs) will flag for manual check. verified = bool(pmid) verified_by = "pubmed+crossref" if (pmid and doi) else ("pubmed" if pmid else "") print(f" verified = {{{'true' if verified else 'false'}}},") if verified_by: print(f" verified_by = {{{verified_by}}},") print(f" verified_on = {{{date.today().isoformat()}}},") print("}") print() if __name__ == "__main__": if len(sys.argv) < 2: print(__doc__) sys.exit(1) mode = sys.argv[1] data = sys.stdin.read() dispatch = { "esearch": parse_esearch, "esummary": parse_esummary, "efetch": parse_efetch, "bibtex": generate_bibtex, } func = dispatch.get(mode) if func is None: print(f"Unknown mode: {mode}. Use: {', '.join(dispatch.keys())}") sys.exit(1) func(data) -
pubmed_eutils.sh 4.2 KB
#!/bin/bash # PubMed E-utilities CLI wrapper # Fallback when PubMed MCP server is unavailable # Usage: bash pubmed_eutils.sh <command> <args...> # # Commands: # search <query> [retmax] -- Search PubMed, return PMIDs # fetch <pmid1,pmid2,...> -- Fetch article metadata (XML) # fetch_json <pmid1,pmid2,...> -- Fetch article summary (JSON, DocSum) # related <pmid> [retmax] -- Find related articles # cite_lookup <title> -- Search by exact title to verify citation # # Environment: # NCBI_API_KEY -- Optional. Increases rate limit from 3/sec to 10/sec. # Register at https://www.ncbi.nlm.nih.gov/account/settings/ # # Rate limiting: 3 requests/second without API key, 10/sec with key. # This script sleeps 350ms between calls for safety. set -euo pipefail BASE="https://eutils.ncbi.nlm.nih.gov/entrez/eutils" TOOL="claude-code-search-lit" EMAIL="noreply@example.com" DB="pubmed" SLEEP=0.35 # Build API key param if available API_KEY_PARAM="" if [ -n "${NCBI_API_KEY:-}" ]; then API_KEY_PARAM="&api_key=${NCBI_API_KEY}" SLEEP=0.1 fi _sleep() { sleep "$SLEEP"; } # Values reach Python as argv, never as source code. A query is free text ("Crohn's disease", # "O'Brien[Author]"): pasted into a Python string literal, the first apostrophe was a SyntaxError # that left the search term silently empty, and a crafted query ran its own Python. _urlencode() { python3 -c 'import sys, urllib.parse; print(urllib.parse.quote(sys.argv[1]))' "$1" } _count() { # _count <value> <name>: retmax must be a whole number or the command stops case "$1" in '' | *[!0-9]*) echo "{\"error\": \"$2 must be a whole number\"}" >&2; return 2 ;; esac } _curl() { local http_code body body=$(curl -sS -w '\n%{http_code}' -A "Mozilla/5.0 (${TOOL})" "$@") http_code=$(echo "$body" | tail -n1) body=$(echo "$body" | sed '$d') if [ "$http_code" -ge 400 ] 2>/dev/null; then echo "{\"error\": \"HTTP ${http_code}\", \"url\": \"$1\"}" >&2 return 1 fi echo "$body" } cmd_search() { local query="${1:?Usage: search <query> [retmax]}" local retmax="${2:-20}" term _count "$retmax" retmax # A separate assignment, so an encoding failure stops the script instead of searching term=. term="$(_urlencode "$query")" local url="${BASE}/esearch.fcgi?db=${DB}&term=${term}&retmax=${retmax}&retmode=json&tool=${TOOL}&email=${EMAIL}${API_KEY_PARAM}" _curl "$url" } cmd_fetch() { local ids="${1:?Usage: fetch <pmid1,pmid2,...>}" local url="${BASE}/efetch.fcgi?db=${DB}&id=${ids}&rettype=xml&retmode=xml&tool=${TOOL}&email=${EMAIL}${API_KEY_PARAM}" _curl "$url" } cmd_fetch_json() { local ids="${1:?Usage: fetch_json <pmid1,pmid2,...>}" local url="${BASE}/esummary.fcgi?db=${DB}&id=${ids}&retmode=json&tool=${TOOL}&email=${EMAIL}${API_KEY_PARAM}" _curl "$url" } cmd_related() { local pmid="${1:?Usage: related <pmid> [retmax]}" local retmax="${2:-10}" _count "$retmax" retmax local url="${BASE}/elink.fcgi?dbfrom=${DB}&db=${DB}&id=${pmid}&cmd=neighbor_score&retmode=json&tool=${TOOL}&email=${EMAIL}${API_KEY_PARAM}" local result result=$(_curl "$url") # Extract linked PMIDs and fetch their summaries local linked_ids linked_ids=$(echo "$result" | python3 -c " import sys, json retmax = int(sys.argv[1]) data = json.load(sys.stdin) links = data.get('linksets', [{}])[0].get('linksetdbs', [{}]) for db in links: if db.get('linkname') == 'pubmed_pubmed': ids = [str(l['id']) for l in db.get('links', [])[:retmax]] print(','.join(ids)) break " "$retmax" 2>/dev/null || echo "") if [ -n "$linked_ids" ]; then _sleep cmd_fetch_json "$linked_ids" else echo '{"error": "No related articles found"}' fi } cmd_cite_lookup() { local title="${1:?Usage: cite_lookup <title>}" # Search by title field for exact verification cmd_search "${title}[Title]" 5 } # Dispatch case "${1:-help}" in search) shift; cmd_search "$@" ;; fetch) shift; cmd_fetch "$@" ;; fetch_json) shift; cmd_fetch_json "$@" ;; related) shift; cmd_related "$@" ;; cite_lookup) shift; cmd_cite_lookup "$@" ;; help|*) echo "Usage: bash pubmed_eutils.sh <command> <args...>" echo "Commands: search, fetch, fetch_json, related, cite_lookup" ;; esac -
snowball.py 14.6 KB
#!/usr/bin/env python3 """ Citation snowballing for search-lit (Phase 2.5 Citation Searching). Expands a seed set of papers along the citation graph and emits API-verified BibTeX candidates, deduplicated against an existing candidate pool. Directions ---------- backward : references the seed papers cite (cited-by-the-seed) forward : papers that cite the seed (citing-the-seed) similar : Semantic Scholar recommendations for the seed all : backward + forward + similar (default) Data source: Semantic Scholar Graph API (deterministic, no model memory). references GET /graph/v1/paper/{id}/references citations GET /graph/v1/paper/{id}/citations recommendations GET /recommendations/v1/papers/forpaper/{id} Anti-hallucination contract --------------------------- Snowball candidates carry `verified=false` + `verified_by=semantic_scholar`. They are NOT cross-checked against PubMed/CrossRef here; downstream `/verify-refs` confirms each entry and upgrades the flag. Nothing is ever generated from memory. Output contract (matches search-lit BibTeX section) --------------------------------------------------- - BibTeX is APPENDED to the candidate pool (default references/library.bib). - NEVER writes to manuscript/_src/refs.bib (that is /lit-sync's sole path). - A PRISMA "records identified through citation searching" line is printed. - Exit 0 only when every seed/direction answered in full. A failed fetch (network error, an error body) or a page the source says continues (`next`, i.e. more than --limit records) is printed as FAILED / TRUNCATED, the PRISMA line ends with INCOMPLETE, and the exit is 1. Usage ----- # Live (network) — expand one DOI in all directions, dedup against pool python3 snowball.py --seed DOI:10.1000/synthetic.example \ --pool references/library.bib --out references/library.bib # Multiple seeds from a file (one id per line), backward only python3 snowball.py --seed @seeds.txt --direction backward # Deterministic / offline (challenge-card verifier): read recorded JSON python3 snowball.py --seed DOI:10.0/seed1 --direction all \ --offline-fixture fixture --as-of 2026-06-14 --stdout Seed id formats accepted: `DOI:10.x/...`, `PMID:123456`, bare `10.x/...` (treated as DOI), bare digits (treated as PMID), or a raw S2 paper id. """ from __future__ import annotations import argparse import json import re import sys import urllib.parse import urllib.request from datetime import date from pathlib import Path S2_GRAPH = "https://api.semanticscholar.org/graph/v1/paper" S2_REC = "https://api.semanticscholar.org/recommendations/v1/papers/forpaper" S2_FIELDS = "title,year,venue,externalIds,authors.name" DIRECTIONS = ("backward", "forward", "similar") # --------------------------------------------------------------------------- # # Seed id normalization # --------------------------------------------------------------------------- # def normalize_seed_id(raw: str) -> str: """Return an S2-acceptable paper id token (DOI:/PMID:/raw).""" s = raw.strip() if not s: return "" up = s.upper() if up.startswith("DOI:") or up.startswith("PMID:") or up.startswith("ARXIV:"): return s if s.startswith("10.") or "doi.org/" in s.lower(): return "DOI:" + s.split("doi.org/")[-1] if s.isdigit(): return "PMID:" + s return s # assume raw S2 id def fixture_slug(seed_id: str) -> str: """Filesystem-safe slug for an offline fixture filename.""" return re.sub(r"[^A-Za-z0-9]+", "_", seed_id).strip("_") # --------------------------------------------------------------------------- # # Fetch (network or offline fixture) # --------------------------------------------------------------------------- # def _http_get_json(url: str) -> dict: req = urllib.request.Request(url, headers={"User-Agent": "medsci-skills/snowball"}) with urllib.request.urlopen(req, timeout=30) as resp: # noqa: S310 (trusted host) return json.loads(resp.read().decode("utf-8")) def fetch_direction(seed_id: str, direction: str, limit: int, fixture_dir: Path | None) -> tuple[list[dict], str | None]: """Return (papers, problem) for one seed+direction. Each paper dict has at least: title, year, venue, externalIds, authors. `problem` is None when the source answered in full, otherwise a one-line reason starting with `FAILED` (the fetch did not run, or its answer was not a result) or `TRUNCATED` (the source says more records exist beyond `--limit`). Either way the count for this seed+direction is a lower bound, not the number that exists, and the caller must say so. """ if fixture_dir is not None: fpath = fixture_dir / f"{fixture_slug(seed_id)}.{direction}.json" if not fpath.exists(): return [], None try: payload = json.loads(fpath.read_text()) except (OSError, ValueError) as exc: return [], f"FAILED {direction} for {seed_id}: unreadable fixture {fpath.name} ({exc})" else: enc = urllib.parse.quote(seed_id, safe="") if direction == "backward": url = f"{S2_GRAPH}/{enc}/references?fields={S2_FIELDS}&limit={limit}" elif direction == "forward": url = f"{S2_GRAPH}/{enc}/citations?fields={S2_FIELDS}&limit={limit}" elif direction == "similar": url = f"{S2_REC}/{enc}?fields={S2_FIELDS}&limit={limit}" else: raise ValueError(f"unknown direction: {direction}") try: payload = _http_get_json(url) except Exception as exc: # noqa: BLE001 — reported to the caller, never read as 0 return [], f"FAILED {direction} for {seed_id}: {exc}" # Normalize payload shapes: # references/citations: {"data": [{"citedPaper"|"citingPaper": {...}}], "next": N?} # recommendations: {"recommendedPapers": [{...}]} or {"data": [{...}]} # A body with neither list (an `error` payload, say) is not "no papers": the search did not run. if not isinstance(payload, dict) or not ( isinstance(payload.get("data"), list) or isinstance(payload.get("recommendedPapers"), list)): detail = payload.get("error") or payload.get("message") if isinstance(payload, dict) else None return [], (f"FAILED {direction} for {seed_id}: response has no 'data' or " f"'recommendedPapers' list" + (f" ({detail})" if detail else "")) rows = payload.get("data") or payload.get("recommendedPapers") or [] papers = [] for row in rows: paper = row.get("citedPaper") or row.get("citingPaper") or row if isinstance(paper, dict) and paper.get("title"): papers.append(paper) # The paginated endpoints carry `next` only when more records exist past this page. if payload.get("next") is not None: return papers, (f"TRUNCATED {direction} for {seed_id}: the source has more records " f"after the first {len(rows)} (next offset {payload.get('next')}); " "raise --limit") return papers, None # --------------------------------------------------------------------------- # # Dedup # --------------------------------------------------------------------------- # def norm_doi(doi: str | None) -> str: if not doi: return "" return doi.strip().lower().replace("https://doi.org/", "") def norm_title(title: str | None) -> str: if not title: return "" return re.sub(r"[^a-z0-9]+", "", title.lower()) def parse_pool_keys(pool_path: Path | None) -> tuple[set[str], set[str]]: """Extract existing DOIs and normalized titles from a BibTeX pool.""" dois: set[str] = set() titles: set[str] = set() if pool_path is None or not pool_path.exists(): return dois, titles text = pool_path.read_text(errors="ignore") for m in re.finditer(r"doi\s*=\s*[{\"]([^}\"]+)[}\"]", text, re.I): dois.add(norm_doi(m.group(1))) for m in re.finditer(r"title\s*=\s*[{\"](.+?)[}\"]\s*,?\s*\n", text, re.I | re.S): titles.add(norm_title(m.group(1))) return dois, titles # --------------------------------------------------------------------------- # # BibTeX emit # --------------------------------------------------------------------------- # def _author_last(name: str) -> str: parts = name.strip().split() return parts[-1] if parts else "Anon" def bibtex_key(paper: dict) -> str: authors = paper.get("authors") or [] last = _author_last(authors[0]["name"]) if authors and authors[0].get("name") else "Anon" last = re.sub(r"[^A-Za-z]", "", last) or "Anon" year = str(paper.get("year") or "ND") title_word = "" for w in re.findall(r"[A-Za-z]{4,}", paper.get("title") or ""): if w.lower() not in {"with", "from", "using", "study", "analysis", "based"}: title_word = w.capitalize() break return f"{last}_{year}_{title_word or 'Snowball'}" def to_bibtex(paper: dict, direction: str, as_of: str, key: str) -> str: ext = paper.get("externalIds") or {} doi = ext.get("DOI", "") pmid = ext.get("PubMed", "") authors = paper.get("authors") or [] author_str = " and ".join( f"{_author_last(a['name'])}, {' '.join(a['name'].split()[:-1])}".strip().rstrip(",") for a in authors if a.get("name") ) or "Unknown" lines = [ f"@article{{{key},", f" author = {{{author_str}}},", f" title = {{{paper.get('title', '').strip()}}},", f" journal = {{{paper.get('venue', '') or ''}}},", f" year = {{{paper.get('year') or ''}}},", ] if doi: lines.append(f" doi = {{{doi}}},") if pmid: lines.append(f" pmid = {{{pmid}}},") lines += [ " verified = {false},", " verified_by = {semantic_scholar},", f" verified_on = {{{as_of}}},", " source = {citation_snowball},", f" snowball_direction = {{{direction}}},", "}", ] return "\n".join(lines) # --------------------------------------------------------------------------- # # Main # --------------------------------------------------------------------------- # def load_seeds(raw: str) -> list[str]: if raw.startswith("@"): lines = Path(raw[1:]).read_text().splitlines() items = [ln for ln in lines if ln.strip() and not ln.strip().startswith("#")] else: items = re.split(r"[,\s]+", raw) return [normalize_seed_id(x) for x in items if x.strip()] def main(argv: list[str] | None = None) -> int: ap = argparse.ArgumentParser(description="Citation snowballing for search-lit.") ap.add_argument("--seed", required=True, help="comma/space-separated ids, or @file with one id per line") ap.add_argument("--direction", default="all", choices=("all", *DIRECTIONS)) ap.add_argument("--pool", default=None, help="existing library.bib to dedup against") ap.add_argument("--out", default="references/library.bib", help="BibTeX append target (NEVER manuscript/_src/refs.bib)") ap.add_argument("--limit", type=int, default=50, help="max per seed per direction") ap.add_argument("--offline-fixture", default=None, help="dir of recorded JSON (<slug>.<direction>.json) for deterministic runs") ap.add_argument("--as-of", default=date.today().isoformat(), help="verified_on date stamp (default today; set for reproducible output)") ap.add_argument("--stdout", action="store_true", help="print BibTeX to stdout instead of appending to --out") args = ap.parse_args(argv) out_path = Path(args.out) if out_path.name == "refs.bib" and "_src" in str(out_path): ap.error("refusing to write manuscript/_src/refs.bib — that is /lit-sync's sole path") fixture_dir = Path(args.offline_fixture) if args.offline_fixture else None pool_dois, pool_titles = parse_pool_keys(Path(args.pool) if args.pool else None) seeds = load_seeds(args.seed) directions = list(DIRECTIONS) if args.direction == "all" else [args.direction] seen_doi = set(pool_dois) seen_title = set(pool_titles) seen_key: set[str] = set() counts = {d: 0 for d in directions} entries: list[str] = [] raw_found = 0 problems: list[str] = [] for seed in seeds: for direction in directions: papers, problem = fetch_direction(seed, direction, args.limit, fixture_dir) if problem: problems.append(problem) for paper in papers: raw_found += 1 ext = paper.get("externalIds") or {} d = norm_doi(ext.get("DOI")) t = norm_title(paper.get("title")) if (d and d in seen_doi) or (t and t in seen_title): continue if d: seen_doi.add(d) if t: seen_title.add(t) key = bibtex_key(paper) base_key, n = key, 1 while key in seen_key: n += 1 key = f"{base_key}{chr(96 + n)}" # _b, _c, ... seen_key.add(key) entries.append(to_bibtex(paper, direction, args.as_of, key)) counts[direction] += 1 new_total = len(entries) bibtex_blob = ("\n\n".join(entries) + "\n") if entries else "" if args.stdout or not entries: sys.stdout.write(bibtex_blob) else: out_path.parent.mkdir(parents=True, exist_ok=True) with out_path.open("a", encoding="utf-8") as fh: if out_path.exists() and out_path.stat().st_size > 0: fh.write("\n") fh.write(bibtex_blob) # PRISMA citation-searching line (stderr so --stdout BibTeX stays clean) breakdown = ", ".join(f"{d}={counts[d]}" for d in directions) pool_n = len(pool_dois) + len(pool_titles) for problem in problems: sys.stderr.write(f"[snowball] {problem}\n") incomplete = ( f" INCOMPLETE: {len(problems)} seed/direction fetch(es) failed or were truncated; " "these counts are a lower bound and must not be recorded in the PRISMA flow." if problems else "" ) sys.stderr.write( f"Records identified through citation searching (snowballing): " f"{raw_found} raw ({breakdown}); after dedup against existing pool: " f"{new_total} new candidates.{incomplete}\n" ) sys.stderr.write( f"[snowball] seeds={len(seeds)} directions={'+'.join(directions)} " f"pool_keys={pool_n} -> {new_total} appended" f"{' (stdout)' if args.stdout else f' to {out_path}'}\n" ) return 1 if problems else 0 if __name__ == "__main__": raise SystemExit(main())
-
-
scripts
-
check_doi_record_match_challenge
-
fixture
-
cache
-
10.1000%2Farticle.008.json 278 B
{ "status": "ok", "message": { "type": "journal-article", "title": [ "Cardiac magnetic resonance in acute myocarditis: a multicentre cohort" ], "updated-by": [ { "type": "correction", "DOI": "10.1000/erratum.007", "label": "Correction" } ] } } -
10.1000%2Ferratum.007.json 292 B
{ "status": "ok", "message": { "type": "journal-article", "title": [ "Correction to: Cardiac magnetic resonance in acute myocarditis: a multicentre cohort" ], "update-to": [ { "type": "correction", "DOI": "10.1000/article.008", "label": "Correction" } ] } } -
10.1000%2Fmarkup.005.json 180 B
{ "status": "ok", "message": { "type": "journal-article", "title": [ "Transfer Learning for <i>Chest Radiograph</i> Classification in Low\u2013Resource Settings" ] } } -
10.1000%2Fmatch.001.json 183 B
{ "status": "ok", "message": { "type": "journal-article", "title": [ "Deep learning for intracranial aneurysm detection on CT angiography: a multicentre validation" ] } } -
10.1000%2Fother.002.json 154 B
{ "status": "ok", "message": { "type": "journal-article", "title": [ "Diffusion tensor imaging of the corticospinal tract after stroke" ] } } -
10.1000%2Fsuppl.003.json 135 B
{ "status": "ok", "message": { "type": "journal-issue", "title": [ "Abstracts of the 2026 Annual Scientific Meeting" ] } } -
10.1000%2Fversion.009.json 265 B
{ "status": "ok", "message": { "type": "posted-content", "title": [ "Photon-counting CT of the temporal bone: a reader study" ], "update-to": [ { "type": "new_version", "DOI": "10.1000/version.009a", "label": "New version" } ] } }
-
-
screening.tsv 615 B · in bundle
-
update_notice.tsv 444 B · in bundle
-
-
verify.sh 5.9 KB
#!/usr/bin/env bash # Deterministic verifier for the DOI-belongs-to-this-record challenge card. Network-free: the # fixture ships recorded Crossref responses and the detector reads them with --cache. # # Six rows, one per way this goes wrong or is falsely accused: # # R-001 the DOI is right SILENT # R-002 the DOI belongs to another study DOI_NOT_THIS_RECORD <- the false "missed paper" # R-003 the DOI is a whole abstract supplement DOI_IS_CONTAINER <- the paper that isn't there # R-004 the DOI does not resolve DOI_UNRESOLVED <- reported, not dropped # R-005 no DOI at all SILENT (nothing to check is not a finding) # R-006 same paper, JATS markup + en dash SILENT <- the one that would kill this check # # R-006 is the load-bearing negative. Crossref returns titles with <i> tags, typographic dashes and # its own capitalisation; a comparison that does not normalise those calls EVERY row a mismatch, # and a check that flags everything is switched off within a day. set -euo pipefail HERE="$(cd "$(dirname "$0")" && pwd)" DET="$HERE/../check_doi_record_match.py" FIX="$HERE/fixture" fail=0 pass() { echo "PASS $1"; } bad() { echo "FAIL $1"; fail=1; } WORK="$(mktemp -d)" trap 'rm -rf "$WORK"' EXIT out="$(python3 "$DET" --table "$FIX/screening.tsv" --cache "$FIX/cache" || true)" want_verdict() { # record, verdict, description if grep -qE "\[$2\] \($1\)" <<<"$out"; then pass "$3" else bad "$3" echo "$out" fi } want_verdict R-002 DOI_NOT_THIS_RECORD "a DOI belonging to another study is caught" want_verdict R-003 DOI_IS_CONTAINER "a DOI resolving to a whole supplement is caught" want_verdict R-004 DOI_UNRESOLVED "an unresolvable DOI is reported, not dropped" for r in R-001 R-005 R-006; do if grep -q "($r)" <<<"$out"; then bad "$r was flagged and should not have been" echo "$out" else case "$r" in R-001) pass "a correct DOI is silent" ;; R-005) pass "a row with no DOI is not a finding (there is nothing to check)" ;; R-006) pass "JATS markup, an en dash and different capitalisation are the SAME title" ;; esac fi done # The mismatch has to be readable, or nobody can adjudicate it: both titles, and the measured # similarity that decided it. if grep -q "resolved title:" <<<"$out" && grep -qE "similarity 0\.[0-9]{2} <" <<<"$out"; then pass "...and the mismatch prints both titles and the similarity that decided it" else bad "the mismatch report is not adjudicable" echo "$out" fi # A cache with no recorded response is UNRESOLVED, never a silent live lookup: a run that is half # recorded and half network is reproducible in neither direction. mkdir -p "$WORK/empty" n_unresolved="$(python3 "$DET" --table "$FIX/screening.tsv" --cache "$WORK/empty" \ | grep -c '\[DOI_UNRESOLVED\]' || true)" if [ "$n_unresolved" -eq 5 ]; then pass "an empty cache turns all 5 DOIs into UNRESOLVED and never reaches the network" else bad "an empty cache produced $n_unresolved UNRESOLVED, expected 5" python3 "$DET" --table "$FIX/screening.tsv" --cache "$WORK/empty" | head -20 fi # Missing columns must stop the run. Reading a table whose doi column is called something else and # reporting "0 findings" is the failure mode this whole card exists to prevent, one level up. set +e python3 "$DET" --table "$FIX/screening.tsv" --cache "$FIX/cache" --doi-col identifier \ >/dev/null 2>"$WORK/err"; rc=$? set -e if [ "$rc" -eq 2 ] && grep -q "identifier" "$WORK/err"; then pass "a wrong --doi-col name exits 2 and names it (not a hollow 'no findings')" else bad "exit was $rc for a column that does not exist" cat "$WORK/err" fi if python3 "$DET" --table "$FIX/screening.tsv" --cache "$FIX/cache" --strict >/dev/null 2>&1; then bad "--strict returned 0 on a table with a wrong DOI" else pass "--strict exits non-zero, so a screening pipeline can gate on it" fi python3 "$DET" --table "$FIX/screening.tsv" --cache "$FIX/cache" --json "$WORK/o.json" >/dev/null python3 - "$WORK/o.json" <<'PY' import json, sys d = json.load(open(sys.argv[1])) assert d["detector"] == "check_doi_record_match", d.get("detector") assert d["source"] == "cache", d.get("source") assert d["rows_with_doi"] == 5, d.get("rows_with_doi") assert d["findings"] and all(f["detector"] == "check_doi_record_match" for f in d["findings"]) PY pass "qc JSON self-identifies and records that it read a cache, not the network" # Update notices. A correction or retraction notice repeats the article's title behind a short # prefix, so its title similarity clears the bar (0.91 here) and a title-only check called the # notice's DOI the article's. Crossref marks the notice with a structured `update-to` list; the # article itself carries `updated-by`, and a new-version update is still the work. # # U-001 article row, DOI of its correction notice DOI_IS_UPDATE_NOTICE # U-002 article row, the article's own (corrected) DOI SILENT # U-003 preprint row, a new-version update record SILENT # U-004 the correction notice's own row and DOI SILENT un="$(python3 "$DET" --table "$FIX/update_notice.tsv" --cache "$FIX/cache" || true)" if grep -qE "\[DOI_IS_UPDATE_NOTICE\] \(U-001\)" <<<"$un" && grep -q "correction -> 10.1000/article.008" <<<"$un"; then pass "a DOI of the article's correction notice is caught and names the DOI it corrects" else bad "a correction notice's DOI on the article's row was not caught" echo "$un" fi for r in U-002 U-003 U-004; do if grep -q "($r)" <<<"$un"; then bad "$r was flagged and should not have been" echo "$un" else case "$r" in U-002) pass "an article that has been corrected (updated-by) is silent" ;; U-003) pass "a new-version update record is the work itself and is silent" ;; U-004) pass "a row that IS the correction notice, with the notice's DOI, is silent" ;; esac fi done [ "$fail" -eq 0 ] || exit 1 echo "----" echo "DOI-record-match challenge: all checks passed"
-
-
check_doi_record_match.py 12.9 KB
#!/usr/bin/env python3 """Does this row's DOI point at this row's paper? A screening table's `doi` column is not always something a database handed over with the record. It is often something a pipeline *worked out* — matched by title similarity against Crossref, at a score somebody chose. When that match is wrong, the DOI is a valid, resolvable identifier for a different paper, and nothing downstream can tell. Every check that follows treats it as ground truth: deduplication, full-text retrieval, the reference list, the eligibility decision. Two failures inside two days, from one such column: * a conference-abstract record carried the DOI of a **different study already in the cohort**, which produced a false limitation — a "missed eligible paper" that did not exist. It reached the supporting material of a submitted abstract. * another record's DOI resolved to an **entire supplement of abstracts** rather than to an article, sending a reader to look for a paper that was never there. Neither is detectable by looking at the DOI. Both are obvious the moment you resolve it and read the title back. DOI_NOT_THIS_RECORD the DOI resolves, and to a different paper than the row describes DOI_IS_CONTAINER the DOI resolves to an issue, a supplement, a book or proceedings — a container, not an article DOI_IS_UPDATE_NOTICE the DOI resolves to a correction, erratum, retraction or similar notice that updates another DOI (Crossref `update-to`), not to the article. Its title is usually the article's own title behind a short prefix, so title similarity alone clears it. DOI_UNRESOLVED the DOI does not resolve at all (reported, never silently dropped) WHAT THIS IS NOT It is not `/verify-refs`, which audits a manuscript's finished reference list against PubMed and Crossref. This runs earlier, on the screening table, where a wrong DOI is still cheap — and where the reference list does not exist yet. It also cannot tell you a DOI is *right*. A title that matches is consistent with the row; it is not proof the record was correctly extracted. It rules out one specific, recurring, and otherwise invisible way of being wrong. DETERMINISM `--cache DIR` reads recorded Crossref responses (`<url-safe-doi>.json`) instead of the network, which is how the challenge card runs offline. With `--cache` and no recorded response, a row is UNRESOLVED — the cache is never quietly topped up from the network mid-run, because a check that is half-recorded and half-live is reproducible in neither direction. Stdlib only. Usage: check_doi_record_match.py --table screening.tsv [--title-col title] [--doi-col doi] [--id-col record_id] [--email you@example.org] [--threshold 0.90] [--cache DIR] [--json out.json] [--strict] Exit: 0 clean (or findings without --strict), 1 findings with --strict, 2 unusable input. """ from __future__ import annotations import argparse import csv import json import re import sys import time import unicodedata import urllib.error import urllib.parse import urllib.request from dataclasses import dataclass, field from difflib import SequenceMatcher from pathlib import Path from typing import Dict, List, Optional DETECTOR = "check_doi_record_match" CROSSREF = "https://api.crossref.org/works/" # Crossref `type` values that are an article-shaped thing a screening row can legitimately BE. # Anything else is a container: resolving to it means the row points at the box, not the paper. ARTICLE_TYPES = { "journal-article", "proceedings-article", "posted-content", "book-chapter", "report", "dissertation", "peer-review", "reference-entry", "other", } DOI_RE = re.compile(r"10\.\d{4,9}/[-._;()/:A-Za-z0-9]+") @dataclass class Finding: detector: str verdict: str record: Optional[str] summary: str evidence: List[str] = field(default_factory=list) def normalise(s: str) -> str: """Compare what a title says, not how a database punctuated it.""" s = unicodedata.normalize("NFKD", s or "").lower() s = re.sub(r"<[^>]+>", " ", s) # Crossref titles carry JATS markup s = re.sub(r"[^a-z0-9\s]", " ", s) return " ".join(s.split()) def similarity(a: str, b: str) -> float: na, nb = normalise(a), normalise(b) if not na or not nb: return 0.0 return SequenceMatcher(None, na, nb).ratio() def clean_doi(raw: str) -> str: """Accept a DOI however a spreadsheet wrote it: bare, prefixed, or as a URL.""" m = DOI_RE.search((raw or "").strip()) return m.group(0).rstrip(".,;)") if m else "" def cache_path(cache: Path, doi: str) -> Path: return cache / (urllib.parse.quote(doi, safe="") + ".json") def resolve(doi: str, email: Optional[str], cache: Optional[Path], pause: float) -> Optional[dict]: """-> the Crossref `message` object, or None if it did not resolve.""" if cache is not None: p = cache_path(cache, doi) if not p.is_file(): return None try: return json.loads(p.read_text(encoding="utf-8")).get("message") except (json.JSONDecodeError, OSError): return None url = CROSSREF + urllib.parse.quote(doi, safe="/") ua = "check_doi_record_match/1.0" if email: ua += f" (mailto:{email})" req = urllib.request.Request(url, headers={"User-Agent": ua, "Accept": "application/json"}) try: with urllib.request.urlopen(req, timeout=30) as r: # noqa: S310 - fixed https host payload = json.loads(r.read().decode("utf-8")) except (urllib.error.URLError, json.JSONDecodeError, TimeoutError, OSError): return None finally: time.sleep(pause) return payload.get("message") def resolved_title(msg: dict) -> str: t = msg.get("title") or [] return t[0] if t else "" def update_notice_targets(msg: dict) -> List[str]: """-> "type -> DOI" for each entry of Crossref's structured `update-to` list. A record that carries `update-to` is a notice about another DOI. A new-version update (a revised record of the same work) is still the work itself, so it is not counted here. """ out: List[str] = [] for u in msg.get("update-to") or []: if not isinstance(u, dict): continue kind = str(u.get("type") or "").lower() if "version" in kind: continue out.append(f"{kind or 'update'} -> {u.get('DOI') or '?'}") return out def audit(rows: List[Dict[str, str]], id_col: str, title_col: str, doi_col: str, threshold: float, email: Optional[str], cache: Optional[Path], pause: float) -> List[Finding]: out: List[Finding] = [] for i, row in enumerate(rows, start=2): # 1 is the header; report the line a person can open rid = (row.get(id_col) or f"row {i}").strip() or f"row {i}" own_title = (row.get(title_col) or "").strip() doi = clean_doi(row.get(doi_col) or "") if not doi: continue if not own_title: out.append(Finding( DETECTOR, "DOI_UNRESOLVED", rid, f"{rid}: carries a DOI but no title, so the DOI cannot be checked against " "anything.", [f"doi: {doi}", "A DOI with nothing to compare it to is an assertion, not a fact."], )) continue msg = resolve(doi, email, cache, pause) if msg is None: out.append(Finding( DETECTOR, "DOI_UNRESOLVED", rid, f"{rid}: the DOI did not resolve.", [f"doi: {doi}", f"row title: {own_title[:80]!r}", "Reported rather than dropped: an unresolvable DOI in a screening table is a " "record nobody can retrieve, which is a finding of its own."], )) continue got = resolved_title(msg) ratio = similarity(own_title, got) kind = (msg.get("type") or "").lower() if kind and kind not in ARTICLE_TYPES: out.append(Finding( DETECTOR, "DOI_IS_CONTAINER", rid, f"{rid}: the DOI resolves to a {kind}, not to an article.", [f"doi: {doi}", f"resolved to: {got[:80]!r}", f"row title: {own_title[:80]!r}", "Following this DOI leads to the volume the paper is in, or to a whole " "supplement of abstracts. Anyone sent there looks for a paper that is not " "separately registered."], )) continue updates = update_notice_targets(msg) if updates and normalise(own_title) != normalise(got): out.append(Finding( DETECTOR, "DOI_IS_UPDATE_NOTICE", rid, f"{rid}: the DOI resolves to a notice that updates another DOI " f"({'; '.join(updates)}), not to the article (title similarity {ratio:.2f}).", [f"doi: {doi}", f"row title: {own_title[:100]!r}", f"resolved title: {got[:100]!r}", "A correction or retraction notice repeats the article's title, so a " "title-similarity match lands on it. The article's own DOI is the one the " "notice points to."], )) continue if ratio < threshold: out.append(Finding( DETECTOR, "DOI_NOT_THIS_RECORD", rid, f"{rid}: the DOI resolves to a different paper (title similarity " f"{ratio:.2f} < {threshold:.2f}).", [f"doi: {doi}", f"row title: {own_title[:100]!r}", f"resolved title: {got[:100]!r}", "If this column was filled by matching titles against Crossref, this is what that " "match got wrong — and downstream it is indistinguishable from a real record."], )) return out def read_table(path: Path) -> List[Dict[str, str]]: text = path.read_text(encoding="utf-8-sig") delim = "\t" if path.suffix.lower() in {".tsv", ".tab"} or "\t" in text.split("\n")[0] else "," return list(csv.DictReader(text.splitlines(), delimiter=delim)) def main() -> int: ap = argparse.ArgumentParser(description=__doc__.split("\n")[0]) ap.add_argument("--table", type=Path, required=True, help="screening table (.tsv or .csv)") ap.add_argument("--id-col", default="record_id") ap.add_argument("--title-col", default="title") ap.add_argument("--doi-col", default="doi") ap.add_argument("--threshold", type=float, default=0.90, help="title similarity below which the DOI is another paper (default 0.90). " "Pipelines that FILL a doi column commonly accept ~0.79, which is how " "the wrong DOIs got in; do not set this to the value that produced them.") ap.add_argument("--email", help="contact address for the Crossref polite pool") ap.add_argument("--cache", type=Path, help="directory of recorded Crossref responses; disables all network access") ap.add_argument("--pause", type=float, default=0.2, help="seconds between live requests") ap.add_argument("--json", type=Path) ap.add_argument("--strict", action="store_true") a = ap.parse_args() if not a.table.is_file(): print(f"cannot read {a.table}", file=sys.stderr) return 2 if a.cache is not None and not a.cache.is_dir(): print(f"cache directory {a.cache} does not exist", file=sys.stderr) return 2 rows = read_table(a.table) if not rows: print(f"{a.table} has no data rows", file=sys.stderr) return 2 missing = [c for c in (a.title_col, a.doi_col) if c not in rows[0]] if missing: print(f"{a.table} has no column(s) named {', '.join(missing)} " f"(found: {', '.join(rows[0])})", file=sys.stderr) return 2 findings = audit(rows, a.id_col, a.title_col, a.doi_col, a.threshold, a.email, a.cache, a.pause) checked = sum(1 for r in rows if clean_doi(r.get(a.doi_col) or "")) if a.json: a.json.parent.mkdir(parents=True, exist_ok=True) a.json.write_text(json.dumps( {"detector": DETECTOR, "table": str(a.table), "rows": len(rows), "rows_with_doi": checked, "threshold": a.threshold, "source": "cache" if a.cache else "crossref", "findings": [f.__dict__ for f in findings]}, indent=2, ensure_ascii=False) + "\n", encoding="utf-8") if not findings: print(f"OK: {checked} DOI(s) each resolve to the record that carries them.") return 0 print(f"{len(findings)} finding(s) across {checked} DOI(s)\n") for f in findings: print(f" [{f.verdict}] ({f.record})") print(f" {f.summary}") for e in f.evidence: print(f" - {e}") print() return 1 if a.strict else 0 if __name__ == "__main__": sys.exit(main())
-
-
tests
-
test_parse_pubmed.sh 3.9 KB
#!/usr/bin/env bash # Regression test: parse_pubmed.py keeps the whole title and abstract when PubMed marks them up. # # PubMed efetch XML carries inline markup inside <ArticleTitle> and <AbstractText>: italic species # names, sub/superscripts. `findtext()` and `.text` return only the text before the first child # element, so "Eradication of <i>Helicobacter pylori</i> infection." became # `title = {Eradication of },` in an entry stamped `verified = {true}`, and an abstract reading # "CO<sub>2</sub> levels rose." became "CO". The pre-fix parser fails checks 1-4; checks 5-6 are # the negative control (a title with no markup is unchanged). # # No network: the XML is synthetic and inline. set -u HERE="$(cd "$(dirname "$0")" && pwd)" PARSER="$HERE/../references/parse_pubmed.py" TMP="$(mktemp -d)" trap 'rm -rf "$TMP"' EXIT cat >"$TMP/markup.xml" <<'XML' <?xml version="1.0"?> <PubmedArticleSet> <PubmedArticle> <MedlineCitation> <PMID>99999991</PMID> <Article> <Journal> <Title>Journal of Synthetic Medicine</Title> <ISOAbbreviation>J Synth Med</ISOAbbreviation> <JournalIssue><Volume>1</Volume><Issue>2</Issue><PubDate><Year>2024</Year></PubDate></JournalIssue> </Journal> <ArticleTitle>Eradication of <i>Helicobacter pylori</i> infection.</ArticleTitle> <Abstract> <AbstractText Label="BACKGROUND">CO<sub>2</sub> levels rose.</AbstractText> </Abstract> <AuthorList><Author><LastName>Researcher</LastName><ForeName>Ada</ForeName></Author></AuthorList> <ELocationID EIdType="doi">10.0/synthetic.markup</ELocationID> </Article> </MedlineCitation> </PubmedArticle> </PubmedArticleSet> XML cat >"$TMP/plain.xml" <<'XML' <?xml version="1.0"?> <PubmedArticleSet> <PubmedArticle> <MedlineCitation> <PMID>99999992</PMID> <Article> <Journal> <Title>Journal of Synthetic Medicine</Title> <JournalIssue><PubDate><Year>2023</Year></PubDate></JournalIssue> </Journal> <ArticleTitle>Plain title without markup.</ArticleTitle> <Abstract><AbstractText>Plain abstract.</AbstractText></Abstract> <AuthorList><Author><LastName>Writer</LastName><ForeName>Dan</ForeName></Author></AuthorList> </Article> </MedlineCitation> </PubmedArticle> </PubmedArticleSet> XML fail=0 ck() { if [ "$2" = "$3" ]; then printf ' PASS %s\n' "$1"; else printf ' FAIL %s (want %s got %s)\n' "$1" "$2" "$3"; fail=$((fail+1)); fi; } has() { grep -qF -- "$2" "$1" && echo yes || echo no; } python3 "$PARSER" bibtex <"$TMP/markup.xml" >"$TMP/markup.bib" python3 "$PARSER" efetch <"$TMP/markup.xml" >"$TMP/markup.md" python3 "$PARSER" bibtex <"$TMP/plain.xml" >"$TMP/plain.bib" python3 "$PARSER" efetch <"$TMP/plain.xml" >"$TMP/plain.md" # 1-2. a title with inline <i> keeps the text inside and after it ck "bibtex title keeps the italic species name" yes \ "$(has "$TMP/markup.bib" 'title = {Eradication of Helicobacter pylori infection.},')" ck "efetch title keeps the italic species name" yes \ "$(has "$TMP/markup.md" '**Title**: Eradication of Helicobacter pylori infection.')" # 3-4. an abstract with inline <sub> keeps the text inside and after it ck "efetch abstract keeps the subscript and the rest of the sentence" yes \ "$(has "$TMP/markup.md" '**Abstract**: BACKGROUND: CO2 levels rose.')" ck " ...and never stops at the first child element" no \ "$(has "$TMP/markup.bib" 'title = {Eradication of },')" # 5-6. negative control: a title with no markup is unchanged ck "plain bibtex title is unchanged" yes "$(has "$TMP/plain.bib" 'title = {Plain title without markup.},')" ck "plain efetch abstract is unchanged" yes "$(has "$TMP/plain.md" '**Abstract**: Plain abstract.')" if [ "$fail" -eq 0 ]; then echo "PASS: parse_pubmed.py keeps titles and abstracts whole through inline markup." else echo "FAIL: $fail check(s) failed." >&2 exit 1 fi -
test_pubmed_eutils.sh 3.7 KB
#!/usr/bin/env bash # Regression test: free-text queries in references/pubmed_eutils.sh are data, never code. # # `search` used to URL-encode the query by pasting it into Python source, # `python3 -c "...quote('${query}')"`. Apostrophes are ordinary in queries (Crohn's disease, # Alzheimer's, O'Brien[Author]). The first one closed the literal and caused a SyntaxError, the # substitution came back empty, and the script sent `esearch?term=&retmax=...` and exited 0: # a silent search for nothing, on the path `cite_lookup` shares. A crafted query could also # close the literal itself and run its own Python, and `related` pasted `retmax` into source # the same way. # # No network: `curl` is a stub on PATH that records each URL and answers with canned JSON. # The pre-fix script fails 9 of the 14 assertions. It sends `term=`, runs both breakout payloads # (their marker files appear) and accepts a non-numeric retmax. Check 5 passes before and after; # it pins that retmax still truncates once it is passed as an argument. set -u HERE="$(cd "$(dirname "$0")" && pwd)" EUTILS="$HERE/../references/pubmed_eutils.sh" TMP="$(mktemp -d)" trap 'rm -rf "$TMP"' EXIT mkdir -p "$TMP/bin" cat >"$TMP/bin/curl" <<'STUB' #!/usr/bin/env bash url="${*: -1}" echo "$url" >>"$CALLS" case "$url" in *elink.fcgi*) printf '%s' '{"linksets":[{"linksetdbs":[{"linkname":"pubmed_pubmed","links":[{"id":111},{"id":222},{"id":333}]}]}]}' ;; *) printf '%s' '{"esearchresult":{"idlist":[]}}' ;; esac printf '\n200\n' STUB chmod +x "$TMP/bin/curl" fail=0 ck() { if [ "$2" = "$3" ]; then printf ' PASS %s\n' "$1"; else printf ' FAIL %s (want %s got %s)\n' "$1" "$2" "$3"; fail=$((fail+1)); fi; } run() { # run <args...> ; prints the exit code, URLs land in $TMP/calls : >"$TMP/calls" ( cd "$TMP" && PATH="$TMP/bin:$PATH" CALLS="$TMP/calls" NCBI_API_KEY= bash "$EUTILS" "$@" >"$TMP/out" 2>"$TMP/err" ) echo $? } sent() { grep -qF -- "$1" "$TMP/calls" && echo yes || echo no; } # 1. an apostrophe in a search query ck "search with an apostrophe exits 0" 0 "$(run search "Crohn's disease[Title] AND MRI" 5)" ck " ...sends the whole encoded term" yes "$(sent 'term=Crohn%27s%20disease%5BTitle%5D%20AND%20MRI&retmax=5&')" ck " ...never sends an empty term" no "$(sent 'term=&')" # 2. cite_lookup goes through the same path ck "cite_lookup with an apostrophe exits 0" 0 "$(run cite_lookup "Crohn's disease: a review")" ck " ...sends the whole [Title] term" yes "$(sent 'term=Crohn%27s%20disease%3A%20a%20review%5BTitle%5D&retmax=5&')" # 3. a query written to break out of the old Python string literal MARK="$TMP/query-ran-code" ck "breakout query exits 0" 0 "$(run search "x')); open('$MARK', 'w'); print(('" 5)" ck " ...runs no code" no "$([ -e "$MARK" ] && echo yes || echo no)" ck " ...is searched as literal text" yes "$(sent 'term=x%27%29%29%3B%20open%28')" # 4. a related retmax written to break out of the old Python slice MARK2="$TMP/retmax-ran-code" rc="$(run related 12345 "1]]; open('$MARK2', 'w'); x=[[0")" ck "breakout retmax is refused" yes "$([ "$rc" -ne 0 ] && echo yes || echo no)" ck " ...runs no code" no "$([ -e "$MARK2" ] && echo yes || echo no)" ck " ...before any request" 0 "$(wc -l <"$TMP/calls" | tr -d ' ')" # 5. related still honours retmax, now passed as an argument ck "related 12345 2 exits 0" 0 "$(run related 12345 2)" ck " ...fetches only the first 2 linked PMIDs" yes "$(sent 'esummary.fcgi?db=pubmed&id=111,222&')" # 6. a non-numeric search retmax rc="$(run search "MRI" "5&api_key=x")" ck "non-numeric search retmax is refused" yes "$([ "$rc" -ne 0 ] && echo yes || echo no)" if [ "$fail" -eq 0 ]; then echo "PASS: pubmed_eutils.sh sends free text as data, never as code." else echo "FAIL: $fail check(s) failed." >&2 exit 1 fi
-
-
SKILL.md 18.5 KB
--- name: search-lit description: Use when finding papers or building a reference list. Searches PubMed, Semantic Scholar and bioRxiv/medRxiv, includes only references verified through an API, and generates BibTeX. Auditing an existing reference list is /verify-refs. metadata: triggers: "literature search, find papers, citation, references, bibliography, PubMed search, related work" --- # Literature Search Skill Every reference you produce must come from a live database result — never generate a citation from memory alone, because a recalled citation can look real and not exist. ## Search Tools: MCP (Primary) + E-utilities (Fallback) ### Primary: MCP Tools (Claude.ai Remote) | Database | MCP Tool | Purpose | |----------|----------|---------| | PubMed | `mcp__claude_ai_PubMed__search_articles` | Search by query, MeSH terms | | PubMed | `mcp__claude_ai_PubMed__get_article_metadata` | Full metadata for a PMID | | PubMed | `mcp__claude_ai_PubMed__find_related_articles` | Related articles for a PMID | | PubMed | `mcp__claude_ai_PubMed__lookup_article_by_citation` | Verify a citation | | PubMed | `mcp__claude_ai_PubMed__convert_article_ids` | Convert between PMID/DOI/PMCID | | Semantic Scholar | `mcp__claude_ai_Scholar_Gateway__semanticSearch` | Semantic search across all fields | | bioRxiv/medRxiv | `mcp__claude_ai_bioRxiv__search_preprints` | Search preprint servers | | bioRxiv/medRxiv | `mcp__claude_ai_bioRxiv__get_preprint` | Full preprint metadata | | CrossRef | WebFetch with `https://api.crossref.org/works/{DOI}` | DOI verification | ### Fallback: NCBI E-utilities (Direct API via Bash) If any `mcp__claude_ai_PubMed__*` call returns an error containing "terminated", "not found", "not available", or "not connected", switch ALL subsequent PubMed calls in this session to the bundled E-utilities scripts. Do not retry MCP after a disconnect — it will not recover within the same conversation. ```bash EUTILS="${CLAUDE_SKILL_DIR}/references/pubmed_eutils.sh" PARSER="${CLAUDE_SKILL_DIR}/references/parse_pubmed.py" # Search PubMed (returns PMIDs) bash "$EUTILS" search "diagnostic test accuracy meta-analysis radiology" 20 \ | python3 "$PARSER" esearch # Get article summaries as markdown table bash "$EUTILS" fetch_json "16168343,16085191,31462531" \ | python3 "$PARSER" esummary # Get detailed metadata bash "$EUTILS" fetch "16168343" \ | python3 "$PARSER" efetch # Generate BibTeX entries bash "$EUTILS" fetch "16168343,16085191" \ | python3 "$PARSER" bibtex # Verify a citation by exact title bash "$EUTILS" cite_lookup "Bivariate analysis of sensitivity and specificity" \ | python3 "$PARSER" esearch # Find related articles for a PMID bash "$EUTILS" related "16168343" 10 \ | python3 "$PARSER" esummary ``` The script sleeps 350 ms between calls (NCBI allows 3 requests/s without an API key, 10/s with `NCBI_API_KEY`); keep batch calls sequential. | MCP Tool | E-utilities Command | Parser Mode | |----------|-------------------|-------------| | `search_articles` | `search <query> [retmax]` | `esearch` | | `get_article_metadata` | `fetch <pmids>` | `efetch` or `bibtex` | | `find_related_articles` | `related <pmid> [retmax]` | `esummary` | | `lookup_article_by_citation` | `cite_lookup <title>` | `esearch` → `fetch` | | `convert_article_ids` | Not available (use CrossRef DOI lookup) | — | --- ## Workflow ### Phase 1: Search Strategy 1. Get the research topic, question, or manuscript section that needs references. 2. Build the query from the key concepts (Population, Intervention/Exposure, Comparison, Outcome), with MeSH terms for PubMed: `(concept1 OR synonym1) AND (concept2 OR synonym2)`. 3. Set scope: date range (default: last 10 years unless the user specifies), article types, and language (default: English). 4. Present the Boolean query, databases, and filters to the user. **Gate:** Wait for user approval before running searches. ### Phase 2: Execute Search 1. Search PubMed (`search_articles`, Boolean query), Semantic Scholar (`semanticSearch`, natural language query), and bioRxiv/medRxiv (`search_preprints`) when preprints are relevant. 2. Deduplicate across databases by DOI or title similarity. 3. Present the results in this table, and write the same rows to `references/search_results.tsv`: ``` | # | Title | Authors (first + last) | Year | Journal | PMID/DOI | Relevance | |---|-------|----------------------|------|---------|----------|-----------| | 1 | ... | Kim J, ... Lee S | 2024 | Radiology | 12345678 | High | ``` 4. Ask the user to select which papers to include. #### Record what the source said existed, not only what you downloaded A wrong haul raises no error, and a PRISMA flow built on a wrong number is fiction that nothing downstream contradicts. Check both of these: - **A count that equals a page cap exactly.** Every source reports a total: `esearchresult.count`, `opensearch:totalResults`, `meta.count`. **Record `api_total` beside `downloaded`, and fail loudly when `downloaded < api_total`, or when `downloaded` equals a page or loop cap exactly** (e.g. a loop's own `if start >= 2000: break` reporting 2,000 records). Print `TRUNCATED` and refuse to write the search record. - **A boolean that was never applied.** OpenAlex's `search=` is a relevance-ranked free-text parameter that **silently ignores AND/OR**; `filter=title_and_abstract.search:` honours them. Run the query once more with one mandatory clause negated. **If the hit count does not drop, the boolean is being ignored** — the engine is ranking, not filtering. PubMed via E-utilities is the one place where the naive pattern is safe. Everywhere else, do both. #### A DOI in a screening row is not necessarily that row's DOI When the pipeline **filled** a `doi` column (matched against Crossref by title similarity) rather than receiving it with the record, a wrong match is a valid, resolvable DOI for a different paper. Resolve it and read the title back before any decision rests on it: ```bash python3 "${CLAUDE_SKILL_DIR}/scripts/check_doi_record_match.py" --table 2_Screening/round3.tsv \ --email <contact> --json qc/doi_record_match.json ``` `DOI_NOT_THIS_RECORD` is a DOI that resolves to another paper; `DOI_IS_CONTAINER` resolves to an issue, supplement or proceedings rather than an article; `DOI_IS_UPDATE_NOTICE` resolves to a correction, erratum or retraction notice (Crossref `update-to`) rather than to the article it updates; `DOI_UNRESOLVED` is reported rather than dropped. This runs at screening, where a wrong DOI is still cheap; `/verify-refs` audits a finished reference list. ### Phase 2.5: Citation Searching (Snowballing) Optional; recommended for systematic reviews and thorough background work (PRISMA item 7, "records identified through citation searching"). Expand a seed set along the citation graph with the Semantic Scholar Graph API helper: ```bash python3 "${CLAUDE_SKILL_DIR}/references/snowball.py" \ --seed DOI:10.1000/synthetic.example,PMID:00000000 \ --direction all \ --pool references/library.bib \ --out references/library.bib ``` - **Directions**: `backward` (references the seeds cite), `forward` (papers citing the seeds), `similar` (S2 recommendations), or `all` (default). Dedup is against the `--pool` and within the harvested set, by DOI and normalized title. - **Trust flag**: candidates are written `verified=false` + `verified_by=semantic_scholar`. Run `/verify-refs` (or Phase 4 verification) on each before citing it. - **Output contract**: appends to `references/library.bib` only; the script hard-refuses `manuscript/_src/refs.bib`. - **PRISMA line**: the script prints `Records identified through citation searching (snowballing): N raw (backward=…, forward=…, similar=…); after dedup against existing pool: M new candidates.` Record M in the PRISMA flow's citation-searching box. - **Incomplete runs**: when any seed/direction fetch fails, or the source says more records exist past `--limit` (a `next` page), the script prints a `FAILED` or `TRUNCATED` line per seed/direction, ends the PRISMA line with `INCOMPLETE`, and exits 1. Those counts are a lower bound. Do not record them; re-run (or raise `--limit`) until the run exits 0. ### Phase 3: Deep Read For each selected paper, retrieve full metadata (`get_article_metadata` for PubMed, `get_preprint` for bioRxiv) and extract the study design, sample size/dataset, key methods, primary findings (with specific numbers), and the authors' stated limitations. For several papers, present a literature matrix for review: ``` | Paper | Design | N | Key Finding | Limitation | Relevance to Our Study | |-------|--------|---|-------------|------------|----------------------| ``` ### Phase 4: Citation Management #### Verification 1. **NEVER fabricate a DOI or PMID.** If you cannot find one, mark the reference `[UNVERIFIED - NEEDS MANUAL CHECK]`. 2. Cross-check every reference against the API result: first and last author, publication year, journal, article title (exact, not paraphrased), and volume/pages when available. Flag each field that does not match. 3. Confirm each DOI resolves with WebFetch `https://api.crossref.org/works/{DOI}` (on CrossRef errors, follow Error Handling). #### BibTeX Generation Generate an entry for every reference, verified or not, with an explicit `verified` flag: ```bibtex @article{FirstAuthorLastName_Year_ShortKey, author = {Last1, First1 and Last2, First2 and Last3, First3}, title = {Full Title As Retrieved From Database}, journal = {Journal Name}, year = {2024}, volume = {310}, number = {2}, pages = {e234567}, doi = {10.1001/jama.2024.12345}, pmid = {12345678}, verified = {true}, verified_by = {pubmed+crossref}, verified_on = {2026-04-24}, } ``` | `verified` (required on every entry) | Meaning | |---|---| | `true` | DOI or PMID confirmed via PubMed/CrossRef; title, authors, year all match | | `false` | Parsed from text, but the API lookup failed or returned a mismatch; the manuscript MUST show `[UNVERIFIED - NEEDS MANUAL CHECK]` | | `manual` | User explicitly added it despite the lookup failure; still unverified | `verified_by` lists the confirming sources (`pubmed`, `crossref`, `semantic_scholar`, or a combination); `verified_on` is the ISO date of the most recent successful verification. No downstream script reads these fields — `/verify-refs` re-checks every entry — so they record trust rather than grant it. **BibTeX key convention**: `FirstAuthorLastName_Year_OneWord` (e.g., `Kim_2024_Validation`). The key is provisional: Better BibTeX assigns the citable key when `/lit-sync` imports the entry into Zotero. #### Output 1. Append the entries to `references/library.bib` (do not overwrite) — the candidate pool `/lit-sync` imports into Zotero. NEVER write to `manuscript/_src/refs.bib`, because `/lit-sync` (via Better BibTeX) is its sole writer. 2. Print a summary with verification status: ``` Verified: 12 references (verified=true) Unverified: 1 reference (verified=false) [NEEDS MANUAL CHECK] Total: 13 references ``` ### Phase 4b: Zotero Library Integration Importing into Zotero belongs to `/lit-sync`, which owns Zotero writes and `references/zotero_collection.json`: hand it `references/library.bib`. If a Zotero MCP server is connected, you may read from it here — `zotero_search_items` (by DOI) to mark candidates already in the library, `zotero_get_annotations` to reference the user's prior reading notes. ### Phase 5: Full-Text Retrieval Delegate to `/fulltext-retrieval`, the single home of the open-access cascade; do **not** re-implement OA fetching here. Pass the verified candidate DOIs from `references/library.bib`: ```bash ENGINE="${CLAUDE_SKILL_DIR}/../fulltext-retrieval/fetch_oa.py" # extract DOI + Title (and PMID/FirstAuthor when available) → worklist.tsv python3 "$ENGINE" worklist.tsv -o pdfs/ -e <contact-email> --report pdfs/retrieval_report.json ``` Put the verified bibliographic title in the worklist rather than a DOI-only list. Keep `source_identity` and `file_sha256` from the retrieval report with the record. Download success and title agreement alone do not verify the PDF: inspect conflicts, unresolved/unavailable evidence, and files whose hashes have changed before citing them. Missing identity fields in older reports mean unassessed; even `consistent` is advisory front-matter corroboration, not verification of the paper's claims. For Zotero-resident PDFs and proxy-aware retrieval, use `/lit-sync` Phase 2.7. For DOIs in `pdfs/manual_needed.txt`, use only institutional access (your library's own subscriptions, proxy or VPN), interlibrary loan, or the corresponding author. Never bypass paywalls or publisher access controls, and do not configure unauthorized PDF mirrors. ### Phase 6: Gap Analysis When called during manuscript writing, extract the manuscript's inline citations, compare them with the search results, and report specific gaps: key papers not cited, outdated references with newer versions, and missing methodological references (statistical methods, reporting guidelines). --- ## Specialized Search Modes ### Mode: Manuscript Paper Reference Pool Supplies a manuscript's reference pool — typically invoked by `/write-paper` Step 7.3c (or `/self-review` Phase 2.5c-2) when the reference-adequacy gate finds the draft under target or a named method uncited; usable directly for an original-research bibliography. For an original-research article, return **25–40** verified candidates, not the ~10 a quick search settles on. If the field is genuinely sparse, say so explicitly rather than returning a thin list silently. Respect a narrower journal reference cap or user scope when one is given. Cover **six candidate categories**: 1. **Background / disease burden / clinical context** — why the question matters. 2. **Gap-defining prior studies** — the work the manuscript extends or contradicts. 3. **Comparator / comparable-design cohorts** — studies the Results will be measured against. 4. **Methods / statistical canonical sources** — the originating reference for every named method, model, score, equation, or diagnostic criterion (e.g. competing-risk model, multiple imputation, E-value, eGFR equation, concordance statistic). This category clears Methods named-method gaps. 5. **Reporting-guideline sources** — STROBE, TRIPOD(+AI), CONSORT, PRISMA(-DTA), STARD, etc. 6. **Interpretation / mechanism / limitation support** — grounds Discussion claims. For each candidate, report **PMID/DOI**, **verification status**, **candidate category**, the **target manuscript section**, and a one-line **why it is needed**. Entries go through Phase 4 into `references/library.bib` only. This mode produces candidates: the user decides inclusion, and it does not insert references into the manuscript bib. ### Mode: Crowding Check Run **before a study is designed**. A background search ("what has been written about this topic") leaves the trap open; ask four narrower questions instead: | Ask of | Verdict | |---|---| | the **research question** | taken / partly taken / open | | the **sampling frame** (what population, which records, which years) | taken / partly taken / open | | the **measurement axis** (what is being coded or measured, and at what granularity) | taken / partly taken / open | | the **target journal** | already published there / adjacent / open | Give each its own verdict: a design can be original on one axis and fully occupied on another, and collapsing the four into one answer hides that. The journal row is not vanity: a design once matched an existing paper on frame, coding axis **and** journal — its own first choice. Most of the time this mode **narrows a claim rather than ending a project** (e.g. from "nobody has looked at this" to "nobody has decomposed it by provenance"), and the narrowed claim survives review. Search the way a competitor would: the exact frame, the exact measure, and the journal's own site, not only the topic. Report the four verdicts and the papers behind each, then let the user decide. ### Mode: Systematic Search For systematic reviews or comprehensive literature sections: 1. Document the full search strategy (PRISMA-compliant). 2. Record: database, date of search, query string, number of results. 3. Track inclusion/exclusion at each screening step. 4. Output a PRISMA flow diagram data summary. ### Mode: Quick Cite For a single reference the user describes ("that 2023 paper by Smith about AI in chest X-ray"): search PubMed and Semantic Scholar with the details, present the top 3 candidates, and generate the BibTeX entry for the one the user confirms. ### Mode: Related Papers From a PMID or DOI, get related papers with `find_related_articles` plus Semantic Scholar citation-based recommendations, ranked by relevance. For a structured, dedup-aware, PRISMA-countable expansion (backward + forward + similar), use **Phase 2.5: Citation Searching** with `references/snowball.py` instead. ### Mode: Embase Browser Automation Embase has no public API. Read `${CLAUDE_SKILL_DIR}/references/embase_browser.md` when the search must include Embase — it has the Chrome-automation export steps, the CSV row format, and the PubMed → Embase query translation. --- ## Error Handling - If a search returns 0 results, broaden the query (remove one concept or use broader MeSH terms) and retry. - **CrossRef HTTP errors (token-saving rules):** - **403 (rate-limited):** Do NOT retry. Skip CrossRef → verify via PubMed title search instead. - **303 (redirect):** Follow the redirect if possible. If not, skip CrossRef → PubMed fallback. - After the first CrossRef 403/303 in a session, skip CrossRef for ALL remaining references and go directly to PubMed title verification, to avoid N×retry token waste. - Do not print raw error messages ("Request failed with status code 403."). Report one summary line at the end: `CrossRef unavailable for {N} references (rate-limited). Verified via PubMed instead.` - If a DOI does not resolve via CrossRef (after the rules above), search PubMed by title to confirm the reference exists. - If a reference cannot be verified by any method, state: "This reference could not be verified. Please check manually before submission." Never silently include an unverified reference. ## Known limits - `check_doi_record_match.py` compares titles at a similarity threshold and reads structured Crossref fields (`type`, `update-to`). A study protocol, or part 2 of a multi-part paper, whose title differs from the row's by a word or a number is not separated from the paper the row describes; no structured field marks it. Treat a silent run as "no mismatch detected", not as proof that every DOI is right. -
skill.yml 2.8 KB
schema_version: 2 name: search-lit layer: A owner_domain: literature_discovery maturity: official when_to_use: - User asks to find papers, related work, citations, or background literature on a topic - Generating a verified candidate citation pool (PubMed / Semantic Scholar / bioRxiv / medRxiv) - Pre-/lit-sync candidate sourcing (search-lit produces candidates → /lit-sync syncs to Zotero + refs.bib) - Building a BibTeX library for a topic without yet committing to inclusion when_NOT_to_use: - Verifying citations already in a manuscript (use /verify-refs) - Syncing to Zotero or writing manuscript/_src/refs.bib (use /lit-sync — sole writer) - Generating references from model memory (forbidden — every entry must be API-verified) inputs: - literature_query outputs: - references/library.bib # search-result candidate pool; NOT the manuscript SSOT bib - references/search_results.tsv deterministic_scripts: - scripts/check_doi_record_match.py # screening-row DOI provenance; challenge card in scripts/check_doi_record_match_challenge/ - references/pubmed_eutils.sh - references/parse_pubmed.py - references/snowball.py # Phase 2.5 citation snowballing (S2 Graph API); challenge card in references/snowball_challenge/ side_effects: - may_call_external_literature_apis downstream_consumers: - verify-refs - lit-sync # confirmed candidates flow through Zotero, then lit-sync refreshes manuscript/_src/refs.bib - write-paper ssot_boundary: - manuscript/_src/refs.bib is OWNED by /lit-sync (Better BibTeX auto-export). search-lit MUST NOT write to that path. forbidden_actions: - generate_references_from_memory - silently_include_unverified_references - write_to_manuscript_refs_bib # SSOT owner is /lit-sync # v2.1 quality card purpose: "Search PubMed, Semantic Scholar, and bioRxiv/medRxiv and generate API-verified BibTeX (anti-hallucination: every reference verified before inclusion)." safety_boundaries: - "Never generates references from memory; unverified references are not silently included." - "Does not write to the manuscript refs.bib (that SSOT belongs to lit-sync)." known_limitations: - "Depends on PubMed/Semantic Scholar availability; rate limits/outages reduce recall." - "Verification confirms existence/metadata, not topical relevance." - "A DOI whose resolved title matches its row is consistent with that row; it is not proof the record was extracted correctly. The check rules out one recurring way of being wrong, not all of them." validation_commands: - "bash references/pubmed_eutils.sh <query>" - "python3 scripts/check_doi_record_match.py --table <screening.tsv> --strict" - "bash scripts/check_doi_record_match_challenge/verify.sh # deterministic, network-free" - "bash references/snowball_challenge/verify.sh # deterministic, network-free" - "/verify-refs --strict" evidence_surface: bundled_script
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.