literature-review-tools
Recommend AND run open-source AI tools, agents, Claude Code / Codex skills, and MCP servers for any stage of a literature review — searching, reading, extracting, synthesizing, screening, citation-checking, and paper writing. Use when the user asks "what tool should I use to..."
Install
npx skills add https://github.com/brycewang-stanford/lit-review-agent-tools/tree/main/skills/literature-review-tools
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install brycewang-stanford-lit-review-agent-tools@llmmart
git clone https://github.com/brycewang-stanford/lit-review-agent-tools.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole brycewang-stanford/lit-review-agent-tools collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Literature Review Tools — Select & Run
A curated, use-case-organized catalog of the strongest open-source AI tools for literature review — plus a launcher that actually installs and runs the top ones. Covers: end-to-end research agents, deep-research / auto-survey generators, autonomous "idea→paper" systems, citation-backed RAG over PDFs, PRISMA screening, MCP servers, Zotero/Obsidian integrations, PDF→structured extraction, citation graphs, and paper-writing / peer-review assistants.
Full source of truth (README, always current star counts): https://github.com/brycewang-stanford/lit-review-agent-tools
Three modes
- Look up — user wants the literature itself: "find papers on X", "get me the PDF for this DOI", "does this citation exist", "what does PubMed have since 2022". Answer it directly — no install, no key. Start with
reference/apis/README.mdfor one-off lookups, or run the bundledpapers-fetch/oa-resolvescripts for anything corpus-sized. - Recommend — user asks "what should I use to …". Route with the tables below; cite the catalog for details.
- Run — user asks to install / run / use a specific tool ("turn this PDF into Markdown with MinerU", "ask PaperQA2 about these papers", "set up the arXiv MCP server"). Drive
scripts/litrun.pyvia Bash — do not hand the user raw pip commands to copy.
Mode 1 is the cheap default. Do not send someone to install PyTorch when they asked for five papers and a PDF.
Look up mode — search without installing anything
Two bundled scripts, both standard-library only: no venv, no pip, no API key.
Run them straight (python3 scripts/fetch_papers.py …) or through the launcher.
# 1. search six indexes at once, deduplicated by DOI
python3 scripts/fetch_papers.py --query "active learning for screening" \
--sources openalex,crossref,semanticscholar,pubmed,europepmc,arxiv \
--max 15 --dedup-titles --outdir ./corpus
# 2. turn those DOIs into full text you may legally read
python3 scripts/resolve_oa.py --from-json ./corpus/results.json --outdir ./corpus
fetch_papers.py writes results.json (normalised records, per-source counts, and the
errors of any source that failed) plus one .txt per paper. resolve_oa.py walks
Unpaywall → OpenAlex → Europe PMC → arXiv → CORE and writes oa_report.json recording
how every DOI resolved — including the ones that stayed closed, and separating those
from DOIs Crossref has never registered (usually an invented citation). On a 47-paper
corpus with zero keys it recovered 33 full texts; nothing in it bypasses a paywall, and
you should not offer to.
For a single lookup — one DOI, one author, one citation check — skip the scripts and
call the API directly (WebFetch/curl). reference/apis/ has a
routing table, the identifier formats, and one page per API with the endpoints and the
failure modes that waste time (Crossref's select 400, OpenAlex's inverted abstracts,
Semantic Scholar's exhausted keyless pool, PubMed's multi-part AbstractText).
Report which indexes you queried and which came back empty. "OpenAlex had no match" is a fact; "that paper doesn't exist" is a much bigger claim than one API can support.
Run mode — how to drive scripts/litrun.py
The launcher installs each supported tool into its own venv under ~/.lit-review-tools/
(uses uv if present, else python -m venv) and reads API keys from one shared
~/.lit-review-tools/.env. Machine-readable recipes: recipes/recipes.json.
Typical flow when the user wants to use a tool:
python3 scripts/litrun.py doctor— check toolchain + which API keys are already set.python3 scripts/litrun.py info <id>— confirm what the tool needs (entry, required env).- If a required key is missing, ask the user for it, then
litrun.py env --set KEY=VALUE(never echo the value back in full). python3 scripts/litrun.py run <id> -- <tool args>— installs on first use, then runs. For PDF tools pass the real file path; e.g.run mineru -- -p paper.pdf -o ./out -b pipeline.- For MCP servers, don't "run" them —
litrun.py mcp <id>prints the client config block to register in Claude Code / Cursor.
Commands: list [--category C] [--kind K] · info <id> · doctor · env [--set K=V] · install <id> · run <id> -- <args> · mcp <id> [--storage PATH] [--client claude|cursor] · ui <id>.
Runnable ids by kind:
- python-cli (auto install+run):
mineru,marker,docling(PDF→Markdown) ·paper-qa(cited Q&A) ·asreview(PRISMA screening UI) - python-script, zero install (stdlib only):
papers-fetch(6-source deduplicated search) ·oa-resolve(DOI → open-access full text) - python-script (bundled, auto install+run):
arxiv-fetch(arXiv PDFs) ·openalex-fetch(OpenAlex; PDF or abstract .txt) ·pubmed-fetch(PubMed abstracts) — all keyless for light use.papers-fetchsupersedes all three when you want coverage rather than one index. - python-lib (install + run example):
gpt-researcher,storm(deep research; need API keys) ·scholarly,pyalex(API clients) - mcp-server (install +
mcpconfig):arxiv-mcp-server,paper-search-mcp,zotero-mcp
For gpt-researcher and storm, litrun.py ui <id> clones the repo and launches the full web UI (GPT Researcher → FastAPI at :8000; STORM → Streamlit at :8501). These are long-running servers — launch them with a background Bash call and tell the user the URL. gpt-researcher's UI needs OPENAI_API_KEY + TAVILY_API_KEY set first (litrun writes them into the repo's .env); STORM takes its keys in the app sidebar.
Chained pipelines
For multi-tool tasks, prefer a named workflow over hand-wiring steps: litrun.py workflow list then litrun.py workflow run <id> [--input PATH] [--query "..."] [--question "..."] [--max N]. Built-ins:
pdf-to-markdown— a PDF/folder → clean Markdown (MinerU)pdf-corpus-qa— a folder of PDFs → citation-backed answer (PaperQA2)pdf-md-then-qa— convert to Markdown and answer a question over the corpustopic-to-pdfs— arXiv query → download top-N PDFs (arxiv-fetch, no key)topic-to-review— arXiv query → download PDFs → citation-backed answer (PaperQA2). The end-to-end "retrieve then review" pipeline; no MCP client needed. NeedsOPENAI_API_KEYfor the QA step.topic-to-review-multi— retrieve from arXiv + OpenAlex into one corpus → citation-backed answer. Broader coverage; resilient if one source is rate-limited.topic-to-related-work— retrieve (arXiv + OpenAlex) → PaperQA2 drafts a cited related-work paragraph synthesizing themes/methods/gaps. NeedsOPENAI_API_KEY.topic-to-corpus— six-source search → deduplicated corpus. Keyless, no install — the default first step for any review.topic-to-fulltext— six-source search → open-access full text for every DOI it can legally get. Keyless.topic-to-fulltext-review— the strongest path: search → dedup → full text → PaperQA2 answers over full texts, not abstracts.OPENAI_API_KEYfor the last step only.
Prefer the topic-to-fulltext* workflows over topic-to-review-multi: same idea, six
sources instead of two, DOI-level deduplication, and it retrieves the papers rather
than the abstracts. Biomedical topics need no special casing — PubMed and Europe PMC
are already in the source set.
Add --dry-run first to show the exact resolved step commands without executing — good for confirming paths with the user before a heavy run. Workflows fail fast if a required API key is missing.
Guardrails: installs and downloads happen under the user's home and hit the network — for a heavy first install (marker/docling pull in PyTorch) say so before running. Never fabricate API keys. If a run fails, show the real error rather than claiming success. Paths in this file (scripts/…, recipes/…) are relative to this skill's directory.
Recommend mode — how to route
- Identify which stage of the lit-review workflow the user is on (search → read → extract → synthesize → screen → cite-check → write/review).
- Match it to a category below and recommend the ⭐ editor's pick first, then 1–2 alternatives.
- For anything beyond the top pick — full star counts, every project in a category, or a category not summarized here — read
reference/catalog.md. Do not guess project names or URLs; pull them from the catalog. - Give a one-line "why this one" tied to the user's constraint (Claude Code vs. standalone, open vs. commercial, privacy/local, medical, etc.). If the pick is a runnable id above, offer to install/run it.
⚡ 30-second picker
Just need the papers themselves (topic / DOI / OA PDF) ──▶ Look up mode — no install ⭐
Use Claude Code, want end-to-end research→paper ──────────▶ academic-research-skills ⭐
Want AI to research a topic → cited report ───────────────▶ GPT Researcher / STORM
Want fully autonomous "idea → submittable paper" ────────▶ AI-Scientist-v2 / AutoResearchClaw
Citation-backed Q&A over a pile of PDFs ──────────────────▶ PaperQA2
Rigorous PRISMA review (thousands of abstracts) ─────────▶ ASReview / prismAId
Clean Markdown from PDFs to feed an LLM ─────────────────▶ MinerU / Docling / marker
Lit capabilities inside Claude / Cursor (MCP) ───────────▶ paper-search-mcp / zotero-mcp
Chat with your library inside Zotero ────────────────────▶ zotero-gpt / PapersGPT
Pre-submission AI peer review ───────────────────────────▶ open_reviewer / ai-peer-review
Categories (top pick per category)
| Category | Editor's pick ⭐ | When |
|---|---|---|
| All-in-one research agents & skills | academic-research-skills | Claude Code user wanting research→write→review→revise, with integrity/citation gates |
| Deep research & auto-survey | STORM / gpt-researcher | Topic → cited survey / report / related-work |
| Autonomous science (idea→paper) | AI-Scientist(-v2) / AutoResearchClaw | Fully automated discovery: lit + hypotheses + experiments + writing |
| Literature Q&A / RAG | paper-qa (PaperQA2) | Citation-backed answers over a PDF corpus |
| Systematic review & screening | ASReview | Active-learning screening of thousands of abstracts (PRISMA) |
| MCP servers | zotero-mcp / arxiv-mcp-server | Wire papers into Claude / Cursor / Cline |
| Zotero / Obsidian integration | zotero-gpt | Chat with your library inside your reference manager |
| PDF → structured extraction | MinerU / docling / marker | Turn PDFs into clean Markdown/JSON for LLMs |
| Citation graphs & API clients | scholarly / pyalex | Citation-network analysis; scripting academic DBs |
| Writing & peer-review assistants | open_reviewer / ai-peer-review | Draft, polish, and pre-submission review |
| Awesome lists | Awesome-Auto-Research-Tools | Browse the whole landscape |
Decision table (map need → recommendation)
| User's need | Recommend |
|---|---|
| Claude Code, end-to-end research→paper | academic-research-skills (most complete, #1 in space) |
| Generic "research this topic for me" agent | GPT Researcher / STORM |
| Wiki/survey-style long-form with citations | STORM / Co-STORM |
| Fully autonomous "idea → submittable paper" | AI-Scientist-v2 / AutoResearchClaw |
| Cited Q&A over many PDFs | PaperQA / PaperQA2 |
| Rigorous PRISMA systematic review | ASReview or prismAId |
| PDF → clean Markdown for an LLM | MinerU / Docling / marker |
| Lit capabilities in an MCP client | paper-search-mcp / zotero-mcp |
| Chat with library inside Zotero | zotero-gpt / PapersGPT |
| AI pre-review before submission | open_reviewer / ai-peer-review |
| Just want to browse the landscape | The Awesome lists section |
Notes & caveats
- Open-source is prioritized. Commercial/closed tools (Elicit, Consensus, Scite, SciSpace, Research Rabbit, Connected Papers) are listed for reference only — see the catalog's commercial section.
- Star counts drift. The catalog's numbers are periodic GitHub-API snapshots — treat as rough popularity signals, not exact. For live numbers, point the user at the repo.
- Match the constraint, not just the task. Privacy/local →
local-deep-research; medical →medsci-skills/paperai; Codex instead of Claude →academic-research-skills-codex.
Full catalog with every project, star count, and one-line description: reference/catalog.md.
Files (lit-review-agent-tools)
-
recipes
-
recipes.json 12.5 KB
{ "_meta": { "description": "Runnable-tool manifest for litrun.py. Commands verified against upstream READMEs/PyPI (2025-2026). Star counts and the full 70+ catalog live in ../reference/catalog.md; only tools that can be installed & driven programmatically are listed here.", "source": "https://github.com/brycewang-stanford/lit-review-agent-tools" }, "tools": [ { "id": "mineru", "name": "MinerU", "category": "pdf-extraction", "kind": "python-cli", "repo": "https://github.com/opendatalab/MinerU", "pip": ["mineru[all]"], "entry": "mineru", "example": "mineru -p input.pdf -o ./output -b pipeline", "env": [], "env_optional": [], "notes": "PDF/Office/image -> LLM-ready Markdown/JSON. No API key. GPU optional; add '-b pipeline' for CPU-only. ~16GB RAM recommended." }, { "id": "marker", "name": "marker", "category": "pdf-extraction", "kind": "python-cli", "repo": "https://github.com/datalab-to/marker", "pip": ["marker-pdf"], "entry": "marker_single", "example": "marker_single /path/to/file.pdf --output_dir ./output", "env": [], "env_optional": ["GOOGLE_API_KEY"], "notes": "Fast PDF/doc -> Markdown/JSON. No key for base use; '--use_llm' boost needs an LLM key (e.g. GOOGLE_API_KEY). Pulls in PyTorch; GPU recommended. Use 'marker <dir>' for batch." }, { "id": "docling", "name": "docling", "category": "pdf-extraction", "kind": "python-cli", "repo": "https://github.com/docling-project/docling", "pip": ["docling"], "entry": "docling", "example": "docling https://arxiv.org/pdf/2206.01062", "env": [], "env_optional": [], "notes": "IBM document parser -> Markdown/JSON/HTML. No API key. Accepts local paths or URLs. Optional VLM pipeline: 'docling --pipeline vlm --vlm-model granite_docling <src>'." }, { "id": "paper-qa", "name": "PaperQA2", "category": "rag-qa", "kind": "python-cli", "repo": "https://github.com/Future-House/paper-qa", "pip": ["paper-qa"], "entry": "pqa", "example": "pqa ask 'What methods does this corpus use?'", "env": ["OPENAI_API_KEY"], "env_optional": ["ANTHROPIC_API_KEY", "GEMINI_API_KEY", "CROSSREF_API_KEY", "SEMANTIC_SCHOLAR_API_KEY"], "notes": "Citation-backed RAG QA. Run pqa from a directory containing your PDFs. Needs an LLM key (OpenAI by default; Anthropic/Gemini/local Ollama also supported). Python 3.11+." }, { "id": "asreview", "name": "ASReview LAB", "category": "systematic-review", "kind": "python-cli", "repo": "https://github.com/asreview/asreview", "pip": ["asreview"], "entry": "asreview", "example": "asreview lab", "env": [], "env_optional": [], "notes": "Active-learning screening for PRISMA/systematic reviews. 'asreview lab' launches a local web UI in the browser. Imports CSV/RIS/XLSX. No API key, no GPU. Python 3.10+." }, { "id": "gpt-researcher", "name": "GPT Researcher", "category": "deep-research", "kind": "python-lib", "repo": "https://github.com/assafelovic/gpt-researcher", "pip": ["gpt-researcher"], "entry": null, "example": "python -c \"import asyncio; from gpt_researcher import GPTResearcher; r=GPTResearcher(query='YOUR TOPIC'); asyncio.run(r.conduct_research()); print(asyncio.run(r.write_report()))\"", "env": ["OPENAI_API_KEY", "TAVILY_API_KEY"], "env_optional": ["OPENAI_BASE_URL"], "notes": "Async deep-research library (no console script). Needs BOTH OPENAI_API_KEY and TAVILY_API_KEY. For the web UI, clone the repo and run 'python -m uvicorn main:app --reload'.", "clone_for_ui": true, "ui": { "requirements": ["requirements.txt"], "write_env": true, "run": ["python", "-m", "uvicorn", "main:app", "--reload"], "url": "http://localhost:8000", "notes": "FastAPI + web frontend. Reads keys from a .env in the repo root (litrun writes it for you). Serves at :8000." } }, { "id": "storm", "name": "STORM / Co-STORM", "category": "deep-research", "kind": "python-lib", "repo": "https://github.com/stanford-oval/storm", "pip": ["knowledge-storm"], "entry": null, "example": "See repo examples/ — instantiate STORMWikiRunner with an LM + retriever, then runner.run(topic=..., do_research=True, do_generate_article=True)", "env": ["OPENAI_API_KEY", "YDC_API_KEY"], "env_optional": ["BING_SEARCH_API_KEY"], "notes": "Wikipedia-style survey generator (library). Needs an LLM key AND a retriever key (YDC_API_KEY from You.com, or Bing). Runnable demos + Streamlit UI live in the cloned repo's examples/frontend.", "clone_for_ui": true, "ui": { "requirements": ["requirements.txt", "frontend/demo_light/requirements.txt"], "write_env": false, "run": ["streamlit", "run", "frontend/demo_light/storm.py"], "url": "http://localhost:8501", "notes": "Streamlit demo (frontend/demo_light). Enter your OpenAI + You.com (YDC) keys in the app's sidebar / secrets. Serves at :8501." } }, { "id": "scholarly", "name": "scholarly", "category": "api-client", "kind": "python-lib", "repo": "https://github.com/scholarly-python-package/scholarly", "pip": ["scholarly"], "entry": null, "example": "python -c \"from scholarly import scholarly; print(next(scholarly.search_author('Steven A Cholewiak')))\"", "env": [], "env_optional": [], "notes": "Google Scholar client (library). No key, but Scholar blocks bare IPs on search_pubs/citedby — you typically need a proxy (built-in ProxyGenerator / ScraperAPI)." }, { "id": "pyalex", "name": "pyalex", "category": "api-client", "kind": "python-lib", "repo": "https://github.com/J535D165/pyalex", "pip": ["pyalex"], "entry": null, "example": "python -c \"import pyalex; pyalex.config.email='you@example.com'; from pyalex import Works; print(Works()['W2741809807']['title'])\"", "env": [], "env_optional": ["OPENALEX_API_KEY"], "notes": "OpenAlex API client (library). Set pyalex.config.email for the polite pool. As of ~2026 OpenAlex asks for an API key for higher limits: pyalex.config.api_key='...'." }, { "id": "arxiv-fetch", "name": "arXiv search & download", "category": "retrieval", "kind": "python-script", "repo": "https://github.com/lukasschwab/arxiv.py", "pip": ["arxiv"], "script": "fetch_arxiv.py", "example": "arxiv-fetch --query 'retrieval augmented generation' --max 10 --outdir ./papers", "env": [], "env_optional": [], "notes": "Bundled fetcher (uses the free 'arxiv' package). Searches arXiv and downloads PDFs + a manifest.json into --outdir. No API key. Query supports cat:, AND/OR. Feeds retrieval workflows into PaperQA2." }, { "id": "openalex-fetch", "name": "OpenAlex search & download", "category": "retrieval", "kind": "python-script", "repo": "https://openalex.org", "pip": ["requests"], "script": "fetch_openalex.py", "example": "openalex-fetch --query 'active learning systematic review' --max 10 --outdir ./papers", "env": [], "env_optional": ["OPENALEX_API_KEY", "OPENALEX_MAILTO"], "notes": "Bundled fetcher over the OpenAlex API (240M+ works, all fields). Downloads the OA PDF when available, else saves title+abstract as .txt (PaperQA2 reads text too). No key needed for light use; set OPENALEX_API_KEY / OPENALEX_MAILTO for higher limits. Appends to manifest.json for multi-source runs." }, { "id": "pubmed-fetch", "name": "PubMed search & abstracts", "category": "retrieval", "kind": "python-script", "repo": "https://www.ncbi.nlm.nih.gov/pmc/tools/developers/", "pip": ["requests"], "script": "fetch_pubmed.py", "example": "pubmed-fetch --query 'CRISPR off-target detection' --max 10 --outdir ./papers", "env": [], "env_optional": ["NCBI_API_KEY", "NCBI_EMAIL"], "notes": "Bundled fetcher over NCBI E-utilities. Saves each article's title+abstract as .txt (+ manifest) — PubMed rarely exposes downloadable PDFs. Best for biomedical corpora. No key needed for light use; set NCBI_API_KEY / NCBI_EMAIL to raise the rate limit." }, { "id": "papers-fetch", "name": "Multi-source paper search (6 APIs, deduplicated)", "category": "retrieval", "kind": "python-script", "repo": "https://github.com/brycewang-stanford/lit-review-agent-tools", "pip": [], "script": "fetch_papers.py", "example": "papers-fetch --query 'retrieval augmented generation' --sources openalex,crossref,semanticscholar --max 20 --outdir ./papers", "env": [], "env_optional": ["OPENALEX_API_KEY", "OPENALEX_MAILTO", "S2_API_KEY", "NCBI_API_KEY", "NCBI_EMAIL", "CROSSREF_MAILTO"], "notes": "Bundled, standard-library only — no install, no API key. Searches any subset of OpenAlex, Crossref, Semantic Scholar, PubMed, Europe PMC and arXiv, merges the hits into one normalised record list deduplicated by DOI (add --dedup-titles to also fold preprint/published pairs), and writes results.json + one .txt per record + manifest.json. A failing source is reported and skipped, never fatal. Use it instead of arxiv-fetch/openalex-fetch/pubmed-fetch when you want coverage rather than one index." }, { "id": "oa-resolve", "name": "Open-access full-text resolver", "category": "retrieval", "kind": "python-script", "repo": "https://github.com/brycewang-stanford/lit-review-agent-tools", "pip": [], "script": "resolve_oa.py", "example": "oa-resolve --from-json ./papers/results.json --outdir ./papers --email you@example.com", "env": [], "env_optional": ["UNPAYWALL_EMAIL", "CORE_API_KEY"], "notes": "Bundled, standard-library only. Turns DOIs into downloadable full text by walking Unpaywall -> OpenAlex -> Europe PMC OA XML -> arXiv -> CORE, stopping at the first hit. Writes PDFs/plain text plus oa_report.json recording how each DOI resolved, separating 'closed' (paywalled) from 'not-in-crossref' (no such record — usually an invented citation). Never bypasses a paywall. This is the step that turns an abstract-only corpus into a full-text one before PaperQA2." }, { "id": "arxiv-mcp-server", "name": "arxiv-mcp-server", "category": "mcp-server", "kind": "mcp-server", "repo": "https://github.com/blazickjp/arxiv-mcp-server", "pip": [], "entry": null, "env": [], "env_optional": [], "notes": "MCP server (stdio) to search/download/analyze arXiv. Runs via uvx (no install needed). Registered in an MCP client, not queried directly.", "mcp": { "key": "arxiv", "launcher": "uvx", "package": "arxiv-mcp-server", "args_template": ["arxiv-mcp-server", "--storage-path", "{storage}"], "env": {} } }, { "id": "paper-search-mcp", "name": "paper-search-mcp", "category": "mcp-server", "kind": "mcp-server", "repo": "https://github.com/openags/paper-search-mcp", "pip": ["paper-search-mcp"], "entry": null, "env": [], "env_optional": ["PAPER_SEARCH_MCP_UNPAYWALL_EMAIL"], "notes": "MCP server to search/download across 20+ sources (arXiv, PubMed, bioRxiv, S2, OpenAlex...). Set an Unpaywall email for OA lookups. Verify exact optional env-var names in the README before relying on them.", "mcp": { "key": "paper-search-mcp", "launcher": "venv-module", "module": "paper_search_mcp.server", "env": {"PAPER_SEARCH_MCP_UNPAYWALL_EMAIL": "your@email.com"} } }, { "id": "zotero-mcp", "name": "zotero-mcp", "category": "mcp-server", "kind": "mcp-server", "repo": "https://github.com/54yyyu/zotero-mcp", "pip": ["zotero-mcp"], "entry": "zotero-mcp", "env": [], "env_optional": ["ZOTERO_API_KEY", "ZOTERO_LIBRARY_ID", "OPENAI_API_KEY"], "notes": "MCP server bridging your Zotero library. Easiest: after install run 'zotero-mcp setup' to auto-write the client config. Local mode (ZOTERO_LOCAL=true) needs the Zotero desktop app running with the local API enabled; Web API mode needs ZOTERO_API_KEY + ZOTERO_LIBRARY_ID.", "mcp": { "key": "zotero", "launcher": "venv-entry", "entry": "zotero-mcp", "env": {"ZOTERO_LOCAL": "true"} } } ] } -
workflows.json 8.5 KB
{ "_meta": { "description": "Named multi-step pipelines for litrun.py. Each step invokes a python-cli tool from recipes.json with placeholder-substituted args. Placeholders: {input}, {question}, {workdir}. Steps run in order, sharing one {workdir} under ~/.lit-review-tools/workspace/runs/<id>.", "source": "https://github.com/brycewang-stanford/lit-review-agent-tools" }, "workflows": [ { "id": "pdf-to-markdown", "name": "PDF(s) → clean Markdown", "description": "Convert a PDF file or a folder of PDFs into LLM-ready Markdown with MinerU.", "params": {"input": "path to a PDF file or a folder of PDFs"}, "steps": [ {"tool": "mineru", "args": ["-p", "{input}", "-o", "{workdir}/markdown", "-b", "pipeline"]} ], "outputs": "Markdown/JSON under {workdir}/markdown" }, { "id": "pdf-corpus-qa", "name": "PDF corpus → cited answer", "description": "Ask a question over a folder of PDFs with PaperQA2; every claim comes back with a citation.", "params": { "input": "folder containing the PDFs to search", "question": "the question to answer over the corpus" }, "steps": [ {"tool": "paper-qa", "cwd": "{input}", "args": ["ask", "{question}"]} ], "outputs": "A citation-backed answer printed to stdout (PaperQA2 caches its index in {input})." }, { "id": "pdf-md-then-qa", "name": "PDFs → Markdown + cited answer", "description": "Two-step: convert PDFs to Markdown (MinerU) for a clean artifact, then answer a question over the same corpus with PaperQA2.", "params": { "input": "folder containing the PDFs", "question": "the question to answer over the corpus" }, "steps": [ {"tool": "mineru", "args": ["-p", "{input}", "-o", "{workdir}/markdown", "-b", "pipeline"]}, {"tool": "paper-qa", "cwd": "{input}", "args": ["ask", "{question}"]} ], "outputs": "Markdown under {workdir}/markdown + a citation-backed answer on stdout." }, { "id": "topic-to-pdfs", "name": "Topic → arXiv PDFs", "description": "Search arXiv for a query and download the top-N PDFs (+ manifest) into a folder.", "params": { "query": "arXiv search query", "max": "how many PDFs to download (default 10)" }, "defaults": {"max": "10"}, "steps": [ {"tool": "arxiv-fetch", "args": ["--query", "{query}", "--max", "{max}", "--outdir", "{workdir}/pdfs"]} ], "outputs": "PDFs + manifest.json under {workdir}/pdfs" }, { "id": "topic-to-review", "name": "Topic → retrieved corpus → cited answer", "description": "End-to-end lightweight review: search arXiv, download the top-N PDFs, then answer a question over them with PaperQA2 (citation-backed). No MCP client needed.", "params": { "query": "arXiv search query defining the corpus", "question": "the question to answer over the retrieved papers", "max": "how many PDFs to retrieve (default 10)" }, "defaults": {"max": "10"}, "steps": [ {"tool": "arxiv-fetch", "args": ["--query", "{query}", "--max", "{max}", "--outdir", "{workdir}/pdfs"]}, {"tool": "paper-qa", "cwd": "{workdir}/pdfs", "args": ["ask", "{question}"]} ], "outputs": "Retrieved PDFs under {workdir}/pdfs + a citation-backed answer on stdout." }, { "id": "topic-to-review-multi", "name": "Topic → multi-source corpus → cited answer", "description": "Retrieve from BOTH arXiv and OpenAlex into one corpus, then answer a question over it with PaperQA2. Broader coverage than a single source; resilient if one source is rate-limited.", "params": { "query": "search query defining the corpus", "question": "the question to answer over the retrieved papers", "max": "how many items per source (default 8)" }, "defaults": {"max": "8"}, "steps": [ {"tool": "arxiv-fetch", "args": ["--query", "{query}", "--max", "{max}", "--outdir", "{workdir}/corpus"]}, {"tool": "openalex-fetch", "args": ["--query", "{query}", "--max", "{max}", "--outdir", "{workdir}/corpus"]}, {"tool": "paper-qa", "cwd": "{workdir}/corpus", "args": ["ask", "{question}"]} ], "outputs": "A merged arXiv+OpenAlex corpus under {workdir}/corpus + a citation-backed answer on stdout." }, { "id": "topic-to-related-work", "name": "Topic → related-work paragraph", "description": "Retrieve papers (arXiv + OpenAlex), then have PaperQA2 draft a citation-backed related-work / literature-review paragraph synthesizing themes, methods, and gaps.", "params": { "query": "the topic / search query for the related-work section", "max": "how many items per source (default 10)" }, "defaults": {"max": "10"}, "steps": [ {"tool": "arxiv-fetch", "args": ["--query", "{query}", "--max", "{max}", "--outdir", "{workdir}/corpus"]}, {"tool": "openalex-fetch", "args": ["--query", "{query}", "--max", "{max}", "--outdir", "{workdir}/corpus"]}, {"tool": "paper-qa", "cwd": "{workdir}/corpus", "args": ["ask", "Write a ~250-word related-work / literature-review paragraph on '{query}'. Synthesize the main themes, methods, and open gaps across these papers, and support every claim with an inline citation. Write in an academic tone suitable for a paper's Related Work section."]} ], "outputs": "A cited related-work paragraph on stdout, grounded in {workdir}/corpus." }, { "id": "topic-to-corpus", "name": "Topic → deduplicated multi-source corpus", "description": "Search six scholarly APIs at once (OpenAlex, Crossref, Semantic Scholar, PubMed, Europe PMC, arXiv), deduplicate by DOI, and save one .txt per unique paper. No install, no API key — the keyless starting point for any review.", "params": { "query": "search query defining the corpus", "max": "how many hits per source before dedup (default 15)" }, "defaults": {"max": "15"}, "steps": [ {"tool": "papers-fetch", "args": ["--query", "{query}", "--sources", "openalex,crossref,semanticscholar,pubmed,europepmc,arxiv", "--max", "{max}", "--dedup-titles", "--outdir", "{workdir}/corpus"]} ], "outputs": "results.json (normalised records + per-source counts) and one .txt per unique paper under {workdir}/corpus" }, { "id": "topic-to-fulltext", "name": "Topic → multi-source search → open-access full text", "description": "Search six APIs, deduplicate, then resolve every DOI to a legally downloadable full text via Unpaywall → OpenAlex → Europe PMC → arXiv → CORE. Turns an abstract-only corpus into a real full-text one. Keyless.", "params": { "query": "search query defining the corpus", "max": "how many hits per source before dedup (default 15)" }, "defaults": {"max": "15"}, "steps": [ {"tool": "papers-fetch", "args": ["--query", "{query}", "--sources", "openalex,crossref,semanticscholar,pubmed,europepmc,arxiv", "--max", "{max}", "--dedup-titles", "--outdir", "{workdir}/corpus"]}, {"tool": "oa-resolve", "args": ["--from-json", "{workdir}/corpus/results.json", "--outdir", "{workdir}/corpus", "--skip-existing"]} ], "outputs": "PDFs / full-text .txt + oa_report.json (how each DOI resolved, including the closed ones) under {workdir}/corpus" }, { "id": "topic-to-fulltext-review", "name": "Topic → full-text corpus → cited answer", "description": "The strongest end-to-end path: six-source search → dedup → open-access full-text resolution → PaperQA2 answers over the full texts rather than the abstracts. Needs OPENAI_API_KEY for the final step only.", "params": { "query": "search query defining the corpus", "question": "the question to answer over the retrieved papers", "max": "how many hits per source before dedup (default 12)" }, "defaults": {"max": "12"}, "steps": [ {"tool": "papers-fetch", "args": ["--query", "{query}", "--sources", "openalex,crossref,semanticscholar,pubmed,europepmc,arxiv", "--max", "{max}", "--dedup-titles", "--outdir", "{workdir}/corpus"]}, {"tool": "oa-resolve", "args": ["--from-json", "{workdir}/corpus/results.json", "--outdir", "{workdir}/corpus", "--skip-existing"]}, {"tool": "paper-qa", "cwd": "{workdir}/corpus", "args": ["ask", "{question}"]} ], "outputs": "A full-text corpus under {workdir}/corpus + a citation-backed answer on stdout." } ] }
-
-
reference
-
apis
-
arxiv.md 2.6 KB
# arXiv 2.4M+ preprints in physics, maths, CS, quantitative biology, statistics, economics — and, uniquely in this set, **every record has a free PDF**. No key. Base URL `http://export.arxiv.org/api/query`. Used by `fetch_papers.py` (search), `fetch_arxiv.py` (search + download) and `resolve_oa.py` (last-resort PDF for anything with an arXiv id). ## Limits One request per **3 seconds**, `max_results` ≤ 2000 per call (use `start` to page, and stay under ~30k results per query). arXiv publishes and enforces this; exceeding it gets your IP throttled. Bulk downloading of PDFs is against their terms — for whole-corpus work they publish an S3 bucket instead. ## Query ``` GET /api/query?search_query=all:transformer&start=0&max_results=20&sortBy=relevance ``` | Prefix | Field | |---|---| | `all:` | everything | | `ti:` `abs:` `au:` | title, abstract, author | | `cat:` | category, e.g. `cat:cs.CL`, `cat:stat.ME` | | `id_list=2103.15348` | fetch specific ids instead of searching | Boolean `AND` / `OR` / `ANDNOT`, quotes for phrases, parentheses for grouping — all must be URL-encoded. `sortBy`: `relevance` · `submittedDate` · `lastUpdatedDate`. **There is no date filter.** The legacy API cannot express "since 2023"; either sort by `submittedDate` and take the head, or over-fetch and filter client-side (what `fetch_papers.py --from-year` does). ## Response — Atom XML, not JSON ```xml <entry> <id>http://arxiv.org/abs/2103.15348v2</id> <published>2021-03-29T…</published><updated>…</updated> <title>…</title><summary>…</summary> <author><name>…</name></author> <arxiv:doi>10.1145/…</arxiv:doi> <arxiv:journal_ref>ACL 2021</arxiv:journal_ref> <arxiv:primary_category term="cs.CL"/> <link title="pdf" href="http://arxiv.org/pdf/2103.15348v2"/> </entry> ``` Namespaces: `{"a": "http://www.w3.org/2005/Atom", "arxiv": "http://arxiv.org/schemas/atom"}`. `title` and `summary` arrive with hard line wrapping — collapse whitespace or every title in your manifest carries newlines. The id ends in a **version suffix** (`v2`); `https://arxiv.org/pdf/{id}` works with or without it. `arxiv:doi` is present only once the preprint is published — most entries have none, which is why arXiv records dedupe by title rather than DOI in `fetch_papers.py`, and why `--dedup-titles` exists to fold the preprint into its published twin. ## Good and bad at Good: guaranteed full text, fast, keyless, strong CS/physics/ML coverage, category filters that actually work. Bad: no peer review (a hit is not evidence), no date filter, no citation counts, and nothing biomedical or social-scientific worth relying on. -
crossref.md 2.8 KB
# Crossref The DOI registration agency's own metadata: ~160M records. This is the authority on "does this DOI exist and what is it", plus journal, funder, license and reference-list metadata. No key. Base URL `https://api.crossref.org`. Used by `fetch_papers.py` (search) and by recipe 03, which verifies citations against it. ## Auth & limits - Public pool ~5 req/s. Adding `?mailto=you@example.com` (`CROSSREF_MAILTO`) puts you in the **polite pool** with roughly double the allowance and human contact if you misbehave. A descriptive `User-Agent` matters here more than on other APIs. ## Endpoints ``` GET /works/{doi} one record — the canonical DOI check GET /works?query.bibliographic=…&rows=20 search (whole-reference matching) GET /works?query.title=…&query.author=… fielded search GET /journals/{issn}/works everything in a journal GET /members/{id}/works | /funders/{id}/works ``` | Param | Notes | |---|---| | `query.bibliographic` | best field for "here is a reference string, find it" | | `filter` | `from-pub-date:2023-01-01`, `type:journal-article`, `has-abstract:true`, `is-update:true` (retraction/correction notices), `license.url:…` | | `select` | trims the response — **list endpoints only** | | `rows` / `offset` | ≤1000 rows; use `cursor=*` beyond 10k | | `sort` / `order` | `relevance`, `published`, `is-referenced-by-count` | ## The 400 that wastes an hour `select` is accepted on `/works?…` but **rejected on `/works/{doi}`**: ``` GET /works?query=prisma&rows=1&select=DOI,title → 200 GET /works/10.1136/bmj.n160?select=DOI → 400 ``` If a single-DOI lookup 400s, drop `select` before concluding the DOI is bad. ## Response shape ```json {"status":"ok","message":{"DOI":"10.1136/bmj.n160","title":["PRISMA 2020 …"], "container-title":["BMJ"],"issued":{"date-parts":[[2021,3,29]]}, "author":[{"given":"Matthew J","family":"Page","ORCID":"…"}], "is-referenced-by-count":9781,"type":"journal-article", "abstract":"<jats:p>…</jats:p>","reference":[{"DOI":"…","unstructured":"…"}], "link":[{"URL":"…","content-type":"application/pdf"}], "update-to":[{"type":"retraction","DOI":"…"}]} ``` - `title` and `container-title` are **arrays**; take `[0]`. - `abstract` is JATS XML when present at all (many publishers deposit none) — strip tags. - `reference` gives you the paper's own bibliography, which is how you walk citations backwards without Semantic Scholar. - `update-to` / `filter=is-update:true` surfaces retraction and correction notices. ## Good and bad at Good: DOI truth, reference strings → records, funder/license metadata, retraction notices. Bad: topical relevance search (it matches strings, not meaning), abstract coverage, and it has no notion of open-access location — pair it with Unpaywall or OpenAlex. -
europepmc.md 2.8 KB
# Europe PMC The most underrated API in this set. It indexes PubMed **plus** preprints, patents, agricultural and theses records, answers in JSON (no XML dance), needs no key, and serves open-access **full text** from one endpoint. Base URL `https://www.ebi.ac.uk/europepmc/webservices/rest`. Used by `fetch_papers.py` (search) and `resolve_oa.py` (OA full-text step). ## Endpoints ``` GET /search?query=…&format=json&pageSize=25&resultType=core GET /search?query=…&cursorMark=* deep pagination (follow nextCursorMark) GET /{PMCID}/fullTextXML OA full text as JATS GET /{source}/{id}/textMinedTerms | /citations | /references ``` `resultType`: `idlist` (ids only) · `lite` (default, no abstract) · `core` (**abstract, author list, journal, full-text URLs — what you almost always want**). ## Query language | Want | Query | |---|---| | Preprints only | `(machine learning) AND SRC:PPR` — 4,642 hits for that example | | Open access only | `… AND OPEN_ACCESS:Y` | | Has full text in PMC | `… AND HAS_FT:Y` | | Date range | `… AND (FIRST_PDATE:[2022-01-01 TO 3000-01-01])` | | By DOI | `DOI:"10.1136/bmj.n160"` | | Fielded | `TITLE:"…"`, `AUTH:"Page MJ"`, `JOURNAL:"BMJ"`, `MESH:"Neoplasms"` | `SRC:PPR` is the practical answer to "search preprints by keyword" — bioRxiv and medRxiv's own APIs cannot do it (see [`open-access.md`](open-access.md)). ## Response shape ```json {"hitCount":4642,"nextCursorMark":"…", "resultList":{"result":[{"id":"33781993","source":"MED","pmid":"33781993", "pmcid":"PMC8005925","doi":"10.1136/bmj.n160","title":"…","abstractText":"…", "authorList":{"author":[{"fullName":"Page MJ"}]}, "journalInfo":{"journal":{"title":"BMJ"}},"pubYear":"2021", "isOpenAccess":"Y","citedByCount":9781, "fullTextUrlList":{"fullTextUrl":[{"documentStyle":"pdf","availabilityCode":"OA","url":"…"}]}}]}} ``` `isOpenAccess` is the string `"Y"`/`"N"`, not a boolean. `source` tells you the sub-corpus: `MED` (PubMed), `PPR` (preprint), `PMC`, `PAT`, `AGR`, `CTX`. `abstractText` may carry inline HTML — strip tags. ## Full text ``` GET /PMC8005925/fullTextXML ``` Returns JATS for the **OA subset only** (`isOpenAccess:"Y"`); otherwise 404. Flattening the whole document with a naive `itertext()` prepends a pile of bibliographic tokens — extract `article-title` + `abstract` + `body` instead, which is what `jats_to_text()` in `resolve_oa.py` does. A 200 response under ~2 KB is a stub, not an article. ## Good and bad at Good: one keyless JSON call for metadata *and* full text; preprint keyword search; citation counts; the widest biomedical net available without credentials. Bad: coverage outside life sciences; relevance ranking skews recent, so pair it with OpenAlex when you want the classic papers in a field. -
open-access.md 4 KB
# Getting the actual file: Unpaywall, CORE, bioRxiv/medRxiv Metadata tells you a paper exists. These APIs tell you whether you may read it. `resolve_oa.py` chains them; this file documents each one and the order, so you can debug a DOI that "should" be open. ## The chain, and why this order ``` 1. Unpaywall every publisher, every field — but needs an email 2. OpenAlex same underlying OA data, keyless, sometimes a landing page not a PDF 3. Europe PMC biomedical OA full text as XML — no PDF needed 4. arXiv anything with an arXiv id, guaranteed PDF 5. CORE repository aggregator, catches institutional copies — needs a key ``` Highest precision first, widest net last. Recorded result on a 47-paper multi-source corpus with **no keys at all** (Unpaywall and CORE skipped): 33/47 resolved — 16 via OpenAlex, 10 via arXiv, 7 via Europe PMC. See [`recipes/06-multi-source-search`](../../../../recipes/06-multi-source-search/). **Nothing here bypasses a paywall.** A closed paper stays closed and is reported as `closed` in `oa_report.json`. Do not route around that with Sci-Hub or scraped mirrors: it is unlawful in most jurisdictions and gets institutions blocked. Interlibrary loan and author-request are the legitimate paths, and both need a human. ## Unpaywall ``` GET https://api.unpaywall.org/v2/{doi}?email=you@example.com ``` - The `email` parameter is **mandatory** — without it you get HTTP 422, not a warning. Never invent one; ask the user, or run without this step (`resolve_oa.py` does). - 100k calls/day. Batch downloads are offered as a data dump instead. - Read `is_oa`, then `best_oa_location.url_for_pdf` (may be null when only a landing page exists), `oa_status` (`gold`/`green`/`hybrid`/`bronze`/`closed`), and `.version` — `submittedVersion` means you got the preprint, not the paper of record. Quote page numbers from a `publishedVersion` only. ## CORE ``` POST https://api.core.ac.uk/v3/search/works Authorization: Bearer $CORE_API_KEY body: {"q":"doi:\"10.…\"","limit":1} ``` 37M+ full texts harvested from institutional and subject repositories — the best source for green OA copies that Unpaywall's publisher-centric view misses, and for theses. Requires free registration; without `CORE_API_KEY` this step is skipped. Use `downloadUrl` from a result, and verify the bytes really are a PDF. ## bioRxiv / medRxiv ``` GET https://api.biorxiv.org/details/biorxiv/10.1101/2020.01.30.927871 GET https://api.biorxiv.org/details/medrxiv/2023-01-01/2023-01-31/0 paged by date GET https://api.biorxiv.org/pubs/biorxiv/{doi} published version ``` **There is no keyword search.** These APIs only browse by DOI or by date window; the `collection` array comes back with `title`, `authors`, `date`, `category`, `jatsxml`. To search preprints by topic, use Europe PMC's `SRC:PPR` filter (see [`europepmc.md`](europepmc.md)) or OpenAlex `type:preprint` — then come back here for the version history, or to `/pubs/` to find out whether the preprint was ever published. That last check is the one people skip: citing a preprint whose published version contradicts it is a real and common failure of automated reviews. ## Verifying what you downloaded - A PDF starts with the bytes `%PDF-`. Publishers serve HTML "access denied" pages with HTTP 200 and `Content-Type: application/pdf` often enough that content-type is not evidence — check the magic bytes (`is_pdf()` in `resolve_oa.py`). - Europe PMC full text under ~2 KB is a stub record, not an article. - A DOI that resolves nowhere is two different findings. `resolve_oa.py` asks Crossref before giving up: `closed` means the paper exists behind a paywall; `not-in-crossref` means no Crossref record exists at all — an invented citation, a typo, or a DOI from another registry (DataCite datasets, some preprint servers). Never report the second as the first. - Record the resolver that produced each file. When a citation later looks wrong, the first question is always "was that the published version or the preprint?", and `oa_report.json` answers it. -
openalex.md 3.2 KB
# OpenAlex Broadest of the free indexes: ~250M works across every field, with authors, institutions, sources, topics, citation counts and OA locations. No key required. Base URL `https://api.openalex.org`. Used by `fetch_papers.py` (search), `resolve_oa.py` (step 2 of the OA chain) and `fetch_openalex.py`. ## Auth & limits - Keyless works. `?api_key=…` (`OPENALEX_API_KEY`) or the legacy polite pool `?mailto=you@example.com` (`OPENALEX_MAILTO`) raise the limit. - 100 req/s ceiling. Single-entity lookups by id/DOI are unmetered; list and search queries draw on a daily free allowance. ## Endpoints ``` GET /works/{id} W2741809807 | doi:10.1136/bmj.n160 | pmid:33781993 GET /works?search=…&per_page=25 full-text search over title+abstract+fulltext GET /works?filter=… comma-separated field:value, all ANDed GET /authors?search= | /sources?search= | /institutions?search= | /topics/{id} ``` Useful parameters: | Param | Notes | |---|---| | `search` | boolean `AND`/`OR`/`NOT` (uppercase), `"phrase"~5`, `wildcar*`, `fuzzy~1` | | `search.semantic` | embedding search, beta — 1 req/s, ≤50 results | | `filter` | `from_publication_date:2023-01-01`, `publication_year:2024`, `type:article`, `is_oa:true`, `cited_by_count:>100`, `authorships.author.id:A…`, `institutions.country_code:us`, `doi:…`. Operators `>` `<` `!` `\|` | | `sort` | `cited_by_count:desc`, `publication_date:desc`, `relevance_score:desc` | | `select` | trim the response — works on both list and single-entity endpoints | | `group_by` | aggregate counts, e.g. `group_by=publication_year` for a topic's trend line | | `cursor=*` | deep pagination past 10,000; follow `meta.next_cursor` until null | ## Response shape (fields worth reading) ```json {"id":"https://openalex.org/W2741809807","doi":"https://doi.org/10.7717/peerj.4375", "display_name":"…","publication_year":2018,"type":"article","is_retracted":false, "cited_by_count":1169, "open_access":{"is_oa":true,"oa_status":"gold","oa_url":"…"}, "best_oa_location":{"pdf_url":"…","license":"cc-by","version":"publishedVersion"}, "primary_location":{"source":{"display_name":"PeerJ","issn_l":"2167-8359"}}, "authorships":[{"author":{"id":"…","display_name":"…"},"institutions":[…]}], "abstract_inverted_index":{"Despite":[0],"growing":[1]}, "referenced_works":["https://openalex.org/W…"], "ids":{"doi":"…","pmid":"…","arxiv":"…"}} ``` Two gotchas that bite every time: - **`title` does not exist** on the work object — the field is `display_name`. - **Abstracts are inverted indexes**, `{word: [positions]}`. Rebuild by placing each word at each position and joining in index order (`_inverted()` in `fetch_papers.py`). - `is_retracted` is a free retraction check — cheaper than a dedicated service, and worth reading before you cite anything (see recipe 03). ## What it is good and bad at Good: coverage, OA links, citation counts, institution/author disambiguation, trend aggregation via `group_by`. Bad: relevance ranking is weaker than Semantic Scholar's for conceptual queries, and `open_access.oa_url` is often a **landing page, not a PDF** — always check the bytes start with `%PDF-` before saving one (`resolve_oa.py` does). -
pubmed-pmc.md 3.6 KB
# PubMed & PMC (NCBI E-utilities) 37M+ biomedical citations with MeSH indexing, plus the PMC full-text archive. The MeSH vocabulary is the reason PubMed still beats general indexes for clinical questions: it is human-curated subject indexing, not string matching. Base URL `https://eutils.ncbi.nlm.nih.gov/entrez/eutils`. Used by `fetch_papers.py` and `fetch_pubmed.py`. ## Auth & limits 3 req/s anonymous, 10 with `NCBI_API_KEY` (`&api_key=…`). Add `&email=` and `&tool=` so NCBI can contact you instead of blocking you. Sleep ~0.34 s between calls when keyless — `fetch_papers.py` does. ## The two-step dance ``` GET /esearch.fcgi?db=pubmed&term=…&retmax=20&retmode=json&sort=relevance → PMIDs GET /efetch.fcgi?db=pubmed&id=1,2,3&retmode=xml → full records GET /esummary.fcgi?db=pubmed&id=…&retmode=json → light metadata GET /elink.fcgi?dbfrom=pubmed&db=pmc&id=… → PMID → PMCID GET /efetch.fcgi?db=pmc&id=PMC8005925&retmode=xml → JATS full text ``` `esummary` returns JSON but no abstract. **`efetch` with `retmode=xml` is the only way to get abstracts**, and it returns XML even though the rest of E-utilities speaks JSON. Query syntax is the same as the PubMed website: ``` ("systematic review"[Publication Type]) AND (machine learning[Title/Abstract]) AND ("2022"[Date - Publication] : "3000"[Date - Publication]) covid-19[MeSH Terms] AND humans[MeSH Terms] AND english[Language] ``` Date filtering via parameters: `&datetype=pdat&mindate=2022&maxdate=3000`. ## Parsing the XML Fields worth pulling from each `PubmedArticle`: | XPath | Field | |---|---| | `.//PMID` | PMID | | `.//ArticleTitle` | title — **may contain inline `<i>`/`<sub>`**, so use `itertext()`, not `.text` | | `.//AbstractText` | abstract, often **several elements** with `Label="METHODS"` etc. — join them | | `./PubmedData/ArticleIdList/ArticleId[@IdType='doi']` | DOI — **scope it exactly like this** | | `.//PubDate/Year` or `.//PubDate/MedlineDate` | year (MedlineDate is free text like `2023 Jan-Feb`) | | `.//Author/ForeName` + `LastName` | authors | | `.//Journal/Title` | venue | | `.//PublicationType` | `Retracted Publication`, `Retraction of Publication` live here | Taking only `AbstractText[0]` silently truncates every structured abstract to its Background section — a classic quiet data-loss bug. The DOI path is the sharper trap. A `PubmedArticle` embeds the article's **entire reference list**, and every reference carries its own `<ArticleId IdType="doi">`. A `.//ArticleId` sweep therefore returns a *cited* paper's DOI — in one recorded run, a stroke-imaging review came back carrying an ICPSR **dataset** DOI from its own bibliography. The record looks perfect and points at the wrong object. Scope the lookup to `./PubmedData/ArticleIdList`, and fall back to `.//Article/ELocationID[@EIdType='doi']`. Written up in [`recipes/06`](../../../../recipes/06-multi-source-search/). ## ID conversion The old `/pmc/utils/idconv/v1.0/` path now 301-redirects; the live endpoint is ``` GET https://pmc.ncbi.nlm.nih.gov/tools/idconv/api/v1/articles/?ids=10.1136/bmj.n160&format=json → {"records":[{"doi":"10.1136/bmj.n160","pmcid":"PMC8005925","pmid":33781993}]} ``` Follow redirects (`curl -L`) or you will parse an HTML 301 page as JSON. ## Good and bad at Good: MeSH-indexed retrieval, publication-type filters (RCT, meta-analysis, retraction), clinical coverage, and PMC full text for the OA subset. Bad: anything non-biomedical, and PDFs — PubMed exposes almost none. For full text prefer Europe PMC (same corpus, one JSON call, no XML dance). -
README.md 5.1 KB
# Scholarly APIs — route the question to the right index The rest of this skill installs tools. This directory does the opposite: it lets you answer a literature question **with one HTTP call and nothing installed**, and it documents the exact endpoints the bundled `fetch_papers.py` / `resolve_oa.py` scripts use, so you can debug or extend them. Reach for this when the user wants *a specific paper, a DOI resolved, an author's output, an OA PDF, or a citation graph* — a venv is overkill for that. Reach for the launcher when the user wants a corpus, a screening run, or an extraction pipeline. > Layout modelled on the excellent [`paper-lookup`](https://github.com/K-Dense-AI) skill > by K-Dense Inc. The endpoints, quirks and failure modes here were re-verified against > live APIs on 2026-08-10 and annotated with what this repo's scripts actually hit. ## Pick a database | The user wants… | Query this | Then | |---|---|---| | Papers on a topic, any field | OpenAlex | Semantic Scholar for citation context | | Papers on a biomedical topic | PubMed | Europe PMC for the full text | | Full text of a biomedical paper | Europe PMC (OA subset) | Unpaywall for non-biomed | | Physics / maths / CS preprints | arXiv | Semantic Scholar for the published version | | Biology / health preprints | Europe PMC (`SRC:PPR`) | bioRxiv/medRxiv APIs *by DOI or date only* | | One paper by DOI | Crossref | OpenAlex for citations, Unpaywall for the PDF | | An open-access PDF for a DOI | Unpaywall → OpenAlex → Europe PMC | `resolve_oa.py` already chains these | | Who cites whom | Semantic Scholar | OpenAlex `referenced_works` | | An author's publications | OpenAlex | Semantic Scholar author endpoint | | Journal / funder / license metadata | Crossref | OpenAlex `sources` | | PMID ↔ PMCID ↔ DOI | NCBI ID Converter | Europe PMC search | **One query, several indexes.** Relevance ranking differs enough between these APIs that the same query returns largely disjoint top-10s — in the run recorded in [`recipes/06-multi-source-search`](../../../../recipes/06-multi-source-search/), six sources returned 50 hits of which only 3 were duplicates. If coverage matters, do not trust one index; run `fetch_papers.py` and let it merge them. ## Identifier formats | Identifier | Shape | Example | Understood by | |---|---|---|---| | DOI | `10.xxxx/…` | `10.1136/bmj.n160` | everything | | PMID | digits | `33781993` | PubMed, Europe PMC, S2 (`PMID:`) | | PMCID | `PMC` + digits | `PMC8005925` | Europe PMC, PMC | | arXiv id | `YYMM.NNNNN` | `2103.15348` | arXiv, S2 (`ARXIV:`), OpenAlex | | OpenAlex id | `W` + digits | `W2741809807` | OpenAlex | | S2 id | 40-char hex | `649def34f8be…` | Semantic Scholar | | ORCID | `0000-…-…-…` | `0000-0001-6187-6610` | OpenAlex, Crossref | Cross-lookup prefixes: OpenAlex takes `/works/doi:10.…` and `/works/pmid:…`; Semantic Scholar takes `DOI:…`, `PMID:…`, `PMCID:…`, `ARXIV:…`. A DOI that 404s in one index is usually alive in another — try before concluding it is fake. ## Keys: what is actually needed Everything below works with **no key**. Keys only buy rate limit. | API | Env var | Without it | |---|---|---| | OpenAlex | `OPENALEX_API_KEY` / `OPENALEX_MAILTO` | works; shared pool | | Crossref | `CROSSREF_MAILTO` | works at ~5 req/s; `mailto` doubles it | | Semantic Scholar | `S2_API_KEY` | **frequently HTTP 429** — the shared pool is often exhausted | | NCBI (PubMed) | `NCBI_API_KEY`, `NCBI_EMAIL` | 3 req/s instead of 10 | | Unpaywall | `UNPAYWALL_EMAIL` | **skipped entirely** — the API requires an email | | Europe PMC | — | no key exists | | arXiv | — | no key; 1 request / 3 s | | CORE | `CORE_API_KEY` | skipped; registration required | Read keys from the environment, then from `~/.lit-review-tools/.env` (`litrun.py env --set KEY=VALUE`). Never invent an email for Unpaywall — ask the user. ## Per-API references | API | File | Used by | |---|---|---| | OpenAlex | [`openalex.md`](openalex.md) | `fetch_papers.py`, `resolve_oa.py`, `fetch_openalex.py` | | Crossref | [`crossref.md`](crossref.md) | `fetch_papers.py`, recipe 03 (citation verification) | | Semantic Scholar | [`semantic-scholar.md`](semantic-scholar.md) | `fetch_papers.py` | | PubMed / PMC (NCBI) | [`pubmed-pmc.md`](pubmed-pmc.md) | `fetch_papers.py`, `fetch_pubmed.py` | | Europe PMC | [`europepmc.md`](europepmc.md) | `fetch_papers.py`, `resolve_oa.py` | | arXiv | [`arxiv.md`](arxiv.md) | `fetch_papers.py`, `fetch_arxiv.py` | | Unpaywall, CORE, bioRxiv/medRxiv | [`open-access.md`](open-access.md) | `resolve_oa.py` | ## Calling them Claude Code: `WebFetch`, or `curl` via Bash when you need the raw bytes or a POST. Always send a `User-Agent` that identifies you; several of these APIs throttle anonymous clients harder. If you get 429, wait ~3 s and retry **once**, then move on to another source rather than hammering. ## Reporting back State which APIs you queried and which returned nothing — a silent omission reads as "no such paper exists", which is a much stronger claim than "OpenAlex had no match". Quote DOIs verbatim, and if a DOI failed to resolve anywhere, say so rather than paraphrasing it into a plausible-looking citation. -
semantic-scholar.md 2.5 KB
# Semantic Scholar (S2 Graph API) ~200M papers with the best free **citation graph** — citations, references, influential citation counts, AI-generated TLDRs, and paper recommendations. Base URL `https://api.semanticscholar.org/graph/v1`. Used by `fetch_papers.py` (search). Its ranking is the most "semantic" of the free indexes, which is why it is in the default source set. ## Auth & limits — read this first The keyless pool is **shared across every anonymous client on the internet and is routinely exhausted**. In this repo's recorded runs, keyless `/paper/search` returned HTTP 429 on both attempts while single-paper lookups succeeded. Treat 429 as normal, not as a bug: retry once, then continue without this source. `S2_API_KEY` (free, request form on their site) goes in the `x-api-key` **header**, not the query string. ## Endpoints ``` GET /paper/search?query=…&limit=20&fields=… relevance search GET /paper/search/bulk?query=… up to 1000/page, no relevance ranking GET /paper/{id}?fields=… DOI:… | PMID:… | PMCID:… | ARXIV:… | CorpusId:… | 40-hex GET /paper/{id}/citations?fields=… who cites this GET /paper/{id}/references?fields=… what this cites GET /paper/{id}/recommendations GET /author/search?query=… | /author/{id}/papers POST /paper/batch {"ids":[…]} up to 500 ids in one call ``` `fields` is mandatory in practice — omit it and you get bare ids. Useful set: ``` title,abstract,year,authors,venue,citationCount,influentialCitationCount, externalIds,openAccessPdf,isOpenAccess,url,tldr,fieldsOfStudy,publicationTypes ``` `year=2020-` filters a range; `openAccessPdf` gives a direct PDF URL when one exists. ## Response shape ```json {"total":1234,"data":[{"paperId":"649def34…","title":"…","abstract":"…","year":2021, "venue":"BMJ","citationCount":9781,"influentialCitationCount":812, "externalIds":{"DOI":"10.1136/bmj.n160","PubMed":"33781993","ArXiv":null}, "openAccessPdf":{"url":"…","status":"GOLD"},"tldr":{"text":"…"}}]} ``` `abstract` is often `null` even when the paper has one — S2 cannot redistribute every publisher's text. Fall back to OpenAlex or Europe PMC for the abstract; that fallback is exactly what `fetch_papers.py`'s merge step does. ## Good and bad at Good: citation graph in both directions, `influentialCitationCount` (a better signal than raw counts for "what actually mattered"), TLDRs, recommendations, batch lookup. Bad: availability without a key, abstract coverage, non-English work.
-
-
catalog.md 15.9 KB
# Full Catalog — AI Literature Review Tools 66 open-source projects, organized by use case. ⭐ = editor's pick. > Generated by `scripts/build.py` from `data/tools.yaml`. Do not edit by hand — > edits here are overwritten and CI will reject them. Metadata refreshed from the GitHub API on **2026-09-21**. Health: 🟢 pushed within 90 days · 🟡 within a year · 🔴 over a year · 🗄️ archived. Licence `none` means the repository ships no licence file (all rights reserved); `CC-BY-NC*` forbids commercial use. Recommend accordingly. Source of truth: <https://github.com/brycewang-stanford/lit-review-agent-tools> --- ## 🌟 All-in-one Research Agents & Skills End-to-end `research → write → review → revise` solutions — mostly Claude Code / Codex skills. | Project | Stars | Health | Licence | Stage | Notes | |---|---|---|---|---|---| | ⭐ [academic-research-skills](https://github.com/Imbad0202/academic-research-skills) | ~48.9k | 🟢 | CC-BY-NC-4.0 | search, synthesize, cite-check, write, review | **The most popular project in this space.** A suite of Claude Code skills running a 10-stage pipeline (research→write→review→revise→finalize) with citation/claim "integrity gates," cross-checked against Semantic Scholar + OpenAlex + Crossref. Philosophy: *"AI is your copilot, not the pilot."* `/plugin install academic-research-skills` | | [academic-research-skills-codex](https://github.com/Imbad0202/academic-research-skills-codex) | ~11.3k | 🟢 | CC-BY-NC-4.0 | search, synthesize, cite-check, write, review | Codex-native sibling of the above, human-in-the-loop research flow | | [Research-Paper-Writing-Skills](https://github.com/Master-cai/Research-Paper-Writing-Skills) | ~7.0k | 🟢 | MIT | write | ML/CV/NLP paper-writing skill pack; works with Codex, Claude Code, and Gemini | | [claude-skills](https://github.com/alirezarezvani/claude-skills) | ~26.2k | 🟢 | MIT | search, synthesize, write | Large skill collection incl. a litreview / grants / deep-research stack across Claude Code / Codex / Gemini / Cursor | | [academic-paper-skills](https://github.com/lishix520/academic-paper-skills) | ~1.3k | 🟡 | MIT | write | Strategist (planning) + Composer (writing) skills with quality checkpoints | | [dr-claw](https://github.com/OpenLAIR/dr-claw) | ~1.1k | 🟢 | custom | search, read, synthesize, write | A "research IDE" with multiple AI-assistant personas | | [ScienceClaw](https://github.com/beita6969/ScienceClaw) | 905 | 🟡 | MIT | search, synthesize, write | Self-evolving AI research colleague, 285 skills, "zero hallucination" claim | | [qinyan-academic-skills](https://github.com/LeonChaoX/qinyan-academic-skills) | 913 | 🟢 | MIT | search, synthesize, write | Multilingual library of 182 installable AI-agent skills across disciplines | | [agent-research-skills](https://github.com/lingzhi227/agent-research-skills) | 351 | 🟡 | none | search, synthesize, cite-check | Claude Code skills for systematic literature review, incl. citation-validation scripts | | [medsci-skills](https://github.com/Aperivue/medsci-skills) | 313 | 🟢 | MIT | search, cite-check, write | Medical-research skills: search, reporting-guideline/citation checks, stats, figures, submission (by a physician-researcher) | ## 🔎 Deep Research & Auto Survey Generation Give it a topic; it searches and produces a cited report / survey / related-work section. | Project | Stars | Health | Licence | Stage | Notes | |---|---|---|---|---|---| | ⭐ [STORM](https://github.com/stanford-oval/storm) | ~31.5k | 🟡 | MIT | search, synthesize, write | Stanford OVAL; retrieval-grounded "pre-writing + writing" stages, produces Wikipedia-style long articles with citations; includes conversational Co-STORM | | ⭐ [gpt-researcher](https://github.com/assafelovic/gpt-researcher) | ~29.6k | 🟢 | Apache-2.0 | search, synthesize, write | Autonomous agent that runs deep research on any topic and outputs a cited report; general-purpose, not academic-only | | [deep-research](https://github.com/dzhng/deep-research) | ~19.7k | 🟡 | MIT | search, synthesize | Minimal iterative deep-research agent (search + scrape + LLM refinement); small and hackable | | [open_deep_research](https://github.com/langchain-ai/open_deep_research) | ~12.7k | 🗄️ | MIT | search, synthesize | LangChain's official open deep-research reference implementation | | [local-deep-research](https://github.com/LearningCircuit/local-deep-research) | ~9.1k | 🟢 | MIT | search, synthesize | Local/private deep-research; 10+ sources incl. arXiv & PubMed, fully local LLMs | | [open-deep-research](https://github.com/nickscamara/open-deep-research) | ~6.3k | 🔴 | Apache-2.0 | search, synthesize | Open deep-research clone reasoning over web data via Firecrawl | | [SurveyX](https://github.com/IAAR-Shanghai/SurveyX) | 990 | 🟡 | none | search, synthesize, write | Automated academic survey-paper generation from a topic | | [LitLLM](https://github.com/LitLLM/LitLLM) | 52 | 🟡 | Apache-2.0 | search, synthesize, write | Toolkit focused on scientific literature review; RAG + prompting to draft related-work fast (TMLR 2025) | | [opendraft](https://github.com/federicodeponte/opendraft) | 448 | 🟢 | MIT | write | Free & open-source AI paper writer; 19 agents collaborate to draft long papers | | [AutoSurveyGPT](https://github.com/a554b554/AutoSurveyGPT) | 156 | 🔴 | MIT | search, synthesize, write | Uses GPT to find & rank Google Scholar papers and auto-generate a survey | ## 🧪 Autonomous Science: idea → paper End-to-end "automated scientific discovery" — lit review + hypotheses + experiments + writing + self-review. The most ambitious category. | Project | Stars | Health | Licence | Stage | Notes | |---|---|---|---|---|---| | ⭐ [AI-Scientist](https://github.com/SakanaAI/AI-Scientist) | ~14.6k | 🟡 | custom | search, synthesize, write, review | Sakana AI; end-to-end automated discovery (lit review→experiments→writing→review); see also [v2](https://github.com/SakanaAI/AI-Scientist-v2) (~6.9k, agentic tree search, workshop-level) | | [AutoResearchClaw](https://github.com/aiming-lab/AutoResearchClaw) | ~14.5k | 🟢 | MIT | search, synthesize, cite-check, write, review | Self-evolving autonomous research: idea → conference-ready LaTeX paper (real lit from OpenAlex/S2/arXiv + sandboxed experiments + multi-agent peer review) | | [Agent-Laboratory](https://github.com/SamuelSchmidgall/AgentLaboratory) | ~5.9k | 🔴 | MIT | search, synthesize, write | End-to-end autonomous workflow: literature review → experimentation → report writing | | [SciAgentsDiscovery](https://github.com/lamm-mit/SciAgentsDiscovery) | 639 | 🔴 | Apache-2.0 | synthesize | Multi-agent (ontologist/scientist/critic) automated hypothesis & discovery system | | [Zochi](https://github.com/IntologyAI/Zochi) | 316 | 🟡 | MIT | search, synthesize, write, review | "Artificial scientist" doing end-to-end discovery to peer-reviewed publication | | [DeepInnovator](https://github.com/HKUDS/DeepInnovator) | 291 | 🟡 | MIT | synthesize | Autonomously generates research ideas, questions, testable hypotheses & experiment designs | ## 📚 Paper Q&A and RAG Grounded, **citation-backed** Q&A and extraction over a corpus of PDFs / papers. | Project | Stars | Health | Licence | Stage | Notes | |---|---|---|---|---|---| | ⭐ [paper-qa](https://github.com/Future-House/paper-qa) | ~9.2k | 🟢 | Apache-2.0 | read, extract, synthesize, cite-check | FutureHouse; high-accuracy RAG for scientific papers, answers **always cite sources**; PaperQA2 claims superhuman literature search | | [paperai](https://github.com/neuml/paperai) | ~1.8k | 🟢 | Apache-2.0 | search, read, synthesize | Semantic search + Q&A over medical & scientific papers | | [openpaper](https://github.com/khoj-ai/openpaper) | 481 | 🟢 | AGPL-3.0 | read, synthesize, cite-check | Research-library workbench: read/annotate papers + AI lit-review assistant with grounded citations | ## 🧮 Systematic Review & Screening For rigorous evidence-based / PRISMA reviews — screen thousands of abstracts efficiently. | Project | Stars | Health | Licence | Stage | Notes | |---|---|---|---|---|---| | ⭐ [ASReview](https://github.com/asreview/asreview) | ~1.0k | 🟢 | Apache-2.0 | screen | Active-learning screener for systematic reviews; interactively ranks papers to cut screening time; well-established in academia | | [LatteReview](https://github.com/PouriaRouzrokh/LatteReview) | 121 | 🟡 | CC-BY-NC-ND-4.0 | screen | Low-code Python package automating SR screening via AI agents (OpenAI/Gemini/Claude/Ollama) | | [prismAId](https://github.com/Open-and-Sustainable/prismAId) | 29 | 🟢 | AGPL-3.0 | screen, extract | Generative-AI, protocol-based systematic-review toolkit; no-code, replicable screening & extraction | | [prisma-review-tool](https://github.com/Black-Lights/prisma-review-tool) | niche | 🟢 | MIT | search, screen | PRISMA 2020 flow with AI-assisted screening via MCP (arXiv/OpenAlex/S2, no API keys) | ## 🔌 MCP Servers Bring literature capabilities into Claude / Cursor / Cline and other Model Context Protocol clients. | Project | Stars | Health | Licence | Stage | Notes | |---|---|---|---|---|---| | ⭐ [zotero-mcp](https://github.com/54yyyu/zotero-mcp) | ~5.1k | 🟢 | MIT | search, read | Connects a Zotero library (local + web API) to AI: semantic search, PDF full-text, citation analysis. The most popular Zotero MCP | | ⭐ [arxiv-mcp-server](https://github.com/blazickjp/arxiv-mcp-server) | ~3.2k | 🟢 | Apache-2.0 | search, extract, read | Search & analyze arXiv papers; downloads and converts PDFs to Markdown for LLM context; ships an `.mcpb` bundle | | [paper-search-mcp](https://github.com/openags/paper-search-mcp) | ~2.7k | 🟢 | MIT | search | Multi-source search/download across 20+ sources (arXiv, PubMed, bioRxiv, S2, OpenAlex, Crossref, CORE…) | | [PubMed-MCP-Server](https://github.com/JackKuo666/PubMed-MCP-Server) | 129 | 🔴 | MIT | search, read | Search, access, and analyze PubMed articles (metadata + deep analysis) | | [alex-mcp](https://github.com/drAbreu/alex-mcp) | 56 | 🔴 | MIT | search | OpenAlex MCP focused on author disambiguation and institution/work lookup | | [openalex-research-mcp](https://github.com/oksure/openalex-research-mcp) | 53 | 🟡 | MIT | search, cite-check | OpenAlex (240M+ works): citation analysis, research-trend tracking, collaboration networks | ## 🗂️ Reference & Knowledge Management (Zotero / Obsidian) Embed AI into the reference-manager / note-taking workflow you already use. | Project | Stars | Health | Licence | Stage | Notes | |---|---|---|---|---|---| | ⭐ [zotero-gpt](https://github.com/MuiseDestiny/zotero-gpt) | ~7.4k | 🟡 | AGPL-3.0 | read, synthesize | GPT integrated into Zotero to chat with your library | | [papersgpt-for-zotero](https://github.com/papersgpt/papersgpt-for-zotero) | ~2.6k | 🟢 | AGPL-3.0 | read, extract, synthesize | Zotero AI + MCP plugin; chat/batch-process PDFs across 30+ LLMs | | [ai-research-assistant](https://github.com/lifan0127/ai-research-assistant) | ~1.7k | 🔴 | AGPL-3.0 | read, synthesize | "Aria" — LLM-powered research assistant inside Zotero | | [paper-note-filler](https://github.com/chauff/paper-note-filler) | 47 | 🟡 | none | read | Obsidian plugin auto-creating notes from arXiv / ACL Anthology / Semantic Scholar | ## 📄 PDF → Structured Data Extraction The invisible infrastructure of lit review: turn PDFs into clean, structured Markdown / JSON for LLMs. | Project | Stars | Health | Licence | Stage | Notes | |---|---|---|---|---|---| | ⭐ [MinerU](https://github.com/opendatalab/MinerU) | ~80.4k | 🟢 | Apache-2.0 | extract | High-accuracy PDF/Office → LLM-ready Markdown/JSON (VLM+OCR, 100+ languages, formulas/tables) | | [docling](https://github.com/docling-project/docling) | ~67.5k | 🟢 | MIT | extract | IBM-origin document parser prepping PDFs/docs for gen-AI/RAG | | [marker](https://github.com/datalab-to/marker) | ~39.9k | 🟢 | Apache-2.0 | extract | Fast PDF/doc → clean Markdown/JSON conversion, scientific-doc friendly | | [PDFMathTranslate](https://github.com/PDFMathTranslate/PDFMathTranslate) | ~37.1k | 🟢 | AGPL-3.0 | read | Layout-preserving scientific-PDF translation (formulas/figures intact) | | [grobid](https://github.com/grobidOrg/grobid) | ~5.1k | 🟢 | Apache-2.0 | extract | ML tool extracting structured TEI/XML (metadata, refs, sections) from scholarly PDFs | | [paperetl](https://github.com/neuml/paperetl) | 697 | 🟡 | Apache-2.0 | extract | ETL pipeline for medical & scientific papers into structured stores | | [scipdf_parser](https://github.com/titipata/scipdf_parser) | 456 | 🔴 | MIT | extract | Python parser for scientific-publication PDFs (content + figures, GROBID-backed) | ## 🕸️ Citation Graphs & API Clients Analyze citation networks, or hit the major scholarly databases straight from code. | Project | Stars | Health | Licence | Stage | Notes | |---|---|---|---|---|---| | [scholarly](https://github.com/scholarly-python-package/scholarly) | ~1.9k | 🟡 | Unlicense | search | Pythonic Google Scholar author/publication retrieval | | [semanticscholar](https://github.com/danielnsilva/semanticscholar) | 480 | 🟢 | MIT | search, cite-check | Unofficial Python client for Semantic Scholar APIs | | [pyalex](https://github.com/J535D165/pyalex) | 413 | 🟢 | MIT | search, cite-check | Lightweight Python interface to the OpenAlex API | | [ArxivDigest](https://github.com/AutoLLM/ArxivDigest) | 467 | 🔴 | MIT | search | Personalized daily arXiv digest with GPT relevancy scoring + email pipeline | | [citegraph](https://github.com/Citegraph/citegraph) | 22 | 🟡 | MIT | search | Open web visualizer of 5M+ papers / citation networks (CS bibliography) | ## ✍️ Writing & Peer-review Assistants Draft, polish, and run an "AI pre-review" before you submit. | Project | Stars | Health | Licence | Stage | Notes | |---|---|---|---|---|---| | [lmms-lab-writer](https://github.com/EvolvingLMMs-Lab/lmms-lab-writer) | 273 | 🟡 | MIT | write | Local-first agentic LaTeX writer for AI-assisted academic writing | | [academic-writing-agents](https://github.com/andrehuang/academic-writing-agents) | 200 | 🟡 | MIT | write, review | Claude Code plugin: 10+ specialist agents for academic writing review, research, drafting, polishing | | [ai-peer-review](https://github.com/poldrack/ai-peer-review) | 154 | 🟢 | MIT | review | Multi-LLM meta-review: independent reviews synthesized into a meta-review | | [open_reviewer](https://github.com/maxidl/openreviewer) | 17 | 🔴 | none | review | Generates high-quality peer reviews of ML/AI conference papers for pre-submission feedback | | [academic-research-plugin](https://github.com/JeanDiable/academic-research-plugin) | 25 | 🟡 | MIT | search, synthesize, cite-check, review | Claude Code plugin: lit surveys, paper reviews, citation management; searches arXiv/S2/DBLP and finds research gaps | ## 📖 Awesome Lists Want the full picture? Start from these community-maintained lists. | Project | Stars | Health | Licence | Stage | Notes | |---|---|---|---|---|---| | [Awesome-LLM-Scientific-Discovery](https://github.com/HKUST-KnowComp/Awesome-LLM-Scientific-Discovery) | 439 | 🟢 | MIT | — | EMNLP 2025 survey list: LLMs in scientific discovery | | [Awesome-Auto-Research-Tools](https://github.com/handsome-rich/Awesome-Auto-Research-Tools) | ~1.2k | 🟢 | CC0-1.0 | — | Automated literature search, paper reading, experiment management, code gen | | [awesome-ai-auto-research](https://github.com/worldbench/awesome-ai-auto-research) | 522 | 🟢 | MIT | — | A survey on AI auto-research | | [LLM4SR](https://github.com/du-nlp-lab/LLM4SR) | 133 | 🔴 | MIT | — | Papers & resources on LLMs for scientific research surveys | | [awesome-ai-research-tools](https://github.com/0x11c11e/awesome-ai-research-tools) | 73 | 🟢 | CC0-1.0 | — | AI tools for lit reviews, reference management, data analysis | | [awesome-evidence-synthesis](https://github.com/evidencesynthesis-tools/awesome-evidence-synthesis) | 27 | 🟢 | CC-BY-4.0 | — | Open-source tools for systematic reviews, meta-analysis & evidence synthesis |
-
-
scripts
-
fetch_arxiv.py 3.9 KB
#!/usr/bin/env python3 """fetch_arxiv — search arXiv and download PDFs into a folder. A small bundled tool (uses the `arxiv` PyPI package) so litrun workflows can do a real retrieval step: topic -> download PDFs -> feed a downstream QA tool. No API key required. Usage: fetch_arxiv.py --query "retrieval augmented generation" --max 10 --outdir ./papers fetch_arxiv.py --query "cat:cs.CL AND graph neural network" --sort date --max 5 --outdir ./p """ import argparse import json import re import sys try: import arxiv import requests # arxiv depends on requests, so it's always present in this env except ImportError as e: sys.exit(f"fetch_arxiv: missing dependency ({e}). Reinstall with: litrun.py install arxiv-fetch") from pathlib import Path _UA = {"User-Agent": "litrun-fetch-arxiv/1.0 (+https://github.com/brycewang-stanford/lit-review-agent-tools)"} def safe_name(s, maxlen=80): s = re.sub(r"[^\w\- ]+", "", s).strip().replace(" ", "_") return s[:maxlen] or "paper" def download_pdf(url, dest): """Download a PDF via requests — independent of the arxiv package's own download API, which changes across major versions (4.0 dropped it).""" with requests.get(url, headers=_UA, stream=True, timeout=60, allow_redirects=True) as resp: resp.raise_for_status() with open(dest, "wb") as f: for chunk in resp.iter_content(chunk_size=1 << 15): if chunk: f.write(chunk) def main(): p = argparse.ArgumentParser(prog="fetch_arxiv.py") p.add_argument("--query", required=True, help="arXiv search query (supports cat:, AND/OR, etc.)") p.add_argument("--max", type=int, default=10, help="max papers to download (default 10)") p.add_argument("--outdir", required=True, help="directory to download PDFs into") p.add_argument("--sort", choices=["relevance", "date"], default="relevance") args = p.parse_args() outdir = Path(args.outdir) outdir.mkdir(parents=True, exist_ok=True) sort = (arxiv.SortCriterion.Relevance if args.sort == "relevance" else arxiv.SortCriterion.SubmittedDate) # Polite client: built-in delay + retries so we don't trip arXiv's rate limit (HTTP 429). client = arxiv.Client(page_size=min(args.max, 100), delay_seconds=3.0, num_retries=3) search = arxiv.Search(query=args.query, max_results=args.max, sort_by=sort) manifest = [] n = 0 print(f"fetch_arxiv: searching arXiv for {args.query!r} (max {args.max}, sort={args.sort})", file=sys.stderr) try: results = list(client.results(search)) except Exception as e: # network / API / rate-limit errors hint = " (arXiv rate limit — wait a minute and retry)" if "429" in str(e) else "" sys.exit(f"fetch_arxiv: search failed: {e}{hint}") for r in results: aid = r.get_short_id() title = getattr(r, "title", aid) pdf_url = getattr(r, "pdf_url", None) or f"https://arxiv.org/pdf/{aid}" fname = f"{aid}_{safe_name(title)}.pdf" try: download_pdf(pdf_url, outdir / fname) n += 1 print(f" ✓ {aid} {title[:70]}", file=sys.stderr) except Exception as e: print(f" ✗ {aid} download failed: {e}", file=sys.stderr) continue manifest.append({ "arxiv_id": aid, "title": title, "authors": [getattr(a, "name", str(a)) for a in getattr(r, "authors", [])], "published": str(getattr(r, "published", "")), "pdf": fname, "url": getattr(r, "entry_id", pdf_url), }) (outdir / "manifest.json").write_text(json.dumps(manifest, indent=2, ensure_ascii=False)) print(f"fetch_arxiv: downloaded {n}/{len(results)} PDF(s) to {outdir} " f"(manifest.json written).", file=sys.stderr) if n == 0: sys.exit("fetch_arxiv: no PDFs downloaded.") if __name__ == "__main__": main() -
fetch_openalex.py 4.5 KB
#!/usr/bin/env python3 """fetch_openalex — search OpenAlex and save papers into a folder. Downloads the open-access PDF when one is available, otherwise writes a `<id>.txt` with the title + reconstructed abstract (PaperQA2 reads text too). Writes a merged manifest.json. Uses the free OpenAlex API (no key needed for light use; set OPENALEX_API_KEY / --mailto for higher limits + the polite pool). Usage: fetch_openalex.py --query "retrieval augmented generation" --max 10 --outdir ./papers """ import argparse import json import os import re import sys from pathlib import Path try: import requests except ImportError: sys.exit("fetch_openalex: 'requests' not installed. Run: litrun.py install openalex-fetch") API = "https://api.openalex.org/works" _UA = {"User-Agent": "litrun-fetch-openalex/1.0 (+https://github.com/brycewang-stanford/lit-review-agent-tools)"} def safe_name(s, maxlen=80): s = re.sub(r"[^\w\- ]+", "", s).strip().replace(" ", "_") return s[:maxlen] or "work" def reconstruct_abstract(inv): """OpenAlex stores abstracts as an inverted index {word: [positions]}.""" if not inv: return "" pos = {} for word, positions in inv.items(): for p in positions: pos[p] = word return " ".join(pos[i] for i in sorted(pos)) def download(url, dest): with requests.get(url, headers=_UA, stream=True, timeout=60, allow_redirects=True) as r: r.raise_for_status() with open(dest, "wb") as f: for chunk in r.iter_content(1 << 15): if chunk: f.write(chunk) def main(): p = argparse.ArgumentParser(prog="fetch_openalex.py") p.add_argument("--query", required=True) p.add_argument("--max", type=int, default=10) p.add_argument("--outdir", required=True) p.add_argument("--mailto", default=os.environ.get("OPENALEX_MAILTO", "")) args = p.parse_args() outdir = Path(args.outdir) outdir.mkdir(parents=True, exist_ok=True) params = {"search": args.query, "per_page": min(args.max, 200)} if args.mailto: params["mailto"] = args.mailto key = os.environ.get("OPENALEX_API_KEY") if key: params["api_key"] = key print(f"fetch_openalex: searching OpenAlex for {args.query!r} (max {args.max})", file=sys.stderr) try: resp = requests.get(API, params=params, headers=_UA, timeout=60) resp.raise_for_status() works = resp.json().get("results", [])[:args.max] except Exception as e: sys.exit(f"fetch_openalex: search failed: {e}") manifest, saved = [], 0 for w in works: oid = (w.get("id") or "").rsplit("/", 1)[-1] or "work" title = w.get("display_name") or w.get("title") or oid oa = w.get("best_oa_location") or w.get("primary_location") or {} pdf_url = oa.get("pdf_url") or (w.get("open_access") or {}).get("oa_url") kind = None if pdf_url: try: dest = outdir / f"{oid}_{safe_name(title)}.pdf" download(pdf_url, dest) kind = "pdf" except Exception as e: print(f" ! {oid} PDF failed ({e}); saving abstract instead", file=sys.stderr) if kind is None: abstract = reconstruct_abstract(w.get("abstract_inverted_index")) dest = outdir / f"{oid}_{safe_name(title)}.txt" dest.write_text(f"{title}\n\n{abstract}\n", encoding="utf-8") kind = "txt" saved += 1 print(f" ✓ {oid} [{kind}] {title[:65]}", file=sys.stderr) manifest.append({ "openalex_id": oid, "title": title, "year": w.get("publication_year"), "authors": [a.get("author", {}).get("display_name") for a in w.get("authorships", [])], "file": dest.name, "kind": kind, "url": w.get("id"), }) _merge_manifest(outdir, manifest) print(f"fetch_openalex: saved {saved}/{len(works)} item(s) to {outdir}.", file=sys.stderr) if saved == 0: sys.exit("fetch_openalex: nothing saved.") def _merge_manifest(outdir, entries): """Append to an existing manifest.json so multi-source workflows accumulate.""" path = outdir / "manifest.json" existing = [] if path.exists(): try: existing = json.loads(path.read_text()) except Exception: existing = [] existing.extend(entries) path.write_text(json.dumps(existing, indent=2, ensure_ascii=False), encoding="utf-8") if __name__ == "__main__": main() -
fetch_papers.py 20.9 KB
#!/usr/bin/env python3 """fetch_papers — one keyless search across six scholarly APIs, deduplicated. Queries any combination of OpenAlex, Crossref, Semantic Scholar, PubMed, Europe PMC and arXiv, merges the hits into one normalised record list (deduplicated by DOI, then by normalised title), and writes: <outdir>/results.json normalised records + per-source errors <outdir>/<key>.txt title + abstract per record (--save abstracts, default) <outdir>/manifest.json appended, same shape the other bundled fetchers write Standard library only — no pip install, no API key. Keys are used if present (S2_API_KEY, NCBI_API_KEY, OPENALEX_API_KEY/OPENALEX_MAILTO, CROSSREF_MAILTO) and simply raise rate limits. Usage: fetch_papers.py --query "retrieval augmented generation" --max 20 --outdir ./papers fetch_papers.py --query "CRISPR off-target" --sources pubmed,europepmc --from-year 2022 \ --outdir ./papers --pdfs """ import argparse import json import os import re import sys import time import urllib.error import urllib.parse import urllib.request import xml.etree.ElementTree as ET from pathlib import Path UA = "litrun-fetch-papers/1.0 (+https://github.com/brycewang-stanford/lit-review-agent-tools)" SOURCES = ["openalex", "crossref", "semanticscholar", "pubmed", "europepmc", "arxiv"] # ------------------------------------------------------------------ http def _get(url, params=None, headers=None, timeout=60, retries=1): """GET a URL, returning raw bytes. Retries once on 429/5xx.""" if params: url = f"{url}?{urllib.parse.urlencode(params, doseq=True)}" hdrs = {"User-Agent": UA, "Accept": "*/*"} hdrs.update(headers or {}) for attempt in range(retries + 1): try: req = urllib.request.Request(url, headers=hdrs) with urllib.request.urlopen(req, timeout=timeout) as r: return r.read() except urllib.error.HTTPError as e: if e.code in (429, 500, 502, 503, 504) and attempt < retries: time.sleep(3) continue raise RuntimeError(f"HTTP {e.code} for {url.split('?')[0]}") from e except Exception as e: # timeouts, DNS, TLS if attempt < retries: time.sleep(2) continue raise RuntimeError(f"{type(e).__name__}: {e}") from e def _get_json(url, params=None, headers=None, timeout=60): return json.loads(_get(url, params, headers, timeout).decode("utf-8", "replace")) # ------------------------------------------------------------------ shaping def norm_doi(doi): if not doi: return "" doi = doi.strip().lower() doi = re.sub(r"^https?://(dx\.)?doi\.org/", "", doi) return doi.rstrip(".") def norm_title(title): return re.sub(r"[^a-z0-9]+", "", (title or "").lower())[:120] def strip_tags(s): return re.sub(r"\s+", " ", re.sub(r"<[^>]+>", " ", s or "")).strip() def record(source, **kw): """Every source funnels into this one shape.""" r = { "source": source, "sources": [source], "id": "", "doi": "", "title": "", "abstract": "", "year": None, "authors": [], "venue": "", "cited_by": None, "is_oa": None, "pdf_url": "", "url": "", } r.update({k: v for k, v in kw.items() if v not in (None, "", [])}) r["doi"] = norm_doi(r["doi"]) return r # ------------------------------------------------------------------ sources def s_openalex(query, limit, from_year): params = {"search": query, "per_page": min(limit, 200)} filters = [] if from_year: filters.append(f"from_publication_date:{from_year}-01-01") if filters: params["filter"] = ",".join(filters) if os.environ.get("OPENALEX_API_KEY"): params["api_key"] = os.environ["OPENALEX_API_KEY"] elif os.environ.get("OPENALEX_MAILTO"): params["mailto"] = os.environ["OPENALEX_MAILTO"] data = _get_json("https://api.openalex.org/works", params) out = [] for w in data.get("results", [])[:limit]: loc = w.get("best_oa_location") or w.get("primary_location") or {} out.append(record( "openalex", id=(w.get("id") or "").rsplit("/", 1)[-1], doi=w.get("doi") or "", title=w.get("display_name") or "", abstract=_inverted(w.get("abstract_inverted_index")), year=w.get("publication_year"), authors=[a.get("author", {}).get("display_name") for a in w.get("authorships", [])[:20]], venue=((w.get("primary_location") or {}).get("source") or {}).get("display_name") or "", cited_by=w.get("cited_by_count"), is_oa=(w.get("open_access") or {}).get("is_oa"), pdf_url=loc.get("pdf_url") or "", url=w.get("id") or "", )) return out def _inverted(inv): """OpenAlex ships abstracts as {word: [positions]}.""" if not inv: return "" pos = {} for word, idxs in inv.items(): for i in idxs: pos[i] = word return " ".join(pos[i] for i in sorted(pos)) def s_crossref(query, limit, from_year): params = {"query.bibliographic": query, "rows": min(limit, 100)} if from_year: params["filter"] = f"from-pub-date:{from_year}-01-01" mail = os.environ.get("CROSSREF_MAILTO") or os.environ.get("UNPAYWALL_EMAIL") if mail: params["mailto"] = mail # polite pool: double the rate limit data = _get_json("https://api.crossref.org/works", params) out = [] for w in data.get("message", {}).get("items", [])[:limit]: parts = (w.get("issued") or {}).get("date-parts") or [[None]] pdf = "" for link in w.get("link") or []: if link.get("content-type") == "application/pdf": pdf = link.get("URL", "") break out.append(record( "crossref", id=w.get("DOI", ""), doi=w.get("DOI", ""), title=(w.get("title") or [""])[0], abstract=strip_tags(w.get("abstract")), year=parts[0][0] if parts and parts[0] else None, authors=[" ".join(filter(None, [a.get("given"), a.get("family")])) for a in (w.get("author") or [])[:20]], venue=(w.get("container-title") or [""])[0], cited_by=w.get("is-referenced-by-count"), pdf_url=pdf, url=w.get("URL", ""), )) return out def s_semanticscholar(query, limit, from_year): fields = "title,abstract,year,authors,venue,citationCount,externalIds,openAccessPdf,url,isOpenAccess" params = {"query": query, "limit": min(limit, 100), "fields": fields} if from_year: params["year"] = f"{from_year}-" headers = {} if os.environ.get("S2_API_KEY"): headers["x-api-key"] = os.environ["S2_API_KEY"] data = _get_json("https://api.semanticscholar.org/graph/v1/paper/search", params, headers) out = [] for w in data.get("data", [])[:limit]: ext = w.get("externalIds") or {} out.append(record( "semanticscholar", id=w.get("paperId", ""), doi=ext.get("DOI", ""), title=w.get("title") or "", abstract=w.get("abstract") or "", year=w.get("year"), authors=[a.get("name") for a in (w.get("authors") or [])[:20]], venue=w.get("venue") or "", cited_by=w.get("citationCount"), is_oa=w.get("isOpenAccess"), pdf_url=(w.get("openAccessPdf") or {}).get("url", ""), url=w.get("url") or "", )) return out def s_pubmed(query, limit, from_year): base = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils" common = {} if os.environ.get("NCBI_API_KEY"): common["api_key"] = os.environ["NCBI_API_KEY"] if os.environ.get("NCBI_EMAIL"): common["email"] = os.environ["NCBI_EMAIL"] params = dict(common, db="pubmed", term=query, retmax=min(limit, 100), retmode="json", sort="relevance") if from_year: params["mindate"], params["maxdate"], params["datetype"] = f"{from_year}", "3000", "pdat" ids = _get_json(f"{base}/esearch.fcgi", params).get("esearchresult", {}).get("idlist", []) if not ids: return [] time.sleep(0.34) # NCBI: 3 req/s without a key xml = _get(f"{base}/efetch.fcgi", dict(common, db="pubmed", id=",".join(ids), retmode="xml")) root = ET.fromstring(xml) out = [] for art in root.findall(".//PubmedArticle"): pmid = art.findtext(".//PMID") or "" # Scope the DOI hunt to the article's own id list. A PubmedArticle embeds # its whole reference list, each entry carrying its own <ArticleId # IdType="doi">, so a `.//ArticleId` sweep silently returns a *cited* # paper's DOI — see recipes/06 for the record this corrupted. doi = "" for aid in art.findall("./PubmedData/ArticleIdList/ArticleId"): if aid.get("IdType") == "doi": doi = aid.text or "" break if not doi: for el in art.findall(".//Article/ELocationID"): if el.get("EIdType") == "doi": doi = el.text or "" break abstract = " ".join(" ".join(t.itertext()) for t in art.findall(".//AbstractText")).strip() year = art.findtext(".//PubDate/Year") or art.findtext(".//PubDate/MedlineDate") or "" title_el = art.find(".//ArticleTitle") # may carry inline <i>/<sub> markup out.append(record( "pubmed", id=pmid, doi=doi, title=re.sub(r"\s+", " ", "".join(title_el.itertext())).strip() if title_el is not None else "", abstract=abstract, year=int(year[:4]) if year[:4].isdigit() else None, authors=[" ".join(filter(None, [a.findtext("ForeName"), a.findtext("LastName")])) for a in art.findall(".//Author")[:20]], venue=art.findtext(".//Journal/Title") or "", url=f"https://pubmed.ncbi.nlm.nih.gov/{pmid}/" if pmid else "", )) return out def s_europepmc(query, limit, from_year): q = query if from_year: q = f"({query}) AND (FIRST_PDATE:[{from_year}-01-01 TO 3000-01-01])" params = {"query": q, "format": "json", "pageSize": min(limit, 100), "resultType": "core"} data = _get_json("https://www.ebi.ac.uk/europepmc/webservices/rest/search", params) out = [] for w in data.get("resultList", {}).get("result", [])[:limit]: pdf = "" for u in ((w.get("fullTextUrlList") or {}).get("fullTextUrl") or []): if u.get("documentStyle") == "pdf": pdf = u.get("url", "") break out.append(record( "europepmc", id=w.get("id", ""), doi=w.get("doi", ""), title=w.get("title") or "", abstract=strip_tags(w.get("abstractText")), year=int(w["pubYear"]) if str(w.get("pubYear", "")).isdigit() else None, authors=[a.get("fullName") for a in ((w.get("authorList") or {}).get("author") or [])[:20]], venue=(w.get("journalInfo") or {}).get("journal", {}).get("title") or w.get("bookOrReportDetails", {}).get("publisher", ""), cited_by=w.get("citedByCount"), is_oa=w.get("isOpenAccess") == "Y", pdf_url=pdf, url=f"https://europepmc.org/article/{w.get('source','MED')}/{w.get('id','')}", )) return out def s_arxiv(query, limit, from_year): # arXiv has no server-side date filter on the legacy API; sort by relevance # and drop anything older than --from-year client-side. params = {"search_query": f"all:{query}", "start": 0, "max_results": min(limit * 2 if from_year else limit, 100), "sortBy": "relevance"} xml = _get("http://export.arxiv.org/api/query", params) ns = {"a": "http://www.w3.org/2005/Atom", "arxiv": "http://arxiv.org/schemas/atom"} out = [] for e in ET.fromstring(xml).findall("a:entry", ns): published = e.findtext("a:published", "", ns) year = int(published[:4]) if published[:4].isdigit() else None if from_year and year and year < int(from_year): continue aid = (e.findtext("a:id", "", ns) or "").rsplit("/", 1)[-1] out.append(record( "arxiv", id=aid, doi=e.findtext("arxiv:doi", "", ns), title=re.sub(r"\s+", " ", e.findtext("a:title", "", ns)).strip(), abstract=re.sub(r"\s+", " ", e.findtext("a:summary", "", ns)).strip(), year=year, authors=[a.findtext("a:name", "", ns) for a in e.findall("a:author", ns)[:20]], venue=e.findtext("arxiv:journal_ref", "", ns) or "arXiv", is_oa=True, pdf_url=f"https://arxiv.org/pdf/{aid}" if aid else "", url=e.findtext("a:id", "", ns), )) if len(out) >= limit: break return out FETCHERS = { "openalex": s_openalex, "crossref": s_crossref, "semanticscholar": s_semanticscholar, "pubmed": s_pubmed, "europepmc": s_europepmc, "arxiv": s_arxiv, } # ------------------------------------------------------------------ merge def merge(records): """Deduplicate on DOI first, then on normalised title. First hit wins; later hits only fill in fields the winner left empty.""" by_key, order = {}, [] for r in records: key = r["doi"] or f"t:{norm_title(r['title'])}" if not key or key == "t:": continue if key not in by_key: by_key[key] = r order.append(key) continue cur = by_key[key] if r["source"] not in cur["sources"]: cur["sources"].append(r["source"]) for f in ("doi", "abstract", "title", "venue", "pdf_url", "url"): if not cur.get(f) and r.get(f): cur[f] = r[f] for f in ("year", "cited_by", "is_oa"): if cur.get(f) in (None, "") and r.get(f) is not None: cur[f] = r[f] if len(r.get("authors") or []) > len(cur.get("authors") or []): cur["authors"] = r["authors"] return [by_key[k] for k in order] def collapse_titles(records): """Second pass: fold records that share a title but carry different DOIs — i.e. a preprint and its published version. Keeps the one with the DOI that looks published (non-preprint prefix), else the first seen.""" PREPRINT = ("10.48550", "10.1101", "10.21203", "10.31234", "10.31219", "10.26434") by_title, order = {}, [] for r in records: k = norm_title(r["title"]) if not k: order.append(id(r)) by_title[id(r)] = r continue if k not in by_title: by_title[k] = r order.append(k) continue cur = by_title[k] def rank(rec): # published DOI > preprint DOI > no DOI at all if not rec["doi"]: return 0 return 1 if rec["doi"].startswith(PREPRINT) else 2 keep, drop = cur, r if rank(r) > rank(cur): keep, drop = r, cur by_title[k] = r for s in drop["sources"]: if s not in keep["sources"]: keep["sources"].append(s) keep.setdefault("also_doi", []) if drop["doi"] and drop["doi"] != keep["doi"]: keep["also_doi"].append(drop["doi"]) if not keep["abstract"] and drop["abstract"]: keep["abstract"] = drop["abstract"] if not keep["pdf_url"] and drop["pdf_url"]: keep["pdf_url"] = drop["pdf_url"] return [by_title[k] for k in order] def safe_name(s, maxlen=70): return (re.sub(r"[^\w\- ]+", "", s or "").strip().replace(" ", "_")[:maxlen]) or "paper" def key_for(r): return safe_name(f"{r['source']}_{(r['doi'] or r['id']).replace('/', '_')}_{r['title']}", 90) def download(url, dest): data = _get(url, timeout=90) if not data: raise RuntimeError("empty response") dest.write_bytes(data) def merge_manifest(outdir, entries): path = outdir / "manifest.json" existing = [] if path.exists(): try: existing = json.loads(path.read_text()) except Exception: existing = [] existing.extend(entries) path.write_text(json.dumps(existing, indent=2, ensure_ascii=False), encoding="utf-8") def main(): p = argparse.ArgumentParser(prog="fetch_papers.py", description=__doc__.split("\n")[0], formatter_class=argparse.RawDescriptionHelpFormatter) p.add_argument("--query", required=True) p.add_argument("--sources", default="openalex,crossref,semanticscholar", help=f"comma-separated subset of: {', '.join(SOURCES)} (default: openalex,crossref,semanticscholar)") p.add_argument("--max", type=int, default=20, help="results per source before dedup") p.add_argument("--from-year", dest="from_year", help="drop anything published before this year") p.add_argument("--outdir", required=True) p.add_argument("--save", choices=["abstracts", "none"], default="abstracts", help="write one .txt per record for PaperQA2/grep (default) or metadata only") p.add_argument("--pdfs", action="store_true", help="also download any PDF the search already exposed") p.add_argument("--open-access-only", action="store_true", dest="oa_only") p.add_argument("--dedup-titles", action="store_true", dest="dedup_titles", help="also fold same-title records with different DOIs (preprint + published)") args = p.parse_args() picked = [s.strip() for s in args.sources.split(",") if s.strip()] unknown = [s for s in picked if s not in FETCHERS] if unknown: sys.exit(f"fetch_papers: unknown source(s): {', '.join(unknown)}. Known: {', '.join(SOURCES)}") outdir = Path(args.outdir) outdir.mkdir(parents=True, exist_ok=True) raw, errors, per_source = [], {}, {} for s in picked: print(f"→ {s}: searching {args.query!r} (max {args.max})", file=sys.stderr) try: hits = FETCHERS[s](args.query, args.max, args.from_year) per_source[s] = len(hits) raw.extend(hits) print(f" ✓ {s}: {len(hits)} hit(s)", file=sys.stderr) except Exception as e: errors[s] = str(e) per_source[s] = 0 print(f" ! {s} failed: {e} — continuing with the other sources", file=sys.stderr) time.sleep(0.5) merged = merge(raw) if args.dedup_titles: before = len(merged) merged = collapse_titles(merged) print(f" · title pass folded {before - len(merged)} preprint/published pair(s)", file=sys.stderr) if args.oa_only: merged = [r for r in merged if r.get("is_oa") or r.get("pdf_url")] manifest = [] for r in merged: r["key"] = key_for(r) fname = "" if args.pdfs and r.get("pdf_url"): try: dest = outdir / f"{r['key']}.pdf" download(r["pdf_url"], dest) fname, r["file_kind"] = dest.name, "pdf" except Exception as e: print(f" ! PDF failed for {r['key'][:40]}: {e}", file=sys.stderr) if not fname and args.save == "abstracts": dest = outdir / f"{r['key']}.txt" body = [r["title"], ""] if r["authors"]: body.append(", ".join(a for a in r["authors"] if a)) if r["venue"] or r["year"]: body.append(f"{r['venue']} ({r['year']})".strip()) if r["doi"]: body.append(f"doi:{r['doi']}") body += ["", r["abstract"] or "(no abstract available from this source)"] dest.write_text("\n".join(body) + "\n", encoding="utf-8") fname, r["file_kind"] = dest.name, "txt" r["file"] = fname if fname: manifest.append({ "title": r["title"], "year": r["year"], "authors": r["authors"], "doi": r["doi"], "file": fname, "kind": r.get("file_kind"), "url": r["url"], "sources": r["sources"], }) if manifest: merge_manifest(outdir, manifest) results = { "query": args.query, "sources": picked, "from_year": args.from_year, "raw_hits": len(raw), "unique": len(merged), "per_source": per_source, "errors": errors, "results": merged, } (outdir / "results.json").write_text(json.dumps(results, indent=2, ensure_ascii=False), encoding="utf-8") dupes = len(raw) - len(merged) multi = sum(1 for r in merged if len(r["sources"]) > 1) with_doi = sum(1 for r in merged if r["doi"]) print(f"\nfetch_papers: {len(raw)} raw → {len(merged)} unique " f"({dupes} duplicate(s) merged, {multi} found by >1 source, {with_doi} with a DOI)", file=sys.stderr) print(f" results.json + {len(manifest)} file(s) in {outdir}", file=sys.stderr) if errors: print(f" failed source(s): {', '.join(errors)}", file=sys.stderr) if not merged: sys.exit("fetch_papers: no results — try fewer/other --sources or a broader query.") if __name__ == "__main__": main() -
fetch_pubmed.py 4.2 KB
#!/usr/bin/env python3 """fetch_pubmed — search PubMed and save each article's title + abstract as text. PubMed rarely exposes downloadable PDFs (full text lives behind publishers / PMC), so this saves `<pmid>.txt` files (title + abstract) that PaperQA2 can ingest, plus a merged manifest.json. Uses NCBI E-utilities (no key needed for light use; set NCBI_API_KEY / --email to raise the rate limit). Usage: fetch_pubmed.py --query "CRISPR off-target detection" --max 10 --outdir ./papers """ import argparse import json import os import re import sys import xml.etree.ElementTree as ET from pathlib import Path try: import requests except ImportError: sys.exit("fetch_pubmed: 'requests' not installed. Run: litrun.py install pubmed-fetch") EUTILS = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils" _UA = {"User-Agent": "litrun-fetch-pubmed/1.0 (+https://github.com/brycewang-stanford/lit-review-agent-tools)"} def safe_name(s, maxlen=80): s = re.sub(r"[^\w\- ]+", "", s).strip().replace(" ", "_") return s[:maxlen] or "article" def _common_params(): p = {} if os.environ.get("NCBI_API_KEY"): p["api_key"] = os.environ["NCBI_API_KEY"] return p def main(): p = argparse.ArgumentParser(prog="fetch_pubmed.py") p.add_argument("--query", required=True) p.add_argument("--max", type=int, default=10) p.add_argument("--outdir", required=True) p.add_argument("--email", default=os.environ.get("NCBI_EMAIL", "")) args = p.parse_args() outdir = Path(args.outdir) outdir.mkdir(parents=True, exist_ok=True) common = _common_params() if args.email: common["email"] = args.email print(f"fetch_pubmed: searching PubMed for {args.query!r} (max {args.max})", file=sys.stderr) try: es = requests.get(f"{EUTILS}/esearch.fcgi", headers=_UA, timeout=60, params={ "db": "pubmed", "term": args.query, "retmax": args.max, "retmode": "json", **common}) es.raise_for_status() ids = es.json().get("esearchresult", {}).get("idlist", []) except Exception as e: sys.exit(f"fetch_pubmed: search failed: {e}") if not ids: sys.exit("fetch_pubmed: no results for that query.") try: ef = requests.get(f"{EUTILS}/efetch.fcgi", headers=_UA, timeout=90, params={ "db": "pubmed", "id": ",".join(ids), "rettype": "abstract", "retmode": "xml", **common}) ef.raise_for_status() root = ET.fromstring(ef.content) except Exception as e: sys.exit(f"fetch_pubmed: fetch failed: {e}") manifest, saved = [], 0 for art in root.findall(".//PubmedArticle"): pmid = (art.findtext(".//PMID") or "").strip() or "unknown" title = (art.findtext(".//ArticleTitle") or "").strip() or pmid # AbstractText may be split into labeled sections. parts = [] for node in art.findall(".//Abstract/AbstractText"): label = node.get("Label") text = "".join(node.itertext()).strip() parts.append(f"{label}: {text}" if label else text) abstract = "\n".join(parts) authors = [] for a in art.findall(".//Author"): ln, fn = a.findtext("LastName"), a.findtext("ForeName") if ln: authors.append(f"{fn} {ln}".strip() if fn else ln) dest = outdir / f"pmid{pmid}_{safe_name(title)}.txt" dest.write_text(f"{title}\n\n{abstract}\n", encoding="utf-8") saved += 1 print(f" ✓ {pmid} {title[:65]}", file=sys.stderr) manifest.append({ "pmid": pmid, "title": title, "authors": authors, "file": dest.name, "kind": "txt", "url": f"https://pubmed.ncbi.nlm.nih.gov/{pmid}/", }) _merge_manifest(outdir, manifest) print(f"fetch_pubmed: saved {saved} abstract(s) to {outdir}.", file=sys.stderr) if saved == 0: sys.exit("fetch_pubmed: nothing saved.") def _merge_manifest(outdir, entries): path = outdir / "manifest.json" existing = [] if path.exists(): try: existing = json.loads(path.read_text()) except Exception: existing = [] existing.extend(entries) path.write_text(json.dumps(existing, indent=2, ensure_ascii=False), encoding="utf-8") if __name__ == "__main__": main() -
litrun.py 21.3 KB
#!/usr/bin/env python3 """litrun — install and run open-source AI literature-review tools by id. A thin, dependency-free launcher over recipes.json. Each tool gets its own isolated virtualenv under ~/.lit-review-tools/envs/<id> so installs never collide. API keys live in one shared ~/.lit-review-tools/.env. Usage: litrun.py list [--category C] [--kind K] litrun.py info <id> litrun.py doctor litrun.py env [--set KEY=VALUE ...] # show / edit the shared .env litrun.py install <id> litrun.py run <id> [-- <tool args...>] litrun.py mcp <id> [--storage PATH] [--client claude|cursor] litrun.py ui <id> # clone & launch a web UI (gpt-researcher, storm) litrun.py workflow list litrun.py workflow run <id> --input PATH [--question "..."] [--dry-run] Design notes for the calling agent: - python-cli tools (mineru, marker, docling, paper-qa, asreview): install then run fully automatically. - python-lib tools (gpt-researcher, storm, scholarly, pyalex): install puts the library in a venv; `run` executes the recipe's example snippet. - mcp-server tools: `install` prepares them, `mcp` prints the client config block to register (they are servers, not one-shot CLIs). """ import argparse import json import os import shutil import subprocess import sys from pathlib import Path BASE = Path(os.environ.get("LITRUN_HOME", Path.home() / ".lit-review-tools")) ENVS = BASE / "envs" WORKSPACE = BASE / "workspace" ENV_FILE = BASE / ".env" RECIPES = Path(__file__).resolve().parent.parent / "recipes" / "recipes.json" WORKFLOWS = Path(__file__).resolve().parent.parent / "recipes" / "workflows.json" SCRIPTS_DIR = Path(__file__).resolve().parent def die(msg, code=1): print(f"litrun: {msg}", file=sys.stderr) sys.exit(code) def load_recipes(): if not RECIPES.exists(): die(f"recipes.json not found at {RECIPES}") data = json.loads(RECIPES.read_text()) return {t["id"]: t for t in data["tools"]} def load_workflows(): if not WORKFLOWS.exists(): die(f"workflows.json not found at {WORKFLOWS}") data = json.loads(WORKFLOWS.read_text()) return {w["id"]: w for w in data["workflows"]} def get_tool(tools, tid): if tid not in tools: matches = [k for k in tools if tid in k] hint = f" Did you mean: {', '.join(matches)}?" if matches else "" die(f"unknown tool id '{tid}'.{hint} Run `litrun.py list`.") return tools[tid] def have(cmd): return shutil.which(cmd) is not None def read_env_file(): env = {} if ENV_FILE.exists(): for line in ENV_FILE.read_text().splitlines(): line = line.strip() if not line or line.startswith("#") or "=" not in line: continue k, v = line.split("=", 1) env[k.strip()] = v.strip() return env def write_env_file(env): BASE.mkdir(parents=True, exist_ok=True) lines = ["# litrun shared environment (API keys). One KEY=VALUE per line.\n"] for k, v in sorted(env.items()): lines.append(f"{k}={v}\n") ENV_FILE.write_text("".join(lines)) def merged_env(): """OS environ overlaid with the shared .env (OS wins if already set).""" e = dict(os.environ) for k, v in read_env_file().items(): e.setdefault(k, v) return e def venv_dir(tid): return ENVS / tid def venv_bin(tid, name): d = venv_dir(tid) binpath = d / ("Scripts" if os.name == "nt" else "bin") return binpath / name def venv_python(tid): return venv_bin(tid, "python.exe" if os.name == "nt" else "python") def ensure_venv(tid): d = venv_dir(tid) if venv_python(tid).exists(): return d.parent.mkdir(parents=True, exist_ok=True) if have("uv"): run_cmd(["uv", "venv", str(d)]) else: run_cmd([sys.executable, "-m", "venv", str(d)]) def pip_install(tid, packages): ensure_venv(tid) if have("uv"): cmd = ["uv", "pip", "install", "--python", str(venv_python(tid)), "-U", *packages] else: cmd = [str(venv_python(tid)), "-m", "pip", "install", "-U", *packages] run_cmd(cmd) def run_cmd(cmd, env=None, check=True, cwd=None): printable = " ".join(str(c) for c in cmd) loc = f"(cd {cwd}) " if cwd else "" print(f"$ {loc}{printable}", file=sys.stderr) return subprocess.run(cmd, env=env, check=check, cwd=cwd) def exec_cli(t, tool_args, env, cwd=None, dry=False): """Run a python-cli entrypoint or a bundled python-script. Shared by run & workflow.""" if t["kind"] == "python-cli": probe = venv_bin(t["id"], t["entry"]) cmd = [str(probe), *tool_args] display = t["entry"] elif t["kind"] == "python-script": # Stdlib-only scripts (no pip list) run on the ambient interpreter — # no venv, no install, nothing to go stale. probe = Path(sys.executable) if not t.get("pip") else venv_python(t["id"]) cmd = [str(probe), str(SCRIPTS_DIR / t["script"]), *tool_args] display = t["script"] else: die(f"'{t['id']}' is {t['kind']} — can't exec it directly.") if dry: loc = f"(cd {cwd}) " if cwd else "" print(f"DRY $ {loc}{display} {' '.join(tool_args)}") return if not probe.exists(): if t.get("pip"): print(f"{t['id']} not installed — installing first.") pip_install(t["id"], t["pip"]) else: die(f"'{display}' not found; try: litrun.py install {t['id']}") run_cmd(cmd, env=env, check=False, cwd=cwd) # ---------------------------------------------------------------- commands def cmd_list(tools, args): rows = [] for t in tools.values(): if args.category and t["category"] != args.category: continue if args.kind and t["kind"] != args.kind: continue no_install_needed = not t.get("pip") # stdlib scripts and uvx MCP servers installed = "✓" if venv_python(t["id"]).exists() or no_install_needed else " " rows.append((installed, t["id"], t["kind"], t["category"], t["name"])) if not rows: print("No tools match that filter.") return w_id = max(len(r[1]) for r in rows) w_kind = max(len(r[2]) for r in rows) w_cat = max(len(r[3]) for r in rows) print(f" {'id'.ljust(w_id)} {'kind'.ljust(w_kind)} {'category'.ljust(w_cat)} name") for inst, tid, kind, cat, name in rows: print(f"{inst} {tid.ljust(w_id)} {kind.ljust(w_kind)} {cat.ljust(w_cat)} {name}") print("\n(✓ = installed / no install needed). litrun.py info <id> for details.") def cmd_info(tools, args): t = get_tool(tools, args.id) print(f"# {t['name']} ({t['id']})") print(f"repo: {t['repo']}") print(f"category: {t['category']} kind: {t['kind']}") if t.get("pip"): print(f"install: litrun.py install {t['id']} (pip: {', '.join(t['pip'])})") if t.get("entry"): print(f"entry: {t['entry']}") if t.get("script"): where = "runs via this tool's venv" if t.get("pip") else "stdlib only — no install, runs on python3" print(f"script: {t['script']} (bundled; {where})") if t.get("example"): print(f"example: {t['example']}") req = t.get("env", []) opt = t.get("env_optional", []) if req or opt: env = merged_env() if req: print("required env:") for k in req: print(f" {'✓' if env.get(k) else '✗'} {k}") if opt: print("optional env: " + ", ".join(opt)) if t["kind"] == "mcp-server": print(f"mcp: litrun.py mcp {t['id']} (prints client config)") if t.get("ui"): print(f"ui: litrun.py ui {t['id']} (clone & launch web UI at {t['ui']['url']})") print(f"\nnotes: {t['notes']}") def cmd_doctor(tools, args): print("== toolchain ==") for c in ("python3", "uv", "git", "pip"): print(f" {'✓' if have(c) else '✗'} {c}") if not have("uv"): print(" (uv not found — falling back to python -m venv + pip. Install uv for speed.)") print(f"\n== base dir ==\n {BASE} ({'exists' if BASE.exists() else 'will be created'})") print(f" env file: {ENV_FILE} ({'present' if ENV_FILE.exists() else 'missing'})") print("\n== installed tools ==") any_installed = False for t in tools.values(): if venv_python(t["id"]).exists(): any_installed = True print(f" ✓ {t['id']}") if not any_installed: print(" (none yet — litrun.py install <id>)") print("\n== API keys seen (OS env + .env) ==") env = merged_env() needed = sorted({k for t in tools.values() for k in t.get("env", []) + t.get("env_optional", [])}) for k in needed: print(f" {'✓' if env.get(k) else '✗'} {k}") missing = [k for k in needed if not env.get(k)] if missing: print(f"\nSet missing keys with: litrun.py env --set {missing[0]}=...") def cmd_env(tools, args): env = read_env_file() if args.set: for pair in args.set: if "=" not in pair: die(f"--set expects KEY=VALUE, got '{pair}'") k, v = pair.split("=", 1) env[k.strip()] = v.strip() write_env_file(env) print(f"Wrote {len(args.set)} key(s) to {ENV_FILE}") if not env: print(f"No keys set yet. Add with: litrun.py env --set OPENAI_API_KEY=sk-...\n(file: {ENV_FILE})") return print(f"# {ENV_FILE}") for k, v in sorted(env.items()): masked = (v[:4] + "…" + v[-2:]) if len(v) > 8 else "***" print(f"{k}={masked}") def cmd_install(tools, args): t = get_tool(tools, args.id) if not t.get("pip"): if t["kind"] == "mcp-server": print(f"{t['id']} needs no install (runs via uvx). Register it with: litrun.py mcp {t['id']}") return if t["kind"] == "python-script": print(f"{t['id']} needs no install — it is standard-library only. " f"Just run it: litrun.py run {t['id']} -- <args>") return die(f"{t['id']} has no pip recipe; see notes: {t['notes']}") print(f"Installing {t['name']} into {venv_dir(t['id'])} ...") pip_install(t["id"], t["pip"]) print(f"\n✓ Installed {t['id']}.") if t.get("entry"): print(f" Run it: litrun.py run {t['id']} -- <args> (e.g. {t.get('example','')})") if t["kind"] == "mcp-server": print(f" Register it: litrun.py mcp {t['id']}") req_missing = [k for k in t.get("env", []) if not merged_env().get(k)] if req_missing: print(f" ⚠ Needs API keys: {', '.join(req_missing)} — set with litrun.py env --set KEY=...") def cmd_run(tools, args): t = get_tool(tools, args.id) env = merged_env() missing = [k for k in t.get("env", []) if not env.get(k)] if missing: die(f"missing required env: {', '.join(missing)}. Set with `litrun.py env --set {missing[0]}=...`") if not venv_python(t["id"]).exists() and t.get("pip"): print(f"{t['id']} not installed yet — installing first.") pip_install(t["id"], t["pip"]) if t["kind"] in ("python-cli", "python-script"): exec_cli(t, args.rest, env) elif t["kind"] == "python-lib": print(f"{t['name']} is a library, not a CLI. Running its example snippet:\n {t.get('example','')}\n") # Execute the example via the tool's own venv python. snippet = t.get("example", "") if snippet.startswith("python -c ") or snippet.startswith('python -c'): code = snippet.split("-c", 1)[1].strip().strip('"') run_cmd([str(venv_python(t["id"])), "-c", code], env=env, check=False) else: print("No auto-runnable one-liner; see notes:") print(f" {t['notes']}") if t.get("clone_for_ui"): print(f" For the full UI, clone {t['repo']} and follow its README.") elif t["kind"] == "mcp-server": print(f"{t['id']} is an MCP server — it is launched by your MCP client, not run directly.") print(f"Register it with: litrun.py mcp {t['id']}") else: die(f"don't know how to run kind '{t['kind']}'") def cmd_workflow(tools, args): wfs = load_workflows() if args.wf_cmd == "list": for w in wfs.values(): print(f"# {w['id']} — {w['name']}") print(f" {w['description']}") print(f" params: {', '.join(w['params'])}") chain = " → ".join(s["tool"] for s in w["steps"]) print(f" steps: {chain}\n") print("Run: litrun.py workflow run <id> --input <path> [--question \"...\"] [--dry-run]") return if not args.id: die("usage: litrun.py workflow run <id> [--input ...] [--question ...]") if args.id not in wfs: die(f"unknown workflow '{args.id}'. Run `litrun.py workflow list`.") w = wfs[args.id] # Collect params from flags. params = {} if args.input: params["input"] = args.input if args.question: params["question"] = args.question if args.query: params["query"] = args.query if args.max: params["max"] = str(args.max) for pair in (args.param or []): if "=" not in pair: die(f"--param expects KEY=VALUE, got '{pair}'") k, v = pair.split("=", 1) params[k] = v for k, v in w.get("defaults", {}).items(): params.setdefault(k, v) workdir = WORKSPACE / "runs" / w["id"] params["workdir"] = str(workdir) missing = [p for p in w["params"] if p not in params] if missing: if args.dry_run: for p in missing: params[p] = f"<{p}>" else: die(f"workflow '{w['id']}' needs params: {', '.join(missing)}. " f"Provide with --input / --question / --param KEY=VALUE.") def sub(s): for k, v in params.items(): s = s.replace("{" + k + "}", v) return s # Validate every step tool exists and is python-cli; gather required env. step_tools = [] needed_env = set() for step in w["steps"]: t = get_tool(tools, step["tool"]) if t["kind"] not in ("python-cli", "python-script"): die(f"workflow step tool '{t['id']}' is {t['kind']}; workflows chain python-cli / python-script tools only.") step_tools.append(t) needed_env.update(t.get("env", [])) if not args.dry_run: env = merged_env() env_missing = [k for k in sorted(needed_env) if not env.get(k)] if env_missing: die(f"missing API keys for this workflow: {', '.join(env_missing)}. " f"Set with `litrun.py env --set {env_missing[0]}=...`") workdir.mkdir(parents=True, exist_ok=True) else: env = merged_env() print(f"{'DRY RUN — ' if args.dry_run else ''}workflow '{w['id']}' ({len(w['steps'])} step(s))") print(f" workdir: {workdir}\n") for i, (step, t) in enumerate(zip(w["steps"], step_tools), 1): step_args = [sub(a) for a in step.get("args", [])] cwd = sub(step["cwd"]) if step.get("cwd") else None print(f"[step {i}/{len(w['steps'])}] {t['name']}") exec_cli(t, step_args, env, cwd=cwd, dry=args.dry_run) print() if args.dry_run: print("(dry run — nothing executed. Drop --dry-run to run for real.)") else: print(f"✓ workflow '{w['id']}' finished. Outputs: {sub(w.get('outputs') or str(workdir))}") def cmd_ui(tools, args): t = get_tool(tools, args.id) ui = t.get("ui") if not ui: die(f"{t['id']} has no web UI. Runnable ids with a UI: gpt-researcher, storm.") env = merged_env() missing = [k for k in t.get("env", []) if not env.get(k)] if missing: print(f"⚠ {t['id']} UI needs API keys: {', '.join(missing)}.") if ui.get("write_env"): die(f"Set them first: litrun.py env --set {missing[0]}=...") else: print(f" Enter them in the app once it opens (or: litrun.py env --set {missing[0]}=...).") clone_dir = WORKSPACE / t["id"] if not clone_dir.exists(): WORKSPACE.mkdir(parents=True, exist_ok=True) run_cmd(["git", "clone", "--depth", "1", t["repo"], str(clone_dir)]) else: print(f"(repo already cloned at {clone_dir})", file=sys.stderr) ensure_venv(t["id"]) for req in ui.get("requirements", []): req_path = clone_dir / req if not req_path.exists(): print(f"⚠ requirements file not found: {req_path} (upstream layout may have changed)", file=sys.stderr) continue if have("uv"): run_cmd(["uv", "pip", "install", "--python", str(venv_python(t["id"])), "-r", str(req_path)]) else: run_cmd([str(venv_python(t["id"])), "-m", "pip", "install", "-r", str(req_path)]) if ui.get("write_env"): keys = t.get("env", []) + t.get("env_optional", []) lines = [f"{k}={env[k]}\n" for k in keys if env.get(k)] if lines: (clone_dir / ".env").write_text("".join(lines)) print(f"Wrote {len(lines)} key(s) to {clone_dir / '.env'}", file=sys.stderr) # Prepend the venv's bin to PATH so `streamlit`/`uvicorn`/`python` resolve to it. binpath = venv_dir(t["id"]) / ("Scripts" if os.name == "nt" else "bin") env["PATH"] = f"{binpath}{os.pathsep}{env.get('PATH', '')}" print(f"\n▶ Launching {t['name']} UI — open {ui['url']} once it's up. Ctrl-C to stop.") if ui.get("notes"): print(f" {ui['notes']}") run_cmd(list(ui["run"]), env=env, check=False) def cmd_mcp(tools, args): t = get_tool(tools, args.id) m = t.get("mcp") if not m: die(f"{t['id']} is not an MCP server.") launcher = m["launcher"] if launcher == "uvx": storage = args.storage or str(WORKSPACE / t["id"]) arglist = [a.replace("{storage}", storage) for a in m["args_template"]] block = {"command": "uvx", "args": arglist} elif launcher == "venv-module": if not venv_python(t["id"]).exists(): print(f"(note: {t['id']} not installed yet — run `litrun.py install {t['id']}` so this command resolves)") block = {"command": str(venv_python(t["id"])), "args": ["-m", m["module"]]} elif launcher == "venv-entry": if not venv_python(t["id"]).exists(): print(f"(note: {t['id']} not installed yet — run `litrun.py install {t['id']}` so this command resolves)") block = {"command": str(venv_bin(t["id"], m["entry"]))} else: die(f"unknown mcp launcher '{launcher}'") if m.get("env"): block["env"] = dict(m["env"]) config = {"mcpServers": {m["key"]: block}} print(f"# Add this to your MCP client config ({args.client}):") if args.client == "claude": print("# Claude Code: ~/.claude.json (or project .mcp.json) | Claude Desktop: claude_desktop_config.json") else: print("# Cursor: ~/.cursor/mcp.json (or .cursor/mcp.json in the project)") print(json.dumps(config, indent=2)) if t.get("env_optional"): print(f"\n# Optional env you may want to fill into the 'env' block: {', '.join(t['env_optional'])}") print(f"\n# notes: {t['notes']}") def main(): p = argparse.ArgumentParser(prog="litrun.py", description="Install & run AI lit-review tools by id.") sub = p.add_subparsers(dest="cmd", required=True) p_list = sub.add_parser("list", help="list runnable tools") p_list.add_argument("--category") p_list.add_argument("--kind") p_info = sub.add_parser("info", help="show a tool's install/run details") p_info.add_argument("id") sub.add_parser("doctor", help="check toolchain, installs, and API keys") p_env = sub.add_parser("env", help="show/edit the shared .env") p_env.add_argument("--set", nargs="*", help="KEY=VALUE pairs to write") p_inst = sub.add_parser("install", help="install a tool into an isolated venv") p_inst.add_argument("id") p_run = sub.add_parser("run", help="run a tool (installs if needed)") p_run.add_argument("id") p_run.add_argument("rest", nargs=argparse.REMAINDER, help="args after `--` are passed to the tool") p_mcp = sub.add_parser("mcp", help="print the MCP client config for an MCP server") p_mcp.add_argument("id") p_mcp.add_argument("--storage", help="storage path (uvx servers)") p_mcp.add_argument("--client", choices=["claude", "cursor"], default="claude") p_ui = sub.add_parser("ui", help="clone & launch a tool's web UI (gpt-researcher, storm)") p_ui.add_argument("id") p_wf = sub.add_parser("workflow", help="run a named multi-step pipeline") p_wf.add_argument("wf_cmd", choices=["list", "run"]) p_wf.add_argument("id", nargs="?") p_wf.add_argument("--input") p_wf.add_argument("--question") p_wf.add_argument("--query") p_wf.add_argument("--max") p_wf.add_argument("--param", action="append", help="extra KEY=VALUE params") p_wf.add_argument("--dry-run", action="store_true", dest="dry_run", help="print resolved step commands without running") args = p.parse_args() # Strip a leading `--` separator from run's REMAINDER. if getattr(args, "rest", None) and args.rest and args.rest[0] == "--": args.rest = args.rest[1:] tools = load_recipes() dispatch = { "list": cmd_list, "info": cmd_info, "doctor": cmd_doctor, "env": cmd_env, "install": cmd_install, "run": cmd_run, "mcp": cmd_mcp, "ui": cmd_ui, "workflow": cmd_workflow, } dispatch[args.cmd](tools, args) if __name__ == "__main__": main() -
resolve_oa.py 13.3 KB
#!/usr/bin/env python3 """resolve_oa — turn DOIs into legally downloadable full text. A search gives you metadata; a citation-backed answer needs the actual paper. This walks an escalating chain of open-access resolvers per DOI and stops at the first one that yields a file: 1. Unpaywall best OA location (needs an email — free, no key) 2. OpenAlex best_oa_location / oa_url (keyless) 3. Europe PMC OA full-text XML → plain text (keyless, biomed) 4. arXiv preprint PDF when the record has an arXiv id 5. CORE OA aggregator (only if CORE_API_KEY is set) Standard library only. Nothing here bypasses a paywall: a DOI nothing can open is reported as `closed` and skipped — or as `not-in-crossref` when Crossref has never heard of it, which means either an invented citation or a DOI registered elsewhere (DataCite datasets, some preprint servers). Either way, do not cite it unchecked. Input is DOIs on the command line, a text file of DOIs, or the results.json / manifest.json that fetch_papers.py wrote. Usage: resolve_oa.py --from-json ./papers/results.json --outdir ./papers --email you@example.com resolve_oa.py --doi 10.7717/peerj.4375 --doi 10.1038/nature12373 --outdir ./papers """ import argparse import json import os import re import sys import time import urllib.error import urllib.parse import urllib.request import xml.etree.ElementTree as ET from pathlib import Path UA = "litrun-resolve-oa/1.0 (+https://github.com/brycewang-stanford/lit-review-agent-tools)" def _get(url, params=None, timeout=60, retries=1): if params: url = f"{url}?{urllib.parse.urlencode(params, doseq=True)}" for attempt in range(retries + 1): try: req = urllib.request.Request(url, headers={"User-Agent": UA}) with urllib.request.urlopen(req, timeout=timeout) as r: return r.read() except urllib.error.HTTPError as e: if e.code in (429, 500, 502, 503) and attempt < retries: time.sleep(3) continue raise RuntimeError(f"HTTP {e.code}") from e except Exception as e: if attempt < retries: time.sleep(2) continue raise RuntimeError(f"{type(e).__name__}: {e}") from e def _get_json(url, params=None, timeout=60): return json.loads(_get(url, params, timeout).decode("utf-8", "replace")) def norm_doi(doi): doi = (doi or "").strip().lower() return re.sub(r"^https?://(dx\.)?doi\.org/", "", doi).rstrip(".") def safe_name(s, maxlen=90): return (re.sub(r"[^\w\-]+", "_", s or "").strip("_")[:maxlen]) or "paper" def is_pdf(blob): return blob[:5] == b"%PDF-" def jats_to_text(xml_bytes): """Flatten PMC's JATS XML into title + abstract + body, dropping the bibliographic front matter that makes a naive itertext() dump unreadable.""" root = ET.fromstring(xml_bytes) chunks = [] title = root.find(".//article-title") if title is not None: chunks.append("".join(title.itertext()).strip()) for tag in ("abstract", "body"): for el in root.findall(f".//{tag}"): chunks.append(" ".join("".join(el.itertext()).split())) return "\n\n".join(c for c in chunks if c) # ------------------------------------------------------------------ resolvers # Each returns (kind, bytes_or_text, source_url) or None. def via_unpaywall(doi, email): if not email: return None d = _get_json(f"https://api.unpaywall.org/v2/{urllib.parse.quote(doi)}", {"email": email}) if not d.get("is_oa"): return None loc = d.get("best_oa_location") or {} url = loc.get("url_for_pdf") or loc.get("url") if not url: return None blob = _get(url, timeout=90) return ("pdf", blob, url) if is_pdf(blob) else None def via_openalex(doi): w = _get_json(f"https://api.openalex.org/works/doi:{urllib.parse.quote(doi)}") loc = w.get("best_oa_location") or {} url = loc.get("pdf_url") or (w.get("open_access") or {}).get("oa_url") ids = w.get("ids") or {} if not url: # An arXiv landing page in `ids` is still a downloadable preprint. return ("arxiv-id", ids.get("arxiv") or "", "") if ids.get("arxiv") else None blob = _get(url, timeout=90) return ("pdf", blob, url) if is_pdf(blob) else None def via_europepmc(doi): params = {"query": f'DOI:"{doi}"', "format": "json", "pageSize": 1, "resultType": "core"} res = _get_json("https://www.ebi.ac.uk/europepmc/webservices/rest/search", params) hits = res.get("resultList", {}).get("result", []) if not hits: return None h = hits[0] pmcid = h.get("pmcid") if pmcid and h.get("isOpenAccess") == "Y": url = f"https://www.ebi.ac.uk/europepmc/webservices/rest/{pmcid}/fullTextXML" try: text = jats_to_text(_get(url, timeout=90)) if len(text) > 2000: # a stub isn't full text return ("txt", text.encode("utf-8"), url) except Exception: pass for u in ((h.get("fullTextUrlList") or {}).get("fullTextUrl") or []): if u.get("documentStyle") == "pdf" and u.get("availabilityCode") in ("OA", "F"): try: blob = _get(u["url"], timeout=90) if is_pdf(blob): return ("pdf", blob, u["url"]) except Exception: continue return None def via_arxiv(arxiv_id): aid = (arxiv_id or "").rsplit("/", 1)[-1] if not aid: return None url = f"https://arxiv.org/pdf/{aid}" blob = _get(url, timeout=90) return ("pdf", blob, url) if is_pdf(blob) else None def doi_registered(doi): """Does this DOI exist at all? Asked only when nothing resolved, so that a fabricated citation is reported as `not-in-crossref` rather than as `closed` — 'behind a paywall' and 'no such record' are very different answers.""" try: _get(f"https://api.crossref.org/works/{urllib.parse.quote(doi)}", timeout=30, retries=0) return True except RuntimeError as e: return False if "HTTP 404" in str(e) else None # None = could not tell def via_core(doi, key): if not key: return None body = json.dumps({"q": f'doi:"{doi}"', "limit": 1}).encode() req = urllib.request.Request( "https://api.core.ac.uk/v3/search/works", data=body, headers={"User-Agent": UA, "Authorization": f"Bearer {key}", "Content-Type": "application/json"}, ) with urllib.request.urlopen(req, timeout=60) as r: res = json.loads(r.read().decode("utf-8", "replace")) for w in res.get("results", []): url = w.get("downloadUrl") if url: try: blob = _get(url, timeout=90) if is_pdf(blob): return ("pdf", blob, url) except Exception: continue return None # ------------------------------------------------------------------ input def load_targets(args): """→ [{doi, title, arxiv_id, pdf_url}] from whichever input was given.""" targets = [] for d in args.doi or []: targets.append({"doi": norm_doi(d), "title": "", "arxiv_id": "", "pdf_url": ""}) if args.from_json: data = json.loads(Path(args.from_json).read_text()) rows = data.get("results", data) if isinstance(data, dict) else data for r in rows: if not isinstance(r, dict): continue aid = r.get("id", "") if r.get("source") == "arxiv" else "" targets.append({"doi": norm_doi(r.get("doi")), "title": r.get("title", ""), "arxiv_id": aid, "pdf_url": r.get("pdf_url", "")}) if args.from_file: for line in Path(args.from_file).read_text().splitlines(): line = line.strip() if line and not line.startswith("#"): targets.append({"doi": norm_doi(line), "title": "", "arxiv_id": "", "pdf_url": ""}) # Keep records that carry *some* handle, deduplicated. seen, out = set(), [] for t in targets: k = t["doi"] or t["arxiv_id"] or t["pdf_url"] if not k or k in seen: continue seen.add(k) out.append(t) return out def main(): p = argparse.ArgumentParser(prog="resolve_oa.py", description=__doc__.split("\n")[0], formatter_class=argparse.RawDescriptionHelpFormatter) p.add_argument("--doi", action="append", help="a DOI (repeatable)") p.add_argument("--from-json", dest="from_json", help="results.json/manifest.json from fetch_papers.py") p.add_argument("--from-file", dest="from_file", help="text file, one DOI per line") p.add_argument("--outdir", required=True) p.add_argument("--email", default=os.environ.get("UNPAYWALL_EMAIL") or os.environ.get("NCBI_EMAIL", ""), help="contact email — required by Unpaywall (else that step is skipped)") p.add_argument("--max", type=int, default=0, help="stop after N DOIs (0 = all)") p.add_argument("--skip-existing", action="store_true", help="don't re-download files already in --outdir") args = p.parse_args() if not (args.doi or args.from_json or args.from_file): sys.exit("resolve_oa: give --doi, --from-json or --from-file.") targets = load_targets(args) if args.max: targets = targets[:args.max] if not targets: sys.exit("resolve_oa: no usable DOIs in the input.") outdir = Path(args.outdir) outdir.mkdir(parents=True, exist_ok=True) core_key = os.environ.get("CORE_API_KEY", "") if not args.email: print("resolve_oa: no --email/UNPAYWALL_EMAIL — skipping Unpaywall, " "the rest of the chain still runs keyless.", file=sys.stderr) report, counts = [], {"pdf": 0, "txt": 0, "closed": 0, "not-in-crossref": 0, "error": 0, "skipped": 0} for i, t in enumerate(targets, 1): doi, label = t["doi"], (t["title"] or t["doi"] or t["arxiv_id"])[:60] stem = safe_name(doi.replace("/", "_") or t["arxiv_id"]) print(f"[{i}/{len(targets)}] {label}", file=sys.stderr) existing = list(outdir.glob(f"{stem}.*")) if args.skip_existing and existing: counts["skipped"] += 1 report.append({"doi": doi, "status": "skipped", "file": existing[0].name}) print(" · already present", file=sys.stderr) continue chain = [] if doi: chain += [("unpaywall", lambda: via_unpaywall(doi, args.email)), ("openalex", lambda: via_openalex(doi)), ("europepmc", lambda: via_europepmc(doi))] if t["arxiv_id"]: chain.append(("arxiv", lambda: via_arxiv(t["arxiv_id"]))) if t["pdf_url"]: chain.append(("search-hit", lambda: (lambda b: ("pdf", b, t["pdf_url"]) if is_pdf(b) else None)( _get(t["pdf_url"], timeout=90)))) if doi and core_key: chain.append(("core", lambda: via_core(doi, core_key))) got, errs = None, [] for name, fn in chain: try: res = fn() except Exception as e: errs.append(f"{name}: {e}") continue if res is None: continue kind, payload, url = res if kind == "arxiv-id": # OpenAlex handed us a preprint id instead of a file try: res2 = via_arxiv(payload) except Exception as e: errs.append(f"arxiv: {e}") continue if not res2: continue kind, payload, url = res2 name = "openalex→arxiv" got = (name, kind, payload, url) break if not got: if not doi: status = "error" else: status = "closed" if doi_registered(doi) is not False else "not-in-crossref" counts[status] += 1 report.append({"doi": doi, "title": t["title"], "status": status, "tried": errs}) why = "DOI is not registered with Crossref" if status == "not-in-crossref" else "no open copy found" print(f" ✗ {why}{' (' + '; '.join(errs[:2]) + ')' if errs else ''}", file=sys.stderr) continue name, kind, payload, url = got dest = outdir / f"{stem}.{'pdf' if kind == 'pdf' else 'txt'}" dest.write_bytes(payload) counts[kind] += 1 report.append({"doi": doi, "title": t["title"], "status": kind, "via": name, "url": url, "file": dest.name, "bytes": len(payload)}) print(f" ✓ {kind} via {name} ({len(payload)//1024} KB)", file=sys.stderr) time.sleep(0.3) out = {"resolved": counts, "total": len(targets), "items": report} (outdir / "oa_report.json").write_text(json.dumps(out, indent=2, ensure_ascii=False), encoding="utf-8") hit = counts["pdf"] + counts["txt"] print(f"\nresolve_oa: {hit}/{len(targets)} open ({counts['pdf']} PDF, {counts['txt']} full-text XML→txt), " f"{counts['closed']} closed, {counts['not-in-crossref']} DOI(s) unknown to Crossref, " f"{counts['error']} errored, {counts['skipped']} skipped", file=sys.stderr) print(f" files + oa_report.json in {outdir}", file=sys.stderr) if hit == 0: sys.exit("resolve_oa: nothing was openly available — the corpus is unchanged.") if __name__ == "__main__": main()
-
-
SKILL.md 14.1 KB
--- name: literature-review-tools description: >- Recommend AND run open-source AI tools, agents, Claude Code / Codex skills, and MCP servers for any stage of a literature review — searching, reading, extracting, synthesizing, screening, citation-checking, and paper writing. Use when the user asks "what tool should I use to..." OR "install/run/use <tool> to ..." for research/lit-review work: automating a survey or related-work section, PDF→Markdown extraction for LLMs (MinerU/marker/docling), PRISMA / systematic review (ASReview), citation-backed Q&A over PDFs (PaperQA2), wiring papers into Claude/Cursor via MCP (arxiv/paper-search/zotero servers), or chatting with a Zotero library. Also answers direct literature lookups — find papers on a topic, resolve a DOI, get an open-access PDF, check a citation, search PubMed / OpenAlex / Crossref / Semantic Scholar / Europe PMC / arXiv — with no install and no API key. Ships a launcher (scripts/litrun.py) that installs each tool in an isolated venv and runs it, plus keyless bundled scripts for multi-source search and open-access full-text retrieval. Curated catalog of 70+ vetted projects. 支持中英文(用于「文献综述工具选型」「文献检索」与「一键安装/运行」)。 --- # Literature Review Tools — Select & Run A curated, use-case-organized catalog of the strongest **open-source** AI tools for literature review — **plus a launcher that actually installs and runs the top ones.** Covers: end-to-end research agents, deep-research / auto-survey generators, autonomous "idea→paper" systems, citation-backed RAG over PDFs, PRISMA screening, MCP servers, Zotero/Obsidian integrations, PDF→structured extraction, citation graphs, and paper-writing / peer-review assistants. Full source of truth (README, always current star counts): <https://github.com/brycewang-stanford/lit-review-agent-tools> ## Three modes - **Look up** — user wants *the literature itself*: "find papers on X", "get me the PDF for this DOI", "does this citation exist", "what does PubMed have since 2022". Answer it directly — no install, no key. Start with [`reference/apis/README.md`](reference/apis/README.md) for one-off lookups, or run the bundled `papers-fetch` / `oa-resolve` scripts for anything corpus-sized. - **Recommend** — user asks "what should I use to …". Route with the tables below; cite the catalog for details. - **Run** — user asks to *install / run / use* a specific tool ("turn this PDF into Markdown with MinerU", "ask PaperQA2 about these papers", "set up the arXiv MCP server"). Drive [`scripts/litrun.py`](scripts/litrun.py) via Bash — do not hand the user raw pip commands to copy. Mode 1 is the cheap default. Do not send someone to install PyTorch when they asked for five papers and a PDF. ## Look up mode — search without installing anything Two bundled scripts, both **standard-library only**: no venv, no pip, no API key. Run them straight (`python3 scripts/fetch_papers.py …`) or through the launcher. ```bash # 1. search six indexes at once, deduplicated by DOI python3 scripts/fetch_papers.py --query "active learning for screening" \ --sources openalex,crossref,semanticscholar,pubmed,europepmc,arxiv \ --max 15 --dedup-titles --outdir ./corpus # 2. turn those DOIs into full text you may legally read python3 scripts/resolve_oa.py --from-json ./corpus/results.json --outdir ./corpus ``` `fetch_papers.py` writes `results.json` (normalised records, per-source counts, and the errors of any source that failed) plus one `.txt` per paper. `resolve_oa.py` walks Unpaywall → OpenAlex → Europe PMC → arXiv → CORE and writes `oa_report.json` recording how every DOI resolved — **including the ones that stayed closed**, and separating those from DOIs Crossref has never registered (usually an invented citation). On a 47-paper corpus with zero keys it recovered 33 full texts; nothing in it bypasses a paywall, and you should not offer to. For a *single* lookup — one DOI, one author, one citation check — skip the scripts and call the API directly (`WebFetch`/`curl`). [`reference/apis/`](reference/apis/) has a routing table, the identifier formats, and one page per API with the endpoints and the failure modes that waste time (Crossref's `select` 400, OpenAlex's inverted abstracts, Semantic Scholar's exhausted keyless pool, PubMed's multi-part `AbstractText`). Report which indexes you queried and which came back empty. "OpenAlex had no match" is a fact; "that paper doesn't exist" is a much bigger claim than one API can support. ## Run mode — how to drive `scripts/litrun.py` The launcher installs each supported tool into its own venv under `~/.lit-review-tools/` (uses `uv` if present, else `python -m venv`) and reads API keys from one shared `~/.lit-review-tools/.env`. Machine-readable recipes: [`recipes/recipes.json`](recipes/recipes.json). Typical flow when the user wants to *use* a tool: 1. `python3 scripts/litrun.py doctor` — check toolchain + which API keys are already set. 2. `python3 scripts/litrun.py info <id>` — confirm what the tool needs (entry, required env). 3. If a required key is missing, ask the user for it, then `litrun.py env --set KEY=VALUE` (never echo the value back in full). 4. `python3 scripts/litrun.py run <id> -- <tool args>` — installs on first use, then runs. For PDF tools pass the real file path; e.g. `run mineru -- -p paper.pdf -o ./out -b pipeline`. 5. For **MCP servers**, don't "run" them — `litrun.py mcp <id>` prints the client config block to register in Claude Code / Cursor. Commands: `list [--category C] [--kind K]` · `info <id>` · `doctor` · `env [--set K=V]` · `install <id>` · `run <id> -- <args>` · `mcp <id> [--storage PATH] [--client claude|cursor]` · `ui <id>`. Runnable ids by kind: - **python-cli (auto install+run):** `mineru`, `marker`, `docling` (PDF→Markdown) · `paper-qa` (cited Q&A) · `asreview` (PRISMA screening UI) - **python-script, zero install (stdlib only):** `papers-fetch` (6-source deduplicated search) · `oa-resolve` (DOI → open-access full text) - **python-script (bundled, auto install+run):** `arxiv-fetch` (arXiv PDFs) · `openalex-fetch` (OpenAlex; PDF or abstract .txt) · `pubmed-fetch` (PubMed abstracts) — all keyless for light use. `papers-fetch` supersedes all three when you want coverage rather than one index. - **python-lib (install + run example):** `gpt-researcher`, `storm` (deep research; need API keys) · `scholarly`, `pyalex` (API clients) - **mcp-server (install + `mcp` config):** `arxiv-mcp-server`, `paper-search-mcp`, `zotero-mcp` For **`gpt-researcher`** and **`storm`**, `litrun.py ui <id>` clones the repo and launches the full web UI (GPT Researcher → FastAPI at :8000; STORM → Streamlit at :8501). These are long-running servers — launch them with a background Bash call and tell the user the URL. gpt-researcher's UI needs `OPENAI_API_KEY` + `TAVILY_API_KEY` set first (litrun writes them into the repo's `.env`); STORM takes its keys in the app sidebar. ### Chained pipelines For multi-tool tasks, prefer a named workflow over hand-wiring steps: `litrun.py workflow list` then `litrun.py workflow run <id> [--input PATH] [--query "..."] [--question "..."] [--max N]`. Built-ins: - `pdf-to-markdown` — a PDF/folder → clean Markdown (MinerU) - `pdf-corpus-qa` — a folder of PDFs → citation-backed answer (PaperQA2) - `pdf-md-then-qa` — convert to Markdown **and** answer a question over the corpus - `topic-to-pdfs` — arXiv query → download top-N PDFs (arxiv-fetch, no key) - `topic-to-review` — arXiv query → download PDFs → citation-backed answer (PaperQA2). The end-to-end "retrieve then review" pipeline; no MCP client needed. Needs `OPENAI_API_KEY` for the QA step. - `topic-to-review-multi` — retrieve from **arXiv + OpenAlex** into one corpus → citation-backed answer. Broader coverage; resilient if one source is rate-limited. - `topic-to-related-work` — retrieve (arXiv + OpenAlex) → PaperQA2 drafts a **cited related-work paragraph** synthesizing themes/methods/gaps. Needs `OPENAI_API_KEY`. - `topic-to-corpus` — six-source search → deduplicated corpus. **Keyless, no install** — the default first step for any review. - `topic-to-fulltext` — six-source search → open-access full text for every DOI it can legally get. Keyless. - `topic-to-fulltext-review` — the strongest path: search → dedup → full text → PaperQA2 answers over **full texts, not abstracts**. `OPENAI_API_KEY` for the last step only. Prefer the `topic-to-fulltext*` workflows over `topic-to-review-multi`: same idea, six sources instead of two, DOI-level deduplication, and it retrieves the papers rather than the abstracts. Biomedical topics need no special casing — PubMed and Europe PMC are already in the source set. Add `--dry-run` first to show the exact resolved step commands without executing — good for confirming paths with the user before a heavy run. Workflows fail fast if a required API key is missing. Guardrails: installs and downloads happen under the user's home and hit the network — for a heavy first install (marker/docling pull in PyTorch) say so before running. Never fabricate API keys. If a `run` fails, show the real error rather than claiming success. Paths in this file (`scripts/…`, `recipes/…`) are relative to this skill's directory. ## Recommend mode — how to route 1. Identify **which stage** of the lit-review workflow the user is on (search → read → extract → synthesize → screen → cite-check → write/review). 2. Match it to a category below and recommend the **⭐ editor's pick first**, then 1–2 alternatives. 3. For anything beyond the top pick — full star counts, every project in a category, or a category not summarized here — read [`reference/catalog.md`](reference/catalog.md). Do **not** guess project names or URLs; pull them from the catalog. 4. Give a one-line "why this one" tied to the user's constraint (Claude Code vs. standalone, open vs. commercial, privacy/local, medical, etc.). If the pick is a runnable id above, offer to install/run it. ## ⚡ 30-second picker ```text Just need the papers themselves (topic / DOI / OA PDF) ──▶ Look up mode — no install ⭐ Use Claude Code, want end-to-end research→paper ──────────▶ academic-research-skills ⭐ Want AI to research a topic → cited report ───────────────▶ GPT Researcher / STORM Want fully autonomous "idea → submittable paper" ────────▶ AI-Scientist-v2 / AutoResearchClaw Citation-backed Q&A over a pile of PDFs ──────────────────▶ PaperQA2 Rigorous PRISMA review (thousands of abstracts) ─────────▶ ASReview / prismAId Clean Markdown from PDFs to feed an LLM ─────────────────▶ MinerU / Docling / marker Lit capabilities inside Claude / Cursor (MCP) ───────────▶ paper-search-mcp / zotero-mcp Chat with your library inside Zotero ────────────────────▶ zotero-gpt / PapersGPT Pre-submission AI peer review ───────────────────────────▶ open_reviewer / ai-peer-review ``` ## Categories (top pick per category) | Category | Editor's pick ⭐ | When | |---|---|---| | All-in-one research agents & skills | **academic-research-skills** | Claude Code user wanting research→write→review→revise, with integrity/citation gates | | Deep research & auto-survey | **STORM** / **gpt-researcher** | Topic → cited survey / report / related-work | | Autonomous science (idea→paper) | **AI-Scientist(-v2)** / **AutoResearchClaw** | Fully automated discovery: lit + hypotheses + experiments + writing | | Literature Q&A / RAG | **paper-qa (PaperQA2)** | Citation-backed answers over a PDF corpus | | Systematic review & screening | **ASReview** | Active-learning screening of thousands of abstracts (PRISMA) | | MCP servers | **zotero-mcp** / **arxiv-mcp-server** | Wire papers into Claude / Cursor / Cline | | Zotero / Obsidian integration | **zotero-gpt** | Chat with your library inside your reference manager | | PDF → structured extraction | **MinerU** / **docling** / **marker** | Turn PDFs into clean Markdown/JSON for LLMs | | Citation graphs & API clients | **scholarly** / **pyalex** | Citation-network analysis; scripting academic DBs | | Writing & peer-review assistants | **open_reviewer** / **ai-peer-review** | Draft, polish, and pre-submission review | | Awesome lists | **Awesome-Auto-Research-Tools** | Browse the whole landscape | ## Decision table (map need → recommendation) | User's need | Recommend | |---|---| | Claude Code, end-to-end research→paper | **academic-research-skills** (most complete, #1 in space) | | Generic "research this topic for me" agent | **GPT Researcher** / **STORM** | | Wiki/survey-style long-form with citations | **STORM / Co-STORM** | | Fully autonomous "idea → submittable paper" | **AI-Scientist-v2** / **AutoResearchClaw** | | Cited Q&A over many PDFs | **PaperQA / PaperQA2** | | Rigorous PRISMA systematic review | **ASReview** or **prismAId** | | PDF → clean Markdown for an LLM | **MinerU / Docling / marker** | | Lit capabilities in an MCP client | **paper-search-mcp / zotero-mcp** | | Chat with library inside Zotero | **zotero-gpt / PapersGPT** | | AI pre-review before submission | **open_reviewer / ai-peer-review** | | Just want to browse the landscape | The **Awesome lists** section | ## Notes & caveats - **Open-source is prioritized.** Commercial/closed tools (Elicit, Consensus, Scite, SciSpace, Research Rabbit, Connected Papers) are listed for reference only — see the catalog's commercial section. - **Star counts drift.** The catalog's numbers are periodic GitHub-API snapshots — treat as rough popularity signals, not exact. For live numbers, point the user at the repo. - **Match the constraint, not just the task.** Privacy/local → `local-deep-research`; medical → `medsci-skills` / `paperai`; Codex instead of Claude → `academic-research-skills-codex`. Full catalog with every project, star count, and one-line description: [`reference/catalog.md`](reference/catalog.md).
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.