Claude Skill

literature-review-tools

Recommend AND run open-source AI tools, agents, Claude Code / Codex skills, and MCP servers for any stage of a literature review — searching, reading, extracting, synthesizing, screening, citation-checking, and paper writing. Use when the user asks "what tool should I use to..."

LLM Mart · 0 points · 19 views 29 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download brycewang-stanford-lit-review-agent-tools-skills_literature-review-tools-df3fe57.zip · 57 KB

Install

skills CLI npx skills add https://github.com/brycewang-stanford/lit-review-agent-tools/tree/main/skills/literature-review-tools
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install brycewang-stanford-lit-review-agent-tools@llmmart
Git git clone https://github.com/brycewang-stanford/lit-review-agent-tools.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole brycewang-stanford/lit-review-agent-tools collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Literature Review Tools — Select & Run

A curated, use-case-organized catalog of the strongest open-source AI tools for literature review — plus a launcher that actually installs and runs the top ones. Covers: end-to-end research agents, deep-research / auto-survey generators, autonomous "idea→paper" systems, citation-backed RAG over PDFs, PRISMA screening, MCP servers, Zotero/Obsidian integrations, PDF→structured extraction, citation graphs, and paper-writing / peer-review assistants.

Full source of truth (README, always current star counts): https://github.com/brycewang-stanford/lit-review-agent-tools

Three modes

  • Look up — user wants the literature itself: "find papers on X", "get me the PDF for this DOI", "does this citation exist", "what does PubMed have since 2022". Answer it directly — no install, no key. Start with reference/apis/README.md for one-off lookups, or run the bundled papers-fetch / oa-resolve scripts for anything corpus-sized.
  • Recommend — user asks "what should I use to …". Route with the tables below; cite the catalog for details.
  • Run — user asks to install / run / use a specific tool ("turn this PDF into Markdown with MinerU", "ask PaperQA2 about these papers", "set up the arXiv MCP server"). Drive scripts/litrun.py via Bash — do not hand the user raw pip commands to copy.

Mode 1 is the cheap default. Do not send someone to install PyTorch when they asked for five papers and a PDF.

Look up mode — search without installing anything

Two bundled scripts, both standard-library only: no venv, no pip, no API key. Run them straight (python3 scripts/fetch_papers.py …) or through the launcher.

# 1. search six indexes at once, deduplicated by DOI
python3 scripts/fetch_papers.py --query "active learning for screening" \
    --sources openalex,crossref,semanticscholar,pubmed,europepmc,arxiv \
    --max 15 --dedup-titles --outdir ./corpus

# 2. turn those DOIs into full text you may legally read
python3 scripts/resolve_oa.py --from-json ./corpus/results.json --outdir ./corpus

fetch_papers.py writes results.json (normalised records, per-source counts, and the errors of any source that failed) plus one .txt per paper. resolve_oa.py walks Unpaywall → OpenAlex → Europe PMC → arXiv → CORE and writes oa_report.json recording how every DOI resolved — including the ones that stayed closed, and separating those from DOIs Crossref has never registered (usually an invented citation). On a 47-paper corpus with zero keys it recovered 33 full texts; nothing in it bypasses a paywall, and you should not offer to.

For a single lookup — one DOI, one author, one citation check — skip the scripts and call the API directly (WebFetch/curl). reference/apis/ has a routing table, the identifier formats, and one page per API with the endpoints and the failure modes that waste time (Crossref's select 400, OpenAlex's inverted abstracts, Semantic Scholar's exhausted keyless pool, PubMed's multi-part AbstractText).

Report which indexes you queried and which came back empty. "OpenAlex had no match" is a fact; "that paper doesn't exist" is a much bigger claim than one API can support.

Run mode — how to drive scripts/litrun.py

The launcher installs each supported tool into its own venv under ~/.lit-review-tools/ (uses uv if present, else python -m venv) and reads API keys from one shared ~/.lit-review-tools/.env. Machine-readable recipes: recipes/recipes.json.

Typical flow when the user wants to use a tool:

  1. python3 scripts/litrun.py doctor — check toolchain + which API keys are already set.
  2. python3 scripts/litrun.py info <id> — confirm what the tool needs (entry, required env).
  3. If a required key is missing, ask the user for it, then litrun.py env --set KEY=VALUE (never echo the value back in full).
  4. python3 scripts/litrun.py run <id> -- <tool args> — installs on first use, then runs. For PDF tools pass the real file path; e.g. run mineru -- -p paper.pdf -o ./out -b pipeline.
  5. For MCP servers, don't "run" them — litrun.py mcp <id> prints the client config block to register in Claude Code / Cursor.

Commands: list [--category C] [--kind K] · info <id> · doctor · env [--set K=V] · install <id> · run <id> -- <args> · mcp <id> [--storage PATH] [--client claude|cursor] · ui <id>.

Runnable ids by kind:

  • python-cli (auto install+run): mineru, marker, docling (PDF→Markdown) · paper-qa (cited Q&A) · asreview (PRISMA screening UI)
  • python-script, zero install (stdlib only): papers-fetch (6-source deduplicated search) · oa-resolve (DOI → open-access full text)
  • python-script (bundled, auto install+run): arxiv-fetch (arXiv PDFs) · openalex-fetch (OpenAlex; PDF or abstract .txt) · pubmed-fetch (PubMed abstracts) — all keyless for light use. papers-fetch supersedes all three when you want coverage rather than one index.
  • python-lib (install + run example): gpt-researcher, storm (deep research; need API keys) · scholarly, pyalex (API clients)
  • mcp-server (install + mcp config): arxiv-mcp-server, paper-search-mcp, zotero-mcp

For gpt-researcher and storm, litrun.py ui <id> clones the repo and launches the full web UI (GPT Researcher → FastAPI at :8000; STORM → Streamlit at :8501). These are long-running servers — launch them with a background Bash call and tell the user the URL. gpt-researcher's UI needs OPENAI_API_KEY + TAVILY_API_KEY set first (litrun writes them into the repo's .env); STORM takes its keys in the app sidebar.

Chained pipelines

For multi-tool tasks, prefer a named workflow over hand-wiring steps: litrun.py workflow list then litrun.py workflow run <id> [--input PATH] [--query "..."] [--question "..."] [--max N]. Built-ins:

  • pdf-to-markdown — a PDF/folder → clean Markdown (MinerU)
  • pdf-corpus-qa — a folder of PDFs → citation-backed answer (PaperQA2)
  • pdf-md-then-qa — convert to Markdown and answer a question over the corpus
  • topic-to-pdfs — arXiv query → download top-N PDFs (arxiv-fetch, no key)
  • topic-to-review — arXiv query → download PDFs → citation-backed answer (PaperQA2). The end-to-end "retrieve then review" pipeline; no MCP client needed. Needs OPENAI_API_KEY for the QA step.
  • topic-to-review-multi — retrieve from arXiv + OpenAlex into one corpus → citation-backed answer. Broader coverage; resilient if one source is rate-limited.
  • topic-to-related-work — retrieve (arXiv + OpenAlex) → PaperQA2 drafts a cited related-work paragraph synthesizing themes/methods/gaps. Needs OPENAI_API_KEY.
  • topic-to-corpus — six-source search → deduplicated corpus. Keyless, no install — the default first step for any review.
  • topic-to-fulltext — six-source search → open-access full text for every DOI it can legally get. Keyless.
  • topic-to-fulltext-review — the strongest path: search → dedup → full text → PaperQA2 answers over full texts, not abstracts. OPENAI_API_KEY for the last step only.

Prefer the topic-to-fulltext* workflows over topic-to-review-multi: same idea, six sources instead of two, DOI-level deduplication, and it retrieves the papers rather than the abstracts. Biomedical topics need no special casing — PubMed and Europe PMC are already in the source set.

Add --dry-run first to show the exact resolved step commands without executing — good for confirming paths with the user before a heavy run. Workflows fail fast if a required API key is missing.

Guardrails: installs and downloads happen under the user's home and hit the network — for a heavy first install (marker/docling pull in PyTorch) say so before running. Never fabricate API keys. If a run fails, show the real error rather than claiming success. Paths in this file (scripts/…, recipes/…) are relative to this skill's directory.

Recommend mode — how to route

  1. Identify which stage of the lit-review workflow the user is on (search → read → extract → synthesize → screen → cite-check → write/review).
  2. Match it to a category below and recommend the ⭐ editor's pick first, then 1–2 alternatives.
  3. For anything beyond the top pick — full star counts, every project in a category, or a category not summarized here — read reference/catalog.md. Do not guess project names or URLs; pull them from the catalog.
  4. Give a one-line "why this one" tied to the user's constraint (Claude Code vs. standalone, open vs. commercial, privacy/local, medical, etc.). If the pick is a runnable id above, offer to install/run it.

⚡ 30-second picker

Just need the papers themselves (topic / DOI / OA PDF) ──▶ Look up mode — no install ⭐
Use Claude Code, want end-to-end research→paper ──────────▶ academic-research-skills ⭐
Want AI to research a topic → cited report ───────────────▶ GPT Researcher / STORM
Want fully autonomous "idea → submittable paper" ────────▶ AI-Scientist-v2 / AutoResearchClaw
Citation-backed Q&A over a pile of PDFs ──────────────────▶ PaperQA2
Rigorous PRISMA review (thousands of abstracts) ─────────▶ ASReview / prismAId
Clean Markdown from PDFs to feed an LLM ─────────────────▶ MinerU / Docling / marker
Lit capabilities inside Claude / Cursor (MCP) ───────────▶ paper-search-mcp / zotero-mcp
Chat with your library inside Zotero ────────────────────▶ zotero-gpt / PapersGPT
Pre-submission AI peer review ───────────────────────────▶ open_reviewer / ai-peer-review

Categories (top pick per category)

Category Editor's pick ⭐ When
All-in-one research agents & skills academic-research-skills Claude Code user wanting research→write→review→revise, with integrity/citation gates
Deep research & auto-survey STORM / gpt-researcher Topic → cited survey / report / related-work
Autonomous science (idea→paper) AI-Scientist(-v2) / AutoResearchClaw Fully automated discovery: lit + hypotheses + experiments + writing
Literature Q&A / RAG paper-qa (PaperQA2) Citation-backed answers over a PDF corpus
Systematic review & screening ASReview Active-learning screening of thousands of abstracts (PRISMA)
MCP servers zotero-mcp / arxiv-mcp-server Wire papers into Claude / Cursor / Cline
Zotero / Obsidian integration zotero-gpt Chat with your library inside your reference manager
PDF → structured extraction MinerU / docling / marker Turn PDFs into clean Markdown/JSON for LLMs
Citation graphs & API clients scholarly / pyalex Citation-network analysis; scripting academic DBs
Writing & peer-review assistants open_reviewer / ai-peer-review Draft, polish, and pre-submission review
Awesome lists Awesome-Auto-Research-Tools Browse the whole landscape

Decision table (map need → recommendation)

User's need Recommend
Claude Code, end-to-end research→paper academic-research-skills (most complete, #1 in space)
Generic "research this topic for me" agent GPT Researcher / STORM
Wiki/survey-style long-form with citations STORM / Co-STORM
Fully autonomous "idea → submittable paper" AI-Scientist-v2 / AutoResearchClaw
Cited Q&A over many PDFs PaperQA / PaperQA2
Rigorous PRISMA systematic review ASReview or prismAId
PDF → clean Markdown for an LLM MinerU / Docling / marker
Lit capabilities in an MCP client paper-search-mcp / zotero-mcp
Chat with library inside Zotero zotero-gpt / PapersGPT
AI pre-review before submission open_reviewer / ai-peer-review
Just want to browse the landscape The Awesome lists section

Notes & caveats

  • Open-source is prioritized. Commercial/closed tools (Elicit, Consensus, Scite, SciSpace, Research Rabbit, Connected Papers) are listed for reference only — see the catalog's commercial section.
  • Star counts drift. The catalog's numbers are periodic GitHub-API snapshots — treat as rough popularity signals, not exact. For live numbers, point the user at the repo.
  • Match the constraint, not just the task. Privacy/local → local-deep-research; medical → medsci-skills / paperai; Codex instead of Claude → academic-research-skills-codex.

Full catalog with every project, star count, and one-line description: reference/catalog.md.

Files (lit-review-agent-tools)
  • recipes
    • recipes.json 12.5 KB
      {
        "_meta": {
          "description": "Runnable-tool manifest for litrun.py. Commands verified against upstream READMEs/PyPI (2025-2026). Star counts and the full 70+ catalog live in ../reference/catalog.md; only tools that can be installed & driven programmatically are listed here.",
          "source": "https://github.com/brycewang-stanford/lit-review-agent-tools"
        },
        "tools": [
          {
            "id": "mineru",
            "name": "MinerU",
            "category": "pdf-extraction",
            "kind": "python-cli",
            "repo": "https://github.com/opendatalab/MinerU",
            "pip": ["mineru[all]"],
            "entry": "mineru",
            "example": "mineru -p input.pdf -o ./output -b pipeline",
            "env": [],
            "env_optional": [],
            "notes": "PDF/Office/image -> LLM-ready Markdown/JSON. No API key. GPU optional; add '-b pipeline' for CPU-only. ~16GB RAM recommended."
          },
          {
            "id": "marker",
            "name": "marker",
            "category": "pdf-extraction",
            "kind": "python-cli",
            "repo": "https://github.com/datalab-to/marker",
            "pip": ["marker-pdf"],
            "entry": "marker_single",
            "example": "marker_single /path/to/file.pdf --output_dir ./output",
            "env": [],
            "env_optional": ["GOOGLE_API_KEY"],
            "notes": "Fast PDF/doc -> Markdown/JSON. No key for base use; '--use_llm' boost needs an LLM key (e.g. GOOGLE_API_KEY). Pulls in PyTorch; GPU recommended. Use 'marker <dir>' for batch."
          },
          {
            "id": "docling",
            "name": "docling",
            "category": "pdf-extraction",
            "kind": "python-cli",
            "repo": "https://github.com/docling-project/docling",
            "pip": ["docling"],
            "entry": "docling",
            "example": "docling https://arxiv.org/pdf/2206.01062",
            "env": [],
            "env_optional": [],
            "notes": "IBM document parser -> Markdown/JSON/HTML. No API key. Accepts local paths or URLs. Optional VLM pipeline: 'docling --pipeline vlm --vlm-model granite_docling <src>'."
          },
          {
            "id": "paper-qa",
            "name": "PaperQA2",
            "category": "rag-qa",
            "kind": "python-cli",
            "repo": "https://github.com/Future-House/paper-qa",
            "pip": ["paper-qa"],
            "entry": "pqa",
            "example": "pqa ask 'What methods does this corpus use?'",
            "env": ["OPENAI_API_KEY"],
            "env_optional": ["ANTHROPIC_API_KEY", "GEMINI_API_KEY", "CROSSREF_API_KEY", "SEMANTIC_SCHOLAR_API_KEY"],
            "notes": "Citation-backed RAG QA. Run pqa from a directory containing your PDFs. Needs an LLM key (OpenAI by default; Anthropic/Gemini/local Ollama also supported). Python 3.11+."
          },
          {
            "id": "asreview",
            "name": "ASReview LAB",
            "category": "systematic-review",
            "kind": "python-cli",
            "repo": "https://github.com/asreview/asreview",
            "pip": ["asreview"],
            "entry": "asreview",
            "example": "asreview lab",
            "env": [],
            "env_optional": [],
            "notes": "Active-learning screening for PRISMA/systematic reviews. 'asreview lab' launches a local web UI in the browser. Imports CSV/RIS/XLSX. No API key, no GPU. Python 3.10+."
          },
          {
            "id": "gpt-researcher",
            "name": "GPT Researcher",
            "category": "deep-research",
            "kind": "python-lib",
            "repo": "https://github.com/assafelovic/gpt-researcher",
            "pip": ["gpt-researcher"],
            "entry": null,
            "example": "python -c \"import asyncio; from gpt_researcher import GPTResearcher; r=GPTResearcher(query='YOUR TOPIC'); asyncio.run(r.conduct_research()); print(asyncio.run(r.write_report()))\"",
            "env": ["OPENAI_API_KEY", "TAVILY_API_KEY"],
            "env_optional": ["OPENAI_BASE_URL"],
            "notes": "Async deep-research library (no console script). Needs BOTH OPENAI_API_KEY and TAVILY_API_KEY. For the web UI, clone the repo and run 'python -m uvicorn main:app --reload'.",
            "clone_for_ui": true,
            "ui": {
              "requirements": ["requirements.txt"],
              "write_env": true,
              "run": ["python", "-m", "uvicorn", "main:app", "--reload"],
              "url": "http://localhost:8000",
              "notes": "FastAPI + web frontend. Reads keys from a .env in the repo root (litrun writes it for you). Serves at :8000."
            }
          },
          {
            "id": "storm",
            "name": "STORM / Co-STORM",
            "category": "deep-research",
            "kind": "python-lib",
            "repo": "https://github.com/stanford-oval/storm",
            "pip": ["knowledge-storm"],
            "entry": null,
            "example": "See repo examples/ — instantiate STORMWikiRunner with an LM + retriever, then runner.run(topic=..., do_research=True, do_generate_article=True)",
            "env": ["OPENAI_API_KEY", "YDC_API_KEY"],
            "env_optional": ["BING_SEARCH_API_KEY"],
            "notes": "Wikipedia-style survey generator (library). Needs an LLM key AND a retriever key (YDC_API_KEY from You.com, or Bing). Runnable demos + Streamlit UI live in the cloned repo's examples/frontend.",
            "clone_for_ui": true,
            "ui": {
              "requirements": ["requirements.txt", "frontend/demo_light/requirements.txt"],
              "write_env": false,
              "run": ["streamlit", "run", "frontend/demo_light/storm.py"],
              "url": "http://localhost:8501",
              "notes": "Streamlit demo (frontend/demo_light). Enter your OpenAI + You.com (YDC) keys in the app's sidebar / secrets. Serves at :8501."
            }
          },
          {
            "id": "scholarly",
            "name": "scholarly",
            "category": "api-client",
            "kind": "python-lib",
            "repo": "https://github.com/scholarly-python-package/scholarly",
            "pip": ["scholarly"],
            "entry": null,
            "example": "python -c \"from scholarly import scholarly; print(next(scholarly.search_author('Steven A Cholewiak')))\"",
            "env": [],
            "env_optional": [],
            "notes": "Google Scholar client (library). No key, but Scholar blocks bare IPs on search_pubs/citedby — you typically need a proxy (built-in ProxyGenerator / ScraperAPI)."
          },
          {
            "id": "pyalex",
            "name": "pyalex",
            "category": "api-client",
            "kind": "python-lib",
            "repo": "https://github.com/J535D165/pyalex",
            "pip": ["pyalex"],
            "entry": null,
            "example": "python -c \"import pyalex; pyalex.config.email='you@example.com'; from pyalex import Works; print(Works()['W2741809807']['title'])\"",
            "env": [],
            "env_optional": ["OPENALEX_API_KEY"],
            "notes": "OpenAlex API client (library). Set pyalex.config.email for the polite pool. As of ~2026 OpenAlex asks for an API key for higher limits: pyalex.config.api_key='...'."
          },
          {
            "id": "arxiv-fetch",
            "name": "arXiv search & download",
            "category": "retrieval",
            "kind": "python-script",
            "repo": "https://github.com/lukasschwab/arxiv.py",
            "pip": ["arxiv"],
            "script": "fetch_arxiv.py",
            "example": "arxiv-fetch --query 'retrieval augmented generation' --max 10 --outdir ./papers",
            "env": [],
            "env_optional": [],
            "notes": "Bundled fetcher (uses the free 'arxiv' package). Searches arXiv and downloads PDFs + a manifest.json into --outdir. No API key. Query supports cat:, AND/OR. Feeds retrieval workflows into PaperQA2."
          },
          {
            "id": "openalex-fetch",
            "name": "OpenAlex search & download",
            "category": "retrieval",
            "kind": "python-script",
            "repo": "https://openalex.org",
            "pip": ["requests"],
            "script": "fetch_openalex.py",
            "example": "openalex-fetch --query 'active learning systematic review' --max 10 --outdir ./papers",
            "env": [],
            "env_optional": ["OPENALEX_API_KEY", "OPENALEX_MAILTO"],
            "notes": "Bundled fetcher over the OpenAlex API (240M+ works, all fields). Downloads the OA PDF when available, else saves title+abstract as .txt (PaperQA2 reads text too). No key needed for light use; set OPENALEX_API_KEY / OPENALEX_MAILTO for higher limits. Appends to manifest.json for multi-source runs."
          },
          {
            "id": "pubmed-fetch",
            "name": "PubMed search & abstracts",
            "category": "retrieval",
            "kind": "python-script",
            "repo": "https://www.ncbi.nlm.nih.gov/pmc/tools/developers/",
            "pip": ["requests"],
            "script": "fetch_pubmed.py",
            "example": "pubmed-fetch --query 'CRISPR off-target detection' --max 10 --outdir ./papers",
            "env": [],
            "env_optional": ["NCBI_API_KEY", "NCBI_EMAIL"],
            "notes": "Bundled fetcher over NCBI E-utilities. Saves each article's title+abstract as .txt (+ manifest) — PubMed rarely exposes downloadable PDFs. Best for biomedical corpora. No key needed for light use; set NCBI_API_KEY / NCBI_EMAIL to raise the rate limit."
          },
          {
            "id": "papers-fetch",
            "name": "Multi-source paper search (6 APIs, deduplicated)",
            "category": "retrieval",
            "kind": "python-script",
            "repo": "https://github.com/brycewang-stanford/lit-review-agent-tools",
            "pip": [],
            "script": "fetch_papers.py",
            "example": "papers-fetch --query 'retrieval augmented generation' --sources openalex,crossref,semanticscholar --max 20 --outdir ./papers",
            "env": [],
            "env_optional": ["OPENALEX_API_KEY", "OPENALEX_MAILTO", "S2_API_KEY", "NCBI_API_KEY", "NCBI_EMAIL", "CROSSREF_MAILTO"],
            "notes": "Bundled, standard-library only — no install, no API key. Searches any subset of OpenAlex, Crossref, Semantic Scholar, PubMed, Europe PMC and arXiv, merges the hits into one normalised record list deduplicated by DOI (add --dedup-titles to also fold preprint/published pairs), and writes results.json + one .txt per record + manifest.json. A failing source is reported and skipped, never fatal. Use it instead of arxiv-fetch/openalex-fetch/pubmed-fetch when you want coverage rather than one index."
          },
          {
            "id": "oa-resolve",
            "name": "Open-access full-text resolver",
            "category": "retrieval",
            "kind": "python-script",
            "repo": "https://github.com/brycewang-stanford/lit-review-agent-tools",
            "pip": [],
            "script": "resolve_oa.py",
            "example": "oa-resolve --from-json ./papers/results.json --outdir ./papers --email you@example.com",
            "env": [],
            "env_optional": ["UNPAYWALL_EMAIL", "CORE_API_KEY"],
            "notes": "Bundled, standard-library only. Turns DOIs into downloadable full text by walking Unpaywall -> OpenAlex -> Europe PMC OA XML -> arXiv -> CORE, stopping at the first hit. Writes PDFs/plain text plus oa_report.json recording how each DOI resolved, separating 'closed' (paywalled) from 'not-in-crossref' (no such record — usually an invented citation). Never bypasses a paywall. This is the step that turns an abstract-only corpus into a full-text one before PaperQA2."
          },
          {
            "id": "arxiv-mcp-server",
            "name": "arxiv-mcp-server",
            "category": "mcp-server",
            "kind": "mcp-server",
            "repo": "https://github.com/blazickjp/arxiv-mcp-server",
            "pip": [],
            "entry": null,
            "env": [],
            "env_optional": [],
            "notes": "MCP server (stdio) to search/download/analyze arXiv. Runs via uvx (no install needed). Registered in an MCP client, not queried directly.",
            "mcp": {
              "key": "arxiv",
              "launcher": "uvx",
              "package": "arxiv-mcp-server",
              "args_template": ["arxiv-mcp-server", "--storage-path", "{storage}"],
              "env": {}
            }
          },
          {
            "id": "paper-search-mcp",
            "name": "paper-search-mcp",
            "category": "mcp-server",
            "kind": "mcp-server",
            "repo": "https://github.com/openags/paper-search-mcp",
            "pip": ["paper-search-mcp"],
            "entry": null,
            "env": [],
            "env_optional": ["PAPER_SEARCH_MCP_UNPAYWALL_EMAIL"],
            "notes": "MCP server to search/download across 20+ sources (arXiv, PubMed, bioRxiv, S2, OpenAlex...). Set an Unpaywall email for OA lookups. Verify exact optional env-var names in the README before relying on them.",
            "mcp": {
              "key": "paper-search-mcp",
              "launcher": "venv-module",
              "module": "paper_search_mcp.server",
              "env": {"PAPER_SEARCH_MCP_UNPAYWALL_EMAIL": "your@email.com"}
            }
          },
          {
            "id": "zotero-mcp",
            "name": "zotero-mcp",
            "category": "mcp-server",
            "kind": "mcp-server",
            "repo": "https://github.com/54yyyu/zotero-mcp",
            "pip": ["zotero-mcp"],
            "entry": "zotero-mcp",
            "env": [],
            "env_optional": ["ZOTERO_API_KEY", "ZOTERO_LIBRARY_ID", "OPENAI_API_KEY"],
            "notes": "MCP server bridging your Zotero library. Easiest: after install run 'zotero-mcp setup' to auto-write the client config. Local mode (ZOTERO_LOCAL=true) needs the Zotero desktop app running with the local API enabled; Web API mode needs ZOTERO_API_KEY + ZOTERO_LIBRARY_ID.",
            "mcp": {
              "key": "zotero",
              "launcher": "venv-entry",
              "entry": "zotero-mcp",
              "env": {"ZOTERO_LOCAL": "true"}
            }
          }
        ]
      }
      
    • workflows.json 8.5 KB
      {
        "_meta": {
          "description": "Named multi-step pipelines for litrun.py. Each step invokes a python-cli tool from recipes.json with placeholder-substituted args. Placeholders: {input}, {question}, {workdir}. Steps run in order, sharing one {workdir} under ~/.lit-review-tools/workspace/runs/<id>.",
          "source": "https://github.com/brycewang-stanford/lit-review-agent-tools"
        },
        "workflows": [
          {
            "id": "pdf-to-markdown",
            "name": "PDF(s) → clean Markdown",
            "description": "Convert a PDF file or a folder of PDFs into LLM-ready Markdown with MinerU.",
            "params": {"input": "path to a PDF file or a folder of PDFs"},
            "steps": [
              {"tool": "mineru", "args": ["-p", "{input}", "-o", "{workdir}/markdown", "-b", "pipeline"]}
            ],
            "outputs": "Markdown/JSON under {workdir}/markdown"
          },
          {
            "id": "pdf-corpus-qa",
            "name": "PDF corpus → cited answer",
            "description": "Ask a question over a folder of PDFs with PaperQA2; every claim comes back with a citation.",
            "params": {
              "input": "folder containing the PDFs to search",
              "question": "the question to answer over the corpus"
            },
            "steps": [
              {"tool": "paper-qa", "cwd": "{input}", "args": ["ask", "{question}"]}
            ],
            "outputs": "A citation-backed answer printed to stdout (PaperQA2 caches its index in {input})."
          },
          {
            "id": "pdf-md-then-qa",
            "name": "PDFs → Markdown + cited answer",
            "description": "Two-step: convert PDFs to Markdown (MinerU) for a clean artifact, then answer a question over the same corpus with PaperQA2.",
            "params": {
              "input": "folder containing the PDFs",
              "question": "the question to answer over the corpus"
            },
            "steps": [
              {"tool": "mineru", "args": ["-p", "{input}", "-o", "{workdir}/markdown", "-b", "pipeline"]},
              {"tool": "paper-qa", "cwd": "{input}", "args": ["ask", "{question}"]}
            ],
            "outputs": "Markdown under {workdir}/markdown + a citation-backed answer on stdout."
          },
          {
            "id": "topic-to-pdfs",
            "name": "Topic → arXiv PDFs",
            "description": "Search arXiv for a query and download the top-N PDFs (+ manifest) into a folder.",
            "params": {
              "query": "arXiv search query",
              "max": "how many PDFs to download (default 10)"
            },
            "defaults": {"max": "10"},
            "steps": [
              {"tool": "arxiv-fetch", "args": ["--query", "{query}", "--max", "{max}", "--outdir", "{workdir}/pdfs"]}
            ],
            "outputs": "PDFs + manifest.json under {workdir}/pdfs"
          },
          {
            "id": "topic-to-review",
            "name": "Topic → retrieved corpus → cited answer",
            "description": "End-to-end lightweight review: search arXiv, download the top-N PDFs, then answer a question over them with PaperQA2 (citation-backed). No MCP client needed.",
            "params": {
              "query": "arXiv search query defining the corpus",
              "question": "the question to answer over the retrieved papers",
              "max": "how many PDFs to retrieve (default 10)"
            },
            "defaults": {"max": "10"},
            "steps": [
              {"tool": "arxiv-fetch", "args": ["--query", "{query}", "--max", "{max}", "--outdir", "{workdir}/pdfs"]},
              {"tool": "paper-qa", "cwd": "{workdir}/pdfs", "args": ["ask", "{question}"]}
            ],
            "outputs": "Retrieved PDFs under {workdir}/pdfs + a citation-backed answer on stdout."
          },
          {
            "id": "topic-to-review-multi",
            "name": "Topic → multi-source corpus → cited answer",
            "description": "Retrieve from BOTH arXiv and OpenAlex into one corpus, then answer a question over it with PaperQA2. Broader coverage than a single source; resilient if one source is rate-limited.",
            "params": {
              "query": "search query defining the corpus",
              "question": "the question to answer over the retrieved papers",
              "max": "how many items per source (default 8)"
            },
            "defaults": {"max": "8"},
            "steps": [
              {"tool": "arxiv-fetch", "args": ["--query", "{query}", "--max", "{max}", "--outdir", "{workdir}/corpus"]},
              {"tool": "openalex-fetch", "args": ["--query", "{query}", "--max", "{max}", "--outdir", "{workdir}/corpus"]},
              {"tool": "paper-qa", "cwd": "{workdir}/corpus", "args": ["ask", "{question}"]}
            ],
            "outputs": "A merged arXiv+OpenAlex corpus under {workdir}/corpus + a citation-backed answer on stdout."
          },
          {
            "id": "topic-to-related-work",
            "name": "Topic → related-work paragraph",
            "description": "Retrieve papers (arXiv + OpenAlex), then have PaperQA2 draft a citation-backed related-work / literature-review paragraph synthesizing themes, methods, and gaps.",
            "params": {
              "query": "the topic / search query for the related-work section",
              "max": "how many items per source (default 10)"
            },
            "defaults": {"max": "10"},
            "steps": [
              {"tool": "arxiv-fetch", "args": ["--query", "{query}", "--max", "{max}", "--outdir", "{workdir}/corpus"]},
              {"tool": "openalex-fetch", "args": ["--query", "{query}", "--max", "{max}", "--outdir", "{workdir}/corpus"]},
              {"tool": "paper-qa", "cwd": "{workdir}/corpus", "args": ["ask", "Write a ~250-word related-work / literature-review paragraph on '{query}'. Synthesize the main themes, methods, and open gaps across these papers, and support every claim with an inline citation. Write in an academic tone suitable for a paper's Related Work section."]}
            ],
            "outputs": "A cited related-work paragraph on stdout, grounded in {workdir}/corpus."
          },
          {
            "id": "topic-to-corpus",
            "name": "Topic → deduplicated multi-source corpus",
            "description": "Search six scholarly APIs at once (OpenAlex, Crossref, Semantic Scholar, PubMed, Europe PMC, arXiv), deduplicate by DOI, and save one .txt per unique paper. No install, no API key — the keyless starting point for any review.",
            "params": {
              "query": "search query defining the corpus",
              "max": "how many hits per source before dedup (default 15)"
            },
            "defaults": {"max": "15"},
            "steps": [
              {"tool": "papers-fetch", "args": ["--query", "{query}", "--sources", "openalex,crossref,semanticscholar,pubmed,europepmc,arxiv", "--max", "{max}", "--dedup-titles", "--outdir", "{workdir}/corpus"]}
            ],
            "outputs": "results.json (normalised records + per-source counts) and one .txt per unique paper under {workdir}/corpus"
          },
          {
            "id": "topic-to-fulltext",
            "name": "Topic → multi-source search → open-access full text",
            "description": "Search six APIs, deduplicate, then resolve every DOI to a legally downloadable full text via Unpaywall → OpenAlex → Europe PMC → arXiv → CORE. Turns an abstract-only corpus into a real full-text one. Keyless.",
            "params": {
              "query": "search query defining the corpus",
              "max": "how many hits per source before dedup (default 15)"
            },
            "defaults": {"max": "15"},
            "steps": [
              {"tool": "papers-fetch", "args": ["--query", "{query}", "--sources", "openalex,crossref,semanticscholar,pubmed,europepmc,arxiv", "--max", "{max}", "--dedup-titles", "--outdir", "{workdir}/corpus"]},
              {"tool": "oa-resolve", "args": ["--from-json", "{workdir}/corpus/results.json", "--outdir", "{workdir}/corpus", "--skip-existing"]}
            ],
            "outputs": "PDFs / full-text .txt + oa_report.json (how each DOI resolved, including the closed ones) under {workdir}/corpus"
          },
          {
            "id": "topic-to-fulltext-review",
            "name": "Topic → full-text corpus → cited answer",
            "description": "The strongest end-to-end path: six-source search → dedup → open-access full-text resolution → PaperQA2 answers over the full texts rather than the abstracts. Needs OPENAI_API_KEY for the final step only.",
            "params": {
              "query": "search query defining the corpus",
              "question": "the question to answer over the retrieved papers",
              "max": "how many hits per source before dedup (default 12)"
            },
            "defaults": {"max": "12"},
            "steps": [
              {"tool": "papers-fetch", "args": ["--query", "{query}", "--sources", "openalex,crossref,semanticscholar,pubmed,europepmc,arxiv", "--max", "{max}", "--dedup-titles", "--outdir", "{workdir}/corpus"]},
              {"tool": "oa-resolve", "args": ["--from-json", "{workdir}/corpus/results.json", "--outdir", "{workdir}/corpus", "--skip-existing"]},
              {"tool": "paper-qa", "cwd": "{workdir}/corpus", "args": ["ask", "{question}"]}
            ],
            "outputs": "A full-text corpus under {workdir}/corpus + a citation-backed answer on stdout."
          }
        ]
      }
      
  • reference
    • apis
      • arxiv.md 2.6 KB
        # arXiv
        
        2.4M+ preprints in physics, maths, CS, quantitative biology, statistics, economics —
        and, uniquely in this set, **every record has a free PDF**. No key.
        Base URL `http://export.arxiv.org/api/query`.
        
        Used by `fetch_papers.py` (search), `fetch_arxiv.py` (search + download) and
        `resolve_oa.py` (last-resort PDF for anything with an arXiv id).
        
        ## Limits
        
        One request per **3 seconds**, `max_results` ≤ 2000 per call (use `start` to page, and
        stay under ~30k results per query). arXiv publishes and enforces this; exceeding it
        gets your IP throttled. Bulk downloading of PDFs is against their terms — for whole-corpus
        work they publish an S3 bucket instead.
        
        ## Query
        
        ```
        GET /api/query?search_query=all:transformer&start=0&max_results=20&sortBy=relevance
        ```
        
        | Prefix | Field |
        |---|---|
        | `all:` | everything |
        | `ti:` `abs:` `au:` | title, abstract, author |
        | `cat:` | category, e.g. `cat:cs.CL`, `cat:stat.ME` |
        | `id_list=2103.15348` | fetch specific ids instead of searching |
        
        Boolean `AND` / `OR` / `ANDNOT`, quotes for phrases, parentheses for grouping — all
        must be URL-encoded. `sortBy`: `relevance` · `submittedDate` · `lastUpdatedDate`.
        
        **There is no date filter.** The legacy API cannot express "since 2023"; either sort by
        `submittedDate` and take the head, or over-fetch and filter client-side (what
        `fetch_papers.py --from-year` does).
        
        ## Response — Atom XML, not JSON
        
        ```xml
        <entry>
          <id>http://arxiv.org/abs/2103.15348v2</id>
          <published>2021-03-29T…</published><updated>…</updated>
          <title>…</title><summary>…</summary>
          <author><name>…</name></author>
          <arxiv:doi>10.1145/…</arxiv:doi>
          <arxiv:journal_ref>ACL 2021</arxiv:journal_ref>
          <arxiv:primary_category term="cs.CL"/>
          <link title="pdf" href="http://arxiv.org/pdf/2103.15348v2"/>
        </entry>
        ```
        
        Namespaces: `{"a": "http://www.w3.org/2005/Atom", "arxiv": "http://arxiv.org/schemas/atom"}`.
        `title` and `summary` arrive with hard line wrapping — collapse whitespace or every
        title in your manifest carries newlines. The id ends in a **version suffix** (`v2`);
        `https://arxiv.org/pdf/{id}` works with or without it.
        
        `arxiv:doi` is present only once the preprint is published — most entries have none,
        which is why arXiv records dedupe by title rather than DOI in `fetch_papers.py`, and
        why `--dedup-titles` exists to fold the preprint into its published twin.
        
        ## Good and bad at
        
        Good: guaranteed full text, fast, keyless, strong CS/physics/ML coverage, category
        filters that actually work.
        Bad: no peer review (a hit is not evidence), no date filter, no citation counts, and
        nothing biomedical or social-scientific worth relying on.
        
      • crossref.md 2.8 KB
        # Crossref
        
        The DOI registration agency's own metadata: ~160M records. This is the authority on
        "does this DOI exist and what is it", plus journal, funder, license and reference-list
        metadata. No key. Base URL `https://api.crossref.org`.
        
        Used by `fetch_papers.py` (search) and by recipe 03, which verifies citations against it.
        
        ## Auth & limits
        
        - Public pool ~5 req/s. Adding `?mailto=you@example.com` (`CROSSREF_MAILTO`) puts you
          in the **polite pool** with roughly double the allowance and human contact if you
          misbehave. A descriptive `User-Agent` matters here more than on other APIs.
        
        ## Endpoints
        
        ```
        GET /works/{doi}                          one record — the canonical DOI check
        GET /works?query.bibliographic=…&rows=20  search (whole-reference matching)
        GET /works?query.title=…&query.author=…   fielded search
        GET /journals/{issn}/works                everything in a journal
        GET /members/{id}/works | /funders/{id}/works
        ```
        
        | Param | Notes |
        |---|---|
        | `query.bibliographic` | best field for "here is a reference string, find it" |
        | `filter` | `from-pub-date:2023-01-01`, `type:journal-article`, `has-abstract:true`, `is-update:true` (retraction/correction notices), `license.url:…` |
        | `select` | trims the response — **list endpoints only** |
        | `rows` / `offset` | ≤1000 rows; use `cursor=*` beyond 10k |
        | `sort` / `order` | `relevance`, `published`, `is-referenced-by-count` |
        
        ## The 400 that wastes an hour
        
        `select` is accepted on `/works?…` but **rejected on `/works/{doi}`**:
        
        ```
        GET /works?query=prisma&rows=1&select=DOI,title   → 200
        GET /works/10.1136/bmj.n160?select=DOI            → 400
        ```
        
        If a single-DOI lookup 400s, drop `select` before concluding the DOI is bad.
        
        ## Response shape
        
        ```json
        {"status":"ok","message":{"DOI":"10.1136/bmj.n160","title":["PRISMA 2020 …"],
         "container-title":["BMJ"],"issued":{"date-parts":[[2021,3,29]]},
         "author":[{"given":"Matthew J","family":"Page","ORCID":"…"}],
         "is-referenced-by-count":9781,"type":"journal-article",
         "abstract":"<jats:p>…</jats:p>","reference":[{"DOI":"…","unstructured":"…"}],
         "link":[{"URL":"…","content-type":"application/pdf"}],
         "update-to":[{"type":"retraction","DOI":"…"}]}
        ```
        
        - `title` and `container-title` are **arrays**; take `[0]`.
        - `abstract` is JATS XML when present at all (many publishers deposit none) — strip tags.
        - `reference` gives you the paper's own bibliography, which is how you walk citations
          backwards without Semantic Scholar.
        - `update-to` / `filter=is-update:true` surfaces retraction and correction notices.
        
        ## Good and bad at
        
        Good: DOI truth, reference strings → records, funder/license metadata, retraction notices.
        Bad: topical relevance search (it matches strings, not meaning), abstract coverage,
        and it has no notion of open-access location — pair it with Unpaywall or OpenAlex.
        
      • europepmc.md 2.8 KB
        # Europe PMC
        
        The most underrated API in this set. It indexes PubMed **plus** preprints, patents,
        agricultural and theses records, answers in JSON (no XML dance), needs no key, and
        serves open-access **full text** from one endpoint.
        Base URL `https://www.ebi.ac.uk/europepmc/webservices/rest`.
        
        Used by `fetch_papers.py` (search) and `resolve_oa.py` (OA full-text step).
        
        ## Endpoints
        
        ```
        GET /search?query=…&format=json&pageSize=25&resultType=core
        GET /search?query=…&cursorMark=*                 deep pagination (follow nextCursorMark)
        GET /{PMCID}/fullTextXML                         OA full text as JATS
        GET /{source}/{id}/textMinedTerms | /citations | /references
        ```
        
        `resultType`: `idlist` (ids only) · `lite` (default, no abstract) · `core`
        (**abstract, author list, journal, full-text URLs — what you almost always want**).
        
        ## Query language
        
        | Want | Query |
        |---|---|
        | Preprints only | `(machine learning) AND SRC:PPR` — 4,642 hits for that example |
        | Open access only | `… AND OPEN_ACCESS:Y` |
        | Has full text in PMC | `… AND HAS_FT:Y` |
        | Date range | `… AND (FIRST_PDATE:[2022-01-01 TO 3000-01-01])` |
        | By DOI | `DOI:"10.1136/bmj.n160"` |
        | Fielded | `TITLE:"…"`, `AUTH:"Page MJ"`, `JOURNAL:"BMJ"`, `MESH:"Neoplasms"` |
        
        `SRC:PPR` is the practical answer to "search preprints by keyword" — bioRxiv and
        medRxiv's own APIs cannot do it (see [`open-access.md`](open-access.md)).
        
        ## Response shape
        
        ```json
        {"hitCount":4642,"nextCursorMark":"…",
         "resultList":{"result":[{"id":"33781993","source":"MED","pmid":"33781993",
           "pmcid":"PMC8005925","doi":"10.1136/bmj.n160","title":"…","abstractText":"…",
           "authorList":{"author":[{"fullName":"Page MJ"}]},
           "journalInfo":{"journal":{"title":"BMJ"}},"pubYear":"2021",
           "isOpenAccess":"Y","citedByCount":9781,
           "fullTextUrlList":{"fullTextUrl":[{"documentStyle":"pdf","availabilityCode":"OA","url":"…"}]}}]}}
        ```
        
        `isOpenAccess` is the string `"Y"`/`"N"`, not a boolean. `source` tells you the
        sub-corpus: `MED` (PubMed), `PPR` (preprint), `PMC`, `PAT`, `AGR`, `CTX`.
        `abstractText` may carry inline HTML — strip tags.
        
        ## Full text
        
        ```
        GET /PMC8005925/fullTextXML
        ```
        
        Returns JATS for the **OA subset only** (`isOpenAccess:"Y"`); otherwise 404. Flattening
        the whole document with a naive `itertext()` prepends a pile of bibliographic tokens —
        extract `article-title` + `abstract` + `body` instead, which is what `jats_to_text()`
        in `resolve_oa.py` does. A 200 response under ~2 KB is a stub, not an article.
        
        ## Good and bad at
        
        Good: one keyless JSON call for metadata *and* full text; preprint keyword search;
        citation counts; the widest biomedical net available without credentials.
        Bad: coverage outside life sciences; relevance ranking skews recent, so pair it with
        OpenAlex when you want the classic papers in a field.
        
      • open-access.md 4 KB
        # Getting the actual file: Unpaywall, CORE, bioRxiv/medRxiv
        
        Metadata tells you a paper exists. These APIs tell you whether you may read it.
        `resolve_oa.py` chains them; this file documents each one and the order, so you can
        debug a DOI that "should" be open.
        
        ## The chain, and why this order
        
        ```
        1. Unpaywall    every publisher, every field — but needs an email
        2. OpenAlex     same underlying OA data, keyless, sometimes a landing page not a PDF
        3. Europe PMC   biomedical OA full text as XML — no PDF needed
        4. arXiv        anything with an arXiv id, guaranteed PDF
        5. CORE         repository aggregator, catches institutional copies — needs a key
        ```
        
        Highest precision first, widest net last. Recorded result on a 47-paper multi-source
        corpus with **no keys at all** (Unpaywall and CORE skipped): 33/47 resolved — 16 via
        OpenAlex, 10 via arXiv, 7 via Europe PMC. See
        [`recipes/06-multi-source-search`](../../../../recipes/06-multi-source-search/).
        
        **Nothing here bypasses a paywall.** A closed paper stays closed and is reported as
        `closed` in `oa_report.json`. Do not route around that with Sci-Hub or scraped mirrors:
        it is unlawful in most jurisdictions and gets institutions blocked. Interlibrary loan
        and author-request are the legitimate paths, and both need a human.
        
        ## Unpaywall
        
        ```
        GET https://api.unpaywall.org/v2/{doi}?email=you@example.com
        ```
        
        - The `email` parameter is **mandatory** — without it you get HTTP 422, not a warning.
          Never invent one; ask the user, or run without this step (`resolve_oa.py` does).
        - 100k calls/day. Batch downloads are offered as a data dump instead.
        - Read `is_oa`, then `best_oa_location.url_for_pdf` (may be null when only a landing
          page exists), `oa_status` (`gold`/`green`/`hybrid`/`bronze`/`closed`), and
          `.version` — `submittedVersion` means you got the preprint, not the paper of record.
          Quote page numbers from a `publishedVersion` only.
        
        ## CORE
        
        ```
        POST https://api.core.ac.uk/v3/search/works
        Authorization: Bearer $CORE_API_KEY      body: {"q":"doi:\"10.…\"","limit":1}
        ```
        
        37M+ full texts harvested from institutional and subject repositories — the best
        source for green OA copies that Unpaywall's publisher-centric view misses, and for
        theses. Requires free registration; without `CORE_API_KEY` this step is skipped.
        Use `downloadUrl` from a result, and verify the bytes really are a PDF.
        
        ## bioRxiv / medRxiv
        
        ```
        GET https://api.biorxiv.org/details/biorxiv/10.1101/2020.01.30.927871
        GET https://api.biorxiv.org/details/medrxiv/2023-01-01/2023-01-31/0     paged by date
        GET https://api.biorxiv.org/pubs/biorxiv/{doi}                          published version
        ```
        
        **There is no keyword search.** These APIs only browse by DOI or by date window; the
        `collection` array comes back with `title`, `authors`, `date`, `category`, `jatsxml`.
        To search preprints by topic, use Europe PMC's `SRC:PPR` filter (see
        [`europepmc.md`](europepmc.md)) or OpenAlex `type:preprint` — then come back here for
        the version history, or to `/pubs/` to find out whether the preprint was ever published.
        
        That last check is the one people skip: citing a preprint whose published version
        contradicts it is a real and common failure of automated reviews.
        
        ## Verifying what you downloaded
        
        - A PDF starts with the bytes `%PDF-`. Publishers serve HTML "access denied" pages with
          HTTP 200 and `Content-Type: application/pdf` often enough that content-type is not
          evidence — check the magic bytes (`is_pdf()` in `resolve_oa.py`).
        - Europe PMC full text under ~2 KB is a stub record, not an article.
        - A DOI that resolves nowhere is two different findings. `resolve_oa.py` asks Crossref
          before giving up: `closed` means the paper exists behind a paywall; `not-in-crossref`
          means no Crossref record exists at all — an invented citation, a typo, or a DOI from
          another registry (DataCite datasets, some preprint servers). Never report the second
          as the first.
        - Record the resolver that produced each file. When a citation later looks wrong, the
          first question is always "was that the published version or the preprint?", and
          `oa_report.json` answers it.
        
      • openalex.md 3.2 KB
        # OpenAlex
        
        Broadest of the free indexes: ~250M works across every field, with authors,
        institutions, sources, topics, citation counts and OA locations. No key required.
        Base URL `https://api.openalex.org`.
        
        Used by `fetch_papers.py` (search), `resolve_oa.py` (step 2 of the OA chain) and
        `fetch_openalex.py`.
        
        ## Auth & limits
        
        - Keyless works. `?api_key=…` (`OPENALEX_API_KEY`) or the legacy polite pool
          `?mailto=you@example.com` (`OPENALEX_MAILTO`) raise the limit.
        - 100 req/s ceiling. Single-entity lookups by id/DOI are unmetered; list and search
          queries draw on a daily free allowance.
        
        ## Endpoints
        
        ```
        GET /works/{id}                     W2741809807 | doi:10.1136/bmj.n160 | pmid:33781993
        GET /works?search=…&per_page=25     full-text search over title+abstract+fulltext
        GET /works?filter=…                 comma-separated field:value, all ANDed
        GET /authors?search= | /sources?search= | /institutions?search= | /topics/{id}
        ```
        
        Useful parameters:
        
        | Param | Notes |
        |---|---|
        | `search` | boolean `AND`/`OR`/`NOT` (uppercase), `"phrase"~5`, `wildcar*`, `fuzzy~1` |
        | `search.semantic` | embedding search, beta — 1 req/s, ≤50 results |
        | `filter` | `from_publication_date:2023-01-01`, `publication_year:2024`, `type:article`, `is_oa:true`, `cited_by_count:>100`, `authorships.author.id:A…`, `institutions.country_code:us`, `doi:…`. Operators `>` `<` `!` `\|` |
        | `sort` | `cited_by_count:desc`, `publication_date:desc`, `relevance_score:desc` |
        | `select` | trim the response — works on both list and single-entity endpoints |
        | `group_by` | aggregate counts, e.g. `group_by=publication_year` for a topic's trend line |
        | `cursor=*` | deep pagination past 10,000; follow `meta.next_cursor` until null |
        
        ## Response shape (fields worth reading)
        
        ```json
        {"id":"https://openalex.org/W2741809807","doi":"https://doi.org/10.7717/peerj.4375",
         "display_name":"…","publication_year":2018,"type":"article","is_retracted":false,
         "cited_by_count":1169,
         "open_access":{"is_oa":true,"oa_status":"gold","oa_url":"…"},
         "best_oa_location":{"pdf_url":"…","license":"cc-by","version":"publishedVersion"},
         "primary_location":{"source":{"display_name":"PeerJ","issn_l":"2167-8359"}},
         "authorships":[{"author":{"id":"…","display_name":"…"},"institutions":[…]}],
         "abstract_inverted_index":{"Despite":[0],"growing":[1]},
         "referenced_works":["https://openalex.org/W…"],
         "ids":{"doi":"…","pmid":"…","arxiv":"…"}}
        ```
        
        Two gotchas that bite every time:
        
        - **`title` does not exist** on the work object — the field is `display_name`.
        - **Abstracts are inverted indexes**, `{word: [positions]}`. Rebuild by placing each
          word at each position and joining in index order (`_inverted()` in `fetch_papers.py`).
        - `is_retracted` is a free retraction check — cheaper than a dedicated service, and
          worth reading before you cite anything (see recipe 03).
        
        ## What it is good and bad at
        
        Good: coverage, OA links, citation counts, institution/author disambiguation, trend
        aggregation via `group_by`.
        Bad: relevance ranking is weaker than Semantic Scholar's for conceptual queries, and
        `open_access.oa_url` is often a **landing page, not a PDF** — always check the bytes
        start with `%PDF-` before saving one (`resolve_oa.py` does).
        
      • pubmed-pmc.md 3.6 KB
        # PubMed & PMC (NCBI E-utilities)
        
        37M+ biomedical citations with MeSH indexing, plus the PMC full-text archive. The
        MeSH vocabulary is the reason PubMed still beats general indexes for clinical
        questions: it is human-curated subject indexing, not string matching.
        Base URL `https://eutils.ncbi.nlm.nih.gov/entrez/eutils`.
        
        Used by `fetch_papers.py` and `fetch_pubmed.py`.
        
        ## Auth & limits
        
        3 req/s anonymous, 10 with `NCBI_API_KEY` (`&api_key=…`). Add `&email=` and `&tool=`
        so NCBI can contact you instead of blocking you. Sleep ~0.34 s between calls when
        keyless — `fetch_papers.py` does.
        
        ## The two-step dance
        
        ```
        GET /esearch.fcgi?db=pubmed&term=…&retmax=20&retmode=json&sort=relevance   → PMIDs
        GET /efetch.fcgi?db=pubmed&id=1,2,3&retmode=xml                            → full records
        GET /esummary.fcgi?db=pubmed&id=…&retmode=json                             → light metadata
        GET /elink.fcgi?dbfrom=pubmed&db=pmc&id=…                                  → PMID → PMCID
        GET /efetch.fcgi?db=pmc&id=PMC8005925&retmode=xml                          → JATS full text
        ```
        
        `esummary` returns JSON but no abstract. **`efetch` with `retmode=xml` is the only way
        to get abstracts**, and it returns XML even though the rest of E-utilities speaks JSON.
        
        Query syntax is the same as the PubMed website:
        
        ```
        ("systematic review"[Publication Type]) AND (machine learning[Title/Abstract])
        AND ("2022"[Date - Publication] : "3000"[Date - Publication])
        covid-19[MeSH Terms] AND humans[MeSH Terms] AND english[Language]
        ```
        
        Date filtering via parameters: `&datetype=pdat&mindate=2022&maxdate=3000`.
        
        ## Parsing the XML
        
        Fields worth pulling from each `PubmedArticle`:
        
        | XPath | Field |
        |---|---|
        | `.//PMID` | PMID |
        | `.//ArticleTitle` | title — **may contain inline `<i>`/`<sub>`**, so use `itertext()`, not `.text` |
        | `.//AbstractText` | abstract, often **several elements** with `Label="METHODS"` etc. — join them |
        | `./PubmedData/ArticleIdList/ArticleId[@IdType='doi']` | DOI — **scope it exactly like this** |
        | `.//PubDate/Year` or `.//PubDate/MedlineDate` | year (MedlineDate is free text like `2023 Jan-Feb`) |
        | `.//Author/ForeName` + `LastName` | authors |
        | `.//Journal/Title` | venue |
        | `.//PublicationType` | `Retracted Publication`, `Retraction of Publication` live here |
        
        Taking only `AbstractText[0]` silently truncates every structured abstract to its
        Background section — a classic quiet data-loss bug.
        
        The DOI path is the sharper trap. A `PubmedArticle` embeds the article's **entire
        reference list**, and every reference carries its own `<ArticleId IdType="doi">`. A
        `.//ArticleId` sweep therefore returns a *cited* paper's DOI — in one recorded run,
        a stroke-imaging review came back carrying an ICPSR **dataset** DOI from its own
        bibliography. The record looks perfect and points at the wrong object. Scope the
        lookup to `./PubmedData/ArticleIdList`, and fall back to
        `.//Article/ELocationID[@EIdType='doi']`. Written up in
        [`recipes/06`](../../../../recipes/06-multi-source-search/).
        
        ## ID conversion
        
        The old `/pmc/utils/idconv/v1.0/` path now 301-redirects; the live endpoint is
        
        ```
        GET https://pmc.ncbi.nlm.nih.gov/tools/idconv/api/v1/articles/?ids=10.1136/bmj.n160&format=json
        → {"records":[{"doi":"10.1136/bmj.n160","pmcid":"PMC8005925","pmid":33781993}]}
        ```
        
        Follow redirects (`curl -L`) or you will parse an HTML 301 page as JSON.
        
        ## Good and bad at
        
        Good: MeSH-indexed retrieval, publication-type filters (RCT, meta-analysis, retraction),
        clinical coverage, and PMC full text for the OA subset.
        Bad: anything non-biomedical, and PDFs — PubMed exposes almost none. For full text
        prefer Europe PMC (same corpus, one JSON call, no XML dance).
        
      • README.md 5.1 KB
        # Scholarly APIs — route the question to the right index
        
        The rest of this skill installs tools. This directory does the opposite: it lets you
        answer a literature question **with one HTTP call and nothing installed**, and it
        documents the exact endpoints the bundled `fetch_papers.py` / `resolve_oa.py` scripts
        use, so you can debug or extend them.
        
        Reach for this when the user wants *a specific paper, a DOI resolved, an author's
        output, an OA PDF, or a citation graph* — a venv is overkill for that. Reach for the
        launcher when the user wants a corpus, a screening run, or an extraction pipeline.
        
        > Layout modelled on the excellent [`paper-lookup`](https://github.com/K-Dense-AI) skill
        > by K-Dense Inc. The endpoints, quirks and failure modes here were re-verified against
        > live APIs on 2026-08-10 and annotated with what this repo's scripts actually hit.
        
        ## Pick a database
        
        | The user wants… | Query this | Then |
        |---|---|---|
        | Papers on a topic, any field | OpenAlex | Semantic Scholar for citation context |
        | Papers on a biomedical topic | PubMed | Europe PMC for the full text |
        | Full text of a biomedical paper | Europe PMC (OA subset) | Unpaywall for non-biomed |
        | Physics / maths / CS preprints | arXiv | Semantic Scholar for the published version |
        | Biology / health preprints | Europe PMC (`SRC:PPR`) | bioRxiv/medRxiv APIs *by DOI or date only* |
        | One paper by DOI | Crossref | OpenAlex for citations, Unpaywall for the PDF |
        | An open-access PDF for a DOI | Unpaywall → OpenAlex → Europe PMC | `resolve_oa.py` already chains these |
        | Who cites whom | Semantic Scholar | OpenAlex `referenced_works` |
        | An author's publications | OpenAlex | Semantic Scholar author endpoint |
        | Journal / funder / license metadata | Crossref | OpenAlex `sources` |
        | PMID ↔ PMCID ↔ DOI | NCBI ID Converter | Europe PMC search |
        
        **One query, several indexes.** Relevance ranking differs enough between these APIs
        that the same query returns largely disjoint top-10s — in the run recorded in
        [`recipes/06-multi-source-search`](../../../../recipes/06-multi-source-search/), six
        sources returned 50 hits of which only 3 were duplicates. If coverage matters, do not
        trust one index; run `fetch_papers.py` and let it merge them.
        
        ## Identifier formats
        
        | Identifier | Shape | Example | Understood by |
        |---|---|---|---|
        | DOI | `10.xxxx/…` | `10.1136/bmj.n160` | everything |
        | PMID | digits | `33781993` | PubMed, Europe PMC, S2 (`PMID:`) |
        | PMCID | `PMC` + digits | `PMC8005925` | Europe PMC, PMC |
        | arXiv id | `YYMM.NNNNN` | `2103.15348` | arXiv, S2 (`ARXIV:`), OpenAlex |
        | OpenAlex id | `W` + digits | `W2741809807` | OpenAlex |
        | S2 id | 40-char hex | `649def34f8be…` | Semantic Scholar |
        | ORCID | `0000-…-…-…` | `0000-0001-6187-6610` | OpenAlex, Crossref |
        
        Cross-lookup prefixes: OpenAlex takes `/works/doi:10.…` and `/works/pmid:…`;
        Semantic Scholar takes `DOI:…`, `PMID:…`, `PMCID:…`, `ARXIV:…`.
        A DOI that 404s in one index is usually alive in another — try before concluding it is fake.
        
        ## Keys: what is actually needed
        
        Everything below works with **no key**. Keys only buy rate limit.
        
        | API | Env var | Without it |
        |---|---|---|
        | OpenAlex | `OPENALEX_API_KEY` / `OPENALEX_MAILTO` | works; shared pool |
        | Crossref | `CROSSREF_MAILTO` | works at ~5 req/s; `mailto` doubles it |
        | Semantic Scholar | `S2_API_KEY` | **frequently HTTP 429** — the shared pool is often exhausted |
        | NCBI (PubMed) | `NCBI_API_KEY`, `NCBI_EMAIL` | 3 req/s instead of 10 |
        | Unpaywall | `UNPAYWALL_EMAIL` | **skipped entirely** — the API requires an email |
        | Europe PMC | — | no key exists |
        | arXiv | — | no key; 1 request / 3 s |
        | CORE | `CORE_API_KEY` | skipped; registration required |
        
        Read keys from the environment, then from `~/.lit-review-tools/.env`
        (`litrun.py env --set KEY=VALUE`). Never invent an email for Unpaywall — ask the user.
        
        ## Per-API references
        
        | API | File | Used by |
        |---|---|---|
        | OpenAlex | [`openalex.md`](openalex.md) | `fetch_papers.py`, `resolve_oa.py`, `fetch_openalex.py` |
        | Crossref | [`crossref.md`](crossref.md) | `fetch_papers.py`, recipe 03 (citation verification) |
        | Semantic Scholar | [`semantic-scholar.md`](semantic-scholar.md) | `fetch_papers.py` |
        | PubMed / PMC (NCBI) | [`pubmed-pmc.md`](pubmed-pmc.md) | `fetch_papers.py`, `fetch_pubmed.py` |
        | Europe PMC | [`europepmc.md`](europepmc.md) | `fetch_papers.py`, `resolve_oa.py` |
        | arXiv | [`arxiv.md`](arxiv.md) | `fetch_papers.py`, `fetch_arxiv.py` |
        | Unpaywall, CORE, bioRxiv/medRxiv | [`open-access.md`](open-access.md) | `resolve_oa.py` |
        
        ## Calling them
        
        Claude Code: `WebFetch`, or `curl` via Bash when you need the raw bytes or a POST.
        Always send a `User-Agent` that identifies you; several of these APIs throttle
        anonymous clients harder. If you get 429, wait ~3 s and retry **once**, then move on
        to another source rather than hammering.
        
        ## Reporting back
        
        State which APIs you queried and which returned nothing — a silent omission reads as
        "no such paper exists", which is a much stronger claim than "OpenAlex had no match".
        Quote DOIs verbatim, and if a DOI failed to resolve anywhere, say so rather than
        paraphrasing it into a plausible-looking citation.
        
      • semantic-scholar.md 2.5 KB
        # Semantic Scholar (S2 Graph API)
        
        ~200M papers with the best free **citation graph** — citations, references, influential
        citation counts, AI-generated TLDRs, and paper recommendations. Base URL
        `https://api.semanticscholar.org/graph/v1`.
        
        Used by `fetch_papers.py` (search). Its ranking is the most "semantic" of the free
        indexes, which is why it is in the default source set.
        
        ## Auth & limits — read this first
        
        The keyless pool is **shared across every anonymous client on the internet and is
        routinely exhausted**. In this repo's recorded runs, keyless `/paper/search` returned
        HTTP 429 on both attempts while single-paper lookups succeeded. Treat 429 as normal,
        not as a bug: retry once, then continue without this source.
        
        `S2_API_KEY` (free, request form on their site) goes in the `x-api-key` **header**,
        not the query string.
        
        ## Endpoints
        
        ```
        GET /paper/search?query=…&limit=20&fields=…       relevance search
        GET /paper/search/bulk?query=…                    up to 1000/page, no relevance ranking
        GET /paper/{id}?fields=…                          DOI:… | PMID:… | PMCID:… | ARXIV:… | CorpusId:… | 40-hex
        GET /paper/{id}/citations?fields=…                who cites this
        GET /paper/{id}/references?fields=…               what this cites
        GET /paper/{id}/recommendations
        GET /author/search?query=… | /author/{id}/papers
        POST /paper/batch  {"ids":[…]}                    up to 500 ids in one call
        ```
        
        `fields` is mandatory in practice — omit it and you get bare ids. Useful set:
        
        ```
        title,abstract,year,authors,venue,citationCount,influentialCitationCount,
        externalIds,openAccessPdf,isOpenAccess,url,tldr,fieldsOfStudy,publicationTypes
        ```
        
        `year=2020-` filters a range; `openAccessPdf` gives a direct PDF URL when one exists.
        
        ## Response shape
        
        ```json
        {"total":1234,"data":[{"paperId":"649def34…","title":"…","abstract":"…","year":2021,
         "venue":"BMJ","citationCount":9781,"influentialCitationCount":812,
         "externalIds":{"DOI":"10.1136/bmj.n160","PubMed":"33781993","ArXiv":null},
         "openAccessPdf":{"url":"…","status":"GOLD"},"tldr":{"text":"…"}}]}
        ```
        
        `abstract` is often `null` even when the paper has one — S2 cannot redistribute every
        publisher's text. Fall back to OpenAlex or Europe PMC for the abstract; that fallback
        is exactly what `fetch_papers.py`'s merge step does.
        
        ## Good and bad at
        
        Good: citation graph in both directions, `influentialCitationCount` (a better signal
        than raw counts for "what actually mattered"), TLDRs, recommendations, batch lookup.
        Bad: availability without a key, abstract coverage, non-English work.
        
    • catalog.md 15.9 KB
      # Full Catalog — AI Literature Review Tools
      
      66 open-source projects, organized by use case. ⭐ = editor's pick.
      
      > Generated by `scripts/build.py` from `data/tools.yaml`. Do not edit by hand —
      > edits here are overwritten and CI will reject them.
      
      Metadata refreshed from the GitHub API on **2026-09-21**.
      Health: 🟢 pushed within 90 days · 🟡 within a year · 🔴 over a year · 🗄️ archived.
      Licence `none` means the repository ships no licence file (all rights reserved);
      `CC-BY-NC*` forbids commercial use. Recommend accordingly.
      
      Source of truth: <https://github.com/brycewang-stanford/lit-review-agent-tools>
      
      ---
      
      ## 🌟 All-in-one Research Agents & Skills
      
      End-to-end `research → write → review → revise` solutions — mostly Claude Code / Codex skills.
      
      | Project | Stars | Health | Licence | Stage | Notes |
      |---|---|---|---|---|---|
      | ⭐ [academic-research-skills](https://github.com/Imbad0202/academic-research-skills) | ~48.9k | 🟢 | CC-BY-NC-4.0 | search, synthesize, cite-check, write, review | **The most popular project in this space.** A suite of Claude Code skills running a 10-stage pipeline (research→write→review→revise→finalize) with citation/claim "integrity gates," cross-checked against Semantic Scholar + OpenAlex + Crossref. Philosophy: *"AI is your copilot, not the pilot."* `/plugin install academic-research-skills` |
      | [academic-research-skills-codex](https://github.com/Imbad0202/academic-research-skills-codex) | ~11.3k | 🟢 | CC-BY-NC-4.0 | search, synthesize, cite-check, write, review | Codex-native sibling of the above, human-in-the-loop research flow |
      | [Research-Paper-Writing-Skills](https://github.com/Master-cai/Research-Paper-Writing-Skills) | ~7.0k | 🟢 | MIT | write | ML/CV/NLP paper-writing skill pack; works with Codex, Claude Code, and Gemini |
      | [claude-skills](https://github.com/alirezarezvani/claude-skills) | ~26.2k | 🟢 | MIT | search, synthesize, write | Large skill collection incl. a litreview / grants / deep-research stack across Claude Code / Codex / Gemini / Cursor |
      | [academic-paper-skills](https://github.com/lishix520/academic-paper-skills) | ~1.3k | 🟡 | MIT | write | Strategist (planning) + Composer (writing) skills with quality checkpoints |
      | [dr-claw](https://github.com/OpenLAIR/dr-claw) | ~1.1k | 🟢 | custom | search, read, synthesize, write | A "research IDE" with multiple AI-assistant personas |
      | [ScienceClaw](https://github.com/beita6969/ScienceClaw) | 905 | 🟡 | MIT | search, synthesize, write | Self-evolving AI research colleague, 285 skills, "zero hallucination" claim |
      | [qinyan-academic-skills](https://github.com/LeonChaoX/qinyan-academic-skills) | 913 | 🟢 | MIT | search, synthesize, write | Multilingual library of 182 installable AI-agent skills across disciplines |
      | [agent-research-skills](https://github.com/lingzhi227/agent-research-skills) | 351 | 🟡 | none | search, synthesize, cite-check | Claude Code skills for systematic literature review, incl. citation-validation scripts |
      | [medsci-skills](https://github.com/Aperivue/medsci-skills) | 313 | 🟢 | MIT | search, cite-check, write | Medical-research skills: search, reporting-guideline/citation checks, stats, figures, submission (by a physician-researcher) |
      
      ## 🔎 Deep Research & Auto Survey Generation
      
      Give it a topic; it searches and produces a cited report / survey / related-work section.
      
      | Project | Stars | Health | Licence | Stage | Notes |
      |---|---|---|---|---|---|
      | ⭐ [STORM](https://github.com/stanford-oval/storm) | ~31.5k | 🟡 | MIT | search, synthesize, write | Stanford OVAL; retrieval-grounded "pre-writing + writing" stages, produces Wikipedia-style long articles with citations; includes conversational Co-STORM |
      | ⭐ [gpt-researcher](https://github.com/assafelovic/gpt-researcher) | ~29.6k | 🟢 | Apache-2.0 | search, synthesize, write | Autonomous agent that runs deep research on any topic and outputs a cited report; general-purpose, not academic-only |
      | [deep-research](https://github.com/dzhng/deep-research) | ~19.7k | 🟡 | MIT | search, synthesize | Minimal iterative deep-research agent (search + scrape + LLM refinement); small and hackable |
      | [open_deep_research](https://github.com/langchain-ai/open_deep_research) | ~12.7k | 🗄️ | MIT | search, synthesize | LangChain's official open deep-research reference implementation |
      | [local-deep-research](https://github.com/LearningCircuit/local-deep-research) | ~9.1k | 🟢 | MIT | search, synthesize | Local/private deep-research; 10+ sources incl. arXiv & PubMed, fully local LLMs |
      | [open-deep-research](https://github.com/nickscamara/open-deep-research) | ~6.3k | 🔴 | Apache-2.0 | search, synthesize | Open deep-research clone reasoning over web data via Firecrawl |
      | [SurveyX](https://github.com/IAAR-Shanghai/SurveyX) | 990 | 🟡 | none | search, synthesize, write | Automated academic survey-paper generation from a topic |
      | [LitLLM](https://github.com/LitLLM/LitLLM) | 52 | 🟡 | Apache-2.0 | search, synthesize, write | Toolkit focused on scientific literature review; RAG + prompting to draft related-work fast (TMLR 2025) |
      | [opendraft](https://github.com/federicodeponte/opendraft) | 448 | 🟢 | MIT | write | Free & open-source AI paper writer; 19 agents collaborate to draft long papers |
      | [AutoSurveyGPT](https://github.com/a554b554/AutoSurveyGPT) | 156 | 🔴 | MIT | search, synthesize, write | Uses GPT to find & rank Google Scholar papers and auto-generate a survey |
      
      ## 🧪 Autonomous Science: idea → paper
      
      End-to-end "automated scientific discovery" — lit review + hypotheses + experiments + writing + self-review. The most ambitious category.
      
      | Project | Stars | Health | Licence | Stage | Notes |
      |---|---|---|---|---|---|
      | ⭐ [AI-Scientist](https://github.com/SakanaAI/AI-Scientist) | ~14.6k | 🟡 | custom | search, synthesize, write, review | Sakana AI; end-to-end automated discovery (lit review→experiments→writing→review); see also [v2](https://github.com/SakanaAI/AI-Scientist-v2) (~6.9k, agentic tree search, workshop-level) |
      | [AutoResearchClaw](https://github.com/aiming-lab/AutoResearchClaw) | ~14.5k | 🟢 | MIT | search, synthesize, cite-check, write, review | Self-evolving autonomous research: idea → conference-ready LaTeX paper (real lit from OpenAlex/S2/arXiv + sandboxed experiments + multi-agent peer review) |
      | [Agent-Laboratory](https://github.com/SamuelSchmidgall/AgentLaboratory) | ~5.9k | 🔴 | MIT | search, synthesize, write | End-to-end autonomous workflow: literature review → experimentation → report writing |
      | [SciAgentsDiscovery](https://github.com/lamm-mit/SciAgentsDiscovery) | 639 | 🔴 | Apache-2.0 | synthesize | Multi-agent (ontologist/scientist/critic) automated hypothesis & discovery system |
      | [Zochi](https://github.com/IntologyAI/Zochi) | 316 | 🟡 | MIT | search, synthesize, write, review | "Artificial scientist" doing end-to-end discovery to peer-reviewed publication |
      | [DeepInnovator](https://github.com/HKUDS/DeepInnovator) | 291 | 🟡 | MIT | synthesize | Autonomously generates research ideas, questions, testable hypotheses & experiment designs |
      
      ## 📚 Paper Q&A and RAG
      
      Grounded, **citation-backed** Q&A and extraction over a corpus of PDFs / papers.
      
      | Project | Stars | Health | Licence | Stage | Notes |
      |---|---|---|---|---|---|
      | ⭐ [paper-qa](https://github.com/Future-House/paper-qa) | ~9.2k | 🟢 | Apache-2.0 | read, extract, synthesize, cite-check | FutureHouse; high-accuracy RAG for scientific papers, answers **always cite sources**; PaperQA2 claims superhuman literature search |
      | [paperai](https://github.com/neuml/paperai) | ~1.8k | 🟢 | Apache-2.0 | search, read, synthesize | Semantic search + Q&A over medical & scientific papers |
      | [openpaper](https://github.com/khoj-ai/openpaper) | 481 | 🟢 | AGPL-3.0 | read, synthesize, cite-check | Research-library workbench: read/annotate papers + AI lit-review assistant with grounded citations |
      
      ## 🧮 Systematic Review & Screening
      
      For rigorous evidence-based / PRISMA reviews — screen thousands of abstracts efficiently.
      
      | Project | Stars | Health | Licence | Stage | Notes |
      |---|---|---|---|---|---|
      | ⭐ [ASReview](https://github.com/asreview/asreview) | ~1.0k | 🟢 | Apache-2.0 | screen | Active-learning screener for systematic reviews; interactively ranks papers to cut screening time; well-established in academia |
      | [LatteReview](https://github.com/PouriaRouzrokh/LatteReview) | 121 | 🟡 | CC-BY-NC-ND-4.0 | screen | Low-code Python package automating SR screening via AI agents (OpenAI/Gemini/Claude/Ollama) |
      | [prismAId](https://github.com/Open-and-Sustainable/prismAId) | 29 | 🟢 | AGPL-3.0 | screen, extract | Generative-AI, protocol-based systematic-review toolkit; no-code, replicable screening & extraction |
      | [prisma-review-tool](https://github.com/Black-Lights/prisma-review-tool) | niche | 🟢 | MIT | search, screen | PRISMA 2020 flow with AI-assisted screening via MCP (arXiv/OpenAlex/S2, no API keys) |
      
      ## 🔌 MCP Servers
      
      Bring literature capabilities into Claude / Cursor / Cline and other Model Context Protocol clients.
      
      | Project | Stars | Health | Licence | Stage | Notes |
      |---|---|---|---|---|---|
      | ⭐ [zotero-mcp](https://github.com/54yyyu/zotero-mcp) | ~5.1k | 🟢 | MIT | search, read | Connects a Zotero library (local + web API) to AI: semantic search, PDF full-text, citation analysis. The most popular Zotero MCP |
      | ⭐ [arxiv-mcp-server](https://github.com/blazickjp/arxiv-mcp-server) | ~3.2k | 🟢 | Apache-2.0 | search, extract, read | Search & analyze arXiv papers; downloads and converts PDFs to Markdown for LLM context; ships an `.mcpb` bundle |
      | [paper-search-mcp](https://github.com/openags/paper-search-mcp) | ~2.7k | 🟢 | MIT | search | Multi-source search/download across 20+ sources (arXiv, PubMed, bioRxiv, S2, OpenAlex, Crossref, CORE…) |
      | [PubMed-MCP-Server](https://github.com/JackKuo666/PubMed-MCP-Server) | 129 | 🔴 | MIT | search, read | Search, access, and analyze PubMed articles (metadata + deep analysis) |
      | [alex-mcp](https://github.com/drAbreu/alex-mcp) | 56 | 🔴 | MIT | search | OpenAlex MCP focused on author disambiguation and institution/work lookup |
      | [openalex-research-mcp](https://github.com/oksure/openalex-research-mcp) | 53 | 🟡 | MIT | search, cite-check | OpenAlex (240M+ works): citation analysis, research-trend tracking, collaboration networks |
      
      ## 🗂️ Reference & Knowledge Management (Zotero / Obsidian)
      
      Embed AI into the reference-manager / note-taking workflow you already use.
      
      | Project | Stars | Health | Licence | Stage | Notes |
      |---|---|---|---|---|---|
      | ⭐ [zotero-gpt](https://github.com/MuiseDestiny/zotero-gpt) | ~7.4k | 🟡 | AGPL-3.0 | read, synthesize | GPT integrated into Zotero to chat with your library |
      | [papersgpt-for-zotero](https://github.com/papersgpt/papersgpt-for-zotero) | ~2.6k | 🟢 | AGPL-3.0 | read, extract, synthesize | Zotero AI + MCP plugin; chat/batch-process PDFs across 30+ LLMs |
      | [ai-research-assistant](https://github.com/lifan0127/ai-research-assistant) | ~1.7k | 🔴 | AGPL-3.0 | read, synthesize | "Aria" — LLM-powered research assistant inside Zotero |
      | [paper-note-filler](https://github.com/chauff/paper-note-filler) | 47 | 🟡 | none | read | Obsidian plugin auto-creating notes from arXiv / ACL Anthology / Semantic Scholar |
      
      ## 📄 PDF → Structured Data Extraction
      
      The invisible infrastructure of lit review: turn PDFs into clean, structured Markdown / JSON for LLMs.
      
      | Project | Stars | Health | Licence | Stage | Notes |
      |---|---|---|---|---|---|
      | ⭐ [MinerU](https://github.com/opendatalab/MinerU) | ~80.4k | 🟢 | Apache-2.0 | extract | High-accuracy PDF/Office → LLM-ready Markdown/JSON (VLM+OCR, 100+ languages, formulas/tables) |
      | [docling](https://github.com/docling-project/docling) | ~67.5k | 🟢 | MIT | extract | IBM-origin document parser prepping PDFs/docs for gen-AI/RAG |
      | [marker](https://github.com/datalab-to/marker) | ~39.9k | 🟢 | Apache-2.0 | extract | Fast PDF/doc → clean Markdown/JSON conversion, scientific-doc friendly |
      | [PDFMathTranslate](https://github.com/PDFMathTranslate/PDFMathTranslate) | ~37.1k | 🟢 | AGPL-3.0 | read | Layout-preserving scientific-PDF translation (formulas/figures intact) |
      | [grobid](https://github.com/grobidOrg/grobid) | ~5.1k | 🟢 | Apache-2.0 | extract | ML tool extracting structured TEI/XML (metadata, refs, sections) from scholarly PDFs |
      | [paperetl](https://github.com/neuml/paperetl) | 697 | 🟡 | Apache-2.0 | extract | ETL pipeline for medical & scientific papers into structured stores |
      | [scipdf_parser](https://github.com/titipata/scipdf_parser) | 456 | 🔴 | MIT | extract | Python parser for scientific-publication PDFs (content + figures, GROBID-backed) |
      
      ## 🕸️ Citation Graphs & API Clients
      
      Analyze citation networks, or hit the major scholarly databases straight from code.
      
      | Project | Stars | Health | Licence | Stage | Notes |
      |---|---|---|---|---|---|
      | [scholarly](https://github.com/scholarly-python-package/scholarly) | ~1.9k | 🟡 | Unlicense | search | Pythonic Google Scholar author/publication retrieval |
      | [semanticscholar](https://github.com/danielnsilva/semanticscholar) | 480 | 🟢 | MIT | search, cite-check | Unofficial Python client for Semantic Scholar APIs |
      | [pyalex](https://github.com/J535D165/pyalex) | 413 | 🟢 | MIT | search, cite-check | Lightweight Python interface to the OpenAlex API |
      | [ArxivDigest](https://github.com/AutoLLM/ArxivDigest) | 467 | 🔴 | MIT | search | Personalized daily arXiv digest with GPT relevancy scoring + email pipeline |
      | [citegraph](https://github.com/Citegraph/citegraph) | 22 | 🟡 | MIT | search | Open web visualizer of 5M+ papers / citation networks (CS bibliography) |
      
      ## ✍️ Writing & Peer-review Assistants
      
      Draft, polish, and run an "AI pre-review" before you submit.
      
      | Project | Stars | Health | Licence | Stage | Notes |
      |---|---|---|---|---|---|
      | [lmms-lab-writer](https://github.com/EvolvingLMMs-Lab/lmms-lab-writer) | 273 | 🟡 | MIT | write | Local-first agentic LaTeX writer for AI-assisted academic writing |
      | [academic-writing-agents](https://github.com/andrehuang/academic-writing-agents) | 200 | 🟡 | MIT | write, review | Claude Code plugin: 10+ specialist agents for academic writing review, research, drafting, polishing |
      | [ai-peer-review](https://github.com/poldrack/ai-peer-review) | 154 | 🟢 | MIT | review | Multi-LLM meta-review: independent reviews synthesized into a meta-review |
      | [open_reviewer](https://github.com/maxidl/openreviewer) | 17 | 🔴 | none | review | Generates high-quality peer reviews of ML/AI conference papers for pre-submission feedback |
      | [academic-research-plugin](https://github.com/JeanDiable/academic-research-plugin) | 25 | 🟡 | MIT | search, synthesize, cite-check, review | Claude Code plugin: lit surveys, paper reviews, citation management; searches arXiv/S2/DBLP and finds research gaps |
      
      ## 📖 Awesome Lists
      
      Want the full picture? Start from these community-maintained lists.
      
      | Project | Stars | Health | Licence | Stage | Notes |
      |---|---|---|---|---|---|
      | [Awesome-LLM-Scientific-Discovery](https://github.com/HKUST-KnowComp/Awesome-LLM-Scientific-Discovery) | 439 | 🟢 | MIT | — | EMNLP 2025 survey list: LLMs in scientific discovery |
      | [Awesome-Auto-Research-Tools](https://github.com/handsome-rich/Awesome-Auto-Research-Tools) | ~1.2k | 🟢 | CC0-1.0 | — | Automated literature search, paper reading, experiment management, code gen |
      | [awesome-ai-auto-research](https://github.com/worldbench/awesome-ai-auto-research) | 522 | 🟢 | MIT | — | A survey on AI auto-research |
      | [LLM4SR](https://github.com/du-nlp-lab/LLM4SR) | 133 | 🔴 | MIT | — | Papers & resources on LLMs for scientific research surveys |
      | [awesome-ai-research-tools](https://github.com/0x11c11e/awesome-ai-research-tools) | 73 | 🟢 | CC0-1.0 | — | AI tools for lit reviews, reference management, data analysis |
      | [awesome-evidence-synthesis](https://github.com/evidencesynthesis-tools/awesome-evidence-synthesis) | 27 | 🟢 | CC-BY-4.0 | — | Open-source tools for systematic reviews, meta-analysis & evidence synthesis |
      
  • scripts
    • fetch_arxiv.py 3.9 KB
      #!/usr/bin/env python3
      """fetch_arxiv — search arXiv and download PDFs into a folder.
      
      A small bundled tool (uses the `arxiv` PyPI package) so litrun workflows can do a
      real retrieval step: topic -> download PDFs -> feed a downstream QA tool. No API
      key required.
      
      Usage:
        fetch_arxiv.py --query "retrieval augmented generation" --max 10 --outdir ./papers
        fetch_arxiv.py --query "cat:cs.CL AND graph neural network" --sort date --max 5 --outdir ./p
      """
      import argparse
      import json
      import re
      import sys
      
      try:
          import arxiv
          import requests  # arxiv depends on requests, so it's always present in this env
      except ImportError as e:
          sys.exit(f"fetch_arxiv: missing dependency ({e}). Reinstall with: litrun.py install arxiv-fetch")
      
      from pathlib import Path
      
      _UA = {"User-Agent": "litrun-fetch-arxiv/1.0 (+https://github.com/brycewang-stanford/lit-review-agent-tools)"}
      
      
      def safe_name(s, maxlen=80):
          s = re.sub(r"[^\w\- ]+", "", s).strip().replace(" ", "_")
          return s[:maxlen] or "paper"
      
      
      def download_pdf(url, dest):
          """Download a PDF via requests — independent of the arxiv package's own
          download API, which changes across major versions (4.0 dropped it)."""
          with requests.get(url, headers=_UA, stream=True, timeout=60, allow_redirects=True) as resp:
              resp.raise_for_status()
              with open(dest, "wb") as f:
                  for chunk in resp.iter_content(chunk_size=1 << 15):
                      if chunk:
                          f.write(chunk)
      
      
      def main():
          p = argparse.ArgumentParser(prog="fetch_arxiv.py")
          p.add_argument("--query", required=True, help="arXiv search query (supports cat:, AND/OR, etc.)")
          p.add_argument("--max", type=int, default=10, help="max papers to download (default 10)")
          p.add_argument("--outdir", required=True, help="directory to download PDFs into")
          p.add_argument("--sort", choices=["relevance", "date"], default="relevance")
          args = p.parse_args()
      
          outdir = Path(args.outdir)
          outdir.mkdir(parents=True, exist_ok=True)
      
          sort = (arxiv.SortCriterion.Relevance if args.sort == "relevance"
                  else arxiv.SortCriterion.SubmittedDate)
          # Polite client: built-in delay + retries so we don't trip arXiv's rate limit (HTTP 429).
          client = arxiv.Client(page_size=min(args.max, 100), delay_seconds=3.0, num_retries=3)
          search = arxiv.Search(query=args.query, max_results=args.max, sort_by=sort)
      
          manifest = []
          n = 0
          print(f"fetch_arxiv: searching arXiv for {args.query!r} (max {args.max}, sort={args.sort})",
                file=sys.stderr)
          try:
              results = list(client.results(search))
          except Exception as e:  # network / API / rate-limit errors
              hint = " (arXiv rate limit — wait a minute and retry)" if "429" in str(e) else ""
              sys.exit(f"fetch_arxiv: search failed: {e}{hint}")
      
          for r in results:
              aid = r.get_short_id()
              title = getattr(r, "title", aid)
              pdf_url = getattr(r, "pdf_url", None) or f"https://arxiv.org/pdf/{aid}"
              fname = f"{aid}_{safe_name(title)}.pdf"
              try:
                  download_pdf(pdf_url, outdir / fname)
                  n += 1
                  print(f"  ✓ {aid}  {title[:70]}", file=sys.stderr)
              except Exception as e:
                  print(f"  ✗ {aid}  download failed: {e}", file=sys.stderr)
                  continue
              manifest.append({
                  "arxiv_id": aid,
                  "title": title,
                  "authors": [getattr(a, "name", str(a)) for a in getattr(r, "authors", [])],
                  "published": str(getattr(r, "published", "")),
                  "pdf": fname,
                  "url": getattr(r, "entry_id", pdf_url),
              })
      
          (outdir / "manifest.json").write_text(json.dumps(manifest, indent=2, ensure_ascii=False))
          print(f"fetch_arxiv: downloaded {n}/{len(results)} PDF(s) to {outdir} "
                f"(manifest.json written).", file=sys.stderr)
          if n == 0:
              sys.exit("fetch_arxiv: no PDFs downloaded.")
      
      
      if __name__ == "__main__":
          main()
      
    • fetch_openalex.py 4.5 KB
      #!/usr/bin/env python3
      """fetch_openalex — search OpenAlex and save papers into a folder.
      
      Downloads the open-access PDF when one is available, otherwise writes a
      `<id>.txt` with the title + reconstructed abstract (PaperQA2 reads text too).
      Writes a merged manifest.json. Uses the free OpenAlex API (no key needed for
      light use; set OPENALEX_API_KEY / --mailto for higher limits + the polite pool).
      
      Usage:
        fetch_openalex.py --query "retrieval augmented generation" --max 10 --outdir ./papers
      """
      import argparse
      import json
      import os
      import re
      import sys
      from pathlib import Path
      
      try:
          import requests
      except ImportError:
          sys.exit("fetch_openalex: 'requests' not installed. Run: litrun.py install openalex-fetch")
      
      API = "https://api.openalex.org/works"
      _UA = {"User-Agent": "litrun-fetch-openalex/1.0 (+https://github.com/brycewang-stanford/lit-review-agent-tools)"}
      
      
      def safe_name(s, maxlen=80):
          s = re.sub(r"[^\w\- ]+", "", s).strip().replace(" ", "_")
          return s[:maxlen] or "work"
      
      
      def reconstruct_abstract(inv):
          """OpenAlex stores abstracts as an inverted index {word: [positions]}."""
          if not inv:
              return ""
          pos = {}
          for word, positions in inv.items():
              for p in positions:
                  pos[p] = word
          return " ".join(pos[i] for i in sorted(pos))
      
      
      def download(url, dest):
          with requests.get(url, headers=_UA, stream=True, timeout=60, allow_redirects=True) as r:
              r.raise_for_status()
              with open(dest, "wb") as f:
                  for chunk in r.iter_content(1 << 15):
                      if chunk:
                          f.write(chunk)
      
      
      def main():
          p = argparse.ArgumentParser(prog="fetch_openalex.py")
          p.add_argument("--query", required=True)
          p.add_argument("--max", type=int, default=10)
          p.add_argument("--outdir", required=True)
          p.add_argument("--mailto", default=os.environ.get("OPENALEX_MAILTO", ""))
          args = p.parse_args()
      
          outdir = Path(args.outdir)
          outdir.mkdir(parents=True, exist_ok=True)
      
          params = {"search": args.query, "per_page": min(args.max, 200)}
          if args.mailto:
              params["mailto"] = args.mailto
          key = os.environ.get("OPENALEX_API_KEY")
          if key:
              params["api_key"] = key
      
          print(f"fetch_openalex: searching OpenAlex for {args.query!r} (max {args.max})", file=sys.stderr)
          try:
              resp = requests.get(API, params=params, headers=_UA, timeout=60)
              resp.raise_for_status()
              works = resp.json().get("results", [])[:args.max]
          except Exception as e:
              sys.exit(f"fetch_openalex: search failed: {e}")
      
          manifest, saved = [], 0
          for w in works:
              oid = (w.get("id") or "").rsplit("/", 1)[-1] or "work"
              title = w.get("display_name") or w.get("title") or oid
              oa = w.get("best_oa_location") or w.get("primary_location") or {}
              pdf_url = oa.get("pdf_url") or (w.get("open_access") or {}).get("oa_url")
              kind = None
              if pdf_url:
                  try:
                      dest = outdir / f"{oid}_{safe_name(title)}.pdf"
                      download(pdf_url, dest)
                      kind = "pdf"
                  except Exception as e:
                      print(f"  ! {oid} PDF failed ({e}); saving abstract instead", file=sys.stderr)
              if kind is None:
                  abstract = reconstruct_abstract(w.get("abstract_inverted_index"))
                  dest = outdir / f"{oid}_{safe_name(title)}.txt"
                  dest.write_text(f"{title}\n\n{abstract}\n", encoding="utf-8")
                  kind = "txt"
              saved += 1
              print(f"  ✓ {oid} [{kind}]  {title[:65]}", file=sys.stderr)
              manifest.append({
                  "openalex_id": oid,
                  "title": title,
                  "year": w.get("publication_year"),
                  "authors": [a.get("author", {}).get("display_name") for a in w.get("authorships", [])],
                  "file": dest.name,
                  "kind": kind,
                  "url": w.get("id"),
              })
      
          _merge_manifest(outdir, manifest)
          print(f"fetch_openalex: saved {saved}/{len(works)} item(s) to {outdir}.", file=sys.stderr)
          if saved == 0:
              sys.exit("fetch_openalex: nothing saved.")
      
      
      def _merge_manifest(outdir, entries):
          """Append to an existing manifest.json so multi-source workflows accumulate."""
          path = outdir / "manifest.json"
          existing = []
          if path.exists():
              try:
                  existing = json.loads(path.read_text())
              except Exception:
                  existing = []
          existing.extend(entries)
          path.write_text(json.dumps(existing, indent=2, ensure_ascii=False), encoding="utf-8")
      
      
      if __name__ == "__main__":
          main()
      
    • fetch_papers.py 20.9 KB
      #!/usr/bin/env python3
      """fetch_papers — one keyless search across six scholarly APIs, deduplicated.
      
      Queries any combination of OpenAlex, Crossref, Semantic Scholar, PubMed,
      Europe PMC and arXiv, merges the hits into one normalised record list
      (deduplicated by DOI, then by normalised title), and writes:
      
        <outdir>/results.json   normalised records + per-source errors
        <outdir>/<key>.txt      title + abstract per record   (--save abstracts, default)
        <outdir>/manifest.json  appended, same shape the other bundled fetchers write
      
      Standard library only — no pip install, no API key. Keys are used if present
      (S2_API_KEY, NCBI_API_KEY, OPENALEX_API_KEY/OPENALEX_MAILTO, CROSSREF_MAILTO)
      and simply raise rate limits.
      
      Usage:
        fetch_papers.py --query "retrieval augmented generation" --max 20 --outdir ./papers
        fetch_papers.py --query "CRISPR off-target" --sources pubmed,europepmc --from-year 2022 \
                        --outdir ./papers --pdfs
      """
      import argparse
      import json
      import os
      import re
      import sys
      import time
      import urllib.error
      import urllib.parse
      import urllib.request
      import xml.etree.ElementTree as ET
      from pathlib import Path
      
      UA = "litrun-fetch-papers/1.0 (+https://github.com/brycewang-stanford/lit-review-agent-tools)"
      SOURCES = ["openalex", "crossref", "semanticscholar", "pubmed", "europepmc", "arxiv"]
      
      
      # ------------------------------------------------------------------ http
      
      
      def _get(url, params=None, headers=None, timeout=60, retries=1):
          """GET a URL, returning raw bytes. Retries once on 429/5xx."""
          if params:
              url = f"{url}?{urllib.parse.urlencode(params, doseq=True)}"
          hdrs = {"User-Agent": UA, "Accept": "*/*"}
          hdrs.update(headers or {})
          for attempt in range(retries + 1):
              try:
                  req = urllib.request.Request(url, headers=hdrs)
                  with urllib.request.urlopen(req, timeout=timeout) as r:
                      return r.read()
              except urllib.error.HTTPError as e:
                  if e.code in (429, 500, 502, 503, 504) and attempt < retries:
                      time.sleep(3)
                      continue
                  raise RuntimeError(f"HTTP {e.code} for {url.split('?')[0]}") from e
              except Exception as e:  # timeouts, DNS, TLS
                  if attempt < retries:
                      time.sleep(2)
                      continue
                  raise RuntimeError(f"{type(e).__name__}: {e}") from e
      
      
      def _get_json(url, params=None, headers=None, timeout=60):
          return json.loads(_get(url, params, headers, timeout).decode("utf-8", "replace"))
      
      
      # ------------------------------------------------------------------ shaping
      
      
      def norm_doi(doi):
          if not doi:
              return ""
          doi = doi.strip().lower()
          doi = re.sub(r"^https?://(dx\.)?doi\.org/", "", doi)
          return doi.rstrip(".")
      
      
      def norm_title(title):
          return re.sub(r"[^a-z0-9]+", "", (title or "").lower())[:120]
      
      
      def strip_tags(s):
          return re.sub(r"\s+", " ", re.sub(r"<[^>]+>", " ", s or "")).strip()
      
      
      def record(source, **kw):
          """Every source funnels into this one shape."""
          r = {
              "source": source, "sources": [source], "id": "", "doi": "", "title": "",
              "abstract": "", "year": None, "authors": [], "venue": "", "cited_by": None,
              "is_oa": None, "pdf_url": "", "url": "",
          }
          r.update({k: v for k, v in kw.items() if v not in (None, "", [])})
          r["doi"] = norm_doi(r["doi"])
          return r
      
      
      # ------------------------------------------------------------------ sources
      
      
      def s_openalex(query, limit, from_year):
          params = {"search": query, "per_page": min(limit, 200)}
          filters = []
          if from_year:
              filters.append(f"from_publication_date:{from_year}-01-01")
          if filters:
              params["filter"] = ",".join(filters)
          if os.environ.get("OPENALEX_API_KEY"):
              params["api_key"] = os.environ["OPENALEX_API_KEY"]
          elif os.environ.get("OPENALEX_MAILTO"):
              params["mailto"] = os.environ["OPENALEX_MAILTO"]
          data = _get_json("https://api.openalex.org/works", params)
          out = []
          for w in data.get("results", [])[:limit]:
              loc = w.get("best_oa_location") or w.get("primary_location") or {}
              out.append(record(
                  "openalex",
                  id=(w.get("id") or "").rsplit("/", 1)[-1],
                  doi=w.get("doi") or "",
                  title=w.get("display_name") or "",
                  abstract=_inverted(w.get("abstract_inverted_index")),
                  year=w.get("publication_year"),
                  authors=[a.get("author", {}).get("display_name") for a in w.get("authorships", [])[:20]],
                  venue=((w.get("primary_location") or {}).get("source") or {}).get("display_name") or "",
                  cited_by=w.get("cited_by_count"),
                  is_oa=(w.get("open_access") or {}).get("is_oa"),
                  pdf_url=loc.get("pdf_url") or "",
                  url=w.get("id") or "",
              ))
          return out
      
      
      def _inverted(inv):
          """OpenAlex ships abstracts as {word: [positions]}."""
          if not inv:
              return ""
          pos = {}
          for word, idxs in inv.items():
              for i in idxs:
                  pos[i] = word
          return " ".join(pos[i] for i in sorted(pos))
      
      
      def s_crossref(query, limit, from_year):
          params = {"query.bibliographic": query, "rows": min(limit, 100)}
          if from_year:
              params["filter"] = f"from-pub-date:{from_year}-01-01"
          mail = os.environ.get("CROSSREF_MAILTO") or os.environ.get("UNPAYWALL_EMAIL")
          if mail:
              params["mailto"] = mail  # polite pool: double the rate limit
          data = _get_json("https://api.crossref.org/works", params)
          out = []
          for w in data.get("message", {}).get("items", [])[:limit]:
              parts = (w.get("issued") or {}).get("date-parts") or [[None]]
              pdf = ""
              for link in w.get("link") or []:
                  if link.get("content-type") == "application/pdf":
                      pdf = link.get("URL", "")
                      break
              out.append(record(
                  "crossref",
                  id=w.get("DOI", ""),
                  doi=w.get("DOI", ""),
                  title=(w.get("title") or [""])[0],
                  abstract=strip_tags(w.get("abstract")),
                  year=parts[0][0] if parts and parts[0] else None,
                  authors=[" ".join(filter(None, [a.get("given"), a.get("family")]))
                           for a in (w.get("author") or [])[:20]],
                  venue=(w.get("container-title") or [""])[0],
                  cited_by=w.get("is-referenced-by-count"),
                  pdf_url=pdf,
                  url=w.get("URL", ""),
              ))
          return out
      
      
      def s_semanticscholar(query, limit, from_year):
          fields = "title,abstract,year,authors,venue,citationCount,externalIds,openAccessPdf,url,isOpenAccess"
          params = {"query": query, "limit": min(limit, 100), "fields": fields}
          if from_year:
              params["year"] = f"{from_year}-"
          headers = {}
          if os.environ.get("S2_API_KEY"):
              headers["x-api-key"] = os.environ["S2_API_KEY"]
          data = _get_json("https://api.semanticscholar.org/graph/v1/paper/search", params, headers)
          out = []
          for w in data.get("data", [])[:limit]:
              ext = w.get("externalIds") or {}
              out.append(record(
                  "semanticscholar",
                  id=w.get("paperId", ""),
                  doi=ext.get("DOI", ""),
                  title=w.get("title") or "",
                  abstract=w.get("abstract") or "",
                  year=w.get("year"),
                  authors=[a.get("name") for a in (w.get("authors") or [])[:20]],
                  venue=w.get("venue") or "",
                  cited_by=w.get("citationCount"),
                  is_oa=w.get("isOpenAccess"),
                  pdf_url=(w.get("openAccessPdf") or {}).get("url", ""),
                  url=w.get("url") or "",
              ))
          return out
      
      
      def s_pubmed(query, limit, from_year):
          base = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils"
          common = {}
          if os.environ.get("NCBI_API_KEY"):
              common["api_key"] = os.environ["NCBI_API_KEY"]
          if os.environ.get("NCBI_EMAIL"):
              common["email"] = os.environ["NCBI_EMAIL"]
          params = dict(common, db="pubmed", term=query, retmax=min(limit, 100), retmode="json", sort="relevance")
          if from_year:
              params["mindate"], params["maxdate"], params["datetype"] = f"{from_year}", "3000", "pdat"
          ids = _get_json(f"{base}/esearch.fcgi", params).get("esearchresult", {}).get("idlist", [])
          if not ids:
              return []
          time.sleep(0.34)  # NCBI: 3 req/s without a key
          xml = _get(f"{base}/efetch.fcgi", dict(common, db="pubmed", id=",".join(ids), retmode="xml"))
          root = ET.fromstring(xml)
          out = []
          for art in root.findall(".//PubmedArticle"):
              pmid = art.findtext(".//PMID") or ""
              # Scope the DOI hunt to the article's own id list. A PubmedArticle embeds
              # its whole reference list, each entry carrying its own <ArticleId
              # IdType="doi">, so a `.//ArticleId` sweep silently returns a *cited*
              # paper's DOI — see recipes/06 for the record this corrupted.
              doi = ""
              for aid in art.findall("./PubmedData/ArticleIdList/ArticleId"):
                  if aid.get("IdType") == "doi":
                      doi = aid.text or ""
                      break
              if not doi:
                  for el in art.findall(".//Article/ELocationID"):
                      if el.get("EIdType") == "doi":
                          doi = el.text or ""
                          break
              abstract = " ".join(" ".join(t.itertext()) for t in art.findall(".//AbstractText")).strip()
              year = art.findtext(".//PubDate/Year") or art.findtext(".//PubDate/MedlineDate") or ""
              title_el = art.find(".//ArticleTitle")  # may carry inline <i>/<sub> markup
              out.append(record(
                  "pubmed",
                  id=pmid, doi=doi,
                  title=re.sub(r"\s+", " ", "".join(title_el.itertext())).strip() if title_el is not None else "",
                  abstract=abstract,
                  year=int(year[:4]) if year[:4].isdigit() else None,
                  authors=[" ".join(filter(None, [a.findtext("ForeName"), a.findtext("LastName")]))
                           for a in art.findall(".//Author")[:20]],
                  venue=art.findtext(".//Journal/Title") or "",
                  url=f"https://pubmed.ncbi.nlm.nih.gov/{pmid}/" if pmid else "",
              ))
          return out
      
      
      def s_europepmc(query, limit, from_year):
          q = query
          if from_year:
              q = f"({query}) AND (FIRST_PDATE:[{from_year}-01-01 TO 3000-01-01])"
          params = {"query": q, "format": "json", "pageSize": min(limit, 100), "resultType": "core"}
          data = _get_json("https://www.ebi.ac.uk/europepmc/webservices/rest/search", params)
          out = []
          for w in data.get("resultList", {}).get("result", [])[:limit]:
              pdf = ""
              for u in ((w.get("fullTextUrlList") or {}).get("fullTextUrl") or []):
                  if u.get("documentStyle") == "pdf":
                      pdf = u.get("url", "")
                      break
              out.append(record(
                  "europepmc",
                  id=w.get("id", ""), doi=w.get("doi", ""),
                  title=w.get("title") or "",
                  abstract=strip_tags(w.get("abstractText")),
                  year=int(w["pubYear"]) if str(w.get("pubYear", "")).isdigit() else None,
                  authors=[a.get("fullName") for a in ((w.get("authorList") or {}).get("author") or [])[:20]],
                  venue=(w.get("journalInfo") or {}).get("journal", {}).get("title") or w.get("bookOrReportDetails", {}).get("publisher", ""),
                  cited_by=w.get("citedByCount"),
                  is_oa=w.get("isOpenAccess") == "Y",
                  pdf_url=pdf,
                  url=f"https://europepmc.org/article/{w.get('source','MED')}/{w.get('id','')}",
              ))
          return out
      
      
      def s_arxiv(query, limit, from_year):
          # arXiv has no server-side date filter on the legacy API; sort by relevance
          # and drop anything older than --from-year client-side.
          params = {"search_query": f"all:{query}", "start": 0,
                    "max_results": min(limit * 2 if from_year else limit, 100),
                    "sortBy": "relevance"}
          xml = _get("http://export.arxiv.org/api/query", params)
          ns = {"a": "http://www.w3.org/2005/Atom", "arxiv": "http://arxiv.org/schemas/atom"}
          out = []
          for e in ET.fromstring(xml).findall("a:entry", ns):
              published = e.findtext("a:published", "", ns)
              year = int(published[:4]) if published[:4].isdigit() else None
              if from_year and year and year < int(from_year):
                  continue
              aid = (e.findtext("a:id", "", ns) or "").rsplit("/", 1)[-1]
              out.append(record(
                  "arxiv",
                  id=aid,
                  doi=e.findtext("arxiv:doi", "", ns),
                  title=re.sub(r"\s+", " ", e.findtext("a:title", "", ns)).strip(),
                  abstract=re.sub(r"\s+", " ", e.findtext("a:summary", "", ns)).strip(),
                  year=year,
                  authors=[a.findtext("a:name", "", ns) for a in e.findall("a:author", ns)[:20]],
                  venue=e.findtext("arxiv:journal_ref", "", ns) or "arXiv",
                  is_oa=True,
                  pdf_url=f"https://arxiv.org/pdf/{aid}" if aid else "",
                  url=e.findtext("a:id", "", ns),
              ))
              if len(out) >= limit:
                  break
          return out
      
      
      FETCHERS = {
          "openalex": s_openalex, "crossref": s_crossref, "semanticscholar": s_semanticscholar,
          "pubmed": s_pubmed, "europepmc": s_europepmc, "arxiv": s_arxiv,
      }
      
      
      # ------------------------------------------------------------------ merge
      
      
      def merge(records):
          """Deduplicate on DOI first, then on normalised title. First hit wins;
          later hits only fill in fields the winner left empty."""
          by_key, order = {}, []
          for r in records:
              key = r["doi"] or f"t:{norm_title(r['title'])}"
              if not key or key == "t:":
                  continue
              if key not in by_key:
                  by_key[key] = r
                  order.append(key)
                  continue
              cur = by_key[key]
              if r["source"] not in cur["sources"]:
                  cur["sources"].append(r["source"])
              for f in ("doi", "abstract", "title", "venue", "pdf_url", "url"):
                  if not cur.get(f) and r.get(f):
                      cur[f] = r[f]
              for f in ("year", "cited_by", "is_oa"):
                  if cur.get(f) in (None, "") and r.get(f) is not None:
                      cur[f] = r[f]
              if len(r.get("authors") or []) > len(cur.get("authors") or []):
                  cur["authors"] = r["authors"]
          return [by_key[k] for k in order]
      
      
      def collapse_titles(records):
          """Second pass: fold records that share a title but carry different DOIs —
          i.e. a preprint and its published version. Keeps the one with the DOI that
          looks published (non-preprint prefix), else the first seen."""
          PREPRINT = ("10.48550", "10.1101", "10.21203", "10.31234", "10.31219", "10.26434")
          by_title, order = {}, []
          for r in records:
              k = norm_title(r["title"])
              if not k:
                  order.append(id(r))
                  by_title[id(r)] = r
                  continue
              if k not in by_title:
                  by_title[k] = r
                  order.append(k)
                  continue
              cur = by_title[k]
      
              def rank(rec):  # published DOI > preprint DOI > no DOI at all
                  if not rec["doi"]:
                      return 0
                  return 1 if rec["doi"].startswith(PREPRINT) else 2
      
              keep, drop = cur, r
              if rank(r) > rank(cur):
                  keep, drop = r, cur
                  by_title[k] = r
              for s in drop["sources"]:
                  if s not in keep["sources"]:
                      keep["sources"].append(s)
              keep.setdefault("also_doi", [])
              if drop["doi"] and drop["doi"] != keep["doi"]:
                  keep["also_doi"].append(drop["doi"])
              if not keep["abstract"] and drop["abstract"]:
                  keep["abstract"] = drop["abstract"]
              if not keep["pdf_url"] and drop["pdf_url"]:
                  keep["pdf_url"] = drop["pdf_url"]
          return [by_title[k] for k in order]
      
      
      def safe_name(s, maxlen=70):
          return (re.sub(r"[^\w\- ]+", "", s or "").strip().replace(" ", "_")[:maxlen]) or "paper"
      
      
      def key_for(r):
          return safe_name(f"{r['source']}_{(r['doi'] or r['id']).replace('/', '_')}_{r['title']}", 90)
      
      
      def download(url, dest):
          data = _get(url, timeout=90)
          if not data:
              raise RuntimeError("empty response")
          dest.write_bytes(data)
      
      
      def merge_manifest(outdir, entries):
          path = outdir / "manifest.json"
          existing = []
          if path.exists():
              try:
                  existing = json.loads(path.read_text())
              except Exception:
                  existing = []
          existing.extend(entries)
          path.write_text(json.dumps(existing, indent=2, ensure_ascii=False), encoding="utf-8")
      
      
      def main():
          p = argparse.ArgumentParser(prog="fetch_papers.py", description=__doc__.split("\n")[0],
                                      formatter_class=argparse.RawDescriptionHelpFormatter)
          p.add_argument("--query", required=True)
          p.add_argument("--sources", default="openalex,crossref,semanticscholar",
                         help=f"comma-separated subset of: {', '.join(SOURCES)} (default: openalex,crossref,semanticscholar)")
          p.add_argument("--max", type=int, default=20, help="results per source before dedup")
          p.add_argument("--from-year", dest="from_year", help="drop anything published before this year")
          p.add_argument("--outdir", required=True)
          p.add_argument("--save", choices=["abstracts", "none"], default="abstracts",
                         help="write one .txt per record for PaperQA2/grep (default) or metadata only")
          p.add_argument("--pdfs", action="store_true", help="also download any PDF the search already exposed")
          p.add_argument("--open-access-only", action="store_true", dest="oa_only")
          p.add_argument("--dedup-titles", action="store_true", dest="dedup_titles",
                         help="also fold same-title records with different DOIs (preprint + published)")
          args = p.parse_args()
      
          picked = [s.strip() for s in args.sources.split(",") if s.strip()]
          unknown = [s for s in picked if s not in FETCHERS]
          if unknown:
              sys.exit(f"fetch_papers: unknown source(s): {', '.join(unknown)}. Known: {', '.join(SOURCES)}")
      
          outdir = Path(args.outdir)
          outdir.mkdir(parents=True, exist_ok=True)
      
          raw, errors, per_source = [], {}, {}
          for s in picked:
              print(f"→ {s}: searching {args.query!r} (max {args.max})", file=sys.stderr)
              try:
                  hits = FETCHERS[s](args.query, args.max, args.from_year)
                  per_source[s] = len(hits)
                  raw.extend(hits)
                  print(f"  ✓ {s}: {len(hits)} hit(s)", file=sys.stderr)
              except Exception as e:
                  errors[s] = str(e)
                  per_source[s] = 0
                  print(f"  ! {s} failed: {e} — continuing with the other sources", file=sys.stderr)
              time.sleep(0.5)
      
          merged = merge(raw)
          if args.dedup_titles:
              before = len(merged)
              merged = collapse_titles(merged)
              print(f"  · title pass folded {before - len(merged)} preprint/published pair(s)", file=sys.stderr)
          if args.oa_only:
              merged = [r for r in merged if r.get("is_oa") or r.get("pdf_url")]
      
          manifest = []
          for r in merged:
              r["key"] = key_for(r)
              fname = ""
              if args.pdfs and r.get("pdf_url"):
                  try:
                      dest = outdir / f"{r['key']}.pdf"
                      download(r["pdf_url"], dest)
                      fname, r["file_kind"] = dest.name, "pdf"
                  except Exception as e:
                      print(f"  ! PDF failed for {r['key'][:40]}: {e}", file=sys.stderr)
              if not fname and args.save == "abstracts":
                  dest = outdir / f"{r['key']}.txt"
                  body = [r["title"], ""]
                  if r["authors"]:
                      body.append(", ".join(a for a in r["authors"] if a))
                  if r["venue"] or r["year"]:
                      body.append(f"{r['venue']} ({r['year']})".strip())
                  if r["doi"]:
                      body.append(f"doi:{r['doi']}")
                  body += ["", r["abstract"] or "(no abstract available from this source)"]
                  dest.write_text("\n".join(body) + "\n", encoding="utf-8")
                  fname, r["file_kind"] = dest.name, "txt"
              r["file"] = fname
              if fname:
                  manifest.append({
                      "title": r["title"], "year": r["year"], "authors": r["authors"],
                      "doi": r["doi"], "file": fname, "kind": r.get("file_kind"),
                      "url": r["url"], "sources": r["sources"],
                  })
      
          if manifest:
              merge_manifest(outdir, manifest)
          results = {
              "query": args.query, "sources": picked, "from_year": args.from_year,
              "raw_hits": len(raw), "unique": len(merged), "per_source": per_source,
              "errors": errors, "results": merged,
          }
          (outdir / "results.json").write_text(json.dumps(results, indent=2, ensure_ascii=False), encoding="utf-8")
      
          dupes = len(raw) - len(merged)
          multi = sum(1 for r in merged if len(r["sources"]) > 1)
          with_doi = sum(1 for r in merged if r["doi"])
          print(f"\nfetch_papers: {len(raw)} raw → {len(merged)} unique "
                f"({dupes} duplicate(s) merged, {multi} found by >1 source, {with_doi} with a DOI)", file=sys.stderr)
          print(f"  results.json + {len(manifest)} file(s) in {outdir}", file=sys.stderr)
          if errors:
              print(f"  failed source(s): {', '.join(errors)}", file=sys.stderr)
          if not merged:
              sys.exit("fetch_papers: no results — try fewer/other --sources or a broader query.")
      
      
      if __name__ == "__main__":
          main()
      
    • fetch_pubmed.py 4.2 KB
      #!/usr/bin/env python3
      """fetch_pubmed — search PubMed and save each article's title + abstract as text.
      
      PubMed rarely exposes downloadable PDFs (full text lives behind publishers /
      PMC), so this saves `<pmid>.txt` files (title + abstract) that PaperQA2 can
      ingest, plus a merged manifest.json. Uses NCBI E-utilities (no key needed for
      light use; set NCBI_API_KEY / --email to raise the rate limit).
      
      Usage:
        fetch_pubmed.py --query "CRISPR off-target detection" --max 10 --outdir ./papers
      """
      import argparse
      import json
      import os
      import re
      import sys
      import xml.etree.ElementTree as ET
      from pathlib import Path
      
      try:
          import requests
      except ImportError:
          sys.exit("fetch_pubmed: 'requests' not installed. Run: litrun.py install pubmed-fetch")
      
      EUTILS = "https://eutils.ncbi.nlm.nih.gov/entrez/eutils"
      _UA = {"User-Agent": "litrun-fetch-pubmed/1.0 (+https://github.com/brycewang-stanford/lit-review-agent-tools)"}
      
      
      def safe_name(s, maxlen=80):
          s = re.sub(r"[^\w\- ]+", "", s).strip().replace(" ", "_")
          return s[:maxlen] or "article"
      
      
      def _common_params():
          p = {}
          if os.environ.get("NCBI_API_KEY"):
              p["api_key"] = os.environ["NCBI_API_KEY"]
          return p
      
      
      def main():
          p = argparse.ArgumentParser(prog="fetch_pubmed.py")
          p.add_argument("--query", required=True)
          p.add_argument("--max", type=int, default=10)
          p.add_argument("--outdir", required=True)
          p.add_argument("--email", default=os.environ.get("NCBI_EMAIL", ""))
          args = p.parse_args()
      
          outdir = Path(args.outdir)
          outdir.mkdir(parents=True, exist_ok=True)
      
          common = _common_params()
          if args.email:
              common["email"] = args.email
      
          print(f"fetch_pubmed: searching PubMed for {args.query!r} (max {args.max})", file=sys.stderr)
          try:
              es = requests.get(f"{EUTILS}/esearch.fcgi", headers=_UA, timeout=60, params={
                  "db": "pubmed", "term": args.query, "retmax": args.max, "retmode": "json", **common})
              es.raise_for_status()
              ids = es.json().get("esearchresult", {}).get("idlist", [])
          except Exception as e:
              sys.exit(f"fetch_pubmed: search failed: {e}")
      
          if not ids:
              sys.exit("fetch_pubmed: no results for that query.")
      
          try:
              ef = requests.get(f"{EUTILS}/efetch.fcgi", headers=_UA, timeout=90, params={
                  "db": "pubmed", "id": ",".join(ids), "rettype": "abstract", "retmode": "xml", **common})
              ef.raise_for_status()
              root = ET.fromstring(ef.content)
          except Exception as e:
              sys.exit(f"fetch_pubmed: fetch failed: {e}")
      
          manifest, saved = [], 0
          for art in root.findall(".//PubmedArticle"):
              pmid = (art.findtext(".//PMID") or "").strip() or "unknown"
              title = (art.findtext(".//ArticleTitle") or "").strip() or pmid
              # AbstractText may be split into labeled sections.
              parts = []
              for node in art.findall(".//Abstract/AbstractText"):
                  label = node.get("Label")
                  text = "".join(node.itertext()).strip()
                  parts.append(f"{label}: {text}" if label else text)
              abstract = "\n".join(parts)
              authors = []
              for a in art.findall(".//Author"):
                  ln, fn = a.findtext("LastName"), a.findtext("ForeName")
                  if ln:
                      authors.append(f"{fn} {ln}".strip() if fn else ln)
              dest = outdir / f"pmid{pmid}_{safe_name(title)}.txt"
              dest.write_text(f"{title}\n\n{abstract}\n", encoding="utf-8")
              saved += 1
              print(f"  ✓ {pmid}  {title[:65]}", file=sys.stderr)
              manifest.append({
                  "pmid": pmid, "title": title, "authors": authors,
                  "file": dest.name, "kind": "txt",
                  "url": f"https://pubmed.ncbi.nlm.nih.gov/{pmid}/",
              })
      
          _merge_manifest(outdir, manifest)
          print(f"fetch_pubmed: saved {saved} abstract(s) to {outdir}.", file=sys.stderr)
          if saved == 0:
              sys.exit("fetch_pubmed: nothing saved.")
      
      
      def _merge_manifest(outdir, entries):
          path = outdir / "manifest.json"
          existing = []
          if path.exists():
              try:
                  existing = json.loads(path.read_text())
              except Exception:
                  existing = []
          existing.extend(entries)
          path.write_text(json.dumps(existing, indent=2, ensure_ascii=False), encoding="utf-8")
      
      
      if __name__ == "__main__":
          main()
      
    • litrun.py 21.3 KB
      #!/usr/bin/env python3
      """litrun — install and run open-source AI literature-review tools by id.
      
      A thin, dependency-free launcher over recipes.json. Each tool gets its own
      isolated virtualenv under ~/.lit-review-tools/envs/<id> so installs never
      collide. API keys live in one shared ~/.lit-review-tools/.env.
      
      Usage:
        litrun.py list [--category C] [--kind K]
        litrun.py info <id>
        litrun.py doctor
        litrun.py env [--set KEY=VALUE ...]     # show / edit the shared .env
        litrun.py install <id>
        litrun.py run <id> [-- <tool args...>]
        litrun.py mcp <id> [--storage PATH] [--client claude|cursor]
        litrun.py ui <id>                       # clone & launch a web UI (gpt-researcher, storm)
        litrun.py workflow list
        litrun.py workflow run <id> --input PATH [--question "..."] [--dry-run]
      
      Design notes for the calling agent:
        - python-cli tools (mineru, marker, docling, paper-qa, asreview): install then
          run fully automatically.
        - python-lib tools (gpt-researcher, storm, scholarly, pyalex): install puts the
          library in a venv; `run` executes the recipe's example snippet.
        - mcp-server tools: `install` prepares them, `mcp` prints the client config block
          to register (they are servers, not one-shot CLIs).
      """
      import argparse
      import json
      import os
      import shutil
      import subprocess
      import sys
      from pathlib import Path
      
      BASE = Path(os.environ.get("LITRUN_HOME", Path.home() / ".lit-review-tools"))
      ENVS = BASE / "envs"
      WORKSPACE = BASE / "workspace"
      ENV_FILE = BASE / ".env"
      RECIPES = Path(__file__).resolve().parent.parent / "recipes" / "recipes.json"
      WORKFLOWS = Path(__file__).resolve().parent.parent / "recipes" / "workflows.json"
      SCRIPTS_DIR = Path(__file__).resolve().parent
      
      
      def die(msg, code=1):
          print(f"litrun: {msg}", file=sys.stderr)
          sys.exit(code)
      
      
      def load_recipes():
          if not RECIPES.exists():
              die(f"recipes.json not found at {RECIPES}")
          data = json.loads(RECIPES.read_text())
          return {t["id"]: t for t in data["tools"]}
      
      
      def load_workflows():
          if not WORKFLOWS.exists():
              die(f"workflows.json not found at {WORKFLOWS}")
          data = json.loads(WORKFLOWS.read_text())
          return {w["id"]: w for w in data["workflows"]}
      
      
      def get_tool(tools, tid):
          if tid not in tools:
              matches = [k for k in tools if tid in k]
              hint = f" Did you mean: {', '.join(matches)}?" if matches else ""
              die(f"unknown tool id '{tid}'.{hint} Run `litrun.py list`.")
          return tools[tid]
      
      
      def have(cmd):
          return shutil.which(cmd) is not None
      
      
      def read_env_file():
          env = {}
          if ENV_FILE.exists():
              for line in ENV_FILE.read_text().splitlines():
                  line = line.strip()
                  if not line or line.startswith("#") or "=" not in line:
                      continue
                  k, v = line.split("=", 1)
                  env[k.strip()] = v.strip()
          return env
      
      
      def write_env_file(env):
          BASE.mkdir(parents=True, exist_ok=True)
          lines = ["# litrun shared environment (API keys). One KEY=VALUE per line.\n"]
          for k, v in sorted(env.items()):
              lines.append(f"{k}={v}\n")
          ENV_FILE.write_text("".join(lines))
      
      
      def merged_env():
          """OS environ overlaid with the shared .env (OS wins if already set)."""
          e = dict(os.environ)
          for k, v in read_env_file().items():
              e.setdefault(k, v)
          return e
      
      
      def venv_dir(tid):
          return ENVS / tid
      
      
      def venv_bin(tid, name):
          d = venv_dir(tid)
          binpath = d / ("Scripts" if os.name == "nt" else "bin")
          return binpath / name
      
      
      def venv_python(tid):
          return venv_bin(tid, "python.exe" if os.name == "nt" else "python")
      
      
      def ensure_venv(tid):
          d = venv_dir(tid)
          if venv_python(tid).exists():
              return
          d.parent.mkdir(parents=True, exist_ok=True)
          if have("uv"):
              run_cmd(["uv", "venv", str(d)])
          else:
              run_cmd([sys.executable, "-m", "venv", str(d)])
      
      
      def pip_install(tid, packages):
          ensure_venv(tid)
          if have("uv"):
              cmd = ["uv", "pip", "install", "--python", str(venv_python(tid)), "-U", *packages]
          else:
              cmd = [str(venv_python(tid)), "-m", "pip", "install", "-U", *packages]
          run_cmd(cmd)
      
      
      def run_cmd(cmd, env=None, check=True, cwd=None):
          printable = " ".join(str(c) for c in cmd)
          loc = f"(cd {cwd}) " if cwd else ""
          print(f"$ {loc}{printable}", file=sys.stderr)
          return subprocess.run(cmd, env=env, check=check, cwd=cwd)
      
      
      def exec_cli(t, tool_args, env, cwd=None, dry=False):
          """Run a python-cli entrypoint or a bundled python-script. Shared by run & workflow."""
          if t["kind"] == "python-cli":
              probe = venv_bin(t["id"], t["entry"])
              cmd = [str(probe), *tool_args]
              display = t["entry"]
          elif t["kind"] == "python-script":
              # Stdlib-only scripts (no pip list) run on the ambient interpreter —
              # no venv, no install, nothing to go stale.
              probe = Path(sys.executable) if not t.get("pip") else venv_python(t["id"])
              cmd = [str(probe), str(SCRIPTS_DIR / t["script"]), *tool_args]
              display = t["script"]
          else:
              die(f"'{t['id']}' is {t['kind']} — can't exec it directly.")
          if dry:
              loc = f"(cd {cwd}) " if cwd else ""
              print(f"DRY  $ {loc}{display} {' '.join(tool_args)}")
              return
          if not probe.exists():
              if t.get("pip"):
                  print(f"{t['id']} not installed — installing first.")
                  pip_install(t["id"], t["pip"])
              else:
                  die(f"'{display}' not found; try: litrun.py install {t['id']}")
          run_cmd(cmd, env=env, check=False, cwd=cwd)
      
      
      # ---------------------------------------------------------------- commands
      
      
      def cmd_list(tools, args):
          rows = []
          for t in tools.values():
              if args.category and t["category"] != args.category:
                  continue
              if args.kind and t["kind"] != args.kind:
                  continue
              no_install_needed = not t.get("pip")  # stdlib scripts and uvx MCP servers
              installed = "✓" if venv_python(t["id"]).exists() or no_install_needed else " "
              rows.append((installed, t["id"], t["kind"], t["category"], t["name"]))
          if not rows:
              print("No tools match that filter.")
              return
          w_id = max(len(r[1]) for r in rows)
          w_kind = max(len(r[2]) for r in rows)
          w_cat = max(len(r[3]) for r in rows)
          print(f"  {'id'.ljust(w_id)}  {'kind'.ljust(w_kind)}  {'category'.ljust(w_cat)}  name")
          for inst, tid, kind, cat, name in rows:
              print(f"{inst} {tid.ljust(w_id)}  {kind.ljust(w_kind)}  {cat.ljust(w_cat)}  {name}")
          print("\n(✓ = installed / no install needed).  litrun.py info <id> for details.")
      
      
      def cmd_info(tools, args):
          t = get_tool(tools, args.id)
          print(f"# {t['name']}  ({t['id']})")
          print(f"repo:     {t['repo']}")
          print(f"category: {t['category']}   kind: {t['kind']}")
          if t.get("pip"):
              print(f"install:  litrun.py install {t['id']}   (pip: {', '.join(t['pip'])})")
          if t.get("entry"):
              print(f"entry:    {t['entry']}")
          if t.get("script"):
              where = "runs via this tool's venv" if t.get("pip") else "stdlib only — no install, runs on python3"
              print(f"script:   {t['script']} (bundled; {where})")
          if t.get("example"):
              print(f"example:  {t['example']}")
          req = t.get("env", [])
          opt = t.get("env_optional", [])
          if req or opt:
              env = merged_env()
              if req:
                  print("required env:")
                  for k in req:
                      print(f"  {'✓' if env.get(k) else '✗'} {k}")
              if opt:
                  print("optional env: " + ", ".join(opt))
          if t["kind"] == "mcp-server":
              print(f"mcp:      litrun.py mcp {t['id']}   (prints client config)")
          if t.get("ui"):
              print(f"ui:       litrun.py ui {t['id']}   (clone & launch web UI at {t['ui']['url']})")
          print(f"\nnotes: {t['notes']}")
      
      
      def cmd_doctor(tools, args):
          print("== toolchain ==")
          for c in ("python3", "uv", "git", "pip"):
              print(f"  {'✓' if have(c) else '✗'} {c}")
          if not have("uv"):
              print("  (uv not found — falling back to python -m venv + pip. Install uv for speed.)")
          print(f"\n== base dir ==\n  {BASE}  ({'exists' if BASE.exists() else 'will be created'})")
          print(f"  env file: {ENV_FILE}  ({'present' if ENV_FILE.exists() else 'missing'})")
          print("\n== installed tools ==")
          any_installed = False
          for t in tools.values():
              if venv_python(t["id"]).exists():
                  any_installed = True
                  print(f"  ✓ {t['id']}")
          if not any_installed:
              print("  (none yet — litrun.py install <id>)")
          print("\n== API keys seen (OS env + .env) ==")
          env = merged_env()
          needed = sorted({k for t in tools.values() for k in t.get("env", []) + t.get("env_optional", [])})
          for k in needed:
              print(f"  {'✓' if env.get(k) else '✗'} {k}")
          missing = [k for k in needed if not env.get(k)]
          if missing:
              print(f"\nSet missing keys with:  litrun.py env --set {missing[0]}=...")
      
      
      def cmd_env(tools, args):
          env = read_env_file()
          if args.set:
              for pair in args.set:
                  if "=" not in pair:
                      die(f"--set expects KEY=VALUE, got '{pair}'")
                  k, v = pair.split("=", 1)
                  env[k.strip()] = v.strip()
              write_env_file(env)
              print(f"Wrote {len(args.set)} key(s) to {ENV_FILE}")
          if not env:
              print(f"No keys set yet. Add with: litrun.py env --set OPENAI_API_KEY=sk-...\n(file: {ENV_FILE})")
              return
          print(f"# {ENV_FILE}")
          for k, v in sorted(env.items()):
              masked = (v[:4] + "…" + v[-2:]) if len(v) > 8 else "***"
              print(f"{k}={masked}")
      
      
      def cmd_install(tools, args):
          t = get_tool(tools, args.id)
          if not t.get("pip"):
              if t["kind"] == "mcp-server":
                  print(f"{t['id']} needs no install (runs via uvx). Register it with: litrun.py mcp {t['id']}")
                  return
              if t["kind"] == "python-script":
                  print(f"{t['id']} needs no install — it is standard-library only. "
                        f"Just run it: litrun.py run {t['id']} -- <args>")
                  return
              die(f"{t['id']} has no pip recipe; see notes: {t['notes']}")
          print(f"Installing {t['name']} into {venv_dir(t['id'])} ...")
          pip_install(t["id"], t["pip"])
          print(f"\n✓ Installed {t['id']}.")
          if t.get("entry"):
              print(f"  Run it:  litrun.py run {t['id']} -- <args>   (e.g. {t.get('example','')})")
          if t["kind"] == "mcp-server":
              print(f"  Register it:  litrun.py mcp {t['id']}")
          req_missing = [k for k in t.get("env", []) if not merged_env().get(k)]
          if req_missing:
              print(f"  ⚠ Needs API keys: {', '.join(req_missing)} — set with litrun.py env --set KEY=...")
      
      
      def cmd_run(tools, args):
          t = get_tool(tools, args.id)
          env = merged_env()
          missing = [k for k in t.get("env", []) if not env.get(k)]
          if missing:
              die(f"missing required env: {', '.join(missing)}. Set with `litrun.py env --set {missing[0]}=...`")
      
          if not venv_python(t["id"]).exists() and t.get("pip"):
              print(f"{t['id']} not installed yet — installing first.")
              pip_install(t["id"], t["pip"])
      
          if t["kind"] in ("python-cli", "python-script"):
              exec_cli(t, args.rest, env)
          elif t["kind"] == "python-lib":
              print(f"{t['name']} is a library, not a CLI. Running its example snippet:\n  {t.get('example','')}\n")
              # Execute the example via the tool's own venv python.
              snippet = t.get("example", "")
              if snippet.startswith("python -c ") or snippet.startswith('python -c'):
                  code = snippet.split("-c", 1)[1].strip().strip('"')
                  run_cmd([str(venv_python(t["id"])), "-c", code], env=env, check=False)
              else:
                  print("No auto-runnable one-liner; see notes:")
                  print(f"  {t['notes']}")
                  if t.get("clone_for_ui"):
                      print(f"  For the full UI, clone {t['repo']} and follow its README.")
          elif t["kind"] == "mcp-server":
              print(f"{t['id']} is an MCP server — it is launched by your MCP client, not run directly.")
              print(f"Register it with:  litrun.py mcp {t['id']}")
          else:
              die(f"don't know how to run kind '{t['kind']}'")
      
      
      def cmd_workflow(tools, args):
          wfs = load_workflows()
          if args.wf_cmd == "list":
              for w in wfs.values():
                  print(f"# {w['id']} — {w['name']}")
                  print(f"    {w['description']}")
                  print(f"    params: {', '.join(w['params'])}")
                  chain = " → ".join(s["tool"] for s in w["steps"])
                  print(f"    steps:  {chain}\n")
              print("Run:  litrun.py workflow run <id> --input <path> [--question \"...\"] [--dry-run]")
              return
      
          if not args.id:
              die("usage: litrun.py workflow run <id> [--input ...] [--question ...]")
          if args.id not in wfs:
              die(f"unknown workflow '{args.id}'. Run `litrun.py workflow list`.")
          w = wfs[args.id]
      
          # Collect params from flags.
          params = {}
          if args.input:
              params["input"] = args.input
          if args.question:
              params["question"] = args.question
          if args.query:
              params["query"] = args.query
          if args.max:
              params["max"] = str(args.max)
          for pair in (args.param or []):
              if "=" not in pair:
                  die(f"--param expects KEY=VALUE, got '{pair}'")
              k, v = pair.split("=", 1)
              params[k] = v
          for k, v in w.get("defaults", {}).items():
              params.setdefault(k, v)
          workdir = WORKSPACE / "runs" / w["id"]
          params["workdir"] = str(workdir)
      
          missing = [p for p in w["params"] if p not in params]
          if missing:
              if args.dry_run:
                  for p in missing:
                      params[p] = f"<{p}>"
              else:
                  die(f"workflow '{w['id']}' needs params: {', '.join(missing)}. "
                      f"Provide with --input / --question / --param KEY=VALUE.")
      
          def sub(s):
              for k, v in params.items():
                  s = s.replace("{" + k + "}", v)
              return s
      
          # Validate every step tool exists and is python-cli; gather required env.
          step_tools = []
          needed_env = set()
          for step in w["steps"]:
              t = get_tool(tools, step["tool"])
              if t["kind"] not in ("python-cli", "python-script"):
                  die(f"workflow step tool '{t['id']}' is {t['kind']}; workflows chain python-cli / python-script tools only.")
              step_tools.append(t)
              needed_env.update(t.get("env", []))
      
          if not args.dry_run:
              env = merged_env()
              env_missing = [k for k in sorted(needed_env) if not env.get(k)]
              if env_missing:
                  die(f"missing API keys for this workflow: {', '.join(env_missing)}. "
                      f"Set with `litrun.py env --set {env_missing[0]}=...`")
              workdir.mkdir(parents=True, exist_ok=True)
          else:
              env = merged_env()
      
          print(f"{'DRY RUN — ' if args.dry_run else ''}workflow '{w['id']}' ({len(w['steps'])} step(s))")
          print(f"  workdir: {workdir}\n")
          for i, (step, t) in enumerate(zip(w["steps"], step_tools), 1):
              step_args = [sub(a) for a in step.get("args", [])]
              cwd = sub(step["cwd"]) if step.get("cwd") else None
              print(f"[step {i}/{len(w['steps'])}] {t['name']}")
              exec_cli(t, step_args, env, cwd=cwd, dry=args.dry_run)
              print()
          if args.dry_run:
              print("(dry run — nothing executed. Drop --dry-run to run for real.)")
          else:
              print(f"✓ workflow '{w['id']}' finished. Outputs: {sub(w.get('outputs') or str(workdir))}")
      
      
      def cmd_ui(tools, args):
          t = get_tool(tools, args.id)
          ui = t.get("ui")
          if not ui:
              die(f"{t['id']} has no web UI. Runnable ids with a UI: gpt-researcher, storm.")
      
          env = merged_env()
          missing = [k for k in t.get("env", []) if not env.get(k)]
          if missing:
              print(f"⚠ {t['id']} UI needs API keys: {', '.join(missing)}.")
              if ui.get("write_env"):
                  die(f"Set them first: litrun.py env --set {missing[0]}=...")
              else:
                  print(f"  Enter them in the app once it opens (or: litrun.py env --set {missing[0]}=...).")
      
          clone_dir = WORKSPACE / t["id"]
          if not clone_dir.exists():
              WORKSPACE.mkdir(parents=True, exist_ok=True)
              run_cmd(["git", "clone", "--depth", "1", t["repo"], str(clone_dir)])
          else:
              print(f"(repo already cloned at {clone_dir})", file=sys.stderr)
      
          ensure_venv(t["id"])
          for req in ui.get("requirements", []):
              req_path = clone_dir / req
              if not req_path.exists():
                  print(f"⚠ requirements file not found: {req_path} (upstream layout may have changed)", file=sys.stderr)
                  continue
              if have("uv"):
                  run_cmd(["uv", "pip", "install", "--python", str(venv_python(t["id"])), "-r", str(req_path)])
              else:
                  run_cmd([str(venv_python(t["id"])), "-m", "pip", "install", "-r", str(req_path)])
      
          if ui.get("write_env"):
              keys = t.get("env", []) + t.get("env_optional", [])
              lines = [f"{k}={env[k]}\n" for k in keys if env.get(k)]
              if lines:
                  (clone_dir / ".env").write_text("".join(lines))
                  print(f"Wrote {len(lines)} key(s) to {clone_dir / '.env'}", file=sys.stderr)
      
          # Prepend the venv's bin to PATH so `streamlit`/`uvicorn`/`python` resolve to it.
          binpath = venv_dir(t["id"]) / ("Scripts" if os.name == "nt" else "bin")
          env["PATH"] = f"{binpath}{os.pathsep}{env.get('PATH', '')}"
      
          print(f"\n▶ Launching {t['name']} UI — open {ui['url']} once it's up. Ctrl-C to stop.")
          if ui.get("notes"):
              print(f"  {ui['notes']}")
          run_cmd(list(ui["run"]), env=env, check=False)
      
      
      def cmd_mcp(tools, args):
          t = get_tool(tools, args.id)
          m = t.get("mcp")
          if not m:
              die(f"{t['id']} is not an MCP server.")
          launcher = m["launcher"]
          if launcher == "uvx":
              storage = args.storage or str(WORKSPACE / t["id"])
              arglist = [a.replace("{storage}", storage) for a in m["args_template"]]
              block = {"command": "uvx", "args": arglist}
          elif launcher == "venv-module":
              if not venv_python(t["id"]).exists():
                  print(f"(note: {t['id']} not installed yet — run `litrun.py install {t['id']}` so this command resolves)")
              block = {"command": str(venv_python(t["id"])), "args": ["-m", m["module"]]}
          elif launcher == "venv-entry":
              if not venv_python(t["id"]).exists():
                  print(f"(note: {t['id']} not installed yet — run `litrun.py install {t['id']}` so this command resolves)")
              block = {"command": str(venv_bin(t["id"], m["entry"]))}
          else:
              die(f"unknown mcp launcher '{launcher}'")
          if m.get("env"):
              block["env"] = dict(m["env"])
          config = {"mcpServers": {m["key"]: block}}
      
          print(f"# Add this to your MCP client config ({args.client}):")
          if args.client == "claude":
              print("#   Claude Code: ~/.claude.json  (or project .mcp.json)   |   Claude Desktop: claude_desktop_config.json")
          else:
              print("#   Cursor: ~/.cursor/mcp.json  (or .cursor/mcp.json in the project)")
          print(json.dumps(config, indent=2))
          if t.get("env_optional"):
              print(f"\n# Optional env you may want to fill into the 'env' block: {', '.join(t['env_optional'])}")
          print(f"\n# notes: {t['notes']}")
      
      
      def main():
          p = argparse.ArgumentParser(prog="litrun.py", description="Install & run AI lit-review tools by id.")
          sub = p.add_subparsers(dest="cmd", required=True)
      
          p_list = sub.add_parser("list", help="list runnable tools")
          p_list.add_argument("--category")
          p_list.add_argument("--kind")
      
          p_info = sub.add_parser("info", help="show a tool's install/run details")
          p_info.add_argument("id")
      
          sub.add_parser("doctor", help="check toolchain, installs, and API keys")
      
          p_env = sub.add_parser("env", help="show/edit the shared .env")
          p_env.add_argument("--set", nargs="*", help="KEY=VALUE pairs to write")
      
          p_inst = sub.add_parser("install", help="install a tool into an isolated venv")
          p_inst.add_argument("id")
      
          p_run = sub.add_parser("run", help="run a tool (installs if needed)")
          p_run.add_argument("id")
          p_run.add_argument("rest", nargs=argparse.REMAINDER,
                             help="args after `--` are passed to the tool")
      
          p_mcp = sub.add_parser("mcp", help="print the MCP client config for an MCP server")
          p_mcp.add_argument("id")
          p_mcp.add_argument("--storage", help="storage path (uvx servers)")
          p_mcp.add_argument("--client", choices=["claude", "cursor"], default="claude")
      
          p_ui = sub.add_parser("ui", help="clone & launch a tool's web UI (gpt-researcher, storm)")
          p_ui.add_argument("id")
      
          p_wf = sub.add_parser("workflow", help="run a named multi-step pipeline")
          p_wf.add_argument("wf_cmd", choices=["list", "run"])
          p_wf.add_argument("id", nargs="?")
          p_wf.add_argument("--input")
          p_wf.add_argument("--question")
          p_wf.add_argument("--query")
          p_wf.add_argument("--max")
          p_wf.add_argument("--param", action="append", help="extra KEY=VALUE params")
          p_wf.add_argument("--dry-run", action="store_true", dest="dry_run",
                            help="print resolved step commands without running")
      
          args = p.parse_args()
          # Strip a leading `--` separator from run's REMAINDER.
          if getattr(args, "rest", None) and args.rest and args.rest[0] == "--":
              args.rest = args.rest[1:]
      
          tools = load_recipes()
          dispatch = {
              "list": cmd_list, "info": cmd_info, "doctor": cmd_doctor, "env": cmd_env,
              "install": cmd_install, "run": cmd_run, "mcp": cmd_mcp, "ui": cmd_ui,
              "workflow": cmd_workflow,
          }
          dispatch[args.cmd](tools, args)
      
      
      if __name__ == "__main__":
          main()
      
    • resolve_oa.py 13.3 KB
      #!/usr/bin/env python3
      """resolve_oa — turn DOIs into legally downloadable full text.
      
      A search gives you metadata; a citation-backed answer needs the actual paper.
      This walks an escalating chain of open-access resolvers per DOI and stops at
      the first one that yields a file:
      
        1. Unpaywall        best OA location   (needs an email — free, no key)
        2. OpenAlex         best_oa_location / oa_url        (keyless)
        3. Europe PMC       OA full-text XML → plain text    (keyless, biomed)
        4. arXiv            preprint PDF when the record has an arXiv id
        5. CORE             OA aggregator      (only if CORE_API_KEY is set)
      
      Standard library only. Nothing here bypasses a paywall: a DOI nothing can open is
      reported as `closed` and skipped — or as `not-in-crossref` when Crossref has never
      heard of it, which means either an invented citation or a DOI registered elsewhere
      (DataCite datasets, some preprint servers). Either way, do not cite it unchecked.
      
      Input is DOIs on the command line, a text file of DOIs, or the results.json /
      manifest.json that fetch_papers.py wrote.
      
      Usage:
        resolve_oa.py --from-json ./papers/results.json --outdir ./papers --email you@example.com
        resolve_oa.py --doi 10.7717/peerj.4375 --doi 10.1038/nature12373 --outdir ./papers
      """
      import argparse
      import json
      import os
      import re
      import sys
      import time
      import urllib.error
      import urllib.parse
      import urllib.request
      import xml.etree.ElementTree as ET
      from pathlib import Path
      
      UA = "litrun-resolve-oa/1.0 (+https://github.com/brycewang-stanford/lit-review-agent-tools)"
      
      
      def _get(url, params=None, timeout=60, retries=1):
          if params:
              url = f"{url}?{urllib.parse.urlencode(params, doseq=True)}"
          for attempt in range(retries + 1):
              try:
                  req = urllib.request.Request(url, headers={"User-Agent": UA})
                  with urllib.request.urlopen(req, timeout=timeout) as r:
                      return r.read()
              except urllib.error.HTTPError as e:
                  if e.code in (429, 500, 502, 503) and attempt < retries:
                      time.sleep(3)
                      continue
                  raise RuntimeError(f"HTTP {e.code}") from e
              except Exception as e:
                  if attempt < retries:
                      time.sleep(2)
                      continue
                  raise RuntimeError(f"{type(e).__name__}: {e}") from e
      
      
      def _get_json(url, params=None, timeout=60):
          return json.loads(_get(url, params, timeout).decode("utf-8", "replace"))
      
      
      def norm_doi(doi):
          doi = (doi or "").strip().lower()
          return re.sub(r"^https?://(dx\.)?doi\.org/", "", doi).rstrip(".")
      
      
      def safe_name(s, maxlen=90):
          return (re.sub(r"[^\w\-]+", "_", s or "").strip("_")[:maxlen]) or "paper"
      
      
      def is_pdf(blob):
          return blob[:5] == b"%PDF-"
      
      
      def jats_to_text(xml_bytes):
          """Flatten PMC's JATS XML into title + abstract + body, dropping the
          bibliographic front matter that makes a naive itertext() dump unreadable."""
          root = ET.fromstring(xml_bytes)
          chunks = []
          title = root.find(".//article-title")
          if title is not None:
              chunks.append("".join(title.itertext()).strip())
          for tag in ("abstract", "body"):
              for el in root.findall(f".//{tag}"):
                  chunks.append(" ".join("".join(el.itertext()).split()))
          return "\n\n".join(c for c in chunks if c)
      
      
      # ------------------------------------------------------------------ resolvers
      # Each returns (kind, bytes_or_text, source_url) or None.
      
      
      def via_unpaywall(doi, email):
          if not email:
              return None
          d = _get_json(f"https://api.unpaywall.org/v2/{urllib.parse.quote(doi)}", {"email": email})
          if not d.get("is_oa"):
              return None
          loc = d.get("best_oa_location") or {}
          url = loc.get("url_for_pdf") or loc.get("url")
          if not url:
              return None
          blob = _get(url, timeout=90)
          return ("pdf", blob, url) if is_pdf(blob) else None
      
      
      def via_openalex(doi):
          w = _get_json(f"https://api.openalex.org/works/doi:{urllib.parse.quote(doi)}")
          loc = w.get("best_oa_location") or {}
          url = loc.get("pdf_url") or (w.get("open_access") or {}).get("oa_url")
          ids = w.get("ids") or {}
          if not url:
              # An arXiv landing page in `ids` is still a downloadable preprint.
              return ("arxiv-id", ids.get("arxiv") or "", "") if ids.get("arxiv") else None
          blob = _get(url, timeout=90)
          return ("pdf", blob, url) if is_pdf(blob) else None
      
      
      def via_europepmc(doi):
          params = {"query": f'DOI:"{doi}"', "format": "json", "pageSize": 1, "resultType": "core"}
          res = _get_json("https://www.ebi.ac.uk/europepmc/webservices/rest/search", params)
          hits = res.get("resultList", {}).get("result", [])
          if not hits:
              return None
          h = hits[0]
          pmcid = h.get("pmcid")
          if pmcid and h.get("isOpenAccess") == "Y":
              url = f"https://www.ebi.ac.uk/europepmc/webservices/rest/{pmcid}/fullTextXML"
              try:
                  text = jats_to_text(_get(url, timeout=90))
                  if len(text) > 2000:  # a stub isn't full text
                      return ("txt", text.encode("utf-8"), url)
              except Exception:
                  pass
          for u in ((h.get("fullTextUrlList") or {}).get("fullTextUrl") or []):
              if u.get("documentStyle") == "pdf" and u.get("availabilityCode") in ("OA", "F"):
                  try:
                      blob = _get(u["url"], timeout=90)
                      if is_pdf(blob):
                          return ("pdf", blob, u["url"])
                  except Exception:
                      continue
          return None
      
      
      def via_arxiv(arxiv_id):
          aid = (arxiv_id or "").rsplit("/", 1)[-1]
          if not aid:
              return None
          url = f"https://arxiv.org/pdf/{aid}"
          blob = _get(url, timeout=90)
          return ("pdf", blob, url) if is_pdf(blob) else None
      
      
      def doi_registered(doi):
          """Does this DOI exist at all? Asked only when nothing resolved, so that a
          fabricated citation is reported as `not-in-crossref` rather than as `closed` —
          'behind a paywall' and 'no such record' are very different answers."""
          try:
              _get(f"https://api.crossref.org/works/{urllib.parse.quote(doi)}", timeout=30, retries=0)
              return True
          except RuntimeError as e:
              return False if "HTTP 404" in str(e) else None  # None = could not tell
      
      
      def via_core(doi, key):
          if not key:
              return None
          body = json.dumps({"q": f'doi:"{doi}"', "limit": 1}).encode()
          req = urllib.request.Request(
              "https://api.core.ac.uk/v3/search/works",
              data=body,
              headers={"User-Agent": UA, "Authorization": f"Bearer {key}", "Content-Type": "application/json"},
          )
          with urllib.request.urlopen(req, timeout=60) as r:
              res = json.loads(r.read().decode("utf-8", "replace"))
          for w in res.get("results", []):
              url = w.get("downloadUrl")
              if url:
                  try:
                      blob = _get(url, timeout=90)
                      if is_pdf(blob):
                          return ("pdf", blob, url)
                  except Exception:
                      continue
          return None
      
      
      # ------------------------------------------------------------------ input
      
      
      def load_targets(args):
          """→ [{doi, title, arxiv_id, pdf_url}] from whichever input was given."""
          targets = []
          for d in args.doi or []:
              targets.append({"doi": norm_doi(d), "title": "", "arxiv_id": "", "pdf_url": ""})
          if args.from_json:
              data = json.loads(Path(args.from_json).read_text())
              rows = data.get("results", data) if isinstance(data, dict) else data
              for r in rows:
                  if not isinstance(r, dict):
                      continue
                  aid = r.get("id", "") if r.get("source") == "arxiv" else ""
                  targets.append({"doi": norm_doi(r.get("doi")), "title": r.get("title", ""),
                                  "arxiv_id": aid, "pdf_url": r.get("pdf_url", "")})
          if args.from_file:
              for line in Path(args.from_file).read_text().splitlines():
                  line = line.strip()
                  if line and not line.startswith("#"):
                      targets.append({"doi": norm_doi(line), "title": "", "arxiv_id": "", "pdf_url": ""})
          # Keep records that carry *some* handle, deduplicated.
          seen, out = set(), []
          for t in targets:
              k = t["doi"] or t["arxiv_id"] or t["pdf_url"]
              if not k or k in seen:
                  continue
              seen.add(k)
              out.append(t)
          return out
      
      
      def main():
          p = argparse.ArgumentParser(prog="resolve_oa.py", description=__doc__.split("\n")[0],
                                      formatter_class=argparse.RawDescriptionHelpFormatter)
          p.add_argument("--doi", action="append", help="a DOI (repeatable)")
          p.add_argument("--from-json", dest="from_json", help="results.json/manifest.json from fetch_papers.py")
          p.add_argument("--from-file", dest="from_file", help="text file, one DOI per line")
          p.add_argument("--outdir", required=True)
          p.add_argument("--email", default=os.environ.get("UNPAYWALL_EMAIL") or os.environ.get("NCBI_EMAIL", ""),
                         help="contact email — required by Unpaywall (else that step is skipped)")
          p.add_argument("--max", type=int, default=0, help="stop after N DOIs (0 = all)")
          p.add_argument("--skip-existing", action="store_true", help="don't re-download files already in --outdir")
          args = p.parse_args()
      
          if not (args.doi or args.from_json or args.from_file):
              sys.exit("resolve_oa: give --doi, --from-json or --from-file.")
      
          targets = load_targets(args)
          if args.max:
              targets = targets[:args.max]
          if not targets:
              sys.exit("resolve_oa: no usable DOIs in the input.")
      
          outdir = Path(args.outdir)
          outdir.mkdir(parents=True, exist_ok=True)
          core_key = os.environ.get("CORE_API_KEY", "")
          if not args.email:
              print("resolve_oa: no --email/UNPAYWALL_EMAIL — skipping Unpaywall, "
                    "the rest of the chain still runs keyless.", file=sys.stderr)
      
          report, counts = [], {"pdf": 0, "txt": 0, "closed": 0, "not-in-crossref": 0, "error": 0, "skipped": 0}
          for i, t in enumerate(targets, 1):
              doi, label = t["doi"], (t["title"] or t["doi"] or t["arxiv_id"])[:60]
              stem = safe_name(doi.replace("/", "_") or t["arxiv_id"])
              print(f"[{i}/{len(targets)}] {label}", file=sys.stderr)
      
              existing = list(outdir.glob(f"{stem}.*"))
              if args.skip_existing and existing:
                  counts["skipped"] += 1
                  report.append({"doi": doi, "status": "skipped", "file": existing[0].name})
                  print("   · already present", file=sys.stderr)
                  continue
      
              chain = []
              if doi:
                  chain += [("unpaywall", lambda: via_unpaywall(doi, args.email)),
                            ("openalex", lambda: via_openalex(doi)),
                            ("europepmc", lambda: via_europepmc(doi))]
              if t["arxiv_id"]:
                  chain.append(("arxiv", lambda: via_arxiv(t["arxiv_id"])))
              if t["pdf_url"]:
                  chain.append(("search-hit", lambda: (lambda b: ("pdf", b, t["pdf_url"]) if is_pdf(b) else None)(
                      _get(t["pdf_url"], timeout=90))))
              if doi and core_key:
                  chain.append(("core", lambda: via_core(doi, core_key)))
      
              got, errs = None, []
              for name, fn in chain:
                  try:
                      res = fn()
                  except Exception as e:
                      errs.append(f"{name}: {e}")
                      continue
                  if res is None:
                      continue
                  kind, payload, url = res
                  if kind == "arxiv-id":  # OpenAlex handed us a preprint id instead of a file
                      try:
                          res2 = via_arxiv(payload)
                      except Exception as e:
                          errs.append(f"arxiv: {e}")
                          continue
                      if not res2:
                          continue
                      kind, payload, url = res2
                      name = "openalex→arxiv"
                  got = (name, kind, payload, url)
                  break
      
              if not got:
                  if not doi:
                      status = "error"
                  else:
                      status = "closed" if doi_registered(doi) is not False else "not-in-crossref"
                  counts[status] += 1
                  report.append({"doi": doi, "title": t["title"], "status": status, "tried": errs})
                  why = "DOI is not registered with Crossref" if status == "not-in-crossref" else "no open copy found"
                  print(f"   ✗ {why}{' (' + '; '.join(errs[:2]) + ')' if errs else ''}", file=sys.stderr)
                  continue
      
              name, kind, payload, url = got
              dest = outdir / f"{stem}.{'pdf' if kind == 'pdf' else 'txt'}"
              dest.write_bytes(payload)
              counts[kind] += 1
              report.append({"doi": doi, "title": t["title"], "status": kind, "via": name,
                             "url": url, "file": dest.name, "bytes": len(payload)})
              print(f"   ✓ {kind} via {name}  ({len(payload)//1024} KB)", file=sys.stderr)
              time.sleep(0.3)
      
          out = {"resolved": counts, "total": len(targets), "items": report}
          (outdir / "oa_report.json").write_text(json.dumps(out, indent=2, ensure_ascii=False), encoding="utf-8")
          hit = counts["pdf"] + counts["txt"]
          print(f"\nresolve_oa: {hit}/{len(targets)} open ({counts['pdf']} PDF, {counts['txt']} full-text XML→txt), "
                f"{counts['closed']} closed, {counts['not-in-crossref']} DOI(s) unknown to Crossref, "
                f"{counts['error']} errored, {counts['skipped']} skipped", file=sys.stderr)
          print(f"  files + oa_report.json in {outdir}", file=sys.stderr)
          if hit == 0:
              sys.exit("resolve_oa: nothing was openly available — the corpus is unchanged.")
      
      
      if __name__ == "__main__":
          main()
      
  • SKILL.md 14.1 KB
    ---
    name: literature-review-tools
    description: >-
      Recommend AND run open-source AI tools, agents, Claude Code / Codex skills, and
      MCP servers for any stage of a literature review — searching, reading,
      extracting, synthesizing, screening, citation-checking, and paper writing. Use
      when the user asks "what tool should I use to..." OR "install/run/use <tool> to
      ..." for research/lit-review work: automating a survey or related-work section,
      PDF→Markdown extraction for LLMs (MinerU/marker/docling), PRISMA / systematic
      review (ASReview), citation-backed Q&A over PDFs (PaperQA2), wiring papers into
      Claude/Cursor via MCP (arxiv/paper-search/zotero servers), or chatting with a
      Zotero library. Also answers direct literature lookups — find papers on a topic,
      resolve a DOI, get an open-access PDF, check a citation, search PubMed / OpenAlex /
      Crossref / Semantic Scholar / Europe PMC / arXiv — with no install and no API key.
      Ships a launcher (scripts/litrun.py) that installs each tool in an isolated venv and
      runs it, plus keyless bundled scripts for multi-source search and open-access
      full-text retrieval. Curated catalog of 70+ vetted projects.
      支持中英文(用于「文献综述工具选型」「文献检索」与「一键安装/运行」)。
    ---
    
    # Literature Review Tools — Select & Run
    
    A curated, use-case-organized catalog of the strongest **open-source** AI tools for
    literature review — **plus a launcher that actually installs and runs the top ones.**
    Covers: end-to-end research agents, deep-research / auto-survey generators, autonomous
    "idea→paper" systems, citation-backed RAG over PDFs, PRISMA screening, MCP servers,
    Zotero/Obsidian integrations, PDF→structured extraction, citation graphs, and
    paper-writing / peer-review assistants.
    
    Full source of truth (README, always current star counts): <https://github.com/brycewang-stanford/lit-review-agent-tools>
    
    ## Three modes
    
    - **Look up** — user wants *the literature itself*: "find papers on X", "get me the PDF for this DOI", "does this citation exist", "what does PubMed have since 2022". Answer it directly — no install, no key. Start with [`reference/apis/README.md`](reference/apis/README.md) for one-off lookups, or run the bundled `papers-fetch` / `oa-resolve` scripts for anything corpus-sized.
    - **Recommend** — user asks "what should I use to …". Route with the tables below; cite the catalog for details.
    - **Run** — user asks to *install / run / use* a specific tool ("turn this PDF into Markdown with MinerU", "ask PaperQA2 about these papers", "set up the arXiv MCP server"). Drive [`scripts/litrun.py`](scripts/litrun.py) via Bash — do not hand the user raw pip commands to copy.
    
    Mode 1 is the cheap default. Do not send someone to install PyTorch when they asked
    for five papers and a PDF.
    
    ## Look up mode — search without installing anything
    
    Two bundled scripts, both **standard-library only**: no venv, no pip, no API key.
    Run them straight (`python3 scripts/fetch_papers.py …`) or through the launcher.
    
    ```bash
    # 1. search six indexes at once, deduplicated by DOI
    python3 scripts/fetch_papers.py --query "active learning for screening" \
        --sources openalex,crossref,semanticscholar,pubmed,europepmc,arxiv \
        --max 15 --dedup-titles --outdir ./corpus
    
    # 2. turn those DOIs into full text you may legally read
    python3 scripts/resolve_oa.py --from-json ./corpus/results.json --outdir ./corpus
    ```
    
    `fetch_papers.py` writes `results.json` (normalised records, per-source counts, and the
    errors of any source that failed) plus one `.txt` per paper. `resolve_oa.py` walks
    Unpaywall → OpenAlex → Europe PMC → arXiv → CORE and writes `oa_report.json` recording
    how every DOI resolved — **including the ones that stayed closed**, and separating those
    from DOIs Crossref has never registered (usually an invented citation). On a 47-paper
    corpus with zero keys it recovered 33 full texts; nothing in it bypasses a paywall, and
    you should not offer to.
    
    For a *single* lookup — one DOI, one author, one citation check — skip the scripts and
    call the API directly (`WebFetch`/`curl`). [`reference/apis/`](reference/apis/) has a
    routing table, the identifier formats, and one page per API with the endpoints and the
    failure modes that waste time (Crossref's `select` 400, OpenAlex's inverted abstracts,
    Semantic Scholar's exhausted keyless pool, PubMed's multi-part `AbstractText`).
    
    Report which indexes you queried and which came back empty. "OpenAlex had no match" is
    a fact; "that paper doesn't exist" is a much bigger claim than one API can support.
    
    ## Run mode — how to drive `scripts/litrun.py`
    
    The launcher installs each supported tool into its own venv under `~/.lit-review-tools/`
    (uses `uv` if present, else `python -m venv`) and reads API keys from one shared
    `~/.lit-review-tools/.env`. Machine-readable recipes: [`recipes/recipes.json`](recipes/recipes.json).
    
    Typical flow when the user wants to *use* a tool:
    
    1. `python3 scripts/litrun.py doctor` — check toolchain + which API keys are already set.
    2. `python3 scripts/litrun.py info <id>` — confirm what the tool needs (entry, required env).
    3. If a required key is missing, ask the user for it, then `litrun.py env --set KEY=VALUE` (never echo the value back in full).
    4. `python3 scripts/litrun.py run <id> -- <tool args>` — installs on first use, then runs. For PDF tools pass the real file path; e.g. `run mineru -- -p paper.pdf -o ./out -b pipeline`.
    5. For **MCP servers**, don't "run" them — `litrun.py mcp <id>` prints the client config block to register in Claude Code / Cursor.
    
    Commands: `list [--category C] [--kind K]` · `info <id>` · `doctor` · `env [--set K=V]` · `install <id>` · `run <id> -- <args>` · `mcp <id> [--storage PATH] [--client claude|cursor]` · `ui <id>`.
    
    Runnable ids by kind:
    - **python-cli (auto install+run):** `mineru`, `marker`, `docling` (PDF→Markdown) · `paper-qa` (cited Q&A) · `asreview` (PRISMA screening UI)
    - **python-script, zero install (stdlib only):** `papers-fetch` (6-source deduplicated search) · `oa-resolve` (DOI → open-access full text)
    - **python-script (bundled, auto install+run):** `arxiv-fetch` (arXiv PDFs) · `openalex-fetch` (OpenAlex; PDF or abstract .txt) · `pubmed-fetch` (PubMed abstracts) — all keyless for light use. `papers-fetch` supersedes all three when you want coverage rather than one index.
    - **python-lib (install + run example):** `gpt-researcher`, `storm` (deep research; need API keys) · `scholarly`, `pyalex` (API clients)
    - **mcp-server (install + `mcp` config):** `arxiv-mcp-server`, `paper-search-mcp`, `zotero-mcp`
    
    For **`gpt-researcher`** and **`storm`**, `litrun.py ui <id>` clones the repo and launches the full web UI (GPT Researcher → FastAPI at :8000; STORM → Streamlit at :8501). These are long-running servers — launch them with a background Bash call and tell the user the URL. gpt-researcher's UI needs `OPENAI_API_KEY` + `TAVILY_API_KEY` set first (litrun writes them into the repo's `.env`); STORM takes its keys in the app sidebar.
    
    ### Chained pipelines
    
    For multi-tool tasks, prefer a named workflow over hand-wiring steps: `litrun.py workflow list` then `litrun.py workflow run <id> [--input PATH] [--query "..."] [--question "..."] [--max N]`. Built-ins:
    - `pdf-to-markdown` — a PDF/folder → clean Markdown (MinerU)
    - `pdf-corpus-qa` — a folder of PDFs → citation-backed answer (PaperQA2)
    - `pdf-md-then-qa` — convert to Markdown **and** answer a question over the corpus
    - `topic-to-pdfs` — arXiv query → download top-N PDFs (arxiv-fetch, no key)
    - `topic-to-review` — arXiv query → download PDFs → citation-backed answer (PaperQA2). The end-to-end "retrieve then review" pipeline; no MCP client needed. Needs `OPENAI_API_KEY` for the QA step.
    - `topic-to-review-multi` — retrieve from **arXiv + OpenAlex** into one corpus → citation-backed answer. Broader coverage; resilient if one source is rate-limited.
    - `topic-to-related-work` — retrieve (arXiv + OpenAlex) → PaperQA2 drafts a **cited related-work paragraph** synthesizing themes/methods/gaps. Needs `OPENAI_API_KEY`.
    - `topic-to-corpus` — six-source search → deduplicated corpus. **Keyless, no install** — the default first step for any review.
    - `topic-to-fulltext` — six-source search → open-access full text for every DOI it can legally get. Keyless.
    - `topic-to-fulltext-review` — the strongest path: search → dedup → full text → PaperQA2 answers over **full texts, not abstracts**. `OPENAI_API_KEY` for the last step only.
    
    Prefer the `topic-to-fulltext*` workflows over `topic-to-review-multi`: same idea, six
    sources instead of two, DOI-level deduplication, and it retrieves the papers rather
    than the abstracts. Biomedical topics need no special casing — PubMed and Europe PMC
    are already in the source set.
    
    Add `--dry-run` first to show the exact resolved step commands without executing — good for confirming paths with the user before a heavy run. Workflows fail fast if a required API key is missing.
    
    Guardrails: installs and downloads happen under the user's home and hit the network — for a heavy first install (marker/docling pull in PyTorch) say so before running. Never fabricate API keys. If a `run` fails, show the real error rather than claiming success. Paths in this file (`scripts/…`, `recipes/…`) are relative to this skill's directory.
    
    ## Recommend mode — how to route
    
    1. Identify **which stage** of the lit-review workflow the user is on (search → read → extract → synthesize → screen → cite-check → write/review).
    2. Match it to a category below and recommend the **⭐ editor's pick first**, then 1–2 alternatives.
    3. For anything beyond the top pick — full star counts, every project in a category, or a category not summarized here — read [`reference/catalog.md`](reference/catalog.md). Do **not** guess project names or URLs; pull them from the catalog.
    4. Give a one-line "why this one" tied to the user's constraint (Claude Code vs. standalone, open vs. commercial, privacy/local, medical, etc.). If the pick is a runnable id above, offer to install/run it.
    
    ## ⚡ 30-second picker
    
    ```text
    Just need the papers themselves (topic / DOI / OA PDF) ──▶ Look up mode — no install ⭐
    Use Claude Code, want end-to-end research→paper ──────────▶ academic-research-skills ⭐
    Want AI to research a topic → cited report ───────────────▶ GPT Researcher / STORM
    Want fully autonomous "idea → submittable paper" ────────▶ AI-Scientist-v2 / AutoResearchClaw
    Citation-backed Q&A over a pile of PDFs ──────────────────▶ PaperQA2
    Rigorous PRISMA review (thousands of abstracts) ─────────▶ ASReview / prismAId
    Clean Markdown from PDFs to feed an LLM ─────────────────▶ MinerU / Docling / marker
    Lit capabilities inside Claude / Cursor (MCP) ───────────▶ paper-search-mcp / zotero-mcp
    Chat with your library inside Zotero ────────────────────▶ zotero-gpt / PapersGPT
    Pre-submission AI peer review ───────────────────────────▶ open_reviewer / ai-peer-review
    ```
    
    ## Categories (top pick per category)
    
    | Category | Editor's pick ⭐ | When |
    |---|---|---|
    | All-in-one research agents & skills | **academic-research-skills** | Claude Code user wanting research→write→review→revise, with integrity/citation gates |
    | Deep research & auto-survey | **STORM** / **gpt-researcher** | Topic → cited survey / report / related-work |
    | Autonomous science (idea→paper) | **AI-Scientist(-v2)** / **AutoResearchClaw** | Fully automated discovery: lit + hypotheses + experiments + writing |
    | Literature Q&A / RAG | **paper-qa (PaperQA2)** | Citation-backed answers over a PDF corpus |
    | Systematic review & screening | **ASReview** | Active-learning screening of thousands of abstracts (PRISMA) |
    | MCP servers | **zotero-mcp** / **arxiv-mcp-server** | Wire papers into Claude / Cursor / Cline |
    | Zotero / Obsidian integration | **zotero-gpt** | Chat with your library inside your reference manager |
    | PDF → structured extraction | **MinerU** / **docling** / **marker** | Turn PDFs into clean Markdown/JSON for LLMs |
    | Citation graphs & API clients | **scholarly** / **pyalex** | Citation-network analysis; scripting academic DBs |
    | Writing & peer-review assistants | **open_reviewer** / **ai-peer-review** | Draft, polish, and pre-submission review |
    | Awesome lists | **Awesome-Auto-Research-Tools** | Browse the whole landscape |
    
    ## Decision table (map need → recommendation)
    
    | User's need | Recommend |
    |---|---|
    | Claude Code, end-to-end research→paper | **academic-research-skills** (most complete, #1 in space) |
    | Generic "research this topic for me" agent | **GPT Researcher** / **STORM** |
    | Wiki/survey-style long-form with citations | **STORM / Co-STORM** |
    | Fully autonomous "idea → submittable paper" | **AI-Scientist-v2** / **AutoResearchClaw** |
    | Cited Q&A over many PDFs | **PaperQA / PaperQA2** |
    | Rigorous PRISMA systematic review | **ASReview** or **prismAId** |
    | PDF → clean Markdown for an LLM | **MinerU / Docling / marker** |
    | Lit capabilities in an MCP client | **paper-search-mcp / zotero-mcp** |
    | Chat with library inside Zotero | **zotero-gpt / PapersGPT** |
    | AI pre-review before submission | **open_reviewer / ai-peer-review** |
    | Just want to browse the landscape | The **Awesome lists** section |
    
    ## Notes & caveats
    
    - **Open-source is prioritized.** Commercial/closed tools (Elicit, Consensus, Scite, SciSpace, Research Rabbit, Connected Papers) are listed for reference only — see the catalog's commercial section.
    - **Star counts drift.** The catalog's numbers are periodic GitHub-API snapshots — treat as rough popularity signals, not exact. For live numbers, point the user at the repo.
    - **Match the constraint, not just the task.** Privacy/local → `local-deep-research`; medical → `medsci-skills` / `paperai`; Codex instead of Claude → `academic-research-skills-codex`.
    
    Full catalog with every project, star count, and one-line description: [`reference/catalog.md`](reference/catalog.md).
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related