Claude Skill

academic-aio

Medical AI paper optimization for AI search engines (Perplexity, ChatGPT web, Elicit, Consensus, SciSpace) and RAG-based literature tools. Applies when drafting or reviewing titles, abstracts, structured summary boxes (Key Points / Research in Context / Plain-Language Summary), m

LLM Mart · 0 points · 3 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download aperivue-medsci-skills-skills_academic-aio-815765c.zip · 46 KB
Part of aperivue/medsci-skills — 47 skills

Install

skills CLI npx skills add https://github.com/Aperivue/medsci-skills/tree/main/skills/academic-aio
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install aperivue-medsci-skills@llmmart
Git git clone https://github.com/Aperivue/medsci-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole aperivue/medsci-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Academic AIO Skill — Medical AI Paper Visibility for AI Search Engines

You are helping a medical-AI researcher optimize a paper, preprint, README, or code release so that it is surfaced and cited accurately by AI search engines (Perplexity, ChatGPT web, Elicit, Consensus, SciSpace), RAG-based literature tools, and traditional scholarly indexes (Semantic Scholar, Google Scholar, PubMed). Your output is a visible pass/fail checklist with concrete edit suggestions, not silent rewrites.

Communication Rules

  • Surface the checklist in the response. Never apply AIO edits silently.
  • Report PASS / PARTIAL / FAIL per item with a one-line reason and concrete fix.
  • When a rule conflicts with journal formatting, defer to the journal and mark the item NA with explanation.
  • Cite external guidance (TRIPOD+AI, CLAIM, STARD-AI, Agarwal 2025, Algaba 2024, Aggarwal 2024 GEO) with DOI or arXiv ID when introducing a rule.
  • Do not hallucinate citations. If unsure, mark as [VERIFY].

When to Invoke

Run this skill when the user is working on any of:

  • Drafting or revising a title, abstract, structured-summary box, or plain-language summary.
  • Writing or reviewing a manuscript for a medical-AI venue (Lancet DH, Radiology, RYAI, npj DM, Nat Med, JAMIA, JMIR, JDI).
  • Preparing a preprint (medRxiv, arXiv, bioRxiv, Research Square).
  • Composing a GitHub README, CITATION.cff, Zenodo archive metadata, Hugging Face model card, or dataset card.
  • Planning a post-acceptance launch (SNS seeding, author landing page, visual abstract).
  • Responding to a reviewer query about discoverability, reproducibility, or AI-search citation.

Pairs with (do not duplicate):

  • write-paper — Phase 6 (draft) and Phase 7 (QC). AIO rules extend the title/abstract/discussion sections.
  • check-reporting — reporting-guideline item audit (TRIPOD+AI, CLAIM, etc.). AIO requires guideline adherence but does not reproduce the audit.
  • self-review — adversarial review. Run AIO after self-review so QC-confirmed claims anchor the checklist.
  • humanize — AI-pattern removal. Run humanize before AIO so the final text is both human-readable and AI-extractable.

Core Thesis

Generative engine optimization research (Aggarwal 2024, arXiv:2311.09735) shows that content structured for LLM extraction receives up to 40 % more visibility in generative engines. In medicine this effect is mediated by three gates:

  1. Open-access full text — tools like Elicit and Consensus cannot extract columns from paywalled PDFs; Perplexity Academic favors OA citations.
  2. Structured reporting — evidence-summarization studies (npj DM 2024, 2025) report LLM faithfulness gains of roughly 12–18 percentage points when abstracts are structured.
  3. Machine-readable artifacts — CITATION.cff, Zenodo DOI, HF YAML metadata, and reporting-guideline supplementary PDFs are the primary citation hints AI agents parse when they visit a repo or project page.

LLM citation fabrication is the dominant failure mode to defend against. Agarwal et al. (Nat Commun 2025, doi:10.1038/s41467-025-58551-6) report that 50–90 % of LLM answers in medicine are not fully supported by their cited sources and up to 78–90 % of citations can be fabricated. The defensive strategy is to surface a paper's DOI and PMID in easy-to-copy form so that LLMs substitute the correct identifier instead of confabulating one.

Section 1 — Title and Abstract Optimization

1.1 Title three-slot rule

Structure: [Task] + [Modality or anatomy] + [Model family or method class]. Include one concrete differentiator (dataset scale, new benchmark, "first …") when defensible. Avoid keyword stuffing (penalized as spam by AI overviews).

Examples:

  • PASS: "Transformer-based segmentation of skull fractures on non-contrast head CT."
  • FAIL: "A novel advanced deep-learning AI machine-learning framework for medical image analysis."

1.2 Structured abstract

Use the journal-required structure (Background / Methods / Findings / Interpretation for Lancet family; Background / Purpose / Materials and Methods / Results / Conclusion for RSNA family; etc.). If the journal allows unstructured, still use an internally structured form. Each section stands alone as a semantic chunk of ≤ 3 sentences so that chunk-boundary splits in RAG indexes do not break the claim.

1.3 Opening and closing sentences

  • First sentence: state the problem AND the contribution in one line. LLM summarizers extract this disproportionately.
  • Last sentence: explicit interpretation ("we show that …", "this implies …"). No hedging-only closes.

1.4 Taxonomy line

Include one sentence that names the field's controlled vocabulary (for example, "diagnostic-accuracy study", "foundation-model evaluation", "LLM-as-judge", "agentic radiology workflow"). Entity linkers in AI indexes use this line.

1.5 Quantified claim

Every abstract must contain at least one numeric primary outcome with confidence interval (for example, "AUC 0.94 [95 % CI 0.91–0.96]" or "sensitivity 88.2 % [95 % CI 85.1–91.0]"). LLM retrievers weight papers with concrete numbers.

1.6 Reporting-guideline anchor

Place the guideline name in the abstract or the opening sentence of Methods: "Reported following TRIPOD+AI (Collins 2024) and CLAIM 2024 (Tejani 2024)". When applicable add STARD-AI 2025, DECIDE-AI, TRIPOD-LLM. This signals structure to LLMs and satisfies reviewer checklists.

AIO-rule ↔ guideline-item mapping: references/reporting_guideline_mapping.md.

1.7 Keyword, MeSH, and RadLex coverage

Title, abstract, and keywords together should cover ≥ 3× the surface area of the concept — no redundancy. Include:

  • Core MeSH terms (verify against the NLM MeSH browser).
  • Radiology-specific RadLex terms where applicable.
  • Modality-synonym coverage ("chest radiograph (CXR)", "non-contrast CT (NCCT)").
  • Both US and UK spellings when relevant.

Royal Society 2024 (doi:10.1098/rspb.2024.1222) reports that 92 % of papers waste keyword real estate by repeating title terms in abstract and keywords; avoid this.

Section 2 — Manuscript-Level AIO

2.1 Summary box

Include the journal-specific summary box verbatim when supported:

  • Lancet family: "Research in context" (Evidence before this study / Added value / Implications).
  • RSNA Radiology and RYAI: "Key Points" — 3 bullets, one claim each.
  • npj Digital Medicine: "Plain-language summary" (150–200 words, 8th-grade reading level).
  • Nature Medicine: editor's summary (supplied by editorial, but draft one proactively).

Deterministic format check. Validate the drafted box against its journal spec with python3 ${CLAUDE_SKILL_DIR}/scripts/check_summary_box.py --manuscript <file> --journal <stem> --strict (reads references/summary_box_specs.json: Key Points bullet count + one-claim-per-bullet, Research-in-context's three sub-blocks, plain-language word band). It catches the wrong-format / wrong-bullet-count box that a production technical check rejects.

These boxes are the fragments Perplexity and ChatGPT web most often copy or paraphrase verbatim; treat them as the paper's canonical citation surface.

Journal-specific templates (USER MUST VERIFY against current IFA): references/journal_summarybox_templates.yaml.

2.2 Declarative section headings

Section and subsection headings should state a claim, not a generic label. "Model underperforms on rare-finding subset" beats "Subgroup analysis".

2.3 Numeric claim compression

In the Methods and in at least one Results paragraph, compress primary-outcome statistics into a single sentence pattern: "On the internal test set (n = 842), the model achieved AUC 0.94 (95 % CI 0.91–0.96), sensitivity 88.2 % (85.1–91.0), specificity 91.4 % (88.7–93.6), at an operating point of 0.37."

This pattern is the canonical shape LLM extractors parse first.

2.4 Reproducibility block

Include a labeled block (typically end of Methods or a standalone Data/Code Availability section) listing: data availability and license, code availability with DOI, model weights and checkpoints, prompts and configuration files, random seeds, compute environment. This block is disproportionately scraped by AI agents when they cite a paper as reproducible.

2.4a Citing an AI-assisted tool by use-class

If the work used an AI-assisted tool (a verification/QA suite, analysis code, or a generative drafting assistant), frame the mention by what it did, not by hiding it — under current journal wariness, a proud in-text citation of a generative use invites suspicion, while the same placement for a verification use reads as rigor (like citing a reference manager or a linter). Split the mention: verification/QA and analysis → a Software / Code-availability statement (citable); generative drafting/humanizing → the journal's AI-use disclosure field, not a citation. A self-citation by the tool's author additionally requires a COI disclosure, and you should cite only the functions the work actually used. Full use-class table and rules: ${CLAUDE_SKILL_DIR}/references/ai_tool_citation_framing.md.

2.5 Limitations enumeration

List limitations explicitly and name each one (generalizability, spectrum bias, dataset shift, single-center training, label noise). Papers with enumerated limitations score higher for trustworthiness in LLM summarization benchmarks.

2.6 Standalone figure captions

Each caption should re-state the claim, the dataset, and the metric. Captions survive in vector databases and image-retrieval indexes when surrounding body text is lost.

Section 3 — Preprint, Channel, and Indexing Strategy

3.1 Preprint versus fast-track

  • Default: post to medRxiv (clinical), arXiv (methods, cs.CV / eess.IV), or bioRxiv on the day of journal submission. Rapid preprinting puts the paper into Semantic Scholar within 24–72 hours and into Perplexity's web index immediately.
  • Exception: if the target journal offers a fast-track review cycle (acceptance → online within roughly 30–60 days) AND the authors prefer a single canonical version, a preprint may be skipped. In that case, compensate by aggressive post-acceptance SNS seeding and PMC deposit.
  • Never skip preprint AND fast-track — this is the discoverability deadzone.

3.2 Journal preprint-policy table (verify before submission)

Most medical-AI venues allow preprints (Radiology, RYAI, Lancet DH, npj DM, Nature Medicine, JAMIA, JMIR, Cell Reports Medicine, Cell Patterns). A few have restrictions or require disclosure. Always verify the current policy on Sherpa Romeo or the journal's instructions-for-authors page before posting.

3.3 Indexing time-lag (2025 baseline)

  • Perplexity Academic / ChatGPT web: real-time web crawl, citable on publication day.
  • Semantic Scholar: 24–72 hours from DOI or preprint.
  • Google Scholar: 1–7 days.
  • PMC (NIH deposit): 2–6 weeks for accepted manuscripts; longer for CC-BY-NC.
  • Elicit and Consensus: follow Semantic Scholar / OpenAlex.
  • LLM training corpora (next model generation): 6–18 months.

Plan launch activities around these windows.

3.4 Open-access choice

Prefer gold OA with CC-BY when budget allows. If not, green OA via preprint plus author-accepted manuscript is acceptable. Closed-access papers without preprint lose roughly 30–50 % of AI-tool citations because Elicit, Consensus, and Perplexity Academic cannot extract from paywalled PDFs.

Funder OA-policy decision tree (Plan S, NIH, UKRI, Gates, Wellcome, NRF, MoHW): references/oac_funding_checklist.yaml.

3.5 Post-acceptance channel checklist

  • Deposit AAM to PMC or Europe PMC.
  • Update ORCID and Google Scholar profile.
  • Post to Threads / X / BlueSky with DOI, one-sentence claim, and figure.
  • Long-form post on LinkedIn (targets different LLM training corpora).
  • Submit to Papers with Code if the paper reports a benchmark.
  • Upload model or dataset to Hugging Face with a model/dataset card.

Section 4 — Review-Paper Strategy

Review articles function as hub nodes in knowledge graphs and accrue "lookup citations" when readers need a canonical reference for a taxonomy. For researchers building a portfolio in medical AI:

  • Target at least one review or taxonomy paper per year in a top-tier venue.
  • Include 5 or more taxonomy tables (model class, dataset, task type, evaluation metric, failure mode). Each table becomes a lookup target.
  • Cite 100 or more primary references for breadth; 150+ for canonical status.
  • Co-author with a consortium of 10+ investigators from multiple institutions when possible — this multiplies social-network reach and citation dispersal.
  • Pair the review with a companion dataset, benchmark, or code artifact on Zenodo or Hugging Face to anchor AI-tool citations.

Empirically, review papers with these properties outperform original research on short-term FWCI while feeding traffic to the authors' original papers through reverse citation.

Section 5 — GitHub, CITATION.cff, Zenodo, Hugging Face

5.1 README canonical 10-slot order

  1. Title + one-line description + badges (license, DOI, arXiv, Hugging Face, paper link).
  2. Paper reference block — BibTeX + APA + two-sentence abstract.
  3. TL;DR — at most 5 bullets: problem, approach, key result, intended users.
  4. Quickstart — pip install or git clone && make demo. Should work in under 5 minutes.
  5. Reproducibility — exact commands that regenerate every figure and table. Pin package versions.
  6. Project structure — a tree with one-line folder descriptions.
  7. Data access — license, download scripts, DUA notes.
  8. FAQ — "How is this different from X?", "Can this be used clinically?", "How do I cite this?". High-value retrieval content.
  9. Acknowledgements, funding, and COI.
  10. License (prefer Apache-2.0 for research code).

5.2 CITATION.cff

Add a CITATION.cff file at repository root. GitHub renders it as a "Cite this repository" button, and AI agents treat it as the primary citation hint. Include authors with ORCID, title, version, DOI (post-Zenodo-archive), repository URL, and license.

5.3 Zenodo DOI

Enable GitHub–Zenodo integration for each release. Cite the version-specific DOI in the paper's Data/Code Availability section. Zenodo deposits appear in Google Scholar and OpenAlex, creating an independent citable artifact.

5.4 Hugging Face model card YAML

Required keys: license, library_name, tags, datasets, base_model (when fine-tuning), pipeline_tag. Required prose sections: Intended use, Training data, Evaluation, Limitations, Ethical considerations, and a clinical-use disclaimer ("This model is not approved for clinical diagnostic use; it is provided for research purposes only").

5.5 Hugging Face dataset card

Required prose: license, PHI and re-identification risk, task, language, splits, annotation process, known biases, ethical review status.

5.6 Web-crawler-friendly formatting

  • Markdown headings are declarative claims.
  • Code blocks are fenced and language-tagged.
  • Tables are plain Markdown, not HTML (survive Markdown-to-vector chunking).
  • Images have descriptive alt text (vision-LLMs read alt text when image retrieval fails).
  • Each README section is under about 300 words to survive fixed-size chunking.
  • Use question-style subheadings when natural ("Why another benchmark?", "How fast is inference?").
  • Embed JSON-LD ScholarlyArticle / SoftwareSourceCode / Dataset / Person markup in repository pages and author landing pages — templates in references/schema_markup_templates/, validated with python scripts/validate_schema.py path/to/file.jsonld.

Section 6 — Authority and E-E-A-T Signals

  • Maintain a personal author landing page (GitHub Pages, personal domain, or institutional page) that lists all papers with DOIs and open-access links. AI indexes weight author-entity pages.
  • Use one consistent affiliation string across papers. Inconsistency fragments the author entity in knowledge graphs and loses citation velocity.
  • Keep ORCID complete and linked to Google Scholar. Re-run author-disambiguation on Semantic Scholar every 6 months.
  • Cross-link related papers by the same group in Discussion sections when defensible. Within-group self-citation increases co-retrieval probability in RAG.
  • Refresh repository and model cards quarterly — articles updated quarterly outperform single-publish articles in AI-overview retention (Conductor 2026 benchmark).

Section 7 — LLM-Citation Fabrication Defense

Given Agarwal et al. Nat Commun 2025 (doi:10.1038/s41467-025-58551-6) findings that up to 78–90 % of LLM medical citations can be fabricated, take the following defensive steps:

  • Surface DOI and PMID in copy-friendly text at the top of the paper's landing page and README (for example, DOI: 10.xxxx/yyyy • PMID: 12345678).
  • Add a "How to cite" section with BibTeX, APA, Vancouver, and the plain-text line in one place.
  • Monitor incorrect citations. Set a Google Scholar alert for the paper's title variant; periodically query Perplexity and ChatGPT web for the paper and record hallucinated bibliographic errors.
  • When responding to a reviewer who cites an LLM-generated reference, verify the DOI and PMID yourself before accepting.

Section 8 — Red Flags

  • Closed code described as "available on reasonable request" — scrapers treat this as "not reproducible" and AI tools demote the paper.
  • Paywall-only with no preprint and no fast-track — invisible to most RAG pipelines.
  • Keyword-stuffed titles ("A deep-learning artificial-intelligence machine-learning system for …") — penalized as spam.
  • Abstracts opening with filler ("In recent years, AI has revolutionized …") — burns the chunk most likely to be extracted.
  • Walls of theory before the README quickstart.
  • "Clinical grade" or "replaces radiologists" overclaims — demoted by LLM trust heuristics and may trigger reviewer rejection.
  • PHI leakage in Hugging Face dataset samples.
  • Inconsistent author affiliations across co-authored papers.

Section 9 — Per-Project Application and Pipeline Integration

When invoked, run in this order:

  1. Read the target artifact (title, abstract, manuscript section, README, or card) and identify its lifecycle phase: pre-draft / drafting / pre-submission / post-acceptance / post-publication.
  2. Apply Sections 1–5 and 10 relevant to that artifact, filtering each rule by its applies_to_phase field in references/checklists/AIO_GENERAL.md. Out-of-phase rules become NA rather than FAIL (e.g., do not surface §11.5 multi-disciplinary roster or §12 launch sequencing as FAIL on a pre-submission audit). Produce a PASS / PARTIAL / FAIL table sorted by expected_lift (high → medium → low). Render via templates/aio_audit_checklist.md.j2 when programmatic.
  3. Honour defers_to annotations to avoid duplicate audits. Items annotated with a defers_to field record only present/absent status here; item-level detail belongs to the linked skill or reference (§1.6 → /check-reporting; §3.4 / §11.3 → references/oac_funding_checklist.yaml). Cross-check reporting-guideline anchor (§1.6) by invoking /check-reporting first when the manuscript has not been audited; the AIO ↔ guideline-item mapping is in references/reporting_guideline_mapping.md.
  4. Apply Section 6 author-authority audit once per submission cycle. Sections 11.1–11.5 are pre-draft rules — applies_to_phase filter auto-NAs them once drafting is complete. Section 12 launch sequencing fires only at post-acceptance / post-publication.
  5. Surface Section 7 citation-defense recommendations at post-acceptance time. For multi-repo or Hugging-Face-card team audits, run scripts/batch_metadata_audit.py.
  6. Output:
    • The phase-filtered checklist (visible).
    • A short deferred-item list with one-line status per defers_to rule.
    • At most 5 concrete edits ranked by expected_lift (high first, then medium, then low). Edits whose underlying rule is low-lift should not appear in the Top 5 unless no high / medium items remain open.

Integration with write-paper

  • Phase 4 (Title and abstract drafting) → apply Section 1 as an inline filter.
  • Phase 6 (Discussion) → apply Section 2.5 (limitations) and Section 6 (cross-linking).
  • Phase 7 (QC) → run AIO after reporting-guideline check and numerical-claim audit.

Output template

## Academic AIO Checklist — [Artifact type]

| # | Item | Status | Note |
|---|------|--------|------|
| 1.1 | Title three-slot | PASS/PARTIAL/FAIL | … |
| 1.2 | Structured abstract | PASS/PARTIAL/FAIL | … |
| ... | ... | ... | ... |

## Top 5 suggested edits
1. …
2. …

Section 10 — Q&A and Entity-Extraction Optimization

Modern RAG indexes parse Q&A blocks more reliably than free-form prose; LLM citation engines preferentially extract claim-restatement pairs. Section 10 augments retrievability by structuring how claims are restated and how entities are linked.

10.1 Four-question Q&A block (Discussion or Appendix)

Add a labeled Q&A block — either as the closing subsection of Discussion, or as a Supplementary Box. Pattern:

  • What was known before this study? — two-sentence restatement of the prior state.
  • What does this study add? — two-sentence statement of the contribution.
  • How might this change clinical practice or research? — one-sentence interpretation; avoid overclaim.
  • Why does this matter? — one-sentence "so what" framing for non-specialists.

This block is the canonical fragment that AI-overview systems extract and cite. Lancet Digital Health "Research in context" already encodes the first two questions; the Q&A block extends them and is parseable by LLM web-search agents.

10.2 Glossary block with entity IDs

Define each domain-specific acronym inline on first use AND list them in a Glossary subsection at end of Methods or Supplementary. Attach the canonical entity ID where possible:

  • MeSH term ID for clinical concepts.
  • RadLex ID for radiology-specific terms.
  • UMLS CUI for cross-vocabulary mapping.
  • Hugging Face model ID for named models.
  • arXiv ID for cited methods.

Entity linkers in Elicit, Consensus, and SciSpace use this metadata to connect a paper to knowledge graphs.

10.3 Inline citation anchor text

Avoid bare reference numbers. Use semantic anchor patterns so LLM extractors bind the citation to the specific claim:

  • WEAK: "Prior work [12] showed efficacy."
  • STRONG: "Smith et al. (DOI: 10.xxxx/yyyy) reported a 12 % accuracy gain on the MIMIC-CXR test set [12]."

When citing one's own prior work, name the cohort or dataset explicitly to enable cross-paper retrieval.

10.4 Explicit challenge statement

Beyond Section 2.5 (limitations enumeration), include a single-paragraph "Why this is hard" challenge statement near the start of Discussion. Pattern:

"Building accurate [task] for [modality/anatomy] is constrained by [data scarcity / label noise / dataset shift / regulatory uncertainty / interpretability]. Each of these has been documented [refs], and our results address [subset]."

LLM web-search systems quote challenge statements as authoritative summaries of field state. The 2025 KJR multimodal-LLM review used this pattern (e.g., "lack of large-scale high-quality multimodal datasets") and was preferentially extracted by Perplexity and ChatGPT web (see references/case_studies/kjr_mllm_2025.md).

Section 11 — First-Mover Timing and Citation-Graph Density

Topic timing is the most under-discussed AIO lever. Reviews and original research published at the peak of a topic's hype curve accrue citations disproportionately; reviews that lag the peak by 6–12 months under-perform regardless of quality.

11.1 Topic peak detection

Signals that a topic is approaching peak (write now, publish in ~6 months):

  • arXiv/medRxiv monthly deposit rate growing > 20 % month-over-month for 3+ consecutive months.
  • Major model release (GPT-4o, Claude 3.5 multimodal, MedGemini) introducing a capability not previously available.
  • Funding agency Request-for-Applications (RFA) addressing the topic.
  • Society guidelines (RSNA, ACR, ESR) calling for evaluation studies.
  • Sustained > 1,000 weekly impressions on Twitter/X/LinkedIn for related papers.

Plan submission so publication lands at peak, not after.

11.2 Editorial-board leverage

If a corresponding author serves on the target journal's editorial board, review-process median time often drops noticeably (KJR: ~4–6 weeks faster; varies by journal). Editor's-pick or issue-highlight selection can also drive Google News indexing within 24 hours of publication.

When recruiting senior co-authors for a review paper, prefer those who hold an editorial role at the target venue. This is a legitimate editorial signal, not a conflict-of-interest issue, provided board members recuse themselves from review of their own submissions per ICMJE guidance.

11.3 PMC-auto-deposit journal preference

Open-access journals that automatically deposit to PubMed Central (PMC) reach LLM crawlers within 4–6 weeks of publication; non-PMC OA journals can take 3–6 months. PMC-auto-deposit journals in radiology/medical-AI (verify per submission, policies change):

  • Korean Journal of Radiology (KJR) — auto-deposit confirmed.
  • Lancet Digital Health — author-funded green OA, PMC-eligible after embargo.
  • Radiology and Radiology: AI — selected articles auto-deposit.
  • npj Digital Medicine — auto-deposit (Nature OA).
  • JAMIA — author-funded OA route.
  • JMIR — auto-deposit (PMC-indexed).

When all else is equal, prefer PMC-auto-deposit journals to compress the LLM-discoverability window.

11.4 Citation-graph anchor strategy

Discussion sections should anchor the paper in 5–10 high-visibility prior works that LLM training corpora already index well. This raises co-citation probability and makes the paper retrievable when users query the seminal works.

  • Identify seminal references via Semantic Scholar's "Highly Influential Citations" filter for the topic.
  • Cite them with semantic predicates (Section 10.3), not as bare lists.
  • Mix recent preprints (currency signal) with 2018–2022 seminal papers (graph anchoring) — corpora-cutoff means 2024–2025-only citation profiles have low LLM retrieval weight.

11.5 Multi-disciplinary author roster

Author-affiliation diversity multiplies indexing entry points. A 10–15 author team spanning 3+ institutions and 2+ disciplines (clinical + computational) creates more author-entity nodes in Google Scholar and Semantic Scholar, each acting as a discovery surface. The 2025 KJR MLLM review used a 15-author team spanning resident + engineer + medical student + faculty across 5 institutions and accrued 64 citations within 7 months (see case study).

Section 12 — Cross-Platform Launch Sequencing

Section 3.5 (post-acceptance channel checklist) is unordered; Section 12 prescribes the timing. The first 30 days after publication are the primary discoverability window for AI-search engines and LLM training-data harvesters.

12.1 Day 0 — publication day (execute simultaneously)

  • GitHub release (tag a stable version; let Zenodo mint a version-specific DOI).
  • Hugging Face model card + dataset card (if applicable); link arXiv ID and DOI.
  • Twitter/X + Threads + Bluesky: 1-sentence claim + key figure + DOI in copy-friendly format.
  • LinkedIn announcement (long-form): hook line + structured claim block + DOI.
  • Author landing-page update with PDF link (OA) or AAM.

12.2 Day 1 — propagation

  • Update ORCID with DOI, abstract, and authorship role.
  • Update Google Scholar (verify auto-detection within 24h; manual add if delayed).
  • Update preprint server with "Accepted" version note + link to published version.
  • Update institutional profile / department news page.

12.3 Week 1 — depth posts

  • LinkedIn second post: long-form interpretation or methods spotlight.
  • Papers with Code submission (if benchmark or model with public weights).
  • ResearchGate upload of AAM (per journal policy).
  • Reddit/Hacker News post if the work has broad appeal (assess fit honestly).

12.4 Weeks 2–4 — refresh signals

  • README and HF card minor update (new badges, new FAQ entries).
  • Follow-up blog or Substack post expanding on one figure or limitation.
  • Respond to reader questions on social platforms — those answers themselves become indexed content.

12.5 Month 1 — monitoring

  • Google Scholar alert for the paper title.
  • Semantic Scholar / Scite citation alerts.
  • Quarterly probe: query Perplexity, ChatGPT web, Elicit, Consensus, SciSpace with 3–5 expected discovery queries; record retrieval position and any hallucinated bibliographic errors.
  • If a fabricated citation appears, update the README "How to cite" block (Section 7) to maximize copy-friendliness of the correct identifier.

External References

  • GEO: Generative Engine Optimization — Aggarwal et al., KDD 2024, arXiv:2311.09735.
  • LLM medical citation fabrication — Agarwal et al., Nat Commun 2025, doi:10.1038/s41467-025-58551-6.
  • LLM citation bias — Algaba et al., 2024, arXiv:2405.15739.
  • ExpertQA attribution — Malaviya et al., 2024, arXiv:2309.07852.
  • TRIPOD+AI — Collins et al., BMJ 2024. EQUATOR Network.
  • CLAIM 2024 — Tejani et al., Radiology: AI 2024, doi:10.1148/ryai.240300.
  • STARD-AI — Sounderajah et al., Nat Med 2025, doi:10.1038/s41591-025-03953-8.
  • TRIPOD-LLM — Gallifant et al., Nat Med 2024, doi:10.1038/s41591-024-03425-5.
  • DECIDE-AI — Vasey et al., Nat Med 2022, doi:10.1038/s41591-022-01772-9.
  • Title, abstract, keywords guide — Royal Society Proc B 2024, doi:10.1098/rspb.2024.1222.
  • GitHub repository citation advantage — Yan et al., Inf Process Manag 2024, doi:10.1016/j.ipm.2023.103569.
  • Semantic Scholar Open Data Platform — Kinney et al., arXiv:2301.10140.

Anti-Hallucination

  • Never fabricate citations, DOIs, arXiv IDs, or reporting-guideline item numbers. Every cited reporting framework (TRIPOD+AI, CLAIM, STARD-AI, TRIPOD-LLM, DECIDE-AI) must map to a verifiable DOI or EQUATOR Network entry. Mark unverified items as [UNVERIFIED - NEEDS MANUAL CHECK].
  • Never invent journal-specific summary-box rules (Lancet Digital Health "Research in context", Radiology "Key Points", npj Digital Medicine). Verify current instructions-to-authors from the journal's website before applying.
  • Never fabricate discoverability metrics (Perplexity/Elicit/Consensus retrieval scores) — only report observed behavior from a recorded probe.
  • Never auto-complete author lists, ORCIDs, or affiliations in CITATION.cff or Zenodo metadata; surface empty slots to the user.
  • If a compliance item, journal policy, or AI-search platform behavior is uncertain, state the uncertainty rather than guessing.

Global-rule references

Some passages in this skill cite a path of the form ~/.claude/rules/<name>.md. Those are the maintainer's personal global rules, kept outside this repository. They are not shipped with this skill and will not exist on your machine; they appear only as provenance for where a convention came from. If one of them looks like it is standing in for an instruction you actually need, that is a bug — please open an issue, because the instruction belongs here.

Files (medsci-skills)
  • references
    • case_studies
      • kjr_mllm_2025.md 6.3 KB
        # Case Study — KJR Multimodal LLM Review (2025)
        
        > Post-mortem of an unexpectedly high-engagement medical-AI review. Used to validate Sections 10–12 of the academic-aio skill. Identifying co-author and institutional details are abstracted; the paper itself is publicly indexed.
        
        ## Paper
        
        - **Title**: Multimodal Large Language Models in Medical Imaging: Current State and Future Directions
        - **Venue**: Korean Journal of Radiology, Vol 26, Issue 10, pp. 900–923 (October 2025)
        - **DOI**: 10.3348/kjr.2025.0599
        - **PMC ID**: PMC12479233
        - **OA status**: Creative Commons (gold OA), KJR + PMC dual indexing
        - **Author roster**: 15-author multi-disciplinary team — resident + medical-AI researcher (first author), software engineer, medical student, faculty across 5 Asian academic medical institutions and one university computational biology department, anchored by a mid-career corresponding author who serves on the target journal's editorial board (Technology section).
        - **Timeline**: Received 2025-05-14 → revised 2025-07-03 → accepted 2025-07-08 → published October 2025 (~5 months end-to-end, fast for a review article).
        
        ## Engagement metrics (May 2026, ~7 months post-publication)
        
        - Page views: 3,016
        - PDF downloads: 628
        - Citations: 64
        - First-author Google Scholar impact (since 2021): h-index 4, i10-index 2, total 101 citations — of which ~63 % are attributable to this single review.
        
        These figures place the paper in the top decile of KJR articles by 12-month citation accrual.
        
        ## Driver analysis (which AIO levers actually operated)
        
        The author's initial hypothesis space included (a) journal visibility, (b) hot topic, (c) corresponding author's reach, (d) multi-disciplinary team, (e) fast turnaround, (f) OA + PMC indexing, and (g) early SNS exposure. Post-hoc evidence:
        
        | Hypothesis | Verdict | Notes |
        |------------|---------|-------|
        | AI-search retrievability (Perplexity / ChatGPT web / Elicit / Consensus / SciSpace) | **Strong (primary driver)** | Paper appears at #3–4 for "MLLM medical imaging review 2025" on Google web; Perplexity preferentially cites it for queries on multimodal radiology AI. |
        | OA + PMC dual indexing | Strong | Full text crawled by AI agents within 4–6 weeks; corresponds to Section 11.3. |
        | Topic peak timing | Strong | Submitted shortly after GPT-4o and Claude 3.5 multimodal launches; published at the peak of the 2025 MLLM hype cycle. Corresponds to Section 11.1. |
        | First-mover review | Strong | Competing reviews (in JBI, Archives of Comp Methods, ScienceDirect) appeared later or in less-discoverable venues. |
        | Editorial-board signal | Moderate | Corresponding author's editorial role plausibly accelerated review and raised editor's-pick probability. Corresponds to Section 11.2. |
        | Multi-disciplinary 15-author team | Moderate | Author-entity diversification across institutions multiplied Google Scholar entry points. Corresponds to Section 11.5. |
        | Early SNS exposure | Weak | No clear evidence that SNS drove the bulk of traffic; KJR/PMC pathway dominated. |
        
        ## Structural features that AI-search systems extracted
        
        The article body exhibits several patterns that map to Sections 10–11 of this skill:
        
        1. **Explicit taxonomy** (2D vs 3D MLLM; Applications vs Barriers) — RAG systems chunk-cite each branch independently.
        2. **Verbatim challenge statements** — phrases such as "lack of large-scale high-quality multimodal datasets" and "hallucinated findings" are quoted by Perplexity and ChatGPT web as authoritative summaries of field state. Corresponds to Section 10.4.
        3. **Declarative subsection headings** — claim-style rather than generic labels (Section 2.2).
        4. **10 figures + 5 tables** — rich pull-quotable structures for retrieval.
        5. **DOI + PMID surfaced cleanly** in KJR landing page header — minimizes citation fabrication (Section 7).
        
        ## Lessons codified into the skill
        
        | Skill update | Source observation |
        |--------------|---------------------|
        | §10.4 Explicit challenge statement | Field-state quotes were the most common Perplexity extraction. |
        | §11.1 Topic peak detection | Submission timing aligned with model-release shockwave. |
        | §11.2 Editorial-board leverage | Review-process speed plausibly editor-mediated. |
        | §11.3 PMC-auto-deposit preference | KJR-PMC pathway delivered LLM crawl in 4–6 weeks. |
        | §11.4 Citation-graph anchor | Discussion seeded with a mix of seminal multimodal-LLM works. |
        | §11.5 Multi-disciplinary roster | 15-author 5-institution composition multiplied entry points. |
        
        ## Replication checklist (apply to next medical-AI review)
        
        - [ ] Identify a topic 6–12 months ahead of expected peak using Section 11.1 signals.
        - [ ] Recruit a corresponding author with editorial-board affiliation at a PMC-auto-deposit OA journal (Section 11.2 + 11.3).
        - [ ] Assemble a 10–15 author team spanning ≥ 3 institutions and ≥ 2 disciplines (Section 11.5).
        - [ ] Draft taxonomy headings as declarative claims (Section 2.2) with explicit 2D-vs-3D-style subdivisions.
        - [ ] Insert a "Why this is hard" challenge paragraph at the start of Discussion (Section 10.4).
        - [ ] Add a four-question Q&A block in Discussion or Appendix (Section 10.1).
        - [ ] Anchor Discussion in 5–10 highly-cited prior works (Section 11.4).
        - [ ] Execute Day-0/Day-1/Week-1 launch sequencing (Section 12).
        - [ ] Probe Perplexity / ChatGPT web / Elicit / Consensus / SciSpace at +30, +60, +90 days post-publication; correct any fabricated citations encountered (Section 7).
        
        ## Caveats
        
        - Single-paper case study — not a controlled experiment. Hot-topic timing alone could explain a large fraction of the engagement; the AIO mechanisms identified here are necessary but not provably sufficient.
        - Citation accrual at 7 months is a leading indicator, not a final one. A second readout at 24 months will distinguish landscape-review citation behavior from short-term hype.
        - Editorial-board involvement is institution-specific. Replicating Section 11.2 requires evaluating the target journal's recusal policy and disclosing it appropriately.
        
        ## References
        
        - KJR landing page: `https://www.kjronline.org/DOIx.php?id=10.3348/kjr.2025.0599`
        - PMC full text: `https://pmc.ncbi.nlm.nih.gov/articles/PMC12479233/`
        - GEO framework: Aggarwal et al., KDD 2024, arXiv:2311.09735.
        - LLM medical citation fabrication: Agarwal et al., Nat Commun 2025, doi:10.1038/s41467-025-58551-6.
        
    • checklists
      • AIO_GENERAL.md 13.2 KB
        # Academic AIO General Checklist
        
        > Machine-readable checklist mirroring `SKILL.md` sections 1–12. Use as the canonical pass/fail audit table for any medical-AI artifact (manuscript, preprint, README, CITATION.cff, HF card). Each item maps to a SKILL.md section so audit findings can be traced back to the rule.
        
        ## How to use
        
        1. Copy this checklist into the working artifact directory as `qc/aio_audit.md`.
        2. Determine the artifact's **lifecycle phase** (`pre-draft`, `drafting`, `pre-submission`, `post-acceptance`, `post-publication`) and filter rules whose `applies_to_phase` does not include the current phase — those become NA (do not surface as FAIL).
        3. Mark each remaining item PASS / PARTIAL / FAIL with a one-line reason.
        4. For FAIL items, generate a concrete edit suggestion ranked by `expected_lift` (`high` first, then `medium`, then `low`).
        5. Re-run after edits until ≥ 90 % of applicable items pass.
        6. Where an item carries a `defers_to` link, item-level detail belongs to the linked skill — record only the high-level status here to avoid duplicate audits.
        
        ## Schema (v2)
        
        ```
        items[].id              : §-numbered rule id
        items[].rule            : one-line description
        items[].applies_to      : artifact types where the rule fires
        items[].applies_to_phase: lifecycle phases where the rule is actionable
        items[].priority        : H | M | L  (severity if missing)
        items[].expected_lift   : high | medium | low  (KJR-case-validated visibility gain)
        items[].defers_to       : optional skill / file owning item-level detail
        ```
        
        ## Checklist
        
        ```yaml
        checklist:
          schema_version: 2
          last_updated: "2026-05-11"
        
          metadata:
            artifact_path: ""
            artifact_type: ""        # manuscript | preprint | readme | citation_cff | hf_card | dataset_card
            artifact_phase: ""       # pre-draft | drafting | pre-submission | post-acceptance | post-publication
            journal_target: ""
            audit_date: ""           # YYYY-MM-DD
            reviewer: ""
        
          items:
            # Section 1 — Title and Abstract Optimization
            - id: 1.1
              rule: Title three-slot ([Task] + [Modality/anatomy] + [Model class])
              applies_to: [title]
              applies_to_phase: [drafting, pre-submission]
              priority: H
              expected_lift: high
            - id: 1.2
              rule: Structured abstract per journal template; chunk-friendly ≤3-sentence sub-blocks
              applies_to: [abstract]
              applies_to_phase: [drafting, pre-submission]
              priority: H
              expected_lift: high
            - id: 1.3
              rule: First sentence states problem + contribution; last sentence is explicit interpretation
              applies_to: [abstract]
              applies_to_phase: [drafting, pre-submission]
              priority: H
              expected_lift: high
            - id: 1.4
              rule: Taxonomy line names controlled vocabulary (e.g., DTA, foundation-model evaluation)
              applies_to: [abstract]
              applies_to_phase: [pre-submission]
              priority: M
              expected_lift: low
            - id: 1.5
              rule: Quantified primary outcome with 95 % CI in abstract
              applies_to: [abstract]
              applies_to_phase: [drafting, pre-submission]
              priority: H
              expected_lift: high
            - id: 1.6
              rule: Reporting-guideline anchor present (TRIPOD+AI / CLAIM / STARD-AI / TRIPOD-LLM / DECIDE-AI / PRISMA-DTA)
              applies_to: [abstract, methods]
              applies_to_phase: [drafting, pre-submission]
              priority: H
              expected_lift: high
              defers_to: /check-reporting   # item-level audit; AIO records anchor-present status only
            - id: 1.7
              rule: MeSH and RadLex coverage; consistent UK/US spelling; no abstract-keyword redundancy
              applies_to: [keywords, abstract]
              applies_to_phase: [pre-submission]
              priority: M
              expected_lift: medium
        
            # Section 2 — Manuscript-Level
            - id: 2.1
              rule: Journal-specific summary box present (Research in context / Key Points / PLS / Main Points)
              applies_to: [manuscript]
              applies_to_phase: [drafting, pre-submission]
              priority: H
              expected_lift: high   # KJR case: top extracted fragment by AI overviews
            - id: 2.2
              rule: Section headings are declarative claims, not generic labels
              applies_to: [manuscript]
              applies_to_phase: [drafting, pre-submission]
              priority: M
              expected_lift: medium
            - id: 2.3
              rule: Numeric claim compression sentence in Methods and Results
              applies_to: [methods, results]
              applies_to_phase: [drafting, pre-submission]
              priority: M
              expected_lift: medium
            - id: 2.4
              rule: Reproducibility block (data, code, weights, prompts, seeds, env)
              applies_to: [methods, data_availability]
              applies_to_phase: [pre-submission]
              priority: H
              expected_lift: high
            - id: 2.5
              rule: Limitations enumerated and named
              applies_to: [discussion]
              applies_to_phase: [drafting, pre-submission]
              priority: M
              expected_lift: medium
            - id: 2.6
              rule: Standalone figure captions restate claim, dataset, and metric
              applies_to: [figures]
              applies_to_phase: [pre-submission]
              priority: M
              expected_lift: low
        
            # Section 3 — Preprint, Channel, Indexing
            - id: 3.1
              rule: Preprint posted on submission day, OR fast-track justified
              applies_to: [submission_strategy]
              applies_to_phase: [pre-submission]
              priority: H
              expected_lift: high
            - id: 3.2
              rule: Journal preprint policy verified (Sherpa Romeo or IFA)
              applies_to: [submission_strategy]
              applies_to_phase: [pre-submission]
              priority: H
              expected_lift: medium
            - id: 3.4
              rule: Open-access route selected (gold OA preferred; green OA acceptable)
              applies_to: [submission_strategy]
              applies_to_phase: [pre-draft, pre-submission]
              priority: H
              expected_lift: high
              defers_to: references/oac_funding_checklist.yaml
            - id: 3.5
              rule: Post-acceptance channel checklist scheduled (PMC, ORCID, Scholar, SNS, HF)
              applies_to: [post_acceptance]
              applies_to_phase: [post-acceptance]
              priority: M
              expected_lift: high
        
            # Section 5 — GitHub / CITATION.cff / Zenodo / HF
            - id: 5.1
              rule: README follows 10-slot canonical order
              applies_to: [readme]
              applies_to_phase: [post-acceptance, post-publication]
              priority: H
              expected_lift: medium
            - id: 5.2
              rule: CITATION.cff present at repo root with ORCID, DOI, version
              applies_to: [citation_cff]
              applies_to_phase: [post-acceptance, post-publication]
              priority: H
              expected_lift: medium
            - id: 5.3
              rule: Zenodo DOI minted via GitHub integration; cited in paper
              applies_to: [zenodo, manuscript]
              applies_to_phase: [pre-submission, post-acceptance]
              priority: H
              expected_lift: high
            - id: 5.4
              rule: Hugging Face model card YAML keys + required prose sections complete
              applies_to: [hf_card]
              applies_to_phase: [post-acceptance, post-publication]
              priority: H
              expected_lift: medium
            - id: 5.5
              rule: Hugging Face dataset card includes PHI/re-identification disclosure
              applies_to: [dataset_card]
              applies_to_phase: [post-acceptance, post-publication]
              priority: H
              expected_lift: medium
            - id: 5.6
              rule: Web-crawler-friendly markdown (declarative headings, alt text, fenced code, JSON-LD)
              applies_to: [readme, hf_card]
              applies_to_phase: [post-acceptance, post-publication]
              priority: M
              expected_lift: medium
        
            # Section 6 — Authority / E-E-A-T
            - id: 6.1
              rule: Personal author landing page lists all papers with DOIs
              applies_to: [author_profile]
              applies_to_phase: [pre-draft, drafting, pre-submission, post-acceptance, post-publication]
              priority: M
              expected_lift: medium
            - id: 6.2
              rule: Affiliation string consistent across papers
              applies_to: [author_profile]
              applies_to_phase: [drafting, pre-submission]
              priority: H
              expected_lift: high
            - id: 6.3
              rule: ORCID complete and linked to Google Scholar
              applies_to: [author_profile]
              applies_to_phase: [pre-draft, drafting, pre-submission, post-acceptance, post-publication]
              priority: H
              expected_lift: high
        
            # Section 7 — LLM-Citation Fabrication Defense
            - id: 7.1
              rule: DOI + PMID surfaced in copy-friendly text on landing page and README
              applies_to: [readme, manuscript]
              applies_to_phase: [post-acceptance, post-publication]
              priority: H
              expected_lift: high
            - id: 7.2
              rule: "How to cite" section with BibTeX, APA, Vancouver, plain-text
              applies_to: [readme]
              applies_to_phase: [post-acceptance, post-publication]
              priority: H
              expected_lift: medium
            - id: 7.3
              rule: Scholar alert configured; Perplexity / ChatGPT probes scheduled
              applies_to: [post_acceptance]
              applies_to_phase: [post-publication]
              priority: M
              expected_lift: low
        
            # Section 10 — Q&A and Entity-Extraction
            - id: 10.1
              rule: Four-question Q&A block (What was known / What this adds / How this changes practice / Why it matters)
              applies_to: [discussion, supplement]
              applies_to_phase: [drafting, pre-submission]
              priority: H
              expected_lift: high   # KJR case: directly correlated with AI-overview extraction
            - id: 10.2
              rule: Glossary block with MeSH / RadLex / UMLS / arXiv IDs
              applies_to: [methods, supplement]
              applies_to_phase: [pre-submission]
              priority: L
              expected_lift: low   # over-engineering for non-NLP / non-foundation-model papers; apply selectively
            - id: 10.3
              rule: Inline citation anchor text uses semantic predicates, not bare numbers
              applies_to: [manuscript]
              applies_to_phase: [drafting, pre-submission]
              priority: M
              expected_lift: medium
            - id: 10.4
              rule: Explicit "Why this is hard" challenge statement at start of Discussion
              applies_to: [discussion]
              applies_to_phase: [drafting, pre-submission]
              priority: H
              expected_lift: high   # KJR case: top quoted fragment by Perplexity / ChatGPT web
        
            # Section 11 — First-Mover Timing and Citation-Graph
            - id: 11.1
              rule: Topic-peak signals checked before submission timing decision
              applies_to: [submission_strategy]
              applies_to_phase: [pre-draft]
              priority: H
              expected_lift: high
            - id: 11.2
              rule: Editorial-board leverage considered (corresponding author affiliation)
              applies_to: [submission_strategy]
              applies_to_phase: [pre-draft]
              priority: M
              expected_lift: medium
            - id: 11.3
              rule: Target journal verified as PMC-auto-deposit when applicable
              applies_to: [submission_strategy]
              applies_to_phase: [pre-draft, pre-submission]
              priority: H
              expected_lift: high
              defers_to: references/oac_funding_checklist.yaml
            - id: 11.4
              rule: Discussion anchored in 5–10 high-visibility seminal works
              applies_to: [discussion]
              applies_to_phase: [drafting]
              priority: M
              expected_lift: medium
            - id: 11.5
              rule: Author roster spans ≥ 3 institutions and ≥ 2 disciplines (review papers)
              applies_to: [authorship]
              applies_to_phase: [pre-draft]
              priority: M
              expected_lift: medium   # not actionable once drafting begins; auto-NA in later phases
        
            # Section 12 — Cross-Platform Launch Sequencing
            - id: 12.1
              rule: Day-0 simultaneous launch (GitHub release, HF card, X/Threads, LinkedIn, landing page)
              applies_to: [post_acceptance]
              applies_to_phase: [post-acceptance]
              priority: H
              expected_lift: high
            - id: 12.2
              rule: Day-1 propagation (ORCID, Scholar, preprint update, institutional page)
              applies_to: [post_acceptance]
              applies_to_phase: [post-acceptance]
              priority: M
              expected_lift: medium
            - id: 12.3
              rule: Week-1 depth posts (LinkedIn long-form, Papers with Code, ResearchGate)
              applies_to: [post_acceptance]
              applies_to_phase: [post-acceptance]
              priority: M
              expected_lift: medium
            - id: 12.4
              rule: Weeks 2–4 refresh signals scheduled
              applies_to: [post_acceptance]
              applies_to_phase: [post-publication]
              priority: L
              expected_lift: low
            - id: 12.5
              rule: Month-1 monitoring (alerts + AI-search probes)
              applies_to: [post_acceptance]
              applies_to_phase: [post-publication]
              priority: M
              expected_lift: medium
        ```
        
        ## Output template (paste into qc/aio_audit.md)
        
        ```markdown
        # AIO Audit — {artifact_path}
        
        **Phase**: {pre-draft | drafting | pre-submission | post-acceptance | post-publication}
        **Top 5 ranked by expected_lift (high → medium → low):**
        
        | Rank | ID | Rule | Status | Reason | Suggested edit |
        |------|----|------|--------|--------|----------------|
        | 1 | {id} | … | FAIL | … | … |
        | ... |
        
        ## Full table (applicable items only — out-of-phase rules auto-NA)
        
        | ID | Rule | Lift | Status | Reason | Suggested edit |
        |----|------|------|--------|--------|----------------|
        | ... |
        
        ## Deferred items (audit elsewhere)
        
        - §1.6 → run `/check-reporting`
        - §3.4 / §11.3 → see `references/oac_funding_checklist.yaml`
        
        ## Summary
        
        - PASS: X / Y applicable items
        - High-lift fixes pending: …
        ```
        
        ## Migration from schema v1
        
        If a prior `qc/aio_audit.md` was generated under schema v1 (no `applies_to_phase` / `expected_lift`), regenerate from this v2 file. The id list is unchanged so prior PASS/FAIL state can be re-imported by id.
        
    • schema_markup_templates
      • CodeRepository.jsonld 990 B · in bundle
      • Dataset.jsonld 1.2 KB · in bundle
      • Person.jsonld 885 B · in bundle
      • README.md 2 KB
        # Schema.org JSON-LD Templates
        
        Embed-ready JSON-LD markup for academic-aio Section 5 (GitHub / CITATION.cff / Zenodo / Hugging Face) and Section 6 (author landing page). Schema.org markup is read by Google Scholar's structured-data parser, by AI overview engines, and by RAG systems that crawl repository pages.
        
        ## Files
        
        - **`ScholarlyArticle.jsonld`** — embed in the paper landing page or repository README. Pairs each paper with DOI, PMID, abstract, license, OA flag, and citation graph anchors.
        - **`CodeRepository.jsonld`** — embed in the repository README or as `.well-known/scholarly.jsonld`. Links code to its parent article via `isPartOf`.
        - **`Dataset.jsonld`** — embed at the Zenodo or Hugging Face dataset landing page. Includes provenance, license, and variable-level metadata.
        - **`Person.jsonld`** — embed at the author landing page or institutional profile. Cross-links ORCID, Scholar, Semantic Scholar, GitHub, Hugging Face, LinkedIn.
        
        ## How to embed
        
        ### In a Markdown README (GitHub renders the HTML script tag inside HTML blocks)
        
        ```html
        <script type="application/ld+json">
        { ... contents of ScholarlyArticle.jsonld ... }
        </script>
        ```
        
        ### In a personal landing page
        
        Place the script tag in the `<head>` of the HTML page. For static-site generators (Hugo, Jekyll, Next.js), generate the JSON-LD at build time from frontmatter.
        
        ### Validation
        
        Run `scripts/validate_schema.py path/to/file.jsonld` to verify syntactic validity and required-field presence before deploy.
        
        ```bash
        python scripts/validate_schema.py references/schema_markup_templates/*.jsonld
        ```
        
        ## Anti-hallucination
        
        - Do not auto-fill placeholder values (`<First Last>`, `0000-0000-0000-0000`, `10.xxxx/yyyy`). Mark unknown fields as `null` or remove the field — partial fabricated identifiers are worse than missing ones.
        - Verify DOI format `10.{prefix}/{suffix}` and ORCID format before commit.
        
        ## Related
        
        - `SKILL.md` Section 5 — README, CITATION.cff, Zenodo, Hugging Face
        - `SKILL.md` Section 6 — Authority and E-E-A-T signals
        - `SKILL.md` Section 7 — Citation-fabrication defense
        
      • ScholarlyArticle.jsonld 1.6 KB · in bundle
    • ai_tool_citation_framing.md 4 KB
      # Citing an AI-assisted research tool safely (framing by use-class)
      
      When a manuscript used an AI-assisted tool — a reference-verification / QA suite, a
      statistical-analysis helper, or a generative drafting assistant — *where and how you
      name it* changes how an editor or reviewer reads it. Under current journal wariness
      about AI, a proud in-text citation of a **generative** use can invite suspicion or a
      desk-reject, while the identical placement for a **verification** use reads as rigor
      (the same way citing R, SPSS, or a reference manager does). The safe move is to frame
      the mention by what the tool actually *did*, not to hide it.
      
      This applies to any AI-assisted tool; it also applies to **self-citation** by a tool's
      author (e.g. citing MedSci Skills in your own paper), which additionally requires a
      conflict-of-interest disclosure.
      
      ## The three use-classes
      
      | Use-class | What it covers | Where it belongs | Citable like software? |
      |---|---|---|---|
      | **Verification / QA** (rigor-signalling) | Reference/citation verification (DOI/PMID against PubMed/CrossRef), reporting-guideline compliance checks, deterministic integrity gates, numerical-consistency checks | **Software / Code-availability statement** (or Methods, as a named tool with version) | **Yes** — cite it plainly, like a reference manager or a linter. It signals rigor. |
      | **Analysis** (neutral) | Statistical-analysis code, figure generation, data-transformation scripts | Methods (named, with version) and/or Code-availability | **Yes** — neutral and citable, like citing R/Python packages. |
      | **Generative** (disclosure, not citation) | Drafting or rewriting prose, "humanizing", summarizing, idea generation | The journal's **AI-use disclosure field / statement** (per its policy), *not* a proud in-text citation | **No** — declare it in the disclosure field; do not farm it into the running text. |
      
      ## Rules of thumb
      
      - **Prefer the deterministic, verifiable functions for the cleanest cite.** A tool's
        reproducible/verifiable capabilities (reference verification, compliance gates,
        analysis code) are the ones that read as rigor and are safest to name in a Software
        statement. Its generative capabilities belong in the disclosure field.
      - **Match the placement to the use-class**, not to the tool. The *same* tool can be
        citable (its QA gate ran your references) *and* disclosure-only (its drafting helper
        touched your prose) in the *same* paper — split the mention accordingly.
      - **Do not self-cite generatively.** If the paper did not materially use the generative
        parts, do not add a self-citation for them to inflate a citation count; cite only the
        functions the work actually used.
      - **Pair a self-citation with a COI disclosure.** If you authored the tool, state the
        relationship (e.g. "Author X develops the cited toolkit") in the COI/competing-interests
        section — the same standard as any other intellectual/financial interest.
      - **Honor the target journal's AI policy first.** Where a journal specifies *where* AI
        use must be declared (Methods vs Acknowledgements vs a dedicated field), that placement
        wins for the generative use-class; the Software-statement guidance above is for the
        verification/analysis classes, which are tool-use, not AI-authorship.
      
      ## Why this is guidance, not a deterministic gate (yet)
      
      A deterministic check ("flag an AI-tool citation placed in running-text Methods for a
      *generative* use") would need both a maintained tool-name allowlist and a reliable
      classifier of *which use-class* a given sentence describes — the latter is high false-
      positive without context the grep cannot see. The reliable, low-FP part (a tool named in
      a Software/Code-availability statement) is already the recommended state, so there is
      nothing to flag. If a bounded, allowlist-driven placement check proves worthwhile on real
      manuscripts, it can be added later; until then this stays advisory. (See
      `~/.claude/rules/manuscript-style-classical.md` §7/§15 for AI-disclosure placement and the
      self-applicability rule.)
      
    • journal_summarybox_templates.yaml 4.7 KB
      ---
      schema_version: 1
      last_updated: "2026-05-10"
      purpose: |
        Templates for journal-specific summary boxes used in medical-AI manuscripts.
        Each entry mirrors the journal's Instructions for Authors (IFA) at the date
        shown. USER MUST VERIFY the current IFA before applying — journal policies
        change without notice.
      
      verification_protocol: |
        Before applying any template:
        1. Open the journal's current IFA URL (ifa_url field).
        2. Compare structure / labels / word targets to this YAML.
        3. If any field has changed since `last_verified`, update this file before
           deriving the manuscript's box.
        4. Treat `last_verified` older than 6 months as STALE — re-verify before
           submission.
      
      journals:
        - id: lancet-digital-health
          name: The Lancet Digital Health
          box_label: Research in context
          ifa_url: https://www.thelancet.com/journals/landig/about
          last_verified: "2026-05-10"
          placement: After Introduction in main text, or as Panel 1.
          structure:
            - id: evidence-before
              label: Evidence before this study
              word_target: 100-200
              notes: |
                Describe the systematic search method (databases, dates, terms) and
                what was known prior to this study.
            - id: added-value
              label: Added value of this study
              word_target: 100-200
              notes: |
                State the contribution and how it advances the field beyond the
                evidence summarized above.
            - id: implications
              label: Implications of all the available evidence
              word_target: 100-200
              notes: |
                How findings (combined with prior evidence) should change clinical
                practice, policy, or future research.
          user_must_verify: true
      
        - id: rsna-radiology
          name: Radiology (RSNA)
          box_label: Key Results
          ifa_url: https://pubs.rsna.org/page/radiology/author-instructions
          last_verified: "2026-05-10"
          placement: First page of article, after Abstract.
          structure:
            - id: key-results
              label: Key Results
              format: 3 bullets, claim-centric, one finding each
              notes: |
                Each bullet should state a primary outcome with point estimate and
                confidence interval where applicable. Avoid hedging language.
          user_must_verify: true
      
        - id: rsna-radiology-ai
          name: Radiology Artificial Intelligence (RSNA)
          box_label: Key Points
          ifa_url: https://pubs.rsna.org/page/ai/author-instructions
          last_verified: "2026-05-10"
          placement: After Abstract.
          structure:
            - id: key-points
              label: Key Points
              format: 3 bullets
              notes: |
                Same conventions as Radiology Key Results — claim-centric, numeric
                where possible.
          user_must_verify: true
      
        - id: npj-digital-medicine
          name: npj Digital Medicine
          box_label: Plain Language Summary
          ifa_url: https://www.nature.com/npjdigitalmed/submission-guidelines
          last_verified: "2026-05-10"
          placement: After Abstract; required for all primary research.
          structure:
            - id: pls
              label: Plain Language Summary
              word_target: 150-200
              reading_level: 8th-grade
              notes: |
                Avoid jargon, define acronyms inline, prefer everyday vocabulary.
                Should be readable by a non-specialist physician or informed lay
                reader. No quantitative thresholds; describe magnitude in plain terms.
          user_must_verify: true
      
        - id: nature-medicine
          name: Nature Medicine
          box_label: Editor's Summary
          ifa_url: https://www.nature.com/nm/for-authors
          last_verified: "2026-05-10"
          placement: Editorial-supplied at acceptance; not author-drafted in main text.
          structure:
            - id: editor-summary
              label: Editor's Summary
              word_target: 100-150
              notes: |
                Supplied by the editorial team after acceptance. Authors should
                proactively draft a 100–150 word version for the cover letter and
                post-acceptance launch materials, so the editor has anchor wording
                to refine.
          user_must_verify: true
      
        - id: jamia
          name: JAMIA (Journal of the American Medical Informatics Association)
          box_label: (none mandated; structured abstract used)
          ifa_url: https://academic.oup.com/jamia/pages/General_Instructions
          last_verified: "2026-05-10"
          placement: Structured abstract serves the role of summary box.
          structure:
            - id: structured-abstract
              label: Structured Abstract
              format: Objective / Materials and Methods / Results / Discussion / Conclusion
              notes: Treat each section as a chunk-friendly semantic block (≤ 3 sentences).
          user_must_verify: true
      
      related:
        - skills/academic-aio/SKILL.md Section 2.1 (manuscript-level summary box)
        - skills/academic-aio/SKILL.md Section 1.6 (reporting-guideline anchor)
      
    • oac_funding_checklist.yaml 4.9 KB
      ---
      schema_version: 1
      last_updated: "2026-05-10"
      purpose: |
        Open-access compliance matrix for major funders that medical-AI researchers
        may encounter. Use to decide gold OA vs green OA vs preprint route per paper.
        Funder policies change; verify the policy_url before applying.
      
      verification_protocol: |
        All funder policies change without notice.
        - Verify policy_url before applying decisions.
        - Treat `last_verified` older than 6 months as STALE.
        - When in doubt, contact the institution's research-administration office.
      
      funders:
        - id: plan-s
          name: cOAlition S (Plan S)
          region: Europe (multiple national funders)
          policy_url: https://www.coalition-s.org/plan-s-principles/
          last_verified: "2026-05-10"
          requirement: Immediate OA on publication; no embargo allowed.
          accepted_routes:
            - Gold OA with CC-BY (preferred)
            - Green OA via repository deposit at publication (no embargo)
          apc_cap: ~ EUR 2,500 typical, varies by funder
          notes: Strict; transformative agreements widely available.
          user_must_verify: true
      
        - id: nih-public-access
          name: US NIH Public Access Policy
          region: United States
          policy_url: https://publicaccess.nih.gov/
          last_verified: "2026-05-10"
          requirement: |
            From 2025 onward NIH policy requires manuscript deposit in PMC at
            publication (no 12-month embargo). Earlier grants may still operate
            under the legacy 12-month rule.
          accepted_routes:
            - Gold OA (publisher deposits to PMC)
            - Green OA (author deposits AAM to PMC via NIHMS)
          apc_cap: No cap; APCs are eligible expenses on most grants.
          user_must_verify: true
      
        - id: ukri
          name: UK Research and Innovation (UKRI)
          region: United Kingdom
          policy_url: https://www.ukri.org/manage-your-award/publishing-your-research-findings/making-your-research-open/
          last_verified: "2026-05-10"
          requirement: Immediate OA; CC-BY preferred.
          accepted_routes:
            - Gold OA via transformative agreement (preferred)
            - Green OA in approved repository at publication
          user_must_verify: true
      
        - id: gates-foundation
          name: Bill & Melinda Gates Foundation
          region: Global
          policy_url: https://openaccess.gatesfoundation.org/
          last_verified: "2026-05-10"
          requirement: Immediate OA, CC-BY, no embargo.
          accepted_routes:
            - Gold OA (mandatory)
          notes: Strictest among major funders. Closed-access publication is non-compliant.
          user_must_verify: true
      
        - id: wellcome
          name: Wellcome Trust
          region: UK / global
          policy_url: https://wellcome.org/grant-funding/guidance/open-access-policy
          last_verified: "2026-05-10"
          requirement: Immediate OA, CC-BY, including for preprints.
          accepted_routes:
            - Gold OA (preferred)
            - Green OA at publication
          user_must_verify: true
      
        - id: nrf-korea
          name: National Research Foundation of Korea (NRF)
          region: South Korea
          policy_url: https://www.nrf.re.kr/
          last_verified: "2026-05-10"
          requirement: |
            No formal mandatory immediate-OA policy as of 2026-Q2. Gold OA is
            encouraged and APCs are eligible expenses on most grants.
          accepted_routes:
            - Gold OA (encouraged)
            - Green OA (acceptable)
          user_must_verify: true
      
        - id: hira-mohw
          name: Korean Ministry of Health & Welfare research grants
          region: South Korea
          policy_url: https://www.mohw.go.kr/
          last_verified: "2026-05-10"
          requirement: Public benefit / repository deposit encouraged; no strict immediate-OA mandate.
          accepted_routes:
            - Gold OA (encouraged)
            - Green OA (acceptable)
          user_must_verify: true
      
      decision_tree:
        - step: 1
          question: Does any funder on this paper require immediate OA?
          yes_action: |
            Choose gold OA at a Plan-S compliant journal, OR green OA with
            repository deposit at publication. Closed-access is not allowed.
          no_action: Continue to step 2.
        - step: 2
          question: Is APC budget available (grant line, institutional waiver, or transformative agreement)?
          yes_action: |
            Gold OA preferred — lower friction, faster PMC deposit, cited
            30-50 % more by AI-search tools (Section 3.4).
          no_action: |
            Green OA via preprint at submission + AAM deposit at acceptance.
            Verify journal's green-OA timeline.
        - step: 3
          question: Does the target journal auto-deposit to PMC (Section 11.3)?
          yes_action: Confirm and proceed. Deposit window typically 4-6 weeks.
          no_action: |
            Plan green OA pathway and timeline. Budget the AAM submission step
            (NIHMS or Europe PMC).
        - step: 4
          question: Is a transformative agreement available via the institution?
          yes_action: Gold OA is effectively free for the author — use it.
          no_action: Apply APC budget directly or fall back to green OA.
      
      related:
        - skills/academic-aio/SKILL.md Section 3.4 (open-access choice)
        - skills/academic-aio/SKILL.md Section 11.3 (PMC-auto-deposit journal preference)
      
    • reporting_guideline_mapping.md 3.5 KB
      # Reporting-Guideline ↔ AIO Rule Mapping
      
      This table maps each AIO rule (sections 1-12 of `SKILL.md`) to the corresponding item(s) in the major medical-AI reporting guidelines. Use it to align the `/check-reporting` audit with the `/academic-aio` audit so the same evidence covers both.
      
      ## Core mapping
      
      | AIO rule | TRIPOD+AI 2024 | CLAIM 2024 | STARD-AI 2025 | TRIPOD-LLM 2024 | DECIDE-AI 2022 | Notes |
      |----------|----------------|-------------|----------------|-----------------|----------------|-------|
      | §1.1 Title three-slot | item 1 (title) | item 1 (title and abstract) | item 1 (title) | item 1a (title) | — | All require keyword presence and study-type identification. |
      | §1.2 Structured abstract | item 2 (abstract) | item 1 (title and abstract) | item 2 (abstract) | item 1b (abstract) | — | Each guideline mandates structured form. |
      | §1.5 Quantified primary outcome with CI | item 16 (model performance) | items 28-30 (performance metrics) | items 23-26 (diagnostic estimates) | item 17 (performance with CI) | item 8 (clinical effect) | CI mandatory for all. |
      | §1.6 Reporting-guideline anchor | (compliance declaration) | (compliance declaration) | (compliance declaration) | (compliance declaration) | (compliance declaration) | Cite guideline + checklist in Methods or supplement. |
      | §2.4 Reproducibility block | item 24 (data sharing) + item 25 (code sharing) | items 33-34 (data and code availability) | item 28 (data and code) | items 22-23 (artifacts) | item 12 (artifacts) | All require explicit data/code statement. |
      | §2.5 Limitations enumeration | item 23 (limitations) | item 41 (limitations) | item 27 (limitations) | item 19 (limitations) | item 11 (limitations) | Enumerate, do not narrate generally. |
      | §10.4 Challenge statement | (implicit in Background) | item 4 (rationale) | item 5 (rationale) | item 4 (rationale) | item 4 (rationale) | "Why this is hard" overlaps with rationale items. |
      
      ## Workflow
      
      1. Run `/check-reporting` first — produces a PRESENT/PARTIAL/MISSING audit per guideline item.
      2. Run `/academic-aio` — produces the AIO PASS/PARTIAL/FAIL checklist.
      3. For each AIO FAIL or PARTIAL row, check this mapping. If the underlying reporting-guideline item is also MISSING/PARTIAL, fix it once and both audits update.
      4. Items present in `/check-reporting` audit but not in this mapping (e.g., randomization details for RCTs, domain-specific safety items) do not have an AIO consequence and can be addressed independently.
      
      ## When the two audits disagree
      
      - AIO PASS + reporting-guideline MISSING — the manuscript looks discoverable but is not formally compliant. Reviewers may still reject. Always fix the reporting-guideline gap.
      - AIO FAIL + reporting-guideline PRESENT — a rule was recorded in compliance form but rendered in a way that LLM extractors cannot parse (e.g., reporting CIs in supplementary instead of inline in the abstract). Move the content into a chunk-friendly location.
      
      ## Source
      
      - TRIPOD+AI: Collins et al. BMJ 2024.
      - CLAIM 2024: Tejani et al. Radiology: AI 2024, doi:10.1148/ryai.240300.
      - STARD-AI 2025: Sounderajah et al. Nat Med 2025, doi:10.1038/s41591-025-03953-8.
      - TRIPOD-LLM 2024: Gallifant et al. Nat Med 2024, doi:10.1038/s41591-024-03425-5.
      - DECIDE-AI 2022: Vasey et al. Nat Med 2022, doi:10.1038/s41591-022-01772-9.
      
      ## Anti-hallucination
      
      Item numbers above are derived from the most recent published version of each guideline. Verify against the EQUATOR Network entry before citing item numbers in a manuscript — guideline updates renumber items.
      
    • summary_box_specs.json 1.2 KB
      {
        "_comment": "Synthesis of PUBLIC, journal-documented structured-summary-box facts (bullet counts, sub-block labels, word bands) — NOT verbatim journal text. Mirrors the facts academic-aio/SKILL.md already states (Lancet 'Research in context' 3 sub-blocks; Radiology/RYAI 'Key Points' 3 one-claim bullets; npj 'Plain-language summary' 150-200 words). Verify current instructions-to-authors at the journal site before relying on a format; these are conventions, not guarantees.",
        "schema_version": 1,
        "formats": {
          "key_points": {
            "label": "Key Points",
            "bullets": 3,
            "one_claim_per_bullet": true,
            "journals": ["radiology", "radiology-ai", "ryai", "rsna"]
          },
          "research_in_context": {
            "label": "Research in context",
            "subblocks": [
              "Evidence before this study",
              "Added value of this study",
              "Implications of all the available evidence"
            ],
            "journals": ["lancet-digital-health", "lancet", "lancet-oncology"]
          },
          "plain_language_summary": {
            "label": "Plain-language summary",
            "word_min": 150,
            "word_max": 200,
            "journals": ["npj-digital-medicine", "npj"]
          }
        }
      }
      
  • scripts
    • batch_metadata_audit.py 5.9 KB
      #!/usr/bin/env python3
      """
      batch_metadata_audit.py — Audit multiple medical-AI repos and Hugging Face cards
      for AIO compliance.
      
      Per repository:
      - README.md present, with DOI link / badge and a How-to-cite or Citation section
      - CITATION.cff present at root with title/authors/version + at least one ORCID
      - LICENSE present
      
      Per Hugging Face card (model or dataset):
      - YAML front matter present with license / library_name / tags
      - Required prose sections: intended use, training data, evaluation, limitations,
        ethical considerations
      - No PHI patterns matched
      
      Usage:
          python batch_metadata_audit.py /path/to/repo1 /path/to/repo2 \\
              --hf-card model_card.md --output qc/aio_batch.json
          python batch_metadata_audit.py --hf-card dataset_card.md --fail-on-issue
      """
      
      from __future__ import annotations
      
      import argparse
      import json
      import re
      import sys
      from pathlib import Path
      
      PHI_PATTERNS = [
          re.compile(r"\b\d{3}-\d{2}-\d{4}\b"),                 # US SSN
          re.compile(r"\b\d{6}-\d{7}\b"),                       # KR resident registration number
          re.compile(r"MRN[:\s]*\d+", re.I),                    # medical record number
          re.compile(r"\bpatient[\s_]?id[:\s]*\d+", re.I),      # patient id
          re.compile(r"\b\d{4}-\d{2}-\d{2}\b.*\bDOB\b", re.I),  # DOB pattern
      ]
      
      CITATION_CFF_REQUIRED = ("title", "authors", "version")
      HF_CARD_REQUIRED_YAML = ("license", "library_name", "tags")
      HF_CARD_REQUIRED_SECTIONS = (
          "intended use",
          "training data",
          "evaluation",
          "limitations",
          "ethical considerations",
      )
      
      
      def _phi_hits(text: str) -> list[str]:
          hits: list[str] = []
          for pat in PHI_PATTERNS:
              m = pat.search(text)
              if m:
                  hits.append(f"Possible PHI pattern matched: {m.group(0)[:30]!r}")
          return hits
      
      
      def check_readme(path: Path) -> dict:
          if not path.exists():
              return {"present": False, "issues": ["README.md missing"]}
          text = path.read_text()
          issues: list[str] = []
          if "DOI" not in text and "doi.org" not in text.lower():
              issues.append("No DOI badge or link in README")
          if "## How to cite" not in text and "## Citation" not in text:
              issues.append('No "How to cite" / "Citation" section')
          if not any(s in text.lower() for s in ("quickstart", "getting started", "installation")):
              issues.append("No quickstart / installation section")
          return {"present": True, "issues": issues}
      
      
      def check_citation_cff(path: Path) -> dict:
          if not path.exists():
              return {"present": False, "issues": ["CITATION.cff missing at repo root"]}
          text = path.read_text()
          issues: list[str] = []
          for key in CITATION_CFF_REQUIRED:
              if f"{key}:" not in text:
                  issues.append(f"CITATION.cff missing key: {key}")
          if "orcid" not in text.lower():
              issues.append("CITATION.cff has no ORCID identifier(s)")
          return {"present": True, "issues": issues}
      
      
      def check_license(path: Path) -> dict:
          return {
              "present": path.exists(),
              "issues": [] if path.exists() else ["LICENSE missing"],
          }
      
      
      def check_hf_card(path: Path) -> dict:
          if not path.exists():
              return {"present": False, "issues": ["HF card missing"]}
          text = path.read_text()
          issues: list[str] = []
          if not text.startswith("---"):
              issues.append("HF card missing YAML front matter")
          else:
              front_parts = text.split("---", 2)
              front = front_parts[1] if len(front_parts) >= 3 else ""
              for key in HF_CARD_REQUIRED_YAML:
                  if f"{key}:" not in front:
                      issues.append(f"HF card YAML missing key: {key}")
          body_lower = text.lower()
          for section in HF_CARD_REQUIRED_SECTIONS:
              if section not in body_lower:
                  issues.append(f"HF card missing section: {section}")
          issues.extend(_phi_hits(text))
          return {"present": True, "issues": issues}
      
      
      def audit_repo(repo: Path) -> dict:
          return {
              "repo": str(repo),
              "readme": check_readme(repo / "README.md"),
              "citation_cff": check_citation_cff(repo / "CITATION.cff"),
              "license": check_license(repo / "LICENSE"),
          }
      
      
      def main(argv: list[str] | None = None) -> int:
          parser = argparse.ArgumentParser(
              description="Audit multiple repos and HF cards for AIO compliance."
          )
          parser.add_argument("paths", nargs="*", type=Path,
                              help="Repository directories to audit.")
          parser.add_argument("--hf-card", action="append", type=Path, default=[],
                              help="Path to a Hugging Face model/dataset card markdown file.")
          parser.add_argument("--output", type=Path, default=None,
                              help="Write JSON report to this path (default: stdout).")
          parser.add_argument("--fail-on-issue", action="store_true",
                              help="Exit 1 if any issues are detected.")
          args = parser.parse_args(argv)
      
          if not args.paths and not args.hf_card:
              parser.error("Provide at least one repo path or --hf-card.")
      
          report: dict = {"repos": [], "hf_cards": []}
          for path in args.paths:
              if path.is_dir():
                  report["repos"].append(audit_repo(path))
              else:
                  report["repos"].append({"repo": str(path),
                                          "issues": ["Not a directory"]})
          for card in args.hf_card:
              report["hf_cards"].append({"path": str(card), **check_hf_card(card)})
      
          out = json.dumps(report, indent=2, ensure_ascii=False)
          if args.output:
              args.output.parent.mkdir(parents=True, exist_ok=True)
              args.output.write_text(out)
          else:
              print(out)
      
          has_issues = any(
              (r.get("readme", {}).get("issues") or
               r.get("citation_cff", {}).get("issues") or
               r.get("license", {}).get("issues") or
               r.get("issues"))
              for r in report["repos"]
          ) or any(c.get("issues") for c in report["hf_cards"])
      
          return 1 if (args.fail_on_issue and has_issues) else 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • check_summary_box.py 7.8 KB
      #!/usr/bin/env python3
      """Structured-summary-box conformance detector (academic-aio).
      
      High-impact medical-AI journals require a structured summary box whose *format*
      is journal-specific, and a production/technical check rejects the wrong one:
      
        - Radiology / Radiology:AI (RSNA): "Key Points" — exactly 3 bullets, one claim each.
        - Lancet family: "Research in context" — three labelled sub-blocks
          (Evidence before this study / Added value of this study / Implications of all
          the available evidence).
        - npj Digital Medicine: "Plain-language summary" — ~150-200 words.
      
      academic-aio already *generates* these boxes; this detector makes the spec
      deterministic so a wrong-bullet-count, missing-sub-block, or over/under-length
      box is caught before submission instead of at the technical check. The spec is
      read from references/summary_box_specs.json (public facts, journal-keyed).
      
      INPUTS
        --manuscript   markdown file containing the summary box (required).
        --journal      journal stem to pick the format (e.g. radiology, lancet-digital-health,
                       npj-digital-medicine). Optional if --format is given.
        --format       force a format: key_points | research_in_context | plain_language_summary.
        --specs        path to summary_box_specs.json (default: alongside this script's skill).
        --out          write a JSON report here (default: qc/summary_box_report.json).
        --strict       exit 1 if the box is non-conformant.
      
      VERDICT
        CONFORMANT          the box matches its format's spec.
        NONCONFORMANT       a hard rule failed (wrong bullet count, missing sub-block,
                            word count outside the band, box absent).
        ADVISORY            only soft rules fired (e.g. a bullet carries >1 claim).
        Exit: 0 conformant/advisory or report-only; 1 NONCONFORMANT under --strict;
              2 input/usage error.
      
      Stdlib-only (csv-free: json / argparse / re / pathlib).
      """
      
      from __future__ import annotations
      
      import argparse
      import json
      import re
      import sys
      from pathlib import Path
      
      
      def _default_specs() -> Path:
          return Path(__file__).resolve().parent.parent / "references" / "summary_box_specs.json"
      
      
      def _err(msg: str) -> int:
          print(f"ERROR: {msg}", file=sys.stderr)
          return 2
      
      
      def load_specs(path: Path) -> dict:
          data = json.loads(path.read_text(encoding="utf-8"))
          return data["formats"]
      
      
      def pick_format(formats: dict, journal: str | None, fmt: str | None) -> str | None:
          if fmt:
              return fmt if fmt in formats else None
          if journal:
              j = journal.strip().lower()
              for key, spec in formats.items():
                  if j in [x.lower() for x in spec.get("journals", [])]:
                      return key
          return None
      
      
      def extract_block(text: str, label: str) -> str | None:
          """Return the lines under a heading/bold label matching `label`, up to the
          next markdown heading or a blank-line-separated next bold label section."""
          lines = text.splitlines()
          label_re = re.compile(
              r"^\s*(?:#{1,6}\s*|\*\*\s*|\*\s*)?" + re.escape(label) + r"\b", re.IGNORECASE
          )
          start = None
          for i, ln in enumerate(lines):
              if label_re.search(ln):
                  start = i
                  break
          if start is None:
              return None
          out: list[str] = []
          for ln in lines[start + 1:]:
              if re.match(r"^\s*#{1,6}\s+\S", ln):  # next heading ends the block
                  break
              out.append(ln)
          return "\n".join(out).strip()
      
      
      def count_bullets(block: str) -> list[str]:
          bullets: list[str] = []
          for ln in block.splitlines():
              m = re.match(r"^\s*(?:[-*+]|\d+[.)])\s+(.*\S)", ln)
              if m:
                  bullets.append(m.group(1).strip())
          return bullets
      
      
      def multi_claim(bullet: str) -> bool:
          """Heuristic: a one-claim bullet should not pack two independent assertions.
          Flags a sentence-final period followed by a capitalized new sentence, or a
          semicolon joining two clauses."""
          if ";" in bullet:
              return True
          return bool(re.search(r"[.!?]\s+[A-Z0-9]", bullet.rstrip(".")))
      
      
      def word_count(block: str) -> int:
          # strip the label line if it leaked in; count remaining words.
          return len(re.findall(r"\b[\w'-]+\b", block))
      
      
      def check(text: str, fmt: str, spec: dict) -> dict:
          label = spec["label"]
          block = extract_block(text, label)
          findings: list[dict] = []
          if block is None:
              return {
                  "format": fmt, "label": label, "verdict": "NONCONFORMANT",
                  "findings": [{"rule": "box_present", "severity": "hard",
                                "detail": f"no '{label}' box found in the manuscript"}],
              }
      
          if fmt == "key_points":
              bullets = count_bullets(block)
              want = spec["bullets"]
              if len(bullets) != want:
                  findings.append({"rule": "bullet_count", "severity": "hard",
                                   "detail": f"found {len(bullets)} bullets, expected {want}"})
              if spec.get("one_claim_per_bullet"):
                  for b in bullets:
                      if multi_claim(b):
                          findings.append({"rule": "one_claim_per_bullet", "severity": "soft",
                                           "detail": f"bullet packs >1 claim: {b[:80]}"})
          elif fmt == "research_in_context":
              low = block.lower()
              for sub in spec["subblocks"]:
                  if sub.lower() not in low:
                      findings.append({"rule": "subblock_present", "severity": "hard",
                                       "detail": f"missing sub-block: '{sub}'"})
          elif fmt == "plain_language_summary":
              wc = word_count(block)
              lo, hi = spec["word_min"], spec["word_max"]
              if wc < lo or wc > hi:
                  findings.append({"rule": "word_band", "severity": "hard",
                                   "detail": f"{wc} words, expected {lo}-{hi}"})
          else:
              return {"format": fmt, "label": label, "verdict": "NONCONFORMANT",
                      "findings": [{"rule": "unknown_format", "severity": "hard",
                                    "detail": f"unknown format '{fmt}'"}]}
      
          hard = any(f["severity"] == "hard" for f in findings)
          verdict = "NONCONFORMANT" if hard else ("ADVISORY" if findings else "CONFORMANT")
          return {"format": fmt, "label": label, "verdict": verdict, "findings": findings}
      
      
      def main() -> int:
          ap = argparse.ArgumentParser(description="Check a structured summary box against its journal format spec.")
          ap.add_argument("--manuscript", required=True)
          ap.add_argument("--journal")
          ap.add_argument("--format")
          ap.add_argument("--specs")
          ap.add_argument("--out")
          ap.add_argument("--strict", action="store_true")
          args = ap.parse_args()
      
          man = Path(args.manuscript)
          if not man.is_file():
              return _err(f"manuscript not found: {man}")
          specs_path = Path(args.specs) if args.specs else _default_specs()
          if not specs_path.is_file():
              return _err(f"specs not found: {specs_path}")
      
          formats = load_specs(specs_path)
          fmt = pick_format(formats, args.journal, args.format)
          if fmt is None:
              return _err("could not select a format — pass --format or a --journal listed in the specs")
      
          text = man.read_text(encoding="utf-8")
          report = check(text, fmt, formats[fmt])
      
          out_path = Path(args.out) if args.out else Path("qc") / "summary_box_report.json"
          out_path.parent.mkdir(parents=True, exist_ok=True)
          out_path.write_text(json.dumps({"detector": "check_summary_box", **report}, indent=2) + "\n", encoding="utf-8")
      
          print("=" * 41)
          print(" Summary-Box Conformance")
          print("=" * 41)
          print(f"format: {fmt}  ({report['label']})")
          print(f"verdict: {report['verdict']}")
          for f in report["findings"]:
              print(f"  [{f['severity']}] {f['rule']}: {f['detail']}")
          print(f"report: {out_path}")
      
          if report["verdict"] == "NONCONFORMANT" and args.strict:
              print("\nSUMMARY_BOX_NONCONFORMANT", file=sys.stderr)
              return 1
          return 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • validate_schema.py 3.8 KB
      #!/usr/bin/env python3
      """
      validate_schema.py — JSON-LD validator for academic-aio schema markup files.
      
      Validates:
      - JSON-LD syntactic validity
      - @context = "https://schema.org"
      - @type matches one of the supported types
      - Required fields present (per schema.org minimal recommendations + medsci-skills policy)
      - Identifier format (DOI, ORCID)
      
      Usage:
          python validate_schema.py path/to/file.jsonld [path/to/another.jsonld ...]
          python validate_schema.py --strict references/schema_markup_templates/*.jsonld
      """
      
      from __future__ import annotations
      
      import argparse
      import json
      import re
      import sys
      from pathlib import Path
      
      REQUIRED_BY_TYPE: dict[str, list[str]] = {
          "ScholarlyArticle": ["headline", "datePublished", "author", "identifier", "url"],
          "SoftwareSourceCode": ["name", "codeRepository", "license", "datePublished", "author"],
          "Dataset": ["name", "description", "license", "creator", "datePublished"],
          "Person": ["name", "identifier"],
      }
      
      DOI_RE = re.compile(r"^10\.\d{4,9}/[-._;()/:A-Za-z0-9]+$")
      ORCID_RE = re.compile(r"^https://orcid\.org/\d{4}-\d{4}-\d{4}-\d{3}[\dX]$")
      
      PLACEHOLDER_TOKENS = ("<", "xxxx", "yyyy", "0000-0000-0000-0000")
      
      
      def _is_placeholder(value: str) -> bool:
          """Return True for template placeholder strings that should skip strict checks."""
          if not isinstance(value, str):
              return False
          lowered = value.lower()
          return any(tok in lowered for tok in PLACEHOLDER_TOKENS)
      
      
      def validate(path: Path) -> list[str]:
          errors: list[str] = []
          try:
              data = json.loads(path.read_text())
          except FileNotFoundError:
              return [f"File not found: {path}"]
          except json.JSONDecodeError as exc:
              return [f"Invalid JSON: {exc}"]
      
          ctx = data.get("@context")
          if ctx != "https://schema.org":
              errors.append(f'@context must be "https://schema.org" (got {ctx!r})')
      
          typ = data.get("@type")
          if typ not in REQUIRED_BY_TYPE:
              errors.append(
                  f"@type {typ!r} not recognized "
                  f"(expected one of {sorted(REQUIRED_BY_TYPE)})"
              )
              return errors
      
          for field in REQUIRED_BY_TYPE[typ]:
              if field not in data or data[field] in (None, "", []):
                  errors.append(f"Missing required field: {field}")
      
          ident = data.get("identifier")
          if isinstance(ident, list):
              for entry in ident:
                  if isinstance(entry, dict) and entry.get("propertyID") == "DOI":
                      value = entry.get("value", "")
                      if value and not _is_placeholder(value) and not DOI_RE.match(value):
                          errors.append(f"DOI does not match canonical format: {value!r}")
      
          if typ == "Person" and isinstance(ident, str):
              if not _is_placeholder(ident) and not ORCID_RE.match(ident):
                  errors.append(f"Person identifier should be an ORCID URL (got {ident!r})")
      
          authors = data.get("author") or data.get("creator") or []
          if isinstance(authors, list):
              for i, a in enumerate(authors):
                  if isinstance(a, dict) and not a.get("name"):
                      errors.append(f"author[{i}] missing 'name'")
      
          return errors
      
      
      def main(argv: list[str] | None = None) -> int:
          parser = argparse.ArgumentParser(
              description="Validate academic-aio Schema.org JSON-LD markup files."
          )
          parser.add_argument("files", nargs="+", type=Path)
          parser.add_argument(
              "--strict",
              action="store_true",
              help="Exit 1 on any error (default behaviour; flag retained for clarity).",
          )
          args = parser.parse_args(argv)
      
          overall_ok = True
          for path in args.files:
              errors = validate(path)
              if errors:
                  overall_ok = False
                  print(f"FAIL  {path}")
                  for e in errors:
                      print(f"  - {e}")
              else:
                  print(f"PASS  {path}")
          return 0 if overall_ok else 1
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
  • templates
    • aio_audit_checklist.md.j2 4 KB · in bundle
  • tests
    • test_batch_metadata_audit.sh 3.1 KB
      #!/usr/bin/env bash
      # Regression test for academic-aio/scripts/batch_metadata_audit.py.
      # Builds synthetic repos / HF cards (no committed data) and asserts: a clean
      # repo reports no issues (exit 0), a repo missing README/CITATION/LICENSE fails
      # under --fail-on-issue (exit 1), and an HF card carrying a PHI-shaped string is
      # flagged. Stdlib-only (json/re), network-free, ASCII-only.
      set -u
      
      HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
      SCRIPT="$HERE/../scripts/batch_metadata_audit.py"
      TMP="$(mktemp -d -t aio_audit_XXXX)"
      trap 'rm -rf "$TMP"' EXIT
      
      fail=0
      check_exit() { local label="$1" want="$2"; shift 2
          "$@" >/dev/null 2>&1; local got=$?
          if [[ "$got" -eq "$want" ]]; then printf '  PASS  %s\n' "$label"
          else printf '  FAIL  %s (exit %s, want %s)\n' "$label" "$got" "$want"; fail=$((fail+1)); fi
      }
      check() { local label="$1"; shift
          if "$@" >/dev/null 2>&1; then printf '  PASS  %s\n' "$label"
          else printf '  FAIL  %s\n' "$label"; fail=$((fail+1)); fi
      }
      
      [[ -f "$SCRIPT" ]] || { echo "ENV-ERR: batch_metadata_audit.py missing" >&2; exit 2; }
      
      # --- Clean repo: README (DOI + citation + quickstart), CITATION.cff, LICENSE ---
      CLEAN="$TMP/clean_repo"; mkdir -p "$CLEAN"
      cat > "$CLEAN/README.md" <<'MD'
      # Synthetic Tool
      [![DOI](https://img.shields.io/badge/DOI-10.5281%2Fzenodo.0000000-blue)](https://doi.org/10.5281/zenodo.0000000)
      ## Installation
      pip install synthetic-tool
      ## How to cite
      See CITATION.cff.
      MD
      cat > "$CLEAN/CITATION.cff" <<'CFF'
      cff-version: 1.2.0
      title: Synthetic Tool
      version: 1.0.0
      authors:
        - family-names: Kim
          given-names: Alice
          orcid: https://orcid.org/0000-0002-1825-0097
      CFF
      echo "MIT License" > "$CLEAN/LICENSE"
      check_exit "clean repo -> exit 0 (no issues, --fail-on-issue)" 0 \
          python3 "$SCRIPT" "$CLEAN" --fail-on-issue
      
      # --- Broken repo: directory exists but empty (all three artifacts missing) ---
      BROKEN="$TMP/broken_repo"; mkdir -p "$BROKEN"
      check_exit "repo missing README/CITATION/LICENSE -> exit 1" 1 \
          python3 "$SCRIPT" "$BROKEN" --fail-on-issue
      # Without --fail-on-issue the same audit still exits 0 (report-only).
      check_exit "report-only mode -> exit 0 even with issues" 0 \
          python3 "$SCRIPT" "$BROKEN"
      
      # --- HF card with a PHI-shaped string (KR resident registration number) ---
      CARD="$TMP/model_card.md"
      cat > "$CARD" <<'MD'
      ---
      license: mit
      library_name: transformers
      tags:
        - medical
      ---
      # Model
      ## Intended use
      Research only.
      ## Training data
      Synthetic records, e.g. subject 900101-1234567 was excluded.
      ## Evaluation
      AUC reported.
      ## Limitations
      Small sample.
      ## Ethical considerations
      De-identified.
      MD
      JSON_OUT="$TMP/report.json"
      python3 "$SCRIPT" --hf-card "$CARD" --output "$JSON_OUT" >/dev/null 2>&1
      check "HF card report written" test -s "$JSON_OUT"
      check "PHI pattern flagged in HF card" python3 -c "
      import json
      d=json.load(open('$JSON_OUT'))
      issues=' '.join(d['hf_cards'][0]['issues'])
      assert 'PHI' in issues, d['hf_cards'][0]['issues']"
      check_exit "HF card with PHI -> exit 1 under --fail-on-issue" 1 \
          python3 "$SCRIPT" --hf-card "$CARD" --fail-on-issue
      
      echo "fail=$fail"; [[ "$fail" -eq 0 ]] && echo "ALL PASS" || echo "FAILURES: $fail"
      exit "$fail"
      
    • test_summary_box.sh 2.9 KB
      #!/usr/bin/env bash
      # Regression/challenge test for check_summary_box.py — deterministic, network-free,
      # synthetic fixtures built at runtime. Covers each journal format + each hard rule.
      set -u
      
      HERE="$(cd "$(dirname "$0")" && pwd)"
      SCRIPT="$HERE/../scripts/check_summary_box.py"
      TMP="$(mktemp -d)"
      trap 'rm -rf "$TMP"' EXIT
      
      pass=0
      fail=0
      ck() {
        local label="$1" expected="$2" actual="$3"
        if [ "$expected" = "$actual" ]; then
          printf '  PASS  %-50s exit=%s\n' "$label" "$actual"
          pass=$((pass + 1))
        else
          printf '  FAIL  %-50s expected=%s actual=%s\n' "$label" "$expected" "$actual"
          fail=$((fail + 1))
        fi
      }
      run() { python3 "$SCRIPT" --out "$TMP/r.json" "$@" > /dev/null 2>&1; echo $?; }
      
      # 1) conformant Key Points (3 one-claim bullets) -> exit 0
      cat > "$TMP/kp_ok.md" <<'EOF'
      ## Key Points
      - The model improved detection sensitivity in an internal test set.
      - Specificity was preserved at the chosen operating threshold.
      - External validation is still required before deployment.
      EOF
      ck "key_points conformant" 0 "$(run --manuscript "$TMP/kp_ok.md" --journal radiology --strict)"
      
      # 2) wrong bullet count (2) -> NONCONFORMANT under --strict
      cat > "$TMP/kp_bad.md" <<'EOF'
      ## Key Points
      - The model improved detection sensitivity.
      - Specificity was preserved.
      EOF
      ck "key_points wrong bullet count fails" 1 "$(run --manuscript "$TMP/kp_bad.md" --journal radiology --strict)"
      
      # 3) Research in context missing a sub-block -> NONCONFORMANT
      cat > "$TMP/ric_bad.md" <<'EOF'
      ## Research in context
      **Evidence before this study** We searched PubMed for prior work.
      **Added value of this study** This study adds an external cohort.
      EOF
      ck "research_in_context missing subblock fails" 1 "$(run --manuscript "$TMP/ric_bad.md" --journal lancet-digital-health --strict)"
      
      # 4) Research in context complete -> CONFORMANT
      cat > "$TMP/ric_ok.md" <<'EOF'
      ## Research in context
      **Evidence before this study** We searched PubMed for prior work.
      **Added value of this study** This study adds an external cohort.
      **Implications of all the available evidence** Findings support a prospective trial.
      EOF
      ck "research_in_context complete conformant" 0 "$(run --manuscript "$TMP/ric_ok.md" --journal lancet-digital-health --strict)"
      
      # 5) Plain-language summary over the band -> NONCONFORMANT
      { echo "## Plain-language summary"; for i in $(seq 1 260); do printf 'word '; done; echo; } > "$TMP/pls_bad.md"
      ck "plain_language over-length fails" 1 "$(run --manuscript "$TMP/pls_bad.md" --journal npj-digital-medicine --strict)"
      
      # 6) absent box -> NONCONFORMANT
      echo "## Abstract" > "$TMP/none.md"
      ck "absent box fails" 1 "$(run --manuscript "$TMP/none.md" --format key_points --strict)"
      
      # 7) without --strict, a nonconformant box is reported but tolerated (exit 0)
      ck "nonconformant tolerated without --strict" 0 "$(run --manuscript "$TMP/kp_bad.md" --journal radiology)"
      
      echo "----"
      echo "test_summary_box: $pass passed, $fail failed"
      [ "$fail" -eq 0 ]
      
    • test_validate_schema.sh 2.6 KB
      #!/usr/bin/env bash
      # Regression test for academic-aio/scripts/validate_schema.py.
      # Builds synthetic JSON-LD fixtures (no committed data) and asserts the
      # validator's contract: a complete ScholarlyArticle passes; wrong @context,
      # unknown @type, a missing required field, and a malformed DOI each fail.
      # Stdlib-only (json/re), network-free, ASCII-only.
      set -u
      
      HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
      SCRIPT="$HERE/../scripts/validate_schema.py"
      TMP="$(mktemp -d -t aio_schema_XXXX)"
      trap 'rm -rf "$TMP"' EXIT
      
      fail=0
      check() { local label="$1" want="$2"; shift 2
          "$@" >/dev/null 2>&1; local got=$?
          if [[ "$got" -eq "$want" ]]; then printf '  PASS  %s\n' "$label"
          else printf '  FAIL  %s (exit %s, want %s)\n' "$label" "$got" "$want"; fail=$((fail+1)); fi
      }
      
      [[ -f "$SCRIPT" ]] || { echo "ENV-ERR: validate_schema.py missing" >&2; exit 2; }
      
      # Valid ScholarlyArticle (all required fields, canonical DOI).
      cat > "$TMP/ok.jsonld" <<'JSON'
      {
        "@context": "https://schema.org",
        "@type": "ScholarlyArticle",
        "headline": "Synthetic Diagnostic Study",
        "datePublished": "2026-01-01",
        "author": [{"@type": "Person", "name": "Alice Kim"}],
        "identifier": [{"@type": "PropertyValue", "propertyID": "DOI", "value": "10.1000/synthetic.2026.001"}],
        "url": "https://example.org/article"
      }
      JSON
      check "valid ScholarlyArticle -> exit 0" 0 python3 "$SCRIPT" "$TMP/ok.jsonld"
      
      # Wrong @context.
      sed 's#https://schema.org#https://example.com#' "$TMP/ok.jsonld" > "$TMP/bad_ctx.jsonld"
      check "wrong @context -> exit 1" 1 python3 "$SCRIPT" "$TMP/bad_ctx.jsonld"
      
      # Unknown @type.
      sed 's/ScholarlyArticle/UnicornType/' "$TMP/ok.jsonld" > "$TMP/bad_type.jsonld"
      check "unknown @type -> exit 1" 1 python3 "$SCRIPT" "$TMP/bad_type.jsonld"
      
      # Missing required field (no "url"; still valid JSON).
      cat > "$TMP/missing.jsonld" <<'JSON'
      {
        "@context": "https://schema.org",
        "@type": "ScholarlyArticle",
        "headline": "Synthetic Diagnostic Study",
        "datePublished": "2026-01-01",
        "author": [{"@type": "Person", "name": "Alice Kim"}],
        "identifier": [{"@type": "PropertyValue", "propertyID": "DOI", "value": "10.1000/synthetic.2026.001"}]
      }
      JSON
      check "missing required field -> exit 1" 1 python3 "$SCRIPT" "$TMP/missing.jsonld"
      
      # Malformed DOI.
      sed 's#10.1000/synthetic.2026.001#not-a-doi#' "$TMP/ok.jsonld" > "$TMP/bad_doi.jsonld"
      check "malformed DOI -> exit 1" 1 python3 "$SCRIPT" "$TMP/bad_doi.jsonld"
      
      # Mixed batch (one bad file) still fails overall.
      check "batch with one bad file -> exit 1" 1 python3 "$SCRIPT" "$TMP/ok.jsonld" "$TMP/bad_ctx.jsonld"
      
      echo "fail=$fail"; [[ "$fail" -eq 0 ]] && echo "ALL PASS" || echo "FAILURES: $fail"
      exit "$fail"
      
  • SKILL.md 32 KB
    ---
    name: academic-aio
    description: Medical AI paper optimization for AI search engines (Perplexity, ChatGPT web, Elicit, Consensus, SciSpace) and RAG-based literature tools. Applies when drafting or reviewing titles, abstracts, structured summary boxes (Key Points / Research in Context / Plain-Language Summary), manuscripts for high-impact medical AI journals (Lancet Digital Health, Radiology, Radiology-AI, npj Digital Medicine, Nature Medicine), preprints (medRxiv/arXiv), GitHub README + CITATION.cff + Zenodo archives, and Hugging Face model/dataset cards. Integrates TRIPOD+AI, CLAIM 2024, STARD-AI, TRIPOD-LLM, DECIDE-AI reporting requirements with generative engine optimization (GEO) principles. Produces a visible pass/fail checklist.
    triggers: AIO, LLMO, GEO, AI search optimization, discoverability, abstract optimization, structured abstract, Key Points, Research in context, plain-language summary, preprint strategy, GitHub README, CITATION.cff, Zenodo DOI, Hugging Face model card, dataset card, Perplexity, Elicit, Consensus, SciSpace, RAG visibility, reporting guideline compliance, TRIPOD-AI, CLAIM, STARD-AI, taxonomy review paper, Radiology Key Points, Lancet Digital Health Research in context, npj Digital Medicine
    tools: Read, Write, Edit, Bash, Grep, Glob
    model: inherit
    ---
    
    # Academic AIO Skill — Medical AI Paper Visibility for AI Search Engines
    
    You are helping a medical-AI researcher optimize a paper, preprint, README, or code release so that it is surfaced and cited accurately by AI search engines (Perplexity, ChatGPT web, Elicit, Consensus, SciSpace), RAG-based literature tools, and traditional scholarly indexes (Semantic Scholar, Google Scholar, PubMed). Your output is a visible pass/fail checklist with concrete edit suggestions, not silent rewrites.
    
    ## Communication Rules
    
    - Surface the checklist in the response. Never apply AIO edits silently.
    - Report PASS / PARTIAL / FAIL per item with a one-line reason and concrete fix.
    - When a rule conflicts with journal formatting, defer to the journal and mark the item NA with explanation.
    - Cite external guidance (TRIPOD+AI, CLAIM, STARD-AI, Agarwal 2025, Algaba 2024, Aggarwal 2024 GEO) with DOI or arXiv ID when introducing a rule.
    - Do not hallucinate citations. If unsure, mark as `[VERIFY]`.
    
    ## When to Invoke
    
    Run this skill when the user is working on any of:
    - Drafting or revising a title, abstract, structured-summary box, or plain-language summary.
    - Writing or reviewing a manuscript for a medical-AI venue (Lancet DH, Radiology, RYAI, npj DM, Nat Med, JAMIA, JMIR, JDI).
    - Preparing a preprint (medRxiv, arXiv, bioRxiv, Research Square).
    - Composing a GitHub README, `CITATION.cff`, Zenodo archive metadata, Hugging Face model card, or dataset card.
    - Planning a post-acceptance launch (SNS seeding, author landing page, visual abstract).
    - Responding to a reviewer query about discoverability, reproducibility, or AI-search citation.
    
    Pairs with (do not duplicate):
    - `write-paper` — Phase 6 (draft) and Phase 7 (QC). AIO rules extend the title/abstract/discussion sections.
    - `check-reporting` — reporting-guideline item audit (TRIPOD+AI, CLAIM, etc.). AIO requires guideline adherence but does not reproduce the audit.
    - `self-review` — adversarial review. Run AIO after self-review so QC-confirmed claims anchor the checklist.
    - `humanize` — AI-pattern removal. Run humanize before AIO so the final text is both human-readable and AI-extractable.
    
    ## Core Thesis
    
    Generative engine optimization research (Aggarwal 2024, arXiv:2311.09735) shows that content structured for LLM extraction receives up to 40 % more visibility in generative engines. In medicine this effect is mediated by three gates:
    
    1. **Open-access full text** — tools like Elicit and Consensus cannot extract columns from paywalled PDFs; Perplexity Academic favors OA citations.
    2. **Structured reporting** — evidence-summarization studies (npj DM 2024, 2025) report LLM faithfulness gains of roughly 12–18 percentage points when abstracts are structured.
    3. **Machine-readable artifacts** — CITATION.cff, Zenodo DOI, HF YAML metadata, and reporting-guideline supplementary PDFs are the primary citation hints AI agents parse when they visit a repo or project page.
    
    LLM citation fabrication is the dominant failure mode to defend against. Agarwal et al. (Nat Commun 2025, doi:10.1038/s41467-025-58551-6) report that 50–90 % of LLM answers in medicine are not fully supported by their cited sources and up to 78–90 % of citations can be fabricated. The defensive strategy is to surface a paper's DOI and PMID in easy-to-copy form so that LLMs substitute the correct identifier instead of confabulating one.
    
    ## Section 1 — Title and Abstract Optimization
    
    ### 1.1 Title three-slot rule
    Structure: `[Task] + [Modality or anatomy] + [Model family or method class]`. Include one concrete differentiator (dataset scale, new benchmark, "first …") when defensible. Avoid keyword stuffing (penalized as spam by AI overviews).
    
    Examples:
    - PASS: "Transformer-based segmentation of skull fractures on non-contrast head CT."
    - FAIL: "A novel advanced deep-learning AI machine-learning framework for medical image analysis."
    
    ### 1.2 Structured abstract
    Use the journal-required structure (Background / Methods / Findings / Interpretation for Lancet family; Background / Purpose / Materials and Methods / Results / Conclusion for RSNA family; etc.). If the journal allows unstructured, still use an internally structured form. Each section stands alone as a semantic chunk of ≤ 3 sentences so that chunk-boundary splits in RAG indexes do not break the claim.
    
    ### 1.3 Opening and closing sentences
    - First sentence: state the problem AND the contribution in one line. LLM summarizers extract this disproportionately.
    - Last sentence: explicit interpretation ("we show that …", "this implies …"). No hedging-only closes.
    
    ### 1.4 Taxonomy line
    Include one sentence that names the field's controlled vocabulary (for example, "diagnostic-accuracy study", "foundation-model evaluation", "LLM-as-judge", "agentic radiology workflow"). Entity linkers in AI indexes use this line.
    
    ### 1.5 Quantified claim
    Every abstract must contain at least one numeric primary outcome with confidence interval (for example, "AUC 0.94 [95 % CI 0.91–0.96]" or "sensitivity 88.2 % [95 % CI 85.1–91.0]"). LLM retrievers weight papers with concrete numbers.
    
    ### 1.6 Reporting-guideline anchor
    Place the guideline name in the abstract or the opening sentence of Methods: "Reported following TRIPOD+AI (Collins 2024) and CLAIM 2024 (Tejani 2024)". When applicable add STARD-AI 2025, DECIDE-AI, TRIPOD-LLM. This signals structure to LLMs and satisfies reviewer checklists.
    
    AIO-rule ↔ guideline-item mapping: `references/reporting_guideline_mapping.md`.
    
    ### 1.7 Keyword, MeSH, and RadLex coverage
    Title, abstract, and keywords together should cover ≥ 3× the surface area of the concept — no redundancy. Include:
    - Core MeSH terms (verify against the NLM MeSH browser).
    - Radiology-specific RadLex terms where applicable.
    - Modality-synonym coverage ("chest radiograph (CXR)", "non-contrast CT (NCCT)").
    - Both US and UK spellings when relevant.
    
    Royal Society 2024 (doi:10.1098/rspb.2024.1222) reports that 92 % of papers waste keyword real estate by repeating title terms in abstract and keywords; avoid this.
    
    ## Section 2 — Manuscript-Level AIO
    
    ### 2.1 Summary box
    Include the journal-specific summary box verbatim when supported:
    - Lancet family: "Research in context" (Evidence before this study / Added value / Implications).
    - RSNA Radiology and RYAI: "Key Points" — 3 bullets, one claim each.
    - npj Digital Medicine: "Plain-language summary" (150–200 words, 8th-grade reading level).
    - Nature Medicine: editor's summary (supplied by editorial, but draft one proactively).
    
    **Deterministic format check.** Validate the drafted box against its journal spec with `python3 ${CLAUDE_SKILL_DIR}/scripts/check_summary_box.py --manuscript <file> --journal <stem> --strict` (reads `references/summary_box_specs.json`: Key Points bullet count + one-claim-per-bullet, Research-in-context's three sub-blocks, plain-language word band). It catches the wrong-format / wrong-bullet-count box that a production technical check rejects.
    
    These boxes are the fragments Perplexity and ChatGPT web most often copy or paraphrase verbatim; treat them as the paper's canonical citation surface.
    
    Journal-specific templates (USER MUST VERIFY against current IFA): `references/journal_summarybox_templates.yaml`.
    
    ### 2.2 Declarative section headings
    Section and subsection headings should state a claim, not a generic label. "Model underperforms on rare-finding subset" beats "Subgroup analysis".
    
    ### 2.3 Numeric claim compression
    In the Methods and in at least one Results paragraph, compress primary-outcome statistics into a single sentence pattern:
    "On the internal test set (n = 842), the model achieved AUC 0.94 (95 % CI 0.91–0.96), sensitivity 88.2 % (85.1–91.0), specificity 91.4 % (88.7–93.6), at an operating point of 0.37."
    
    This pattern is the canonical shape LLM extractors parse first.
    
    ### 2.4 Reproducibility block
    Include a labeled block (typically end of Methods or a standalone Data/Code Availability section) listing: data availability and license, code availability with DOI, model weights and checkpoints, prompts and configuration files, random seeds, compute environment. This block is disproportionately scraped by AI agents when they cite a paper as reproducible.
    
    ### 2.4a Citing an AI-assisted tool by use-class
    If the work used an AI-assisted tool (a verification/QA suite, analysis code, or a generative drafting assistant), frame the mention by *what it did*, not by hiding it — under current journal wariness, a proud in-text citation of a **generative** use invites suspicion, while the same placement for a **verification** use reads as rigor (like citing a reference manager or a linter). Split the mention: **verification/QA and analysis** → a **Software / Code-availability statement** (citable); **generative** drafting/humanizing → the journal's **AI-use disclosure field**, not a citation. A self-citation by the tool's author additionally requires a COI disclosure, and you should cite only the functions the work actually used. Full use-class table and rules: `${CLAUDE_SKILL_DIR}/references/ai_tool_citation_framing.md`.
    
    ### 2.5 Limitations enumeration
    List limitations explicitly and name each one (generalizability, spectrum bias, dataset shift, single-center training, label noise). Papers with enumerated limitations score higher for trustworthiness in LLM summarization benchmarks.
    
    ### 2.6 Standalone figure captions
    Each caption should re-state the claim, the dataset, and the metric. Captions survive in vector databases and image-retrieval indexes when surrounding body text is lost.
    
    ## Section 3 — Preprint, Channel, and Indexing Strategy
    
    ### 3.1 Preprint versus fast-track
    - Default: post to medRxiv (clinical), arXiv (methods, cs.CV / eess.IV), or bioRxiv on the day of journal submission. Rapid preprinting puts the paper into Semantic Scholar within 24–72 hours and into Perplexity's web index immediately.
    - Exception: if the target journal offers a fast-track review cycle (acceptance → online within roughly 30–60 days) AND the authors prefer a single canonical version, a preprint may be skipped. In that case, compensate by aggressive post-acceptance SNS seeding and PMC deposit.
    - Never skip preprint AND fast-track — this is the discoverability deadzone.
    
    ### 3.2 Journal preprint-policy table (verify before submission)
    Most medical-AI venues allow preprints (Radiology, RYAI, Lancet DH, npj DM, Nature Medicine, JAMIA, JMIR, Cell Reports Medicine, Cell Patterns). A few have restrictions or require disclosure. Always verify the current policy on Sherpa Romeo or the journal's instructions-for-authors page before posting.
    
    ### 3.3 Indexing time-lag (2025 baseline)
    - Perplexity Academic / ChatGPT web: real-time web crawl, citable on publication day.
    - Semantic Scholar: 24–72 hours from DOI or preprint.
    - Google Scholar: 1–7 days.
    - PMC (NIH deposit): 2–6 weeks for accepted manuscripts; longer for CC-BY-NC.
    - Elicit and Consensus: follow Semantic Scholar / OpenAlex.
    - LLM training corpora (next model generation): 6–18 months.
    
    Plan launch activities around these windows.
    
    ### 3.4 Open-access choice
    Prefer gold OA with CC-BY when budget allows. If not, green OA via preprint plus author-accepted manuscript is acceptable. Closed-access papers without preprint lose roughly 30–50 % of AI-tool citations because Elicit, Consensus, and Perplexity Academic cannot extract from paywalled PDFs.
    
    Funder OA-policy decision tree (Plan S, NIH, UKRI, Gates, Wellcome, NRF, MoHW): `references/oac_funding_checklist.yaml`.
    
    ### 3.5 Post-acceptance channel checklist
    - Deposit AAM to PMC or Europe PMC.
    - Update ORCID and Google Scholar profile.
    - Post to Threads / X / BlueSky with DOI, one-sentence claim, and figure.
    - Long-form post on LinkedIn (targets different LLM training corpora).
    - Submit to Papers with Code if the paper reports a benchmark.
    - Upload model or dataset to Hugging Face with a model/dataset card.
    
    ## Section 4 — Review-Paper Strategy
    
    Review articles function as hub nodes in knowledge graphs and accrue "lookup citations" when readers need a canonical reference for a taxonomy. For researchers building a portfolio in medical AI:
    
    - Target at least one review or taxonomy paper per year in a top-tier venue.
    - Include 5 or more taxonomy tables (model class, dataset, task type, evaluation metric, failure mode). Each table becomes a lookup target.
    - Cite 100 or more primary references for breadth; 150+ for canonical status.
    - Co-author with a consortium of 10+ investigators from multiple institutions when possible — this multiplies social-network reach and citation dispersal.
    - Pair the review with a companion dataset, benchmark, or code artifact on Zenodo or Hugging Face to anchor AI-tool citations.
    
    Empirically, review papers with these properties outperform original research on short-term FWCI while feeding traffic to the authors' original papers through reverse citation.
    
    ## Section 5 — GitHub, CITATION.cff, Zenodo, Hugging Face
    
    ### 5.1 README canonical 10-slot order
    1. Title + one-line description + badges (license, DOI, arXiv, Hugging Face, paper link).
    2. Paper reference block — BibTeX + APA + two-sentence abstract.
    3. TL;DR — at most 5 bullets: problem, approach, key result, intended users.
    4. Quickstart — `pip install` or `git clone && make demo`. Should work in under 5 minutes.
    5. Reproducibility — exact commands that regenerate every figure and table. Pin package versions.
    6. Project structure — a tree with one-line folder descriptions.
    7. Data access — license, download scripts, DUA notes.
    8. FAQ — "How is this different from X?", "Can this be used clinically?", "How do I cite this?". High-value retrieval content.
    9. Acknowledgements, funding, and COI.
    10. License (prefer Apache-2.0 for research code).
    
    ### 5.2 CITATION.cff
    Add a `CITATION.cff` file at repository root. GitHub renders it as a "Cite this repository" button, and AI agents treat it as the primary citation hint. Include authors with ORCID, title, version, DOI (post-Zenodo-archive), repository URL, and license.
    
    ### 5.3 Zenodo DOI
    Enable GitHub–Zenodo integration for each release. Cite the version-specific DOI in the paper's Data/Code Availability section. Zenodo deposits appear in Google Scholar and OpenAlex, creating an independent citable artifact.
    
    ### 5.4 Hugging Face model card YAML
    Required keys: `license`, `library_name`, `tags`, `datasets`, `base_model` (when fine-tuning), `pipeline_tag`. Required prose sections: Intended use, Training data, Evaluation, Limitations, Ethical considerations, and a clinical-use disclaimer ("This model is not approved for clinical diagnostic use; it is provided for research purposes only").
    
    ### 5.5 Hugging Face dataset card
    Required prose: license, PHI and re-identification risk, task, language, splits, annotation process, known biases, ethical review status.
    
    ### 5.6 Web-crawler-friendly formatting
    - Markdown headings are declarative claims.
    - Code blocks are fenced and language-tagged.
    - Tables are plain Markdown, not HTML (survive Markdown-to-vector chunking).
    - Images have descriptive alt text (vision-LLMs read alt text when image retrieval fails).
    - Each README section is under about 300 words to survive fixed-size chunking.
    - Use question-style subheadings when natural ("Why another benchmark?", "How fast is inference?").
    - Embed JSON-LD `ScholarlyArticle` / `SoftwareSourceCode` / `Dataset` / `Person` markup in repository pages and author landing pages — templates in `references/schema_markup_templates/`, validated with `python scripts/validate_schema.py path/to/file.jsonld`.
    
    ## Section 6 — Authority and E-E-A-T Signals
    
    - Maintain a personal author landing page (GitHub Pages, personal domain, or institutional page) that lists all papers with DOIs and open-access links. AI indexes weight author-entity pages.
    - Use one consistent affiliation string across papers. Inconsistency fragments the author entity in knowledge graphs and loses citation velocity.
    - Keep ORCID complete and linked to Google Scholar. Re-run author-disambiguation on Semantic Scholar every 6 months.
    - Cross-link related papers by the same group in Discussion sections when defensible. Within-group self-citation increases co-retrieval probability in RAG.
    - Refresh repository and model cards quarterly — articles updated quarterly outperform single-publish articles in AI-overview retention (Conductor 2026 benchmark).
    
    ## Section 7 — LLM-Citation Fabrication Defense
    
    Given Agarwal et al. Nat Commun 2025 (doi:10.1038/s41467-025-58551-6) findings that up to 78–90 % of LLM medical citations can be fabricated, take the following defensive steps:
    
    - Surface DOI and PMID in copy-friendly text at the top of the paper's landing page and README (for example, `DOI: 10.xxxx/yyyy • PMID: 12345678`).
    - Add a "How to cite" section with BibTeX, APA, Vancouver, and the plain-text line in one place.
    - Monitor incorrect citations. Set a Google Scholar alert for the paper's title variant; periodically query Perplexity and ChatGPT web for the paper and record hallucinated bibliographic errors.
    - When responding to a reviewer who cites an LLM-generated reference, verify the DOI and PMID yourself before accepting.
    
    ## Section 8 — Red Flags
    
    - Closed code described as "available on reasonable request" — scrapers treat this as "not reproducible" and AI tools demote the paper.
    - Paywall-only with no preprint and no fast-track — invisible to most RAG pipelines.
    - Keyword-stuffed titles ("A deep-learning artificial-intelligence machine-learning system for …") — penalized as spam.
    - Abstracts opening with filler ("In recent years, AI has revolutionized …") — burns the chunk most likely to be extracted.
    - Walls of theory before the README quickstart.
    - "Clinical grade" or "replaces radiologists" overclaims — demoted by LLM trust heuristics and may trigger reviewer rejection.
    - PHI leakage in Hugging Face dataset samples.
    - Inconsistent author affiliations across co-authored papers.
    
    ## Section 9 — Per-Project Application and Pipeline Integration
    
    When invoked, run in this order:
    
    1. Read the target artifact (title, abstract, manuscript section, README, or card) and **identify its lifecycle phase**: `pre-draft` / `drafting` / `pre-submission` / `post-acceptance` / `post-publication`.
    2. Apply Sections 1–5 and 10 relevant to that artifact, **filtering each rule by its `applies_to_phase` field in `references/checklists/AIO_GENERAL.md`**. Out-of-phase rules become NA rather than FAIL (e.g., do not surface §11.5 multi-disciplinary roster or §12 launch sequencing as FAIL on a `pre-submission` audit). Produce a PASS / PARTIAL / FAIL table **sorted by `expected_lift` (high → medium → low)**. Render via `templates/aio_audit_checklist.md.j2` when programmatic.
    3. **Honour `defers_to` annotations to avoid duplicate audits.** Items annotated with a `defers_to` field record only present/absent status here; item-level detail belongs to the linked skill or reference (§1.6 → `/check-reporting`; §3.4 / §11.3 → `references/oac_funding_checklist.yaml`). Cross-check reporting-guideline anchor (§1.6) by invoking `/check-reporting` first when the manuscript has not been audited; the AIO ↔ guideline-item mapping is in `references/reporting_guideline_mapping.md`.
    4. Apply Section 6 author-authority audit once per submission cycle. Sections 11.1–11.5 are `pre-draft` rules — `applies_to_phase` filter auto-NAs them once drafting is complete. Section 12 launch sequencing fires only at `post-acceptance` / `post-publication`.
    5. Surface Section 7 citation-defense recommendations at `post-acceptance` time. For multi-repo or Hugging-Face-card team audits, run `scripts/batch_metadata_audit.py`.
    6. Output:
       - The phase-filtered checklist (visible).
       - A short deferred-item list with one-line status per `defers_to` rule.
       - At most 5 concrete edits **ranked by `expected_lift`** (high first, then medium, then low). Edits whose underlying rule is `low`-lift should not appear in the Top 5 unless no `high` / `medium` items remain open.
    
    ### Integration with `write-paper`
    - Phase 4 (Title and abstract drafting) → apply Section 1 as an inline filter.
    - Phase 6 (Discussion) → apply Section 2.5 (limitations) and Section 6 (cross-linking).
    - Phase 7 (QC) → run AIO after reporting-guideline check and numerical-claim audit.
    
    ### Output template
    
    ```
    ## Academic AIO Checklist — [Artifact type]
    
    | # | Item | Status | Note |
    |---|------|--------|------|
    | 1.1 | Title three-slot | PASS/PARTIAL/FAIL | … |
    | 1.2 | Structured abstract | PASS/PARTIAL/FAIL | … |
    | ... | ... | ... | ... |
    
    ## Top 5 suggested edits
    1. …
    2. …
    ```
    
    ## Section 10 — Q&A and Entity-Extraction Optimization
    
    Modern RAG indexes parse Q&A blocks more reliably than free-form prose; LLM citation engines preferentially extract claim-restatement pairs. Section 10 augments retrievability by structuring how claims are restated and how entities are linked.
    
    ### 10.1 Four-question Q&A block (Discussion or Appendix)
    
    Add a labeled Q&A block — either as the closing subsection of Discussion, or as a Supplementary Box. Pattern:
    
    - **What was known before this study?** — two-sentence restatement of the prior state.
    - **What does this study add?** — two-sentence statement of the contribution.
    - **How might this change clinical practice or research?** — one-sentence interpretation; avoid overclaim.
    - **Why does this matter?** — one-sentence "so what" framing for non-specialists.
    
    This block is the canonical fragment that AI-overview systems extract and cite. Lancet Digital Health "Research in context" already encodes the first two questions; the Q&A block extends them and is parseable by LLM web-search agents.
    
    ### 10.2 Glossary block with entity IDs
    
    Define each domain-specific acronym inline on first use AND list them in a Glossary subsection at end of Methods or Supplementary. Attach the canonical entity ID where possible:
    
    - MeSH term ID for clinical concepts.
    - RadLex ID for radiology-specific terms.
    - UMLS CUI for cross-vocabulary mapping.
    - Hugging Face model ID for named models.
    - arXiv ID for cited methods.
    
    Entity linkers in Elicit, Consensus, and SciSpace use this metadata to connect a paper to knowledge graphs.
    
    ### 10.3 Inline citation anchor text
    
    Avoid bare reference numbers. Use semantic anchor patterns so LLM extractors bind the citation to the specific claim:
    
    - WEAK: "Prior work [12] showed efficacy."
    - STRONG: "Smith et al. (DOI: 10.xxxx/yyyy) reported a 12 % accuracy gain on the MIMIC-CXR test set [12]."
    
    When citing one's own prior work, name the cohort or dataset explicitly to enable cross-paper retrieval.
    
    ### 10.4 Explicit challenge statement
    
    Beyond Section 2.5 (limitations enumeration), include a single-paragraph "Why this is hard" challenge statement near the start of Discussion. Pattern:
    
    "Building accurate [task] for [modality/anatomy] is constrained by [data scarcity / label noise / dataset shift / regulatory uncertainty / interpretability]. Each of these has been documented [refs], and our results address [subset]."
    
    LLM web-search systems quote challenge statements as authoritative summaries of field state. The 2025 KJR multimodal-LLM review used this pattern (e.g., "lack of large-scale high-quality multimodal datasets") and was preferentially extracted by Perplexity and ChatGPT web (see `references/case_studies/kjr_mllm_2025.md`).
    
    ## Section 11 — First-Mover Timing and Citation-Graph Density
    
    Topic timing is the most under-discussed AIO lever. Reviews and original research published at the peak of a topic's hype curve accrue citations disproportionately; reviews that lag the peak by 6–12 months under-perform regardless of quality.
    
    ### 11.1 Topic peak detection
    
    Signals that a topic is approaching peak (write now, publish in ~6 months):
    
    - arXiv/medRxiv monthly deposit rate growing > 20 % month-over-month for 3+ consecutive months.
    - Major model release (GPT-4o, Claude 3.5 multimodal, MedGemini) introducing a capability not previously available.
    - Funding agency Request-for-Applications (RFA) addressing the topic.
    - Society guidelines (RSNA, ACR, ESR) calling for evaluation studies.
    - Sustained > 1,000 weekly impressions on Twitter/X/LinkedIn for related papers.
    
    Plan submission so publication lands at peak, not after.
    
    ### 11.2 Editorial-board leverage
    
    If a corresponding author serves on the target journal's editorial board, review-process median time often drops noticeably (KJR: ~4–6 weeks faster; varies by journal). Editor's-pick or issue-highlight selection can also drive Google News indexing within 24 hours of publication.
    
    When recruiting senior co-authors for a review paper, prefer those who hold an editorial role at the target venue. This is a legitimate editorial signal, not a conflict-of-interest issue, provided board members recuse themselves from review of their own submissions per ICMJE guidance.
    
    ### 11.3 PMC-auto-deposit journal preference
    
    Open-access journals that automatically deposit to PubMed Central (PMC) reach LLM crawlers within 4–6 weeks of publication; non-PMC OA journals can take 3–6 months. PMC-auto-deposit journals in radiology/medical-AI (verify per submission, policies change):
    
    - Korean Journal of Radiology (KJR) — auto-deposit confirmed.
    - Lancet Digital Health — author-funded green OA, PMC-eligible after embargo.
    - Radiology and Radiology: AI — selected articles auto-deposit.
    - npj Digital Medicine — auto-deposit (Nature OA).
    - JAMIA — author-funded OA route.
    - JMIR — auto-deposit (PMC-indexed).
    
    When all else is equal, prefer PMC-auto-deposit journals to compress the LLM-discoverability window.
    
    ### 11.4 Citation-graph anchor strategy
    
    Discussion sections should anchor the paper in 5–10 high-visibility prior works that LLM training corpora already index well. This raises co-citation probability and makes the paper retrievable when users query the seminal works.
    
    - Identify seminal references via Semantic Scholar's "Highly Influential Citations" filter for the topic.
    - Cite them with semantic predicates (Section 10.3), not as bare lists.
    - Mix recent preprints (currency signal) with 2018–2022 seminal papers (graph anchoring) — corpora-cutoff means 2024–2025-only citation profiles have low LLM retrieval weight.
    
    ### 11.5 Multi-disciplinary author roster
    
    Author-affiliation diversity multiplies indexing entry points. A 10–15 author team spanning 3+ institutions and 2+ disciplines (clinical + computational) creates more author-entity nodes in Google Scholar and Semantic Scholar, each acting as a discovery surface. The 2025 KJR MLLM review used a 15-author team spanning resident + engineer + medical student + faculty across 5 institutions and accrued 64 citations within 7 months (see case study).
    
    ## Section 12 — Cross-Platform Launch Sequencing
    
    Section 3.5 (post-acceptance channel checklist) is unordered; Section 12 prescribes the timing. The first 30 days after publication are the primary discoverability window for AI-search engines and LLM training-data harvesters.
    
    ### 12.1 Day 0 — publication day (execute simultaneously)
    
    - GitHub release (tag a stable version; let Zenodo mint a version-specific DOI).
    - Hugging Face model card + dataset card (if applicable); link arXiv ID and DOI.
    - Twitter/X + Threads + Bluesky: 1-sentence claim + key figure + DOI in copy-friendly format.
    - LinkedIn announcement (long-form): hook line + structured claim block + DOI.
    - Author landing-page update with PDF link (OA) or AAM.
    
    ### 12.2 Day 1 — propagation
    
    - Update ORCID with DOI, abstract, and authorship role.
    - Update Google Scholar (verify auto-detection within 24h; manual add if delayed).
    - Update preprint server with "Accepted" version note + link to published version.
    - Update institutional profile / department news page.
    
    ### 12.3 Week 1 — depth posts
    
    - LinkedIn second post: long-form interpretation or methods spotlight.
    - Papers with Code submission (if benchmark or model with public weights).
    - ResearchGate upload of AAM (per journal policy).
    - Reddit/Hacker News post if the work has broad appeal (assess fit honestly).
    
    ### 12.4 Weeks 2–4 — refresh signals
    
    - README and HF card minor update (new badges, new FAQ entries).
    - Follow-up blog or Substack post expanding on one figure or limitation.
    - Respond to reader questions on social platforms — those answers themselves become indexed content.
    
    ### 12.5 Month 1 — monitoring
    
    - Google Scholar alert for the paper title.
    - Semantic Scholar / Scite citation alerts.
    - Quarterly probe: query Perplexity, ChatGPT web, Elicit, Consensus, SciSpace with 3–5 expected discovery queries; record retrieval position and any hallucinated bibliographic errors.
    - If a fabricated citation appears, update the README "How to cite" block (Section 7) to maximize copy-friendliness of the correct identifier.
    
    ## External References
    
    - GEO: Generative Engine Optimization — Aggarwal et al., KDD 2024, arXiv:2311.09735.
    - LLM medical citation fabrication — Agarwal et al., Nat Commun 2025, doi:10.1038/s41467-025-58551-6.
    - LLM citation bias — Algaba et al., 2024, arXiv:2405.15739.
    - ExpertQA attribution — Malaviya et al., 2024, arXiv:2309.07852.
    - TRIPOD+AI — Collins et al., BMJ 2024. EQUATOR Network.
    - CLAIM 2024 — Tejani et al., Radiology: AI 2024, doi:10.1148/ryai.240300.
    - STARD-AI — Sounderajah et al., Nat Med 2025, doi:10.1038/s41591-025-03953-8.
    - TRIPOD-LLM — Gallifant et al., Nat Med 2024, doi:10.1038/s41591-024-03425-5.
    - DECIDE-AI — Vasey et al., Nat Med 2022, doi:10.1038/s41591-022-01772-9.
    - Title, abstract, keywords guide — Royal Society Proc B 2024, doi:10.1098/rspb.2024.1222.
    - GitHub repository citation advantage — Yan et al., Inf Process Manag 2024, doi:10.1016/j.ipm.2023.103569.
    - Semantic Scholar Open Data Platform — Kinney et al., arXiv:2301.10140.
    
    ## Anti-Hallucination
    
    - **Never fabricate citations, DOIs, arXiv IDs, or reporting-guideline item numbers.** Every cited reporting framework (TRIPOD+AI, CLAIM, STARD-AI, TRIPOD-LLM, DECIDE-AI) must map to a verifiable DOI or EQUATOR Network entry. Mark unverified items as `[UNVERIFIED - NEEDS MANUAL CHECK]`.
    - **Never invent journal-specific summary-box rules** (Lancet Digital Health "Research in context", Radiology "Key Points", npj Digital Medicine). Verify current instructions-to-authors from the journal's website before applying.
    - **Never fabricate discoverability metrics** (Perplexity/Elicit/Consensus retrieval scores) — only report observed behavior from a recorded probe.
    - **Never auto-complete author lists, ORCIDs, or affiliations** in CITATION.cff or Zenodo metadata; surface empty slots to the user.
    - If a compliance item, journal policy, or AI-search platform behavior is uncertain, state the uncertainty rather than guessing.
    
    ## Global-rule references
    
    Some passages in this skill cite a path of the form `~/.claude/rules/<name>.md`. Those are the
    maintainer's personal global rules, kept outside this repository. They are **not shipped with
    this skill** and will not exist on your machine; they appear only as provenance for where a
    convention came from. If one of them looks like it is standing in for an instruction you actually
    need, that is a bug — please open an issue, because the instruction belongs here.
    
  • skill.yml 1.6 KB
    schema_version: 2
    name: academic-aio
    layer: C
    owner_domain: manuscript_optimization
    maturity: official
    
    when_to_use: "Optimize a medical-AI manuscript's title, abstract, and structured summary boxes for AI search engines and RAG tools, integrating TRIPOD+AI/CLAIM/STARD-AI reporting requirements."
    when_NOT_to_use: "Drafting a manuscript from scratch (use write-paper); removing AI writing tells (use humanize)."
    
    inputs:
      - "draft manuscript / abstract / title (Markdown)"
    outputs:
      - "AIO-optimized title, abstract, structured summary boxes"
      - "visible pass/fail checklist"
    deterministic_scripts:
      - scripts/validate_schema.py
      - scripts/batch_metadata_audit.py
    side_effects:
      - writes_project_artifacts
    downstream_consumers:
      - self-review
      - check-reporting
    forbidden_actions:
      - silently_rewrite_without_user_review
      - fabricate_reporting_guideline_compliance
    
    # v2.1 quality card
    purpose: "Make a medical-AI manuscript discoverable and citable by AI search/RAG tools without sacrificing reporting-guideline compliance."
    safety_boundaries:
      - "Off by default in autonomous pipelines; rewrites require user review (no silent rewrite)."
      - "Reporting-guideline claims are checked, never asserted without the underlying item being present."
    known_limitations:
      - "GEO/AIO heuristics evolve with each engine; recommendations are point-in-time."
      - "Does not draft new scientific content; optimizes existing approved text only."
    validation_commands:
      - "python3 scripts/validate_schema.py <summary-box-file>"
      - "python3 scripts/check_summary_box.py --manuscript <file> --journal <stem> --strict"
      - "bash tests/test_summary_box.sh"
    evidence_surface: bundled_script
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related