check-reporting
Check manuscript compliance with medical research reporting guidelines. Supports 49 guidelines including STROBE, STROBE-MR, RECORD, REMARK (prognostic tumor-marker studies), TARGET (target trial emulation), GATHER (burden-of-disease / health-estimate modeling), CONSORT, CONSORT-A
Install
npx skills add https://github.com/Aperivue/medsci-skills/tree/main/skills/check-reporting
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install aperivue-medsci-skills@llmmart
git clone https://github.com/Aperivue/medsci-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole aperivue/medsci-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Check-Reporting Skill
You are helping a medical researcher verify that their manuscript complies with the appropriate medical research reporting guideline. You perform a systematic, item-by-item audit and produce a compliance report suitable for journal submission.
Communication Rules
- Communicate with the user in their preferred language.
- Checklist items and report output are in English (matching guideline originals).
- Medical terminology is always in English.
Reference Files
- Checklists (bundled, open license):
${CLAUDE_SKILL_DIR}/references/checklists/STROBE.md-- observational studies (CC BY)STROBE_MR.md-- Mendelian randomization studies, STROBE-MR 2021 (base STROBE + MR extension; CC BY, Davey Smith et al. BMJ 2021)STARD.md-- diagnostic accuracy studies (CC BY 4.0)STARD_AI.md-- AI diagnostic accuracy studies (CC BY, Sounderajah et al. Nat Med 2025)TRIPOD.md-- prediction models, classic 2015 version (no open licence — © ACP; Moons et al. Ann Intern Med 2015)TRIPOD_AI.md-- prediction models with AI/ML (CC BY 4.0, Collins et al. BMJ 2024)TRIPOD_LLM.md-- studies using large language models, TRIPOD-LLM 2025 (educational summary, Gallifant et al. Nat Med 2025)PGS_RS.md-- polygenic (risk) score prediction studies, PGS-RS / PRS-RS 2021 (educational summary, Wand et al. Nature 2021)CHEERS_2022.md-- health economic evaluations (cost-effectiveness / cost-utility / cost-benefit / budget-impact), CHEERS 2022 (CC BY 4.0, Husereau et al. BMJ 2022)RECORD.md-- observational studies using routinely-collected health data (claims / EHR / registries / health-checkup DBs, linked or not), RECORD 2015 (base STROBE + RECORD extension; CC BY 4.0, Benchimol et al. PLoS Med 2015; RECORD-PE for drug studies)CROSS.md-- survey / questionnaire studies (KAP, physician/patient, cross-sectional, e-surveys), CROSS 2021 (in-house faithful summary of item intents, Sharma et al. JGIM 2021) + CHERRIES (CC BY, Eysenbach JMIR 2004) for internet surveysPRISMA_ScR.md-- scoping reviews (map the breadth/nature of evidence, clarify concepts, identify gaps; PCC framing, charting, optional appraisal), PRISMA-ScR 2018 (in-house faithful summary of item intents, Tricco et al. Ann Intern Med 2018; DOI 10.7326/M18-0850)SRQR.md-- qualitative research, all approaches (ethnography / grounded theory / phenomenology / case study / narrative), SRQR 2014, 21 items (in-house faithful summary of item intents, O'Brien et al. Acad Med 2014; DOI 10.1097/ACM.0000000000000388)COREQ.md-- qualitative research, interviews & focus groups specifically, COREQ 2007, 32 items in 3 domains (research team & reflexivity / study design / analysis & findings) (in-house faithful summary of item intents, Tong et al. Int J Qual Health Care 2007; DOI 10.1093/intqhc/mzm042)REMARK.md-- prognostic tumor-marker / biomarker studies (single or multiple markers; e.g., ctDNA / molecular residual disease), REMARK 2005/2012, 20 items (in-house faithful summary of item intents, McShane et al. Br J Cancer 2005 + Altman et al. PLoS Med 2012)TARGET.md-- observational studies emulating a target trial (causal / comparative-effectiveness questions on routinely-collected / registry / EHR data), TARGET 2025, 21 items (in-house faithful summary of item intents, Cashin/Hansford/Hernán et al. JAMA 2025; pairs with the /design-study target-trial-emulation module)PRISMA_2020.md-- systematic reviews (CC BY)PRISMA_2020_Abstracts.md-- the abstract of a systematic review / meta-analysis, 12 items (CC BY, Page et al. BMJ 2021). A separate instrument from the 27-item checklist, not a subset: item 2 of the main checklist defers to it. Score it with its own denominator.ARRIVE_2.md-- animal studies (CC0)PRISMA_DTA.md-- DTA systematic reviews (no open licence — © AMA; McInnes et al. JAMA 2018)QUADAS3.md-- diagnostic accuracy risk of bias, current recommended version (no open licence -- (c) ACP; Whiting et al. Ann Intern Med 2026)QUADAS2.md-- diagnostic accuracy risk of bias (no open licence — © ACP; Whiting et al. Ann Intern Med 2011)RoB2.md-- RCT risk of bias (CC BY, Sterne et al. BMJ 2019)ROBINS_I.md-- non-randomised studies risk of bias (CC BY-NC 3.0 — non-commercial; Sterne et al. BMJ 2016)PROBAST.md-- prediction model risk of bias (no open licence — © ACP; Wolff et al. Ann Intern Med 2019)NOS.md-- observational study quality (public domain, Ottawa Hospital)CONSORT.md-- randomised controlled trials, CONSORT 2025 (CC BY 4.0, Hopewell et al. BMJ 2025)CONSORT_AI.md-- AI clinical-trial reports, CONSORT-AI 2020 (CC BY 4.0, Liu et al. Nat Med 2020)CARE.md-- case reports, CARE 2013 (no confirmed open licence — Elsevier TDM only; Gagnier et al. J Clin Epidemiol 2014)SPIRIT.md-- clinical trial protocols, SPIRIT 2025 (CC BY 4.0, Chan et al. BMJ 2025)SPIRIT_AI.md-- AI clinical-trial protocols, SPIRIT-AI 2020 (CC BY 4.0, Cruz Rivera et al. Nat Med 2020)CLAIM_2024.md-- AI/ML in clinical imaging, CLAIM 2024 Update (RSNA open access, Tejani et al. Radiol Artif Intell 2024)DECIDE_AI.md-- early-stage clinical evaluation of AI decision-support systems, DECIDE-AI 2022 (educational summary, CC BY-NC, Vasey et al. Nat Med 2022)MI_CLEAR_LLM.md-- LLM accuracy studies in healthcare (CC BY-NC 4.0, Park et al. KJR 2024; 2025 update)SQUIRE_2.md-- quality improvement in healthcare/education (no open licence — Crossref returns none; Ogrinc et al. BMJ Qual Saf 2016)CLEAR.md-- radiomics studies (CC BY 4.0, Kocak et al. Insights Imaging 2023)MOOSE.md-- meta-analysis of observational studies (Stroup et al. JAMA 2000)GRRAS.md-- reliability and agreement studies (Kottner et al. J Clin Epidemiol 2011)QUADAS_C.md-- comparative DTA risk of bias, extension to QUADAS-2 (no open licence — © ACP; Yang et al. Ann Intern Med 2021)ROBINS_E.md-- non-randomised exposure studies risk of bias (CC BY-NC-ND 4.0, Higgins et al. Environ Int 2024)ROBIS.md-- risk of bias in systematic reviews (Whiting et al. J Clin Epidemiol 2016)ROB_ME.md-- risk of bias due to missing evidence in meta-analysis (no open licence — BMJ TDM policy only; Page et al. BMJ 2023)PROBAST_AI.md-- prediction model risk of bias, updated for AI/ML (Moons et al. BMJ 2025)COSMIN_RoB.md-- reliability/measurement error risk of bias (Mokkink et al. BMC Med Res Methodol 2020)RoB_NMA.md-- risk of bias in network meta-analysis (Lunny et al. 2024)AMSTAR2.md-- quality of systematic reviews (Shea et al. BMJ 2017)PRISMA_P.md-- systematic review protocols (Shamseer et al. BMJ 2015)SWiM.md-- synthesis without meta-analysis reporting (Campbell et al. BMJ 2020)GATHER.md-- health-estimate / burden-of-disease modeling studies (GBD and GBD-satellite, comparative-risk / population-attributable-fraction, cause-of-death and prevalence/incidence estimation, with or without forecasts), GATHER 2016 (in-house faithful summary; CC BY, Stevens et al. Lancet 2016;388:e19-23 / PLoS Med 2016;13(6):e1002056). Pairs with/analyze-statsreferences/analysis_guides/burden_decomposition_forecasting.mdfor the analytic methods.
- Fail-fast contract: if a routed guideline has no vendored checklist file, the skill does not silently construct items from memory. It halts with a
MISSING_CHECKLIST_CONTRACT_VIOLATIONand surfaces the gap. A from-memory assessment is allowed only with the explicit--allow-from-memoryopt-in, and that report must be clearly labelled NON-AUTHORITATIVE. See Step 2 andscripts/check_checklist_exists.py. - Critical-item floor:
${CLAUDE_SKILL_DIR}/references/critical_item_floor.md-- the small set of non-waivable items per study type (presence outranks the headline %), plus the AI/radiomics methodological-quality / risk-of-bias instruments (PROBAST+AI, METRICS/RQS, APPRAISE-AI) kept distinct from their reporting counterparts. Loaded in Step 4f.
Workflow
Step 0: Existing-checklist staleness pre-check
If a checklist already exists for this project (qc/reporting_checklist.json or a prior .md report), verify it targets the current manuscript before reusing it — a checklist generated against an older version carries stale section/line references and a stale version label that a reviewer who cross-checks will catch:
python3 "${CLAUDE_SKILL_DIR}/scripts/check_checklist_version.py" \
--checklist qc/reporting_checklist.json --manuscript manuscript_v8.md
A non-zero exit means the existing checklist is stale (older target_version, changed source_sha256, different target_manuscript) or pre-dates the version contract — regenerate it against the current manuscript (Steps 1–5) rather than reusing it. Every report you generate must carry the target_manuscript / target_version / source_sha256 fields (Part A header + Part D JSON) so this check works next round.
Step 1: Select Guideline
Determine the appropriate reporting guideline. Auto-detect from the manuscript type or accept user specification.
Auto-detection mapping:
| Study Type | Primary Guideline | AI Extension |
|---|---|---|
| Observational study | STROBE | -- |
| Mendelian randomization study | STROBE-MR (base STROBE + MR extension) | -- |
| Health economic evaluation (cost-effectiveness / cost-utility / cost-benefit / budget-impact) | CHEERS 2022 | -- |
| Observational study using routinely-collected data (claims / EHR / registry / health-checkup DB) | RECORD (base STROBE + RECORD extension; RECORD-PE for drug studies) | -- |
| Survey / questionnaire study (KAP, physician/patient, cross-sectional, e-survey) | CROSS (+ CHERRIES for internet surveys) | -- |
| Scoping review (maps breadth/nature of evidence, clarifies concepts, identifies gaps — not a focused effectiveness/accuracy question) | PRISMA-ScR (base PRISMA + scoping-review extension) | -- |
| Qualitative study (interviews, focus groups, ethnography, grounded theory, phenomenology, document analysis) | SRQR (all qualitative approaches); COREQ (interviews/focus groups specifically) | -- |
| Randomized controlled trial | CONSORT 2025 | CONSORT-AI |
| Diagnostic accuracy study | STARD 2015 | STARD-AI |
| Prediction model (development/validation) | TRIPOD | TRIPOD+AI |
| Polygenic (risk) score prediction study | PGS-RS (with TRIPOD / TRIPOD+AI) | -- |
| Prognostic tumor-marker / biomarker study (single or multiple markers; e.g., ctDNA / molecular residual disease) | REMARK (pair with STROBE for the observational-design items; TRIPOD / TRIPOD+AI if a prognostic model is developed) | -- |
| Causal / comparative-effectiveness question emulated on observational data (treatment vs treatment, screening vs none, drug A vs B on registry / EHR / claims data) | TARGET (pair with the /design-study target-trial-emulation module for design; RECORD / STROBE for the routinely-collected-data items) | -- |
| Health-estimate / burden-of-disease modeling study (GBD or GBD-satellite, comparative-risk / population-attributable-fraction, cause-of-death or prevalence/incidence estimation, with or without forecasts) | GATHER (pair with /analyze-stats burden-decomposition-forecasting guide for the analytic layer) |
-- |
| Systematic review / meta-analysis | PRISMA 2020 | PRISMA 2020 for Abstracts (run on the abstract, scored separately) |
| DTA systematic review / meta-analysis | PRISMA-DTA | PRISMA 2020 for Abstracts (run on the abstract, scored separately) |
| Meta-analysis of observational studies | MOOSE | PRISMA 2020 (use both) |
| Risk of bias (DTA studies) | QUADAS-3 (current recommended version) | QUADAS-2 only when appraising or reproducing a review that used it |
| Risk of bias (RCTs) | RoB 2 | -- |
| Risk of bias (non-randomised intervention studies) | ROBINS-I | -- |
| Risk of bias (non-randomised exposure studies) | ROBINS-E | -- |
| Risk of bias (comparative DTA studies) | QUADAS-C | QUADAS-3 (use both; apply the E&E's adaptation — see Using QUADAS-C with QUADAS-3 in QUADAS3.md) |
| Risk of bias (prediction models) | PROBAST | PROBAST+AI |
| Risk of bias (systematic reviews) | ROBIS | AMSTAR 2 |
| Risk of bias (missing evidence in MA) | ROB-ME | -- |
| Risk of bias (network meta-analysis) | RoB NMA | -- |
| Risk of bias (measurement properties) | COSMIN RoB | -- |
| Quality assessment (observational) | NOS | -- |
| Case report | CARE | -- |
| Study protocol | SPIRIT 2025 | SPIRIT-AI |
| Animal study | ARRIVE 2.0 | -- |
| AI/ML study in clinical imaging | CLAIM 2024 | -- |
| Study using a large language model (develop/fine-tune/prompt/evaluate an LLM) | TRIPOD-LLM | MI-CLEAR-LLM (use alongside when LLM accuracy is an outcome) |
| Early-stage / live clinical evaluation of an AI decision-support system (human factors, workflow, safety) | DECIDE-AI | -- |
| LLM accuracy evaluation in healthcare | MI-CLEAR-LLM | STARD-AI or CLAIM 2024 (use alongside) |
| Reliability / agreement study | GRRAS | -- |
| SR protocol | PRISMA-P | -- |
| Synthesis without meta-analysis | SWiM | PRISMA 2020 (use both) |
| Quality of systematic reviews | AMSTAR 2 | ROBIS |
| Radiomics study | CLEAR | CLAIM 2024 (if deep learning component) |
| Educational / QI study | SQUIRE 2.0 | -- |
| Generative AI images ARE the study object (realism / real-vs-synthetic reader study / model-vs-model quality) | (no single guideline -- assemble) | see decision aid below |
QUADAS-3 has two protocol-stage phases, and this skill usually runs too late for them. Phase 1 (state the synthesis question) and phase 2 (define the ideal test accuracy trial each judgement is made against) are review-level and belong in the protocol, alongside the review-specific guidance for answering each signalling question. Reaching them for the first time during manuscript QC means writing the comparator after seeing the results. If they are missing, say so as a limitation rather than reconstructing them — and route the protocol work to
/meta-analysisPhase 1. Phases 3–6 are what a QC pass can genuinely run.
Rules:
- If the study involves AI/ML, always apply the AI extension in addition to the base guideline.
- Exception — TRIPOD: TRIPOD+AI 2024 (Collins et al., BMJ 2024) is a complete rewrite, not an addendum to TRIPOD 2015 (Moons et al., Ann Intern Med 2015). For non-AI prediction models, use TRIPOD 2015 only. For AI/ML prediction models, use TRIPOD+AI 2024 only. Do NOT apply both simultaneously.
- STARD-AI (Sounderajah et al., Nat Med 2025) extends STARD 2015 with 14 new and 4 modified items (40 total). For AI diagnostic accuracy studies, use STARD-AI (which incorporates all STARD 2015 items). Do NOT apply both STARD 2015 and STARD-AI simultaneously — STARD-AI supersedes STARD 2015 for AI studies.
- TRIPOD-LLM (Gallifant et al., Nat Med 2025) is the reporting guideline for studies that develop, fine-tune, prompt, or evaluate a large language model for a clinical/biomedical task. It extends the TRIPOD family (TRIPOD 2015 → TRIPOD+AI 2024 → TRIPOD-LLM 2025); name the base instrument and the extension and cite each. It is modular — task-specific items (Annotation, Prompting, Summarization, Instruction-tuning) are N/A when that component is absent. Use TRIPOD-LLM for LLM studies in place of TRIPOD+AI; pair with MI-CLEAR-LLM when LLM accuracy is an evaluated outcome. The vendored checklist is an educational summary (own-words paraphrase of item intent); complete the official instrument for a submission checklist.
- MI-CLEAR-LLM is a supplementary checklist (8 item categories in the 2025 update; the 2024 original had 6), not a standalone reporting guideline. Always pair it with the study's primary guideline (e.g., STARD-AI for AI diagnostic accuracy, CLAIM for imaging AI). Apply MI-CLEAR-LLM whenever the study evaluates LLM accuracy as an outcome — do NOT apply it merely because the manuscript was written with LLM assistance. Its scope is LLM accuracy studies (including VLMs interpreting images); it does not apply at study level to studies where a generative model produces the images under study (see next bullet).
- Generative-AI images as the study object (a generative model synthesizes images and the study evaluates their realism, controllability, real-vs-synthetic distinguishability, or model-vs-model quality) has no single dominant checklist. Assemble: CLAIM 2024 (imaging-AI umbrella; model-development items N/A when commercial models are used as-is) + FUTURE-AI traceability + MI-CLEAR-LLM transparency items only (prompt/model/version/params/runs — for generation provenance, not study-level compliance) on the generator side; STARD-AI (for real-vs-synthetic detection) + GRRAS (reader reliability) + MRMC reporting on the evaluation side. Map applicable items and cite base + extension; never claim wholesale compliance. Full decision aid:
${CLAUDE_SKILL_DIR}/references/genai_image_study_object_decision_aid.md. - If multiple guidelines apply (e.g., a diagnostic accuracy study that is also an AI study), check against all relevant guidelines and merge into one report.
- If the user requests a specific guideline, use that one regardless of auto-detection.
Step 2: Load Checklist
Run the fail-fast guard first for every guideline you intend to apply:
python "${CLAUDE_SKILL_DIR}/scripts/check_checklist_exists.py" --guideline "STARD-AI"- Exit 0 → the vendored checklist exists; read it from
${CLAUDE_SKILL_DIR}/references/checklists/and proceed. - Exit 1 (
MISSING_CHECKLIST_CONTRACT_VIOLATION) → the guideline is routed but no checklist file is vendored. Do not construct items from memory. Halt, report the violation to the user, and stop unless they explicitly opt in (next bullet). - Exit 2 (
UNKNOWN_GUIDELINE) → the name is not recognised; confirm the correct guideline with the user.
- Exit 0 → the vendored checklist exists; read it from
No silent fallback. A from-memory checklist is permitted only when the user explicitly accepts it — re-run the guard with
--allow-from-memory(exit 0 + a NON-AUTHORITATIVE warning). In that case the output report MUST carry a prominent banner that the assessment was constructed from model knowledge and is not backed by a vendored checklist, andsubmission_safemust not be asserted on its basis.
Step 3: Scan Manuscript
Read all sections of the manuscript thoroughly:
- Title and abstract
- Introduction
- Methods (all subsections)
- Results (all subsections)
- Discussion
- Tables, figures, and their captions
- Supplemental materials (if available)
- References (for registration numbers, protocol references)
Gather context from the full document before starting the item-by-item assessment.
Step 4: Assess Each Item
For every checklist item, determine:
| Status | Criteria |
|---|---|
| PRESENT | The item is fully addressed with sufficient detail. |
| PARTIAL | The item is mentioned or partially addressed but lacks required detail. |
| MISSING | The item is not found anywhere in the manuscript. |
| N/A | The item does not apply to this particular study (justify why). |
For each item, record:
- Status: PRESENT / PARTIAL / MISSING / N/A
- Location: Section name and paragraph or approximate position (e.g., "Methods, paragraph 3")
- Notes: What was found (if PRESENT/PARTIAL) or what should be added (if MISSING)
What is appraised is the source paper's reporting — never your convenience in using it. This holds for every instrument here, reporting checklists and risk-of-bias / quality tools alike, and it is easiest to lose in a systematic review, where you read each paper in order to extract from it. An item asking "are the results clearly reported?" is not asking "were they reported in the unit my pool needs".
A scorer working a case-series quality tool marks a paper down on the outcome-reporting item because its analysis unit does not match the pool's — treatment-level results against a patient-level denominator. The correction is one sentence, that is a limit of our extraction, not a defect in their reporting, and the score goes back up. Single-scorer appraisal is where this happens, because there is nobody to say it.
So: if a downgrade's stated reason turns on a denominator, an analysis unit, a subgroup you needed and they did not report separately, or a format you could not parse, it is an extraction note, not a scoring reason. Record it in a separate extraction-note column and restore the score.
Both belong in the table. An extraction limitation is a real constraint on your synthesis and often belongs in your limitations paragraph — it just is not evidence about the paper being appraised, and folding it into the score makes the appraisal unreproducible: another assessor with a different pool would score the same paper differently.
Step 4b: Section Boundary Check
In addition to checklist items, verify that:
- Results section contains only factual findings: no interpretation, no "why" explanations, no prior literature comparisons, no evaluative adjectives without numbers.
- Discussion section does not introduce new data not presented in Results.
- Flag any boundary violation as a separate finding in Part C Action Items with the label
[BOUNDARY].
Step 4c: Registration / Protocol Timing Consistency Check
Applies to: systematic reviews, meta-analyses, and intervention studies with prospective registration (PRISMA 2020, PRISMA-DTA, PRISMA-P, MOOSE, CONSORT, SPIRIT).
Why this step exists: the registration identifier is a single checklist item and can pass Step 4 even when the manuscript is internally inconsistent about when the registration or its amendments occurred relative to the analysis. An undisclosed post-hoc amendment is a common rejection trigger.
Five audit items (summary): (1) registration identifier present in Methods, Abstract, and cover letter; (2) initial registration date precedes — or is explicitly disclosed as post-dating — the extraction milestone; (3) amendment dates appear in Methods, the described change is visible in Methods, analysis was re-run if amendment post-dates the lock, and no amendment post-dates submission; (4) cross-artifact agreement between Methods and the registry record (PROSPERO PDF, ClinicalTrials.gov export) — silent discrepancy is a finding; (5) retrospective-registration disclosure paragraph when evidence suggests post-extraction filing.
Registration-ID format gate: a PROSPERO ID is CRD42 + 9 digits = 14 characters
(^CRD42\d{9}$, e.g. CRD42024500001). Run grep -oE 'CRD42[0-9]+' manuscript.md and
assert each match is 14 characters long; a 15-character ID (a stray inserted digit) is a
transcription error logged as [REGISTRATION-TIMING] (fixable_by_ai: false — verify against
the live PROSPERO record, do not guess the correct digit).
Flagging: any failure is logged in Part C Action Items with label
[REGISTRATION-TIMING]. fixable_by_ai: false when reconciliation requires an external
amendment filing; true only when the fix is a Methods-text insertion of a date already
disclosed elsewhere. Part D JSON includes a registration_timing object
(registry, id, initial_registration_date, amendments[], timing_consistency, findings[]).
Load-on-demand procedural detail (exact item-by-item procedure, JSON schema,
flagging edge cases): ${CLAUDE_SKILL_DIR}/references/step4c_registration_timing.md.
Step 4d: PRISMA Figure 1 Arithmetic & Cross-Reference Audit
Applies to: systematic reviews and meta-analyses using PRISMA 2020 / PRISMA-DTA / PRISMA-P. Triggers when Item 16a (flow diagram) is PRESENT.
Why this step exists: the flow diagram is a single checklist item and can pass Step 4 visually while still containing arithmetic errors (records screened ≠ identified − duplicates; sought-for-retrieval ≠ screened − excluded) or text↔figure number disagreements. Senior MA reviewers commonly require strict PRISMA 2020 diagram conformance and explicit body↔ figure number agreement; reviewers who detect these mismatches lose confidence in the study's data integrity immediately.
Four arithmetic checks:
- records screened = records identified − duplicates removed
- records sought-for-retrieval = records screened − records excluded (screening)
- reports retrieved = sought − reports not retrieved
- studies included = reports assessed for eligibility − reports excluded (with reasons)
Two cross-reference checks:
- Body text PRISMA numbers (e.g., "315 records identified, 122 duplicates removed, 186 records screened") match Figure 1 box labels 1:1.
- Reasons for exclusion (Methods + Figure legend) agree on counts and category names.
Procedure:
Run the deterministic implementation first — it performs steps 1, 4, 5, and 6 below
automatically (same keyword regex, the four arithmetic equations, the body↔figure
cross-reference) and writes qc/prisma_figure_audit.json:
python3 ${CLAUDE_SKILL_DIR}/scripts/check_prisma_figure.py \
--md <manuscript.md> --figure <Figure 1 source: .md manifest / caption / text export> \
--out qc/prisma_figure_audit.json
Exit 1 = an arithmetic or cross-reference MISMATCH (log a Part C Action Item labelled
[PRISMA-FIGURE], fixable_by_ai: false — the author must reconcile the numbers); exit
2 = missing/unparsable input. The manual algorithm below documents exactly what the
script checks and is the fallback when Figure 1 numbers live only in a PNG/SVG that must
be transcribed by hand:
- Extract numbers from manuscript Results / PRISMA flow paragraph (regex: integers near
keywords
identified,duplicates,screened,excluded,sought,retrieved,assessed,included). - Extract numbers from Figure 1 source — preferred order: (a)
analysis/figures/Figure1_PRISMA.mdmarkdown manifest, (b) caption text inmanuscript.md, (c) PPTX text run if.pptxexists, (d) manual entry from PNG/SVG. - Cross-check
analysis/figures/_figure_manifest.md(produced by/make-figures): verify that the row whoseType = prisma(orType = prisma-dta) points at the same file path used as the audit source, and that the row'sCriticfield isyesorpartial(notno). A missing manifest row, mismatched path, orCritic = noflag logs[MANIFEST-XREF](advisory) — the arithmetic check still runs against the source identified in step 2. Skip this sub-step if_figure_manifest.mddoes not exist (older projects). - Run 4 arithmetic checks; emit PRESENT / MISSING / MISMATCH per equation.
- Run 2 cross-reference checks; emit PRESENT / MISSING / MISMATCH per number.
- Output
qc/prisma_figure_audit.jsonand a short table.
Flagging: any MISMATCH or arithmetic failure logs a Part C Action Item with label
[PRISMA-FIGURE]. fixable_by_ai: false (numbers must be reconciled by the author).
Load-on-demand procedural detail (exact regex set, JSON schema, edge cases —
duplicates handled across databases, citation searching strand, dual-reviewer screening):
${CLAUDE_SKILL_DIR}/references/step4d_prisma_figure_audit.md.
Cross-cutting: integrates with ~/.claude/rules/numerical-safety.md (PRISMA 5-way
consistency: text ↔ Figure ↔ extraction CSV ↔ analysis script ↔ supplementary).
Step 4e: Reporting-Framework Naming Audit
Applies to: any manuscript that invokes an AI/extension reporting framework (PROBAST+AI, STARD-AI, TRIPOD+AI, TRIPOD-LLM, CONSORT-AI, SPIRIT-AI, PRISMA-DTA, QUADAS-C).
Why this step exists: a base reporting tool and its extension are distinct instruments
with separate citations (manuscript-style-classical §14). Step 1 routes to the right
checklist but does not police how the framework is named in prose. The recurring failures
are: invoking an extension without ever naming or citing the base instrument it extends;
mixing +AI and -AI hyphenation for one family within a single document; coining item
labels like "12-AI"; and waving at "recent guidance" instead of naming the framework.
Run the deterministic gate:
python3 "${CLAUDE_SKILL_DIR}/scripts/check_framework_naming.py" \
--manuscript manuscript.md --out qc/framework_naming.json --strict
Verdicts: BASE_MISSING (extension used, base instrument never named standalone) is a
Major and logs [FRAMEWORK-NAMING] in Part C with fixable_by_ai: true (insert the base
name + its citation). HYPHEN_MIX, CITE_MISSING, SELF_COINED_LABEL, and VAGUE_GUIDANCE
are Minor (fixable_by_ai: true). Part D JSON includes a framework_naming object mirroring
the script's claims[].
Step 4f: Critical-item floor cross-check
Applies to: every guideline assessment for which the floor defines a row (load and
check only those; do not invent a floor for an unlisted guideline). After the item-by-item
table, load ${CLAUDE_SKILL_DIR}/references/critical_item_floor.md and check the small set
of non-waivable items for this study type. A MISSING critical item is surfaced as a
Critical gap and becomes the report's headline regardless of the overall percentage —
a high percentage with a missing critical item (undefined reference standard, no
leakage-controlled partition, calibration absent for a prediction model, an unreconciled
flow diagram) is not "broadly acceptable."
For AI/ML and radiomics manuscripts, also confirm the chosen methodological-quality /
risk-of-bias instrument (PROBAST+AI, METRICS/RQS, APPRAISE-AI) and its non-waivable
concerns — a fully reported paper can still be at high risk of bias. For radiomics, the
fuller METRICS breakdown (9 categories / 30 weighted items) is in
${CLAUDE_SKILL_DIR}/references/appraisal_tools/METRICS.md (an appraisal reference, not a counted
reporting checklist). Keep these distinct
from the reporting counterparts (CLEAR, DECIDE-AI), which route through the normal checklist
flow. Do not assert a numeric journal desk-reject threshold; the hard signals are a missing
critical item and the journal's own required elements.
Step 5: Generate Report
Produce a structured compliance report in four parts.
This report is an internal working audit — it carries auto-fix annotations, a
machine-readable JSON block (compliance_pct, fixable_by_ai, …), and Action Items. It is
NOT the official reporting checklist a journal expects (that is the blank guideline form with
Item | Recommendation | Reported in page/section, which the authors fill in). Never submit
this report as the submission checklist. So that the file is self-identifying and cannot be
reused by filename into a later submission package, the report MUST begin with this banner as
its very first line:
<!-- INTERNAL AUDIT — NOT FOR SUBMISSION. This is the /check-reporting working
report, not the official journal checklist. Do not upload to a submission portal. -->
(/sync-submission's check_checklist_dump_leak gate also catches this dump if it ever lands in
a submission directory — but the banner is what makes it catchable.)
The four parts — literal templates in ${CLAUDE_SKILL_DIR}/references/report_templates.md:
- Part A — Summary. Header (manuscript file, version token, guideline, date), the
PRESENT/PARTIAL/MISSING/N-A count table, and overall compliance. The headline is the critical
items (Step 4f), not the percentage: report
{present}/{total}and name every missing critical item with the section it belongs in. - Part B — Item-by-item checklist. One row per item:
# | Section | Item | Status | Location | Notes. - Part C — Action items (MISSING and PARTIAL only), ordered by: items most journals enforce strictly (ethics approval, registration, sample size) → items in Methods (easiest to fix) → everything else.
- Part D — Machine-readable JSON, appended as a fenced block. MUST be present under
--jsonor when called from/write-paperPhase 7, which parses it.
JSON field contract (the part other skills depend on — get these right):
compliance_pct—present / (total_items - na) * 100, one decimal.action_items— MISSING and PARTIAL only; PRESENT and N/A are excluded.fixable_by_ai—truewhen the fix inserts or expands text using information already in the manuscript or inferable from it;falsewhen it needs external facts the author alone holds (registration number, IRB approval number, protocol details).suggested_fix— concrete draft text, insertable as written.source_sha256— first 12 hex chars of the SHA-256 of the manuscript bytes, so a stale report cannot be silently attributed to a newer manuscript.
Read on demand:
| File | Read it when | Cost if read blindly |
|---|---|---|
references/report_templates.md |
you have finished the audit and are writing the report | ~1,900 tokens of pure output format — it informs no part of the assessment itself |
Assessment Standards
Be Strict
- PARTIAL means the item is mentioned but lacks specificity. For example:
- "We used appropriate statistical tests" = PARTIAL (which tests?)
- "We used the Mann-Whitney U test for continuous variables and Fisher's exact test for categorical variables" = PRESENT
- A vague reference does not count as PRESENT. The detail level must match what the guideline expects.
Be Specific in Suggestions
- For MISSING items, provide a draft sentence the user can insert.
- For PARTIAL items, point to the exact gap and suggest specific additions.
- Reference the specific manuscript section where the addition should go.
Common Gaps to Watch For
These items are frequently missing in medical manuscripts:
- Study registration number (CONSORT, PRISMA, STARD)
- Registration / amendment date consistency (PRISMA 2020, PRISMA-DTA, CONSORT, SPIRIT) — run Step 4c whenever a registration identifier is present
- Sample size justification (CONSORT, STROBE, STARD)
- Missing data handling (all guidelines)
- Blinding details (CONSORT, STARD)
- Funding and conflicts of interest (all guidelines)
- Ethics approval with committee name and approval number (all guidelines)
- Data availability statement (increasingly required)
- AI-specific: training/validation/test split details (TRIPOD+AI, CLAIM, STARD-AI)
- AI-specific: model architecture and hyperparameters (TRIPOD+AI, CLAIM, STARD-AI)
- AI-specific: failure mode analysis (CLAIM, STARD-AI)
- AI-specific: fairness/bias assessment (STARD-AI)
- AI-specific: commercial interests and data/code availability (STARD-AI)
- Power-aware framing of a null result (STROBE 16a / 18 / 20) — for an observational study whose headline is a non-significant association, a flat "X was not associated with Y" overreads the data when the analysis is not powered to exclude a clinically meaningful effect. Mark item 18/20 PARTIAL unless the manuscript states the precision as an exclusion (e.g., "the 95% CI excluded an eGFR difference larger than ~1.7") or reports a minimum detectable effect — "no effect" vs "could not exclude an effect of size X" are different claims, and a negative conclusion needs the latter.
- Confounder-selection rationale, not "adjust for everything that differs" (STROBE 16a explicitly asks which confounders were adjusted for and why) — flag a kitchen-sink adjustment set chosen because variables differ in Table 1. The Methods must give a causal rationale (DAG / prior literature) and must not adjust for a mediator or consequence of the outcome (over-adjustment, e.g. serum uric acid in an eGFR model); both an unjustified inclusion and an unjustified omission are item-16a gaps.
PRISMA Cascade Arithmetic Auto-Verify
PRISMA 2020 flow diagrams chain a cascade of subtractions (database
records → after dedup → title/abstract screened → full-text reviewed →
included in synthesis). Off-by-one errors in the prose cascade are a
high-frequency reviewer red flag (e.g., 151 + 108 + 39 + 1 + 1 + 4 = 304 followed by a prose summary "305" four lines later).
When PRISMA 2020 or PRISMA-DTA is selected and round-by-round screening TSV artifacts are available, run the cascade auto-verify:
python "${CLAUDE_SKILL_DIR}/scripts/prisma_cascade_check.py" \
--round1 2_Screening/round1.tsv \
--round2 2_Screening/round2.tsv \
--round3 2_Screening/round3_adjudication.tsv \
--manuscript manuscript.md \
--out qc/prisma_cascade.json
The script:
- Reads the round TSVs and counts
INCLUDE/EXCLUDE/MAYBEdecisions per round. - Computes the cascade arithmetic from raw decisions (no prose).
- Optionally grep the manuscript for matching stage-count claims and emits per-stage drift when the prose disagrees.
Treat any manuscript_drift entry as a P0 blocker — fix the prose to
match the computed cascade and re-run.
Submission Checklist Export
Many journals require a filled reporting checklist to be submitted alongside the manuscript. When the user asks for a submission-ready checklist, format the output as:
{Guideline Name} Checklist
Manuscript title: {title}
Date: {YYYY-MM-DD}
| Item # | Checklist Item | Reported on Page # | Reported in Section |
|--------|---------------|-------------------|-------------------|
| 1 | {item text} | {page or N/A} | {section} |
| 2 | {item text} | {page or N/A} | {section} |
| ... | ... | ... | ... |
Page numbers should be filled in by the user after final formatting. Use section names as placeholders.
Skill Interactions
| When | Call | Purpose |
|---|---|---|
| During manuscript writing | /write-paper Phase 7 |
Final compliance check |
| Need to add Methods text | /write-paper Phase 3 |
Draft missing Methods content |
| Need statistical details | /analyze-stats |
Generate missing statistical reporting |
| Need flow diagram | /make-figures |
Generate CONSORT/STARD/PRISMA diagram |
Error Handling
- If the manuscript file cannot be read, ask the user for the correct path.
- If the study type is ambiguous, ask the user to confirm before selecting a guideline.
- If a checklist item is genuinely unclear in its applicability, mark as N/A with justification.
- This is a pre-screening tool. Always remind the user that final compliance should be verified by all co-authors and ideally by a methodologist.
Language
- Checklist content and compliance report: English
- Communication with user: Match user's preferred language
- Medical terms: English only
Anti-Hallucination
- Never fabricate references. All citations must be verified via
/search-litwith confirmed DOI or PMID. Mark unverified references as[UNVERIFIED - NEEDS MANUAL CHECK]. - Never invent clinical definitions, diagnostic criteria, or guideline recommendations. If uncertain, flag with
[VERIFY]and ask the user.
Gates
| Gate | Severity | Trigger | Action on fail |
|---|---|---|---|
| Mandatory items present | ENFORCED at submission | < 100% of guideline-mandatory items marked PRESENT | Auto-fix MISSING items where text exists; otherwise route to /write-paper Phase 7 for re-draft |
| Step 4d PRISMA Figure 1 arithmetic & cross-reference audit (PRISMA / PRISMA-DTA only) | ENFORCED for SR/MA | flow numbers don't sum (e.g., screened ≠ included + excluded), or in-text counts mismatch flow diagram | HALT; reconcile against extraction artifacts |
| Optional items (e.g., supplementary AI declarations) | ADVISORY | < 80% of optional items present | warn; user accepts |
| Cross-reporting-guideline routing (study type → guideline) | ENFORCED | study type undeclared or guideline missing | Ask user; do not silently default |
Global-rule references
Some passages in this skill cite a path of the form ~/.claude/rules/<name>.md. Those are the
maintainer's personal global rules, kept outside this repository. They are not shipped with
this skill and will not exist on your machine; they appear only as provenance for where a
convention came from. If one of them looks like it is standing in for an instruction you actually
need, that is a bug — please open an issue, because the instruction belongs here.
Files (medsci-skills)
-
references
-
appraisal_tools
-
METRICS.md 4.3 KB
# METRICS — radiomics methodological-quality appraisal **METhodological RadiomICs Score (METRICS)** · EuSoMII-endorsed quality-scoring tool Reference: Kocak B, et al. Insights Imaging 2024;15:8. doi:10.1186/s13244-023-01572-w (CC BY 4.0) Tool / calculator: https://metricsscore.github.io/metrics/METRICS.html > **This is an appraisal (methodological-quality) tool, NOT a reporting guideline.** It answers > "was the radiomics study *done well* (low risk of bias)?", which is distinct from "was it > *reported*?" (that is CLEAR — a reporting checklist). It therefore lives under `appraisal_tools/` > and does **not** count toward the reporting-guideline catalog. Use it alongside, not instead of, > the reporting checklist. Educational summary in our own words from the CC BY 4.0 source; complete > the official weighted calculator for a submission-ready score and cite Kocak et al. 2024. ## Scope and how it scores - Covers radiomics research from **handcrafted** features to **fully deep-learning** pipelines. - **30 items across 9 categories.** Items carry **condition-dependent weights** and some categories are **conditional** (apply only when that step is present — e.g., manual segmentation, explicit feature selection); the official tool produces a weighted **percentage** with quality bands. - Pair with the reporting checklist (CLEAR) and, for prediction-model risk of bias, PROBAST+AI — a fully *reported* radiomics paper can still be at high methodological risk of bias. ## Categories (9) and the concerns each covers | # | Category (items) | Key methodological concerns | |---|---|---| | 1 | Study design (3) | Adherence to radiomics/ML guidance; eligibility describing a representative population; a high-quality reference standard. | | 2 | Imaging data (4) | Multi-centre data; clinical translatability of the setting; imaging-protocol parameters reported; relevant temporal intervals documented. | | 3 | Segmentation (3, *conditional*) | Transparent segmentation methodology; evaluation of automated segmentation; segmentation/masks available for the test set. | | 4 | Image processing & feature extraction (3) | Preprocessing described; standardised/validated extraction software; transparent extraction parameters. | | 5 | Feature processing (4, *conditional*) | Removal of non-robust features; removal of redundant features; dimensionality appropriate to sample size; robustness assessment for deep-learning features. | | 6 | Preparation for modeling (2) | Proper data partitioning (no leakage between train/tune/test); handling of confounding factors. | | 7 | Metrics & comparison (6) | Appropriate performance metrics; uncertainty (CIs); **calibration**; comparison against uni-parametric imaging, against non-radiomic predictors, and against classical/clinical models. | | 8 | Testing (2) | Internal validation **and** external (independent) validation. | | 9 | Open science (3) | Availability of data, code, and the trained model. | ## Non-waivable concerns (surface as a Critical gap if absent) Consistent with the critical-item floor's appraisal note, the highest-yield METRICS concerns are: - **Feature reproducibility / stability** — test–retest, inter-observer/segmentation stability, ICC-based filtering (category 5; segmentation, category 3). - **Internal *and* external validation, with multiplicity control** — a single internal split is not enough for a generalisation claim (categories 6–8). - **Calibration, not discrimination only** — a probability that drives a decision must be calibrated (category 7). - **Leakage-controlled partitioning** — feature selection / preprocessing fit on pooled data, or the same data used for tuning and testing, inflates performance (category 6). ## How to use in the report - This note backs the **Step 4f** appraisal cross-check (`critical_item_floor.md`, METRICS/RQS row). After the reporting item-by-item table, confirm the manuscript's chosen methodological-quality instrument and whether the non-waivable concerns above are met. - Do **not** fold the METRICS score into the reporting compliance percentage — keep appraisal (risk of bias / quality) and reporting (completeness) separate, and do not assert a journal desk-reject threshold from either. - For the exact item wording, weights, and quality bands, use the official EuSoMII calculator. -
METRICS_RELOADED.md 2.3 KB
# Metrics Reloaded — metric-selection appraisal reference An **appraisal / selection** reference (deliberately **not** a counted reporting checklist — the same treatment as `METRICS.md`). It summarises the metric-selection guidance of **Metrics Reloaded** (Maier-Hein, Reinke et al., "Metrics reloaded: recommendations for image analysis validation," *Nature Methods* 2024, CC BY) and its **pitfalls** companion (Reinke et al., *Nature Methods* 2024), for choosing the right validation metric for an image-analysis task. Consumed by `/model-evaluation` (`check_metric_reporting.py`), `/model-validation` (MD6), and the `model_development` probe. > Verify wording against the papers before quoting them as a formal instrument; the points > below are the load-bearing recommendations, phrased for medical imaging. ## Pick the metric from the problem, not habit Metrics Reloaded frames metric choice by the **problem category** (image-level classification, semantic segmentation, instance segmentation, object detection) and the **domain interest** (is boundary accuracy important? are small structures clinically critical? is the data imbalanced?). The metric should reflect what a clinical error costs. ## Common pitfalls it warns against - **Segmentation**: Dice/IoU alone is overlap-only and **insensitive to boundary error** and unstable on **small structures**; pair it with a **boundary metric** (HD95 / NSD) and report **per structure**. Define behaviour for **empty references** (Dice 0/0 is undefined). - **Classification under imbalance**: **accuracy is misleading**; use threshold-independent discrimination (**AUROC**) plus **AUPRC** (minority class), and prevalence-dependent **PPV/NPV at the deployment base rate** — not a balanced set. - **Object detection**: report **FROC / mAP with the IoU match criterion stated**; an unstated match threshold makes the metric undefined. - **Aggregation**: a single global mean hides per-case/per-structure failure; report the distribution and handle missing values explicitly. ## Use in the lane - `/model-evaluation` computes the recommended metric set with CIs and gates the report with `check_metric_reporting.py`. - `/model-validation` MD6 and the `model_development` probe flag a metric-vs-task mismatch. - Reporting compliance of the manuscript stays with CLAIM 2024 / TRIPOD+AI in `/check-reporting`.
-
-
checklists
-
AMSTAR2.md 5.7 KB
# AMSTAR 2 Checklist **A MeaSurement Tool to Assess systematic Reviews, version 2** Version: AMSTAR 2 (2017) Source: Shea BJ et al. BMJ 2017;358:j4008. doi: 10.1136/bmj.j4008 Licence: *BMJ* — Crossref returns no Creative Commons licence for this article. Verification: the 16 item stems, the critical-domain list and the overall-confidence scheme were compared against the official AMSTAR 2 checklist and Boxes 1–2 of the statement (Europe PMC full text, PMC5833365, plus the article's own supplementary appendix `sheb036104.wf1.pdf`). Stems 1–6 and 11–16 matched verbatim; the stems of items 7–10 were not recoverable from the appendix's column layout, so for those four the **response criteria** were compared instead and matched. Critical domains 2, 4, 7, 9, 11, 13 and 15 and the four confidence levels matched Boxes 1 and 2. Two omissions were corrected: the **Partial Yes** response category, which the file did not mention at all, and item 4's "search within 24 months" criterion. ## Response Options Each item is answered **Yes** / **No**, and for items 2, 4, 7, 8 and 9 also **Partial Yes** — a partial answer means the minimum criteria were met but not the fuller set. Items 11, 12 and 15 additionally allow **No meta-analysis**. There is no numeric score. ## Checklist Items (16 items) ### Items | # | Item | Description | Critical? | |---|------|-------------|-----------| | 1 | PICO components | Did the research questions and inclusion criteria for the review include the components of PICO? | No | | 2 | Protocol registered | Did the report of the review contain an explicit statement that the review methods were established prior to the conduct of the review and did the report justify any significant deviations from the protocol? | Yes | | 3 | Study design selection | Did the review authors explain their selection of the study designs for inclusion in the review? | No | | 4 | Comprehensive search | Did the review authors use a comprehensive literature search strategy? *Partial Yes*: searched at least 2 databases relevant to the question, provided key words and/or search strategy, justified publication restrictions. *Yes* also requires: searched reference lists of included studies, searched trial/study registries, included or consulted content experts, searched grey literature where relevant, and conducted the search within 24 months of completing the review. | Yes | | 5 | Duplicate selection | Did the review authors perform study selection in duplicate? | No | | 6 | Duplicate extraction | Did the review authors perform data extraction in duplicate? | No | | 7 | Excluded studies | Did the review authors provide a list of excluded studies and justify the exclusions? | Yes | | 8 | Study descriptions | Did the review authors describe the included studies in adequate detail? (PICO elements, follow-up period, study design, country, setting) | No | | 9 | RoB assessment | Did the review authors use a satisfactory technique for assessing the risk of bias (RoB) in individual studies that were included in the review? (For RCTs: randomization, blinding, missing data, selective reporting. For NRSI: confounding, selection, measurement) | Yes | | 10 | Funding sources | Did the review authors report on the sources of funding for the studies included in the review? | No | | 11 | Statistical methods | If meta-analysis was performed, did the review authors use appropriate methods for statistical combination of results? (effect measures, model choice, heterogeneity assessment) | Yes | | 12 | RoB impact on MA | If meta-analysis was performed, did the review authors assess the potential impact of RoB in individual studies on the results of the meta-analysis or other evidence synthesis? | No | | 13 | RoB in interpretation | Did the review authors account for RoB in individual studies when interpreting/discussing the results of the review? | Yes | | 14 | Heterogeneity | Did the review authors provide a satisfactory explanation for, and discussion of, any heterogeneity observed in the results of the review? | No | | 15 | Publication bias | If they performed quantitative synthesis did the review authors carry out an adequate investigation of publication bias (small study bias) and discuss its likely impact on the results of the review? | Yes | | 16 | Conflicts of interest | Did the review authors report any potential sources of conflict of interest, including any funding they received for conducting the review? | No | --- ## Overall Confidence Rating AMSTAR 2 does NOT generate a numerical score. Instead, rate overall confidence: | Rating | Criteria | |--------|----------| | **High** | No or one non-critical weakness: the systematic review provides an accurate and comprehensive summary of the results | | **Moderate** | More than one non-critical weakness (but no critical flaws): the review provides an accurate summary but may have some weaknesses | | **Low** | One critical flaw with or without non-critical weaknesses: the review may not provide an accurate and comprehensive summary | | **Critically Low** | More than one critical flaw with or without non-critical weaknesses: the review should not be relied on to provide an accurate and comprehensive summary | ## Critical Domains (7 of 16) Items 2, 4, 7, 9, 11, 13, 15 are considered **critical domains**. A flaw in any critical domain results in at least "Low" confidence. ## Notes for Assessors - AMSTAR 2 replaces the original AMSTAR (2007) - Designed for systematic reviews of **interventions** (RCTs and/or NRSI) - Not designed for DTA reviews (use ROBIS or domain-specific tools) - Cannot be used to assess individual primary studies - The tool should NOT be used to generate an overall score — use the confidence rating scheme above - For reviews including NRSI: Item 9 should assess confounding, selection bias, and information bias -
ARRIVE_2.md 10.4 KB
# ARRIVE 2.0 Checklist — Animal Research **Reference:** Percie du Sert N et al. The ARRIVE guidelines 2.0: Updated guidelines for reporting animal research. PLoS Biol. 2020;18(7):e3000410. PMID: 32663219 **Website:** https://arriveguidelines.org --- Version: ARRIVE 2.0 (2020) — Essential 10 + Recommended Set Source: Percie du Sert N, Hurst V, Ahluwalia A, Alam S, Avey MT, Baker M, et al. The ARRIVE guidelines 2.0: updated guidelines for reporting animal research. *PLoS Biol* 2020;18(7):e3000410 (DOI 10.1371/journal.pbio.3000410). Licence: CC0 1.0 (public domain dedication) — confirmed via Crossref. Item text below is reproduced from the published Tables 1 and 2. Verification: all 21 items were extracted from the article's own Table 1 (Essential 10) and Table 2 (Recommended Set) via the Europe PMC full text (PMC7360023) and compared item by item. **This file previously carried an invented item 19 "Limitations", renumbered items 20–21, and omitted item 21 "Declaration of interests" entirely**; the item text below is now the published text. ## How to Use This Checklist ARRIVE 2.0 has two tiers: - **Essential 10** (Items 1–10): the minimum that must be reported for a reader to assess the reliability of the findings. - **Recommended Set** (Items 11–21): add context and completeness. Report these too whenever possible. For each item: **PRESENT** / **PARTIAL** / **MISSING** Where a "How to satisfy it" note appears below, it is **our guidance, not part of the instrument** — the **Description** is the published item text. --- ## ESSENTIAL 10 ### Item 1 — Study design **Description:** For each experiment, provide brief details of study design including: a. The groups being compared, including control groups. If no control group has been used, the rationale should be stated. b. The experimental unit (e.g., a single animal, litter, or cage of animals). *How to satisfy it:* name the design (parallel group, crossover, factorial, dose–response), the groups and their sizes, and state explicitly what the experimental unit was — animal, litter, or cage. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING **Location:** ___ --- ### Item 2 — Sample size **Description:** a. Specify the exact number of experimental units allocated to each group, and the total number in each experiment. Also indicate the total number of animals used. b. Explain how the sample size was decided. Provide details of any a priori sample size calculation, if done. *How to satisfy it:* a formal power calculation states the effect size, α, power, test, and software; a pragmatic justification (animal availability, pilot study) is acceptable but must be stated as such. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING **Location:** ___ --- ### Item 3 — Inclusion and exclusion criteria **Description:** a. Describe any criteria used for including and excluding animals (or experimental units) during the experiment, and data points during the analysis. Specify if these criteria were established a priori. If no criteria were set, state this explicitly. b. For each experimental group, report any animals, experimental units, or data points not included in the analysis and explain why. If there were no exclusions, state so. c. For each analysis, report the exact value of n in each experimental group. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING **Location:** ___ --- ### Item 4 — Randomisation **Description:** a. State whether randomisation was used to allocate experimental units to control and treatment groups. If done, provide the method used to generate the randomisation sequence. b. Describe the strategy used to minimise potential confounders such as the order of treatments and measurements, or animal/cage location. If confounders were not controlled, state this explicitly. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING **Location:** ___ --- ### Item 5 — Blinding **Description:** Describe who was aware of the group allocation at the different stages of the experiment (during the allocation, the conduct of the experiment, the outcome assessment, and the data analysis). *How to satisfy it:* state the position at each of the four stages, including where blinding was impossible and why. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING **Location:** ___ --- ### Item 6 — Outcome measures **Description:** a. Clearly define all outcome measures assessed (e.g., cell death, molecular markers, or behavioural changes). b. For hypothesis-testing studies, specify the primary outcome measure, i.e., the outcome measure that was used to determine the sample size. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING **Location:** ___ --- ### Item 7 — Statistical methods **Description:** a. Provide details of the statistical methods used for each analysis, including software used. b. Describe any methods used to assess whether the data met the assumptions of the statistical approach, and what was done if the assumptions were not met. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING **Location:** ___ --- ### Item 8 — Experimental animals **Description:** a. Provide species-appropriate details of the animals used, including species, strain and substrain, sex, age or developmental stage, and, if relevant, weight. b. Provide further relevant information on the provenance of animals, health/immune status, genetic modification status, genotype, and any previous procedures. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING **Location:** ___ --- ### Item 9 — Experimental procedures **Description:** For each experimental group, including controls, describe the procedures in enough detail to allow others to replicate them, including: a. What was done, how it was done, and what was used. b. When and how often. c. Where (including detail of any acclimatisation periods). d. Why (provide rationale for procedures). *How to satisfy it:* anaesthesia and analgesia agents with dose and route, monitoring, equipment make/model/settings, and anything else that would change the result if varied. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING **Location:** ___ --- ### Item 10 — Results **Description:** For each experiment conducted, including independent replications, report: a. Summary/descriptive statistics for each experimental group, with a measure of variability where applicable (e.g., mean and SD, or median and range). b. If applicable, the effect size with a confidence interval. Adverse events belong to item 16b, not here. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING **Location:** ___ --- ## RECOMMENDED SET ### Item 11 — Abstract **Description:** Provide an accurate summary of the research objectives, animal species, strain and sex, key methods, principal findings, and study conclusions. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### Item 12 — Background **Description:** a. Include sufficient scientific background to understand the rationale and context for the study, and explain the experimental approach. b. Explain how the animal species and model used address the scientific objectives and, where appropriate, the relevance to human biology. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### Item 13 — Objectives **Description:** Clearly describe the research question, research objectives and, where appropriate, specific hypotheses being tested. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### Item 14 — Ethical statement **Description:** Provide the name of the ethical review committee or equivalent that has approved the use of animals in this study, and any relevant licence or protocol numbers (if applicable). If ethical approval was not sought or granted, provide a justification. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### Item 15 — Housing and husbandry **Description:** Provide details of housing and husbandry conditions, including any environmental enrichment. *How to satisfy it:* cage type and dimensions, animals per cage, temperature, humidity, light cycle, food and water access and diet, acclimatisation period, enrichment. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### Item 16 — Animal care and monitoring **Description:** a. Describe any interventions or steps taken in the experimental protocols to reduce pain, suffering, and distress. b. Report any expected or unexpected adverse events. c. Describe the humane endpoints established for the study, the signs that were monitored, and the frequency of monitoring. If the study did not have humane endpoints, state this. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### Item 17 — Interpretation/scientific implications ⚠️ limitations belong here **Description:** a. Interpret the results, taking into account the study objectives and hypotheses, current theory, and other relevant studies in the literature. b. Comment on the study limitations, including potential sources of bias, limitations of the animal model, and imprecision associated with the results. ARRIVE 2.0 has **no standalone "Limitations" item** — limitations are sub-item 17b. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### Item 18 — Generalisability/translation **Description:** Comment on whether, and how, the findings of this study are likely to generalise to other species or experimental conditions, including any relevance to human biology (where appropriate). **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### Item 19 — Protocol registration **Description:** Provide a statement indicating whether a protocol (including the research question, key design features, and analysis plan) was prepared before the study, and if and where this protocol was registered. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### Item 20 — Data access **Description:** Provide a statement describing if and where study data are available. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### Item 21 — Declaration of interests **Description:** a. Declare any potential conflicts of interest, including financial and nonfinancial. If none exist, this should be stated. b. List all funding sources (including grant identifier) and the role of the funder(s) in the design, analysis, and reporting of the study. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING --- ## Summary | Category | PRESENT | PARTIAL | MISSING | |----------|---------|---------|---------| | Essential 10 | /10 | /10 | /10 | | Recommended Set | /11 | /11 | /11 | | **TOTAL** | /21 | /21 | /21 | **Verdict:** [ ] All Essential 10 PRESENT → proceed to submission [ ] Any Essential 10 MISSING → MUST REVISE before submission -
CARE.md 6.4 KB
# CARE Checklist **CAse REports (CARE) guidelines** Version: CARE 2013 Source: https://www.care-statement.org Reference: Gagnier JJ, Kienle G, Altman DG, Moher D, Sox H, Riley D. The CARE guidelines: consensus-based clinical case report guideline development. J Clin Epidemiol 2014;67(1):46-51. Source: Gagnier JJ, Kienle G, Altman DG, Moher D, Sox H, Riley D, et al. The CARE guidelines: consensus-based clinical case report guideline development. *J Clin Epidemiol* 2014;67(1):46-51 (DOI 10.1016/j.jclinepi.2013.08.003). Licence: Crossref returns only an Elsevier text-and-data-mining licence. A CC BY-NC 4.0 claim circulates for the CARE materials; it is **not confirmed** here. Treat as non-open. Verification: all 13 topics and their 23 sub-items were compared against the official CARE 2013 checklist published at care-statement.org, and cross-checked against the checklist table of the CC BY-NC co-publication in *Global Advances in Health and Medicine* (Europe PMC full text, PMC3833570). The topics, their order and their sub-elements match. Wording stays paraphrased — the licence is unconfirmed and treated as non-open. ## Checklist Items (13 topics) ### Title and Key Words | # | Item | Description | |---|------|-------------| | 1 | Title | The words "case report" should appear in the title along with the diagnosis or intervention of primary focus. | | 2 | Key Words | Two to five key words that identify topics in this case report, including "case report". | ### Abstract | # | Item | Description | |---|------|-------------| | 3a | Abstract — Introduction | What is unique about this case and what does it add to the scientific literature? | | 3b | Abstract — Main concerns | The patient's main concerns and important clinical findings. | | 3c | Abstract — Diagnoses, interventions, outcomes | The primary diagnoses, interventions, and outcomes. | | 3d | Abstract — Conclusion | What are one or more "take-away" lessons from this case report? | ### Introduction | # | Item | Description | |---|------|-------------| | 4 | Introduction | One or two paragraphs summarizing why this case is unique (may include references). | ### Patient Information | # | Item | Description | |---|------|-------------| | 5a | Patient Information | De-identified demographic and other patient information. | | 5b | Patient Information | Main concerns and symptoms of the patient. | | 5c | Patient Information | Medical, family, and psychosocial history including relevant genetic information. | | 5d | Patient Information | Relevant past interventions and their outcomes. | ### Clinical Findings | # | Item | Description | |---|------|-------------| | 6 | Clinical Findings | Describe the significant physical examination and other clinical findings. | ### Timeline | # | Item | Description | |---|------|-------------| | 7 | Timeline | Historical and current information from this episode of care organized as a timeline (figure or table). | ### Diagnostic Assessment | # | Item | Description | |---|------|-------------| | 8a | Diagnostic Assessment | Diagnostic methods (e.g., physical examination, laboratory testing, imaging, questionnaires). | | 8b | Diagnostic Assessment | Diagnostic challenges (e.g., access to testing, financial, or cultural). | | 8c | Diagnostic Assessment | The diagnosis, including other diagnoses that were considered. | | 8d | Diagnostic Assessment | Prognostic characteristics (e.g., staging) where applicable. | ### Therapeutic Intervention | # | Item | Description | |---|------|-------------| | 9a | Therapeutic Intervention | Types of therapeutic intervention (e.g., pharmacologic, surgical, preventive, self-care). | | 9b | Therapeutic Intervention | Administration of therapeutic intervention (e.g., dosage, strength, duration). | | 9c | Therapeutic Intervention | Changes in therapeutic intervention (with rationale). | ### Follow-up and Outcomes | # | Item | Description | |---|------|-------------| | 10a | Follow-up and Outcomes | Clinician- and patient-assessed outcomes (when appropriate). | | 10b | Follow-up and Outcomes | Important follow-up diagnostic and other test results. | | 10c | Follow-up and Outcomes | Intervention adherence and tolerability (and how this was assessed). | | 10d | Follow-up and Outcomes | Adverse and unanticipated events. | ### Discussion | # | Item | Description | |---|------|-------------| | 11a | Discussion | A scientific discussion of the strengths and limitations associated with this case report. | | 11b | Discussion | Discussion of the relevant medical literature with references. | | 11c | Discussion | The scientific rationale for any conclusions (including assessment of possible causes). | | 11d | Discussion | The primary "take-away" lessons of this case report (without references) in a one-paragraph conclusion. | ### Patient Perspective | # | Item | Description | |---|------|-------------| | 12 | Patient Perspective | The patient should share their perspective on the treatment(s) they received, when appropriate. | ### Informed Consent | # | Item | Description | |---|------|-------------| | 13 | Informed Consent | Did the patient give informed consent? Provide if requested. | --- ## Applying CARE to common case-report subtypes CARE covers the single-patient narrative. Two frequent subtypes need additions CARE alone does not name: - **Adverse drug / device / contrast reaction (pharmacovigilance).** Beyond item 8 (diagnostic assessment), report a **named causality instrument** (Naranjo or WHO-UMC) with score and tier, document **dechallenge** (withdrawal → resolution) and the exposure-to-onset latency, and consider **severity** (e.g., Modified Hartwig–Siegel) and **preventability** (e.g., Schumock–Thornton). Locate the event against a denominator and note safety-reporting. Causality asserted without an instrument, dechallenge, or exclusion of alternatives is a reporting gap. - **Case series (n≥2).** CARE is single-patient; a series additionally needs **cohort-style methods** (design, setting, case-identification source, eligibility, protocol) and an **all-cases summary table**. Report counts, not rates — a selected/referral series cannot estimate prevalence — and state selection/ascertainment with the screened pool size. See `/write-paper` `paper_types/case_series.md`. *Educational summary of the CARE 2013 checklist (CC BY-NC 4.0). Cite the original guideline (Gagnier et al. 2014) and consult https://www.care-statement.org for the authoritative, full checklist with explanation and elaboration.* -
CHEERS_2022.md 8.3 KB
# CHEERS 2022 Checklist **Consolidated Health Economic Evaluation Reporting Standards 2022** Version: CHEERS 2022 (28 items; replaces CHEERS 2013). Source: Husereau D, Drummond M, Augustovski F, et al. *BMJ* 2022;376:e067975 (the CHEERS 2022 statement), co-published simultaneously across BMJ, *Value in Health*, *PharmacoEconomics*, *Int J Technol Assess Health Care* and others. CC BY 4.0. https://www.equator-network.org/reporting-guidelines/cheers/ · ISPOR CHEERS Task Force. Verification: all 28 items were compared against Table 1 of the published statement (Europe PMC full text, PMC8749494); 28/28 match, with no item missing and none invented. Apply when the manuscript is a **health economic evaluation** — a comparative analysis of costs and consequences of two or more courses of action: cost-effectiveness (CEA), cost-utility (CUA), cost-benefit (CBA), or cost-minimisation analysis, whether trial-based or decision-model-based (decision tree, Markov/state-transition, discrete-event simulation), including budget-impact and HTA submissions. For the design/validity review of the same study, pair with the HE1–HE8 domain probes in `peer-review` / `self-review` `references/domain-probes/health_economic_evaluation.md`; for the analysis, with `analyze-stats` `references/analysis_guides/health_economic_evaluation.md`. Source: Husereau D, Drummond M, Augustovski F, de Bekker-Grob E, Briggs AH, Carswell C, et al. Consolidated Health Economic Evaluation Reporting Standards 2022 (CHEERS 2022) statement. *BMJ* 2022;376:e067975 (DOI 10.1136/bmj-2021-067975). ## Checklist Items (28 items) ### Title | # | Item | Description | |---|------|-------------| | 1 | Title | Identify the study as an economic evaluation and specify the interventions being compared. | ### Abstract | # | Item | Description | |---|------|-------------| | 2 | Abstract | Provide a structured summary that highlights context, key methods, results, and alternative analyses. | ### Introduction | # | Item | Description | |---|------|-------------| | 3 | Background and objectives | Give the context for the study, the study question, and its practical relevance for decision making in policy or practice. | ### Methods | # | Item | Description | |---|------|-------------| | 4 | Health economic analysis plan | Indicate whether a health economic analysis plan was developed and where available. | | 5 | Study population | Describe characteristics of the study population (such as age range, demographics, socioeconomic, or clinical characteristics). | | 6 | Setting and location | Provide relevant contextual information that may influence findings. | | 7 | Comparators | Describe the interventions or strategies being compared and why chosen. | | 8 | Perspective | State the perspective(s) adopted by the study and why chosen. | | 9 | Time horizon | State the time horizon for the study and why appropriate. | | 10 | Discount rate | Report the discount rate(s) and reason chosen. | | 11 | Selection of outcomes | Describe what outcomes were used as the measure(s) of benefit(s) and harm(s). | | 12 | Measurement of outcomes | Describe how outcomes used to capture benefit(s) and harm(s) were measured. | | 13 | Valuation of outcomes | Describe the population and methods used to measure and value outcomes. | | 14 | Measurement and valuation of resources and costs | Describe how costs were valued. | | 15 | Currency, price date, and conversion | Report the dates of the estimated resource quantities and unit costs, plus the currency and year of conversion. | | 16 | Rationale and description of model | If modelling is used, describe in detail and why used. Report whether the model is publicly available and where. | | 17 | Analytics and assumptions | Describe any methods for analysing or statistically transforming data, any extrapolation methods, and approaches for validating any model used. | | 18 | Characterising heterogeneity | Describe any methods used for estimating how the results of the study vary for subgroups. | | 19 | Characterising distributional effects | Describe how impacts are distributed across different individuals or whether adjustments were made to reflect priority populations. | | 20 | Characterising uncertainty | Describe methods to characterise any sources of uncertainty in the analysis. | | 21 | Approach to engagement with patients and others affected by the study | Describe any approaches to engage patients or service recipients, the general public, communities, or stakeholders (such as clinicians or payers) in the design of the study. | ### Results | # | Item | Description | |---|------|-------------| | 22 | Study parameters | Report all analytic inputs (such as values, ranges, references) including uncertainty or distributional assumptions. | | 23 | Summary of main results | Report the mean values for the main categories of costs and outcomes of interest and summarise them in the most appropriate overall measure (e.g. the incremental cost-effectiveness ratio, ICER). | | 24 | Effect of uncertainty | Describe how uncertainty about analytic judgments, inputs, or projections affect findings. Report the effect of choice of discount rate and time horizon, if relevant. | | 25 | Effect of engagement with patients and others affected by the study | Report on any difference patient/service recipient, general public, community, or stakeholder involvement made to the approach or findings of the study. | ### Discussion | # | Item | Description | |---|------|-------------| | 26 | Study findings, limitations, generalisability, and current knowledge | Report key findings, limitations, ethical or equity considerations, and how these could affect patients, policy, or practice. | ### Other relevant information | # | Item | Description | |---|------|-------------| | 27 | Source of funding | Describe how the study was funded and the role of the funder in the identification, design, conduct, and reporting of the analysis. Describe other non-monetary sources of support. | | 28 | Conflicts of interest | Describe any potential for conflict of interest among study contributors in accordance with journal policy. In the absence of a journal policy, we recommend authors comply with International Committee of Medical Journal Editors (ICMJE) recommendations. | --- ## Notes for Assessors - **Highest-yield items** (where economic evaluations most often fail review): **8** (perspective stated and consistent with the costs counted — productivity/informal-care costs belong only to a societal perspective), **9** (time horizon long enough to capture all relevant differential costs and effects — a lifetime horizon for a chronic condition; a truncated horizon flatters whichever arm has early benefit), **10** (both costs *and* outcomes discounted at a stated, justified rate for any horizon beyond ~1 year), **15** (currency *and* price year stated, with the conversion method for multi-source costs), **16–17** (model type/structure justified and validated; structural assumptions and extrapolation declared), and **20 / 24** (uncertainty characterised by *probabilistic* sensitivity analysis — a cost-effectiveness acceptability curve / plane — not a single deterministic ICER). An evaluation that reports a point-estimate ICER with no probabilistic sensitivity analysis is non-compliant on items 20/24. - **CHEERS 2022 replaced CHEERS 2013**; cite the 2022 statement (do not cite the 2013 version as current). CHEERS 2022 added explicit items on the analysis plan (4), distributional/equity effects (19), and patient/stakeholder engagement (21/25); the engagement items are reported as "not done" rather than omitted when no engagement occurred. - ICER interpretation is not a CHEERS item per se but follows from items 23–24: incremental costs and effects must be reported (not just the ratio), dominance/extended dominance resolved, and the ICER interpreted against a *stated, justified* cost-effectiveness threshold (willingness-to-pay) rather than an arbitrary one. The HE1–HE8 design probes cover these judgments. - This checklist was authored as a faithful summary of the CHEERS 2022 statement (Husereau D, et al. *BMJ* 2022;376:e067975, **CC BY 4.0** — the item list and Explanation & Elaboration are reusable/adaptable with attribution) for item-by-item assessment; verify against the published statement and its Explanation & Elaboration for full item wording. Verified 2026-06-29. -
CLAIM_2024.md 7.1 KB
# CLAIM 2024 Checklist **Checklist for Artificial Intelligence in Medical Imaging** Version: CLAIM 2024 Update Source: https://pubs.rsna.org/doi/10.1148/ryai.240300 Reference: Tejani AS, Klontzas ME, Gatti AA, Mongan JT, Moy L, Park SH, Kahn CE Jr. Checklist for Artificial Intelligence in Medical Imaging (CLAIM): 2024 Update. Radiol Artif Intell 2024;6(4):e240300. > Note: The 2024 update replaces "ground truth" with "reference standard" and discourages "validation" in favour of "internal/external testing". Each item is answered Yes / No / Not Applicable with the manuscript location cited. Licence: © RSNA, open access. Consult RSNA for reuse terms; Crossref returns no Creative Commons licence. Verification: all 44 items were compared, by number and content, against the item-by-item text of the published update (PubMed Central record PMC11304031, whose body carries each item in full). 44/44 match, with no item missing and none invented. Item labels below are our own short names; item 21's official name is "Intended sample size" and it does concern the testing set, as our description says. Wording stays paraphrased — Crossref returns no Creative Commons licence. ## Checklist Items (44 items) ### Title and Abstract | # | Item | Description | |---|------|-------------| | 1 | Title | Identify the study as employing AI methodology and name the specific technology category (e.g., deep learning). | | 2 | Abstract | Structured summary including study design, methods, results, and conclusions; population details, data partitions, prospective/retrospective status, statistical analysis, outcomes, and availability of resources. | ### Introduction | # | Item | Description | |---|------|-------------| | 3 | Background | Scientific and clinical background, current practice, intended use, and clinical role of the AI approach. | | 4 | Objectives | Study aims, objectives, and hypotheses (if not data-driven). | ### Methods — Study Design | # | Item | Description | |---|------|-------------| | 5 | Study design | Prospective or retrospective study. | | 6 | Study goal | Goal of the study (e.g., model creation, feasibility, trial type, intended use for the classification task). | ### Methods — Data | # | Item | Description | |---|------|-------------| | 7 | Data sources | State data sources; provide links to publicly available datasets. | | 8 | Eligibility | Inclusion and exclusion criteria and selection methodology for the data. | | 9 | Preprocessing | Data preprocessing steps (e.g., normalization, resampling, window/level adjustment). | | 10 | Subset selection | Selection of data subsets and training of personnel involved. | | 11 | De-identification | De-identification methods meeting HIPAA/GDPR/AI Act standards. | | 12 | Missing data | How missing data were handled and potential biases from imputation. | | 13 | Acquisition protocol | Image acquisition protocol parameters (e.g., manufacturer, sequences, resolution). | ### Methods — Reference Standard | # | Item | Description | |---|------|-------------| | 14 | Reference standard definition | Method for obtaining the reference standard, with precise and replicable definitions. | | 15 | Reference standard rationale | Rationale for choosing the reference standard versus alternatives. | | 16 | Annotators | Source, qualifications, and training materials of annotators. | | 17 | Annotation procedures | Test-set annotation procedures, software version, and any NLP/automated approaches. | | 18 | Annotation variability | Measurement of inter- and intra-rater variability and method of discrepancy resolution. | ### Methods — Data Partitions | # | Item | Description | |---|------|-------------| | 19 | Partition assignment | Partition assignment (train/tune/test), proportions, justification, and class-imbalance handling. | | 20 | Partition disjointness | Level of partition disjointness (patient-, series-, or image-level). | ### Methods — Testing Data | # | Item | Description | |---|------|-------------| | 21 | Test set size | Testing-set size derived from a power calculation or AUC-based estimation. | ### Methods — Model | # | Item | Description | |---|------|-------------| | 22 | Model architecture | Complete model architecture (inputs, outputs, layers, pooling, normalization). | | 23 | Software | Software libraries, frameworks, packages, and version numbers. | | 24 | Initialization | Parameter initialization; transfer-learning sources, if used. | ### Methods — Training | # | Item | Description | |---|------|-------------| | 25 | Training procedures | Training procedures, data augmentation, convergence monitoring, and all hyperparameters. | | 26 | Model selection | Method and metrics for selecting the best-performing model. | | 27 | Ensembling | If an ensemble approach is used, details of each model and how outputs are combined. | ### Methods — Evaluation | # | Item | Description | |---|------|-------------| | 28 | Performance metrics | Performance metrics and comparison to published models. | | 29 | Uncertainty | Measures of uncertainty (e.g., standard deviation, confidence intervals) and statistical significance tests. | | 30 | Robustness | Robustness or sensitivity analysis. | | 31 | Explainability | If applied, explainability/interpretability methods and their validation. | | 32 | Internal testing | Internal-data evaluation and consistency between training and test performance. | | 33 | External testing | External-data testing, or justification for its omission. | | 34 | Trial registration | If applicable, compliance with ICMJE clinical-trial registration requirements. | ### Results — Data | # | Item | Description | |---|------|-------------| | 35 | Inclusion/exclusion numbers | Numbers of patients/examinations included and excluded, with a flowchart. | | 36 | Demographics | Demographic and clinical characteristics per partition; identify potential sources of bias. | ### Results — Model Performance | # | Item | Description | |---|------|-------------| | 37 | Performance reporting | Final model performance benchmarked against the reference standard across partitions and subgroups. | | 38 | Accuracy estimates | Diagnostic accuracy estimates with 95% confidence intervals; ROC analysis; address class imbalance. | | 39 | Failure analysis | Failure analysis with a confusion matrix; examples of incorrect classifications in the medical context. | ### Discussion | # | Item | Description | |---|------|-------------| | 40 | Limitations | Study limitations (methods, materials, biases, generalisability). | | 41 | Implications | Clinical implications, intended use, practice changes, and barriers to translation. | ### Other Information | # | Item | Description | |---|------|-------------| | 42 | Full protocol | Reference to the full protocol or technical details if exceeding journal word limits. | | 43 | Availability | Availability of software, model, and data, and access conditions. | | 44 | Funding | Funding sources and the role of funders. | --- *Educational summary of the CLAIM 2024 Update checklist (© RSNA, open access). Cite the original (Tejani et al., Radiol Artif Intell 2024;6(4):e240300) and consult the RSNA article for the authoritative, full checklist.* -
CLEAR.md 7.8 KB
# CLEAR Checklist **CheckList for EvaluAtion of Radiomics research** - **Version:** CLEAR 2023 - **Citation:** Kocak B, Baessler B, Bakas S, et al. *CheckList for EvaluAtion of Radiomics research (CLEAR): a step-by-step reporting guideline for authors and reviewers endorsed by ESR and EuSoMII.* Insights Imaging. 2023;14(1):75. - **DOI:** 10.1186/s13244-023-01415-8 - **Source:** https://pmc.ncbi.nlm.nih.gov/articles/PMC10160267/ · official item list: https://clearchecklist.github.io/clear_checklist/CLEAR.html - **Licence:** CC BY 4.0. Item wording below is reproduced faithfully from the published statement with attribution. CLEAR is a **58-item** step-by-step reporting guideline for radiomics research, ordered by **manuscript section** to follow a paper from title to open science: Title (1), Abstract (2), Keywords (3), Introduction (4–6), Methods (7–43), Results (44–48), Discussion (49–52), and Open Science (53–58). The Methods block is subdivided into Study design (7–12), Data (13–18), Segmentation (19–20), Pre-processing (21–24), Feature extraction (25–28), Data preparation (29–33), Modeling (34–37), and Evaluation (38–43). Two items — **53** and **58** — are marked **[n/e]** ("not essential" but recommended). All other items are essential; score a missing essential item as a reporting gap, and a missing [n/e] item as N/A only when justified. CLEAR is written for **hand-crafted radiomics**; for deep-learning pipelines without radiomic features, CLAIM 2024 or TRIPOD+AI may fit better (a study that does both should be assessed against both). Verification: all 58 items were compared against Table 1 of the published checklist (Europe PMC full text, PMC10160267); 58/58 match, with no item missing and none invented. ## Checklist Items ### Title | Item | Checklist item | |---|---| | 1 | Relevant title, specifying the radiomic methodology (generally identifying the study as radiomics-related). | ### Abstract | Item | Checklist item | |---|---| | 2 | Structured summary with relevant information (with a structured or unstructured summary presenting key information). | ### Keywords | Item | Checklist item | |---|---| | 3 | Relevant keywords for radiomics (providing keywords most relevant to the topic). | ### Introduction | Item | Checklist item | |---|---| | 4 | Scientific or clinical background (mentioning the current scientific or clinical background). | | 5 | Rationale for using a radiomic approach (explaining the rationale for using a radiomic approach). | | 6 | Study objective(s) (stating the study objectives, hypotheses, or aims). | ### Methods — Study design | Item | Checklist item | |---|---| | 7 | Adherence to guidelines or checklists (e.g., CLEAR checklist). | | 8 | Ethical details (e.g., approval, consent, data protection). | | 9 | Sample size calculation (with a statistical power analysis, if performed). | | 10 | Study nature (e.g., retrospective, prospective). | | 11 | Eligibility criteria (with inclusion and exclusion criteria). | | 12 | Flowchart for technical pipeline (presenting a technical pipeline flowchart). | ### Methods — Data | Item | Checklist item | |---|---| | 13 | Data source (e.g., private, public). | | 14 | Data overlap (declaring any data overlap with previous studies). | | 15 | Data split methodology (describing how the data were split, e.g., training/validation/test). | | 16 | Imaging protocol (i.e., image acquisition and processing). | | 17 | Definition of non-radiomic predictor variables. | | 18 | Definition of the reference standard (i.e., outcome variable). | ### Methods — Segmentation | Item | Checklist item | |---|---| | 19 | Segmentation strategy (2D/3D, manual/automatic, software, region of interest). | | 20 | Details of operators performing segmentation (number, experience, qualifications). | ### Methods — Pre-processing | Item | Checklist item | |---|---| | 21 | Image pre-processing details. | | 22 | Resampling method and its parameters. | | 23 | Discretization method and its parameters (e.g., fixed bin width or count). | | 24 | Image types (e.g., original, filtered, transformed). | ### Methods — Feature extraction | Item | Checklist item | |---|---| | 25 | Feature extraction method (software and version). | | 26 | Feature classes (e.g., shape, first-order, texture). | | 27 | Number of features (extracted per region and in total). | | 28 | Default configuration statement for remaining parameters. | ### Methods — Data preparation | Item | Checklist item | |---|---| | 29 | Handling of missing data. | | 30 | Details of class imbalance. | | 31 | Details of segmentation reliability analysis (e.g., inter-/intra-observer agreement). | | 32 | Feature scaling details (e.g., normalization, standardization). | | 33 | Dimension reduction details (e.g., feature selection). | ### Methods — Modeling | Item | Checklist item | |---|---| | 34 | Algorithm details (name and characteristics of the modeling algorithm[s]). | | 35 | Training and tuning details (including hyperparameter optimization). | | 36 | Handling of confounders. | | 37 | Model selection strategy. | ### Methods — Evaluation | Item | Checklist item | |---|---| | 38 | Testing technique (e.g., internal, external). | | 39 | Performance metrics and rationale for choosing. | | 40 | Uncertainty evaluation and measures (e.g., confidence intervals). | | 41 | Statistical performance comparison (e.g., DeLong's test). | | 42 | Comparison with non-radiomic and combined methods. | | 43 | Interpretability and explainability methods. | ### Results | Item | Checklist item | |---|---| | 44 | Baseline demographic and clinical characteristics (across data partitions). | | 45 | Flowchart for eligibility criteria (participant flow). | | 46 | Feature statistics (e.g., reproducibility, feature selection). | | 47 | Model performance evaluation (with the pre-specified metrics). | | 48 | Comparison with non-radiomic and combined approaches. | ### Discussion | Item | Checklist item | |---|---| | 49 | Overview of important findings. | | 50 | Previous works with differences from the current study. | | 51 | Practical implications. | | 52 | Strengths and limitations (e.g., bias and generalizability issues). | ### Open Science — Data availability | Item | Checklist item | |---|---| | 53 | Sharing images along with segmentation data **[n/e]**. | | 54 | Sharing radiomic feature data. | ### Open Science — Code availability | Item | Checklist item | |---|---| | 55 | Sharing pre-processing scripts or settings. | | 56 | Sharing source code for modeling. | ### Open Science — Model availability | Item | Checklist item | |---|---| | 57 | Sharing final model files. | | 58 | Sharing a ready-to-use system **[n/e]**. | --- ## Notes for assessors - **Order is by manuscript section, not by topic.** CLEAR numbers items in the order they appear in a paper (Title → Abstract → … → Open Science). When you cite a CLEAR item, cite it by this official number — item 1 is the title, item 44 is baseline demographics in the Results, item 58 is sharing a ready-to-use system. - **Items 53 and 58 are the only non-essential ([n/e]) items** — recommended best practice for open science. Every other item is essential; a missing essential item is a reporting gap, not an optional extra. - **Segmentation items (19–20)** may be N/A for studies using fully automated, atlas-based segmentation with no reader involvement — note this rather than scoring MISSING. - **Open Science (53–58)** is where a radiomics study's reproducibility is assessed and is increasingly required by journals; missing items here commonly draw reviewer comments. - CLEAR was endorsed by the European Society of Radiology (ESR) and the European Society of Medical Imaging Informatics (EuSoMII). For methodological quality (as opposed to reporting completeness), pair CLEAR with METRICS (Kocak et al. 2024). -
CONSORT.md 7.1 KB
# CONSORT 2025 Checklist **Consolidated Standards of Reporting Trials** Version: CONSORT 2025 Source: https://www.consort-spirit.org Reference: Hopewell S, Chan AW, Collins GS, et al. CONSORT 2025 statement: updated guideline for reporting randomised trials. BMJ 2025;389:e081123 (published simultaneously in BMJ, JAMA, Lancet, Nature Medicine, PLoS Medicine). > Note: CONSORT 2025 supersedes CONSORT 2010. It is a 30-item checklist (seven new items, three revised, one deleted) restructured with a new Open Science section. Source: Hopewell S, Chan AW, Collins GS, Hróbjartsson A, Moher D, Schulz KF, et al. CONSORT 2025 statement: updated guideline for reporting randomised trials. *BMJ* 2025;389:e081123 (DOI 10.1136/bmj-2024-081123). Licence: CC BY 4.0 — confirmed via Crossref. Verification: all 42 sub-items were compared against Table 1 of the published statement (Europe PMC full text, PMC11995449); 42/42 match. Item 13's label and its second sentence, which were missing, have been restored. ## Checklist Items (30 items) ### Title and Abstract | # | Item | Description | |---|------|-------------| | 1a | Title | Identification as a randomised trial. | | 1b | Abstract | Structured summary of the trial design, methods, results, and conclusions. | ### Open Science | # | Item | Description | |---|------|-------------| | 2 | Trial registration | Name of trial registry, identifying number (with URL) and date of registration. | | 3 | Protocol and SAP | Where the trial protocol and statistical analysis plan can be accessed. | | 4 | Data, code, materials | Where and how the individual de-identified participant data (including data dictionary), statistical code and any other materials can be accessed. | | 5a | Funding | Sources of funding and other support (e.g., supply of drugs), and role of funders in the design, conduct, analysis and reporting of the trial. | | 5b | Conflicts of interest | Financial and other conflicts of interest of the manuscript authors. | ### Introduction | # | Item | Description | |---|------|-------------| | 6 | Background | Scientific background and rationale. | | 7 | Objectives | Specific objectives related to benefits and harms. | ### Methods | # | Item | Description | |---|------|-------------| | 8 | Patient and public involvement | Details of patient or public involvement in the design, conduct and reporting of the trial. | | 9 | Trial design | Description of trial design including type of trial (e.g., parallel group, crossover), allocation ratio, and framework. | | 10 | Changes to trial | Important changes to the trial after it commenced including any outcomes or analyses that were not prespecified, with reason. | | 11 | Settings and locations | Settings (e.g., community, hospital) and locations (e.g., countries, sites) where the trial was conducted. | | 12a | Eligibility — participants | Eligibility criteria for participants. | | 12b | Eligibility — sites/deliverers | If applicable, eligibility criteria for sites and for individuals delivering the interventions. | | 13 | Intervention and comparator | Intervention and comparator with sufficient details to allow replication. If relevant, where additional materials describing the intervention and comparator (e.g., intervention manual) can be accessed. | | 14 | Outcomes | Prespecified primary and secondary outcomes, including the specific measurement variable, analysis metric, method of aggregation, and time point for each outcome. | | 15 | Harms | How harms were defined and assessed (e.g., systematically, non-systematically). | | 16a | Sample size | How sample size was determined, including all assumptions supporting the sample size calculation. | | 16b | Interim analyses | Explanation of any interim analyses and stopping guidelines. | | 17a | Randomisation — sequence | Who generated the random allocation sequence and the method used. | | 17b | Randomisation — restriction | Type of randomisation and details of any restriction (e.g., stratification, blocking and block size). | | 18 | Allocation concealment | Mechanism used to implement the random allocation sequence (e.g., central computer/telephone; sequentially numbered, opaque, sealed containers). | | 19 | Implementation | Whether the personnel who enrolled and those who assigned participants to the interventions had access to the random allocation sequence. | | 20a | Blinding — who | Who was blinded after assignment to interventions (e.g., participants, care providers, outcome assessors, data analysts). | | 20b | Blinding — how | If blinded, how blinding was achieved and description of the similarity of interventions. | | 21a | Statistical methods | Statistical methods used to compare groups for primary and secondary outcomes, including harms. | | 21b | Analysis populations | Definition of who is included in each analysis (e.g., all randomised participants), and in which group. | | 21c | Missing data | How missing data were handled in the analysis. | | 21d | Additional analyses | Methods for any additional analyses (e.g., subgroup and sensitivity analyses), distinguishing prespecified from post hoc. | ### Results | # | Item | Description | |---|------|-------------| | 22a | Participant flow — numbers | For each group, the numbers of participants who were randomly assigned, received intended intervention, and were analysed for the primary outcome. | | 22b | Participant flow — losses | For each group, losses and exclusions after randomisation, together with reasons. | | 23a | Recruitment — dates | Dates defining the periods of recruitment and follow-up for outcomes of benefits and harms. | | 23b | Recruitment — stopping | If relevant, why the trial ended or was stopped. | | 24a | Intervention as administered | Intervention and comparator as they were actually administered (e.g., where appropriate, who delivered the intervention/comparator, how participants adhered, whether they were delivered as intended). | | 24b | Concomitant care | Concomitant care received during the trial for each group. | | 25 | Baseline data | A table showing baseline demographic and clinical characteristics for each group. | | 26 | Outcomes and estimation | For each primary and secondary outcome, by group: the number of participants included in the analysis, the number with available data at the outcome time point, result for each group, and the estimated effect size and its precision. | | 27 | Harms | All harms or unintended events in each group. | | 28 | Ancillary analyses | Any other analyses performed, including subgroup and sensitivity analyses, distinguishing pre-specified from post hoc. | ### Discussion | # | Item | Description | |---|------|-------------| | 29 | Interpretation | Interpretation consistent with results, balancing benefits and harms, and considering other relevant evidence. | | 30 | Limitations | Trial limitations, addressing sources of potential bias, imprecision, generalisability, and, if relevant, multiplicity of analyses. | --- *Educational summary of the CONSORT 2025 checklist (CC BY 4.0). Cite the original statement (Hopewell et al., BMJ 2025) and consult https://www.consort-spirit.org for the authoritative, full checklist with explanation and elaboration.* -
CONSORT_AI.md 5.1 KB
# CONSORT-AI Checklist **Consolidated Standards of Reporting Trials -- Artificial Intelligence Extension** Version: CONSORT-AI 2020 (extends CONSORT 2010) Source: https://www.consort-spirit.org · EQUATOR Network Reference: Liu X, Cruz Rivera S, Moher D, et al. Nat Med 2020;26(9):1364-1374. doi:10.1038/s41591-020-1034-x (CC BY 4.0) > Educational summary, authored in our own words from the CC BY 4.0 source. Use the official > CONSORT-AI checklist for a submission-ready form and cite Liu et al. 2020. > **Verified against the published statement.** All 14 AI items were enumerated from the article's > own checklist table via the PMC XML and compared label-by-label: **11 Extensions and > 3 Elaborations**, matching exactly with no item missing and none invented. > **Extension vs Elaboration — the statement distinguishes them, and it matters.** An **Extension** > is a *new* reporting requirement that CONSORT 2010 does not contain. An **Elaboration** clarifies how an > existing CONSORT 2010 item applies when the intervention involves AI; the requirement already existed, > the guidance is what is new. Both are assessed, but only the Extensions are additional obligations — > do not report an Elaboration as though the base instrument had been silent on it. Elaborations are > marked **(E)** below. ## Naming and scope (read first) - CONSORT-AI is an **extension** of **CONSORT 2010**, for reports of randomized clinical trials of interventions that **include an AI/ML component**. Apply **both**: every base CONSORT 2010 item plus the AI-specific items below; name and cite both instruments (manuscript-style-classical §14). - It is the **reports** counterpart of **SPIRIT-AI** (trial protocols). For a trial protocol use SPIRIT-AI; for the completed-trial report use CONSORT-AI. - The AI items elaborate existing CONSORT items (numbered to match), so assess them alongside the parent item. ## AI-specific extension items Status each PRESENT / PARTIAL / MISSING / N/A. ### Title and Abstract | # | Item | Description (intent) | |---|------|----------------------| | 1a,b (i) **(E)** | AI identification | State in the title/abstract that the intervention involves AI/ML and specify the type of model. | | 1a,b (ii) **(E)** | Intended use | State the intended use of the AI intervention in the title/abstract. | ### Introduction — Background and objectives | # | Item | Description (intent) | |---|------|----------------------| | 2a (i) | Intended use in context | Explain the intended use of the AI intervention in the context of the clinical pathway, including its purpose and the intended user. | ### Methods | # | Item | Description (intent) | |---|------|----------------------| | 4a (i) **(E)** | Participant eligibility | State the participant-level inclusion and exclusion criteria. | | 4a (ii) | Input-data eligibility | State the inclusion and exclusion criteria at the level of the **input data** to the AI system. | | 4b | Setting integration | Describe how the AI intervention was integrated into the trial setting, including any onsite or offsite requirements. | | 5 (i) | Algorithm version | State which version of the AI algorithm was used. | | 5 (ii) | Input acquisition | Describe how the input data were acquired and selected for the AI intervention. | | 5 (iii) | Poor/unavailable input | Describe how poor-quality or unavailable input data were assessed and handled. | | 5 (iv) | Human–AI interaction | Specify whether there is human–AI interaction in the handling of the input data, and the expertise required of the user. | | 5 (v) | AI output | Specify the output of the AI intervention. | | 5 (vi) | Output to decision | Explain how the AI intervention's outputs contributed to decision-making or other elements of clinical practice. | ### Results | # | Item | Description (intent) | |---|------|----------------------| | 19 | Performance errors | Describe the results of any analysis of performance errors and how errors were identified; if none was done, justify why. | ### Other Information | # | Item | Description (intent) | |---|------|----------------------| | 25 | Code/intervention access | State whether and how the AI intervention and/or its code can be accessed, including any restrictions on access or reuse. | --- ## Notes for Assessors - Apply CONSORT-AI **with** all base CONSORT 2010 items — the AI items do not replace them. - **Algorithm version (5 (i))** is non-waivable: a trial result tied to an unspecified model version is not interpretable or reproducible. Mark MISSING if absent. - **Input-data eligibility (4a (ii))** is distinct from participant eligibility — a frequent omission; a trial can enroll eligible patients yet feed the AI out-of-distribution inputs. - **Human–AI interaction (5 (iv))** and **how outputs fed decisions (5 (vi))** determine whether the trial evaluated the AI as used in practice; vague "the model assisted clinicians" is PARTIAL. - **Performance-error analysis (19)** and **code accessibility (25)** are commonly dropped and are where AI-specific safety/reproducibility live. - Use **SPIRIT-AI** for the trial protocol; CONSORT-AI is for the completed-trial report. -
COREQ.md 7.2 KB
# COREQ Checklist (qualitative — interviews & focus groups) **Consolidated criteria for reporting qualitative research** Version: COREQ 2007 — 32 items in 3 domains. **Specific to in-depth interviews and focus groups** (the dominant qualitative data-collection methods in health research). For other qualitative approaches, use the broader **SRQR** (`SRQR.md`). Source: Tong A, Sainsbury P, Craig J. *Int J Qual Health Care* 2007;19(6):349–357 (the COREQ statement; DOI 10.1093/intqhc/mzm042). EQUATOR Network. Apply when the manuscript reports an **interview or focus-group** qualitative study. For the design/conduct review of the same study, pair with the QL1–QL8 domain probes in `peer-review` / `self-review` `references/domain-probes/qualitative_research.md`; for a non-interview qualitative approach (ethnography, document analysis, etc.), use `SRQR.md`. > Licensing note: COREQ is published in the *International Journal for Quality in Health Care* (© Oxford University Press), with **no Creative Commons licence**. The items below are an **in-house, faithful summary of the criteria (facts/intents, paraphrased — not the verbatim COREQ wording)** for item-by-item assessment; consult the published article (DOI 10.1093/intqhc/mzm042) for exact item text. ## Criteria (grouped by domain) ### Domain 1 — Research team and reflexivity **Personal characteristics** | # | Item | What to check is reported | |---|------|---------------------------| | 1 | Interviewer / facilitator | Which author(s) conducted the interviews or focus groups. | | 2 | Credentials | The interviewer's/facilitator's credentials (e.g. PhD, MD, RN). | | 3 | Occupation | Their occupation/role at the time of the study. | | 4 | Gender | The researcher's gender (where relevant to the dynamic). | | 5 | Experience & training | The interviewer's experience and training in qualitative methods. | **Relationship with participants** | # | Item | What to check is reported | |---|------|---------------------------| | 6 | Relationship established | Whether a relationship existed between researcher and participants **before** the study. | | 7 | Participant knowledge of the interviewer | What participants knew about the researcher (e.g. personal goals, reasons for the research). | | 8 | Interviewer characteristics | Reported researcher characteristics — assumptions, interests, reasons for doing the research. | ### Domain 2 — Study design **Theoretical framework** | # | Item | What to check is reported | |---|------|---------------------------| | 9 | Methodological orientation & theory | The methodological orientation/theory underpinning the study (e.g. grounded theory, content analysis, phenomenology). | **Participant selection** | # | Item | What to check is reported | |---|------|---------------------------| | 10 | Sampling | How participants were selected (purposive, convenience, consecutive, snowball). | | 11 | Method of approach | How participants were approached (face-to-face, telephone, mail, email). | | 12 | Sample size | The number of participants. | | 13 | Non-participation | How many declined or dropped out, and why (where known). | **Setting** | # | Item | What to check is reported | |---|------|---------------------------| | 14 | Setting of data collection | Where the data were collected (home, clinic, workplace). | | 15 | Presence of non-participants | Whether anyone besides participants and researchers was present. | | 16 | Description of sample | The sample's key characteristics (e.g. demographics, dates). | **Data collection** | # | Item | What to check is reported | |---|------|---------------------------| | 17 | Interview guide | Whether questions/prompts/guides were provided, and whether they were pilot-tested. | | 18 | Repeat interviews | Whether any interviews were repeated, and how many. | | 19 | Audio/visual recording | Whether the data were audio- or video-recorded. | | 20 | Field notes | Whether field notes were made during/after the interview or focus group. | | 21 | Duration | The duration of the interviews or focus groups. | | 22 | Data saturation | Whether **data saturation** was discussed/reached. | | 23 | Transcripts returned | Whether participants received the transcripts to review, comment on, or correct. | ### Domain 3 — Analysis and findings **Data analysis** | # | Item | What to check is reported | |---|------|---------------------------| | 24 | Number of data coders | How many coders coded the data. | | 25 | Description of the coding tree | Whether a description of the coding tree/framework is provided. | | 26 | Derivation of themes | Whether themes were identified in advance or derived from the data. | | 27 | Software | Any software used to manage/analyse the data. | | 28 | Participant checking | Whether participants provided feedback on the findings (**member checking**). | **Reporting** | # | Item | What to check is reported | |---|------|---------------------------| | 29 | Quotations presented | Whether participant **quotations** illustrate the themes/findings, and whether each is identified (e.g. participant number). | | 30 | Data & findings consistent | Whether the reported findings cohere with the underlying data shown. | | 31 | Clarity of major themes | Whether the major themes are presented clearly in the results. | | 32 | Clarity of minor themes | Whether deviant/diverse cases and minor themes are addressed, not only the dominant ones. | --- ## Notes for Assessors - The **highest-yield** checks (where interview/focus-group studies most often fail review): **Domain 1 reflexivity** (items 1–8 — who interviewed, their relationship to participants, and their assumptions; the single most-omitted COREQ domain), **item 9** (a named methodological orientation — not "themes emerged" with no method), **items 10/22** (a stated sampling approach and a **saturation** discussion), **items 24–28** (the coding/analysis process — how many coders, the coding framework, software, member checking), and **item 29** (themes substantiated by **identified participant quotations**). - **Do not apply quantitative criteria**: a small purposive sample is appropriate, "generalizability" is **transferability**, and there are no power/effect-size/p-value requirements. A "sample too small / not generalizable" comment is mis-calibrated for an interview study. - COREQ is interview/focus-group-specific; for ethnography, document analysis, or other qualitative approaches use `SRQR.md`. This is an **in-house faithful summary of the COREQ criteria (paraphrased intents, not verbatim)**; map the manuscript's content to the items rather than to exact wording, and consult the published checklist (Tong et al. *Int J Qual Health Care* 2007; DOI 10.1093/intqhc/mzm042) for the exact text. - Verification: all 32 items were compared, by number, domain and topic, against the official COREQ checklist hosted by the EQUATOR Network (the Souza et al. 2021 Brazilian-Portuguese translation, `VERSÃO-FINAL-COREQ.pdf`, which carries the full numbered table). 32/32 match across the three domains, with no item missing and none invented. The English original is behind a paywall and the OUP PDF blocks automated retrieval, so item *wording* has been checked only through that translation; the numbering, domains and topics are confirmed. -
COSMIN_RoB.md 8.5 KB
# COSMIN Risk of Bias Assessment Guide COnsensus-based Standards for the selection of health Measurement INstruments — Risk of Bias tool for reliability and measurement error. Reference: Mokkink LB et al. BMC Medical Research Methodology 2020;20:293. Website: https://www.cosmin.nl Version: COSMIN Risk of Bias tool for reliability and measurement error (2020 Delphi study) Source: Mokkink LB, Boers M, van der Vleuten CPM, Bouter LM, Alonso J, Patrick DL, et al. COSMIN Risk of Bias tool to assess the quality of studies on reliability or measurement error of outcome measurement instruments: a Delphi study. *BMC Med Res Methodol* 2020;20:293 (DOI 10.1186/s12874-020-01179-5). The four-point rating system and the 'worst score counts' principle: Mokkink LB, de Vet HCW, Prinsen CAC, Patrick DL, Alonso J, Bouter LM, Terwee CB. COSMIN Risk of Bias checklist for systematic reviews of Patient-Reported Outcome Measures. *Qual Life Res* 2018;27(5):1171-1179 (DOI 10.1007/s11136-017-1765-4). Review workflow: Prinsen CAC, et al. *Qual Life Res* 2018;27(5):1147-1157 (DOI 10.1007/s11136-018-1798-3). Licence: CC BY 4.0 for all three — confirmed via Crossref. Verification: the seven Part A elements and every Part B standard were compared against Tables 4–7 of the 2020 Delphi paper (Europe PMC full text, PMC7712525). **The counts were right and the standards were not**: this file previously carried invented standards for *Missing data*, *Sample size* ("minimum 30 recommended, 50+ preferred") and *Reporting* — the words "missing" and "sample size" appear nowhere in the source, and the 2018 paper states that standards concerning reporting only were deleted. Official design standard 6 and statistical standards 8 and 9 were absent. The SOURCE line also cited Prinsen 2018, which does not contain these standards. ## Purpose The COSMIN Risk of Bias tool assesses the methodological quality of studies on **reliability** and **measurement error** of outcome measurement instruments (e.g., questionnaires, imaging measurements, lab tests). ## Structure Two parts: - **Part A**: Understanding how the study informs on reliability/measurement error (7 elements of a comprehensive research question) - **Part B**: Assessing quality using standards (9 for reliability, 8 for measurement error) Quality rating per standard: Very good / Adequate / Doubtful / Inadequate. Each standard also carries **NA**. Overall quality uses the **worst score counts** principle — the lowest rating of any standard in the box (Mokkink 2018). There is no averaging and no summary score. Note what is **not** here: this tool has no sample-size standard, no missing-data standard and no reporting standard. Standards that concerned reporting only were deliberately removed when the COSMIN checklist became a risk-of-bias checklist. ## Part A: Elements of a Comprehensive Research Question Extract these 7 elements from the study: | # | Element | |---|---------| | 1 | Name of the outcome measurement instrument | | 2 | Version or operationalization of the measurement protocol | | 3 | Construct measured by the instrument | | 4 | Reliability parameter (ICC, kappa, etc.) or measurement error parameter (SEM, LoA, SDC) | | 5 | Components of the instrument that will be repeated | | 6 | Source(s) of variation that will be varied (time, rater, machine, etc.) | | 7 | Patient population studied | ## Components of Outcome Measurement Instruments Element 5 of the research question asks which **components** are repeated. The Delphi panel agreed on five, in two variants. ### Without biological sampling 1. **Equipment** — all equipment used in preparation, administration, and assigning scores 2. **Preparatory actions** — 'first time only' general actions (required expertise or training) and actions repeated for each measurement 3. **Unprocessed data collection** — what the patient and/or professional(s) actually do to obtain the unprocessed data 4. **Data processing and storage** — all actions on the unprocessed data that allow a score to be assigned 5. **Assignment of the score** — methods used to transform processed data into a final score ### With biological sampling 1. **Equipment** — all equipment used in preparation, administration, and determination of values 2. **Preparatory actions preceding sample collection** — by professionals, patients, and others as applicable 3. **Collection of the biological sample** — all actions to collect the sample, before any processing 4. **Biological sample processing and storage** — preserving, transporting and storing the sample for determination 5. **Determination of the value of the sample** — methods used to count or quantify the substance or entity of interest ## Part B: Standards for Studies on Reliability Design requirements (standards 1–6) are the same for reliability and measurement error; only the statistical standards differ. ### Design requirements (standards 1–6) | # | Standard | |---|----------| | 1 | Were patients stable in the time between the repeated measurements on the construct to be measured? | | 2 | Was the time interval between the repeated measurements appropriate? | | 3 | Were the measurement conditions similar for the repeated measurements — except for the condition being evaluated as a source of variation? | | 4 | Did the professional(s) administer the measurement without knowledge of scores or values of other repeated measurement(s) in the same patients? | | 5 | Did the professional(s) assign the scores or determine the values without knowledge of the scores or values of other repeated measurement(s) in the same patients? | | 6 | Were there any other important flaws in the design or statistical methods of the study? | Standard 6 is rated in the opposite direction: **No** = very good, minor methodological flaws = doubtful, **Yes** = inadequate. ### Preferred statistical methods — reliability (standards 7–9) | # | Standard | Very good | |---|----------|-----------| | 7 | For continuous scores: was an Intraclass Correlation Coefficient (ICC) calculated? | ICC calculated; the model or formula was described and matches the study design and the data | | 8 | For ordinal scores: was a (weighted) Kappa calculated? | Kappa calculated; the weighting scheme was described and matches the study design and the data | | 9 | For dichotomous/nominal scores: was Kappa calculated for each category against the other categories combined? | Kappa calculated for each category against the other categories combined | For standard 7, a Pearson or Spearman correlation **without** evidence that no systematic difference occurred between measurements rates *doubtful*; an ICC whose model or formula is not described rates *adequate*. ## Part B: Standards for Studies on Measurement Error ### Design requirements (standards 1–6) Identical to the reliability design requirements above. ### Preferred statistical methods — measurement error / agreement (standards 7–8) | # | Standard | Very good | |---|----------|-----------| | 7 | For continuous scores: was the Standard Error of Measurement (SEM), Smallest Detectable Change (SDC), Limits of Agreement (LoA) or Coefficient of Variation (CV) calculated? | SEM, SDC, LoA or CV calculated; the model or formula for the SEM/SDC is described and matches the study design and the data | | 8 | For dichotomous/nominal/ordinal scores: was the percentage specific (e.g. positive and negative) agreement calculated? | Percentage *specific* agreement calculated (percentage agreement alone rates only adequate) | A SEM calculated from Cronbach's alpha, or using the SD from another population, rates *inadequate*. Total: **9 standards for reliability** (6 design + 3 statistical) and **8 for measurement error** (6 design + 2 statistical). ## Overall Quality Rating Uses the **worst-score-counts** principle: - Rate each standard as: Very Good / Adequate / Doubtful / Inadequate - Overall rating = lowest rating across all applicable standards | Rating | Interpretation | |--------|---------------| | Very Good | Study design and methods are optimal for this measurement property | | Adequate | Study design and methods are acceptable | | Doubtful | Study design or methods raise some concerns | | Inadequate | Study design or methods are clearly flawed | ## When to Use - Systematic reviews of measurement properties of health measurement instruments - Selecting outcome measurement instruments for clinical trials or research - Developing core outcome sets (COS) - Evaluating reliability/agreement of imaging measurements, scoring systems, clinical tests - Typically used alongside other COSMIN boxes (content validity, structural validity, etc.) -
CROSS.md 9.5 KB
# CROSS Checklist (survey studies) **Consensus-Based Checklist for Reporting of Survey Studies** Version: CROSS 2021 (40 reportable elements across ~7 sections). For internet/electronic surveys, pair with **CHERRIES** (Checklist for Reporting Results of Internet E-Surveys). Sources: Sharma A, Minh Duc NT, Luu Lam Thang T, et al. *J Gen Intern Med* 2021;36:3179–3187 (the CROSS statement; DOI 10.1007/s11606-021-06737-1). Eysenbach G. *J Med Internet Res* 2004;6(3):e34 (CHERRIES; CC BY). EQUATOR Network. Apply when the manuscript is a **self-report survey / questionnaire study** — knowledge-attitudes-practices (KAP), physician or patient surveys, cross-sectional questionnaires, and web/e-surveys. CROSS covers the reportable elements of design, sampling, instrument development, administration, and analysis; for an **internet/e-survey**, the CHERRIES items (open-vs-closed sample, denominator/completion definition, voluntariness/incentives, duplicate-submission control) also apply. For the design/validity review of the same study, pair with the SV1–SV8 domain probes in `peer-review` / `self-review` `references/domain-probes/survey_research.md`; for scale reliability, see `analyze-stats` `Survey/Likert` + `survey_weighted.md`. > Licensing note: the CROSS statement is © Society of General Internal Medicine (not a Creative Commons licence). The items below are an **in-house, faithful summary of the reportable elements (facts/intents, paraphrased — not the verbatim CROSS wording)** for item-by-item assessment; consult the published CROSS article for exact item text. The CHERRIES e-survey items are grounded in the CC BY JMIR source. Verification: every item was compared against Table 1 of the published statement (PubMed Central record PMC8481359). **This file previously carried 19 items of its own devising** — the count happened to look plausible, but from item 8 onwards nothing lined up with CROSS, and the header's "40 reportable elements" was not what the file listed. The numbering below is now the statement's: 19 topics carrying 40 sub-items. The CHERRIES items, which CROSS does not contain, are kept in a clearly separate section. ## Reportable elements (19 topics, 40 sub-items) ### Title and abstract | # | Topic | What to check is reported | |---|-------|---------------------------| | 1a | Title and abstract | The word "survey", with a commonly used design term, appears in the title or abstract. | | 1b | Title and abstract | The abstract summarises background, objectives, methods, findings, interpretation and conclusions. | ### Introduction | # | Topic | What to check is reported | |---|-------|---------------------------| | 2 | Background | Rationale for the survey, what has been done before, and why this survey is needed. | | 3 | Purpose/aim | Specific purposes, aims, goals or objectives. | ### Methods | # | Topic | What to check is reported | |---|-------|---------------------------| | 4 | Study design | The design named in Methods with a commonly used term (cross-sectional, longitudinal). | | 5a | Data collection methods | The questionnaire itself — number of sections, number of questions, number and names of instruments used. | | 5b | Data collection methods | Each instrument used to measure a concept: target population, reported validity and reliability, scoring/classification procedure, reference links. | | 5c | Data collection methods | Pretesting, if performed: method, how many rounds, number and demographics of pretest participants, and how similar they were to the sample population. | | 5d | Data collection methods | The questionnaire provided in full where possible, in the article or as an appendix/online supplement. | | 6a | Sample characteristics | The study population — background, locations, eligibility criteria, exclusion criteria. | | 6b | Sample characteristics | The sampling technique (single- or multistage, simple random, stratified, cluster, convenience), with the locations of participants when clustered sampling was used. | | 6c | Sample characteristics | Sample size, with the details of its calculation. | | 6d | Sample characteristics | How representative the sample is of the study (or target) population — particularly for population-based surveys. | | 7a | Survey administration | Modes of administration: type and number of contacts, and where the survey was conducted (clinic room, an online tool). | | 7b | Survey administration | The survey's time frame — recruitment, exposure and follow-up periods. | | 7c | Survey administration | The entry process: for non-web surveys, how human data-entry error was minimised; for web surveys, how multiple participation was prevented. | | 8 | Study preparation | Any preparation before fielding — interviewer training, advertising the survey. | | 9a | Ethical considerations | Ethical approval where obtained: informed consent, IRB approval, Helsinki declaration, GCP declaration as appropriate. | | 9c | Ethical considerations | Anonymity and confidentiality, and the mechanisms used to prevent unauthorised access. | | 10a | Statistical analysis | Statistical methods and analytical approach, and the software used. | | 10b | Statistical analysis | Any modification of variables used in the analysis, with a reference where available. | | 10c | Statistical analysis | Missing data: rate of missing items, the missing-data mechanism (MCAR / MAR / MNAR), and the methods used to handle it. | | 10d | Statistical analysis | How non-response error was addressed. | | 10e | Statistical analysis | For longitudinal surveys, how loss to follow-up was addressed. | | 10f | Statistical analysis | Whether weighting or propensity scores were used to adjust for non-representativeness. | | 10g | Statistical analysis | Any sensitivity analysis conducted. | *The published table numbers ethical considerations 9a and 9c with no 9b. That gap is the statement's; do not renumber to close it.* ### Results | # | Topic | What to check is reported | |---|-------|---------------------------| | 11a | Respondent characteristics | Numbers of individuals at each stage of the study; a flow diagram where possible. | | 11b | Respondent characteristics | Reasons for non-participation at each stage, where possible. | | 11c | Respondent characteristics | The response rate, **with the definition or formula used to calculate it**. | | 11d | Respondent characteristics | How unique visitors were determined, and their number with the relevant proportions (view, participation, completion). | | 12 | Descriptive results | Characteristics of participants, plus information on potential confounders and assessed outcomes. | | 13a | Main findings | Unadjusted estimates and, where applicable, confounder-adjusted estimates with 95% confidence intervals and p values. | | 13b | Main findings | For multivariable analysis: the model-building process, model fit statistics and model assumptions. | | 13c | Main findings | Any sensitivity analysis; with substantial missing data, a comparison of complete cases against the imputed dataset. | ### Discussion | # | Topic | What to check is reported | |---|-------|---------------------------| | 14 | Limitations | Sources of potential bias and imprecision — non-representativeness of the sample, the design, important uncontrolled confounders. | | 15 | Interpretations | A cautious overall interpretation given those biases and imprecisions, and areas for future research. | | 16 | Generalizability | The external validity of the results. | ### Other sections | # | Topic | What to check is reported | |---|-------|---------------------------| | 17 | Role of the funding source | Whether any funding organisation had a role in the survey's design, implementation or analysis. | | 18 | Conflict of interest | Any potential conflict of interest. | | 19 | Acknowledgements | Organisations/persons acknowledged, with their contribution. | ## CHERRIES companion items (internet/e-surveys) **Not part of CROSS.** For a web or email survey, the CHERRIES items add: open versus closed (invited) survey; how the denominator and completion/response were defined; voluntariness and any incentive; duplicate-submission control (IP address, cookie, log-in); and the use of adaptive or mandatory questions and their effect on completeness. Source: Eysenbach G. *J Med Internet Res* 2004;6(3):e34 (CC BY). CROSS covers some of this ground at 7c and 11d; the rest is CHERRIES only. ## Notes for Assessors - The **highest-yield** items (where surveys most often fail review): **6b/6d** (the sampling technique named, probability versus convenience, and how representative the sample is — a self-selected web panel generalised to "clinicians"/"patients" is the commonest over-reach), **11c** (a response rate **with its definition or formula** — "N responded" with no denominator is non-compliant), **11b/10d** (non-response), **5b/5c** (instrument validity, reliability and pretesting — a novel unvalidated questionnaire carrying the headline), **10f** (weighting for a skewed sample), and the CHERRIES denominator/duplicate-control items for e-surveys. - A survey reporting only a raw count of respondents, with no sampling description, no defined-denominator response rate, no non-response assessment, and an unvalidated instrument, is non-compliant on 6a–6d, 5b, 10d and 11c, and its population-level claims should be downgraded to "among respondents." - This is an **in-house faithful summary of the CROSS reportable elements (paraphrased intents, not verbatim)**; the numbering is the statement's. Consult the published CROSS article (Sharma et al. *JGIM* 2021) for the exact item text. -
DECIDE_AI.md 6 KB
# DECIDE-AI Checklist **Developmental and Exploratory Clinical Investigations of DEcision support systems driven by Artificial Intelligence** Version: DECIDE-AI 2022 Source: https://www.equator-network.org · https://www.decide-ai.org Reference: Vasey B, et al. Nat Med 2022;28(5):924-933. doi:10.1038/s41591-022-01772-9 > Educational summary, authored in our own words. DECIDE-AI materials are CC BY-NC; this file > paraphrases the *intent* of each item and copies no verbatim guideline wording. Complete the > official DECIDE-AI checklist for a submission-ready form and cite Vasey et al. 2022. Verification: all 17 AI-specific items and all 10 generic items were compared, by number, label and order, against Table 2 of the *BMJ* co-publication of the same guideline (Europe PMC full text, PMC9116198). 17/17 and 10/10 match, and the 28 AI-specific sub-items were counted from that table. The generic items had been glossed with a list that named things the guideline does not contain and omitted two that it does; they are now listed. Wording stays paraphrased: DECIDE-AI is CC BY-**NC**, which cannot be redistributed under this repository's MIT licence. ## Naming and scope (read first) - DECIDE-AI reports the **early-stage (live) clinical evaluation** of an AI decision-support system — the development-to-implementation gap *between* offline model validation (TRIPOD+AI / STARD-AI / CLAIM) and a definitive trial (CONSORT-AI). It is **not** a model-accuracy guideline; it is about how the AI behaves, is used, and is kept safe in real clinical workflow during first-in-human use. - It has **17 AI-specific items** (28 subitems) plus **10 generic** reporting items (study identifiers, objectives, setting, sample-size rationale, statistics, funding/COI, registration, etc. — apply these like any reporting guideline). The AI-specific items below are the core. - Distinct from CONSORT-AI (definitive RCT report) and SPIRIT-AI (trial protocol); use DECIDE-AI for the exploratory/early live-clinical evaluation stage. ## AI-specific items (17) Status each PRESENT / PARTIAL / MISSING / N/A. ### Title / Abstract | # | Item | Description (intent) | |---|------|----------------------| | 1 | Title | Identify the study as an early-stage clinical evaluation of an AI/ML-driven decision-support system. | ### Introduction | # | Item | Description (intent) | |---|------|----------------------| | 2 | Intended use | Describe the targeted medical condition(s), the decision the system supports, and the intended users and use context. | ### Methods | # | Item | Description (intent) | |---|------|----------------------| | 3 | Participants | Describe how patients were recruited and the inclusion/exclusion criteria at both patient and data level, how the number was arrived at, and the corresponding information for the **users** (clinicians). | | 4 | AI system | Describe the system: algorithm type, training data and provenance, inputs, outputs, and version. | | 5 | Implementation | Describe how the system was integrated into the **clinical workflow** and the evaluation settings (how/where it was used in practice). | | 6 | Safety and errors | Pre-define what counts as a significant error/malfunction and how such events were identified and captured. | | 7 | Human factors | Describe the human-factors approach: tools, methods/frameworks, and the users involved (usability evaluation plan). | | 8 | Ethics | Describe whether specific methodologies were used to fulfil an ethics-related goal (such as algorithmic fairness), and their rationale. | ### Results | # | Item | Description (intent) | |---|------|----------------------| | 9 | Participants | Report baseline characteristics of patients/users and data missingness. | | 10 | Implementation | Report user exposure to the system and adherence to the intended use (how it was actually used). | | 11 | Modifications | Report any changes made to the AI system during the study (versioning, retraining, threshold changes). | | 12 | Human–computer agreement | Report how often and how users agreed with / overrode the AI recommendations. | | 13 | Safety and errors | List significant errors, malfunctions, and patient-safety events observed. | | 14 | Human factors | Report usability results and **learning curves** over the evaluation. | ### Discussion | # | Item | Description (intent) | |---|------|----------------------| | 15 | Support for intended use | Discuss whether the results support the stated intended clinical use, within the early-stage limits. | | 16 | Safety and errors | Discuss the safety profile and the implications of the observed errors for wider deployment. | ### Other | # | Item | Description (intent) | |---|------|----------------------| | 17 | Data availability | Disclose availability of data and code (with access constraints). | --- ## Notes for Assessors - Apply the **10 generic items** as well. They are, in the guideline's own order: **I** Abstract, **II** Objectives, **III** Research governance, **IV** Outcomes, **V** Analysis, **VI** Patient involvement, **VII** Main results, **VIII** Subgroups analysis, **IX** Strengths and limitations, **X** Conflicts of interest. DECIDE-AI assumes standard reporting underneath. - **Human factors / learning curve (7, 14)** and **human–computer agreement / override (12)** are what make DECIDE-AI distinct from offline-accuracy guidelines; a study that reports only model metrics and no real-use human-interaction data is PARTIAL/MISSING on the core of DECIDE-AI. - **Safety and errors (6/13/16)** must run end-to-end: pre-defined → captured → discussed. A safety claim with no pre-defined error capture is PARTIAL. - **Implementation (5/10)** is about *workflow integration and actual use*, not just deployment intent. - **Modifications (11)**: silent mid-study model/threshold changes without disclosure undermine the evaluation — mark MISSING if changes are implied but not reported. - Use **CONSORT-AI** for a definitive trial and **TRIPOD+AI / STARD-AI / CLAIM** for offline model development/accuracy; DECIDE-AI covers the early live-clinical evaluation between them. -
GATHER.md 7.5 KB
# GATHER Checklist **Guidelines for Accurate and Transparent Health Estimates Reporting** Version: GATHER 2016 (18 items). Source: Stevens GA, Alkema L, Black RE, et al. *The Lancet* 2016;388(10062):e19–e23, published simultaneously in *PLoS Medicine* 2016;13(6):e1002056, https://doi.org/10.1371/journal.pmed.1002056 (CC BY 4.0). https://gather-statement.org · EQUATOR Network. The item text below is reproduced from the published checklist under CC BY 4.0 with attribution; cite the source statement. Apply when the manuscript **reports population health estimates produced by a statistical or mathematical model that synthesizes multiple data sources** — Global Burden of Disease (GBD/IHME) analyses and satellite papers, WHO/UN-agency burden estimates, attributable-burden (comparative-risk / population-attributable-fraction) studies, cause-of-death modeling, prevalence/incidence/mortality estimation, disability-adjusted or quality-adjusted life-year estimation, and their forecasts. GATHER is the reporting standard those estimates are held to; it is orthogonal to STROBE/RECORD (which govern primary and routinely-collected-data studies of *individuals*). A single-institution cohort that only **contextualizes** its finding against a published burden number does not itself trigger GATHER, but a paper that **re-estimates or re-projects** burden does. Pair the analytic methods with `/analyze-stats` `references/analysis_guides/burden_decomposition_forecasting.md` (decomposition, joinpoint/AAPC, forecasting, PAF); pair the reproducibility items (8, 14, 15) with `/verify-refs` and the project's data/code-availability discipline. Source: Stevens GA, Alkema L, Black RE, Boerma JT, Collins GS, Ezzati M, et al. Guidelines for Accurate and Transparent Health Estimates Reporting: the GATHER statement. *Lancet* 2016;388(10062):e19-e23 (DOI 10.1016/S0140-6736(16)30388-9); also *PLoS Med* 2016;13(6):e1002056 (DOI 10.1371/journal.pmed.1002056). Verification: all 18 items were extracted from the checklist table of the *PLoS Medicine* version (Europe PMC full text, PMC4924581) and compared item by item. **The count matched while almost nothing else did**: this file previously invented four items (Sampling, Bias/misclassification correction, Comparability adjustments, Citable results file), dropped four real ones (7, 8, 17, 18), renumbered everything from official item 9 onwards, and rewrote item 15's requirement. ## Checklist Items (18 items) ### Objectives and funding | # | Item | Description | |---|------|-------------| | 1 | Objectives | Define the indicator(s), populations (including age, sex, and geographic entities), and time period(s) for which estimates were made. | | 2 | Funding | List the funding sources for the work. | ### Data inputs *Items 3–6 apply to all data inputs from multiple sources that are synthesized as part of the study.* | # | Item | Description | |---|------|-------------| | 3 | Data identification and access | Describe how the data were identified and how the data were accessed. | | 4 | Inclusion/exclusion criteria | Specify the inclusion and exclusion criteria. Identify all ad-hoc exclusions. | | 5 | Source characteristics | Provide information about all included data sources and their main characteristics. For each data source used, report reference information or contact name/institution, population represented, data collection method, year(s) of data collection, sex and age range, diagnostic criteria or measurement method, and sample size, as relevant. | | 6 | Input-data bias | Identify and describe any categories of input data that have potentially important biases (e.g., based on characteristics listed in item 5). | *Item 7 applies to data inputs that contribute to the analysis but were not synthesized as part of the study.* | # | Item | Description | |---|------|-------------| | 7 | Other data inputs | Describe and give sources for any other data inputs. | *Item 8 applies to all data inputs.* | # | Item | Description | |---|------|-------------| | 8 | Data inputs in an extractable format | Provide all data inputs in a file format from which data can be efficiently extracted (e.g., a spreadsheet rather than a PDF), including all relevant meta-data listed in item 5. For any data inputs that cannot be shared because of ethical or legal reasons, such as third-party ownership, provide a contact name or the name of the institution that retains the right to the data. | ### Data analysis | # | Item | Description | |---|------|-------------| | 9 | Analysis overview | Provide a conceptual overview of the data analysis method. A diagram may be helpful. | | 10 | Analysis detail | Provide a detailed description of all steps of the analysis, including mathematical formulae. This description should cover, as relevant, data cleaning, data pre-processing, data adjustments and weighting of data sources, and mathematical or statistical model(s). | | 11 | Model selection | Describe how candidate models were evaluated and how the final model(s) were selected. | | 12 | Model performance | Provide the results of an evaluation of model performance, if done, as well as the results of any relevant sensitivity analysis. | | 13 | Uncertainty methods | Describe methods of calculating uncertainty of the estimates. State which sources of uncertainty were, and were not, accounted for in the uncertainty analysis. | | 14 | Source-code access | State how analytic or statistical source code used to generate estimates can be accessed. | ### Results and discussion | # | Item | Description | |---|------|-------------| | 15 | Estimates in an extractable format | Provide published estimates in a file format from which data can be efficiently extracted. | | 16 | Quantitative uncertainty | Report a quantitative measure of the uncertainty of the estimates (e.g., uncertainty intervals). | | 17 | Interpretation | Interpret results in light of existing evidence. If updating a previous set of estimates, describe the reasons for changes in estimates. | | 18 | Limitations | Discuss limitations of the estimates. Include a discussion of any modelling assumptions or data limitations that affect interpretation of the estimates. | ## MedSci application notes - **Uncertainty intervals, not confidence intervals (items 13, 16).** Model-based estimates report 95% **uncertainty intervals (UIs)** from posterior/Monte-Carlo draws (commonly 250–500), taken as the 2.5th–97.5th percentiles. A UI crossing the null is "insufficient evidence for direction," not a non-significant test — do not translate it into *P*-value language. - **Robustness is uncertainty propagation plus honest disclosure (items 12, 13, 18).** Burden papers rarely carry a confounding-control toolkit (DAG, E-value, negative controls); that toolkit belongs to individual-level cohorts (STROBE/RECORD). For an estimate paper, the substitute is a propagated UI at every modeling step **plus an itemized limitations paragraph** stating which biases were and were not addressed. State it explicitly rather than implying a sensitivity suite that was not run. - **Reproducibility by pointer (items 8, 14, 15).** Estimate papers satisfy code/data availability by pointing to the standing pipeline's public repository (for GBD: a GHDx data citation and the IHME/analysis GitHub), not by curating a study-specific dataset. - **Provenance of any add-on layer.** If the contribution is a forecast, a decomposition, or a policy-stratified re-slice on top of an existing platform's estimates, report items 9–14 for that added layer specifically — the base platform's methods do not document it. -
GRRAS.md 6.2 KB
# GRRAS Checklist **Guidelines for Reporting Reliability and Agreement Studies** Version: GRRAS 2011 Source: https://doi.org/10.1016/j.jclinepi.2010.03.002 Reference: Kottner J, Audige L, Brorson S, et al. Guidelines for Reporting Reliability and Agreement Studies (GRRAS) were proposed. J Clin Epidemiol. 2011;64(1):96-106. doi:10.1016/j.jclinepi.2010.03.002 Licence: Crossref returns no Creative Commons licence (Elsevier). Verification: all 15 items were compared against the checklist table of the published article (Table 1, *J Clin Epidemiol* 2011;64(1):96-106), read from the article PDF obtained through an institutional subscription. The table renders several rows as mirrored glyphs, so each item was also recovered independently from the article's own section-3 headings ("Item N: ...") and the two readings agreed. **The count matched and almost nothing else did**: 13 of 15 items were wrong. Only items 1 and 14 survived. This file previously invented a *Limitations* item that GRRAS does not have, and omitted item 2 (name and describe the measurement device) entirely; items 2-13 were a different checklist wearing GRRAS's numbering. The source article is not redistributable (Elsevier, no open licence; **OpenAlex and Unpaywall both report no open-access copy of either the *J Clin Epidemiol* or the *Int J Nurs Stud* version**). The item text below therefore states what each item asks **in our own words**. Complete the official instrument for anything you report. ## Checklist Items (15 items) ### Title and abstract | # | Item | Description | |---|------|-------------| | 1 | Identification | Make clear in the title or the abstract whether interrater or intrarater **reliability** or **agreement** was the subject of the study. | ### Introduction | # | Item | Description | |---|------|-------------| | 2 | Measurement device | Name the diagnostic or measurement device under study and describe it explicitly. | | 3 | Subject population | Specify the population of subjects (or objects) the study is about. | | 4 | Rater population | Specify the population of raters the study is about, where raters are involved. | | 5 | Existing knowledge and rationale | Describe what is already known about the reliability and agreement of this measurement, and give the rationale for doing this study. | ### Methods | # | Item | Description | |---|------|-------------| | 6 | Sample size | Explain how the sample size was arrived at, and state the number of raters, of subjects/objects, and of replicate observations that were decided on. | | 7 | Sampling method | Describe how subjects/objects (and raters, where applicable) were sampled. | | 8 | Measurement/rating process | Describe how measurements or ratings were carried out — including the time interval between repeated measurements, what clinical information was available to raters, and any blinding. | | 9 | Independence | State whether the repeated measurements or ratings were made **independently** of one another. | | 10 | Statistical analysis | Describe the statistical analysis. | ### Results | # | Item | Description | |---|------|-------------| | 11 | Numbers actually included | State the number of raters and of subjects/objects that were actually included, and the number of replicate observations actually made. | | 12 | Sample characteristics | Describe the characteristics of the raters and of the subjects — for raters, things such as training and experience. | | 13 | Reliability and agreement estimates | Report the reliability and/or agreement estimates **together with a measure of statistical uncertainty** (e.g. a confidence interval or standard error). | ### Discussion | # | Item | Description | |---|------|-------------| | 14 | Practical relevance | Discuss the practical relevance of the results. | ### Auxiliary material | # | Item | Description | |---|------|-------------| | 15 | Detailed results | Provide detailed results where possible — for example as online supplementary material. | **GRRAS has no Limitations item.** It also has no separate Objectives item; the study's purpose is carried by items 3-5. Do not add either when scoring. ## Notes for Assessors - GRRAS applies to all reliability and agreement studies, including inter-rater, intra-rater, test-retest, and method-comparison designs across any medical discipline. - **Reliability vs Agreement**: These are distinct concepts. Reliability (e.g., ICC, kappa) reflects the ability to distinguish between subjects; agreement (e.g., Bland-Altman limits of agreement, SEM) reflects how close scores are on repeated measurements. A study may report one or both. - Item 10 (Statistical analysis): the ICC model (one-way random, two-way random, two-way mixed) and type (single measures, average measures) must match the study design. Flag mismatches as a Major Comment. GRRAS states this item in one line and leaves the detail to the analyst — the article's Table 2 maps statistical methods onto the level of measurement and onto reliability versus agreement. - Items 4 and 12 (rater population, rater characteristics): rater characteristics strongly influence reliability estimates. A study using expert raters only will overestimate reliability for general clinical practice. Note it when the raters are not representative of the population item 4 names. - Item 9 (Independence) is the one most often skipped, and it is narrow: it asks whether the repeated measurements were made independently — not whether their *order* was randomised. Order randomisation and blinding belong to item 8. Both matter for memory and learning effects; score them where GRRAS puts them. - Item 13: for an ICC, report the model, type and definition (consistency vs absolute agreement); for a kappa, whether it is weighted and with which scheme. A bare estimate with no measure of statistical uncertainty fails the item outright, since the uncertainty is part of what item 13 asks for. - Common in radiology: inter-reader agreement for measurements (Bland-Altman, ICC), segmentation agreement (Dice, ICC of volumes), and diagnostic classification agreement (kappa, percent agreement). - GRRAS is complementary to study-type checklists. For example, a diagnostic accuracy study reporting inter-reader agreement should follow STARD for the main analysis and GRRAS for the agreement component. -
MI_CLEAR_LLM.md 8.7 KB
# MI-CLEAR-LLM Checklist — Reporting LLM Accuracy Studies in Healthcare **Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare** - **Version:** MI-CLEAR-LLM 2025 (8 item categories) — updates and expands the 2024 original (6 items). - **Citation (2025 update):** Park SH, Suh CH, Lee JH, Tejani AS, You SC, Kahn CE Jr, Moy L. *MI-CLEAR-LLM: 2025 Updates.* Korean J Radiol. 2025;26(12):1123-1132. PMID: 41199132. - **DOI:** 10.3348/kjr.2025.1522 - **Citation (2024 original):** Park SH, Suh CH, Lee JH, Kahn CE Jr, Moy L. Korean J Radiol. 2024;25(10):865-868. PMID: 39344542. - **Licence:** CC BY-NC 4.0 (Korean Society of Radiology). > **Educational summary, authored in our own words.** This file paraphrases the *intent* of each MI-CLEAR-LLM > 2025 item to drive an item-by-item audit; it does not reproduce the guideline's verbatim wording. For a > submission-ready checklist, complete the official instrument at the source above and cite Park et al. 2025. **The 2025 update expanded the original 6 items into 8 item categories.** The three items that are new or promoted to the top level in 2025 — **Access mode (2)**, **Input data type (3)**, and **Adaptation strategy used (4)** — were prompted by the growth of API access and self-managed (often open-source) LLMs, which the 2024 six-item version folded into a single "model identification" item. Cite items by the 2025 numbering. **Scope:** Studies evaluating the accuracy of LLMs in healthcare tasks (diagnosis, triage, clinical decision support, medical question answering, information extraction, report generation, etc.). This is NOT for disclosing LLM use in manuscript *writing* — for that, see ICMJE/COPE policies and the write-paper skill's LLM disclosure feature. MI-CLEAR-LLM supplements, and does not replace, a primary reporting guideline (STARD / STARD-AI, CLAIM 2024, or TRIPOD+AI) chosen for the study design. Verification: all 8 item categories were compared, by name and order, against Table 1 of the published statement (Europe PMC full text, PMC12683746); 8/8 match. The wording below stays our own condensed summary because the source is CC BY-**NC**, which cannot be redistributed under this repository's licence. ## Checklist Items (8 items) | # | Item category | What must be reported | |---|---|---| | 1 | Model identification | Model name, version, developer, proprietary/open-source status, access date(s), and training-data cutoff. | | 2 | Access mode | Whether a web chatbot, an API, or a self-managed local deployment was used, and why; any known system-level features beyond the LLM itself (system prompts, intersession memory); and, for a local deployment, key computational-environment details. | | 3 | Input data type | The type and format of data supplied with the prompt, in enough detail for a reader to replicate. | | 4 | Adaptation strategy used | Whether model weights were altered (e.g., fine-tuning) or non-parametric methods (e.g., prompting, RAG) were used — or neither. | | 5 | Prompt optimization procedures | How prompts were created and optimized, the rationale, and the full executable prompt text. | | 6 | Prompt execution setup | How queries were submitted (session handling, batching, post-processing); provide the experiment scripts where feasible. | | 7 | Stochasticity management | Temperature and sampling settings, number of attempts per item, how repeats were synthesized, and reliability across repeats. | | 8 | Independence of test data | Any overlap between the test data and the model's training data or the data used for prompt development/adaptation. | ## Item detail ### 1. Model identification ⚠️ commonly missed Fully identify the LLM: model name and version (e.g., GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro), developer (OpenAI, Anthropic, Google, etc.), whether it is proprietary or open-source, the date(s) the queries were run, and the training-data cutoff if disclosed. LLMs are frequently updated behind the same name, so a name alone (e.g., "GPT-4") is insufficient — the exact version and access date are what make the study reproducible. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### 2. Access mode ⚠️ new/expanded in 2025 State how the model was accessed and why: a web-based chatbot interface, a programmatic API, or a self-managed local deployment. For API access, give the specific API version or endpoint. For self-managed/open-source models, give the weights version, any quantization, and the hardware. Access mode changes what is controllable (e.g., a web chatbot may not expose temperature) and what is reproducible. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### 3. Input data type ⚠️ new/expanded in 2025 Describe the type and format of the data supplied to the model with the prompt — free text, structured fields, tables, images (for vision-language models), or file attachments — in enough detail that a reader could reconstruct the input. Where images or clinical data are used, describe how they were encoded and presented. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### 4. Adaptation strategy used ⚠️ new/expanded in 2025 State whether the base model was adapted, and how. Distinguish **parametric** adaptation (model weights were changed, e.g., fine-tuning, instruction-tuning) from **non-parametric** methods that leave weights unchanged (e.g., prompting strategies, retrieval-augmented generation, tool use) — or state that no adaptation was applied. Adaptation is a first-class methodological choice, not a footnote. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### 5. Prompt optimization procedures ⚠️ commonly missed Document how prompts were developed and optimized: who wrote them, how they were iterated, whether prompt engineering or automated tuning was used, and how many variants were compared. Provide the **full prompt text** (system, user, and any few-shot examples), preserving exact wording, punctuation, and formatting — in supplementary material if lengthy. Any dataset used to optimize prompts must be independent of the test data (see item 8); undisclosed optimization on test-overlapping data inflates apparent performance. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### 6. Prompt execution setup ⚠️ commonly missed Describe how prompts were run operationally: whether each query was an independent session or a continuing conversation, whether queries were batched or sequential, whether prior outputs could carry over as context, and any post-processing applied to outputs (e.g., parsing structured answers from free text). For API studies, give batch size, rate limiting, and failure handling. Provide the experiment scripts where feasible. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### 7. Stochasticity management ⚠️ commonly missed LLM outputs are stochastic. Report the number of query attempts per item, how multiple outputs were combined (majority vote, best-of-N, mean score), and a reliability analysis across repeats (e.g., agreement rate, kappa). Report the parameters that control randomness — temperature, top-p, top-k, penalties — or, if they were not controllable (e.g., a web chatbot), state that explicitly. A single-attempt study cannot characterize reliability and should be scored PARTIAL at best. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ### 8. Independence of test data ⚠️ commonly missed Confirm that the test data were separate from all other data used, and address training-data contamination: whether the test items could have been in the model's pretraining corpus (a particular risk for published exam questions and public benchmarks), and any mitigation (unpublished cases, temporal or institution-specific data). Because most LLM training corpora are undisclosed, contamination cannot be ruled out by inspection and must be discussed. **Status:** [ ] PRESENT [ ] PARTIAL [ ] MISSING ## Notes for assessors - **Apply MI-CLEAR-LLM when the study's outcome is LLM accuracy or performance** (answering board questions, clinical reasoning, triage, information extraction, report generation). Do NOT apply it when an LLM is merely a pipeline tool whose accuracy is not the outcome, or when the manuscript only discloses writing assistance. - **"GPT-4" is insufficient (item 1)** — require the exact version (e.g., gpt-4-0613, gpt-4-turbo-2024-04-09) and the access date. **Single-run studies (item 7)** are PARTIAL at best. **Published exam questions (item 8)** demand an explicit contamination discussion. - **Co-application:** use MI-CLEAR-LLM alongside STARD/STARD-AI (diagnostic accuracy against a reference standard), CLAIM 2024 (an LLM processing medical images), or TRIPOD+AI (an LLM used as a prediction model). The 8 items here supplement — they do not replace — the primary reporting guideline. -
MOOSE.md 6.5 KB
# MOOSE Checklist **Meta-analysis Of Observational Studies in Epidemiology** Version: MOOSE 2000 — 35 items in six groups (Background 6, Search strategy 10, Methods 8, Results 4, Discussion 3, Conclusions 4). Source: Stroup DF, Berlin JA, Morton SC, Olkin I, Williamson GD, Rennie D, et al. Meta-analysis of observational studies in epidemiology: a proposal for reporting. *JAMA* 2000;283(15):2008-2012 (DOI 10.1001/jama.283.15.2008). Licence: *JAMA* (© American Medical Association) — **no open licence**. The descriptions below are an in-house summary of what each item asks, in our own words, not the published wording. Verification: all 35 items were compared against the checklist table published in the statement. The item set, the six groups, their counts and their order all match; nothing is missing and nothing has been added. ## Checklist Items (35 items) ### Reporting of Background | # | Item | Description | |---|------|-------------| | 1 | Problem definition | Define the clinical or public health problem being addressed. | | 2 | Hypothesis statement | State the hypothesis to be tested. | | 3 | Study outcome(s) | Describe the study outcome(s) of interest. | | 4 | Type of exposure or intervention | Describe the type of exposure or intervention used. | | 5 | Type of study designs | Describe the type of study designs included (cohort, case-control, cross-sectional, etc.). | | 6 | Study population | Describe the study population. | ### Reporting of Search Strategy | # | Item | Description | |---|------|-------------| | 7 | Qualifications of searchers | Report the qualifications of the searchers (e.g., librarian, investigator). | | 8 | Search strategy | Describe the search strategy, including time period included and keywords used. | | 9 | Effort to include all available studies | Describe efforts to include all available studies, including contact with authors. | | 10 | Databases and registries searched | List all databases and registries searched. | | 11 | Search software used | Report the search software used, including name and version. | | 12 | Use of hand searching | Describe any use of hand searching (e.g., reference lists, conference proceedings). | | 13 | List of citations located and excluded | Provide a list of citations located and those excluded, including justifications for exclusions. | | 14 | Method for non-English articles | Describe the method of addressing articles published in languages other than English. | | 15 | Method for unpublished studies | Describe the method of handling abstracts and unpublished studies. | | 16 | Description of author contact | Describe any contact with authors for additional data or clarification. | ### Reporting of Methods | # | Item | Description | |---|------|-------------| | 17 | Study relevance | Describe the relevance or appropriateness of studies assembled for assessing the hypothesis to be tested. | | 18 | Rationale for data selection and coding | Provide the rationale for the selection and coding of data (e.g., sound clinical principles, convenience, etc.). | | 19 | Data classification | Document how data were classified and coded (e.g., multiple raters, blinding, inter-rater reliability). | | 20 | Assessment of confounding | Describe the assessment of confounding (e.g., comparability of cases and controls, adjustment for confounders). | | 21 | Assessment of study quality | Describe the assessment of study quality, including blinding of quality assessors; stratification or regression on possible predictors of study results. | | 22 | Assessment of heterogeneity | Describe the assessment of heterogeneity. | | 23 | Statistical methods | Describe statistical methods (e.g., fixed vs random effects, meta-regression, cumulative meta-analysis) in sufficient detail to be replicated. | | 24 | Tables and graphics | Provide appropriate tables and graphics. | ### Reporting of Results | # | Item | Description | |---|------|-------------| | 25 | Graphic summary | Provide a graphic summarizing individual study estimates and the overall estimate (e.g., forest plot). | | 26 | Descriptive table | Provide a table giving descriptive information for each study included (design, sample size, outcome measures, effect sizes, confounders adjusted for). | | 27 | Sensitivity testing | Report results of sensitivity testing (e.g., subgroup analysis, influence analysis, varying inclusion criteria). | | 28 | Statistical uncertainty | Indicate statistical uncertainty of findings (e.g., 95% confidence intervals). | ### Reporting of Discussion | # | Item | Description | |---|------|-------------| | 29 | Quantitative assessment of bias | Provide a quantitative assessment of bias (e.g., publication bias via funnel plot, Egger test, trim-and-fill). | | 30 | Justification for exclusion | Justify any exclusions from the meta-analysis. | | 31 | Assessment of quality of included studies | Assess the quality of included studies and its impact on the overall findings. | ### Reporting of Conclusions | # | Item | Description | |---|------|-------------| | 32 | Alternative explanations | Consider alternative explanations for observed results. | | 33 | Generalization | Discuss the generalization of the conclusions (i.e., external validity). | | 34 | Guidelines for future research | Provide guidelines for future research. | | 35 | Disclosure of funding | Disclose the funding source for the meta-analysis. | --- ## Notes for Assessors - MOOSE is designed specifically for meta-analyses of **observational** studies (cohort, case-control, cross-sectional). For meta-analyses of RCTs, use PRISMA 2020. - MOOSE and PRISMA are complementary. Many journals require adherence to both for observational MAs. When both apply, audit against both checklists separately. - Items 7-16 (Search Strategy) overlap substantially with PRISMA but include observational-specific items such as handling of confounding and study quality assessment. - Item 20 (Assessment of confounding) is particularly important for observational MAs, as confounding is the primary threat to validity -- unlike RCTs where randomization addresses this. - Item 21 (Assessment of study quality) should describe a formal tool. Common tools include NOS (Newcastle-Ottawa Scale), ROBINS-I, or the Joanna Briggs Institute checklists. - Item 29 (Quantitative assessment of bias) is critical. Publication bias assessment is expected in all meta-analyses, but is especially important for observational studies where positive-result bias is well documented. - MOOSE was published before the PRISMA era (2009/2020). Some journals now accept PRISMA with MOOSE-specific additions. If in doubt, check journal instructions. -
NOS.md 4.3 KB
# Newcastle-Ottawa Scale (NOS) Assessment Guide Quality assessment tool for non-randomised studies in meta-analyses. Reference: Wells GA et al. Ottawa Hospital Research Institute. Version: Newcastle-Ottawa Scale (current web version) Source: Wells GA, Shea B, O'Connell D, Peterson J, Welch V, Losos M, Tugwell P. The Newcastle-Ottawa Scale (NOS) for assessing the quality of nonrandomised studies in meta-analyses. Ottawa Hospital Research Institute. **No DOI: the tool is distributed from the institute's website and has never been issued one.**. Licence: No formal licence is published with the tool. Verification: both scales were compared item by item and response-option by response-option against the official NOS rating sheet distributed by the Ottawa Hospital Research Institute (`nosgen.pdf`), with the accompanying manual (`nos_manual.pdf`) checked for the scoring rules. The items, their order, the star allocations and the two-star comparability rule all match. **Two numbers did not**: the scale leaves the follow-up percentages blank for the reviewer to set, and this file had hardcoded them; and the Good/Fair/Poor star bands appear nowhere in the NOS documents. Both are corrected below. ## Structure NOS uses a "star system" (maximum 9 stars) across 3 categories. Higher stars = higher quality. ## Cohort Studies (max 9 stars) ### Selection (max 4 stars) 1. **Representativeness of the exposed cohort** (1 star) - a) Truly representative of the average [describe] in the community * - b) Somewhat representative * - c) Selected group of users - d) No description 2. **Selection of the non-exposed cohort** (1 star) - a) Drawn from the same community as the exposed * - b) Drawn from a different source - c) No description 3. **Ascertainment of exposure** (1 star) - a) Secure record (e.g., surgical records) * - b) Structured interview * - c) Written self-report - d) No description 4. **Demonstration that outcome was not present at start** (1 star) - a) Yes * - b) No ### Comparability (max 2 stars) 5. **Comparability of cohorts on the basis of design or analysis** (up to 2 stars) - a) Study controls for [most important factor] * - b) Study controls for any additional factor * ### Outcome (max 3 stars) 6. **Assessment of outcome** (1 star) - a) Independent blind assessment * - b) Record linkage * - c) Self-report - d) No description 7. **Was follow-up long enough for outcomes to occur?** (1 star) - a) Yes (select adequate follow-up period) * - b) No 8. **Adequacy of follow-up of cohorts** (1 star) - a) Complete follow-up (all subjects accounted for) * - b) Subjects lost to follow-up unlikely to introduce bias — small number lost, above a threshold **you select**, or description provided of those lost * - c) Follow-up rate below the threshold **you select**, and no description of those lost - d) No statement The scale prints these two percentages as blanks ("select an adequate %"). NOS does not supply a number; pre-specify yours in the protocol and report it. ## Case-Control Studies (max 9 stars) ### Selection (max 4 stars) 1. Is the case definition adequate? 2. Representativeness of the cases 3. Selection of controls 4. Definition of controls ### Comparability (max 2 stars) 5. Comparability of cases and controls (same 2-star system) ### Exposure (max 3 stars) 6. Ascertainment of exposure 7. Same method of ascertainment for cases and controls 8. Non-response rate ## Interpretation **The NOS does not define quality bands.** Neither the rating sheet nor the manual maps a star total onto "good", "fair" or "poor". The bands below are a widely used external convention (they come from AHRQ-derived practice, not from the scale) and are reproduced here only because reviews so often cite them: | Stars | Convention | |-------|------------| | 7-9 | Good | | 4-6 | Fair | | 0-3 | Poor | If you use them, say where they came from and pre-specify them in the protocol. A different threshold is equally defensible, and the scale's authors leave the choice to you. ## When to Use - Observational cohort studies in intervention or exposure meta-analyses - Case-control studies - Simpler alternative to ROBINS-I when full domain-level assessment is not needed - Note: NOS does not provide domain-level judgments -- only an aggregate score -
PGS_RS.md 6.4 KB
# PGS-RS (PRS-RS) Checklist **Polygenic Score / Polygenic Risk Score Reporting Standards** Version: PGS-RS 2021 (also cited as PRS-RS) Source: Wand H, Lambert SA, Tamburro C, et al. *Improving reporting standards for polygenic scores in risk prediction studies.* Nature 2021;591:211–219. Endorsed by the Clinical Genome Resource (ClinGen) Complex Disease Working Group and the Polygenic Score (PGS) Catalog. Apply when the manuscript develops, validates, or applies a **polygenic risk score / polygenic score (PRS / PGS)** as a predictor or risk-stratifier. For the design-validity review of the same study, pair with the PG1–PG8 domain probes in `peer-review` / `self-review` `references/domain-probes/polygenic_risk_score.md`; for the analysis, with `analyze-stats` `analysis_guides/polygenic_risk_score.md`. PGS-RS is a domain extension of the TRIPOD prediction-model reporting principles — for the general prediction-model items also consult `TRIPOD.md` / `TRIPOD_AI.md`. Source: Wand H, Lambert SA, Tamburro C, Iacocca MA, O'Sullivan JW, Sillari C, et al. Improving reporting standards for polygenic scores in risk prediction studies. *Nature* 2021;591(7849):211-219 (DOI 10.1038/s41586-021-03243-6). Licence: Crossref returns only a Springer Nature text-and-data-mining licence; no Creative Commons licence. ## Checklist Items (22 items) ### Background | # | Item | Description | |---|------|-------------| | 1 | Study type | Specify the development and/or validation status of the polygenic score and provide identifiers (e.g., a PGS Catalog ID) where available. | | 2 | Risk model purpose and predicted outcome | Define the intended use of the risk model and the outcome it predicts. | ### Study Population and Data | # | Item | Description | |---|------|-------------| | 3 | Study design and recruitment | Describe the cohort/study design, eligibility criteria, and recruitment period for the development and the evaluation samples. | | 4 | Participant demographic and clinical characteristics | Report age, sex, and relevant clinical/phenotypic characteristics of each sample. | | 5 | Ancestry | Report the ancestral background of the GWAS (base), training/tuning, and evaluation (target) samples using a standardized ancestry framework; do not assume transferability across ancestries. | | 6 | Genetic data | Describe genotyping/sequencing methods, imputation, and quality control. | | 7 | Non-genetic variables | Define any non-genetic variables included in the (integrated) risk model. | | 8 | Outcome of interest | Define and report how the predicted outcome(s) were ascertained. | | 9 | Missing data | Explain how missing data (genetic and non-genetic) were handled. | ### Risk Model Development and Application | # | Item | Description | |---|------|-------------| | 10 | Polygenic risk score construction and estimation | Detail the source GWAS, variant selection, weighting/shrinkage method (e.g., P+T, LDpred, PRS-CS), reference panel, and any tuning of hyperparameters (and the data on which tuning was performed). | | 11 | Risk model type | Describe the statistical method(s) used to estimate risk from the score. | | 12 | Integrated risk model(s) description and fitting | If the PGS is combined with non-genetic predictors (e.g., a clinical risk model), describe the integrated-model development and fitting procedure. | ### Risk Model Evaluation | # | Item | Description | |---|------|-------------| | 13 | PRS distribution | Describe the distribution of the continuous polygenic score (and any standardization or quantile categorization, with the reference group). | | 14 | Risk model predictive ability | Report variance explained (e.g., R², liability-scale) and effect estimates (e.g., OR/HR per SD, per quantile) with measures of uncertainty. | | 15 | Risk model discrimination | Report discrimination (AUROC/C-index; AUPRC where prevalence is low), with confidence intervals — and, for an incremental claim, the change versus the established clinical model (ΔC-statistic, NRI/IDI). | | 16 | Risk model calibration | Describe how calibration of absolute risk was assessed (calibration plot, slope/intercept, observed vs expected), especially in the target population and across ancestries. | | 17 | Subgroup analyses | Report performance for relevant subgroups (e.g., ancestry, sex, age) rather than aggregate metrics alone. | ### Limitations and Clinical Implications | # | Item | Description | |---|------|-------------| | 18 | Risk model interpretation | Summarize predictive performance and the incremental value over existing risk factors / clinical models; do not equate a quantile relative-risk gradient with demonstrated clinical actionability. | | 19 | Limitations | Outline study restrictions and their impact on interpretation. | | 20 | Generalizability | Discuss applicability to target populations, including ancestries and settings not represented in development/evaluation. | | 21 | Risk model intended uses | Discuss clinical utility and deployment readiness, scaled to the evidence (calibration, incremental value, and ideally decision-analytic or trial evidence). | | 22 | Data transparency and availability | Make the PRS parameters (variants and weights) and evaluation results publicly available (e.g., deposit in the PGS Catalog) so the score is reproducible. | --- ## Notes for Assessors - The highest-yield items are **5 / 17 / 20** (ancestry reporting and per-ancestry/subgroup performance — the central PGS transferability problem), **15** (incremental value over the clinical model, not score-alone discrimination), **16** (absolute-risk calibration, not discrimination alone), and **22** (reproducibility via deposited variants+weights / a PGS Catalog ID). - PGS-RS is a domain extension of general prediction-model reporting; for items on model development/validation also apply `TRIPOD.md` / `TRIPOD_AI.md`, and cite both the base instrument and PGS-RS rather than PGS-RS alone. - The guideline is abbreviated both **PGS-RS** and **PRS-RS** in the literature; they refer to the same Wand et al. 2021 standard. - Verification: all 22 items were compared, by label and order, against Table 1 of the published standard (PubMed Central record PMC8609771). 22/22 match, with no item missing and none invented. The wording stays a condensed summary — Crossref returns only a Springer Nature text-and-data-mining licence, so the item text cannot be reproduced here. Consult the published Table 1 for the full wording. -
PRISMA_2020.md 10.9 KB
# PRISMA 2020 Checklist **Preferred Reporting Items for Systematic Reviews and Meta-Analyses** Version: PRISMA 2020 — 27 items / 42 sub-items. Source: Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. *BMJ* 2021;372:n71 (DOI 10.1136/bmj.n71). Official checklist: https://www.prisma-statement.org Licence: CC BY 4.0. Item text is reproduced verbatim from the official checklist with attribution. > **Fidelity**: every one of the 42 sub-item cells below has been compared, normalised and > character-for-character, against the official PRISMA 2020 checklist PDF. Any editorial comment > of ours lives in *Notes for Assessors*, never inside an item cell — an assessor scoring a > manuscript must read the guideline's words, not ours. ## Checklist Items (27 items) ### Title | # | Item | Description | |---|------|-------------| | 1 | Title | Identify the report as a systematic review. | ### Abstract | # | Item | Description | |---|------|-------------| | 2 | Abstract | See the PRISMA 2020 for Abstracts checklist. | ### Introduction | # | Item | Description | |---|------|-------------| | 3 | Rationale | Describe the rationale for the review in the context of existing knowledge. | | 4 | Objectives | Provide an explicit statement of the objective(s) or question(s) the review addresses. | ### Methods | # | Item | Description | |---|------|-------------| | 5 | Eligibility criteria | Specify the inclusion and exclusion criteria for the review and how studies were grouped for the syntheses. | | 6 | Information sources | Specify all databases, registers, websites, organisations, reference lists, and other sources searched or consulted to identify studies. Specify the date when each source was last searched or consulted. | | 7 | Search strategy | Present the full search strategies for all databases, registers, and websites, including any filters and limits used. | | 8 | Selection process | Specify the methods used to decide whether a study met the inclusion criteria of the review, including how many reviewers screened each record and each report retrieved, whether they worked independently, and if applicable, details of automation tools used in the process. | | 9 | Data collection process | Specify the methods used to collect data from reports, including how many reviewers collected data from each report, whether they worked independently, any processes for obtaining or confirming data from study investigators, and if applicable, details of automation tools used in the process. | | 10a | Data items | List and define all outcomes for which data were sought. Specify whether all results that were compatible with each outcome domain in each study were sought (e.g., for all measures, time points, analyses), and if not, the methods used to decide which results to collect. | | 10b | Data items | List and define all other variables for which data were sought (e.g., participant and intervention characteristics, funding sources). Describe any assumptions made about any missing or unclear information. | | 11 | Study risk of bias assessment | Specify the methods used to assess risk of bias in the included studies, including details of the tool(s) used, how many reviewers assessed each study and whether they worked independently, and if applicable, details of automation tools used in the process. | | 12 | Effect measures | Specify for each outcome the effect measure(s) (e.g., risk ratio, mean difference) used in the synthesis or presentation of results. | | 13a | Synthesis methods | Describe the processes used to decide which studies were eligible for each synthesis (e.g., tabulating the study intervention characteristics and comparing against the planned groups for each synthesis). | | 13b | Synthesis methods | Describe any methods required to prepare the data for presentation or synthesis, such as handling of missing summary statistics, or data conversions. | | 13c | Synthesis methods | Describe any methods used to tabulate or visually display results of individual studies and syntheses. | | 13d | Synthesis methods | Describe any methods used to synthesize results and provide a rationale for the choice(s). If meta-analysis was performed, describe the model(s), method(s) to identify the presence and extent of statistical heterogeneity, and software package(s) used. | | 13e | Synthesis methods | Describe any methods used to explore possible causes of heterogeneity among study results (e.g., subgroup analysis, meta-regression). | | 13f | Synthesis methods | Describe any sensitivity analyses conducted to assess robustness of the synthesized results. | | 14 | Reporting bias assessment | Describe any methods used to assess risk of bias due to missing results in a synthesis (arising from reporting biases). | | 15 | Certainty assessment | Describe any methods used to assess certainty (or confidence) in the body of evidence for an outcome. | ### Results | # | Item | Description | |---|------|-------------| | 16a | Study selection | Describe the results of the search and selection process, from the number of records identified in the search to the number of studies included in the review, ideally using a flow diagram. | | 16b | Study selection | Cite studies that might appear to meet the inclusion criteria, but which were excluded, and explain why they were excluded. | | 17 | Study characteristics | Cite each included study and present its characteristics. | | 18 | Risk of bias in studies | Present assessments of risk of bias for each included study. | | 19 | Results of individual studies | For all outcomes, present, for each study: (a) summary statistics for each group (where appropriate) and (b) an effect estimate and its precision (e.g. confidence/credible interval), ideally using structured tables or plots. | | 20a | Results of syntheses | For each synthesis, briefly summarise the characteristics and risk of bias among contributing studies. | | 20b | Results of syntheses | Present results of all statistical syntheses conducted. If meta-analysis was done, present for each the summary estimate and its precision (e.g., confidence/credible interval) and measures of statistical heterogeneity. If comparing groups, describe the direction of the effect. | | 20c | Results of syntheses | Present results of all investigations of possible causes of heterogeneity among study results. | | 20d | Results of syntheses | Present results of all sensitivity analyses conducted to assess the robustness of the synthesized results. | | 21 | Reporting biases | Present assessments of risk of bias due to missing results (arising from reporting biases) for each synthesis assessed. | | 22 | Certainty of evidence | Present assessments of certainty (or confidence) in the body of evidence for each outcome assessed. | ### Discussion | # | Item | Description | |---|------|-------------| | 23a | Discussion | Provide a general interpretation of the results in the context of other evidence. | | 23b | Discussion | Discuss any limitations of the evidence included in the review. | | 23c | Discussion | Discuss any limitations of the review processes used. | | 23d | Discussion | Discuss implications of the results for practice, policy, and future research. | ### Other Information | # | Item | Description | |---|------|-------------| | 24a | Registration and protocol | Provide registration information for the review, including register name and registration number, or state that the review was not registered. | | 24b | Registration and protocol | Indicate where the review protocol can be accessed, or state that a protocol was not prepared. | | 24c | Registration and protocol | Describe and explain any amendments to information provided at registration or in the protocol. | | 25 | Support | Describe sources of financial or non-financial support for the review, and the role of the funders or sponsors in the review. | | 26 | Competing interests | Declare any competing interests of review authors. | | 27 | Availability of data, code, and other materials | Report which of the following are publicly available and where they can be found: template data collection forms; data extracted from included studies; data used for all analyses; analytic code; any other materials used in the review. | --- ## PRISMA 2020 Flow Diagram The PRISMA flow diagram (item 16a) should include four phases: ``` IDENTIFICATION Records identified from databases (n = ?) Records identified from other sources (n = ?) | v Records removed before screening: Duplicate records (n = ?) Records marked as ineligible by automation tools (n = ?) Records removed for other reasons (n = ?) | v SCREENING Records screened (n = ?) Records excluded (n = ?) | v Reports sought for retrieval (n = ?) Reports not retrieved (n = ?) | v Reports assessed for eligibility (n = ?) Reports excluded (n = ?), with reasons: Reason 1 (n = ?) Reason 2 (n = ?) Reason 3 (n = ?) | v INCLUDED Studies included in review (n = ?) Reports of included studies (n = ?) ``` --- ## Notes for Assessors - PRISMA 2020 expanded from the original 27-item checklist; several items now have sub-items (e.g., 10a/10b, 13a-13f, 16a/16b). - Item 2 delegates to a **separate 12-item abstract checklist** (`PRISMA_2020_Abstracts.md`). The statement's own item 2 is that one sentence and nothing more. Run the abstract checklist as its own pass and report its score with its own denominator — a manuscript can satisfy every main-text item and still fail most of the twelve, and a high main-text percentage carries no information about it. - **Item 4 does not require PICO.** This file previously read "using the PICO framework or similar", which is PRISMA *2009* wording; the 2020 item asks only for an explicit statement of the objective(s) or question(s). Do not score a manuscript down for stating its objectives without a PICO frame. - **Item 13b is about missing summary statistics and data conversions**, not about multi-arm studies or multiple outcome measures — this file previously asked for the latter, which is a different requirement and sent assessors looking for the wrong thing. - Item 7 (full search strategy): the complete strategy for at least one database must be provided, either in the manuscript or supplementary materials. - Item 15 (certainty assessment): GRADE is the most common framework; mark as MISSING if no certainty assessment is reported. - Item 24 (registration): PROSPERO is the standard registry for systematic reviews. If not registered, the authors should explicitly state this. - Item 27 (data availability): increasingly required; check if authors state where extracted data and analytic code can be accessed. - For systematic reviews of diagnostic test accuracy, also consider PRISMA-DTA extension. - For reviews that include meta-analysis, items 13d (synthesis model) and 20b (heterogeneity measures) are critical. -
PRISMA_2020_Abstracts.md 5.8 KB
# PRISMA 2020 for Abstracts Checklist **Preferred Reporting Items for Systematic Reviews and Meta-Analyses — abstract checklist** Version: PRISMA 2020 for Abstracts — 12 items, grouped by the sections of a structured abstract. Source: Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. *BMJ* 2021;372:n71 (DOI 10.1136/bmj.n71). Published under CC BY 4.0; item text is reproduced with attribution. Verification: all 12 items were transcribed from, and checked against, the abstract checklist published in that statement. Apply to the **abstract of a systematic review or meta-analysis**. This is a separate instrument from the 27-item main checklist in `PRISMA_2020.md`, not a subset of it: the main checklist devotes a single item (item 2) to the abstract and defers the detail here. A manuscript can satisfy all 27 main items — all 42 with sub-items — and still fail most of these twelve, because the abstract is usually written last, to a word limit, from a template the journal supplied. That is not a hypothetical. In a published audit of 24 systematic reviews and meta-analyses in a radiology journal (Park HY et al. *Korean J Radiol* 2022;23(3):355–369; PMID:35213097), **item 3 (eligibility criteria) and item 12 (registration) were each reported by 0 of 24 abstracts**, and item 5 (risk-of-bias methods) by 3 of 23. Run this checklist as its own pass — a compliance percentage computed over the main checklist says nothing about it. For a DTA review the abstract items still apply; pair them with `PRISMA_DTA.md` for the main text. For a scoping review use the abstract item inside `PRISMA_ScR.md` instead. Word limits are real: where an item genuinely cannot fit, that is a judgement to record, not an item to score PRESENT. ## Reporting items (grouped by abstract section) ### Title | # | Item | What to check is reported | |---|------|---------------------------| | 1 | Title | Identify the report as a systematic review. | ### Background | # | Item | What to check is reported | |---|------|---------------------------| | 2 | Objectives | Provide an explicit statement of the main objective(s) or question(s) the review addresses. | ### Methods | # | Item | What to check is reported | |---|------|---------------------------| | 3 | Eligibility criteria | Specify the inclusion and exclusion criteria for the review. | | 4 | Information sources | Specify the information sources (e.g. databases, registers) used to identify studies and the date when each was last searched. | | 5 | Risk of bias | Specify the methods used to assess risk of bias in the included studies. | | 6 | Synthesis of results | Specify the methods used to present and synthesise results. | ### Results | # | Item | What to check is reported | |---|------|---------------------------| | 7 | Included studies | Give the total number of included studies and participants and summarise relevant characteristics of studies. | | 8 | Synthesis of results | Present results for main outcomes, preferably indicating the number of included studies and participants for each. If meta-analysis was done, report the summary estimate and confidence/credible interval. If comparing groups, indicate the direction of the effect (i.e. which group is favoured). | | 9 | Limitations of evidence | Provide a brief summary of the limitations of the evidence included in the review (e.g. study risk of bias, inconsistency and imprecision). | ### Discussion | # | Item | What to check is reported | |---|------|---------------------------| | 10 | Interpretation | Provide a general interpretation of the results and important implications. | ### Other | # | Item | What to check is reported | |---|------|---------------------------| | 11 | Funding | Specify the primary source of funding for the review. | | 12 | Registration | Provide the register name and registration number. | ## Notes for Assessors - **Score this checklist separately and report its own denominator.** Folding twelve abstract items into a 42-item main-text total hides them: they are a small fraction of the sum and they fail together. - **Item 3 is not the same as item 7.** Naming the studies that ended up included ("24 studies of adults undergoing CT") does not state the criteria that decided inclusion. The item asks what *would have been excluded*. - **Item 4** needs both halves — which sources, and the date each was last searched. A database list with no search date is PARTIAL. - **Item 5** asks for the risk-of-bias *method* (the named tool), not its result. Reporting "most studies were at low risk of bias" without naming QUADAS-2 / RoB 2 / ROBINS-I leaves item 5 MISSING and item 9 partially addressed. - **Item 8** is the one that carries the numbers: a summary estimate with its interval, and, when groups are compared, which group the effect favours. An abstract that reports a pooled estimate with no interval fails this item even though it looks quantitative. - **Item 9 is about the evidence, not the review.** "Only English-language studies were searched" is a limitation of the review process (main checklist item 23c); "the included studies were at high risk of bias for flow and timing" is the limitation this item wants. - **Item 12**: register name *and* number. "The review was registered in PROSPERO" without the CRD number is PARTIAL. If the review was not registered, an explicit statement to that effect is the honest answer — but check whether the target journal's abstract structure has anywhere to put it before scoring it MISSING. - **Structured vs unstructured abstracts**: journals that require an unstructured abstract still expect the content; assess by content, not by the presence of the headings. - Conference abstracts of the same review are out of scope for this checklist. -
PRISMA_DTA.md 10 KB
# PRISMA-DTA Checklist (diagnostic test accuracy systematic reviews) **Preferred Reporting Items for Systematic Reviews and Meta-Analyses — diagnostic test accuracy extension** Version: PRISMA-DTA 2018 — 27 items. Two PRISMA items were **deleted** (15 and 22) and two DTA-specific items were **added** (D1 and D2), so the count matches the original while the numbering skips 15 and 22. A separate abstract checklist accompanies it (below). Source: McInnes MDF, Moher D, Thombs BD, McGrath TA, Bossuyt PM, the PRISMA-DTA Group, et al. Preferred Reporting Items for a Systematic Review and Meta-analysis of Diagnostic Test Accuracy Studies: The PRISMA-DTA Statement. *JAMA* 2018;319(4):388-396 (DOI 10.1001/jama.2017.19163). Correction: *JAMA* 2020;323(6):580 — in item 20, "receiver operating characteristic curve" should read "receiver operating characteristic **plot**"; the correction is reflected below. > **Fidelity and licence.** The statement is published in *JAMA* (© American Medical Association) under > **no open licence** — free to read is not free to reuse. The descriptions below are therefore an > **in-house summary of what each item asks, in our own words — not the published wording**. What *is* > verified against the statement is the structure: which items exist, their numbering (including the > deleted 15 and 22 and the added D1 and D2), their section placement, and their topic labels. > **Complete the official checklist from the statement or EQUATOR for anything you submit.** > **What this file used to be.** Until this revision it listed items 1–27 sequentially — including 15 > and 22, which PRISMA-DTA deletes, and **without D1 and D2, the two DTA-specific items the extension > exists to add**. Its item 2 asked a diagnostic-accuracy review to summarise "interventions". It was > PRISMA 2009 renumbered, presented as PRISMA-DTA. Any assessment made with it scored two items the > guideline had removed and never scored the two it had introduced. ## Main checklist (27 items) ### Title | # | Topic | What to check is reported | |---|-------|---------------------------| | 1 | Title | The report is identified as a systematic review, or meta-analysis, **of DTA studies**. | ### Abstract | # | Topic | What to check is reported | |---|-------|---------------------------| | 2 | Abstract | Assessed with the **abstract checklist below**, not by a single item here. | ### Introduction | # | Topic | What to check is reported | |---|-------|---------------------------| | 3 | Rationale | Why the review was done, set against what is already known. | | **D1** | **Clinical role of index test** | The scientific and clinical background: what the index test is **intended to be used for** and its **role in the care pathway**, plus — where relevant — the reasoning behind a minimally acceptable accuracy (or a minimum difference in accuracy, for a comparative design). **DTA-specific; has no counterpart in PRISMA.** | | 4 | Objectives | An explicit statement of the question in terms of **participants, index test, and target conditions**. | ### Methods | # | Topic | What to check is reported | |---|-------|---------------------------| | 5 | Protocol and registration | Where the protocol can be read (e.g. a web address), and the registration number where one exists. | | 6 | Eligibility criteria | The study characteristics used to decide eligibility — participants, setting, index test, reference standards, target conditions, study design — and the report characteristics (years, language, publication status), **with the reasoning for them**. | | 7 | Information sources | Every source searched, including databases with their coverage dates and any contact with study authors, and the date each was last searched. | | 8 | Search | The **full** search strategy for every database and other source, with any limits, in enough detail to repeat it. | | 9 | Study selection | How studies were selected — screening, eligibility, inclusion in the review, and where applicable inclusion in the meta-analysis. | | 10 | Data collection process | How data were extracted from reports (piloted forms, independent extraction, duplicate extraction) and how data were obtained or confirmed from investigators. | | 11 | Definitions for data extraction | The definitions used when extracting and classifying **target conditions, index tests, reference standards** and other characteristics such as study design and clinical setting. | | 12 | Risk of bias and applicability | The methods used to assess risk of bias in each study **and** concerns about applicability to the review question. | | 13 | Diagnostic accuracy measures | Which accuracy measures are the principal ones (sensitivity, specificity, and so on) **and the unit of assessment** — per patient or per lesion. | | 14 | Synthesis of results | How the data were handled and combined and how between-study variability was described. The statement names six cases to address: multiple definitions of the target condition; multiple positivity thresholds; multiple index-test readers; indeterminate results; grouping and comparing tests; and differing reference standards. | | **D2** | **Meta-analysis** | The statistical methods used for the meta-analysis, where one was performed. **DTA-specific; has no counterpart in PRISMA.** | | 16 | Additional analyses | The methods of any additional analyses — sensitivity, subgroup, meta-regression — **and which of them were prespecified**. | ### Results | # | Topic | What to check is reported | |---|-------|---------------------------| | 17 | Study selection | The count at every stage of selection — records screened, those whose eligibility was assessed, studies entering the review, and studies entering any meta-analysis — together with why studies dropped out at each stage, ideally drawn as a flow diagram. | | 18 | Study characteristics | For each included study, a citation and its key characteristics: participant presentation and prior testing, clinical setting, study design, target-condition definition, index test, reference standard, sample size, and funding source. | | 19 | Risk of bias and applicability | The risk-of-bias and applicability assessment **for each study**. | | 20 | Results of individual studies | For every analysis in every study — each unique combination of index test, reference standard and positivity threshold — the **2 × 2 data (TP, FP, FN, TN)** with accuracy estimates and confidence intervals, ideally shown as a forest plot or a receiver operating characteristic **plot**. | | 21 | Synthesis of results | Test accuracy including its variability; where a meta-analysis was done, the results with confidence intervals. | | 23 | Additional analyses | The results of any additional analyses, including sensitivity, subgroup and meta-regression analyses, and analyses of the index test such as failure rates, the proportion of inconclusive results, and adverse events. | ### Discussion | # | Topic | What to check is reported | |---|-------|---------------------------| | 24 | Summary | The main findings, **with the strength of the evidence**. | | 25 | Limitations | Limitations of the included studies (risk of bias, applicability) **and** limitations of the review process itself, such as incomplete retrieval. | | 26 | Conclusions | A general interpretation against other evidence, and implications for future research and for clinical practice — including the intended use and clinical role of the index test. | ### Funding | # | Topic | What to check is reported | |---|-------|---------------------------| | 27 | Funding | Sources of funding and other support for the review, and the role of the funders. | ## Abstract checklist PRISMA-DTA carries its own abstract checklist; item 2 above defers to it. Score it separately, with its own denominator — the same discipline as `PRISMA_2020_Abstracts.md`. Its items, in our own words: | # | Topic | What to check is reported | |---|-------|---------------------------| | 1 | Title | Identified as a systematic review, or meta-analysis, of DTA studies. | | 2 | Objectives | The research question, with its components — participants, index test, target conditions. | | 3 | Eligibility criteria | The study characteristics used to decide eligibility. | | 4 | Information sources | The key databases searched and the search dates. | | 5 | Risk of bias and applicability | The methods used to assess risk of bias and applicability. | | **A1** | **Synthesis of results** (methods) | The methods used for the data synthesis. **DTA-specific.** | | 6 | Included studies | How many studies and of what type, and the participants and relevant study characteristics, including the reference standard. | | 7 | Synthesis of results (results) | The diagnostic-accuracy results, preferably with the number of studies and participants; accuracy and its variability, and where a meta-analysis was done, summary results with confidence intervals. | The remaining abstract items (interpretation, funding, registration) follow the PRISMA pattern; consult Table 3 of the statement for their exact wording before scoring an abstract you intend to report. ## Notes for Assessors - **Do not score items 15 or 22.** PRISMA-DTA deletes them. An assessment that reports 27 sequential items has not used this instrument. - **D1 and D2 are the point of the extension.** A review that never states the index test's intended clinical role (D1), or that meta-analyses without reporting the statistical methods used (D2), is missing the DTA-specific requirements — and these are exactly the items a generic PRISMA checklist cannot surface. - **Item 13's unit of assessment** (per patient vs per lesion) is the item most often skipped in imaging reviews, and it changes what every downstream accuracy estimate means. - **Item 20 wants the 2 × 2 cells**, not just sensitivity and specificity. Without TP/FP/FN/TN a reader cannot recompute the estimates or pool them. - For the risk-of-bias assessment referenced by items 12 and 19, see `QUADAS2.md`. - For a review that is not diagnostic test accuracy, use `PRISMA_2020.md` with `PRISMA_2020_Abstracts.md`. -
PRISMA_P.md 5.2 KB
# PRISMA-P Checklist **Preferred Reporting Items for Systematic review and Meta-Analysis Protocols** Version: PRISMA-P 2015 Source: Shamseer L et al. BMJ 2015;349:g7647. Source: Shamseer L, Moher D, Clarke M, Ghersi D, Liberati A, Petticrew M, et al. Preferred reporting items for systematic review and meta-analysis protocols (PRISMA-P) 2015: elaboration and explanation. *BMJ* 2015;350:g7647 (DOI 10.1136/bmj.g7647). Licence: Crossref returns no Creative Commons licence for this article. Verification: all 26 items were compared against Table 3 of the CC BY *Systematic Reviews* co-publication of the same statement (Europe PMC full text, PMC4320440); 26/26 match, with no item missing and none invented. ## Checklist Items (17 items, 26 sub-items) ### Administrative Information | # | Item | Description | |---|------|-------------| | 1a | Identification | Identify the report as a protocol of a systematic review | | 1b | Update | If the protocol is for an update of a previous systematic review, identify as such | | 2 | Registration | If registered, provide the name of the registry (such as PROSPERO) and registration number | | 3a | Contact | Provide name, institutional affiliation, e-mail address of all protocol authors; provide physical mailing address of corresponding author | | 3b | Contributions | Describe contributions of protocol authors and identify the guarantor of the review | | 4 | Amendments | If the protocol represents an amendment of a previously completed or published protocol, identify as such and list changes; otherwise, state plan for documenting important protocol amendments | | 5a | Sources | Indicate sources of financial or other support for the review | | 5b | Sponsor | Provide name for the review funder and/or sponsor | | 5c | Role of sponsor or funder | Describe roles of funder(s), sponsor(s), and/or institution(s), if any, in developing the protocol | ### Introduction | # | Item | Description | |---|------|-------------| | 6 | Rationale | Describe the rationale for the review in the context of what is already known | | 7 | Objectives | Provide an explicit statement of the question(s) the review will address with reference to participants, interventions, comparators, and outcomes (PICO) | ### Methods | # | Item | Description | |---|------|-------------| | 8 | Eligibility criteria | Specify the study characteristics (such as PICO, study design, setting, time frame) and report characteristics (such as years considered, language, publication status) to be used as criteria for eligibility for the review | | 9 | Information sources | Describe all intended information sources (such as electronic databases, contact with study authors, trial registers or other grey literature sources) with planned dates of coverage | | 10 | Search strategy | Present draft of search strategy to be used for at least one electronic database, including planned limits, such that it could be repeated | | 11a | Data management | Describe the mechanism(s) that will be used to manage records and data throughout the review | | 11b | Selection process | State the process that will be used for selecting studies (such as two independent reviewers) through each phase of the review (that is, screening, eligibility and inclusion in meta-analysis) | | 11c | Data collection process | Describe planned method of extracting data from reports (such as piloting forms, done independently, in duplicate), any processes for obtaining and confirming data from investigators | | 12 | Data items | List and define all variables for which data will be sought (such as PICO items, funding sources), any pre-planned data assumptions and simplifications | | 13 | Outcomes and prioritization | List and define all outcomes for which data will be sought, including prioritization of main and additional outcomes, with rationale | | 14 | Risk of bias in individual studies | Describe anticipated methods for assessing risk of bias of individual studies, including whether this will be done at the outcome or study level, or both; state how this information will be used in data synthesis | | 15a | Data synthesis | Describe criteria under which study data will be quantitatively synthesised | | 15b | Data synthesis | If data are appropriate for quantitative synthesis, describe planned summary measures, methods of handling data and methods of combining data from studies, including any planned exploration of consistency (such as I^2, Kendall's tau) | | 15c | Data synthesis | Describe any proposed additional analyses (such as sensitivity or subgroup analyses, meta-regression) | | 15d | Data synthesis | If quantitative synthesis is not appropriate, describe the type of summary planned | | 16 | Meta-bias(es) | Specify any planned assessment of meta-bias(es) (such as publication bias across studies, selective reporting within studies) | | 17 | Confidence in cumulative evidence | Describe how the strength of the body of evidence will be assessed (such as GRADE) | --- ## Notes for Assessors - This checklist should be read in conjunction with the PRISMA-P Explanation and Elaboration document - Amendments to a review protocol should be tracked and dated - The copyright for PRISMA-P is held by the PRISMA-P Group and is distributed under a Creative Commons Attribution Licence 4.0 -
PRISMA_ScR.md 9 KB
# PRISMA-ScR Checklist (scoping reviews) **Preferred Reporting Items for Systematic reviews and Meta-Analyses extension for Scoping Reviews** Version: PRISMA-ScR 2018 — the checklist keeps PRISMA's **1–27 numbering**. Of those 27 rows, **5 are "Not applicable for scoping reviews"** (13, 15, 16, 22, 23) and **2 are optional, phrased "If done"** (**12** and **19**, critical appraisal), leaving **20 essential items**. Source: Tricco AC, Lillie E, Zarin W, O'Brien KK, Colquhoun H, Levac D, et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. *Ann Intern Med* 2018;169(7):467-473 (DOI 10.7326/M18-0850). EQUATOR Network. Scoping-review conduct guidance: JBI (Peters et al.); Arksey & O'Malley; Levac et al. Apply when the manuscript is a **scoping review** — a review that *maps* the breadth of evidence on a topic, characterises concepts and definitions, and identifies gaps, rather than answering a focused effectiveness or accuracy question (that is a systematic review → `PRISMA_2020.md` / `PRISMA_DTA.md`). For the design/conduct review of the same study, pair with the SC1–SC8 domain probes in `peer-review` / `self-review` `references/domain-probes/scoping_review.md`; for the flow diagram use `make-figures` (PRISMA-style). > **Fidelity and licence.** Published in *Annals of Internal Medicine* (© American College of > Physicians) under **no open licence**; one secondary source lists CC BY-NC-SA 3.0, which is also not > permissively reusable. The descriptions below are an **in-house summary of item intent in our own > words — not the published wording**. The structure — the 1–27 numbering, which rows are N/A, which > are optional, and the section grouping — is **verified against the statement's table**. Consult the > published article for exact item text, and complete the official checklist for anything you submit. > **What this file used to be.** It numbered the items **1–22**: it had deleted the five "Not > applicable" rows and closed the gaps, so every item from 13 onward carried the wrong number — 10 of > 22. Critical appraisal within sources of evidence sat at 16 instead of **19**; Limitations at 20 > instead of **25**; Funding at 22 instead of **27**. The header and the assessor notes named the > optional pair as "12 and 16", but 16 is *Additional analyses* and is **N/A**, not optional. Item > numbers are how authors, reviewers and editors refer to a checklist, so a renumbered instrument > silently breaks every reference to it. ## Reporting items (grouped by section) ### Title | # | Item | What to check is reported | |---|------|---------------------------| | 1 | Title | The report is identified as a **scoping review** in the title. | ### Abstract | # | Item | What to check is reported | |---|------|---------------------------| | 2 | Structured summary | A structured abstract that, as relevant, states the background and objectives, the eligibility criteria and the evidence sources searched, the **charting** approach, and the headline results and conclusions tied to the review questions. | ### Introduction | # | Item | What to check is reported | |---|------|---------------------------| | 3 | Rationale | The rationale situated against existing knowledge, and **why the question suits a scoping (mapping) approach** rather than a systematic review. | | 4 | Objectives | Explicit questions/objectives stated with their key elements — **Population, Concept, Context (PCC)** or equivalent — used to frame the review. | ### Methods | # | Item | What to check is reported | |---|------|---------------------------| | 5 | Protocol and registration | Whether an a-priori **protocol** exists, where it can be accessed, and any registration details. (Note: PROSPERO does **not** register scoping reviews; OSF/Figshare or a published protocol is used.) | | 6 | Eligibility criteria | Characteristics used as eligibility criteria (years, language, publication status, source types) **with a rationale**. | | 7 | Information sources | All information sources searched (databases with coverage dates, grey-literature sources, author contact) and the date the most recent search was run. | | 8 | Search | A reproducible search strategy for **at least one database** (search terms and any limits), detailed enough to be repeated by another team. | | 9 | Selection of sources of evidence | The process for selecting **sources of evidence** (screening and eligibility) into the review. | | 10 | Data charting process | The **data-charting** methods (calibrated/team-tested form; whether charting was independent/duplicate; iterative refinement) and any process for obtaining/confirming data from investigators. "Charting," not "extraction." | | 11 | Data items | The data items charted — every variable sought — plus any assumptions or simplifications made. | | **12** | Critical appraisal of individual sources of evidence — **OPTIONAL** | *If done*, the rationale for critical appraisal, the methods used, and how the information was used in any data synthesis. Scoping reviews are **not required** to appraise; "critical appraisal" is used instead of "risk of bias." | | 13 | Summary measures | **Not applicable for scoping reviews.** | | 14 | Synthesis of results | The methods for **handling and summarising (charting/mapping)** the charted data. | | 15 | Risk of bias across studies | **Not applicable for scoping reviews.** | | 16 | Additional analyses | **Not applicable for scoping reviews.** | ### Results | # | Item | What to check is reported | |---|------|---------------------------| | 17 | Selection of sources of evidence | Numbers screened, assessed for eligibility, and included, with reasons for exclusion at each stage — ideally a **flow diagram**. | | 18 | Characteristics of sources of evidence | Per-source characteristics for which data were charted, with citations. | | **19** | Critical appraisal within sources of evidence — **OPTIONAL** | *If done* (see item 12), the critical-appraisal data per source. | | 20 | Results of individual sources of evidence | For each included source, the relevant charted data that relate to the review questions/objectives. | | 21 | Synthesis of results | The charting results **summarised/mapped** in relation to the review questions/objectives — a **map/characterisation, not a pooled effect estimate**. | | 22 | Risk of bias across studies | **Not applicable for scoping reviews.** | | 23 | Additional analyses | **Not applicable for scoping reviews.** | ### Discussion | # | Item | What to check is reported | |---|------|---------------------------| | 24 | Summary of evidence | The main results summarised (overview of concepts, themes, evidence types), linked to the questions/objectives and their relevance to key groups. | | 25 | Limitations | Limitations of the **scoping-review process**. | | 26 | Conclusions | A general interpretation against the questions/objectives, with potential implications and/or next steps (e.g. whether a systematic review is warranted). | ### Funding | # | Item | What to check is reported | |---|------|---------------------------| | 27 | Funding | Sources of funding for the included sources of evidence **and** for the scoping review, and the funders' role. | --- ## Notes for Assessors - **Score against 1–27, not 1–20.** The five N/A rows stay in the table so the numbering matches the guideline; mark them N/A rather than deleting them. The denominator for compliance is the 20 essential items, plus the two optional items only where the review says it appraised. - **The optional pair is 12 and 19**, not 12 and 16. Item 16 is *Additional analyses* and is N/A. Why they are optional: a scoping review may include documents, blogs, websites, interviews and opinions, and is not required to examine risk of bias in what it includes. - The **highest-yield** checks (where scoping reviews most often go wrong): **item 4** (objectives framed as PCC mapping, not a focused PICO effectiveness question — a study asking "is X effective/accurate?" should be a systematic review), **item 21** (results presented as a **map** — counts, categories, themes, evidence gaps — **not** a pooled OR/RR/AUC or a definitive effectiveness conclusion), and **items 12/19** (do not penalise a scoping review for omitting critical appraisal, and do not let it claim GRADE-style certainty it never derived). - A "scoping review" that reports pooled effect estimates or a meta-analysis, or concludes an intervention is effective or a test accurate, has overstepped its design — its synthesis (21) and conclusions (26) should be downgraded to descriptive mapping, or the study reframed as a systematic review. - **Terminology is deliberate and the statement says so**: **charting** rather than extraction, **sources of evidence** rather than studies, and **critical appraisal** rather than risk of bias — the last because risk of bias suits systematic reviews of interventions, while a scoping review takes in quantitative and qualitative research, expert opinion and policy documents. A scoping review claiming PROSPERO registration is in error; PROSPERO excludes them. -
PROBAST.md 5 KB
# PROBAST Assessment Guide Prediction model Risk Of Bias ASsessment Tool. Version: PROBAST 2019 — 20 signalling questions across 4 domains (2 / 3 / 6 / 9). Source: Wolff RF, Moons KGM, Riley RD, Whiting PF, Westwood M, Collins GS, et al. PROBAST: A Tool to Assess the Risk of Bias and Applicability of Prediction Model Studies. *Ann Intern Med* 2019;170(1):51-58 (PMID 30596875; DOI 10.7326/M18-1376). Explanation and Elaboration: *Ann Intern Med* 2019;170(1):W1-W33. > **Fidelity and licence.** PROBAST is published in *Annals of Internal Medicine* (© American College > of Physicians) and carries **no open licence**. The signalling questions below are therefore an > **in-house summary of what each question asks — not the published wording**. The numbering, > the domain structure and the count (2 / 3 / 6 / 9) are verified against the statement; the phrasing > is ours. **Complete the official tool from probast.org or the statement for any assessment you > report.** Use this file to organise a first pass. > **PROBAST 2019 is superseded.** PROBAST+AI (2025) replaces it for all new assessments, covering > regression and AI/ML models alike — see `PROBAST_AI.md`. Use this file only when appraising against > the 2019 instrument specifically (for example, reproducing an earlier review). ## Structure PROBAST assesses 4 domains, each for Risk of Bias AND Applicability. - **Signalling questions**: Yes / Probably yes / No / Probably no / No information - **Domain judgment**: Low / High / Unclear - **Overall judgment**: High if any domain is high; Low only if all domains are low ## Domain 1: Participants ### Signalling questions (risk of bias) — what each asks 1.1 Whether the data source suits the question — a cohort, a randomised trial, or nested case–control data, rather than a design that distorts the sampling. 1.2 Whether every inclusion and exclusion applied to participants was appropriate. ### Applicability - Do the participants and setting match the review question? ## Domain 2: Predictors ### Signalling questions (risk of bias) — what each asks 2.1 Whether predictors were defined and assessed the same way for every participant. 2.2 Whether predictors were assessed without knowledge of the outcome. 2.3 Whether every predictor is available at the moment the model is meant to be used. ### Applicability - Do the predictors, their assessment, and timing match the review question? ## Domain 3: Outcome ### Signalling questions (risk of bias) — what each asks 3.1 Whether the outcome was determined by an appropriate method. 3.2 Whether the outcome definition was prespecified or a standard one. 3.3 Whether predictors were kept out of the outcome definition. 3.4 Whether the outcome was defined and determined the same way for every participant. 3.5 Whether the outcome was determined without knowledge of predictor information. 3.6 Whether the interval between predictor assessment and **outcome determination** was appropriate — the question is about when the outcome was *determined*, not merely when it occurred. ### Applicability - Does the outcome and its definition/timing match the review question? ## Domain 4: Analysis ### Signalling questions (risk of bias) — what each asks 4.1 Whether the number of participants with the outcome was reasonable. 4.2 Whether continuous and categorical predictors were handled appropriately. 4.3 Whether every enrolled participant was included in the analysis. 4.4 Whether participants with missing data were handled appropriately. 4.5 Whether selection of predictors on the basis of univariable analysis was avoided. 4.6 Whether complexities in the data were accounted for — the statement names **censoring, competing risks, and the sampling of control participants** as the cases to look for. 4.7 Whether relevant measures of model performance were evaluated appropriately. 4.8 Whether **overfitting, underfitting, and optimism** in model performance were accounted for. All three, not overfitting alone. 4.9 Whether the predictors and their assigned weights in the final model correspond to the **results from** the reported multivariable analysis. ### For validation studies (additional) - Were the model and its performance evaluated appropriately? ## For AI / machine-learning models Use **`PROBAST_AI.md`** (PROBAST+AI 2025; Moons KGM et al. *BMJ* 2025;388:e082505, DOI 10.1136/bmj-2024-082505), which carries the instrument's own 16 development and 18 evaluation signalling questions. Do not improvise AI addenda on top of the 2019 questions: an earlier version of this file listed four invented bullets under a heading that implied they were part of the extension, which they were not. ## When to Use - Diagnostic prediction models (e.g., AI classifiers for imaging findings) - Prognostic prediction models (e.g., risk scores, survival prediction) - Both development AND validation studies - For any model developed with machine learning or deep learning, and for new assessments generally, use PROBAST+AI (`PROBAST_AI.md`) rather than this file -
PROBAST_AI.md 5.9 KB
# PROBAST+AI Assessment Guide Prediction model Risk Of Bias ASsessment Tool — updated for AI/ML methods. Reference: Moons KGM et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ 2025;388:e082505. doi: 10.1136/bmj-2024-082505. Website: https://www.probast.org Version: PROBAST+AI (2025) — 16 development / 18 evaluation signalling questions Licence: CC BY-NC 4.0 (non-commercial) — confirmed via Crossref (DOI 10.1136/bmj-2024-082505). Not redistributable verbatim under this repository's MIT licence. Verification: the development and evaluation signalling-question counts and the four-step process were checked against the statement and its supplements during this audit. ## Purpose PROBAST+AI is the updated version of PROBAST (2019) that extends the original tool to cover prediction models developed using **machine learning and artificial intelligence** methods, in addition to traditional regression-based models. It replaces PROBAST-2019 for all new assessments. ## What Changed from PROBAST-2019 - Updated signalling questions to address AI/ML-specific methodological considerations - Expanded from 20 to **16 targeted signalling questions** for model development and **18 for model evaluation** - Covers all prediction model types: regression, ML, DL, ensemble methods - Applicable regardless of whether the model uses statistical or AI/ML techniques - Original PROBAST Explanation and Elaboration document still provides useful background ## Structure PROBAST+AI has two distinct parts: 1. **Model development** assessment: quality and applicability 2. **Model evaluation** assessment: risk of bias and applicability Both parts assess 4 domains: - Domain 1: Participants - Domain 2: Predictors - Domain 3: Outcome - Domain 4: Analysis Signalling questions answered: Yes / Probably Yes / No / Probably No / No Information ## Four Steps of Assessment 1. Specify the intended purpose of the prediction model assessment or systematic review 2. Classify the type of prediction model study (development or evaluation or both) 3. Assess quality/applicability (development) or risk of bias/applicability (evaluation) for each domain 4. Assess overall quality (development) or risk of bias (evaluation) ## Domain 1: Participants ### For Model Development (Quality) - Were participants representative of the target population? - Was the study setting appropriate? - Were inclusion/exclusion criteria clearly defined? ### For Model Evaluation (Risk of Bias) - Were participants selected appropriately for the evaluation? - Were participants representative of the intended target population? ### Applicability - Do the participants match the intended target population for the model? ## Domain 2: Predictors ### For Model Development (Quality) - Were predictors defined and assessed in a standardized way? - Were predictors available at the intended moment of use? - Were all candidate predictors assessed appropriately? ### For Model Evaluation (Risk of Bias) - Were predictors assessed in the same way as in the development study? - Were predictors assessed without knowledge of the outcome? ### Applicability - Do the predictor definitions and assessments match the intended use? ## Domain 3: Outcome ### For Model Development (Quality) - Was the outcome clearly defined? - Was the outcome determined appropriately? - Was the time horizon appropriate? ### For Model Evaluation (Risk of Bias) - Was the outcome determined in the same way as in development? - Was the outcome determined without knowledge of predictor information? ### Applicability - Does the outcome definition and assessment match the intended target? ## Domain 4: Analysis ### For Model Development (Quality — key AI/ML considerations) - Was the sample size adequate for the modeling approach? - Were missing data handled appropriately? - Was the choice of algorithm/method appropriate? - Were hyperparameters tuned appropriately (for ML/DL models)? - Was overfitting addressed (regularization, early stopping, cross-validation)? - Was model performance assessed using appropriate metrics? - Was internal validation performed adequately (bootstrapping, cross-validation)? - Were relevant performance measures reported (discrimination, calibration)? ### For Model Evaluation (Risk of Bias — key AI/ML considerations) - Was the evaluation dataset independent from the development dataset? - Were appropriate performance measures used (discrimination AND calibration)? - Was the statistical analysis appropriate? - Were model updates or recalibration described if performed? ### Applicability - Does the analysis approach and validation strategy support the intended clinical use? ## Overall Judgment ### For Model Development: Overall Quality - **High quality**: Low concern across all domains - **Unclear quality**: Unclear in at least one domain - **Low quality**: High concern in at least one domain ### For Model Evaluation: Overall Risk of Bias - **Low risk**: Low risk in all domains - **Unclear risk**: Unclear in at least one domain, no high risk in any - **High risk**: High risk in at least one domain ## Key Differences from Original PROBAST-2019 | Aspect | PROBAST-2019 | PROBAST+AI | |--------|-------------|------------| | Scope | Regression models | Regression + AI/ML models | | Development assessment | Risk of bias | **Quality** (more appropriate term) | | Evaluation assessment | Risk of bias | Risk of bias (unchanged) | | AI-specific items | None | Hyperparameter tuning, overfitting mitigation, algorithm choice | | Signalling questions | 20 per assessment | 16 (development) / 18 (evaluation) | ## When to Use - Systematic reviews of prediction model studies (development, validation, or both) - Applies to ALL prediction models regardless of algorithm (logistic regression, random forest, deep learning, etc.) - Replaces PROBAST-2019 for new assessments - Use alongside TRIPOD+AI for reporting quality assessment -
QUADAS2.md 8.1 KB
# QUADAS-2 Assessment Guide Quality Assessment of Diagnostic Accuracy Studies, version 2. Version: QUADAS-2 (2011) — 4 domains, **10 signalling questions** (3 / 2 / 2 / 3), risk of bias for every domain and applicability for the first three only. Source: Whiting PF, Rutjes AWS, Westwood ME, Mallett S, Deeks JJ, Reitsma JB, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. *Ann Intern Med* 2011;155(8):529-536 (DOI 10.7326/0003-4819-155-8-201110180-00009). The tool itself, training material and worked examples were distributed from www.quadas.org; that domain no longer resolves, and the QUADAS group now distributes its tools from https://www.bristol.ac.uk/population-health-sciences/projects/quadas/ > **Fidelity and licence.** QUADAS-2 is published in *Annals of Internal Medicine* (© American College > of Physicians) under **no open licence**. **Verified against the published article**: the four > domains, the 3 / 2 / 2 / 3 split of signalling questions, what each question asks, the answer > options, the judgement levels, and the restriction of applicability to the first three domains all > match. The descriptions below state what each question asks **in our own words** rather than > reproducing the published wording. **Complete the official QUADAS-2 form from www.quadas.org for > any assessment you report.** > > **QUADAS-3 has superseded this tool.** The QUADAS group states that "QUADAS-3 is the current > iteration of the QUADAS tool and is the current recommended version" — it extends QUADAS-2 by > introducing an explicit ideal-test-accuracy-trial comparator and by moving assessment **from the > study level to the level of individual accuracy estimates**. This file documents QUADAS-2, which > remains what most published reviews used. For a new review use **`QUADAS3.md`**; > `QUADAS_C.md` should now be paired with QUADAS-3 rather than with this tool. ## How the tool is applied — four phases QUADAS-2 is not a fixed questionnaire. The statement specifies four phases, and skipping the second is the most common misuse: 1. **State the review question** — patients, index test, reference standard and target condition. Describe patients by setting, the intended use of the index test, presentation, and prior testing, because accuracy depends on where in the diagnostic pathway the test sits. 2. **Tailor the tool to the review.** Add or omit signalling questions and write review-specific guidance on how each is to be judged. (For an objective index test, for example, the question about blinding the interpreter to the reference standard may not apply.) Avoid piling on extra questions. At least two people should pilot the tailored tool; refine it if agreement is poor. 3. **Draw the study's flow diagram** — use the published one, or draw it yourself if it is absent or inadequate. It need not be reported; it exists to make the flow-and-timing judgements possible. 4. **Judge bias and applicability**, recording the information each judgement rests on. ## Answers and judgements - **Signalling questions**: Yes / No / Unclear, phrased so that **Yes indicates low risk of bias**. - **Risk of bias**: Low / High / Unclear. - **Applicability concern** (domains 1–3 only): Low / High / Unclear. - **Unclear is only for insufficient reporting**, not for a difficult judgement. **How a "No" is handled.** If every signalling question in a domain is Yes, risk of bias can be judged low. A **No does not automatically make the domain High** — it establishes that *potential for bias exists*, and the reviewer then applies the review-specific guidance written in phase 2 to reach the judgement. A file or workflow that maps "any No → High" has removed the judgement the tool asks for. ## Domain 1: Patient Selection **Risk of bias — could patient selection have introduced bias?** 1. Whether enrolment took a consecutive or random sample of eligible patients. 2. Whether a case–control design was avoided. 3. Whether the study avoided inappropriate exclusions. *Why it matters*: enrolling patients whose diagnosis is already confirmed, or excluding the difficult-to-diagnose, inflates apparent accuracy; excluding patients with obvious signs of the target condition can deflate it. **Applicability** — whether the included patients and the setting differ from the review question (severity, demographics, comorbidity and differential diagnosis, setting, prior testing). ## Domain 2: Index Test **Risk of bias — could the conduct or interpretation of the index test have introduced bias?** 1. Whether the index test was interpreted without knowledge of the reference standard result. 2. Whether a threshold, if one was used, was prespecified. *Why it matters*: this is the diagnostic equivalent of blinding, and its force depends on how subjective the index test is and on the order of testing. Choosing the threshold after the fact to maximise sensitivity or specificity overstates performance that will not hold in a new sample. **Applicability** — whether the index test, how it was carried out, or how it was interpreted differs from the review question. ## Domain 3: Reference Standard **Risk of bias — could the reference standard, its conduct or its interpretation have introduced bias?** 1. Whether the reference standard is likely to classify the target condition correctly. 2. Whether the reference standard was interpreted without knowledge of the index test result. *Why it matters*: accuracy estimates assume the reference standard is correct, so that any disagreement is the index test's error. **Applicability** — whether the target condition **as the reference standard defines it** differs from the one in the review question. ## Domain 4: Flow and Timing **Risk of bias — could patient flow have introduced bias?** 1. Whether the interval between index test and reference standard was appropriate. 2. Whether all patients received the same reference standard. 3. Whether all patients were included in the analysis. *Why it matters*: an interval long enough for the condition to change causes misclassification, and how long is too long depends on the condition. Verifying only some patients, or verifying different patients with different reference standards, biases the estimate. Patients lost between enrolment and the 2 × 2 table differ systematically from those who remain. **This domain has no applicability judgement.** ## Reporting the assessment - **Do not produce a summary quality score.** The statement is explicit about this, for the reasons well established in the literature on quality scores. - A study low on every domain may be called low risk of bias, or low concern for applicability, overall. High or unclear on one or more domains means it is at risk of bias, or of concern. - At minimum, report the assessment across all included studies — how many were low, high or unclear per domain — and consider highlighting signalling questions on which studies consistently do badly. - Restricting the primary analysis to low-risk studies is legitimate, but it is often better to include everything and then investigate heterogeneity — by subgroup, sensitivity analysis, or by entering domains as covariates in meta-regression. - QUADAS-2 does **not** cover studies comparing multiple index tests. The development group considered it and concluded the evidence base was insufficient. ## Common Issues in DTA Studies - **Partial verification bias**: Not all patients receive the reference standard (especially when invasive, e.g., biopsy) - **Differential verification**: Different reference standards used for different patients - **Incorporation bias**: Index test forms part of the reference standard - **Review bias**: Knowledge of index test results influences reference standard interpretation - **Clinical review bias**: Additional clinical information available during index test interpretation - **Uninterpretable results**: Exclusion of technically inadequate or indeterminate results ## Related - `PRISMA_DTA.md` items 12 and 19 are the reporting counterparts: the methods used for this assessment, and its result presented **for each study**. -
QUADAS3.md 12.9 KB
# QUADAS-3 Assessment Guide Quality Assessment of Diagnostic Accuracy Studies, version 3 — **the current recommended version**. Version: QUADAS-3 tool v1.2 — 6 phases, 4 domains, **20 signalling questions** (4 / 4 / 8 / 4). Source: Whiting PF, Tomlinson E, Rutjes AWS, Davenport C, Yang B, Westwood M, et al. QUADAS-3: a revised tool for the quality assessment of diagnostic test accuracy studies. *Ann Intern Med* 2026;179(4):548-555 (DOI 10.7326/ANNALS-25-02104). The tool itself, the Explanation & Elaboration report and an introductory video are distributed by the QUADAS group at https://www.bristol.ac.uk/population-health-sciences/projects/quadas/quadas-3/ Explanation & Elaboration: Davenport CF, Rutjes AWS, Mallett S, Tomlinson E, Yang B, et al. QUADAS-3 explanation and elaboration: guidance for quality assessment of diagnostic test accuracy studies. *Ann Intern Med* 2026;179(4):e2504943 (DOI 10.7326/ANNALS-25-04943). Resource site: **www.quadas.info** (the older www.quadas.org no longer resolves). > **Fidelity and licence.** QUADAS-3 is published in *Annals of Internal Medicine* (© American > College of Physicians) under **no open licence** — Crossref returns only ACP's text-and-data-mining > policy. The descriptions below state what each question asks **in our own words** rather than > reproducing the published wording. **Complete the official QUADAS-3 form (`QUADAS-3 1.2.docx`) > from the page above for any assessment you report, and read the E&E report before using it.** > > Verification: the six phases and when each is completed, the four domains, all 20 signalling > questions, the response options, the domain-level rule, which domains carry an applicability > judgement, and the overall-judgement rules were compared against **the official tool document > v1.2** distributed by the QUADAS group. All matched. The *Using QUADAS-C with QUADAS-3*, > *Tailoring* and *no "moderate" grade* sections below are taken from the **Explanation and > Elaboration paper**, read directly. ## QUADAS-3 supersedes QUADAS-2 The QUADAS group states QUADAS-3 "is the current version of QUADAS and the tool that we recommend." For a new review, use this file. `QUADAS2.md` documents the 2011 tool, which is what most published reviews used and what you will still be reading in them. What changed, in the group's own framing: | | QUADAS-2 | QUADAS-3 | |---|---|---| | Unit of assessment | the **study** | **each set of accuracy estimates** | | Comparator for judging | implicit | an explicit **ideal test accuracy trial**, defined per synthesis question | | Synthesis questions | one, implicit | **multiple, defined up front** | | Overall judgment | none | **a formal phase (6)** | | Phases | 4 | **6** | | Domains | Patient Selection, Index Test, Reference Standard, Flow and Timing | **Participants, Index Test, Target Condition, Analysis** | | Signalling questions | 10 (3/2/2/3) | **20 (4/4/8/4)** | | Third judgement level | "unclear" | **"insufficient information" (II)** | Note the domain rename: QUADAS-2's *Flow and Timing* is gone. Timing moved into **Target Condition** (the index-test-to-reference-standard interval), and participant exclusions, missing data and the unit of analysis moved into the new **Analysis** domain. **Comparative accuracy reviews**: the group recommends using **QUADAS-C in addition to QUADAS-3** (`QUADAS_C.md`). QUADAS-C was written against QUADAS-2 and needs adaptation — see the next section. ## Using QUADAS-C with QUADAS-3 For comparative accuracy studies — where two or more index tests are compared — the guideline says to use **QUADAS-C alongside QUADAS-3**, because such studies carry additional sources of bias (confounding between tests, and interference of one test with another). QUADAS-C was written as an extension of QUADAS-2, so it needs adapting. **An updated QUADAS-C is in development**; what follows are the E&E's own *preliminary* modifications, not a finished tool. ### What you assess The unit is a **comparative measure** — for example the difference in sensitivity or in specificity between two tests. Most primary studies report only the separate estimate for each index test, so **you will usually have to compute the comparative measure yourself.** Specify which estimate a QUADAS-C assessment refers to, exactly as you do in QUADAS-3 phase 4. ### Domains are renamed onto QUADAS-3's | QUADAS-C (as published) | becomes | |---|---| | Patient Selection | **Participants** | | Index Test | Index Test *(unchanged)* | | Reference Standard | **Target Condition** | | Flow and Timing | **Analysis** | ### Three signalling questions change (E&E Table 8) | QUADAS-C question | Change | |---|---| | **C3.2** Did the reference standard avoid incorporating any of the index tests? | **Removed** — it overlaps QUADAS-3 item 3.4 | | **C4.2** Was there an appropriate interval between the index tests? | **Moved to domain 2** (Index Test) | | **C4.3** Was the same reference standard used for all index tests? | **Moved to domain 3** (Target Condition) | Everything else carries over unchanged. ### Answers and judgements follow QUADAS-3 Signalling questions take **Y / PY / PN / N / NI**; each domain is judged **low / high / insufficient information**. The overall judgement for the comparative estimate: - **low** if all domains are low - **high** if at least one domain is high - **insufficient information** if at least one domain is II and none is high ## The six phases | Phase | What | How often | |---|---|---| | 1 | State the systematic review synthesis question(s) | once per review | | 2 | Define the **ideal test accuracy trial** for each synthesis question | once per review | | 3 | Draw a flow diagram | once per study | | 4 | Identify which accuracy estimates to assess | once per study | | 5 | Assess risk of bias and applicability | for each selected estimate | | 6 | Overall judgment | for each selected estimate | Phases 1 and 2 are review-level and **belong in the review protocol**. Phases 3–4 are study-level. Phases 5–6 run once per selected set of estimates. **Phase 1** — a review may address more than one synthesis question. Specify each with its population, index test(s) and target condition, and pre-specify them in the protocol. **Phase 2** — the ideal test accuracy trial is the study that would answer the synthesis question with minimum bias and maximum applicability. Define it per question across: objective, participants, index test(s), definition of the target condition, and analysis. Every later judgement is made **against this trial**, not against an unstated ideal. **Phase 4** — a single primary study usually yields several two-by-two tables. Assess only the estimates relevant to a synthesis question. Record, for each: the synthesis question, the numerical result, participants, index test and threshold, target condition, reference standard, unit of analysis, and the analysis method. After the first estimate, **only the domains whose characteristics differ between estimates need reassessing**. ## Phase 5 — signalling questions Signalling questions: **Y / PY / PN / N / NI**. Domain risk-of-bias judgement: **low / high / insufficient information (II)**. ### Domain 1: Participants (4) | # | Signalling question | |---|---------------------| | 1.1 | Was a single-gate design used? | | 1.2 | Were participants prospectively enrolled? | | 1.3 | Was a consecutive or random sample of participants included? | | 1.4 | Is the study group a representative sample of the intended-use population? | *Applicability*: does the included population match the ideal trial's? Participants who dropped out or were excluded because they did not receive the index test or the reference standard belong in **Domain 4 (Analysis)**, not here. ### Domain 2: Index Test (4) | # | Signalling question | |---|---------------------| | 2.1 | Was the index test conducted and interpreted according to the recommended instructions? | | 2.2 | Were the index test results interpreted without knowledge of the reference standard results? | | 2.3 | Were the index test results interpreted with the same information that would be available when the test is used in practice? | | 2.4 | If an index test threshold was used, was it standard or pre-specified? | *Applicability*: does the index test, its conduct and its interpretation match the ideal trial's? 2.3 is new relative to QUADAS-2 and cuts both ways — a reader given **more** information than they would have in practice is as much a problem as one given less. ### Domain 3: Target Condition (8) | # | Signalling question | |---|---------------------| | 3.1 | Does the reference standard adequately identify those with and without the target condition? | | 3.2 | Was the target condition assessed in all participants? | | 3.3 | Was the target condition assessed in the same way in all participants? | | 3.4 | Did the reference standard avoid incorporating the index test? | | 3.5 | Was the reference standard conducted and interpreted according to the recommended instructions? | | 3.6 | Were the reference standard results interpreted without knowledge of the index test results? | | 3.7 | If a reference standard threshold was used, was it standard or pre-specified? | | 3.8 | Was there an appropriate time interval between index test and reference standard? | *Applicability*: does the target condition as defined by the reference standard match the ideal trial's? This domain absorbs QUADAS-2's Reference Standard domain **and** its verification and timing questions. 3.2 and 3.3 are partial and differential verification; 3.8 is the interval that used to sit in Flow and Timing. ### Domain 4: Analysis (4) | # | Signalling question | |---|---------------------| | 4.1 | Were all participants included in the analysis? | | 4.2 | Were missing data handled appropriately? | | 4.3 | Does the unit of analysis match the ideal test accuracy trial? | | 4.4 | Were the estimates of sensitivity and specificity calculated appropriately? | **No applicability judgement** — applicability is assessed for the first three domains only. 4.3 is where a lesion-level or sample-level analysis meets a participant-level synthesis question. That mismatch had no home in QUADAS-2. ## Judgement rules **Domain level.** If all signalling questions in a domain are answered *yes* or *probably yes*, risk of bias can be judged **low**. A *no* or *probably no* **flags potential** for bias — it does not settle it. Reviewers then apply their judgement and their review-specific guidance to decide whether the issue is likely to have influenced the accuracy estimates. > **A study can still be at low risk of bias with one or more signalling questions answered "no."** > The tool says this explicitly. Do not implement "any No → High" as a rule; that replaces the > judgement the tool asks for. Use **insufficient information** only when too little is reported to permit a judgement. It is not a middle rating between low and high. **Overall (phase 6)**, per estimate, done separately for risk of bias and for applicability: - any domain **high** → overall **high** - all domains **low** → overall **low** - any domain **insufficient information** and none high → overall **insufficient information** Record a rationale naming the major limitations behind the overall judgement. **Do not add a "moderate" grade.** Reviewers sometimes want one, to separate a study that is high risk in a single domain from one that is high risk in several. The E&E says plainly that the authors *do not support* this: if an estimate is high risk for one domain, it is high risk, whatever the other domains say. ## Tailoring the tool to your review Phase 2's tailoring is the step most often skipped, and the E&E is specific about it. - Do it **at the protocol stage, alongside phases 1 and 2** — not when you reach the studies. - Write review-specific guidance on how to answer each signalling question, adapting the general guidance tables. Publish it as a web appendix so the application is auditable. - Draw on **both clinical and methodological** expertise in the review area. - **Do not remove signalling questions.** Keep them even when they cannot bite: if two-gate designs were excluded, every study answers "yes" to 1.1, and recording that shows the issue was considered. Deleting the question hides that. - If you add a question, it must: address **one** issue only; concern **risk of bias, not reporting quality**; and be **factual**, phrased so that "yes"/"probably yes" means bias is absent. ## When to Use - Systematic reviews assessing the accuracy of tests used for **diagnosis, screening or staging** - New reviews — QUADAS-3 is the current recommended version - Alongside **QUADAS-C** when the review compares the accuracy of two or more index tests - Read the **Explanation & Elaboration report** before first use, and tailor the signalling questions and their guidance to your review (that tailoring is the step most often skipped) - Not for prediction models (PROBAST), non-randomised intervention studies (ROBINS-I), or randomised trials (RoB 2) -
QUADAS_C.md 9.9 KB
# QUADAS-C Assessment Guide Quality Assessment of Diagnostic Accuracy Studies — Comparative (extension to QUADAS-2). Reference: Yang B et al. *Guidance on how to use QUADAS-C*. Official tool, guidance document and templates are distributed by the QUADAS group at https://www.bristol.ac.uk/population-health-sciences/projects/quadas/quadas-c/ Version: QUADAS-C (2021) Source: Yang B, Mallett S, Takwoingi Y, Davenport CF, Hyde CJ, Whiting PF, et al. QUADAS-C: a tool for assessing risk of bias in comparative diagnostic accuracy studies. *Ann Intern Med* 2021;174(11):1592-1599 (DOI 10.7326/M21-2234). Licence: *Annals of Internal Medicine* (© American College of Physicians) — no open licence. Verification: every comparative signalling question (C1.1–C1.5, C2.1–C2.5, C3.1–C3.3, C4.1–C4.5), the five comparative study designs, the four completion phases, the domain judgement rules and the five optional Table 4 questions were compared against the **official QUADAS-C tool (`QUADAS-C tool.docx`) and guidance document**, both distributed by the QUADAS group at the University of Bristol. All matched. The article itself is paywalled and the author manuscripts at two university repositories sit behind a browser challenge; the group's own distribution made those unnecessary. > **QUADAS-3 is now the current recommended version of the base tool.** The QUADAS group states > that QUADAS-C "cannot be used alone and must be used alongside the main QUADAS tool", and now > recommends pairing it with **QUADAS-3** rather than QUADAS-2 — with adaptation, for which the > QUADAS-3 E&E document gives guidance. This file documents QUADAS-C as published against > QUADAS-2. **`QUADAS3.md`** carries the E&E's own adaptation: the four domains are renamed onto > QUADAS-3's (Flow and Timing becomes **Analysis**), **C3.2 is removed** as overlapping QUADAS-3 > item 3.4, **C4.2 moves to Index Test** and **C4.3 moves to Target Condition**, and the unit > assessed is a **comparative measure** you will usually have to compute yourself. An updated > QUADAS-C is in development. `QUADAS2.md` is the tool this extension was written against. ## Purpose QUADAS-C is an extension (add-on) to QUADAS-2 for assessing risk of bias in **comparative diagnostic test accuracy studies** — studies comparing two or more index tests in the same population. QUADAS-C should be used **together with** QUADAS-2: - QUADAS-2 assesses risk of bias for each individual test - QUADAS-C assesses additional risk of bias for the **comparison** between tests ## Structure QUADAS-C assesses the same 4 domains as QUADAS-2, with comparative signalling questions: - **Signalling questions**: answered Yes / No / Unclear - **Risk of bias judgment**: Low / High / Unclear (for the comparison) - **No applicability assessment** (inferred from QUADAS-2 judgments) ## Comparative Study Designs 1. **#1 Fully paired**: All participants receive all index tests (most robust) 2. **#2 Randomized**: Participants randomly allocated to one index test 3. **#3 Partially paired, random subset**: Some randomly selected for multiple tests 4. **#4 Partially paired, nonrandom subset**: Some receive multiple tests by clinical decision 5. **#5 Unpaired nonrandomized**: Different participants receive different tests Designs #1, #2, and #3 protect against confounding. Designs #4 and #5 have serious confounding potential. ## Four Phases of Completion 1. State the review question (which tests are compared, for what condition, in which population) 2. Tailor the tool (add/omit signalling questions per review) 3. Review the flow diagram (must show how participants were assigned to index tests) 4. Judge risk of bias (and applicability, based on QUADAS-2) ## Domain 1: Patient Selection ### Signalling Questions 1. **(C1.1)** Was the risk of bias for each index test judged 'low' for this domain? - If any index test is judged 'unclear' or 'high' in QUADAS-2, answer 'no' 2. **(C1.2)** Was a fully paired or randomized design used? - Designs #1, #2, and #3 (partially paired, random subset) → 'yes' or imply low risk - A 'no' answer should almost always prompt 'high risk' for this domain 3. **(C1.3)** Was the allocation sequence random? *(only applicable to randomized designs)* - Random: computer-generated numbers, random number tables, drawing lots - Non-random: alternation, dates, clinician preference 4. **(C1.4)** Was the allocation sequence concealed until patients were enrolled and assigned? *(only applicable to randomized designs)* - Appropriate: central randomization, telephone/internet service, sealed opaque envelopes ### Risk of Bias Judgment — **C1.5** Could the selection of patients have introduced bias in the comparison? - **Low**: 'Yes' to all applicable signalling questions - **High**: 'No' to C1.2 (not fully paired/randomized) almost always → high risk; 'No' to other questions raises concern - **Unclear**: Insufficient information to judge ## Domain 2: Index Test ### Signalling Questions 1. **(C2.1)** Was the risk of bias for each index test judged 'low' for this domain? - If any test is 'high' or 'unclear' in QUADAS-2, the comparison is also at risk 2. **(C2.2)** Were the index test results interpreted without knowledge of the results of the other index test(s)? *(only for paired/partially paired designs #1, #3, #4)* - Consider: subjectivity of interpretation, order of testing, whether output is automated 3. **(C2.3)** Is undergoing one index test unlikely to affect the performance of the other index test(s)? *(only for paired/partially paired designs #1, #3, #4)* - Consider: learning effects, fatigue, tissue distortion, sample depletion 4. **(C2.4)** Were the index tests conducted and interpreted without advantaging one of the tests? - Tests should be performed under similar conditions; differences in sample handling, equipment generation, or operator experience may bias the comparison ### Risk of Bias Judgment — **C2.5** Could the conduct or interpretation of the index tests have introduced bias in the comparison? - **Low**: 'Yes' to all applicable signalling questions - **High**: 'No' to any question suggesting the comparison is unfairly biased - **Unclear**: Insufficient information ## Domain 3: Reference Standard ### Signalling Questions 1. **(C3.1)** Was the risk of bias for each index test judged 'low' for this domain? - Reference standard misclassification biases both individual and comparative accuracy 2. **(C3.2)** Did the reference standard avoid incorporating any of the index tests? - If one index test is part of the reference standard, its accuracy is artificially inflated relative to the other test ### Risk of Bias Judgment — **C3.3** Could the reference standard, its conduct, or its interpretation have introduced bias in the comparison? - **Low**: 'Yes' to all signalling questions - **High**: 'No' to any, particularly if incorporation affects tests asymmetrically - **Unclear**: Insufficient information ## Domain 4: Flow and Timing ### Signalling Questions 1. **(C4.1)** Was the risk of bias for each index test judged 'low' for this domain? - Timing, verification, and exclusion issues for individual tests also bias the comparison 2. **(C4.2)** Was there an appropriate interval between the index tests? - Simultaneous or near-simultaneous testing avoids disease progression bias - What is 'appropriate' depends on the disease (acute vs. chronic) 3. **(C4.3)** Was the same reference standard used for all index tests? - Different reference standards across test groups (e.g., surgery vs. follow-up) introduce differential verification bias - If reference standards are exchangeable (detect same condition equally), a 'no' may not imply high risk 4. **(C4.4)** Are the proportions and reasons for missing data similar across index tests? - Differential missing data between test groups can bias the comparison - Consider both proportion and mechanism of missingness ### Risk of Bias Judgment — **C4.5** Could the patient flow have introduced bias in the comparison? - **Low**: 'Yes' to all signalling questions - **High**: 'No' to any question, particularly C4.3 (differential verification) or C4.4 (differential missing data) - **Unclear**: Insufficient information 'Unclear' does not mean 'moderate' risk of bias — it means there is insufficient information to judge. ## Optional Additional Signalling Questions (Table 4) These are not part of the core tool but can be added to tailor for specific reviews: | Domain | Question | Applicable to | |--------|----------|---------------| | Patient Selection | Were the same patient selection criteria used for those assigned to each index test? | Unpaired nonrandomized studies | | Patient Selection | If patients received all index tests, was the decision to use all tests made before participants were recruited? | Paired studies | | Patient Selection | Did the study avoid using prior tests as inclusion criteria that were correlated with only one of the index tests? | All | | Index Test | Did the study avoid using index test thresholds that are likely to advantage some of the index tests? | Studies with nonbinary tests | | Reference Standard | Is the mechanistic basis of one index test more closely shared with the reference standard than the other? | All | ## Presenting QUADAS-C Results - Present QUADAS-2 and QUADAS-C results side by side - QUADAS-C results are specific to each test comparison (present separately if multiple comparisons) - Use traffic-light tables: + (low risk, green), - (high risk, red), ? (unclear, yellow) - The guidance document points to www.quadas.org for templates; that domain no longer resolves, and the tabular and graphical templates are distributed from the QUADAS group's Bristol page (see the Reference line above) ## When to Use - Systematic reviews comparing two or more diagnostic tests head-to-head - Should always be used alongside QUADAS-2 (not as a replacement) - Primarily designed for fully paired (#1) and randomized (#2) designs - Can be adapted for other comparative designs with appropriate tailoring -
RECORD.md 6.7 KB
# RECORD Checklist **REporting of studies Conducted using Observational Routinely-collected health Data** Version: RECORD 2015 (13 items; a STROBE extension). The pharmacoepidemiology extension is **RECORD-PE** (Langan et al. *BMJ* 2018). Source: Benchimol EI, Smeeth L, Guttmann A, Harron K, Moher D, Petersen I, et al. The REporting of studies Conducted using Observational Routinely-collected health Data (RECORD) statement. *PLoS Med* 2015;12(10):e1001885 (DOI 10.1371/journal.pmed.1001885). https://www.record-statement.org Licence: CC BY 4.0 — confirmed via Crossref. Verification: all 13 items were compared against the checklist published in the statement; the id set — 1.1, 1.2, 1.3, 6.1, 6.2, 6.3, 7.1, 12.1, 12.2, 12.3, 13.1, 19.1, 22.1 — matches exactly. Apply when the manuscript is an **observational study conducted using routinely-collected health data** — administrative claims, electronic health records (EHR), disease/population registries, health-administrative or health-checkup databases, or linked versions of these — i.e. data **not collected for the purpose of the specific study**. RECORD extends the base **STROBE** items with reporting specific to secondary-use data: database identity, the codes/algorithms used to define the population and the variables, data linkage and its quality, and the limitations of analysing data collected for another purpose. For a drug safety/effectiveness study in such data, also apply **RECORD-PE**. For the design/validity review of the same study, pair with the RD1–RD8 domain probes in `peer-review` / `self-review` `references/domain-probes/record_routinely_collected_data.md`, and with the observational-confounding probes (`observational_confounding.md`). ## Checklist Items (13 items, extending STROBE) ### Title and Abstract (STROBE item 1) | # | Item | Description | |---|------|-------------| | 1.1 | Data type | The type of data used should be specified in the title or abstract. When possible, the name(s) of the database(s) used should be stated. | | 1.2 | Geography and timeframe | The geographic region and timeframe within which the study took place should be reported in the title or abstract. | | 1.3 | Linkage | If linkage between databases was conducted for the study, this should be clearly stated in the title or abstract. | ### Methods — Setting / Participants (STROBE item 6) | # | Item | Description | |---|------|-------------| | 6.1 | Population selection | The methods of study population selection (such as codes or algorithms used to identify subjects) should be listed in detail. If this is not possible, an explanation should be provided. | | 6.2 | Validation of codes | Any validation studies of the codes or algorithms used to select the population should be referenced. If validation was conducted for this study and not published elsewhere, detailed methods and results should be provided. | | 6.3 | Linkage diagram | If the study involved linkage of databases, consider use of a flow diagram or other graphical display to demonstrate the data linkage process, including the number of individuals with linked data at each stage. | ### Methods — Variables (STROBE item 7) | # | Item | Description | |---|------|-------------| | 7.1 | Codes for variables | A complete list of codes and algorithms used to classify exposures, outcomes, confounders, and effect modifiers should be provided. If these cannot be reported, an explanation should be provided. | ### Methods — Data access and cleaning (STROBE item 12) | # | Item | Description | |---|------|-------------| | 12.1 | Data access | Authors should describe the extent to which the investigators had access to the database population used to create the study population. | | 12.2 | Data cleaning | Authors should provide information on the data cleaning methods used in the study. | | 12.3 | Linkage methods | State whether the study included person-level, institutional-level, or other data linkage across two or more databases. The methods of linkage and methods of linkage quality evaluation should be provided. | ### Results — Participants (STROBE item 13) | # | Item | Description | |---|------|-------------| | 13.1 | Selection of included persons | Describe in detail the selection of the persons included in the study (i.e. study population selection) including filtering based on data quality, data availability and linkage. The selection of included persons can be described in the text and/or by means of the study flow diagram. | ### Discussion — Limitations (STROBE item 19) | # | Item | Description | |---|------|-------------| | 19.1 | Secondary-data limitations | Discuss the implications of using data that were not created or collected to answer the specific research question(s). Include discussion of misclassification bias, unmeasured confounding, missing data, and changing eligibility over time, as they pertain to the study being reported. | ### Other Information — Data access / cleaning (STROBE item 22) | # | Item | Description | |---|------|-------------| | 22.1 | Supplemental access | Authors should provide information on how to access any supplemental information such as the study protocol, raw data, or programming code. | --- ## Notes for Assessors - RECORD is an **extension of STROBE**; for the non-RECORD-specific items the base `STROBE.md` guidance also applies. Report both the base instrument and the extension when describing methods (do not cite RECORD as if it replaced STROBE). - The **highest-yield** items are **6.1 / 7.1** (the actual code lists / phenotype algorithms used to define the population, exposures, outcomes, and confounders — the single most common omission; "we identified diabetes from the database" with no codes is non-compliant), **6.2** (whether those algorithms were *validated*, and where), **12.3 / 6.3** (linkage method and linkage-quality evaluation, with a person-flow at each linkage stage), **13.1** (a participant-selection flow that includes data-quality/availability/linkage filtering — not only clinical eligibility), and **19.1** (the limitations specific to secondary-use data: misclassification from codes, unmeasured confounding, informative missingness, and eligibility drift over time). - For a **drug safety/effectiveness** study in routinely-collected data, also apply **RECORD-PE** (Langan et al. *BMJ* 2018;363:k3532), which adds items on exposure definition (drug codes, dose, duration, exposure windows), the comparator and new-user/active-comparator design, and immortal-time/protopathic bias. - This checklist was authored as a faithful summary of the RECORD statement (Benchimol EI, et al. *PLoS Med* 2015;12(10):e1001885, **CC BY 4.0**) for item-by-item assessment; verify against the published statement and its explanation-and-elaboration document for full item wording. Verified 2026-06-29. -
REMARK.md 7.1 KB
# REMARK Checklist **REporting recommendations for tumour MARKer prognostic studies** Version: REMARK 2005 (McShane et al.), with the 2012 Explanation and Elaboration (Altman et al.) Source: Item text reproduced from Table 1 of the CC BY explanation-and-elaboration paper. McShane LM, Altman DG, Sauerbrei W, Taube SE, Gion M, Clark GM. Br J Cancer 2005;93(4):387-391. Explanation and Elaboration: Altman DG, McShane LM, Sauerbrei W, Taube SE. PLoS Med 2012;9(5):e1001216. Complete the official REMARK instrument for a submission checklist. Source: McShane LM, Altman DG, Sauerbrei W, Taube SE, Gion M, Clark GM. Reporting recommendations for tumour marker prognostic studies (REMARK). *Br J Cancer* 2005;93(4):387-391. Explanation and elaboration: Altman DG, McShane LM, Sauerbrei W, Taube SE. *PLoS Med* 2012;9(5):e1001216 (DOI 10.1371/journal.pmed.1001216). Licence: The *PLoS Medicine* explanation-and-elaboration paper is CC BY 4.0 (confirmed via Crossref); the original *Br J Cancer* statement is not openly licensed. Verification: all 20 items were extracted from Table 1 of the CC BY explanation-and-elaboration paper (Europe PMC full text, PMC3362085) and compared item by item. **The count matched while the second half did not**: this file previously replaced item 14 with "Assay performance" and item 17 with "report all endpoints", dropped item 18 (further investigations) entirely, invented item 19 "Analyses reported", and shifted items 18–20. ## Checklist Items (20 items) ### Introduction | # | Item | Description | |---|------|-------------| | 1 | Marker and objectives | State the marker examined, the study objectives, and any pre-specified hypotheses. | ### Materials and Methods #### Patients | # | Item | Description | |---|------|-------------| | 2 | Patients | Describe the characteristics (e.g., disease stage, comorbidities) of the study patients, including their source and the inclusion and exclusion criteria. | | 3 | Treatments | Describe treatments received and how treatment was chosen (e.g., randomized or rule-based). | #### Specimen characteristics | # | Item | Description | |---|------|-------------| | 4 | Specimens | Describe the type of biological material used (including control samples) and the methods of preservation and storage. | #### Assay methods | # | Item | Description | |---|------|-------------| | 5 | Assay methods | Specify the assay method and provide (or reference) a detailed protocol, including specific reagents or kits, quality-control procedures, reproducibility assessments, quantitation methods, scoring and reporting protocols, and whether assays were performed blinded to the study endpoint. | #### Study design | # | Item | Description | |---|------|-------------| | 6 | Study design | State the method of case selection, whether the study was prospective or retrospective, and whether stratification or matching (e.g., by stage or age) was used; specify the time period from which cases were taken, the end of the follow-up period, and the median follow-up time. | | 7 | Endpoints | Precisely define all clinical endpoints examined. | | 8 | Candidate variables | List all candidate variables initially examined or considered for inclusion in models. | | 9 | Sample size | Give the rationale for the sample size; if the study was designed to detect a specified effect size, state the target power and effect size. | #### Statistical analysis methods | # | Item | Description | |---|------|-------------| | 10 | Statistical methods | Specify all statistical methods, including any variable-selection procedures and other model-building details, how model assumptions were verified, and how missing data were handled. | ### Results #### Data | # | Item | Description | |---|------|-------------| | 11 | Marker values in analysis | Clarify how marker values were handled in the analyses; if relevant, describe methods used for cutpoint determination. | | 12 | Patient flow | Describe the flow of patients through the study, including the number of patients included in each stage of the analysis (a diagram may be helpful) and reasons for dropout. Specifically, both overall and for each subgroup extensively examined, report the number of patients and the number of events. | | 13 | Baseline distributions | Report distributions of basic demographic characteristics (at least age and sex), standard (disease-specific) prognostic variables, and tumour marker, including numbers of missing values. | #### Analysis and presentation | # | Item | Description | |---|------|-------------| | 14 | Marker vs standard variables | Show the relation of the marker to standard prognostic variables. | | 15 | Univariable analyses | Present univariable analyses showing the relation between the marker and outcome, with the estimated effect (for example, hazard ratio and survival probability). Preferably provide similar analyses for all other variables being analysed. For the effect of a tumour marker on a time-to-event outcome, a Kaplan-Meier plot is recommended. | | 16 | Multivariable analyses | For key multivariable analyses, report estimated effects (for example, hazard ratio) with confidence intervals for the marker and, at least for the final model, all other variables in the model. | | 17 | Marker with standard variables | Among reported results, provide estimated effects with confidence intervals from an analysis in which the marker and standard prognostic variables are included, regardless of their statistical significance. | | 18 | Further investigations | If done, report results of further investigations, such as checking assumptions, sensitivity analyses, and internal validation. | ### Discussion | # | Item | Description | |---|------|-------------| | 19 | Interpretation | Interpret the results in the context of the pre-specified hypotheses and other relevant studies; include a discussion of limitations of the study. | | 20 | Implications | Discuss implications for future research and clinical value. | --- ## Notes for Assessors - REMARK primarily targets **single-marker** prognostic studies, but most items apply equally to studies of **multiple markers**, studies that **develop a prognostic model**, and studies that **predict response to treatment**. Apply the relevant items rather than marking them N/A by default. - The items most often MISSING in review are **item 11** (how marker values were handled, including cutpoint determination) and **item 17** (effects with confidence intervals from the analysis containing the marker *and* the standard prognostic variables, **regardless of statistical significance**) — flag these explicitly. - For a time-to-event marker study, expect a **Kaplan-Meier plot** (item 15) and a **multivariable model adjusted for established prognostic variables** (item 16), not a marker-alone analysis. - Pair REMARK with **STROBE** for the observational-design items (setting, eligibility, bias, missing data), and with **TRIPOD / TRIPOD+AI** when the study develops or validates a prognostic model. Name the base instrument and any extension and cite each (Step 4e). - Item text is reproduced from Table 1 of the CC BY explanation-and-elaboration paper with attribution. Complete the official REMARK instrument for a submission checklist. -
RoB2.md 8.1 KB
# RoB 2 Assessment Guide Revised Cochrane Risk-of-Bias tool for Randomised Trials. Version: RoB 2 (22 August 2019). Full guidance and the current tool: https://www.riskofbias.info Source: Sterne JAC, Savović J, Page MJ, Elbers RG, Blencowe NS, Boutron I, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. *BMJ* 2019;366:l4898 (DOI 10.1136/bmj.l4898; PMID 31462531). > **Fidelity and licence.** **No open licence was found for the article** — Crossref returns only > BMJ's text-and-data-mining policy, and `LICENSES.md` previously claimed CC BY for it on no > evidence. The signalling questions below were checked against the **official RoB 2 template and > full guidance document** distributed at riskofbias.info (version of 22 August 2019); their > wording is short-form and paraphrased. **Complete the official RoB 2 form for any assessment you > report.** > > Verification: every signalling question, its conditionality, the response options and the > domain- and overall-judgement rules were compared against that template and guidance. Four > problems were found and corrected: two questions were **missing** from the assignment variant of > domain 2 (2.5, 2.7), question 2.5 of the adherence variant had **inverted polarity** (the tool > asks about *non-adherence*), domain 5's three questions had been **collapsed into two**, and the > missing-data judgement carried a **">95%" threshold that the tool does not define** — the > guidance defines "nearly all" qualitatively and gives no percentage. Reference: Sterne JAC et al. BMJ 2019;366:l4898. PMID: 31462531. ## Structure RoB 2 is applied **to a specific result** — one outcome, one numerical result, in one trial. Record the experimental and comparator interventions, the outcome, and the numerical result before you start. - **Signalling question responses**: Yes / Probably yes / Probably no / No / No information (and **NA** where a question is conditional and its condition was not met) - **Domain judgement**: Low risk of bias / Some concerns / High risk of bias - **Optional, per domain and overall**: the predicted **direction** of bias — NA, favours experimental, favours comparator, towards null, away from null, or unpredictable Questions written as "If … to N.n" are **conditional**: ask them only when the stated answer was given to the earlier question. ## Before domain 2: state the effect of interest The review team must declare whether the aim for this result is to assess **the effect of assignment** to intervention (the intention-to-treat effect) or **the effect of adhering** to intervention. Domain 2 has a different set of questions for each; do not mix them. ## Domain 1: Bias arising from the randomisation process | # | Signalling question | |---|---------------------| | 1.1 | Was the allocation sequence random? | | 1.2 | Was the allocation sequence concealed until participants were enrolled and assigned to interventions? | | 1.3 | Did baseline differences between intervention groups suggest a problem with the randomisation process? | The tool does not aim to identify baseline imbalances that arose by chance; a small number of "statistically significant" differences at 0.05 is usually compatible with chance. ## Domain 2: Bias due to deviations from the intended interventions ### Variant A — effect of **assignment** to intervention | # | Signalling question | |---|---------------------| | 2.1 | Were participants aware of their assigned intervention during the trial? | | 2.2 | Were carers and people delivering the interventions aware of participants' assigned intervention during the trial? | | 2.3 | *If Y/PY/NI to 2.1 or 2.2:* Were there deviations from the intended intervention that arose because of the trial context? | | 2.4 | *If Y/PY to 2.3:* Were these deviations likely to have affected the outcome? | | 2.5 | *If Y/PY/NI to 2.4:* Were these deviations from intended intervention balanced between groups? | | 2.6 | Was an appropriate analysis used to estimate the effect of assignment to intervention? | | 2.7 | *If N/PN/NI to 2.6:* Was there potential for a substantial impact (on the result) of the failure to analyse participants in the group to which they were randomised? | ### Variant B — effect of **adhering** to intervention | # | Signalling question | |---|---------------------| | 2.1 | Were participants aware of their assigned intervention during the trial? | | 2.2 | Were carers and people delivering the interventions aware of participants' assigned intervention during the trial? | | 2.3 | *If applicable, and if Y/PY/NI to 2.1 or 2.2:* Were important non-protocol interventions balanced across intervention groups? | | 2.4 | *If applicable:* Were there failures in implementing the intervention that could have affected the outcome? | | 2.5 | *If applicable:* Was there **non-adherence** to the assigned intervention regimen that could have affected participants' outcomes? | | 2.6 | *If N/PN/NI to 2.3, or Y/PY/NI to 2.4 or 2.5:* Was an appropriate analysis used to estimate the effect of adhering to the intervention? | ## Domain 3: Bias due to missing outcome data | # | Signalling question | |---|---------------------| | 3.1 | Were data for this outcome available for all, or nearly all, participants randomised? | | 3.2 | *If N/PN/NI to 3.1:* Is there evidence that the result was not biased by missing outcome data? | | 3.3 | *If N/PN to 3.2:* Could missingness in the outcome depend on its true value? | | 3.4 | *If Y/PY/NI to 3.3:* Is it likely that missingness in the outcome depended on its true value? | **"Nearly all" is not a percentage.** The guidance defines it as: the number of participants with missing outcome data is so small that their outcomes, whatever they were, could have made no important difference to the estimated effect. Do not substitute a 95% or 80% rule — RoB 2 states none. **Low risk** requires any one of: (i) outcome data available for all, or nearly all, randomised participants; **or** (ii) evidence that the result was not biased by missing outcome data; **or** (iii) missingness in the outcome could not depend on its true value. ## Domain 4: Bias in measurement of the outcome | # | Signalling question | |---|---------------------| | 4.1 | Was the method of measuring the outcome inappropriate? | | 4.2 | Could measurement or ascertainment of the outcome have differed between intervention groups? | | 4.3 | *If N/PN/NI to 4.1 and 4.2:* Were outcome assessors aware of the intervention received by study participants? | | 4.4 | *If Y/PY/NI to 4.3:* Could assessment of the outcome have been influenced by knowledge of intervention received? | | 4.5 | *If Y/PY/NI to 4.4:* Is it likely that assessment of the outcome was influenced by knowledge of intervention received? | ## Domain 5: Bias in selection of the reported result | # | Signalling question | |---|---------------------| | 5.1 | Were the data that produced this result analysed in accordance with a pre-specified analysis plan that was finalised before unblinded outcome data were available for analysis? | | 5.2 | Is the numerical result being assessed likely to have been selected, on the basis of the results, from multiple eligible outcome **measurements** (e.g. scales, definitions, time points) within the outcome domain? | | 5.3 | Is the numerical result being assessed likely to have been selected, on the basis of the results, from multiple eligible **analyses** of the data? | 5.2 and 5.3 are separate questions. Selecting a measurement and selecting an analysis are different acts, and a result can be at risk from one and not the other. ## Overall risk of bias - **Low risk of bias**: low risk of bias for all domains - **Some concerns**: some concerns in at least one domain, but not high risk of bias in any domain - **High risk of bias**: high risk of bias in at least one domain, **or** some concerns for multiple domains in a way that substantially lowers confidence in the result ## When to Use - Use for **individually randomised, parallel-group trials** (default) - Variants are published for cluster-randomised trials and crossover trials — use the matching variant - Do NOT use for non-randomised studies (use ROBINS-I instead) -
ROBINS_E.md 8.9 KB
# ROBINS-E Assessment Guide Risk Of Bias In Non-randomized Studies — of Exposures. Reference: Higgins JPT et al. Environment International 2024;186:108602. doi: 10.1016/j.envint.2024.108602. Website: https://www.riskofbias.info/welcome/robins-e-tool Version: ROBINS-E (2024) Licence: CC BY-NC 4.0 (non-commercial) — confirmed via Crossref. Not redistributable verbatim under this repository's MIT licence. Verification: the seven bias domains and their order were compared against Table 1 of the article (Europe PMC full text, PMC11098530), and the preliminary-considerations parts A–E, the four domain-judgement levels, the three per-domain outputs and the domain-1 label "Low risk of bias (except for concerns about uncontrolled confounding)" against its section 4–6 text. All matched. Item wording stays paraphrased: the source is CC BY-**NC**, which cannot be redistributed under this repository's MIT licence. Complete the official instrument for anything you report. ## Purpose ROBINS-E assesses the risk of **material bias** in individual observational studies examining the effect of an **exposure** on an outcome. Designed for follow-up (cohort) studies. Material bias = bias sufficient to affect the direction of the estimated effect or impact the ability to draw conclusions. ## Structure ROBINS-E has: - **Planning stage**: List confounders and consider appropriateness - **Preliminary considerations** (per study): Sections A-E - **7 bias domains** with signalling questions - **Overall judgment** across domains Signalling question responses: Yes / Probably yes / Probably no / No / No information (Some questions have domain-specific response options) Domain judgments: Low risk of bias / Some concerns / High risk of bias / Very high risk of bias Three outputs per domain: 1. Risk of bias judgment (algorithmic, overridable) 2. Predicted direction of bias 3. Whether bias threatens conclusions (Yes / No / Cannot tell) ## Planning Stage ### P1: List Important Confounding Factors Specify confounders relevant to all or most studies on this topic. State whether they are particular to specific exposure-outcome combinations. ### P2: Appropriateness Assessment Will the review use the ROBINS-E assessment of appropriateness (study sensitivity)? → Yes / No If Yes, complete Appendix 1 (Parts I, II, III). ## Preliminary Considerations (Per Study) ### A. Specify the Result Being Assessed - ROBINS-E is specific to a particular study result - Different results from the same study may have different risks of bias - Specify the numerical result (e.g., RR=1.52, 95% CI 0.83 to 2.77) ### B. Decide Whether to Proceed - Some study characteristics may directly indicate very high risk of bias - Screening section to identify such situations before detailed assessment ### C. Specify the Analysis - Gather information about participants, exposure measure(s), outcome, and analysis methods ### D. Specify the Causal Effect - Define the causal effect of exposure being estimated - This is essential: the causal effect defines what the result would be in the absence of bias ### E. Specify Important Confounding Factors - Specify the known important confounding factors likely to influence the exposure–outcome association - Identify them by reviewing the literature **and** consulting experts, including members of the review team - Parts C and D may be completed in either order, but both come before part E ## Domain 1: Risk of Bias Due to Confounding ### Key Signalling Questions - Were all important confounding domains measured? - Were all important confounding domains adequately controlled for? - Was the analysis adjusted for all important confounders, or was matching/restriction used? - Were appropriate methods used to control confounding (regression, propensity score, IP weighting)? ### Judgment - **Low risk of bias (except for concerns about uncontrolled confounding)**: All critical confounders well addressed, but residual uncontrolled confounding cannot be eliminated in observational studies - **Some concerns**: Minor concerns about residual confounding - **High risk of bias**: Important confounders not adequately controlled - **Very high risk of bias**: Critical confounding renders the estimate unreliable *Note: For Domain 1, 'Low risk' is expressed as "Low risk of bias (except for concerns about uncontrolled confounding)" because uncontrolled confounding can never be fully excluded in observational studies.* ## Domain 2: Risk of Bias Arising from Measurement of the Exposure ### Key Signalling Questions - Was the exposure clearly defined and consistently measured? - Was exposure assessment valid and reliable? - Was exposure measured at the appropriate time point? - Was exposure classification differential with respect to the outcome? ### Judgment - **Low**: Exposure well-defined, measured validly, non-differential misclassification - **Some concerns**: Minor measurement issues unlikely to materially affect results - **High**: Exposure measurement likely to introduce material bias - **Very high**: Exposure so poorly measured that no useful estimate possible ## Domain 3: Risk of Bias in Selection of Participants ### Key Signalling Questions - Was selection into the study (or into the analysis) related to both exposure and outcome? - Was start of follow-up aligned with exposure assessment? - Were adjustments made for selection effects? ### Judgment - **Low**: Selection not related to both exposure and outcome - **Some concerns**: Minor selection issues - **High**: Selection bias likely to materially affect results - **Very high**: Extreme selection bias ## Domain 4: Risk of Bias Due to Post-exposure Interventions ### Key Signalling Questions - Were there post-exposure interventions that could have affected the outcome? - Were these interventions differential across exposure groups? - Were they likely to affect the estimated exposure effect? ### Judgment - **Low**: No important post-exposure interventions, or balanced across groups - **Some concerns**: Some differential post-exposure interventions - **High**: Important post-exposure interventions likely bias the result - **Very high**: Post-exposure interventions dominate the observed effect ## Domain 5: Risk of Bias Due to Missing Data ### Key Signalling Questions - Were outcome data available for all or nearly all participants? - Were participants excluded due to missing data on exposure or other variables? - Was the proportion of missing data similar across exposure groups? - Were appropriate methods used to handle missing data? ### Judgment - **Low**: Complete data or minimal, non-differential missingness - **Some concerns**: Some missing data but unlikely to materially bias results - **High**: Missing data pattern likely to introduce material bias - **Very high**: Extensive missing data with differential patterns ## Domain 6: Risk of Bias Arising from Measurement of the Outcome ### Key Signalling Questions - Could outcome measurement have been influenced by knowledge of exposure status? - Were outcome assessors blinded to exposure? - Were outcome measures comparable across exposure groups? - Was the outcome defined and assessed using validated methods? ### Judgment - **Low**: Objective outcome or blinded assessment, non-differential measurement - **Some concerns**: Minor measurement concerns - **High**: Outcome measurement likely differentially affected by exposure knowledge - **Very high**: Outcome measurement severely compromised ## Domain 7: Risk of Bias in Selection of the Reported Result ### Key Signalling Questions - Were multiple outcome measurements reported? - Were multiple analyses performed (different adjustments, subgroups, models)? - Is the reported result likely selected from among multiple measurements or analyses? ### Judgment - **Low**: Pre-specified analysis plan or single planned analysis - **Some concerns**: Minor concerns about selective reporting - **High**: Reported result likely selected to favor a particular conclusion - **Very high**: Clear evidence of selective reporting ## Overall Risk of Bias Three overall judgments are derived from the domain judgments: 1. **Overall risk of bias**: Most conservative across domains - **Low**: Low in all domains - **Some concerns**: Some concerns in at least one domain, but no high/very high - **High**: High in at least one domain, but not very high in any - **Very high**: Very high in at least one domain 2. **Overall predicted direction of bias**: Considering all domains together 3. **Overall: does bias threaten conclusions?**: Yes / No / Cannot tell Justification should be provided when overriding algorithm-suggested judgments. ## When to Use - Systematic reviews of observational studies examining exposure effects (environmental, occupational, behavioral, dietary) - Currently designed for cohort (follow-up) studies - Not for case-control studies (extension planned) - Not for intervention studies (use ROBINS-I instead) - Complementary to GRADE for rating certainty of evidence from observational studies -
ROBINS_I.md 5.6 KB
# ROBINS-I Assessment Guide Risk Of Bias In Non-randomised Studies - of Interventions. Version: ROBINS-I (2016), the original version. Tool home: https://www.riskofbias.info Source: Sterne JAC, Hernán MA, Reeves BC, Savović J, Berkman ND, Viswanathan M, et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. *BMJ* 2016;355:i4919 (DOI 10.1136/bmj.i4919). > **Fidelity and licence.** The source article is **CC BY-NC 3.0** — non-commercial. This repository > is MIT-licensed and redistributed without restriction, so the tool's wording cannot be carried > verbatim here. This file is an **in-house summary of the tool's structure**: the seven domains, > their order, the answer options and the judgement levels, all of which were checked against the > article. The per-domain questions below are abbreviated and are **not** the tool's signalling > questions. **Complete the official ROBINS-I form from riskofbias.info for any assessment you > report.** > > **Verification.** The seven domains, their order, their pre-/at-/post-intervention grouping and > the five judgement levels with their across-domain criteria were compared against Tables 1 and 2 > of the article (Europe PMC full text, PMC5062054). All matched. The **No information** overall > judgement, which the file omitted, has been added. > > **A version 2 exists and is still in draft.** ROBINS-I V2 adds algorithms mapping signalling-question > answers onto domain judgements, and covers bias due to immortal time, which the 2016 version omits. > A revised draft was posted in November 2025 and is subject to change. Check riskofbias.info before > choosing which version to appraise against; this file documents the 2016 version. Reference: Sterne JAC et al. BMJ 2016;355:i4919. ## Structure ROBINS-I assesses 7 domains + overall judgment. The article groups them by when the bias arises: **pre-intervention** (domains 1–2, where assessment is mainly distinct from randomised trials), **at intervention** (domain 3, also mainly distinct), and **post-intervention** (domains 4–7, which overlap substantially with assessments of randomised trials). - **Signalling questions**: Yes / Probably yes / Probably no / No / No information - **Domain judgment**: Low / Moderate / Serious / Critical / No information - **Overall judgment**: Lowest of all domain judgments (most conservative) ## Pre-assessment Requirements Before applying ROBINS-I, specify: 1. The target trial (what RCT would ideally answer this question?) 2. The effect of interest (assignment to intervention vs starting and adhering) 3. Confounders to be controlled ## Domain 1: Bias Due to Confounding ### Key Questions - Is there potential for confounding not accounted for? - Did the authors use appropriate methods to control confounding (matching, regression, propensity score)? ### Judgment - **Low**: All critical confounders appropriately controlled - **Moderate**: Minor concerns about residual confounding - **Serious**: Important confounders not adequately controlled - **Critical**: Confounding so severe that no useful estimate possible ## Domain 2: Bias in Selection of Participants into the Study ### Key Questions - Was selection into the study related to both intervention and outcome? - Was start of follow-up and intervention aligned? - Were adjustments made for different start times? ## Domain 3: Bias in Classification of Interventions ### Key Questions - Were intervention groups clearly defined? - Was information used to classify interventions recorded at the start of the intervention? - Could classification of intervention status have been affected by knowledge of the outcome? ## Domain 4: Bias Due to Deviations from Intended Interventions ### Key Questions - Were there deviations from intended intervention beyond what would be expected? - Were these deviations unbalanced between groups and likely to affect outcomes? - Were important co-interventions balanced across groups? ## Domain 5: Bias Due to Missing Data ### Key Questions - Were outcome data available for all or nearly all participants? - Were participants excluded due to missing data on intervention or other variables? - Was the proportion of missing data similar across groups? - Were appropriate methods used to handle missing data? ## Domain 6: Bias in Measurement of Outcomes ### Key Questions - Could outcome measurement have been influenced by knowledge of intervention? - Were outcome assessors blinded? - Were outcome measures comparable across groups? ## Domain 7: Bias in Selection of the Reported Result ### Key Questions - Were multiple outcome measurements reported? - Were multiple analyses performed? - Is the reported result likely selected from among multiple measurements or analyses? ## Overall Risk of Bias The overall judgment follows the criteria in the article's Table 2: - **Low**: the study is judged at low risk of bias for all domains — comparable to a well performed randomised trial - **Moderate**: low or moderate for all domains — sound evidence for a non-randomised study, but not comparable to a well performed randomised trial - **Serious**: serious in at least one domain, but not critical in any - **Critical**: critical in at least one domain — too problematic to provide useful evidence, and should not be included in any synthesis - **No information**: no clear indication that the study is at serious or critical risk of bias, and information is lacking in one or more key domains ## Recommendation for Synthesis - Studies at **critical** risk of bias should be excluded from meta-analysis - Present critical studies in a separate table for completeness - Conduct sensitivity analysis excluding serious risk of bias studies -
ROBIS.md 7.2 KB
# ROBIS Assessment Guide Risk Of Bias In Systematic Reviews, version 1.2. Reference: Whiting P et al. J Clin Epidemiol 2016;69:225-234. Version: ROBIS (2016) — 4 domains in 3 phases Source: Whiting P, Savović J, Higgins JPT, Caldwell DM, Reeves BC, Shea B, et al. ROBIS: a new tool to assess risk of bias in systematic reviews was developed. *J Clin Epidemiol* 2016;69:225-234 (DOI 10.1016/j.jclinepi.2015.06.005). Licence: CC BY 4.0 — confirmed via Crossref. (`LICENSES.md` previously recorded no open licence for ROBIS; that was wrong.) Verification: all 24 questions (21 phase-2 signalling questions plus phase-3 A/B/C) were compared word by word against the **official ROBIS 1.2 tool PDF** distributed by the University of Bristol (`ROBIS 1.2 Clean.pdf`), and the judgement guidance against the statement's full text (Europe PMC, PMC4687950). The questions matched 24/24. **The judgement guidance did not**: this file previously carried a "Judgment Algorithm" stating that high concern in any domain means high risk of bias, which the statement contradicts, and omitted the domain rule that the statement does give. ## Purpose ROBIS assesses the risk of bias **in systematic reviews themselves** (not in individual studies). It evaluates whether the review process was conducted appropriately and whether the conclusions are trustworthy. Applicable to reviews of: interventions, aetiology, diagnosis, and prognosis. ## Structure ROBIS has 3 phases: - **Phase 1** (optional): Assessing relevance of the review to the target question - **Phase 2**: Identifying concerns with the review process (4 domains, signalling questions) - **Phase 3**: Judging overall risk of bias in the review Signalling questions answered: Y (Yes) / PY (Probably Yes) / PN (Probably No) / N (No) / NI (No Information) Domain concerns rated: LOW / HIGH / UNCLEAR ## Phase 1: Assessing Relevance (Optional) State the target question and compare it with the question addressed by the review. For **intervention reviews**, specify: Patients/Population, Intervention(s), Comparator(s), Outcome(s) For **aetiology reviews**, specify: Patients/Population, Exposure(s) and Comparator(s), Outcome(s) For **DTA reviews**, specify: Patients, Index test(s), Reference standard, Target condition For **prognostic reviews**, specify: Patients, Outcome to be predicted, Intended use of model, Intended moment in time Final question: Does the question addressed by the review match the target question? → YES / NO / UNCLEAR ## Phase 2: Identifying Concerns with the Review Process ### Domain 1: Study Eligibility Criteria Describe the study eligibility criteria, any restrictions, and whether objectives and criteria were pre-specified. #### Signalling Questions 1.1 Did the review adhere to pre-defined objectives and eligibility criteria? 1.2 Were the eligibility criteria appropriate for the review question? 1.3 Were eligibility criteria unambiguous? 1.4 Were any restrictions in eligibility criteria based on study characteristics appropriate (e.g., date, sample size, study quality, outcomes measured)? 1.5 Were any restrictions in eligibility criteria based on sources of information appropriate (e.g., publication status or format, language, availability of data)? **Concerns regarding specification of study eligibility criteria:** LOW / HIGH / UNCLEAR ### Domain 2: Identification and Selection of Studies Describe methods of study identification and selection (e.g., number of reviewers involved). #### Signalling Questions 2.1 Did the search include an appropriate range of databases/electronic sources for published and unpublished reports? 2.2 Were methods additional to database searching used to identify relevant reports? 2.3 Were the terms and structure of the search strategy likely to retrieve as many eligible studies as possible? 2.4 Were restrictions based on date, publication format, or language appropriate? 2.5 Were efforts made to minimise error in selection of studies? **Concerns regarding methods used to identify and/or select studies:** LOW / HIGH / UNCLEAR ### Domain 3: Data Collection and Study Appraisal Describe methods of data collection, what data were extracted, how risk of bias was assessed (e.g., number of reviewers), and the tool used. #### Signalling Questions 3.1 Were efforts made to minimise error in data collection? 3.2 Were sufficient study characteristics available for both review authors and readers to be able to interpret the results? 3.3 Were all relevant study results collected for use in the synthesis? 3.4 Was risk of bias (or methodological quality) formally assessed using appropriate criteria? 3.5 Were efforts made to minimise error in risk of bias assessment? **Concerns regarding methods used to collect data and appraise studies:** LOW / HIGH / UNCLEAR ### Domain 4: Synthesis and Findings Describe synthesis methods. #### Signalling Questions 4.1 Did the synthesis include all studies that it should? 4.2 Were all pre-defined analyses reported or departures explained? 4.3 Was the synthesis appropriate given the nature and similarity in the research questions, study designs and outcomes across included studies? 4.4 Was between-study variation (heterogeneity) minimal or addressed in the synthesis? 4.5 Were the findings robust e.g. as demonstrated through funnel plot or sensitivity analyses? 4.6 Were biases in primary studies minimal or addressed in the synthesis? **Concerns regarding the synthesis and findings:** LOW / HIGH / UNCLEAR ## Phase 3: Judging Risk of Bias in the Review Describe whether conclusions were supported by the evidence. ### Signalling Questions A. Did the interpretation of findings address all of the concerns identified in Domains 1 to 4? B. Was the relevance of identified studies to the review's research question appropriately considered? C. Did the reviewers avoid emphasizing results on the basis of their statistical significance? **Risk of bias in the review:** LOW / HIGH / UNCLEAR ### How the judgements are reached ROBIS gives a rule for **domain concern** and deliberately declines to give one for overall risk of bias — a design goal of the tool was "an overall risk of bias rating for a single review without using summary quality scores." **Phase 2, per domain.** If the answers to all signalling questions for a domain are *yes* or *probably yes*, the level of concern can be judged **low**. If any signalling question is answered *no* or *probably no*, potential for concern exists. Use *no information* only when insufficient data are reported to permit a judgment. **Phase 3.** Concerns identified in phase 2 do **not** automatically make the review high risk. The statement is explicit: where concerns were identified in earlier domains but "these were appropriately considered when interpreting results and drawing conclusions, then this may also be rated as 'yes,' and depending on the rating of the other signaling questions, the review may still be rated as 'low risk of bias.'" The phase-3 judgement is the assessor's, recorded with a rationale. ## When to Use - Assessing quality of existing systematic reviews (overview of reviews) - Evaluating competing systematic reviews on the same topic - Guideline development (appraising evidence base) - Critical appraisal teaching and journal clubs - Not for assessing individual primary studies (use QUADAS-2, RoB 2, ROBINS-I, etc.) -
ROB_ME.md 7.7 KB
# ROB-ME Assessment Guide Risk Of Bias due to Missing Evidence in a meta-analysis. Reference: Page MJ et al. BMJ 2023;383:e076754; ROB-ME cribsheet Version 1, October 2023. Website: https://www.riskofbias.info/welcome/rob-me-tool Version: ROB-ME (2023) Source: Page MJ, Sterne JAC, Boutron I, Hróbjartsson A, Kirkham JJ, Li T, et al. ROB-ME: a tool for assessing risk of bias due to missing evidence in systematic reviews with meta-analysis. *BMJ* 2023;383:e076754 (DOI 10.1136/bmj-2023-076754). Licence: Crossref returns only a BMJ text-and-data-mining policy; no Creative Commons licence. Verification: the five steps, the Results Matrix symbols, all eleven signalling questions (3.1–3.3, 4.1–4.8) with their conditions and response options, and the judgement scale were compared word by word against the **official ROB-ME cribsheet Version 1 (October 2023)** and the completion template distributed at riskofbias.info. All matched; only the truncated citation on line 4 was wrong. > A newer cribsheet exists, generalised from "meta-analysis" to "**synthesis**". Its question 4.5 > takes Y / PY / PN / N rather than Y / N, and 4.7's condition widens to "Y to 4.1 or 4.3 **or > Y/PY to 4.5**". This file documents Version 1, which its citation names; check riskofbias.info > before appraising against a version. ## Purpose ROB-ME assesses the risk of bias in a **pairwise meta-analysis result** due to missing evidence — encompassing both non-reporting biases (selective reporting of results) and non-publication biases (unpublished studies). It is applied at the synthesis level (not individual study level). ## Structure ROB-ME has 5 steps applied to **each meta-analysis** in a systematic review: 1. Select and define meta-analyses to assess 2. Determine which studies have missing results (Results Matrix) 3. Consider the potential for missing studies across the systematic review 4. Assess risk of bias due to missing evidence (signalling questions) 5. Overall risk of bias judgment Response options: Y (Yes) / PY (Probably Yes) / PN (Probably No) / N (No) / NI (No Information) / NA (Not Applicable) Green-underlined responses are potential markers for **low** risk of bias. Red responses are potential markers for **a** risk of bias. ## Step 1: Select and Define Meta-analyses For each meta-analysis, specify: - The PICO: Participants, Intervention, Comparator, Outcome - Eligible study designs - Eligible outcome definitions (measures, metrics, time points) - Eligible methods of analysis (analysis populations, crude vs adjusted estimates) ## Step 2: Determine Which Studies Have Missing Results (Results Matrix) For each study meeting inclusion criteria, create a Results Matrix using these symbols: | Symbol | Meaning | |--------|---------| | check (green) | A study result is available for inclusion in the meta-analysis | | ~ (yellow) | No study result available, for a reason **unrelated** to the P value, magnitude or direction | | ? (orange) | Unclear whether an eligible result was generated | | X (red) | No study result available, **likely because of** the P value, magnitude or direction | Record: Study ID, Source(s) used, Number of participants analysed, and availability status for each meta-analysis. Also record any known information about the results (direction of effect, statistical significance, or narrative descriptions). ## Step 3: Consider the Potential for Missing Studies Answer the following questions **once** for the systematic review as a whole: ### Signalling Questions 3.1 Were prospectively registered studies or studies identified for a prospective meta-analysis the only type of study eligible for inclusion? - **Y**: inception cohort → lower risk of missing studies - **N**: proceed to 3.2 3.2 If N to 3.1: Would you expect every eligible study to be identifiable regardless of its results? - **NA/Y/PY**: lower risk - **PN/N**: higher risk — proceed to 3.3 3.3 If Y/PY to 3.2: Were you likely to have found all eligible studies regardless of their results? - Consider: trial registers searched, search strategy designed to retrieve regardless of outcomes - **NA/Y/PY**: lower risk - **PN/N**: higher risk — check the box indicating potential for missing studies ## Step 4: Assess Risk of Bias (per meta-analysis) Specify the meta-analysis details: description, summary effect estimate (95% CI), number of included studies and participants. ### Within-study Assessment ('Known Unknowns') 4.1 Of the studies identified, was there any for which no result was available for inclusion in the meta-analysis, **likely because of the P value, magnitude or direction** of the result generated (refer to Step 2)? - Answer 'Yes' if any study has an 'X' in the Results Matrix for this meta-analysis - **Y/N** 4.2 If Y to 4.1: Is it likely that there would be a notable change to the summary effect estimate if the omitted results had been included? - Consider: amount of missing evidence relative to total; direction of effect in omitted studies; fixed-effect vs random-effects model - **Trivial missing → 'No/Probably no'**; Non-trivial with differing direction → 'Yes/Probably yes' - **NA/Y/PY/PN/N/NI** ### Within-study Assessment ('Unknown Unknowns') 4.3 Of the studies identified, was there any for which it was unclear whether an eligible result was generated (refer to Step 2)? - Answer 'Yes' if any study has a '?' in the Results Matrix - **Y/N** 4.4 If Y to 4.3: Is it likely that there would be a notable change to the summary effect estimate if the potentially omitted results had been included? - Consider same factors as 4.2 but for 'potentially' missing evidence - **NA/Y/PY/PN/N/NI** ### Across-study Assessment 4.5 Do circumstances (identified in Step 3) indicate potential for some eligible studies not being identified because of the P value, magnitude or direction of the results generated? - Answer 'Yes' if the checkbox in Step 3 was checked - **Y/N** 4.6 If Y to 4.5: Is it likely that studies not identified had results that were eligible for inclusion in the meta-analysis? - Consider: core outcome sets, scope of outcome domain, whether missing studies would have measured this outcome - **NA/Y/PY/PN/N** 4.7 If Y to 4.1, 4.3, or 4.5: Does the pattern of observed study results suggest that the meta-analysis is likely to be missing results that were systematically different (in terms of P value, magnitude or direction) from those observed? - Methods: (1) funnel plot inspection, (2) funnel plot asymmetry test, (3) compare fixed-effect vs random-effects estimates, (4) inspect P values/direction in forest plot - Distinguish small-study effects from non-reporting biases - **NA/Y/PY/PN/N** 4.8 If Y/PY/NI to 4.2, 4.4, 4.6, or 4.7: Did sensitivity analyses suggest that the summary effect estimate was biased due to missing results? - Consider: selection models, regression-based adjustment, restricting to largest studies - **NA/Y/PY/PN/N** ## Step 5: Risk of Bias Judgment ### Judgment Scale - **Low**: Little or no concern about missing evidence biasing the meta-analysis result - **Some concerns**: Some concern but not sufficient to judge high risk - **High**: The meta-analysis result is likely biased due to missing evidence ### Optional: Predicted Direction of Bias - Favours experimental / Favours comparator / Towards null / Away from null / Unpredictable ## When to Use - Every pairwise meta-analysis in a systematic review (PRISMA 2020 recommends this) - Replaces informal funnel plot interpretation with a structured assessment - Complements individual study RoB tools (RoB 2, ROBINS-I, QUADAS-2) - Not designed for network meta-analysis (use RoB NMA instead) - Should be completed by someone familiar with the meta-analysis methods and included studies -
RoB_NMA.md 6.4 KB
# RoB NMA Assessment Guide Risk of Bias in Network Meta-Analysis tool. Website: https://www.riskofbias.info Version: RoB NMA (2025) — 17 items in 3 domains Source: Lunny C, Higgins JPT, White IR, Dias S, Hutton B, Pham B, et al. Risk of Bias in Network Meta-Analysis (RoB NMA) tool. *BMJ* 2025;388:e079839 (DOI 10.1136/bmj-2024-079839). Licence: CC BY 4.0 — confirmed via Crossref. Verification: all 17 signalling statements were extracted from the guidance article's own item sections (Europe PMC full text, PMC11915405) and compared statement by statement, together with the domain and overall judgement rules. **This file previously carried 18 invented items**; only statement 1.1 survived the comparison. ## Purpose The RoB NMA tool assesses the risk of bias in **a single network meta-analysis** by identifying limitations in how the NMA was conducted, including how the evidence was assembled, that could bias the NMA's results or conclusions. It is **not** a tool for the primary studies inside the network (use RoB 2 or ROBINS-I) and **not** a tool for the systematic review that contains the NMA (use ROBIS or AMSTAR 2). It is meant to be used alongside one of the latter. One review yields one ROBIS assessment but as many RoB NMA assessments as there are NMAs. ## Structure **17 items** as signalling statements, in 3 domains: 1. Interventions and network geometry (4 statements) 2. Effect modifiers (4 statements) 3. Statistical synthesis (9 statements) Response options: **true / probably true / probably false / false / no information**. *True* indicates the lowest risk of bias. Item 3.9 may also be answered *not applicable*. Statements 2.4 and 3.8 are **conditional** — they are considered only when a preceding statement was answered false or probably false. Domain-level judgment: **low risk of bias / some concerns / high risk of bias**, supported by written justification and quotes from the NMA manuscript. ## Domain 1: Interventions and Network Geometry How the interventions were selected and grouped, and whether they are an appropriate set for performing an NMA. | # | Signalling statement | |---|----------------------| | 1.1 | All interventions and their comparators included in the NMA are reasonable alternatives for the whole target population | | 1.2 | All eligible interventions were included in the network | | 1.3 | Interventions were appropriately grouped into nodes in the network | | 1.4 | All compared interventions were connected through a suitable chain of within study comparisons | ## Domain 2: Effect Modifiers Whether the studies contributing to different direct comparisons are similar enough on the characteristics that modify the intervention effect (the transitivity requirement). | # | Signalling statement | |---|----------------------| | 2.1 | Outcome definitions and time points were similar across direct comparisons in the network | | 2.2 | Effect modifying participant characteristics were similar across direct comparisons in the network | | 2.3 | Effect modifying study characteristics were similar across direct comparisons in the network | | 2.4 | *(only if 2.1, 2.2 or 2.3 was false or probably false)* The analysis appropriately looked at the differences in effect modifiers across the network | ## Domain 3: Statistical Synthesis Non-reporting biases, biases within the primary studies, statistical methods, and conflict between direct and indirect evidence. | # | Signalling statement | |---|----------------------| | 3.1 | The analysis respected within study randomisation | | 3.2 | No publication bias or other selective non-reporting biases were suspected | | 3.3 | All predefined analyses, and only those analyses, were reported, or discrepancies were explained | | 3.4 | Biases in primary studies were minimal or addressed in the synthesis | | 3.5 | Appropriate methods were used to handle multi-arm studies | | 3.6 | Appropriate assumptions were made about homogeneity or heterogeneity of effects within comparisons | | 3.7 | No evidence of conflict between direct and indirect estimates of the same effect | | 3.8 | *(only if 3.7 was false or probably false)* Conflicting results between direct and indirect evidence were adequately dealt with | | 3.9 | If a bayesian analysis was performed, the choice of prior distributions was appropriate | ## Overall Judgment An overall judgment may be made about the **results** of the NMA, its **conclusions**, or both — whichever the assessor intends to use. Risk of bias at the systematic-review level (ROBIS or AMSTAR 2) is combined with the three RoB NMA domain judgments. When RoB NMA is used with ROBIS, ROBIS phase 3 is omitted and only its first three domains are considered. **Bias in the results of the NMA** — low risk of bias / some concerns / high risk of bias. If all domains were judged low risk, a judgment of low risk should generally be made; otherwise the assessor decides between some concerns and high risk. Some concerns in multiple domains may warrant an overall judgment of high risk. **Bias in the conclusions of the NMA** — concerns / no concerns. The question is whether the NMA authors dealt with all the limitations identified. Items 3.5, 3.6 and 3.8 should be reconsidered here, because inappropriate modelling choices can lead to uncertainty being under- or overestimated. The two judgments can differ: estimated effects may be at high risk of bias because of the primary studies while the conclusions are at low risk, if that risk was carefully taken into account. If the effects are at high risk because of how the NMA itself was conducted, the conclusions are unlikely to be at low risk. Focus on bias in results when the NMA results feed a decision model; focus on bias in conclusions when the conclusions are used for decision making. ## Note on treatment rankings There is **no separate item for ranking probabilities**. The tool's authors state that the factors affecting rankings — unequal numbers of studies per comparison, study sample sizes, network configuration, effect sizes — are already covered by the items above. Assessors should, however, consider potential bias from rankings when judging the conclusions, since overinterpretation of rankings can bias them. ## When to Use - Assessing the risk of bias of a published network meta-analysis - Guideline development involving multiple treatment comparisons - Overviews of NMAs - Paired with ROBIS or AMSTAR 2 for the review, and RoB 2 / ROBINS-I for the primary studies -
SPIRIT.md 10.8 KB
# SPIRIT 2025 Checklist **Standard Protocol Items: Recommendations for Interventional Trials** Version: SPIRIT 2025 Source: https://www.consort-spirit.org Reference: Chan AW, Hopewell S, Moher D, et al. SPIRIT 2025 statement: updated guideline for protocols of randomised trials. BMJ 2025;389:e081477 (published simultaneously in BMJ, JAMA, Lancet, Nature Medicine, PLoS Medicine). > Note: SPIRIT 2025 supersedes SPIRIT 2013. It is a 34-item checklist (two new items, five revised, five deleted/merged) restructured with a new Open Science section. Use for clinical trial *protocols* (CONSORT is for the completed trial report). Source: Chan AW, Boutron I, Hopewell S, Moher D, Schulz KF, Collins GS, et al. SPIRIT 2025 statement: updated guideline for protocols of randomised trials. *BMJ* 2025;389:e081477 (DOI 10.1136/bmj-2024-081477). Licence: CC BY 4.0 — confirmed via Crossref. Verification: all 53 sub-items (1a–34) were compared against Table 1 of the published statement (Europe PMC full text, PMC12035670); 53/53 match, with no item missing and none invented. Two labels carried over from SPIRIT 2013 have been corrected to their 2025 names. ## Checklist Items (34 items) ### Administrative Information | # | Item | Description | |---|------|-------------| | 1a | Title | Title stating the trial design, population, and interventions, with identification as a protocol. | | 1b | Abstract | Structured summary of trial design and methods, including items from the World Health Organization Trial Registration Data Set. | | 2 | Protocol version | Version date and identifier. | | 3a | Roles — contributors | Names, affiliations, and roles of protocol contributors. | | 3b | Roles — sponsor contact | Name and contact information for the trial sponsor. | | 3c | Roles — sponsor role | Role of trial sponsor and funders in design, conduct, analysis, and reporting of trial; including any authority over these activities. | | 3d | Roles — committees | Composition, roles, and responsibilities of the coordinating site, steering committee, endpoint adjudication committee, data management team, and other individuals or groups overseeing the trial, if applicable. | ### Open Science | # | Item | Description | |---|------|-------------| | 4 | Trial registration | Name of trial registry, identifying number (with URL), and date of registration. If not yet registered, name of intended registry. | | 5 | Protocol and SAP | Where the trial protocol and statistical analysis plan can be accessed. | | 6 | Data, code, materials | Where and how the individual de-identified participant data (including data dictionary), statistical code, and any other materials will be accessible. | | 7a | Funding | Sources of funding and other support (e.g., supply of drugs). | | 7b | Conflicts of interest | Financial and other conflicts of interest for principal investigators and steering committee members. | | 8 | Dissemination | Plans to communicate trial results to participants, healthcare professionals, the public, and other relevant groups (e.g., reporting in trial registry, plain language summary, publication). | ### Introduction | # | Item | Description | |---|------|-------------| | 9a | Background | Scientific background and rationale, including summary of relevant studies (published and unpublished) examining benefits and harms for each intervention. | | 9b | Comparator choice | Explanation for choice of comparator. | | 10 | Objectives | Specific objectives related to benefits and harms. | ### Methods: Patient and Public Involvement, Trial Design | # | Item | Description | |---|------|-------------| | 11 | Patient and public involvement | Details of, or plans for, patient or public involvement in the design, conduct, and reporting of the trial. | | 12 | Trial design | Description of trial design including type of trial (e.g., parallel group, crossover), allocation ratio, and framework (e.g., superiority, equivalence, non-inferiority, exploratory). | ### Methods: Participants, Interventions, and Outcomes | # | Item | Description | |---|------|-------------| | 13 | Trial setting | Settings (e.g., community, hospital) and locations (e.g., countries, sites) where the trial will be conducted. | | 14a | Eligibility — participants | Eligibility criteria for participants. | | 14b | Eligibility — sites/deliverers | If applicable, eligibility criteria for sites and for individuals who will deliver the interventions (e.g., surgeons, physiotherapists). | | 15a | Interventions — description | Intervention and comparator with sufficient details to allow replication including how, when, and by whom they will be administered. If relevant, where additional materials describing the intervention and comparator (e.g., intervention manual) can be accessed. | | 15b | Interventions — modifications | Criteria for discontinuing or modifying allocated intervention/comparator for a trial participant (e.g., drug dose change in response to harms, participant request, or improving/worsening disease). | | 15c | Interventions — adherence | Strategies to improve adherence to intervention/comparator protocols, if applicable, and any procedures for monitoring adherence (e.g., drug tablet return, sessions attended). | | 15d | Interventions — concomitant care | Concomitant care that is permitted or prohibited during the trial. | | 16 | Outcomes | Primary and secondary outcomes, including the specific measurement variable (e.g., systolic blood pressure), analysis metric (e.g., change from baseline, final value, time to event), method of aggregation (e.g., median, proportion), and time point for each outcome. | | 17 | Harms | How harms are defined and will be assessed (e.g., systematically, non-systematically). | | 18 | Participant timeline | Time schedule of enrolment, interventions (including any run-ins and washouts), assessments, and visits for participants. A schematic diagram is highly recommended. | | 19 | Sample size | How sample size was determined, including all assumptions supporting the sample size calculation. | | 20 | Recruitment | Strategies for achieving adequate participant enrolment to reach target sample size. | ### Methods: Assignment of Interventions | # | Item | Description | |---|------|-------------| | 21a | Allocation — sequence | Who will generate the random allocation sequence and the method used. | | 21b | Allocation — restriction | Type of randomisation (simple or restricted) and details of any factors for stratification. To reduce predictability, other details of any planned restriction (e.g., blocking) should be provided in a separate document unavailable to those who enrol participants or assign interventions. | | 22 | Allocation concealment | Mechanism used to implement the random allocation sequence (e.g., central computer/telephone; sequentially numbered, opaque, sealed containers), describing any steps to conceal the sequence until interventions are assigned. | | 23 | Implementation | Whether the personnel who will enrol and those who will assign participants to the interventions will have access to the random allocation sequence. | | 24a | Blinding — who | Who will be blinded after assignment to interventions (e.g., participants, care providers, outcome assessors, data analysts). | | 24b | Blinding — how | If blinded, how blinding will be achieved and description of the similarity of interventions. | | 24c | Blinding — unblinding | If blinded, circumstances under which unblinding is permissible, and procedure for revealing a participant's allocated intervention during the trial. | ### Methods: Data Collection, Management, and Analysis | # | Item | Description | |---|------|-------------| | 25a | Data collection | Plans for assessment and collection of trial data, including any related processes to promote data quality (e.g., duplicate measurements, training of assessors) and a description of trial instruments (e.g., questionnaires, laboratory tests) along with their reliability and validity, if known. Reference to where data collection forms can be accessed, if not in the protocol. | | 25b | Retention | Plans to promote participant retention and complete follow-up, including list of any outcome data to be collected for participants who discontinue or deviate from intervention protocols. | | 26 | Data management | Plans for data entry, coding, security, and storage, including any related processes to promote data quality (e.g., double data entry; range checks for data values). Reference to where details of data management procedures can be accessed, if not in the protocol. | | 27a | Statistical methods | Statistical methods used to compare groups for primary and secondary outcomes, including harms. | | 27b | Analysis populations | Definition of who will be included in each analysis (e.g., all randomised participants), and in which group. | | 27c | Missing data | How missing data will be handled in the analysis. | | 27d | Additional analyses | Methods for any additional analyses (e.g., subgroup and sensitivity analyses). | ### Methods: Monitoring | # | Item | Description | |---|------|-------------| | 28a | Data monitoring committee | Composition of data monitoring committee (DMC); summary of its role and reporting structure; statement of whether it is independent from the sponsor and funder; conflicts of interest and reference to where further details about its charter can be found, if not in the protocol. Alternatively, an explanation of why a DMC is not needed. | | 28b | Interim analyses | Explanation of any interim analyses and stopping guidelines, including who will have access to these interim results and make the final decision to terminate the trial. | | 29 | Trial monitoring | Frequency and procedures for monitoring trial conduct. If there is no monitoring, give explanation. | ### Ethics | # | Item | Description | |---|------|-------------| | 30 | Research ethics approval | Plans for seeking research ethics committee/institutional review board approval. | | 31 | Protocol amendments | Plans for communicating important protocol modifications to relevant parties. | | 32a | Consent | Who will obtain informed consent or assent from potential trial participants or authorised proxies, and how. | | 32b | Ancillary consent | Additional consent provisions for collection and use of participant data and biological specimens in ancillary studies, if applicable. | | 33 | Confidentiality | How personal information about potential and enrolled participants will be collected, shared, and maintained in order to protect confidentiality before, during, and after the trial. | | 34 | Ancillary and post-trial care | Provisions, if any, for ancillary and post-trial care, and for compensation to those who suffer harm from trial participation. | --- *Educational summary of the SPIRIT 2025 checklist (CC BY 4.0). Cite the original statement (Chan et al., BMJ 2025) and consult https://www.consort-spirit.org for the authoritative, full checklist with explanation and elaboration.* -
SPIRIT_AI.md 4.9 KB
# SPIRIT-AI Checklist **Standard Protocol Items: Recommendations for Interventional Trials -- Artificial Intelligence Extension** Version: SPIRIT-AI 2020 (extends SPIRIT 2013) Source: https://www.consort-spirit.org · EQUATOR Network Reference: Cruz Rivera S, Liu X, Chan AW, et al. Nat Med 2020;26(9):1351-1363. doi:10.1038/s41591-020-1037-7 (CC BY 4.0) > Educational summary, authored in our own words from the CC BY 4.0 source. Use the official > SPIRIT-AI checklist for a submission-ready form and cite Cruz Rivera et al. 2020. > **Verified against the published statement.** All 15 AI items were enumerated from the article's > own checklist table via the PMC XML and compared label-by-label: **12 Extensions and > 3 Elaborations**, matching exactly with no item missing and none invented. > **Extension vs Elaboration — the statement distinguishes them, and it matters.** An **Extension** > is a *new* reporting requirement that SPIRIT 2013 does not contain. An **Elaboration** clarifies how an > existing SPIRIT 2013 item applies when the intervention involves AI; the requirement already existed, > the guidance is what is new. Both are assessed, but only the Extensions are additional obligations — > do not report an Elaboration as though the base instrument had been silent on it. Elaborations are > marked **(E)** below. ## Naming and scope (read first) - SPIRIT-AI is an **extension** of **SPIRIT 2013**, for **protocols** of randomized clinical trials of interventions that **include an AI/ML component**. Apply **both**: every base SPIRIT 2013 item plus the AI-specific items below; name and cite both instruments (manuscript-style-classical §14). - It is the **protocol** counterpart of **CONSORT-AI** (completed-trial reports). For a finished trial report use CONSORT-AI. - The AI items elaborate existing SPIRIT items (numbered to match); assess each alongside its parent. ## AI-specific extension items Status each PRESENT / PARTIAL / MISSING / N/A. ### Administrative Information | # | Item | Description (intent) | |---|------|----------------------| | 1 (i) **(E)** | AI identification | State in the title that the intervention involves AI/ML and specify the type of model. | | 1 (ii) **(E)** | Intended use | State the intended use of the AI intervention. | ### Introduction — Background and rationale | # | Item | Description (intent) | |---|------|----------------------| | 6a (i) | Role in pathway | Explain the intended use of the AI intervention within the clinical pathway, including its purpose and the intended users. | | 6a (ii) | Prior validation | Describe any pre-existing evidence for the AI intervention (e.g., prior validation studies). | ### Methods — Participants, interventions, outcomes | # | Item | Description (intent) | |---|------|----------------------| | 9 | Setting integration | Describe the onsite and offsite requirements needed to integrate the AI intervention into the trial setting. | | 10 (i) **(E)** | Participant eligibility | State the participant-level inclusion and exclusion criteria. | | 10 (ii) | Input-data eligibility | State the inclusion and exclusion criteria at the level of the **input data**. | | 11a (i) | Algorithm version | Specify which version of the AI algorithm will be used. | | 11a (ii) | Input acquisition | Specify how the input data will be acquired and selected. | | 11a (iii) | Poor-quality input | Specify how poor-quality or unavailable input data will be assessed and handled. | | 11a (iv) | Human–AI interaction | Specify any human–AI interaction in handling input data and the expertise required of the user. | | 11a (v) | AI output | Specify the output of the AI intervention. | | 11a (vi) | Output to decision | Explain how the AI intervention's outputs will contribute to decision-making or other elements of clinical practice. | ### Methods — Monitoring | # | Item | Description (intent) | |---|------|----------------------| | 22 | Error analysis plan | Specify the plans to identify and analyze performance errors, and any planned mitigation. | ### Ethics and Dissemination | # | Item | Description (intent) | |---|------|----------------------| | 29 | Code/intervention access | State whether and how the AI intervention and/or its code can be accessed, including any restrictions. | --- ## Notes for Assessors - Apply SPIRIT-AI **with** all base SPIRIT 2013 items — the AI items do not replace them. - **Algorithm version (11a (i))** and an explicit **error-analysis plan (22)** are the protocol-stage commitments that make the eventual CONSORT-AI report verifiable; mark MISSING if absent. - **Input-data eligibility (10 (ii))** is distinct from participant eligibility — a common gap. - **Setting integration (9)** and **human–AI interaction (11a (iv))** define whether the protocol evaluates the AI as it will actually be used; vague phrasing is PARTIAL. - Use **CONSORT-AI** for the completed-trial report; SPIRIT-AI is for the protocol. -
SQUIRE_2.md 6.1 KB
# SQUIRE 2.0 Checklist **Standards for QUality Improvement Reporting Excellence 2.0** Version: SQUIRE 2.0 (2015) Source: http://squire-statement.org Reference: Ogrinc G, Davies L, Goodman D, Batalden P, Davidoff F, Stevens D. SQUIRE 2.0 (Standards for QUality Improvement Reporting Excellence): revised publication guidelines from a detailed consensus process. BMJ Qual Saf. 2016;25(12):986-992. doi:10.1136/bmjqs-2015-004411 Licence: Crossref returns no Creative Commons licence. Verification: all 18 items were compared against Table 1 of the published guidelines (Europe PMC full text, PMC5256233); 18/18 match, with no item missing and none invented. Item 2's abstract specification, which had been condensed away, has been restored. ## Checklist Items (18 items) ### Title and Abstract | # | Item | Description | |---|------|-------------| | 1 | Title | Indicate that the manuscript concerns an initiative to improve healthcare (broadly defined to include the quality, safety, or value of care). | | 2 | Abstract | a) Provide adequate information to aid in searching and indexing. b) Summarise all key information from various sections of the text using the abstract format of the intended publication or a structured summary such as: background, local problem, methods, interventions, results, conclusions. | ### Introduction | # | Item | Description | |---|------|-------------| | 3 | Problem description | Nature and significance of the local problem. | | 4 | Available knowledge | Summary of what is currently known about the problem, including relevant previous studies. | | 5 | Rationale | Informal or formal frameworks, models, concepts, and/or theories used to explain the problem, any reasons or assumptions that were used to develop the intervention(s), and reasons why the intervention(s) was expected to work. | | 6 | Specific aims | Purpose of the project and of this report. | ### Methods | # | Item | Description | |---|------|-------------| | 7 | Context | Contextual elements considered important at the outset of introducing the intervention(s). | | 8 | Intervention(s) | a) Description of the intervention(s) in sufficient detail that others could reproduce it. b) Specifics of the team involved in the work. | | 9 | Study of the intervention(s) | a) Approach chosen for assessing the impact of the intervention(s). b) Approach used to establish whether the observed outcomes were due to the intervention(s). | | 10 | Measures | a) Measures chosen for studying processes and outcomes of the intervention(s), including rationale for choosing them, their operational definitions, and their validity and reliability. b) Description of the approach to the ongoing assessment of contextual elements that contributed to the success, failure, efficiency, and cost. c) Methods employed for assessing completeness and accuracy of data. | | 11 | Analysis | a) Qualitative and quantitative methods used to draw inferences from the data. b) Methods for understanding variation within the data, including the effects of time as a variable. | | 12 | Ethical considerations | Ethical aspects of implementing and studying the intervention(s) and how they were addressed, including, but not limited to, formal ethics review and potential conflict(s) of interest. | ### Results | # | Item | Description | |---|------|-------------| | 13 | Results | a) Initial steps of the intervention(s) and their evolution over time, including modifications made to the intervention during the project. b) Details of the process measures and outcome. c) Contextual elements that interacted with the intervention(s). d) Observed associations between outcomes, interventions, and relevant contextual elements. e) Unintended consequences such as unexpected benefits, problems, failures, or costs associated with the intervention(s). f) Details about missing data. | ### Discussion | # | Item | Description | |---|------|-------------| | 14 | Summary | a) Key findings, including relevance to the rationale and specific aims. b) Particular strengths of the project. | | 15 | Interpretation | a) Nature of the association between the intervention(s) and the outcomes. b) Comparison of results with findings from other publications. c) Impact of the project on people and systems. d) Reasons for any differences between observed and anticipated outcomes, including the influence of context. e) Costs and strategic trade-offs, including opportunity costs. | | 16 | Limitations | a) Limits to the generalizability of the work. b) Factors that might have limited internal validity such as confounding, bias, or imprecision in the design, methods, measurement, or analysis. c) Efforts made to minimize and adjust for limitations. | | 17 | Conclusions | a) Usefulness of the work. b) Sustainability. c) Potential for spread to other contexts. d) Implications for practice and for further study in the field. e) Suggested next steps. | ### Other Information | # | Item | Description | |---|------|-------------| | 18 | Funding | Sources of funding that supported this work. Role, if any, of the funding organization in the design, implementation, interpretation, and reporting. | --- ## Notes for Assessors - SQUIRE 2.0 is designed for studies evaluating quality improvement interventions in healthcare, including medical education QI initiatives. - Not all items will be relevant to every QI study. Items should be assessed in the context of the specific improvement effort. - Item 13 (Results) has 6 sub-items (a-f). Assess each sub-item separately for comprehensive evaluation. - Items 5 (Rationale) and 9 (Study of the intervention) are critical for distinguishing QI reports from simple descriptions of implementation. - For educational quality improvement studies, "intervention" refers to the educational program, curriculum change, or training initiative being evaluated. - SQUIRE 2.0 applies to systematic efforts to improve the quality, safety, and value of healthcare; it is not intended for basic science or clinical trials (use CONSORT for RCTs). - The original SQUIRE guidelines (1.0) were published in 2008. SQUIRE 2.0 (2015) is the current version and should be used for all new submissions. -
SRQR.md 7 KB
# SRQR Checklist (qualitative research) **Standards for Reporting Qualitative Research** Version: SRQR 2014 — 21 items across Title/Abstract, Introduction, Methods, Results/Findings, Discussion, and Other. **Broad** — applies to all qualitative approaches (ethnography, grounded theory, phenomenology, case study, narrative research), not only interviews/focus groups. Source: O'Brien BC, Harris IB, Beckman TJ, Reed DA, Cook DA. *Acad Med* 2014;89(9):1245–1251 (the SRQR statement; DOI 10.1097/ACM.0000000000000388). EQUATOR Network. Apply when the manuscript is a **qualitative study** of any approach — interviews, focus groups, observation/ethnography, document analysis, grounded theory, phenomenology, narrative research. For **interview / focus-group** studies specifically, the more granular **COREQ** (`COREQ.md`) is the better-fit companion. For the design/conduct review of the same study, pair with the QL1–QL8 domain probes in `peer-review` / `self-review` `references/domain-probes/qualitative_research.md`. > Licensing note: SRQR is published in *Academic Medicine* (© AAMC), with **no Creative Commons licence**. The items below are an **in-house, faithful summary of the reporting items (facts/intents, paraphrased — not the verbatim SRQR wording)** for item-by-item assessment; consult the published article (DOI 10.1097/ACM.0000000000000388) for exact item text. ## Reporting items (grouped by section) ### Title and Abstract | # | Item | What to check is reported | |---|------|---------------------------| | 1 | Title | A concise description of the study's nature/topic that **identifies it as qualitative** and (recommended) names the approach (e.g. ethnography, grounded theory) or data-collection method (e.g. interviews, focus groups). | | 2 | Abstract | A summary of the key elements in the journal's abstract format — typically background, purpose, methods, results, conclusions. | ### Introduction | # | Item | What to check is reported | |---|------|---------------------------| | 3 | Problem formulation | The problem/phenomenon studied and its significance, with a review of relevant theory and prior empirical work, and a problem statement. | | 4 | Purpose or research question | The study's purpose and its specific objectives or questions. | ### Methods | # | Item | What to check is reported | |---|------|---------------------------| | 5 | Qualitative approach & research paradigm | The chosen approach (ethnography / grounded theory / case study / phenomenology / narrative) and the guiding paradigm (e.g. postpositivist, constructivist/interpretivist), **with a rationale**. | | 6 | Researcher characteristics & **reflexivity** | The researchers' attributes, qualifications/experience, relationship with participants, and assumptions, and how these may have interacted with the questions, methods, findings, or transferability. | | 7 | Context | The setting/site and salient contextual factors, with a rationale. | | 8 | Sampling strategy | How and why participants/documents/events were selected, and the criterion for stopping sampling (e.g. **saturation**), with a rationale. | | 9 | Ethical issues (human subjects) | Ethics-board approval and participant consent (or an explanation for their absence), and confidentiality / data-security handling. | | 10 | Data collection methods | The data types and collection procedures — dates, iterative process, triangulation of sources/methods, and any procedure changes during the study — with a rationale. | | 11 | Data collection instruments & technologies | The instruments (e.g. interview guides, questionnaires) and devices (e.g. audio recorders), and whether/how they changed over the study. | | 12 | Units of study | The number and relevant characteristics of the participants/documents/events, and their level of participation. | | 13 | Data processing | How data were processed before/during analysis — transcription, data entry/management/security, integrity checks, coding, and de-identification of excerpts. | | 14 | Data analysis | The process by which inferences/themes were identified and developed, who analysed the data, and the referenced paradigm/approach, with a rationale. | | 15 | Techniques to enhance **trustworthiness** | The techniques used to enhance trustworthiness/credibility (e.g. member checking, audit trail, triangulation), with a rationale. | ### Results / Findings | # | Item | What to check is reported | |---|------|---------------------------| | 16 | Synthesis & interpretation | The main findings (interpretations, inferences, themes), and any theory/model development or integration with prior research. | | 17 | Links to empirical data | Evidence — participant **quotations**, field notes, text excerpts, images — that substantiates each analytic finding. | ### Discussion | # | Item | What to check is reported | |---|------|---------------------------| | 18 | Integration, implications, transferability, contribution | A short summary of findings; how they connect to / extend / challenge prior scholarship; the scope of application / **transferability**; and the study's unique contribution. | | 19 | Limitations | The trustworthiness and limitations of the findings. | ### Other | # | Item | What to check is reported | |---|------|---------------------------| | 20 | Conflicts of interest | Potential sources of influence on the study and how they were managed. | | 21 | Funding | Sources of funding/support and the role of the funders in data collection, interpretation, and reporting. | --- ## Notes for Assessors - The **highest-yield** checks (where qualitative studies most often go wrong on review): **item 6** (researcher **reflexivity** — the researchers' position, assumptions, and relationship to participants is a core qualitative-rigor requirement, frequently omitted), **item 8** (a stated **sampling strategy and stopping criterion** — purposive logic and saturation, not a convenience sample with no rationale), **item 15** (explicit **trustworthiness** techniques — member checking / audit trail / triangulation), and **item 17** (analytic claims grounded in **quoted data**, not asserted). - **Do not apply quantitative criteria** to a qualitative study: a small purposive sample is not a flaw, "generalizability" is **transferability** (not statistical external validity), and there are no power calculations, effect sizes, or p-values to demand. Mis-calibrated "the sample is too small / not generalizable / no p-value" comments are inappropriate here. - This is an **in-house faithful summary of the SRQR items (paraphrased intents, not verbatim)**; map the manuscript's content to the items rather than to exact wording. For interview/focus-group designs, COREQ (`COREQ.md`) gives finer-grained items. - Verification: all 21 items were compared, by number, name and definition, against the statement's own **Supplemental Digital Appendix 3** (item-by-item explanations and examples, `links.lww.com/ACADMED/A218`). 21/21 match, with no item missing and none invented. The supplement is marked "unauthorized reproduction is prohibited", so the wording here stays paraphrased. -
STARD.md 5.7 KB
# STARD 2015 Checklist **Standards for Reporting Diagnostic Accuracy Studies** Version: STARD 2015 — **30 items in 34 rows**; items 10, 12, 13 and 21 each split into **a** and **b**. Source: Bossuyt PM, Reitsma JB, Bruns DE, Gatsonis CA, Glasziou PP, Irwig L, et al. STARD 2015: an updated list of essential items for reporting diagnostic accuracy studies. *BMJ* 2015;351:h5527 (DOI 10.1136/bmj.h5527). Licence: CC BY 4.0 — confirmed via the PubMed Central record (PMC4623764). Item text below is reproduced from the published checklist with attribution. Verification: all 34 rows were extracted from the published checklist and compared programmatically; the id set matches exactly. > **What this file used to be.** It listed 30 rows ending at item 28. It **invented a split at item 8** > (8a/8b, where the statement has a single item), **collapsed the a/b pairs at 12, 13 and 21** into > single rows, and **omitted items 29 and 30 entirely** — where the full study protocol can be > accessed, and sources of funding and the role of funders. A reviewer scoring against it never asked > for either. ## Title or abstract | # | Checklist item | |---|----------------| | 1 | Identification as a study of diagnostic accuracy using at least one measure of accuracy (such as sensitivity, specificity, predictive values, or AUC) Abstract | ## Abstract | # | Checklist item | |---|----------------| | 2 | Structured summary of study design, methods, results, and conclusions (for specific guidance, see STARD for Abstracts) Introduction | ## Introduction | # | Checklist item | |---|----------------| | 3 | Scientific and clinical background, including the intended use and clinical role of the index test | | 4 | Study objectives and hypotheses | ## Methods — Study design | # | Checklist item | |---|----------------| | 5 | Whether data collection was planned before the index test and reference standard were performed (prospective study) or after (retrospective study) | ## Methods — Participants | # | Checklist item | |---|----------------| | 6 | Eligibility criteria | | 7 | On what basis potentially eligible participants were identified (such as symptoms, results from previous tests, inclusion in registry) | | 8 | Where and when potentially eligible participants were identified (setting, location, and dates) | | 9 | Whether participants formed a consecutive, random, or convenience series | ## Methods — Test methods | # | Checklist item | |---|----------------| | 10a | Index test, in sufficient detail to allow replication | | 10b | Reference standard, in sufficient detail to allow replication | | 11 | Rationale for choosing the reference standard (if alternatives exist) | | 12a | Definition of and rationale for test positivity cut-offs or result categories of the index test, distinguishing pre-specified from exploratory | | 12b | Definition of and rationale for test positivity cut-offs or result categories of the reference standard, distinguishing pre-specified from exploratory | | 13a | Whether clinical information and reference standard results were available to the performers or readers of the index test | | 13b | Whether clinical information and index test results were available to the assessors of the reference standard | ## Methods — Analysis | # | Checklist item | |---|----------------| | 14 | Methods for estimating or comparing measures of diagnostic accuracy | | 15 | How indeterminate index test or reference standard results were handled | | 16 | How missing data on the index test and reference standard were handled | | 17 | Any analyses of variability in diagnostic accuracy, distinguishing pre-specified from exploratory | | 18 | Intended sample size and how it was determined | ## Results — Participants | # | Checklist item | |---|----------------| | 19 | Flow of participants, using a diagram | | 20 | Baseline demographic and clinical characteristics of participants | | 21a | Distribution of severity of disease in those with the target condition | | 21b | Distribution of alternative diagnoses in those without the target condition | | 22 | Time interval and any clinical interventions between index test and reference standard | ## Results — Test results | # | Checklist item | |---|----------------| | 23 | Cross tabulation of the index test results (or their distribution) by the results of the reference standard | | 24 | Estimates of diagnostic accuracy and their precision (such as 95% confidence intervals) | | 25 | Any adverse events from performing the index test or the reference standard | ## Discussion | # | Checklist item | |---|----------------| | 26 | Study limitations, including sources of potential bias, statistical uncertainty, and generalisability | | 27 | Implications for practice, including the intended use and clinical role of the index test | ## Other information | # | Checklist item | |---|----------------| | 28 | Registration number and name of registry | | 29 | Where the full study protocol can be accessed | | 30 | Sources of funding and other support; role of funders | ## Notes for Assessors - **Score 34 rows, not 30.** The a/b halves of items 10, 12, 13 and 21 are separate requirements: the index test and the reference standard are described, cut-offs defined, blinding stated, and severity distributions reported **for each side separately**. A file that collapses them lets a study satisfy half an item and be marked complete. - **Items 29 and 30 are commonly dropped** — where the protocol can be accessed, and funding with the role of funders. - **Item 2** defers to STARD for Abstracts for the abstract's own requirements. - For a systematic review of diagnostic accuracy studies use `PRISMA_DTA.md`; for risk of bias in the included studies, `QUADAS2.md`. -
STARD_AI.md 14.2 KB
# STARD-AI Checklist **Standards for Reporting of Diagnostic Accuracy Studies — Artificial Intelligence** Version: STARD-AI 2025 Source: https://doi.org/10.1038/s41591-025-03953-8 **Reference:** STARD-AI Steering Committee, STARD-AI Consensus Group, Sounderajah V, Guni A, Liu X, Collins GS, Karthikesalingam A, Markar SR, Golub RM, Denniston AK, Shetty S, Moher D, Bossuyt PM, Darzi A, Ashrafian H. The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nat Med. 2025;31:3283-3289. PMID: 40954311 **License:** Creative Commons Attribution (CC BY) **Scope:** AI-centred diagnostic test accuracy studies. Applies to studies evaluating the diagnostic accuracy of AI systems including machine learning, deep learning models, natural language processing tools, and foundation models that generate or support diagnostic outputs. Does NOT apply to static/manually programmed rule-based systems or simple decision trees. **Relationship to STARD 2015:** STARD-AI extends STARD 2015 (Bossuyt et al. BMJ 2015). Of the 40 items, 22 are UNCHANGED from STARD 2015, 4 are MODIFIED, and 14 are NEW. Items are numbered to maintain alignment with STARD 2015 where possible. **Relationship to other guidelines:** - CONSORT-AI: for clinical trials of AI interventions - SPIRIT-AI: for trial protocols involving AI - TRIPOD+AI: for AI prediction/prognostic models - CLAIM 2024: for AI/ML in medical imaging - Use STARD-AI when diagnostic accuracy of an AI system is the primary focus --- ## Checklist Items (40 items; items 15, 16, 17, 26, and 40 have sub-items) ### Title and Abstract | # | Item | Description | Status vs STARD 2015 | |---|------|-------------|---------------------| | 1 | Title | Identification as a study reporting AI-centered diagnostic accuracy and reporting at least one measure of accuracy within title or abstract. | MODIFIED | | 2 | Abstract | Structured summary of study design, methods, results and conclusions (for specific guidance, see STARD for Abstracts). | UNCHANGED | ### Introduction | # | Item | Description | Status vs STARD 2015 | |---|------|-------------|---------------------| | 3 | Scientific background | Scientific and clinical background, including the intended use of the index test, whether it is novel or an established index test and its integration into an existing or new workflow, if applicable. | MODIFIED | | 4 | Objectives | Study objectives and hypotheses. | UNCHANGED | ### Methods #### Study Design | # | Item | Description | Status vs STARD 2015 | |---|------|-------------|---------------------| | 5 | Study design | Whether data collection was planned before the index test and reference standard were performed (prospective study) or after (retrospective study). | UNCHANGED | #### Ethics | # | Item | Description | Status vs STARD 2015 | |---|------|-------------|---------------------| | 6 | Ethics approval | Formal approval from an ethics committee. If not required, justify why. | NEW | #### Participants | # | Item | Description | Status vs STARD 2015 | |---|------|-------------|---------------------| | 7 | Eligibility criteria | Eligibility criteria: listing separate inclusion and exclusion criteria in the order that they are applied at both participant level and data level. | MODIFIED | | 8 | Participant identification basis | On what basis potentially eligible participants were identified (such as symptoms, results from previous tests and inclusion in registry). | UNCHANGED | | 9 | Setting and dates | Where and when potentially eligible participants were identified (setting, location and dates). | UNCHANGED | | 10 | Participant series | Whether participants formed a consecutive, random or convenience series. | UNCHANGED | #### Dataset | # | Item | Description | Status vs STARD 2015 | |---|------|-------------|---------------------| | 11 | Data source | Source of the data and whether they have been routinely collected, specifically collected for the purpose of the study or acquired from an open-source repository. | NEW | | 12 | Dataset annotation | Who undertook the annotations for the dataset (including experience levels and background) and how (within the same clinical context or in a post hoc fashion), if applicable. | NEW | | 13 | Data capture devices and software | Devices (manufacturer and model) that were used to capture data; software (with version number) used to engineer the index test, highlighting the intended use. | NEW | | 14 | Data acquisition and preprocessing | Data acquisition protocols (for example, contrast protocol or reconstruction method for medical images) and details of data preprocessing, in sufficient detail to allow replication. | NEW | #### Test Methods | # | Item | Description | Status vs STARD 2015 | |---|------|-------------|---------------------| | 15a | Index test | Index test, in sufficient detail to allow replication. | UNCHANGED | | 15b | Index test development | How the index test was developed, including any training, validation, testing and external evaluation, detailing sample sizes, when applicable. | NEW | | 15c | Index test cut-offs | Definition of and rationale for test positivity cutoffs or result categories of the index test, distinguishing prespecified from exploratory. | UNCHANGED | | 15d | End-user specification | The specified end-user of the index test and the level of expertise required of users. | NEW | | 16a | Reference standard | Reference standard, in sufficient detail to allow replication. | UNCHANGED | | 16b | Reference standard rationale | Rationale for choosing the reference standard (if alternatives exist). | UNCHANGED | | 16c | Reference standard cut-offs | Definition of and rationale for test positivity cutoffs or result categories of the reference standard, distinguishing prespecified from exploratory. | UNCHANGED | | 17a | Blinding (index test) | Whether clinical information and reference standard results were available to the performers or readers of the index test. | UNCHANGED | | 17b | Blinding (reference standard) | Whether clinical information and index test results were available to the assessors of the reference standard. | UNCHANGED | #### Analysis | # | Item | Description | Status vs STARD 2015 | |---|------|-------------|---------------------| | 18 | Statistical methods | Methods for estimating or comparing measures of diagnostic accuracy. | UNCHANGED | | 19 | Indeterminate results | How indeterminate index test or reference standard results were handled. | UNCHANGED | | 20 | Missing data | How missing data on the index test and reference standard were handled. | UNCHANGED | | 21 | Variability analyses | Any analyses of variability in diagnostic accuracy, distinguishing prespecified from exploratory. | UNCHANGED | | 22 | Sample size | Intended sample size and how it was determined. | UNCHANGED | | 23 | Fairness assessment (methods) | Details of any performance error analysis and algorithmic bias and fairness assessments, if undertaken. | NEW | ### Results #### Participants and Dataset | # | Item | Description | Status vs STARD 2015 | |---|------|-------------|---------------------| | 24 | Flow diagram | Flow of participants, using a diagram. | UNCHANGED | | 25 | Baseline characteristics | Baseline demographic, clinical and technical characteristics of training, validation and test sets, if applicable. | MODIFIED | | 26a | Distribution of severity | Distribution of severity of disease in those with the target condition. | UNCHANGED | | 26b | Alternative diagnoses | Distribution of alternative diagnoses in those without the target condition. | UNCHANGED | | 27 | Time interval | Time interval and any clinical interventions between index test and reference standard. | UNCHANGED | | 28 | Test set representativeness | Whether the datasets represent the distribution of the target condition that one would expect from the intended use population. | NEW | | 29 | External evaluation differences | For external evaluation on an independent dataset, an assessment of how this differs from the training, validation and test sets. | NEW | #### Test Results | # | Item | Description | Status vs STARD 2015 | |---|------|-------------|---------------------| | 30 | Cross-tabulation | Cross-tabulation of the index test results (or their distribution) by the results of the reference standard. | UNCHANGED | | 31 | Accuracy estimates | Estimates of diagnostic accuracy and their precision (such as 95% confidence intervals). | UNCHANGED | | 32 | Adverse events | Any adverse events from performing the index test or the reference standard. | UNCHANGED | ### Discussion | # | Item | Description | Status vs STARD 2015 | |---|------|-------------|---------------------| | 33 | Study limitations | Study limitations, including sources of potential bias, statistical uncertainty and generalizability. | UNCHANGED | | 34 | Clinical applicability | Implications for practice, including the intended use and clinical role of the index test. | UNCHANGED | | 35 | Ethical considerations and fairness | Ethical considerations and adherence to ethical standards associated with the use of the index test and issues of fairness. | NEW | ### Other Important Information | # | Item | Description | Status vs STARD 2015 | |---|------|-------------|---------------------| | 36 | Registration | Registration number and name of registry. | UNCHANGED | | 37 | Protocol | Where the full study protocol can be accessed. | UNCHANGED | | 38 | Funding | Sources of funding and other support; role of funders. | UNCHANGED | | 39 | Commercial interests | Commercial interests, if applicable. | NEW | | 40a | Data and code availability | Availability of datasets and code, detailing any restrictions on their reuse and repurposing. | NEW | | 40b | Audit and evaluation | Whether outputs are stored, auditable and available for evaluation, if necessary. | NEW | --- ## Summary of Changes from STARD 2015 ### MODIFIED Items (4) | STARD-AI Item | Change Summary | |---------------|---------------| | 1 (Title) | Added requirement to identify the study as AI-centered diagnostic accuracy | | 3 (Introduction) | Added requirement for intended use, novelty, and workflow integration | | 7 (Eligibility criteria) | Expanded to include eligibility at both participant and data level | | 25 (Baseline characteristics) | Expanded to include characteristics of training, validation, and test sets | ### NEW Items (14) | STARD-AI Item | Topic | |---------------|-------| | 6 | Ethics approval | | 11 | Data source and collection method | | 12 | Dataset annotation | | 13 | Data capture devices and software | | 14 | Data acquisition and preprocessing | | 15b | Index test development (training, validation, testing) | | 15d | End-user specification and expertise level | | 23 | Fairness assessment methods | | 28 | Test set representativeness | | 29 | External evaluation differences | | 35 | Ethical considerations and fairness | | 39 | Commercial interests | | 40a | Data and code availability | | 40b | Audit and evaluation of outputs | --- ## Notes for Assessors ### When to Use STARD-AI vs Other AI Guidelines - **STARD-AI**: When the study evaluates diagnostic accuracy of an AI system as the primary outcome. Includes imaging AI, LLM-based diagnostic tools, pathology AI, EHR-based diagnostic systems. - **TRIPOD+AI / TRIPOD-LLM**: When the study develops or validates a prediction/prognostic model using AI/ML. - **CLAIM 2024**: When the study develops or validates an AI model specifically for medical imaging. - **CONSORT-AI**: When the study is a clinical trial of an AI intervention. - Authors may refer to multiple checklists and select the one most aligned with the study's primary aim. ### Key Assessment Points 1. **Item 1 (Title)**: The title must explicitly identify the study as reporting AI-centered diagnostic accuracy. Merely mentioning "machine learning" or "deep learning" without connecting it to diagnostic accuracy is PARTIAL. 2. **Item 7 (Eligibility)**: STARD-AI requires eligibility criteria at BOTH participant level (e.g., age, diagnosis) AND data level (e.g., image quality, scanner type). Reporting only participant-level criteria is PARTIAL. 3. **Item 12 (Annotation)**: Must include who annotated, their experience levels and background, and how (within clinical context or post hoc). Missing any of these elements is PARTIAL. 4. **Item 14 (Data acquisition)**: Must describe data acquisition protocols and preprocessing in sufficient detail to allow replication. 5. **Item 23 & 35 (Fairness/Ethics)**: These are paired items — methods/assessment in Item 23, ethical considerations and fairness discussion in Item 35. Both must be present. If the study claims fairness was not assessed, this should be explicitly stated with justification. 6. **Item 25 (Baseline characteristics)**: For AI studies, characteristics must be reported for EACH dataset partition (training, validation, test), not just overall. 7. **Items 39-40b (Other information)**: Commercial interests (39), data/code availability (40a), and audit/evaluation (40b) are NEW and commonly missing. These are increasingly required by journals. ### Relationship to MI-CLEAR-LLM If the AI system under evaluation is an LLM (e.g., GPT-4, Claude, Gemini), apply MI-CLEAR-LLM (6 items) alongside STARD-AI. MI-CLEAR-LLM addresses LLM-specific concerns (stochasticity, prompt documentation, data contamination) that are not covered by STARD-AI. ### Common Gaps in AI Diagnostic Accuracy Studies Based on the systematic review that informed STARD-AI development (Aggarwal et al. npj Digit Med 2021): 1. Missing annotation process details 2. Missing data preprocessing description 3. No fairness/bias assessment 4. No external validation or test set representativeness discussion 5. Missing model architecture and training details 6. No commercial interest disclosure 7. No data/code availability statement --- ## Verification Note This checklist was verified against the published Table 2 of the STARD-AI paper in Nature Medicine (2025;31:3283-3289, DOI: 10.1038/s41591-025-03953-8). Item numbering, descriptions, and NEW/MODIFIED/UNCHANGED classifications match the published version. Verified 2026-04-11 via Nature Medicine online full-size Table 2 (https://www.nature.com/articles/s41591-025-03953-8/tables/2) and STARD 2015 checklist (EQUATOR Network). The STARD 2015 original checklist (Bossuyt et al. BMJ 2015) was used to confirm UNCHANGED items. -
STROBE.md 6.8 KB
# STROBE Checklist **Strengthening the Reporting of Observational Studies in Epidemiology** Version: STROBE 2007 (combined checklist for cohort, case-control, and cross-sectional studies) Source: https://www.strobe-statement.org Source: von Elm E, Altman DG, Egger M, Pocock SJ, Gøtzsche PC, Vandenbroucke JP. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement. *PLoS Med* 2007;4(10):e296 (DOI 10.1371/journal.pmed.0040296). Licence: CC BY 4.0 — confirmed via the PubMed Central record (PMC2020495). Verification: all 34 sub-items (1a–22) were compared against the official combined STROBE checklist PDF distributed by the EQUATOR Network (`STROBE_checklist_v4_combined.pdf`), parsed sub-item by sub-item. 34/34 match, with no item missing and none invented. The statement's own PubMed Central record carries its checklist as an image, not as text, which is why the audit used the distributed checklist rather than the article. ## Checklist Items (22 items) ### Title and Abstract | # | Item | Description | |---|------|-------------| | 1a | Title | Indicate the study's design with a commonly used term in the title or abstract. | | 1b | Abstract | Provide in the abstract an informative and balanced summary of what was done and what was found. | ### Introduction | # | Item | Description | |---|------|-------------| | 2 | Background/rationale | Explain the scientific background and rationale for the investigation being reported. | | 3 | Objectives | State specific objectives, including any prespecified hypotheses. | ### Methods | # | Item | Description | |---|------|-------------| | 4 | Study design | Present key elements of study design early in the paper. | | 5 | Setting | Describe the setting, locations, and relevant dates, including periods of recruitment, exposure, follow-up, and data collection. | | 6a | Participants | *Cohort study*: Give the eligibility criteria, and the sources and methods of selection of participants. Describe methods of follow-up. *Case-control study*: Give the eligibility criteria, and the sources and methods of case ascertainment and control selection. Give the rationale for the choice of cases and controls. *Cross-sectional study*: Give the eligibility criteria, and the sources and methods of selection of participants. | | 6b | Participants | *Cohort study*: For matched studies, give matching criteria and number of exposed and unexposed. *Case-control study*: For matched studies, give matching criteria and the number of controls per case. | | 7 | Variables | Clearly define all outcomes, exposures, predictors, potential confounders, and effect modifiers. Give diagnostic criteria, if applicable. | | 8 | Data sources/measurement | For each variable of interest, give sources of data and details of methods of assessment (measurement). Describe comparability of assessment methods if there is more than one group. | | 9 | Bias | Describe any efforts to address potential sources of bias. | | 10 | Study size | Explain how the study size was arrived at. | | 11 | Quantitative variables | Explain how quantitative variables were handled in the analyses. If applicable, describe which groupings were chosen and why. | | 12a | Statistical methods | Describe all statistical methods, including those used to control for confounding. | | 12b | Statistical methods | Describe any methods used to examine subgroups and interactions. | | 12c | Statistical methods | Explain how missing data were addressed. | | 12d | Statistical methods | *Cohort study*: If applicable, explain how loss to follow-up was addressed. *Case-control study*: If applicable, explain how matching of cases and controls was addressed. *Cross-sectional study*: If applicable, describe analytical methods taking account of sampling strategy. | | 12e | Statistical methods | Describe any sensitivity analyses. | ### Results | # | Item | Description | |---|------|-------------| | 13a | Participants | Report numbers of individuals at each stage of study -- e.g., numbers potentially eligible, examined for eligibility, confirmed eligible, included in the study, completing follow-up, and analysed. | | 13b | Participants | Give reasons for non-participation at each stage. | | 13c | Participants | Consider use of a flow diagram. | | 14a | Descriptive data | Give characteristics of study participants (e.g., demographic, clinical, social) and information on exposures and potential confounders. | | 14b | Descriptive data | Indicate number of participants with missing data for each variable of interest. | | 14c | Descriptive data | *Cohort study*: Summarise follow-up time (e.g., average and total amount). | | 15 | Outcome data | *Cohort study*: Report numbers of outcome events or summary measures over time. *Case-control study*: Report numbers in each exposure category, or summary measures of exposure. *Cross-sectional study*: Report numbers of outcome events or summary measures. | | 16a | Main results | Give unadjusted estimates and, if applicable, confounder-adjusted estimates and their precision (e.g., 95% confidence interval). Make clear which confounders were adjusted for and why they were included. | | 16b | Main results | Report category boundaries when continuous variables were categorized. | | 16c | Main results | If relevant, consider translating estimates of relative risk into absolute risk for a meaningful time period. | | 17 | Other analyses | Report other analyses done -- e.g., analyses of subgroups and interactions, and sensitivity analyses. | ### Discussion | # | Item | Description | |---|------|-------------| | 18 | Key results | Summarise key results with reference to study objectives. | | 19 | Limitations | Discuss limitations of the study, taking into account sources of potential bias or imprecision. Discuss both direction and magnitude of any potential bias. | | 20 | Interpretation | Give a cautious overall interpretation of results considering objectives, limitations, multiplicity of analyses, results from similar studies, and other relevant evidence. | | 21 | Generalisability | Discuss the generalisability (external validity) of the study results. | ### Other Information | # | Item | Description | |---|------|-------------| | 22 | Funding | Give the source of funding and the role of the funders for the present study and, if applicable, for the original study on which the present article is based. | --- ## Notes for Assessors - Items 6b, 12d, and 14c have study-design-specific wording. Choose the version matching the study design. - Item 13c (flow diagram) is a recommendation, not a strict requirement, but strongly encouraged. - For studies involving AI/ML, consider also checking against TRIPOD+AI or CLAIM as applicable. - STROBE extensions exist for specific designs: STROBE-ME (molecular epidemiology), STROBE-NI (neonatal infections), RECORD (routinely collected health data). -
STROBE_MR.md 10.4 KB
# STROBE-MR Checklist **Strengthening the Reporting of Observational Studies in Epidemiology using Mendelian Randomization** Version: STROBE-MR 2021 — **20 items and 30 sub-items** across six sections. Source (the statement): Skrivankova VW, Richmond RC, Woolf BAR, Yarmolinsky J, Davies NM, Swanson SA, et al. Strengthening the Reporting of Observational Studies in Epidemiology Using Mendelian Randomization: The STROBE-MR Statement. *JAMA* 2021;326(16):1614-1621 (DOI 10.1001/jama.2021.18236). Explanation and Elaboration: Skrivankova VW, et al. *BMJ* 2021;375:n2233. Also: https://www.strobe-mr.org > **Fidelity and licence.** The statement is published in *JAMA* (© American Medical Association) under > **no open licence**. The descriptions below state what each item asks **in our own words**; the > structure — 20 items, which sub-items exist and under which item, the section grouping — is > **verified against the statement**. **Complete the official checklist for anything you submit**, and > read the Explanation and Elaboration for the rationale and worked examples. > **What this file used to be.** It cited "Skidmore ME / Davey Smith G, Davies NM, et al. *BMJ* > 2021;375:n2233" as the statement — a non-existent first author, and the E&E paper rather than the > statement. It carried **items 1–20 with none of the 30 sub-items**, which is where every > MR-specific requirement lives. From item 17 on it was shifted by one (Funding at 17 instead of 18) > and it **invented an item 20, "Other"**, where the statement has Conflicts of interest. It also > claimed the statement was CC BY, and carried a "Verified" stamp. None of that survived contact with > the published table. **Scope.** Covers 1-sample and 2-sample MR, single or multiple exposures and outcomes, and MR following a GWAS reported in the same article. Does **not** cover GWAS themselves (use STREGA), sequencing or expression studies, or ordinary observational epidemiology (use `STROBE.md`). For MR that does not use instrumental-variable estimation — some gene-by-environment interaction studies — some items will not apply. **Relationship to STROBE.** A stand-alone extension: STROBE has 22 items and 18 sub-items, STROBE-MR has 20 items and 30 sub-items. Every item and sub-item was modified for MR **except sub-item 6d** (missing data), which is unchanged. Name both instruments in Methods. **Not a quality instrument.** The statement says explicitly that the checklist is not to be used to evaluate the quality of MR research. ## Title and Abstract | # | Item | What to check is reported | |---|------|---------------------------| | 1 | Title and abstract | MR is named as the study design in the title and/or abstract, where that is a main purpose of the study. | ## Introduction | # | Item | What to check is reported | |---|------|---------------------------| | 2 | Background | The scientific background and rationale; what the exposure is; whether a causal exposure–outcome relationship is plausible; and **a justification of why MR helps answer this question**. | | 3 | Objectives | Specific objectives, including **prespecified causal hypotheses** if any, and a statement that MR estimates causal effects only under stated assumptions. | ## Methods | # | Item | What to check is reported | |---|------|---------------------------| | 4 | Study design and data sources | Key design elements early in the article; consider a table of data sources for every phase. Then, **for each data source**: | | 4a | — Setting | The design and underlying population; setting, locations and relevant dates, including recruitment, exposure, follow-up and data collection. | | 4b | — Participants | Eligibility criteria, how participants were selected, sample size, and whether any power or sample-size calculation was done **before** the main analysis. | | 4c | — Genetic variants | **Measurement, quality control and selection of the genetic variants.** | | 4d | — Variables | Assessment methods and diagnostic criteria for each exposure, outcome and other relevant variable. | | 4e | — Ethics | Ethics committee approval and participant informed consent, where relevant. | | 5 | Assumptions | **The three core IV assumptions stated explicitly** — relevance, independence, exclusion restriction — plus the assumptions of any additional or sensitivity analysis. | | 6 | Statistical methods: main analysis | The statistical methods and statistics used: | | 6a | — Quantitative variables | How they were handled — scale, units, model. | | 6b | — Genetic variants | How variants were handled and, if applicable, how their weights were chosen. | | 6c | — MR estimator | **Which estimator** (two-stage least squares, Wald ratio, …) and its related statistics; the covariates included; and for 2-sample MR whether the same covariate set was used in both samples. | | 6d | — Missing data | How missing data were addressed. *(The only sub-item unchanged from STROBE.)* | | 6e | — Multiple testing | How multiple testing was addressed, if applicable. | | 7 | Assessment of assumptions | The methods or prior knowledge used to **assess** the assumptions or justify their validity. | | 8 | Sensitivity and additional analyses | Any sensitivity or additional analyses — comparison of estimates from different approaches, independent replication, bias-analytic techniques, instrument validation, simulations. | | 9 | Software and preregistration | | | 9a | — Software | Statistical software and packages, **with version and settings**. | | 9b | — Preregistration | **Whether the protocol and details were preregistered**, and when and where. | ## Results | # | Item | What to check is reported | |---|------|---------------------------| | 10 | Descriptive data | | | 10a | — Flow | Numbers of individuals at each stage of the included studies and reasons for exclusion; consider a flow diagram. | | 10b | — Summary statistics | For phenotypic exposures, outcomes and other relevant variables — means, SDs, proportions. | | 10c | — Heterogeneity | Where data sources include meta-analyses of previous studies, the assessments of heterogeneity across them. | | 10d | — 2-sample MR | (i) **Justification that the variant–exposure associations are similar** between the exposure and outcome samples; (ii) **the number of individuals overlapping** between the two studies. | | 11 | Main results | | | 11a | — Associations | The variant–exposure and variant–outcome associations, preferably on an interpretable scale. | | 11b | — MR estimates | The MR estimate of the exposure–outcome relationship with its uncertainty, on an interpretable scale such as an odds ratio or relative risk per SD. | | 11c | — Absolute risk | Where relevant, relative risk translated into absolute risk over a meaningful period. | | 11d | — Plots | Consider plots — forest plot, or variant–outcome against variant–exposure associations. | | 12 | Assessment of assumptions | | | 12a | — Validity | The assessment of the validity of the assumptions. | | 12b | — Statistics | Additional statistics — heterogeneity across variants (I², Q) or an E-value. | | 13 | Sensitivity and additional analyses | | | 13a | — Robustness | Sensitivity analyses testing robustness to violations of the assumptions. | | 13b | — Other | Results of other sensitivity or additional analyses. | | 13c | — Direction | **Any assessment of the direction of the causal relationship**, e.g. bidirectional MR. | | 13d | — Non-MR comparison | Where relevant, comparison with estimates from non-MR analyses. | | 13e | — Plots | Consider additional plots, e.g. leave-one-out analyses. | ## Discussion | # | Item | What to check is reported | |---|------|---------------------------| | 14 | Key results | The key results summarised against the study objectives. | | 15 | Limitations | Limitations, taking in the validity of the IV assumptions, other sources of bias, and imprecision — with **the direction and magnitude** of any potential bias and what was done about it. | | 16 | Interpretation | | | 16a | — Meaning | A cautious overall interpretation in the light of the limitations and of other studies. | | 16b | — Mechanism | The biological mechanisms that could drive the relationship, and **whether the gene-environment equivalence assumption is reasonable**; causal language used carefully, making clear that IV estimates are causal only under certain assumptions. | | 16c | — Clinical relevance | Whether the results have clinical or public-policy relevance, and what they imply about the size of possible interventions. | | 17 | Generalizability | Generalisability of the results **(a) to other populations, (b) across other exposure periods or timings, and (c) across other levels of exposure**. | ## Other Information | # | Item | What to check is reported | |---|------|---------------------------| | 18 | Funding | Sources of funding and the role of funders for this study and, where applicable, for the databases and original studies it rests on. | | 19 | Data and data sharing | The data used, or where and how it can be accessed, referenced in the article; and **the statistical code needed to reproduce the results**, or where it is publicly accessible. | | 20 | Conflicts of interest | Declared by **all** authors. | ## Notes for Assessors - **The sub-items are the extension.** Items 1–20 alone are close to a re-lettered STROBE. What makes a report an MR report is 4c (variant measurement, QC and selection), 6c (the estimator), 9b (preregistration), 10d (2-sample similarity and **participant overlap**), 12a/12b (assumption validity, heterogeneity across variants), 13c (direction of causation), and 16b (gene-environment equivalence). Score them. - **Item 5 is the one most often skipped**: the three IV assumptions named explicitly — relevance, independence, exclusion restriction — not gestured at. - A study reporting only an inverse-variance-weighted estimate, with no assessment of the assumptions (7, 12) and no sensitivity analyses (8, 13), is non-compliant regardless of how well it reads. - **Name both instruments.** STROBE-MR is a stand-alone extension of STROBE; cite each. See `STROBE.md` for the base items. - Authors are expected to address every item and sub-item, using supplementary material where space is short. - For the design-validity review of the same study, pair with the MR domain probes in `peer-review` / `self-review` `references/domain-probes/mendelian_randomization.md`; for the analysis, with `analyze-stats` `analysis_guides/mendelian_randomization.md`. -
SWiM.md 4.1 KB
# SWiM Checklist **Synthesis Without Meta-analysis (SWiM) in systematic reviews: reporting guideline** Version: SWiM 2020 Source: Campbell M et al. BMJ 2020;368:l6890. doi: 10.1136/bmj.l6890 Website: https://swim.sphsu.gla.ac.uk/ Licence: CC BY 4.0 — confirmed via Crossref. Verification: all 9 items were compared against Table 1 of the published statement (Europe PMC full text, PMC7190266). The nine labels matched; **six descriptions did not** and have been replaced with the published text — item 5 previously stated a different requirement altogether, item 7 dropped its second element, and items 2, 4 and 8 dropped required elements (rationale for the metric, justification for the prioritisation criteria, certainty of the findings). ## Checklist Items (9 items) ### Reporting Items | # | Item | Description | |---|------|-------------| | 1 | Grouping studies for synthesis | **1a)** Provide a description of, and rationale for, the groups used in the synthesis (e.g., groupings of populations, interventions, outcomes, study design). **1b)** Detail and provide rationale for any changes made subsequent to the protocol in the groups used in the synthesis. | | 2 | Describe the standardised metric and transformation methods used | Describe the standardised metric for each outcome. Explain why the metric(s) was chosen and describe any methods used to transform the intervention effects, as reported in the study, to the standardised metric, citing any methodological guidance consulted. | | 3 | Describe the synthesis methods | Describe and justify the methods used to synthesise the effects for each outcome when it was not possible to undertake a meta-analysis of effect estimates. | | 4 | Criteria used to prioritise results for summary and synthesis | Where applicable, provide the criteria used, with supporting justification, to select the particular studies, or a particular study, for the main synthesis or to draw conclusions from the synthesis (e.g., based on study design, risk of bias assessments, directness in relation to the review question). | | 5 | Investigation of heterogeneity in reported effects | State the method(s) used to examine heterogeneity in reported effects when it was not possible to undertake a meta-analysis of effect estimates and its extensions to investigate heterogeneity. | | 6 | Certainty of evidence | Describe the methods used to assess the certainty of the synthesis findings. | | 7 | Data presentation methods | Describe the graphical and tabular methods used to present the effects (e.g., tables, forest plots, harvest plots). Specify key study characteristics (e.g., study design, risk of bias) used to order the studies, in the text and any tables or graphs, clearly referencing the studies included. | | 8 | Reporting results | For each comparison and outcome, provide a description of the synthesised findings and the certainty of the findings. Describe the result in language that is consistent with the question the synthesis addresses, and indicate which studies contribute to the synthesis. | | 9 | Limitations of the synthesis | Report the limitations of the synthesis methods used and/or the groupings used in the synthesis and how these affect the conclusions that can be drawn in relation to the original review question. | The synthesis methods item 3 refers to include vote counting based on direction of effect, combining P values, calculating the median effect size, and combining confidence intervals — the statement's Table 2 maps which of these can answer which question, given the available data. --- ## Notes for Assessors - SWiM complements PRISMA 2020 — it provides additional guidance for reporting when meta-analysis is not performed - We estimate that 35% of health-related systematic reviews do not do meta-analysis - Items currently available in PRISMA (Items 13, 14, and 21) and RAMESES (Items 15 and 19) also apply - SWiM items should accompany PRISMA items, not replace them - The guideline applies to systematic reviews using alternative synthesis methods: vote counting based on direction, combining P values, narrative synthesis, range of effects, etc. - Not applicable to qualitative evidence syntheses (use ENTREQ instead) -
TARGET.md 7.9 KB
# TARGET Checklist **Transparent Reporting of Observational Studies Emulating a Target Trial** Version: TARGET 2025 (21 items across 6 sections; items 6 and 7 pair the target-trial *specification* with its *emulation* in the data) Source: In-house faithful summary of the TARGET item intents (own-words paraphrase, not verbatim). Cashin AG, Hansford HJ, Hernán MA, et al. Transparent Reporting of Observational Studies Emulating a Target Trial: The TARGET Statement. JAMA 2025;334(12):1084-1093. DOI 10.1001/jama.2025.13350. Official checklist: https://target-guideline.org. Complete the official TARGET instrument for a submission checklist. Pairs with the `/design-study` target-trial-emulation design module. Licence: *JAMA* (© American Medical Association) — no open licence. Verification: all 39 sub-items across the 21 numbered items were compared, by number and order, against the checklist tables of the published statement (PubMed Central record PMC13084563). 39/39 are present, including the paired 6a–h specification and 7a–7h(ii) emulation columns. Three items dropped a required clause (6c, 6d, 7d) and are corrected. Wording stays paraphrased — the statement is © American Medical Association with no open licence. ## Checklist Items (21 items) ### Title and Abstract | # | Item | Description | |---|------|-------------| | 1a | Study type | Identify that the study attempts to emulate a target trial using observational data. | | 1b | Data sources | Report the data sources used for the emulation. | | 1c | Key elements | Summarize the key assumptions, statistical methods, findings, and conclusions. | ### Introduction | # | Item | Description | |---|------|-------------| | 2 | Background | Describe the scientific background and the gap in knowledge the study addresses. | | 3 | Causal question | Summarize the causal question specified by the target-trial protocol. | | 4 | Rationale | Describe the rationale for emulating a target trial with the available data. | ### Methods | # | Item | Description | |---|------|-------------| | 5 | Data sources | Cite the data sources and describe their original purpose, type, geographic locations, setting, and time period. | #### Target-trial specification (the protocol you would run) | # | Item | Description | |---|------|-------------| | 6a | Eligibility criteria | Describe the eligibility criteria defining the target population. | | 6b | Treatment strategies | Describe the treatment strategies to be compared, in sufficient detail (e.g., dose, duration, start/stop rules). | | 6c | Assignment | Report that eligible individuals would be randomly assigned to the treatment strategies, and may be aware of their treatment allocation. | | 6d | Follow-up | Clarify that follow-up would start at the time of assignment to the treatment strategies, and specify when follow-up would end. | | 6e | Outcomes | Describe the outcomes, including their measurement and timing. | | 6f | Causal contrasts | Describe the causal contrasts of interest, including the effect measures. | | 6g | Identifying assumptions | Describe the assumptions that would be made to identify each causal estimand. | | 6h | Data analysis plan | For each causal estimand, describe the data-analysis procedures and statistical models. | #### Target-trial emulation (mapping to the observational data) | # | Item | Description | |---|------|-------------| | 7a | Eligibility (emulation) | Describe how the eligibility criteria were operationalized with the data. | | 7b | Treatment strategies (emulation) | Describe how the treatment strategies were operationalized with the data. | | 7c | Assignment (emulation) | Describe how assignment to treatment strategies was operationalized with the data. | | 7d | Follow-up (emulation) | Clarify that follow-up starts at the time individuals were assigned to the treatment strategies, and describe how the end of follow-up was operationalized with the data. *(Misaligning eligibility, assignment and start of follow-up is what introduces immortal-time bias.)* | | 7e | Outcomes (emulation) | Describe how the outcomes were operationalized with the data. | | 7f | Causal contrasts (emulation) | Describe how the causal contrasts were operationalized with the data. | | 7g(i) | Identifying assumptions (emulation) | For each causal estimand, describe the assumptions made, including baseline confounding. | | 7g(ii) | Assumption variables | Describe how the variables related to those assumptions were operationalized. | | 7h(i) | Data analysis (emulation) | Describe modifications to the data-analysis methods needed for the observational emulation. | | 7h(ii) | Sensitivity analyses | Describe any additional analyses assessing the sensitivity of results to the operationalization choices. | ### Results | # | Item | Description | |---|------|-------------| | 8 | Participant selection | Report the numbers of individuals assessed for eligibility, eligible, and assigned to each treatment strategy. | | 9 | Baseline data | Describe the distribution of baseline characteristics of individuals, by treatment strategy. | | 10 | Follow-up | Summarize the length of follow-up and describe the reasons for its end. | | 11 | Missing data | Describe the frequency of missing data in all variables, by treatment strategy. | | 12 | Outcomes | Describe the frequency or distribution of each outcome, by treatment strategy. | | 13 | Effect estimates | Report the effect estimate for each causal contrast, with its corresponding measure of precision. | | 14 | Additional analyses | Report the results of all analyses assessing the sensitivity of the estimates to the choices made. | ### Discussion | # | Item | Description | |---|------|-------------| | 15 | Interpretation | Provide an interpretation of the key findings in the context of the causal question. | | 16 | Limitations | Discuss limitations, considering differences between the target trial and its emulation. | ### Other Information | # | Item | Description | |---|------|-------------| | 17 | Ethics | Provide the institutional review board or ethics committee approval information. | | 18 | Registration | State whether, when, and where the study protocol was registered. | | 19 | Data sharing | State whether the data, analytic code, and materials are accessible. | | 20 | Funding | Provide the sources of funding and detail the role of the funders. | | 21 | Conflicts of interest | State any conflicts of interest and financial disclosures for all authors. | --- ## Notes for Assessors - The distinctive TARGET structure is the paired **specification (item 6)** and **emulation (item 7)**: for each protocol element — eligibility, treatment strategies, assignment, start of follow-up, outcomes, causal contrast, identifying assumptions, analysis — the study must state both the target-trial version and how it was operationalized in the data. A study that reports the emulation without ever specifying the target trial it emulates is a gap. - The single most consequential defect this catches is **time-zero misalignment → immortal-time bias** (items 6d / 7d): eligibility, treatment assignment, and start of follow-up must coincide. - Items 6g / 7g require an **explicit causal estimand and its identifying assumptions (including baseline confounding)** — an association reported with no stated estimand or assumptions is a gap, not merely thin reporting. - TARGET is the **reporting** side; pair it with the `/design-study` **target-trial-emulation module** (the design side, which enforces the same seven-component protocol before data extraction). Use **RECORD / STROBE** for the routinely-collected-data and general observational items not specific to the emulation. - Not needed for a purely descriptive, prevalence, or diagnostic-accuracy study — TARGET applies to a **causal / comparative-effectiveness** question emulated on observational data. - The vendored checklist is an educational own-words summary of item intent; complete the official TARGET instrument (target-guideline.org) for a submission checklist. -
TRIPOD.md 9.2 KB
# TRIPOD Checklist **Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis** Version: TRIPOD 2015 Source: Moons KGM et al. Ann Intern Med. 2015;162:55-63. https://www.tripod-statement.org Source: Collins GS, Reitsma JB, Altman DG, Moons KGM. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD). *Ann Intern Med* 2015;162(1):55-63 (DOI 10.7326/M14-0697). Licence: *Annals of Internal Medicine* (© American College of Physicians) — no open licence. Verification: all 37 sub-items and every D / V / D;V designation were compared against the official TRIPOD checklist for prediction-model development **and** validation, distributed at tripod-statement.org (`Tripod-Checklist-Prediction-Model-Development-and-Validation-PDF.pdf`). 37/37 items and 37/37 designations match. Item 17 stated the wrong parenthetical and item 16 presented our own guidance as though it were the item text; both are corrected. ## Applicability Items apply to different study types: - **D** = Development study only - **V** = Validation study only - **DV** = Both development and validation studies ## Checklist Items ### Title and Abstract | # | Item | Applies | Description | |---|------|---------|-------------| | 1 | Title | DV | Identify the study as developing and/or validating a multivariable prediction model, the target population, and the outcome to be predicted. | | 2 | Abstract | DV | Provide a summary of objectives, study design, setting, participants, sample size, predictors, outcome, statistical analysis, results, and conclusions. | ### Introduction | # | Item | Applies | Description | |---|------|---------|-------------| | 3a | Background | DV | Explain the medical context (including whether diagnostic or prognostic) and rationale for developing or validating the multivariable prediction model, including references to existing models. | | 3b | Objectives | DV | Specify the objectives, including whether the study describes the development or validation of the model or both. | ### Methods #### Source of Data | # | Item | Applies | Description | |---|------|---------|-------------| | 4a | Source of data | DV | Describe the study design or source of data (e.g., randomized trial, cohort, or registry data), separately for the development and validation data sets, if applicable. | | 4b | Source of data | DV | Specify the key study dates, including start of accrual; end of accrual; and, if applicable, end of follow-up. | #### Participants | # | Item | Applies | Description | |---|------|---------|-------------| | 5a | Participants | DV | Specify key elements of the study setting (e.g., primary care, secondary care, general population) including number and location of centres. | | 5b | Participants | DV | Describe eligibility criteria for participants. | | 5c | Participants | DV | Give details of treatments received, if relevant. | #### Outcome | # | Item | Applies | Description | |---|------|---------|-------------| | 6a | Outcome | DV | Clearly define the outcome that is predicted by the prediction model, including how and when assessed. | | 6b | Outcome | DV | Report any actions to blind assessment of the outcome to be predicted. | #### Predictors | # | Item | Applies | Description | |---|------|---------|-------------| | 7a | Predictors | DV | Clearly define all predictors used in developing or validating the multivariable prediction model, including how and when they were measured. | | 7b | Predictors | DV | Report any actions to blind assessment of predictors for the outcome and other predictors. | #### Sample Size | # | Item | Applies | Description | |---|------|---------|-------------| | 8 | Sample size | DV | Explain how the study size was arrived at. | #### Missing Data | # | Item | Applies | Description | |---|------|---------|-------------| | 9 | Missing data | DV | Describe how missing data were handled (e.g., complete-case analysis, single imputation, multiple imputation) with details of any imputation method. | #### Statistical Analysis Methods | # | Item | Applies | Description | |---|------|---------|-------------| | 10a | Model building | D | Describe how predictors were handled in the analyses. | | 10b | Model building | D | Specify type of model, all model-building procedures (including any predictor selection), and method for internal validation. | | 10c | Validation | V | For validation, describe how the predictions were calculated. | | 10d | Model performance | DV | Specify all measures used to assess model performance and, if relevant, to compare multiple models. | | 10e | Model updating | V | Describe any model updating (e.g., recalibration) arising from the validation, if done. | #### Risk Groups | # | Item | Applies | Description | |---|------|---------|-------------| | 11 | Risk groups | DV | Provide details on how risk groups were created, if done. | #### Development vs. Validation | # | Item | Applies | Description | |---|------|---------|-------------| | 12 | Development vs. validation | V | For validation, identify any differences from the development data in setting, eligibility criteria, outcome, and predictors. | ### Results #### Participants | # | Item | Applies | Description | |---|------|---------|-------------| | 13a | Participants | DV | Describe the flow of participants through the study, including the number of participants with and without the outcome and, if applicable, a summary of the follow-up time. A diagram may be helpful. | | 13b | Participants | DV | Describe the characteristics of the participants (basic demographics, clinical features, available predictors), including the number of participants with missing data for predictors and outcome. | | 13c | Participants | V | For validation, show a comparison with the development data of the distribution of important variables (demographics, predictors, and outcome). | #### Model Development | # | Item | Applies | Description | |---|------|---------|-------------| | 14a | Model development | D | Specify the number of participants and outcome events in each analysis. | | 14b | Model development | D | If done, report the unadjusted association between each candidate predictor and outcome. | #### Model Specification | # | Item | Applies | Description | |---|------|---------|-------------| | 15a | Model specification | D | Present the full prediction model to allow predictions for individuals (i.e., all regression coefficients, and model intercept or baseline survival at a given time point). | | 15b | Model specification | D | Explain how to use the prediction model. | #### Model Performance | # | Item | Applies | Description | |---|------|---------|-------------| | 16 | Model performance | DV | Report performance measures (with CIs) for the prediction model. *(In practice this means discrimination — e.g., C-statistic/AUC — and calibration — e.g., calibration plot, calibration slope and intercept; the item itself does not enumerate them.)* | #### Model Updating | # | Item | Applies | Description | |---|------|---------|-------------| | 17 | Model updating | V | If done, report the results from any model updating (i.e., model specification, model performance). | ### Discussion | # | Item | Applies | Description | |---|------|---------|-------------| | 18 | Limitations | DV | Discuss any limitations of the study (such as nonrepresentative sample, few events per predictor, missing data). | | 19a | Interpretation | V | For validation, discuss the results with reference to performance in the development data, and any other validation data. | | 19b | Interpretation | DV | Give an overall interpretation of the results, considering objectives, limitations, results from similar studies, and other relevant evidence. | | 20 | Implications | DV | Discuss the potential clinical use of the model and implications for future research. | ### Other Information | # | Item | Applies | Description | |---|------|---------|-------------| | 21 | Supplementary information | DV | Provide information about the availability of supplementary resources, such as study protocol, Web calculator, and data sets. | | 22 | Funding | DV | Give the source of funding and the role of the funders for the present study. | --- ## Notes for Assessors - TRIPOD 2015 applies to **non-AI/ML prediction models only** (logistic regression, Cox regression, etc.). For AI/ML prediction models, use TRIPOD+AI 2024 instead. Do NOT apply both simultaneously. - Items 10a, 10b, 14a, 14b, 15a, 15b are specific to **development studies**. Mark as N/A for validation-only studies. - Items 10c, 10e, 13c, 17, 19a are specific to **validation studies**. Mark as N/A for development-only studies. - Item 16 must include BOTH discrimination (C-statistic/AUC) and calibration (calibration plot, Hosmer-Lemeshow, or calibration slope/intercept). A study reporting only AUC without calibration assessment is PARTIAL on this item. - Item 8 (sample size): the events per variable (EPV) ratio should be reported. EPV < 10 should be flagged as a limitation. - Item 15a (full model specification): the complete regression equation must be presented, not just significant predictors. Omitting non-significant predictors from the final model presentation is a common gap. - For studies that combine development and validation (e.g., internal-external validation), all items apply. -
TRIPOD_AI.md 14.1 KB
# TRIPOD+AI 2024 Checklist **Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis — Artificial Intelligence extension** - **Version:** TRIPOD+AI 2024 - **Citation:** Collins GS, Moons KGM, Dhiman P, et al. *TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods.* BMJ 2024;385:e078378. - **DOI:** 10.1136/bmj-2023-078378 - **Source:** https://www.tripod-statement.org — official expanded checklist: https://www.tripod-statement.org/wp-content/uploads/2024/04/TRIPODAI-Supplement.pdf - **Licence:** CC BY 4.0. Item wording below is reproduced from the published statement with attribution. - **Verified:** all 52 sub-items compared against the official supplements (ST1 expanded checklist, ST2 fillable checklist). 52/52 match in substance; the D/E designations match the statement exactly. Differences are orthographic only. **TRIPOD+AI 2024 supersedes and replaces TRIPOD 2015.** It is not TRIPOD 2015 plus AI addenda — it is a complete rewrite, applicable to prediction-model studies using **either** regression **or** machine-learning methods. For any prediction-model study (regression or AI/ML), use this checklist; do **not** apply the 2015 checklist alongside it. The checklist has **27 main items** and **52 checklist subitems**, in eight sections: Title (1), Abstract (2), Introduction (3–4), Methods (5–17), Open science (18), Patient and public involvement (19), Results (20–24), and Discussion (25–27). **Applicability column:** `D` = relevant only to model **development**; `E` = relevant only to model **evaluation**; `D;E` = applicable to **both**. ## Checklist Items ### Title | Item | Topic | D/E | Checklist item | |---|---|---|---| | 1 | Title | D;E | Identify the study as developing or evaluating the performance of a multivariable prediction model, the target population, and the outcome to be predicted. | ### Abstract | Item | Topic | D/E | Checklist item | |---|---|---|---| | 2 | Abstract | D;E | See TRIPOD+AI for Abstracts checklist. | ### Introduction | Item | Topic | D/E | Checklist item | |---|---|---|---| | 3a | Background | D;E | Explain the healthcare context (including whether diagnostic or prognostic) and rationale for developing or evaluating the prediction model, including references to existing models. | | 3b | Background | D;E | Describe the target population and the intended purpose of the prediction model in the context of the care pathway, including its intended users (e.g., healthcare professionals, patients, public). | | 3c | Background | D;E | Describe any known health inequalities between sociodemographic groups. | | 4 | Objectives | D;E | Specify the study objectives, including whether the study describes the development or validation of a prediction model (or both). | ### Methods | Item | Topic | D/E | Checklist item | |---|---|---|---| | 5a | Data | D;E | Describe the sources of data separately for the development and evaluation datasets (e.g., randomised trial, cohort, routine care or registry data), the rationale for using these data, and representativeness of the data. | | 5b | Data | D;E | Specify the dates of the collected participant data, including start and end of participant accrual; and, if applicable, end of follow-up. | | 6a | Participants | D;E | Specify key elements of the study setting (e.g., primary care, secondary care, general population) including the number and location of centres. | | 6b | Participants | D;E | Describe the eligibility criteria for study participants. | | 6c | Participants | D;E | Give details of any treatments received, and how they were handled during model development or evaluation, if relevant. | | 7 | Data preparation | D;E | Describe any data pre-processing and quality checking, including whether this was similar across relevant sociodemographic groups. | | 8a | Outcome | D;E | Clearly define the outcome that is being predicted and the time horizon, including how and when assessed, the rationale for choosing this outcome, and whether the method of outcome assessment is consistent across sociodemographic groups. | | 8b | Outcome | D;E | If outcome assessment requires subjective interpretation, describe the qualifications and demographic characteristics of the outcome assessors. | | 8c | Outcome | D;E | Report any actions to blind assessment of the outcome to be predicted. | | 9a | Predictors | D | Describe the choice of initial predictors (e.g., literature, previous models, all available predictors) and any pre-selection of predictors before model building. | | 9b | Predictors | D;E | Clearly define all predictors, including how and when they were measured (and any actions to blind assessment of predictors for the outcome and other predictors). | | 9c | Predictors | D;E | If predictor measurement requires subjective interpretation, describe the qualifications and demographic characteristics of the predictor assessors. | | 10 | Sample size | D;E | Explain how the study size was arrived at (separately for development and evaluation), and justify that the study size was sufficient to answer the research question. Include details of any sample size calculation. | | 11 | Missing data | D;E | Describe how missing data were handled. Provide reasons for omitting any data. | | 12a | Analytical methods | D | Describe how the data were used (e.g., for development and evaluation of model performance) in the analysis, including whether the data were partitioned, considering any sample size requirements. | | 12b | Analytical methods | D | Depending on the type of model, describe how predictors were handled in the analyses (functional form, rescaling, transformation, or any standardisation). | | 12c | Analytical methods | D | Specify the type of model, rationale, all model-building steps, including any hyperparameter tuning, and method for internal validation. | | 12d | Analytical methods | D;E | Describe if and how any heterogeneity in estimates of model parameter values and model performance was handled and quantified across clusters (e.g., hospitals, countries). See TRIPOD-Cluster for additional considerations. | | 12e | Analytical methods | D;E | Specify all measures and plots used (and their rationale) to evaluate model performance (e.g., discrimination, calibration, clinical utility) and, if relevant, to compare multiple models. | | 12f | Analytical methods | E | Describe any model updating (e.g., recalibration) arising from the model evaluation, either overall or for particular sociodemographic groups or settings. | | 12g | Analytical methods | E | For model evaluation, describe how the model predictions were calculated (e.g., formula, code, object, application programming interface). | | 13 | Class imbalance | D;E | If class imbalance methods were used, state why and how this was done, and any subsequent methods to recalibrate the model or the model predictions. | | 14 | Fairness | D;E | Describe any approaches that were used to address model fairness and their rationale. | | 15 | Model output | D | Specify the output of the prediction model (e.g., probabilities, classification). Provide details and rationale for any classification and how the thresholds were identified. | | 16 | Training versus evaluation | D;E | Identify any differences between the development and evaluation data in healthcare setting, eligibility criteria, outcome, and predictors. | | 17 | Ethical approval | D;E | Name the institutional research board or ethics committee that approved the study and describe the participant informed consent or the ethics committee waiver of informed consent. | ### Open science | Item | Topic | D/E | Checklist item | |---|---|---|---| | 18a | Funding | D;E | Give the source of funding and the role of the funders for the present study. | | 18b | Conflicts of interest | D;E | Declare any conflicts of interest and financial disclosures for all authors. | | 18c | Protocol | D;E | Indicate where the study protocol can be accessed or state that a protocol was not prepared. | | 18d | Registration | D;E | Provide registration information for the study, including register name and registration number, or state that the study was not registered. | | 18e | Data sharing | D;E | Provide details of the availability of the study data. | | 18f | Code sharing | D;E | Provide details of the availability of the analytical code. | ### Patient and public involvement | Item | Topic | D/E | Checklist item | |---|---|---|---| | 19 | Patient and public involvement | D;E | Provide details of any patient and public involvement during the design, conduct, reporting, interpretation, or dissemination of the study or state no involvement. | ### Results | Item | Topic | D/E | Checklist item | |---|---|---|---| | 20a | Participants | D;E | Describe the flow of participants through the study, including the number of participants with and without the outcome and, if applicable, a summary of the follow-up time. A diagram may be helpful. | | 20b | Participants | D;E | Report the characteristics overall and, where applicable, for each data source or setting, including the key dates, key predictors (including demographics), treatments received, sample size, number of outcome events, follow-up time, and amount of missing data. A table may be helpful. Report any differences across key demographic groups. | | 20c | Participants | E | For model evaluation, show a comparison with the development data of the distribution of important predictors (demographics, predictors, and outcome). | | 21 | Model development | D;E | Specify the number of participants and outcome events in each analysis (e.g., for model development, hyperparameter tuning, model evaluation). | | 22 | Model specification | D | Provide details of the full prediction model (e.g., formula, code, object, application programming interface) to allow predictions in new individuals and to enable third-party evaluation and implementation, including any restrictions to access or re-use (e.g., freely available, proprietary). | | 23a | Model performance | D;E | Report model performance estimates with confidence intervals, including for any key subgroups (e.g., sociodemographic). Consider plots to aid presentation. | | 23b | Model performance | D;E | If examined, report results of any heterogeneity in model performance across clusters. See TRIPOD-Cluster for additional details. | | 24 | Model updating | E | Report the results from any model updating, including the updated model and subsequent performance. | ### Discussion | Item | Topic | D/E | Checklist item | |---|---|---|---| | 25 | Interpretation | D;E | Give an overall interpretation of the main results, including issues of fairness in the context of the objectives and previous studies. | | 26 | Limitations | D;E | Discuss any limitations of the study (such as a non-representative sample, sample size, overfitting, missing data) and their effects on any biases, statistical uncertainty, and generalizability. | | 27a | Usability of the model in the context of current care | D | Describe how poor quality or unavailable input data (e.g., predictor values) should be assessed and handled when implementing the prediction model. | | 27b | Usability of the model in the context of current care | D | Specify whether users will be required to interact in the handling of the input data or use of the model, and what level of expertise is required of users. | | 27c | Usability of the model in the context of current care | D;E | Discuss any next steps for future research, with a specific view to applicability and generalizability of the model. | --- ## MedSci supplemental checks — NOT official TRIPOD+AI items The following are **not** TRIPOD+AI checklist items. They are engineering-reproducibility prompts this toolkit adds because they matter for an AI/ML model to be rebuildable, and they extend (do not replace) the official items noted. Assess them **only as supplements**; never report them as canonical TRIPOD+AI item numbers, and never let them substitute for an official item. | Supplemental check | Extends official item | What to look for | |---|---|---| | Model architecture | 12c | Network type, depth, layers, activation and loss functions — enough detail to rebuild the model. | | Training configuration | 12c | Optimizer, learning-rate schedule, batch size, epochs, early-stopping criterion, regularisation. | | Software and hardware | 18f, 22 | Libraries with version numbers, language, and the compute used (e.g., GPU type for deep learning). | | Reproducibility | 18e, 18f | Random seeds, the exact data split, and whether code + weights are available to reproduce results. | Official item **14 (Fairness)**, **18f (Code sharing)**, **22 (Model specification / API)**, and **7 (Data preparation)** already cover fairness, code/model availability, and preprocessing — assess those under their official items, not as extras here. ## Notes for assessors - **Footnotes carried from the official checklist.** Item 12c is assessed **separately for every model-building approach** used, not once for the study. Item 18f (code sharing) refers to the *analysis* code — data cleaning, feature engineering, model building, evaluation — while item 22 refers to the code needed to *implement* the model for a new individual; a study can satisfy one and not the other. - **Applicability.** Mark items labelled `D` (development-only) or `E` (evaluation-only) as N/A when they do not apply to the study design; `D;E` items apply to both. - **Open science (18a–18f) and PPI (19) are official, mandatory items** — a compliance report that skips them is incomplete. Data sharing (18e) and code sharing (18f) are where an AI/ML study's reproducibility is assessed. - **Performance (23a) with confidence intervals is required**, including for key subgroups; heterogeneity across clusters (23b) when examined. A study reporting a point estimate without a CI is PARTIAL on 23a. - **Fairness is an official item (14)** and is reported across demographic groups in 20b and 23a — not an optional AI add-on. - Do **not** apply TRIPOD 2015 alongside this checklist; TRIPOD+AI 2024 supersedes it for both regression and AI/ML prediction models. - For a study that is both a diagnostic-accuracy study and a prediction-model study, cross-reference STARD. -
TRIPOD_LLM.md 11.4 KB
# TRIPOD-LLM Checklist **Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis -- Large Language Models** Version: TRIPOD-LLM 2025 (living guideline) Source: https://www.tripod-statement.org · interactive checklist: https://tripod-llm.vercel.app Reference: Gallifant J, ..., Bitterman DS. Nat Med 2025;31(1):60-69. doi:10.1038/s41591-024-03425-5 > **Educational summary, authored in our own words.** This file paraphrases the *intent* of each > TRIPOD-LLM item to drive an item-by-item audit; it does **not** reproduce the guideline's verbatim > wording. The published article is subscription-access. For a submission-ready checklist, complete > the official instrument at the source above and cite Gallifant et al. 2025. > > **Item numbers below follow the official TRIPOD-LLM scheme: 19 main items (1–19) with letter > subitems (≈50 subitems).** 14 main items / 32 subitems apply to every LLM study; 5 main items / > 18 subitems are **task- or design-specific** (marked *task-specific*) and are N/A when that > component is absent — justify the N/A. Licence: The published *Nature Medicine* article returns no Creative Commons licence via Crossref (DOI 10.1038/s41591-024-03425-5); the author-accepted manuscript is CC BY 4.0 through institutional rights retention. Verification: all 49 sub-items were compared, by number and order, against Table 2 of the published guideline (PubMed Central record PMC12104976). 49/49 are now present. **This file previously collapsed item 16's four sub-items into a single row**, which is the shape that lets an assessor tick one box while three requirements go unchecked; they are expanded below. Wording stays paraphrased — Crossref returns only a Springer Nature text-and-data-mining licence for this article. ## Naming and scope (read first) - TRIPOD-LLM is an **extension** of the **TRIPOD** family (TRIPOD 2015 → TRIPOD+AI 2024 → TRIPOD-LLM 2025). In Methods, name **both** the base instrument and the extension and cite each (manuscript-style-classical §14): "reported per TRIPOD (Collins et al. 2015) and its large-language-model extension TRIPOD-LLM (Gallifant et al. 2025)." - **Applies to** studies that develop, fine-tune, prompt, or evaluate **large language models** for a biomedical/clinical task (classification, extraction, summarization, generation, question answering, etc.) — not only to risk-prediction models. - **Pairs with, does not replace, MI-CLEAR-LLM** (the 6-item LLM-accuracy supplement). Apply MI-CLEAR-LLM alongside when LLM accuracy is an outcome. ## Checklist Items (official numbering 1–19) Status each PRESENT / PARTIAL / MISSING / N/A. ### Title and Abstract | # | Item | Description (intent) | |---|------|----------------------| | 1 | Title | Identify the study as developing, fine-tuning, prompting, and/or evaluating an LLM, and name the clinical task and target setting. | | 2 | Abstract | Provide a structured summary (objectives, data, LLM and version, task, evaluation approach, human oversight, key results, limitations). Follow the TRIPOD-LLM-for-Abstracts items. | ### Introduction | # | Item | Description (intent) | |---|------|----------------------| | 3a | Background — context | Explain the healthcare problem, the clinical context, and the rationale for using an LLM. | | 3b | Background — population & use | Describe the target population/users and the intended use of the LLM. | | 4 | Objectives | State the specific study objectives and the LLM task(s) addressed. | ### Methods — Data (item 5, a–e) | # | Item | Description (intent) | |---|------|----------------------| | 5a | Data sources | Describe all data sources and input data types (clinical text, notes, structured fields, image-to-text). | | 5b | Data description / distribution | Describe the dataset and how data were partitioned (train / tune / validation / held-out test, internal vs external), and steps to prevent train–test contamination and leakage of evaluation data into pretraining/prompts. | | 5c | Study dates | Specify key dates (data accrual start/end, and the model knowledge-cutoff relative to the data). | | 5d | Preprocessing | Describe text preprocessing, de-identification, and any filtering. | | 5e | Missing / inadequate / imbalanced data | Describe handling of missing, truncated, out-of-context, or class-imbalanced inputs. | ### Methods — Analytical / LLM methods (item 6, a–e) | # | Item | Description (intent) | |---|------|----------------------| | 6a | LLM identity and version | Name the model, exact version/snapshot or weights, provider/access route (API vs local), and date of access — versions drift, so this is essential for reproducibility. | | 6b | Development / adaptation | Describe how the LLM was developed or adapted (zero/few-shot prompting, retrieval augmentation, fine-tuning, instruction-tuning) in enough detail to reproduce. | | 6c | Text generation settings | Report decoding/generation parameters (temperature, top-p, max tokens, stop criteria, seed/determinism where available). | | 6d | Output | Define the expected output format and how free-text outputs were mapped to the study endpoint. | | 6e | Output handling / classification | Describe parsing, constraint enforcement, and any human or rule-based post-processing/classification of outputs. | ### Methods — LLM output evaluation (item 7, a–e) | # | Item | Description (intent) | |---|------|----------------------| | 7a | Quality / performance metrics | Specify all performance/quality metrics, including task-specific measures. | | 7b | Downstream / clinical relevance | Describe how the metrics relate to downstream clinical relevance. | | 7c | Outcome definition | Define the outcome / reference standard and who set it. | | 7d | Subjective / human assessment | Describe any human rating: rubric, anchors, number and expertise of raters, blinding, and inter-rater agreement. | | 7e | Comparisons | Describe comparators (clinicians, prior models, guidelines) and ensure same-data, same-task comparison. | ### Methods — Annotation (item 8, a–c) | # | Item | Description (intent) | |---|------|----------------------| | 8a | Labeling process | Describe the labeling/annotation process and adjudication. | | 8b | Annotator count | Report the number of annotators. | | 8c | Annotator background | Report annotator expertise/background. | ### Methods — task/design-specific components (N/A if absent; justify) | # | Item | Description (intent) | |---|------|----------------------| | 9a | Prompting — design *(task-specific)* | If prompting was used: describe prompt design/templates and the prompt-selection procedure. | | 9b | Prompting — data *(task-specific)* | Describe the data used to develop prompts (kept separate from test data). | | 10 | Summarization preprocessing *(task-specific)* | If summarization was a task: describe inputs/preprocessing and how summaries were assessed for faithfulness/omission. | | 11 | Instruction tuning / alignment *(task-specific)* | If instruction-tuning/alignment was applied: describe the instructions and the data and procedure used. | ### Methods — Compute and Ethics | # | Item | Description (intent) | |---|------|----------------------| | 12 | Compute | Report computational resources (hardware, and for fine-tuning, scale/cost relevant to reproducibility and environmental reporting). | | 13 | Ethical approval | Report IRB/ethics approval (or exemption) and consent/data-governance for the data used. | ### Open Science (item 14, a–f) | # | Item | Description (intent) | |---|------|----------------------| | 14a | Funding | Source of funding and role of funders. | | 14b | Conflicts of interest | Declare conflicts, including relationships with model providers. | | 14c | Protocol | State protocol availability. | | 14d | Registration | State any study registration. | | 14e | Data availability | State availability of data (and access constraints for clinical text). | | 14f | Code / prompt availability | State availability of code, prompts, and (where applicable) model weights or access instructions. | ### Patient and Public Involvement | # | Item | Description (intent) | |---|------|----------------------| | 15 | PPI | Describe patient/public involvement in design/conduct, or state that there was none. | ### Results | # | Item | Description (intent) | |---|------|----------------------| | 16a | Participants — data flow *(patient/EHR data)* | Describe the flow of text/EHR/patient data through the study, including the number of documents/questions/participants with and without the outcome/label, and follow-up where applicable. | | 16b | Participants — characteristics *(patient/EHR data)* | Report characteristics overall and for each data source or setting, and for the development/evaluation splits, including key dates and key characteristics. | | 16c | Participants — distribution comparison *(evaluations with clinical outcomes)* | Show a comparison of the distribution of important clinical variables that may be associated with the outcome, between development and evaluation data. | | 16d | Participants — analysis sample sizes *(patient/EHR data)* | Specify the number of participants and outcome events in each analysis — LLM development, hyperparameter tuning, and LLM evaluation. | | 17 | Performance | Report performance/quality results with appropriate uncertainty, including human-evaluation results and, where relevant, subgroup/fairness performance. | | 18 | LLM updating | If the model or prompts were updated during the study, report results before and after. | ### Discussion (item 19, a–g) | # | Item | Description (intent) | |---|------|----------------------| | 19a | Interpretation | Interpret results in light of objectives, comparators, and the clinical task. | | 19b | Limitations | Discuss limitations (data, evaluation, generalizability, version drift, hallucination/safety). | | 19c | Usability and context | Discuss the deployment context and conditions required for safe use. | | 19d | Intended use | State the intended use and avoid claims beyond the evidence. | | 19e | Input-data / poor-quality handling | Discuss handling of poor-quality inputs and observed failure modes. | | 19f | User expertise / oversight | Discuss the user expertise and human oversight required for safe use. | | 19g | Future research | Outline implications and future work. | --- ## Notes for Assessors - The numbering follows the official TRIPOD-LLM scheme (19 main items, ~50 subitems). Use the official interactive checklist for the submission form and the separate TRIPOD-LLM-for-Abstracts sub-checklist (item 2). - **Version reporting (item 6a) is non-waivable** for an LLM study: a result tied to an unnamed or undated model snapshot is not reproducible. Mark MISSING if the exact version/access date is absent. - **Leakage/contamination (item 5b)** is the LLM-specific analogue of train–test separation: evaluation data must not have entered pretraining, fine-tuning, or prompt-development. Probe explicitly. - **Human evaluation (item 7d)** that drives a quality claim needs a rubric with anchors, rater expertise, blinding, and inter-rater agreement — a bare "two physicians reviewed the outputs" is PARTIAL. - Task-specific items (Prompting 9, Summarization 10, Instruction-tuning 11, Results-participants 16) are **N/A** when the component is absent — record a one-line justification rather than a blank. - TRIPOD-LLM is a **living guideline**; confirm the current item set at the source before a formal submission checklist, and pair with MI-CLEAR-LLM when LLM accuracy is an evaluated outcome.
-
-
critical_item_floor.md 5.2 KB
# Critical-item floor — presence outranks the headline percentage A compliance percentage is a coarse summary: two manuscripts at the same percentage can differ enormously in reviewability because **not all items carry equal weight**. After the item-by-item table (Step 4), use this floor so a MISSING **critical** item is surfaced as a headline gap regardless of the overall percentage. ## The principle 1. **Critical items are pass/fail.** For each study type a small set of items is treated as non-waivable by methods reviewers/editors. If one is MISSING, surface it as a **Critical gap** even when the overall percentage is high — a 90%-compliant manuscript whose ground truth is undefined or which cannot be reproduced is not "90% acceptable." 2. **The percentage is a secondary signal** ("broadly thorough vs broadly thin"), not a desk-reject line. Do not assert a numeric floor a journal has not published. 3. **The journal's own required elements are hard** (a missing required panel, the wrong checklist for the design) — verify against the author guidelines, never invent them. ## Critical items by guideline (MISSING → Critical gap) | Guideline (study type) | Critical items | |---|---| | **CLAIM / STARD-AI** (imaging AI / AI diagnostic) | Ground-truth/reference-standard definition; data-partition with leakage control; evaluation metrics with uncertainty; model + training procedure (when a model is developed — for a locked/already-trained model evaluated as a diagnostic test, the version/provenance instead) | | **TRIPOD+AI** (prediction model) | Input data + outcome definition/timing; development and validation; performance with **calibration**, not discrimination only | | **STARD** (diagnostic accuracy) | Reference standard + rationale; reader blinding (index vs reference); participant flow (2×2 recoverable); indeterminate-result handling | | **STROBE** (observational) | Eligibility/selection; exposure/outcome definitions; **missing-data handling**; participant flow that reconciles | | **PRISMA / PRISMA-DTA** (systematic review) | Full search strategy (≥1 database verbatim); flow diagram that reconciles; per-study risk of bias; registration/protocol **status or reference** (an explicit "not registered / no protocol" satisfies this — a reference is not mandatory) | | **REMARK** (prognostic tumor-marker study) | Marker definition + pre-specified vs exploratory hypotheses; **cutpoint justification** when a continuous marker is categorized; multivariable effect **adjusted for established prognostic variables** (not a marker-alone analysis); all examined endpoints reported, not only the significant ones | | **TARGET** (observational study emulating a target trial) | Explicit target-trial protocol specification (eligibility, treatment strategies, assignment, time zero, outcome, causal contrast) **and** its emulation mapping in the data; **time-zero alignment** of eligibility / assignment / start of follow-up (immortal-time control); stated causal estimand + identifying assumptions including baseline confounding | Where an extension genuinely extends a base instrument (e.g., STARD-AI extends STARD), a manuscript that names only the extension and skips the base items is a gap — but defer this to Step 4e (framework naming), which already distinguishes true extensions from standalone rewrites (TRIPOD+AI is a complete rewrite, not an add-on to TRIPOD 2015, so it implies no separate TRIPOD-2015 floor). Do not manufacture a base-citation Critical gap here. ## AI / radiomics: appraisal is separate from reporting The checklists above answer "was it **reported**?" For AI/ML and radiomics, also confirm the manuscript's chosen **methodological-quality / risk-of-bias** instrument and its non-waivable concerns — a fully *reported* paper can still be at high risk of bias. | Appraisal / RoB instrument | Non-waivable concerns | |---|---| | **PROBAST+AI** (prediction-model RoB, regression or ML/AI) | Train/test independence + leakage; sample size / overfitting-and-optimism control; calibration, not discrimination only | | **METRICS / RQS** (radiomics methodological quality, EuSoMII) | Feature reproducibility/stability (test–retest, segmentation, ICC); internal–external validation + multiplicity; calibration | | **APPRAISE-AI** (clinical-AI study quality) | Independent/external validation; robustness / error analysis; reproducibility (data + code) | Keep these distinct from their **reporting** counterparts — **CLEAR** (radiomics reporting) and **DECIDE-AI** (early clinical-evaluation reporting) are reporting guidelines, not RoB tools; route them through the normal checklist flow, not this appraisal note. For the fuller **METRICS** breakdown (9 categories / 30 weighted, condition-dependent items, with the non-waivable concerns), see `appraisal_tools/METRICS.md`. It is an appraisal reference, not a counted reporting checklist. ## How to use in the report - After the item table, list **Critical items: X / Y present**, naming each MISSING critical item and the section where it should appear. - If any critical item is MISSING, the report's headline is the **Critical gap**, not the percentage. (This floor complements the reviewer-side judgment layer; it does not assert a journal desk-reject threshold.) -
genai_image_study_object_decision_aid.md 4.2 KB
# Decision aid — reporting studies where generative AI images ARE the study object **When this applies:** the study evaluates **images that a generative AI model synthesized** (realism, controllability/steerability, whether human readers can distinguish synthetic from real, or model-vs-model quality). The generative model's *output* is the object under study. **When this does NOT apply:** a model (incl. a vision-language model) *interprets* images and you measure its diagnostic accuracy — that is an AI-accuracy study; use the relevant accuracy guideline directly (e.g., STARD-AI, CLAIM, TRIPOD+AI, MI-CLEAR-LLM). ## There is no single dominant checklist for this study type Generative-image-as-study-object work (e.g., RSNA reader studies on AI-synthesized or "deepfake" medical images) is reported by **assembling** existing guidelines plus a precedent bar. Do not claim wholesale compliance with any one checklist; map applicable items and cite the base guideline together with any AI extension (verify each item against the published source — never invent items). ### Generator / provenance side - **CLAIM 2024** — medical-imaging-AI umbrella; the 2024 revision covers generative/foundation models. If commercial models are used **as-is** (no training/fine-tuning by the authors), the model-development / training / validation-split items are **N/A**; report data sources, reference/real comparators, evaluation, transparency, and limitations. - **FUTURE-AI** — use the **Traceability** principle: persist verbatim prompts, a generation manifest, model + version + access date, and parameters (a prompt/generation registry). - **MI-CLEAR-LLM — transparency *items* only, not study-level compliance.** MI-CLEAR-LLM is scoped to **LLM *accuracy* studies in healthcare** (including VLMs interpreting images); it is **not** a guideline for generative-output studies. Borrow its reporting *items* for prompt-driven foundation models — verbatim prompt(s), model name + version + access date, access channel/API, sampling parameters, number of runs, handling of non-determinism, responsible party — to document generation provenance. Cite it as the basis for prompt logging, not as the study's reporting guideline. ### Reader / evaluation side - **STARD 2015 + STARD-AI** — if the reader task is **real-vs-synthetic discrimination**, that is a diagnostic-accuracy structure: report the reference standard (what counts as "real"/"synthetic"), reader blinding, flow, and accuracy with intervals. Cite base STARD **and** the STARD-AI extension. - **GRRAS** (Guidelines for Reporting Reliability and Agreement Studies) — for inter-reader feature/quality ratings: number and qualification of readers, blinding, the agreement statistic (ICC / weighted kappa) with 95% CI, and separate reporting of any anchor/control items. - **MRMC reporting** — for multi-reader multi-case designs: a-priori power, per-reader randomization/seed, and a real-control arm matched on non-content attributes (resolution, cropping, compression) so a format-only classifier cannot rival the readers. ### Precedent bar (de-facto standard for this study type) Match the methodological bar set by published generative-image-as-study-object reader studies in high-impact radiology venues: a-priori power, MRMC reader platform with per-reader seeds, real-control matching on non-content attributes, and **explicit, pre-specified handling of failed / low-quality generations** (count them rather than silently excluding survivors). ## Cross-cutting cautions - **No overclaim:** state which items of which guideline the study satisfies, verified against the published checklist; do not assert blanket "reported per [guideline]". - **Manuscript's own AI-use disclosure** (writing assistance) is separate from the study-object reporting above — see ICMJE/COPE and the write-paper LLM-disclosure feature. - **Pre-registration** of the primary estimand, frequency/realism references, and the fresh-only firewall (pilot/calibration images excluded from the confirmatory set) belongs in a study registry (e.g., OSF) for non-clinical reader studies — not PROSPERO (systematic reviews) or a clinical-trial registry (no health-outcome intervention). -
LICENSES.md 5.4 KB
# Checklist Licenses Attribution and licence status for the bundled reporting-guideline and risk-of-bias checklists. **How the Licence column is established.** Each entry is resolved from the article's DOI through the Crossref `license` field and, where Crossref carries only a text-and-data-mining policy, through the PubMed Central record's `<license>` element. A row reads **verified** only when one of those two returned an explicit Creative Commons or public-domain URL. Where neither did, the row says so — an absent licence statement is **not** evidence of permissive licensing, and several publishers in this table (ACP, JAMA Network) do not publish these instruments under an open licence at all. This distinction is load-bearing. This repository is MIT-licensed and is redistributed through npm, GitHub and a classroom ZIP without restriction. A checklist whose source is CC BY-**NC** cannot be carried verbatim under those terms, and one with no open licence cannot be carried verbatim at all. ## Verified permissive | File | Guideline | Reference | Licence | Verified via | |------|-----------|-----------|---------|--------------| | STROBE.md | STROBE 2007 | von Elm E et al. PLoS Med 2007 | CC BY 4.0 | PMC2020495 | | STARD.md | STARD 2015 | Bossuyt PM et al. BMJ 2015;351:h5527 | CC BY 4.0 | PMC4623764 | | TRIPOD_AI.md | TRIPOD+AI 2024 | Collins GS et al. BMJ 2024;385:e078378 | CC BY 4.0 | Crossref | | PRISMA_2020.md | PRISMA 2020 | Page MJ et al. BMJ 2021;372:n71 | CC BY 4.0 | Crossref | | PRISMA_2020_Abstracts.md | PRISMA 2020 for Abstracts | Page MJ et al. BMJ 2021;372:n71 | CC BY 4.0 | Crossref | | CONSORT.md | CONSORT 2025 | Hopewell S et al. BMJ 2025;389:e081123 | CC BY 4.0 | Crossref | | SPIRIT.md | SPIRIT 2025 | Chan AW et al. BMJ 2025;389:e081477 | CC BY 4.0 | Crossref | | ARRIVE_2.md | ARRIVE 2.0 | Percie du Sert N et al. PLoS Biol 2020 | CC0 1.0 | Crossref | ## Non-commercial — NOT covered by this repository's MIT licence Free to copy and redistribute **with attribution, for non-commercial purposes**. Commercial use requires permission from the rights holder. These files must remain summaries of item *intent* in our own words rather than reproductions. | File | Guideline | Reference | Licence | Verified via | |------|-----------|-----------|---------|--------------| | ROBINS_I.md | ROBINS-I 2016 | Sterne JAC et al. BMJ 2016;355:i4919 | **CC BY-NC 3.0** | PMC5062054 | | CARE.md | CARE 2013 | Gagnier JJ et al. J Clin Epidemiol 2014;67(1):46-51 | CC BY-NC 4.0 | publisher statement | | MI_CLEAR_LLM.md | MI-CLEAR-LLM | Park SH et al. Korean J Radiol 2024;25(10):865-868; 2025 update KJR 2025;26(12):1123-1132 | CC BY-NC 4.0 | publisher statement | | DECIDE_AI.md | DECIDE-AI 2022 | Vasey B et al. Nat Med 2022;28(5):924-933 | CC BY-NC 4.0 (DECIDE-AI materials) | publisher statement | ## No open licence found — summaries only, never reproductions Neither Crossref nor PMC returned a Creative Commons or public-domain licence for these. The publishers are subscription-access and do not place these instruments under an open licence. The bundled files must express item *intent* in our own words, cite the source, and direct the reader to complete the official instrument. | File | Guideline | Reference | Status | Verified via | |------|-----------|-----------|--------|--------------| | QUADAS3.md | QUADAS-3 | Whiting PF et al. Ann Intern Med 2026;179(4):548-555 | © ACP — no open licence | Crossref (TDM policy only) | | QUADAS2.md | QUADAS-2 | Whiting PF et al. Ann Intern Med 2011;155(8):529-536 | © ACP — no open licence | Crossref (TDM policy only) | | PROBAST.md | PROBAST 2019 | Wolff RF et al. Ann Intern Med 2019;170(1):51-58 | © ACP — no open licence | Crossref (TDM policy only) | | RoB2.md | RoB 2 2019 | Sterne JAC et al. BMJ 2019;366:l4898 | no CC licence found | Crossref (TDM policy only) | | PRISMA_DTA.md | PRISMA-DTA 2018 | McInnes MDF et al. JAMA 2018;319(4):388-396 | © AMA — no open licence | Crossref (no licence field) | | TRIPOD_LLM.md | TRIPOD-LLM 2025 | Gallifant J et al. Nat Med 2025;31(1):60-69 | published version subscription-access; author-accepted manuscript CC BY 4.0 via rights retention | publisher statement | | CLAIM_2024.md | CLAIM 2024 Update | Tejani AS et al. Radiol Artif Intell 2024;6(4):e240300 | © RSNA, open access — consult RSNA for reuse terms | publisher statement | | NOS.md | Newcastle-Ottawa Scale | Wells GA et al. Ottawa Hospital Research Institute | no formal licence published | — | ## Not yet resolved Every other file under `checklists/` is not listed above because its licence has **not** been resolved. That is a gap in this table, not a finding of permissiveness. Treat an unlisted file as "unknown licence, summarise only" until it appears here. ## Corrections made to this table Five rows previously claimed **CC BY** on no evidence: ROBINS-I (actually CC BY-**NC** 3.0), RoB 2, QUADAS-2, PROBAST, and PRISMA-DTA (no open licence found for any of the four). The claim was inherited rather than checked. It mattered: an NC restriction is incompatible with redistributing a verbatim reproduction under this repository's MIT licence, and two of the instruments come from a publisher that licenses none of this material openly. --- All files here are educational summaries that cite their source; they do not relicense the underlying guidelines. Any manuscript that uses one should cite the original instrument, and any assessment that is reported should be completed against the official document. -
report_templates.md 5.6 KB
# Step 5 — Report templates (Parts A–D) Load-on-demand companion to `/check-reporting` Step 5. SKILL.md keeps the NOT-FOR-SUBMISSION rule, the part list, and the JSON field contract; this file carries the four literal output templates — Part A (summary), Part B (item-by-item checklist), Part C (action items), and Part D (the machine-readable JSON block). Read it when you are writing the report. Everything here is an output format: none of it informs the audit itself. **The banner is not optional.** The report MUST begin with the NOT-FOR-SUBMISSION comment as its very first line — this is an internal working audit, never the official journal checklist an author fills in and uploads. Produce a structured compliance report in two parts. This report is an **internal working audit** — it carries auto-fix annotations, a machine-readable JSON block (`compliance_pct`, `fixable_by_ai`, …), and Action Items. It is **NOT** the official reporting checklist a journal expects (that is the blank guideline form with `Item | Recommendation | Reported in page/section`, which the authors fill in). Never submit this report as the submission checklist. To make the file self-identifying so it cannot be reused by filename into a later submission package, **the report MUST begin with the NOT-FOR-SUBMISSION banner below** as its very first line. (`/sync-submission`'s `check_checklist_dump_leak` gate also catches this dump if it ever lands in a submission directory.) #### Part A: Summary ``` <!-- INTERNAL AUDIT — NOT FOR SUBMISSION. This is the /check-reporting working report, not the official journal checklist. Do not upload to a submission portal. --> ## Reporting Guideline Compliance Report Manuscript: {title} Target manuscript file: {manuscript filename, e.g. manuscript_v8.md} Target version: {version token from the filename or frontmatter, e.g. v8} Guideline: {name and version} Date: {YYYY-MM-DD} Assessed by: Claude (automated pre-screening) ### Summary | Status | Count | Percentage | |--------|-------|------------| | PRESENT | {n} | {%} | | PARTIAL | {n} | {%} | | MISSING | {n} | {%} | | N/A | {n} | {%} | | **Total** | **{n}** | **100%** | Overall compliance: {PRESENT count}/{applicable count} ({%}) Critical items (Step 4f): {present}/{total} present.{ if any missing: " Critical gap — " + each MISSING critical item with the section it belongs in. This, not the percentage, is the headline.} ``` #### Part B: Item-by-Item Checklist ``` ### Detailed Checklist | # | Section | Item | Status | Location | Notes | |---|---------|------|--------|----------|-------| | 1 | Title/Abstract | {item text} | PRESENT | Title | {notes} | | 2 | Introduction | {item text} | MISSING | -- | {suggestion} | | ... | ... | ... | ... | ... | ... | ``` #### Part C: Action Items (for MISSING and PARTIAL) ``` ### Action Items (Priority Order) 1. **[MISSING] Item {N}: {item name}** - Required: {what needs to be added} - Suggested location: {section, paragraph} - Example text: "{draft sentence or phrase}" 2. **[PARTIAL] Item {N}: {item name}** - Current: {what was found} - Needed: {what additional detail is required} - Suggested revision: "{draft revision}" ``` Order action items by: 1. Items most journals enforce strictly (e.g., ethics approval, registration, sample size) 2. Items in the Methods section (easiest to fix) 3. Items in other sections #### Part D: Machine-Readable JSON Summary Append a fenced JSON block at the end of the report. This enables `/write-paper` Phase 7 and `/orchestrate` to parse compliance results programmatically. This block **MUST** be present when invoked with `--json` flag or when called from `/write-paper` Phase 7. It SHOULD also be present in standard invocations (appended after Part C). ```json { "check_reporting_version": "1.1", "manuscript_title": "...", "target_manuscript": "manuscript_v8.md", "target_version": "v8", "source_sha256": "<first 12 hex chars of sha256 of the manuscript file bytes>", "guideline": "STARD-AI", "guideline_version": "2025", "date": "YYYY-MM-DD", "total_items": 40, "present": 32, "partial": 4, "missing": 3, "na": 1, "compliance_pct": 88.9, "action_items": [ { "item_number": 12, "section": "Methods", "item_name": "Sample size justification", "status": "MISSING", "suggested_location": "Methods, after participant description", "suggested_fix": "Add: 'The sample size was determined based on [rationale]. A minimum of [N] cases was required to achieve [target] precision for the primary endpoint.'", "fixable_by_ai": true }, { "item_number": 7, "section": "Methods", "item_name": "Blinding of index test to reference standard", "status": "PARTIAL", "current_text": "Readers were blinded", "needed": "Specify what readers were blinded to (reference standard results, clinical information, other reader results)", "suggested_fix": "Expand to: 'Readers interpreted [index test] images blinded to the reference standard results, clinical information, and other readers' assessments.'", "fixable_by_ai": true } ] } ``` **Field definitions:** - `compliance_pct`: `present / (total_items - na) * 100`, rounded to one decimal - `action_items`: Array of MISSING and PARTIAL items only (PRESENT and N/A excluded) - `fixable_by_ai`: `true` if the fix involves inserting or expanding text with information available in the manuscript or inferable from context; `false` if it requires external information (e.g., registration number, IRB approval number, specific protocol details only the author knows) - `suggested_fix`: Concrete draft text that can be inserted or used to expand an existing sentence --- -
step4c_registration_timing.md 4.1 KB
# Step 4c Reference — Registration / Protocol Timing Consistency Check Load this reference when running Step 4c during `/check-reporting`. The SKILL.md body carries only the scope + trigger summary; full item-by-item procedure lives here. ## Applies To Systematic reviews, meta-analyses, and intervention studies with prospective registration (PRISMA 2020, PRISMA-DTA, PRISMA-P, MOOSE, CONSORT, SPIRIT). ## Motivation The registration identifier itself is a single checklist item in most guidelines, so it can pass the Step 4 audit even when the manuscript is internally inconsistent about *when* the registration or its amendments occurred relative to the analysis. Reviewers and editors scrutinize this timing — an undisclosed post-hoc amendment is a common rejection trigger. ## Items to Check ### 1. Registration identifier present Confirm the registration number (e.g., PROSPERO CRD####, ClinicalTrials.gov NCT####) is cited in Methods, the Abstract, and (where required) the cover letter. ### 2. Initial registration date vs. manuscript milestone dates Extract from the manuscript: - Search start / end date (e.g., "databases were searched from 2010-01-01 to 2025-12-31"). - Screening / data extraction completion date (often in the PRISMA flow caption or in Methods). - If the Methods states an explicit registration date, that date must predate — or at minimum not contradict — the screening/extraction milestone. A registration date *after* data extraction completion must be disclosed as retrospective and justified. ### 3. Amendment date(s) consistency If the manuscript references an amendment to the registered protocol: - The amendment date must appear in Methods (typical phrasing: "The registered protocol was amended on YYYY-MM-DD to …"). - The described amendment content must match what Methods actually reports (e.g., an eligibility refinement, subgroup reassignment, outcome addition). A Methods-stated amendment that does not correspond to any visible methods change is a red flag. - If the amendment post-dates the analysis lock, Methods must state that the analysis was re-run after the amendment — otherwise the amendment is a post-hoc rationalization. - The amendment date must not post-date the manuscript submission date (when the latter is known from the cover letter or file metadata). ### 4. Cross-artifact agreement When the author provides the registry record (PROSPERO PDF, ClinicalTrials.gov export) as a supplement or accompanying document: - Primary outcome(s), eligibility criteria, and analysis plan described in Methods must agree with the registry entry, or explicit amendment citations must reconcile any difference. - A silent discrepancy between registry and Methods is a `[REGISTRATION-TIMING]` finding, reported in Part C with `fixable_by_ai: false` (requires author action — file an amendment or correct Methods text). ### 5. Retrospective registration disclosure If any evidence suggests the registration was filed after data extraction began (registration date later than the reported extraction start, or the registry record's "current stage" indicates post-extraction filing), Methods must contain a disclosure paragraph. The absence of such disclosure in a retrospective-registered review is a `[REGISTRATION-TIMING]` finding. ## Flagging Rules - Any failure among items 1 – 5 is reported in Part C Action Items with the label `[REGISTRATION-TIMING]`. - Mark `fixable_by_ai: false` when reconciliation requires an external amendment filing or an author-supplied date. - Mark `fixable_by_ai: true` only when the fix is a Methods-text insertion of a known registration identifier or amendment date already disclosed elsewhere in the manuscript. ## JSON Field (Part D) Include a `registration_timing` object when this step runs: ```json "registration_timing": { "registry": "PROSPERO", "registration_id": "CRD########", "initial_registration_date": "YYYY-MM-DD or null", "amendments": [ { "date": "YYYY-MM-DD", "described_change": "..." } ], "timing_consistency": "CONSISTENT | DISCREPANCY | INCOMPLETE", "findings": ["free-text list of specific issues"] } ``` -
step4d_prisma_figure_audit.md 6.5 KB
# Step 4d — PRISMA Figure 1 Arithmetic & Cross-Reference Audit (Procedural Detail) Load-on-demand from `SKILL.md` Step 4d. Applies to PRISMA 2020 / PRISMA-DTA / PRISMA-P systematic reviews and meta-analyses where Item 16a (flow diagram) is PRESENT. ## Inputs | Source | Path | Required | |--------|------|----------| | Manuscript body | `manuscript/manuscript.md` (or path provided) | yes | | Figure 1 manifest | `analysis/figures/Figure1_PRISMA.md` (preferred) | one of | | Figure 1 PPTX | `analysis/figures/Figure1_PRISMA.pptx` (text-extractable) | one of | | Figure 1 caption | embedded in `manuscript.md` | one of | | Figure 1 PNG/SVG | `analysis/figures/Figure1_PRISMA.{png,svg}` | fallback (manual entry) | If no machine-readable Figure source exists, Step 4d emits `MISSING` for the cross-reference checks and asks the user to supply numbers manually. ## Number extraction (regex) ```python KEYWORDS = { "identified": r"(\d[\d,]*)\s+(?:records?|reports?)\s+identified", "duplicates": r"(\d[\d,]*)\s+(?:records?|reports?|duplicates?)\s+(?:removed|duplicates? removed)", "screened": r"(\d[\d,]*)\s+(?:records?|reports?)\s+screened", "excluded_screening": r"(\d[\d,]*)\s+(?:records?|reports?)\s+excluded(?:\s+(?:at|after|during)\s+screening)?", "sought": r"(\d[\d,]*)\s+(?:reports?|records?)\s+sought(?:\s+for\s+retrieval)?", "not_retrieved": r"(\d[\d,]*)\s+(?:reports?|records?)\s+(?:not\s+retrieved|unobtainable)", "retrieved": r"(\d[\d,]*)\s+(?:reports?|records?)\s+retrieved", "assessed": r"(\d[\d,]*)\s+(?:reports?|records?)\s+assessed(?:\s+for\s+eligibility)?", "excluded_eligibility": r"(\d[\d,]*)\s+(?:reports?|records?)\s+excluded(?:\s+with\s+reasons?)?", "included": r"(\d[\d,]*)\s+(?:studies|records?|reports?)\s+included", } ``` Apply to body text and Figure source independently → produce two dictionaries `body_numbers` and `figure_numbers`. ## Arithmetic checks (4) ```python def check_arithmetic(n: dict) -> list[dict]: results = [] if all(k in n for k in ["identified", "duplicates", "screened"]): ok = n["identified"] - n["duplicates"] == n["screened"] results.append({"eq": "screened = identified - duplicates", "lhs": n["screened"], "rhs": n["identified"] - n["duplicates"], "status": "PRESENT" if ok else "MISMATCH"}) if all(k in n for k in ["screened", "excluded_screening", "sought"]): ok = n["screened"] - n["excluded_screening"] == n["sought"] results.append({"eq": "sought = screened - excluded_screening", "lhs": n["sought"], "rhs": n["screened"] - n["excluded_screening"], "status": "PRESENT" if ok else "MISMATCH"}) if all(k in n for k in ["sought", "not_retrieved", "retrieved"]): ok = n["sought"] - n["not_retrieved"] == n["retrieved"] results.append({"eq": "retrieved = sought - not_retrieved", "lhs": n["retrieved"], "rhs": n["sought"] - n["not_retrieved"], "status": "PRESENT" if ok else "MISMATCH"}) if all(k in n for k in ["assessed", "excluded_eligibility", "included"]): ok = n["assessed"] - n["excluded_eligibility"] == n["included"] results.append({"eq": "included = assessed - excluded_eligibility", "lhs": n["included"], "rhs": n["assessed"] - n["excluded_eligibility"], "status": "PRESENT" if ok else "MISMATCH"}) return results ``` Run independently on `body_numbers` and `figure_numbers`. ## Cross-reference check (body ↔ figure) For each key in `KEYWORDS`: - Both present + agree → `PRESENT` - Both present + disagree → `MISMATCH` (record both values) - Only one source has it → `MISSING` (record which side) - Neither → skip ## JSON schema (`qc/prisma_figure_audit.json`) ```json { "manuscript": "manuscript/manuscript.md", "figure_source": "analysis/figures/Figure1_PRISMA.md", "body_numbers": { "identified": 315, "duplicates": 122, "screened": 186, "...": "..." }, "figure_numbers": { "identified": 315, "duplicates": 122, "screened": 186, "...": "..." }, "arithmetic_body": [ {"eq": "screened = identified - duplicates", "lhs": 186, "rhs": 193, "status": "MISMATCH"} ], "arithmetic_figure": [], "cross_reference": [ {"key": "screened", "body": 186, "figure": 186, "status": "PRESENT"} ], "audit_safe": false, "action_items": [ "[PRISMA-FIGURE] body arithmetic 'screened = identified - duplicates' fails: 186 vs 315-122=193" ] } ``` `audit_safe: true` ⟺ all arithmetic_body, arithmetic_figure, cross_reference rows are `PRESENT`. Anything else → `false` and Step 5 must surface action_items. ## Edge cases - **Multi-database identification**: PRISMA 2020 supports separate boxes for database vs register vs other methods. Sum across boxes for `identified` total — extraction regex must handle `321 records (213 from databases, 108 from citation searching)`. - **Citation searching strand**: separate flow on right side of PRISMA 2020 diagram. If present, run arithmetic checks on each strand independently + a combined `total identified` check. - **Dual-reviewer screening**: numbers should reflect post-consensus counts. If body reports pre/post-consensus separately, use post-consensus. - **Reports vs records**: PRISMA 2020 distinguishes records (citations) from reports (full-text). Regex captures both; treat them as the same key for arithmetic but flag inconsistent terminology (`[PRISMA-FIGURE-TERMINOLOGY]`). - **Duplicates split across stages**: some manuscripts list duplicates removed automatically (deduplication tool) separately from manual deduplication. Sum. - **Reasons for exclusion**: Step 4d does not enforce specific reasons but checks that the count of "with reasons" categories sums to `excluded_eligibility`. Add as optional check `excluded_eligibility = sum(reason_counts)`. ## Cross-cutting rules - `~/.claude/rules/numerical-safety.md`: PRISMA 5-way consistency (text ↔ Figure ↔ extraction CSV ↔ analysis script ↔ supplementary). Step 4d covers text ↔ Figure; extraction CSV ↔ script ↔ supplementary belong to `/meta-analysis` Phase 6 and `/write-paper` Step 7.3a. - `~/.claude/rules/manuscript-style-classical.md`: number formatting (Arabic numerals, thousands separator consistent with journal style). ## Related - `/check-reporting prisma` (this step caller) - `/write-paper` Step 7.3a (Numerical Claim Audit — different scope: pooled estimates, not flow diagram) - `/make-figures` (PRISMA flow diagram generation — produces `Figure1_PRISMA.md` manifest this step consumes)
-
-
scripts
-
check_checklist_exists.py 6.9 KB
#!/usr/bin/env python3 """Fail-fast guard for /check-reporting checklist routing. The /check-reporting skill assesses a manuscript against a vendored reporting guideline checklist (references/checklists/*.md). Historically, when a routed checklist file was absent the skill silently fell back to constructing the checklist "from its knowledge of the guideline" — a fabrication path that is exactly what the rest of the skill exists to prevent. This guard makes that case loud: - The requested guideline resolves to a vendored file that exists -> exit 0. - It resolves to a known guideline whose file is ABSENT -> exit 1, prints MISSING_CHECKLIST_CONTRACT_VIOLATION. - The name is not a recognised guideline -> exit 2, prints UNKNOWN_GUIDELINE. `--allow-from-memory` is the single explicit opt-in escape hatch: it downgrades exit 1/2 to exit 0 but always emits a NON-AUTHORITATIVE warning, so a construct-from-memory assessment can never happen silently. Source of truth for "what is vendored" is the checklists directory on disk, not this file — the alias map only translates display names to filename stems. Stdlib-only. """ from __future__ import annotations import argparse import re import sys from pathlib import Path # Display/guideline name (normalised) -> checklist filename stem. # Keys are normalised by _norm(): lowercased, with spaces/hyphens/dots/plus/ # underscores and a trailing 4-digit year stripped. The stem is the expected # file under references/checklists/<stem>.md. A stem whose file is absent is a # contract violation, surfaced loudly rather than silently fabricated. ALIAS_TO_STEM = { "strobe": "STROBE", "consort": "CONSORT", "consortai": "CONSORT_AI", "stard": "STARD", "stardai": "STARD_AI", "tripod": "TRIPOD", "tripodai": "TRIPOD_AI", "tripodllm": "TRIPOD_LLM", "prisma": "PRISMA_2020", "prismadta": "PRISMA_DTA", "prismap": "PRISMA_P", "prismascr": "PRISMA_ScR", "scopingreview": "PRISMA_ScR", "scoping": "PRISMA_ScR", "arrive": "ARRIVE_2", "care": "CARE", "spirit": "SPIRIT", "spiritai": "SPIRIT_AI", "claim": "CLAIM_2024", "decideai": "DECIDE_AI", "miclearllm": "MI_CLEAR_LLM", "squire": "SQUIRE_2", "clear": "CLEAR", "moose": "MOOSE", "grras": "GRRAS", "swim": "SWiM", "amstar": "AMSTAR2", "amstar2": "AMSTAR2", "quadas": "QUADAS2", "quadas2": "QUADAS2", "quadasc": "QUADAS_C", "rob2": "RoB2", "robinsi": "ROBINS_I", "robinse": "ROBINS_E", "robis": "ROBIS", "robme": "ROB_ME", "robnma": "RoB_NMA", "probast": "PROBAST", "probastai": "PROBAST_AI", "nos": "NOS", "cosmin": "COSMIN_RoB", "cosminrob": "COSMIN_RoB", "strobemr": "STROBE_MR", "pgsrs": "PGS_RS", "prsrs": "PGS_RS", "cheers": "CHEERS_2022", "cheers2022": "CHEERS_2022", "record": "RECORD", "remark": "REMARK", "tumormarker": "REMARK", "target": "TARGET", "targettrial": "TARGET", "targettrialemulation": "TARGET", "tte": "TARGET", "cross": "CROSS", "srqr": "SRQR", "coreq": "COREQ", "qualitative": "SRQR", } EXIT_OK = 0 EXIT_MISSING = 1 EXIT_UNKNOWN = 2 def _norm(name: str) -> str: """Normalise a guideline name to an alias key. Lowercase, drop a trailing 4-digit year, then strip spaces/hyphens/dots/ plus/underscores. "STARD-AI" -> "stardai"; "AMSTAR 2" -> "amstar2"; "CONSORT 2010" -> "consort"; "TRIPOD+AI" -> "tripodai". """ n = name.strip().lower() n = re.sub(r"\b(19|20)\d{2}\b", "", n) # drop version year n = re.sub(r"[\s\-_.+]", "", n) # strip spaces/-/_/./+ return n def checklist_dir(skill_dir: Path) -> Path: return skill_dir / "references" / "checklists" def resolve_stem(guideline: str) -> str | None: return ALIAS_TO_STEM.get(_norm(guideline)) def available_stems(cdir: Path) -> list[str]: if not cdir.is_dir(): return [] return sorted(p.stem for p in cdir.glob("*.md")) def main() -> int: parser = argparse.ArgumentParser( description="Fail-fast guard: confirm a routed reporting-guideline checklist is vendored." ) parser.add_argument("--guideline", help="Guideline name, e.g. 'STROBE', 'STARD-AI', 'CONSORT 2010'.") parser.add_argument( "--skill-dir", default=str(Path(__file__).resolve().parent.parent), help="check-reporting skill root (default: inferred from this script).", ) parser.add_argument( "--allow-from-memory", action="store_true", help="Explicit opt-in: downgrade a missing/unknown checklist to exit 0 with a " "NON-AUTHORITATIVE warning instead of failing. Never silent.", ) parser.add_argument( "--simulate-missing-checklist", action="store_true", help="Force the missing-file path (for contract tests).", ) parser.add_argument("--list", action="store_true", help="List vendored checklist files and exit.") args = parser.parse_args() skill_dir = Path(args.skill_dir).resolve() cdir = checklist_dir(skill_dir) if args.list: stems = available_stems(cdir) print(f"{len(stems)} vendored checklists in {cdir}:") for s in stems: print(f" {s}") return EXIT_OK if args.simulate_missing_checklist: print("MISSING_CHECKLIST_CONTRACT_VIOLATION: simulated missing checklist " "(guideline routed but no vendored file).", file=sys.stderr) if args.allow_from_memory: print("WARNING: --allow-from-memory set; a from-memory (NON-AUTHORITATIVE) " "assessment would proceed. The report MUST be marked non-vendored.") return EXIT_OK return EXIT_MISSING if not args.guideline: parser.error("--guideline is required (or use --list / --simulate-missing-checklist)") stem = resolve_stem(args.guideline) if stem is None: print(f"UNKNOWN_GUIDELINE: '{args.guideline}' is not a recognised reporting guideline.", file=sys.stderr) if args.allow_from_memory: print("WARNING: --allow-from-memory set; proceeding from memory for an unrecognised " "guideline (NON-AUTHORITATIVE). The report MUST be marked non-vendored.") return EXIT_OK return EXIT_UNKNOWN path = cdir / f"{stem}.md" if path.is_file(): print(f"OK: {args.guideline} -> references/checklists/{stem}.md") return EXIT_OK print(f"MISSING_CHECKLIST_CONTRACT_VIOLATION: '{args.guideline}' resolves to " f"'{stem}.md' which is not vendored under {cdir}.", file=sys.stderr) if args.allow_from_memory: print(f"WARNING: --allow-from-memory set; proceeding from memory for '{args.guideline}' " f"(NON-AUTHORITATIVE). The report MUST be marked non-vendored.") return EXIT_OK return EXIT_MISSING if __name__ == "__main__": sys.exit(main()) -
check_checklist_version.py 7.1 KB
#!/usr/bin/env python3 """ check_checklist_version.py — detect a reporting checklist that is stale relative to the current manuscript (A4b). A reporting checklist is frequently generated against an *older* manuscript version: its section/line references and version label no longer match the submitted manuscript, and a reviewer who cross-checks the line numbers sees the mismatch. This detector compares an existing checklist's target metadata against the current manuscript and flags a regenerate-needed condition. It reads the version contract emitted by check-reporting (v1.1+): `target_manuscript`, `target_version`, `source_sha256` — from a JSON checklist (`qc/reporting_checklist.json`) or from the text header / embedded JSON of a Markdown report. Comparison precedence: 1. `source_sha256` present and != current manuscript hash → STALE (content changed) 2. else `target_version` present and != current version → STALE (version bump) 3. else `target_manuscript` present and != current filename → STALE (different file) 4. no version metadata at all → UNVERIFIABLE (regen w/ v1.1) Exit: 0 = in sync, 1 = stale / unverifiable, 2 = usage/error. Stdlib-only. Usage: python3 check_checklist_version.py --checklist qc/reporting_checklist.json \ --manuscript manuscript_v8.md [--manuscript-version v8] \ [--out qc/checklist_version.json] [--quiet] """ from __future__ import annotations import argparse import hashlib import json import re import sys from pathlib import Path FILENAME_VERSION_RE = re.compile(r"[_\-.]v(\d{1,3})\b", re.IGNORECASE) HEADER_TARGET_FILE_RE = re.compile(r"Target manuscript file:\s*([^\n]+)", re.IGNORECASE) HEADER_TARGET_VER_RE = re.compile(r"Target version:\s*v?(\d{1,3})\b", re.IGNORECASE) JSON_BLOCK_RE = re.compile(r"```json\s*(\{.*?\})\s*```", re.DOTALL) def manuscript_identity(path: Path, explicit_version: str | None) -> dict: data = path.read_bytes() sha = hashlib.sha256(data).hexdigest()[:12] version = None if explicit_version: m = re.search(r"\d+", explicit_version) version = int(m.group(0)) if m else None if version is None: m = FILENAME_VERSION_RE.search(path.name) version = int(m.group(1)) if m else None return {"file": path.name, "version": version, "sha256": sha} def parse_checklist(path: Path) -> dict: """Extract target_manuscript / target_version / source_sha256 from a JSON or Markdown checklist. Returns a dict with those keys (values may be None).""" text = path.read_text(encoding="utf-8", errors="replace") out = {"target_manuscript": None, "target_version": None, "source_sha256": None} obj = None if path.suffix.lower() == ".json": try: obj = json.loads(text) except Exception: obj = None if obj is None: m = JSON_BLOCK_RE.search(text) if m: try: obj = json.loads(m.group(1)) except Exception: obj = None if isinstance(obj, dict): out["target_manuscript"] = obj.get("target_manuscript") out["source_sha256"] = obj.get("source_sha256") tv = obj.get("target_version") if tv is not None: mm = re.search(r"\d+", str(tv)) out["target_version"] = int(mm.group(0)) if mm else None # Markdown header fallbacks (only fill what JSON did not provide) if out["target_manuscript"] is None: m = HEADER_TARGET_FILE_RE.search(text) if m: out["target_manuscript"] = m.group(1).strip() if out["target_version"] is None: m = HEADER_TARGET_VER_RE.search(text) if m: out["target_version"] = int(m.group(1)) return out def evaluate(current: dict, checklist: dict) -> list[dict]: findings: list[dict] = [] ck_sha = checklist.get("source_sha256") ck_ver = checklist.get("target_version") ck_file = checklist.get("target_manuscript") if ck_sha and current["sha256"] and ck_sha != current["sha256"]: findings.append({ "type": "checklist_content_stale", "severity": "stale", "detail": f"checklist source_sha256 {ck_sha} != current {current['sha256']} " "— manuscript content changed since the checklist was generated"}) elif ck_ver is not None and current["version"] is not None and ck_ver != current["version"]: findings.append({ "type": "checklist_version_stale", "severity": "stale", "detail": f"checklist targets v{ck_ver} but current is v{current['version']}"}) elif ck_file and ck_file != current["file"]: findings.append({ "type": "checklist_file_mismatch", "severity": "stale", "detail": f"checklist targets '{ck_file}' but current is '{current['file']}'"}) if ck_sha is None and ck_ver is None and ck_file is None: findings.append({ "type": "checklist_no_version_metadata", "severity": "unverifiable", "detail": "checklist carries no target_manuscript/target_version/source_sha256 " "(pre-v1.1 contract) — cannot verify; regenerate with current check-reporting"}) return findings def main(argv: list[str] | None = None) -> int: ap = argparse.ArgumentParser( description="Flag a reporting checklist that is stale vs the current manuscript.") ap.add_argument("--checklist", type=Path, required=True, help="qc/reporting_checklist.json or the .md report.") ap.add_argument("--manuscript", type=Path, required=True, help="Current manuscript file.") ap.add_argument("--manuscript-version", default=None, help="Current version token (else inferred from filename).") ap.add_argument("--out", type=Path, default=None, help="Write JSON report here.") ap.add_argument("--quiet", action="store_true", help="Suppress stdout summary.") args = ap.parse_args(argv) if not args.checklist.is_file(): print(f"ERROR: --checklist not a file: {args.checklist}", file=sys.stderr) return 2 if not args.manuscript.is_file(): print(f"ERROR: --manuscript not a file: {args.manuscript}", file=sys.stderr) return 2 current = manuscript_identity(args.manuscript, args.manuscript_version) checklist = parse_checklist(args.checklist) findings = evaluate(current, checklist) safe = not findings report = {"submission_safe": safe, "current": current, "checklist": checklist, "findings": findings} if args.out is not None: args.out.parent.mkdir(parents=True, exist_ok=True) args.out.write_text(json.dumps({"detector": "check_checklist_version", **report}, indent=2), encoding="utf-8") if not args.quiet: if safe: print(f"PASS: checklist matches current manuscript (v{current['version']}, " f"{current['sha256']}).") else: print("FAIL: checklist is stale or unverifiable:") for f in findings: print(f" - [{f['severity']}] {f['type']}: {f['detail']}") return 0 if safe else 1 if __name__ == "__main__": sys.exit(main()) -
check_framework_naming.py 8.3 KB
#!/usr/bin/env python3 """Reporting-framework naming discipline audit (check-reporting Step 4e). A base reporting tool and its AI/extension are distinct instruments with separate citations. Manuscripts routinely (a) invoke an extension (PROBAST+AI, STARD-AI, TRIPOD+AI, PRISMA-DTA) without ever naming or citing the base instrument it extends, (b) mix hyphenation for the same family within one document (PROBAST+AI 13x next to PROBAST-AI 2x), (c) coin item labels like "12-AI", or (d) wave at "recent guidance" instead of naming the framework. Each is a reviewer red flag and each is a deterministic grep. INPUTS --manuscript manuscript markdown/text (required). CHECKS BASE_MISSING (Major) an extension is used but its base instrument name never appears standalone in the document. HYPHEN_MIX (Minor) both "<FAMILY>+AI" and "<FAMILY>-AI" occur — inconsistent. CITE_MISSING (Minor) the sentence first invoking an extension carries no citation marker. SELF_COINED_LABEL (Minor) a self-coined "<digits>-AI" item label. VAGUE_GUIDANCE (Minor) "adapted per recent guidance"-style wording with no named framework. OUTPUT A reconciliation table (stdout) and, with --out, a JSON artifact: {manuscript, claims[{verdict, severity, detail, where}], summary} Exit 1 (with --strict) when any Major-severity claim exists. Stdlib-only (json / re / argparse / pathlib). Exit codes: 0 clean (or report-only), 1 Major claim(s) found (with --strict), 2 input/usage error. """ from __future__ import annotations import argparse import json import re import sys from pathlib import Path # Extension regex -> base instrument name. The base must appear standalone (not as a # prefix of the extension) somewhere in the document. EXT_TO_BASE = [ (r"PROBAST\s*\+\s*AI", "PROBAST"), (r"PROBAST-AI", "PROBAST"), (r"TRIPOD\s*\+\s*AI", "TRIPOD"), (r"TRIPOD-AI", "TRIPOD"), (r"TRIPOD-LLM", "TRIPOD"), (r"STARD-AI", "STARD"), (r"STARD\s*\+\s*AI", "STARD"), (r"CONSORT-AI", "CONSORT"), (r"SPIRIT-AI", "SPIRIT"), (r"PRISMA-DTA", "PRISMA"), (r"QUADAS-C", "QUADAS"), ] HYPHEN_FAMILIES = ("PROBAST", "TRIPOD", "STARD", "CONSORT", "SPIRIT", "QUADAS", "DECIDE") CITE_MARKER = re.compile(r"\[\d|\[@|\bet al\.?|\(\s*[A-Z][A-Za-z]+,?\s+\d{4}|\b\d{4}[a-z]?\)") SELF_COINED = re.compile(r"\b\d{1,2}-AI\b") VAGUE = re.compile( r"\b(?:adapted|adjusted|modified|aligned|updated|following|per|in line with)\b" r"[\w\s,]{0,20}?" r"\b(?:recent|current|emerging|latest|evolving)\b\s+" r"(?:best[- ]practice|guidance|practice|recommendations?|standards?|guidelines?)", re.I) # VAGUE only counts inside a reporting-framework context; otherwise "recent # best-practice recommendations" about a method is a false positive (F05). REPORTING_CUE = re.compile( r"\b(?:report(?:ing|ed)?|checklist|EQUATOR|reporting\s+standard|" r"reporting\s+framework)\b", re.I) def _sentences(text: str) -> list[str]: units = [] for para in re.split(r"\n[ \t]*\n", text): flat = re.sub(r"\s*\n\s*", " ", para).strip() units.extend(re.split(r"(?<=[.;])\s+", flat)) return [u for u in units if u.strip()] def check(text: str) -> list[dict]: claims = [] # BASE_MISSING + first-use CITE_MISSING flagged_base = set() for ext_re, base in EXT_TO_BASE: m = re.search(ext_re, text, re.I) if not m: continue standalone = re.search(rf"\b{base}\b(?!\s*[-+]\s*(?:AI|LLM))(?!-(?:AI|LLM|DTA|C\b))", text, re.I) if not standalone and base not in flagged_base: flagged_base.add(base) claims.append({ "verdict": "BASE_MISSING", "severity": "Major", "detail": (f"the extension '{m.group(0)}' is used but the base instrument " f"'{base}' is never named standalone (name and cite both)"), "where": m.group(0), }) # citation near first use sent = next((s for s in _sentences(text) if re.search(ext_re, s, re.I)), "") if sent and not CITE_MARKER.search(sent): claims.append({ "verdict": "CITE_MISSING", "severity": "Minor", "detail": (f"first use of '{m.group(0)}' has no citation marker in its " f"sentence"), "where": sent.strip()[:160], }) # HYPHEN_MIX for fam in HYPHEN_FAMILIES: plus = len(re.findall(rf"{fam}\s*\+\s*AI", text, re.I)) hyph = len(re.findall(rf"{fam}-AI", text, re.I)) if plus and hyph: claims.append({ "verdict": "HYPHEN_MIX", "severity": "Minor", "detail": (f"'{fam}+AI' ({plus}x) and '{fam}-AI' ({hyph}x) both used — " f"pick one hyphenation"), "where": fam, }) # SELF_COINED_LABEL for m in dict.fromkeys(SELF_COINED.findall(text)): claims.append({ "verdict": "SELF_COINED_LABEL", "severity": "Minor", "detail": f"self-coined AI item label '{m}' — use the framework's own item numbering", "where": m, }) # VAGUE_GUIDANCE — only when the sentence is clearly about REPORTING (a reporting # guideline / checklist) yet names no specific framework. Gating on a reporting cue # prevents firing on method-level wording like "external validation following recent # best-practice recommendations", which is not a reporting-framework claim at all. for s in _sentences(text): m = VAGUE.search(s) if m and REPORTING_CUE.search(s): claims.append({ "verdict": "VAGUE_GUIDANCE", "severity": "Minor", "detail": (f"vague wording '{m.group(0).strip()}' in a reporting context — name " f"the specific framework and cite it"), "where": m.group(0).strip()[:80], }) return claims def analyze(manuscript: str) -> dict: p = Path(manuscript) if not p.is_file(): sys.stderr.write(f"ERROR: manuscript not found: {manuscript}\n") sys.exit(2) claims = check(p.read_text(encoding="utf-8")) n_major = sum(1 for c in claims if c["severity"] == "Major") return { "manuscript": str(p), "claims": claims, "summary": { "n_claims": len(claims), "n_major": n_major, "n_flag": len(claims) - n_major, "verdict": "MAJOR_CANDIDATE" if n_major else ("FLAG" if claims else "OK"), }, } def render(result: dict) -> str: lines = ["| Check | Severity | Detail |", "|---|---|---|"] for c in result["claims"]: lines.append(f"| {c['verdict']} | {c['severity']} | {c['detail']} |") if len(lines) == 2: lines.append("| (none) | — | reporting-framework naming is disciplined |") return "\n".join(lines) def main() -> int: ap = argparse.ArgumentParser(description="Reporting-framework naming audit (Step 4e).") ap.add_argument("--manuscript", required=True, help="manuscript markdown/text") ap.add_argument("--out", help="write JSON artifact to this path") ap.add_argument("--strict", action="store_true", help="exit 1 if any Major claim exists") ap.add_argument("--quiet", action="store_true", help="suppress stdout table") args = ap.parse_args() result = analyze(args.manuscript) if not args.quiet: print("=" * 41) print(" Framework Naming (Step 4e)") print("=" * 41) print(render(result)) print() s = result["summary"] if s["n_major"]: print(f"MAJOR candidate: {s['n_major']} base-instrument naming gap(s).") elif s["n_flag"]: print(f"FLAG: {s['n_flag']} naming/citation inconsistency(ies).") else: print("OK: reporting-framework naming is disciplined.") if args.out: Path(args.out).parent.mkdir(parents=True, exist_ok=True) Path(args.out).write_text(json.dumps({"detector": "check_framework_naming", **result}, indent=2), encoding="utf-8") if not args.quiet: print(f"\nwrote {args.out}") return 1 if (args.strict and result["summary"]["n_major"]) else 0 if __name__ == "__main__": sys.exit(main()) -
check_prisma_figure.py 7.6 KB
#!/usr/bin/env python3 """PRISMA Figure 1 Arithmetic & Cross-Reference Audit. Loaded by `/check-reporting prisma` Step 4d. Inputs: --md manuscript markdown (body text PRISMA numbers). --figure Figure 1 source: markdown manifest, caption .md, or text export. --out output JSON (default: qc/prisma_figure_audit.json). Outputs: qc/prisma_figure_audit.json + console table. Exit codes: 0 audit_safe (all PRESENT) 1 arithmetic or cross-reference MISMATCH / MISSING 2 invalid input (file missing, parse failure) """ from __future__ import annotations import argparse import json import re import sys from pathlib import Path KEYWORDS = { "identified": r"(\d[\d,]*)\s+(?:records?|reports?)\s+identified", "duplicates": r"(\d[\d,]*)\s+(?:records?|reports?|duplicates?)\s+(?:duplicates?\s+)?removed", "screened": r"(\d[\d,]*)\s+(?:records?|reports?)\s+screened", # Explicit suffix required to avoid collision with excluded_eligibility. "excluded_screening": r"(\d[\d,]*)\s+(?:records?|reports?)\s+excluded\s+(?:at|after|during)\s+screening", "sought": r"(\d[\d,]*)\s+(?:reports?|records?)\s+sought(?:\s+for\s+retrieval)?", "not_retrieved": r"(\d[\d,]*)\s+(?:reports?|records?)\s+(?:not\s+retrieved|unobtainable)", # Match both "reports retrieved" and "retrieved 186 reports". "retrieved": r"(?:(\d[\d,]*)\s+(?:records?|reports?)\s+retrieved|retrieved\s+(\d[\d,]*)\s+(?:records?|reports?))", "assessed": r"(\d[\d,]*)\s+(?:reports?|records?)\s+assessed(?:\s+for\s+eligibility)?", # Explicit suffix required to avoid collision with excluded_screening. "excluded_eligibility": r"(\d[\d,]*)\s+(?:records?|reports?)\s+excluded\s+with\s+reasons?", "included": r"(\d[\d,]*)\s+(?:studies|records?|reports?)\s+included", } def extract_numbers(text: str) -> dict[str, int]: """Return {key: int} for keywords that match. First match wins per key.""" out: dict[str, int] = {} for key, pattern in KEYWORDS.items(): m = re.search(pattern, text, flags=re.IGNORECASE) if m: num = next((g for g in m.groups() if g), None) if num is None: continue try: out[key] = int(num.replace(",", "")) except ValueError: continue return out def check_arithmetic(n: dict[str, int]) -> list[dict]: eqs = [ ("screened = identified - duplicates", "screened", "identified", "duplicates"), ("sought = screened - excluded_screening", "sought", "screened", "excluded_screening"), ("retrieved = sought - not_retrieved", "retrieved", "sought", "not_retrieved"), ("included = assessed - excluded_eligibility", "included", "assessed", "excluded_eligibility"), ] out = [] for label, lhs_key, a_key, b_key in eqs: if all(k in n for k in (lhs_key, a_key, b_key)): lhs = n[lhs_key] rhs = n[a_key] - n[b_key] out.append({ "eq": label, "lhs": lhs, "rhs": rhs, "status": "PRESENT" if lhs == rhs else "MISMATCH", }) else: missing = [k for k in (lhs_key, a_key, b_key) if k not in n] out.append({ "eq": label, "lhs": None, "rhs": None, "status": "MISSING", "missing_keys": missing, }) return out def cross_reference(body: dict[str, int], figure: dict[str, int]) -> list[dict]: out = [] keys = set(body) | set(figure) for key in sorted(keys): b = body.get(key) f = figure.get(key) if b is None and f is None: continue if b is None: status = "MISSING" note = "body lacks number" elif f is None: status = "MISSING" note = "figure lacks number" elif b == f: status = "PRESENT" note = "" else: status = "MISMATCH" note = f"body={b}, figure={f}" out.append({"key": key, "body": b, "figure": f, "status": status, "note": note}) return out def render_table(rows: list[dict], cols: list[str]) -> str: if not rows: return " (none)" widths = {c: max(len(c), max(len(str(r.get(c, ""))) for r in rows)) for c in cols} head = " " + " ".join(c.ljust(widths[c]) for c in cols) sep = " " + " ".join("-" * widths[c] for c in cols) body = "\n".join( " " + " ".join(str(r.get(c, "")).ljust(widths[c]) for c in cols) for r in rows ) return "\n".join([head, sep, body]) def main() -> int: ap = argparse.ArgumentParser(description=__doc__) ap.add_argument("--md", required=True, help="manuscript markdown path") ap.add_argument("--figure", required=True, help="Figure 1 source path (.md, .txt)") ap.add_argument("--out", default="qc/prisma_figure_audit.json", help="output JSON path") args = ap.parse_args() md_path = Path(args.md) fig_path = Path(args.figure) if not md_path.exists(): print(f"ERROR: manuscript not found: {md_path}", file=sys.stderr) return 2 if not fig_path.exists(): print(f"ERROR: figure source not found: {fig_path}", file=sys.stderr) return 2 body_text = md_path.read_text(encoding="utf-8") fig_text = fig_path.read_text(encoding="utf-8") body_numbers = extract_numbers(body_text) figure_numbers = extract_numbers(fig_text) arith_body = check_arithmetic(body_numbers) arith_figure = check_arithmetic(figure_numbers) xref = cross_reference(body_numbers, figure_numbers) audit_safe = ( all(r["status"] == "PRESENT" for r in arith_body) and all(r["status"] == "PRESENT" for r in arith_figure) and all(r["status"] == "PRESENT" for r in xref) ) action_items: list[str] = [] for r in arith_body: if r["status"] == "MISMATCH": action_items.append( f"[PRISMA-FIGURE] body arithmetic '{r['eq']}' fails: {r['lhs']} vs {r['rhs']}" ) for r in arith_figure: if r["status"] == "MISMATCH": action_items.append( f"[PRISMA-FIGURE] figure arithmetic '{r['eq']}' fails: {r['lhs']} vs {r['rhs']}" ) for r in xref: if r["status"] == "MISMATCH": action_items.append( f"[PRISMA-FIGURE] cross-ref '{r['key']}' MISMATCH: body={r['body']}, figure={r['figure']}" ) result = { "manuscript": str(md_path), "figure_source": str(fig_path), "body_numbers": body_numbers, "figure_numbers": figure_numbers, "arithmetic_body": arith_body, "arithmetic_figure": arith_figure, "cross_reference": xref, "audit_safe": audit_safe, "action_items": action_items, } out_path = Path(args.out) out_path.parent.mkdir(parents=True, exist_ok=True) out_path.write_text(json.dumps({"detector": "check_prisma_figure", **result}, indent=2, ensure_ascii=False), encoding="utf-8") print(f"== PRISMA Figure Audit — {md_path.name} vs {fig_path.name} ==\n") print("Body arithmetic:") print(render_table(arith_body, ["eq", "lhs", "rhs", "status"])) print("\nFigure arithmetic:") print(render_table(arith_figure, ["eq", "lhs", "rhs", "status"])) print("\nCross-reference (body ↔ figure):") print(render_table(xref, ["key", "body", "figure", "status", "note"])) print(f"\naudit_safe: {audit_safe}") print(f"output: {out_path}") if action_items: print("\nAction items:") for a in action_items: print(f" - {a}") return 0 if audit_safe else 1 if __name__ == "__main__": sys.exit(main()) -
prisma_cascade_check.py 9.2 KB
#!/usr/bin/env python3 """ prisma_cascade_check.py — PRISMA flow cascade arithmetic auto-verify. PRISMA 2020 flow diagrams chain a cascade of subtractions: [identified across databases] → [after dedup] → [title/abstract screened] → [full-text reviewed] → [included in qualitative synthesis] → [included in quantitative synthesis] Each transition has a corresponding "excluded" count. The arithmetic is trivial in principle but reviewers and editors find off-by-one errors at high frequency: the prose cascade `151 + 108 + 39 + 1 + 1 + 4 = 304` is followed by a prose summary "305" four lines later. The desk-reject follows immediately because the prose is presented as fact. This script: 1. Reads round-by-round screening TSV artifacts (`round1.tsv`, `round2.tsv`, `round3_adjudication.tsv`). 2. Computes the canonical flow chain from raw decisions. 3. Optionally cross-checks against a manuscript markdown body and emits a per-stage drift report. Usage ===== python prisma_cascade_check.py \\ --round1 2_Screening/round1.tsv \\ --round2 2_Screening/round2.tsv \\ --round3 2_Screening/round3_adjudication.tsv \\ --out qc/prisma_cascade.json python prisma_cascade_check.py \\ --round1 ... --round2 ... --round3 ... \\ --manuscript manuscript.md \\ --out qc/prisma_cascade.json Decision column defaults: round 1 / round 2: `decision` ∈ {INCLUDE, EXCLUDE, MAYBE} round 3: `round3_decision` ∈ {INCLUDE, EXCLUDE} Output JSON: { "submission_safe": false, "stage_counts": { "round1_total": 458, "round1_include": 220, "round2_include": 87, "round3_include": 42, "round3_exclude": 45 }, "cascade_arithmetic": { "round1_to_round2": {"excluded": 238, "checked": true}, "round2_to_round3": {"excluded": 133, "checked": true} }, "manuscript_drift": [ {"stage": "round3_include", "computed": 42, "manuscript": 43, "manuscript_line": 67} ] } Exit codes: 0 — no drift (and manuscript agrees if supplied) 1 — drift between computed cascade and manuscript prose 2 — invocation error """ from __future__ import annotations import argparse import csv import json import re import sys from dataclasses import dataclass, field from pathlib import Path def count_decisions(tsv_path: Path, decision_col: str) -> dict[str, int]: """Return {decision: count} for the given TSV.""" delim = "," if tsv_path.suffix.lower() == ".csv" else "\t" counts: dict[str, int] = {} total = 0 with tsv_path.open("r", encoding="utf-8", newline="") as fh: reader = csv.DictReader(fh, delimiter=delim) if reader.fieldnames is None or decision_col not in reader.fieldnames: print( f"ERROR: decision column {decision_col!r} not in {tsv_path}", file=sys.stderr, ) sys.exit(2) for row in reader: decision = (row.get(decision_col) or "").strip().upper() if decision: counts[decision] = counts.get(decision, 0) + 1 total += 1 counts["_total"] = total return counts def search_manuscript_stage(text: str, stage_name: str) -> tuple[int | None, int | None]: """Best-effort grep for a stage count and its line number. Stage names map to a small dictionary of cascade phrases. Returns (value, lineno) or (None, None) when no match found. """ patterns = { "round1_include": [ r"(\d{1,3}(?:,\d{3})*)\s+records?\s+(?:were\s+)?(?:included\s+)?after\s+title.{0,40}abstract", r"(\d{1,3}(?:,\d{3})*)\s+records?\s+screened\s+by\s+title", ], "round2_include": [ r"(\d{1,3}(?:,\d{3})*)\s+records?\s+(?:moved\s+forward\s+)?to\s+full.?text", r"(\d{1,3}(?:,\d{3})*)\s+records?\s+(?:were\s+)?retrieved\s+for\s+full.?text", ], "round3_include": [ r"(\d{1,3}(?:,\d{3})*)\s+studies?\s+(?:were\s+)?included\s+in\s+(?:the\s+)?(?:final\s+)?(?:qualitative|quantitative)\s+synthesis", r"(\d{1,3}(?:,\d{3})*)\s+studies?\s+(?:were\s+)?included\s+in\s+(?:the\s+)?(?:meta-analysis|review)", r"(?:included|comprised|finally\s+included)\s+(\d{1,3}(?:,\d{3})*)\s+studies?", ], } pats = patterns.get(stage_name) if not pats: return None, None for lineno, line in enumerate(text.splitlines(), start=1): for pat in pats: m = re.search(pat, line, flags=re.IGNORECASE) if m: try: val = int(m.group(1).replace(",", "")) except ValueError: continue return val, lineno return None, None @dataclass class CascadeReport: submission_safe: bool stage_counts: dict[str, int] = field(default_factory=dict) cascade_arithmetic: dict[str, dict] = field(default_factory=dict) manuscript_drift: list[dict] = field(default_factory=list) def build_report( round1: Path, round2: Path, round3: Path, manuscript: Path | None, r1_col: str, r2_col: str, r3_col: str, ) -> CascadeReport: r1 = count_decisions(round1, r1_col) r2 = count_decisions(round2, r2_col) r3 = count_decisions(round3, r3_col) stage_counts = { "round1_total": r1.get("_total", 0), "round1_include": r1.get("INCLUDE", 0) + r1.get("MAYBE", 0), "round2_include": r2.get("INCLUDE", 0) + r2.get("MAYBE", 0), "round3_include": r3.get("INCLUDE", 0), "round3_exclude": r3.get("EXCLUDE", 0), } cascade: dict[str, dict] = { "round1_to_round2": { "excluded": stage_counts["round1_total"] - stage_counts["round2_include"], "checked": True, }, "round2_to_round3": { "excluded": ( stage_counts["round2_include"] - (stage_counts["round3_include"] + stage_counts["round3_exclude"]) ), "checked": True, }, } drifts: list[dict] = [] if manuscript is not None and manuscript.is_file(): text = manuscript.read_text(encoding="utf-8") for stage in ("round1_include", "round2_include", "round3_include"): val, lineno = search_manuscript_stage(text, stage) if val is None: continue if val != stage_counts[stage]: drifts.append( { "stage": stage, "computed": stage_counts[stage], "manuscript": val, "manuscript_line": lineno, } ) submission_safe = not drifts return CascadeReport( submission_safe=submission_safe, stage_counts=stage_counts, cascade_arithmetic=cascade, manuscript_drift=drifts, ) def main(argv: list[str] | None = None) -> int: parser = argparse.ArgumentParser(description="PRISMA flow cascade arithmetic auto-verify.") parser.add_argument("--round1", type=Path, required=True) parser.add_argument("--round2", type=Path, required=True) parser.add_argument("--round3", type=Path, required=True) parser.add_argument("--manuscript", type=Path, default=None) parser.add_argument("--r1-col", default="decision") parser.add_argument("--r2-col", default="decision") parser.add_argument("--r3-col", default="round3_decision") parser.add_argument("--out", type=Path, default=Path("qc/prisma_cascade.json")) parser.add_argument("--quiet", action="store_true") args = parser.parse_args(argv) for p, label in ( (args.round1, "--round1"), (args.round2, "--round2"), (args.round3, "--round3"), ): if not p.is_file(): print(f"ERROR: {label} not a file: {p}", file=sys.stderr) return 2 report = build_report( args.round1, args.round2, args.round3, args.manuscript, args.r1_col, args.r2_col, args.r3_col, ) args.out.parent.mkdir(parents=True, exist_ok=True) args.out.write_text( json.dumps( { "submission_safe": report.submission_safe, "stage_counts": report.stage_counts, "cascade_arithmetic": report.cascade_arithmetic, "manuscript_drift": report.manuscript_drift, }, indent=2, ), encoding="utf-8", ) if not args.quiet: if report.submission_safe: print( "PASS: cascade computed. Stages: " + " → ".join( f"{k}={v}" for k, v in report.stage_counts.items() if not k.startswith("_") ) ) else: print( f"FAIL: {len(report.manuscript_drift)} manuscript-prose " "drift(s) vs computed cascade." ) for d in report.manuscript_drift: print( f" {d['stage']}: computed={d['computed']} " f"manuscript={d['manuscript']} (line {d['manuscript_line']})" ) return 0 if report.submission_safe else 1 if __name__ == "__main__": sys.exit(main()) -
verify_checklist_fidelity.py 8.5 KB
#!/usr/bin/env python3 """A bundled checklist must match the official instrument it claims to be. Issue #352 (an external report, 2026-07-21): the file labelled "TRIPOD+AI 2024" was actually TRIPOD 2015 with separately-numbered `-AI` additions — the older section sequence, non-canonical item identifiers (`1-AI`, `10-AI-a`, …), and no Open Science or Patient-and-Public-Involvement items. The official TRIPOD+AI 2024 (Collins et al., BMJ 2024;385:e078378) is a **rewrite**: 27 main items, 52 subitems, with Open science (18) and PPI (19) as first-class items. Nothing caught it. `check_checklist_exists` verifies the file is *present*; `check_framework_naming` verifies it *names* its base instrument. Neither compares the file's item inventory against the official one — so a checklist could silently drift from the guideline it claims to reproduce, and an audit could report "TRIPOD+AI compliant" while checking a different, older instrument. This is that check. It is manifest-driven, so it generalises: each entry states the official inventory (item count, required sections, forbidden structures, a source token) and this script holds the bundled file to it. Add a guideline by adding an EXPECTED entry, not code. The same audit (2026-07-21) found two more of the #352 class and they are covered here: CLEAR had been regrouped into seven invented topical "domains" (item 1 = "Study hypothesis") when the official instrument is numbered by manuscript section (item 1 = Title, item 44 = baseline demographics), and MI-CLEAR-LLM carried the 2024 six-item body under a "Version 2025" label when the official 2025 update has eight item categories. Not named `check_*` on purpose — it is a fidelity regression, run in CI, not one of the manuscript integrity detectors in the published count. Usage: verify_checklist_fidelity.py [--strict] [--root PATH] Stdlib only. """ from __future__ import annotations import argparse import re import sys from pathlib import Path ROOT = Path(__file__).resolve().parents[3] CHECKLISTS = "skills/check-reporting/references/checklists" # The official inventory each bundled checklist must reproduce. Sourced from the published statement, # not from the file being checked — that is the whole point. EXPECTED = { "TRIPOD_AI.md": { "official_section_start": "## Checklist Items", "official_section_end": "## MedSci supplemental", # supplemental checks are ours, exempt "main_items": list(range(1, 28)), # 1..27 "subitem_rows": 52, "required_headings": ["### Open science", "### Patient and public involvement"], "forbidden_in_official": r"\b\d+-AI\b", # non-canonical identifiers "must_contain": ["10.1136/bmj-2023-078378", "supersedes and replaces TRIPOD 2015"], "source": "Collins GS et al. BMJ 2024;385:e078378 (TRIPOD+AI 2024)", }, # Issue-#352 class, found by the same fidelity audit (2026-07-21): the bundled CLEAR invented a # 7-topical-domain taxonomy (item 1 = "Study hypothesis") — but official CLEAR is numbered by # MANUSCRIPT SECTION (item 1 = Title, 2 = Abstract, 44 = baseline demographics), and its only two # non-essential items are 53 and 58 (the file wrongly said 17 and 57). Every cited item number was wrong. "CLEAR.md": { "official_section_start": "## Checklist Items", "official_section_end": "## Notes for assessors", "main_items": list(range(1, 59)), # 1..58 "subitem_rows": 58, "required_headings": ["### Title", "### Abstract", "### Results", "### Discussion"], "forbidden_in_official": r"Domain \d", # the topical-domain regrouping tell "must_contain": ["10.1186/s13244-023-01415-8", "specifying the radiomic methodology"], "source": "Kocak B et al. Insights Imaging 2023;14(1):75 (CLEAR 2023)", }, # Issue-#352 class (2026-07-21): the file was labelled "Version 2025" but carried the 2024 SIX-item # body. The official 2025 update has EIGHT item categories, promoting Access mode, Input data type, # and Adaptation strategy to first-class items. "MI_CLEAR_LLM.md": { "official_section_start": "## Checklist Items", "official_section_end": "## Notes for assessors", "main_items": list(range(1, 9)), # 1..8 "subitem_rows": 8, "required_headings": ["### 2. Access mode", "### 3. Input data type", "### 4. Adaptation strategy used"], "must_contain": ["10.3348/kjr.2025.1522", "8 item categories"], "source": "Park SH et al. Korean J Radiol 2025;26(12):1123-1132 (MI-CLEAR-LLM 2025)", }, # GATHER 2016 (Stevens et al.): 18 items in four sections (Objectives and funding; # Data inputs; Data analysis; Results and discussion). Registered so the burden-of-disease # reporting checklist cannot silently drift from the 18-item statement. "GATHER.md": { "official_section_start": "## Checklist Items", "official_section_end": "## MedSci application notes", "main_items": list(range(1, 19)), # 1..18 "subitem_rows": 18, "required_headings": ["### Objectives and funding", "### Data inputs", "### Data analysis", "### Results and discussion"], "must_contain": ["10.1371/journal.pmed.1002056", "Health Estimates Reporting"], "source": "Stevens GA et al. Lancet 2016;388:e19-23 / PLoS Med 2016;13(6):e1002056 (GATHER)", }, } def official_slice(text: str, spec: dict) -> str: s = text.find(spec["official_section_start"]) e = text.find(spec["official_section_end"]) if s < 0: return text return text[s : e if e > s else len(text)] def check_one(path: Path, spec: dict) -> list[str]: out: list[str] = [] if not path.is_file(): return [f"{path.name}: file missing"] text = path.read_text(encoding="utf-8") official = official_slice(text, spec) # main item numbers present in the official section nums = sorted({int(m) for m in re.findall(r"^\|\s*(\d+)[a-g]?\s*\|", official, re.MULTILINE)}) want = spec["main_items"] if nums != want: missing = [n for n in want if n not in nums] extra = [n for n in nums if n not in want] out.append(f"{path.name}: main items are {nums or '[]'} — " f"expected {want[0]}..{want[-1]}" + (f"; missing {missing}" if missing else "") + (f"; unexpected {extra}" if extra else "")) rows = len(re.findall(r"^\|\s*\d+[a-g]?\s*\|", official, re.MULTILINE)) if rows != spec["subitem_rows"]: out.append(f"{path.name}: {rows} checklist rows in the official section — expected " f"{spec['subitem_rows']} subitems ({spec['source']}).") for h in spec["required_headings"]: if h not in text: out.append(f"{path.name}: missing required section '{h}' — it is an official item, not optional.") forbidden = spec.get("forbidden_in_official") if forbidden: bad = re.findall(forbidden, official) if bad: out.append(f"{path.name}: forbidden pattern in the official section: " f"{sorted(set(bad))} — this marks a structure the official instrument does not use " f"(non-canonical identifiers, or a grouping the guideline does not have).") for tok in spec["must_contain"]: if tok not in text: out.append(f"{path.name}: missing required marker {tok!r} (source DOI / version / framing).") return out def main() -> int: ap = argparse.ArgumentParser(description=__doc__.split("\n")[0]) ap.add_argument("--strict", action="store_true") ap.add_argument("--root", type=Path, default=ROOT) a = ap.parse_args() problems: list[str] = [] for name, spec in EXPECTED.items(): problems += check_one(a.root / CHECKLISTS / name, spec) if not problems: print(f"OK: {len(EXPECTED)} bundled checklist(s) match their official item inventory.") return 0 print(f"CHECKLIST_FIDELITY: {len(problems)} discrepanc(ies) — a bundled checklist does not match " f"the official instrument it claims to be.\n") for p in problems: print(f" - {p}") print( "\nA checklist labelled with an official guideline's name must reproduce that guideline's item\n" "inventory, or a compliance audit checks a different instrument than the one it reports.\n" ) return 1 if a.strict else 0 if __name__ == "__main__": sys.exit(main())
-
-
tests
-
fixtures
-
framework_bad.md 247 B
# Methods Risk of bias was assessed with PROBAST+AI. Some studies used PROBAST-AI signalling questions instead. Diagnostic accuracy reporting followed STARD-AI as adapted per recent guidance. We applied a 12-AI item subset to score the models. -
framework_clean.md 242 B
# Methods Risk of bias was assessed using PROBAST (Wolff 2019) together with its PROBAST+AI extension (Collins 2024) [12]. Diagnostic accuracy reporting followed STARD 2015 (Bossuyt 2015) and its STARD-AI extension (Sounderajah 2025) [13]. -
prisma_body.md 381 B
## PRISMA flow A total of 1000 records identified through database searching. After 200 duplicates removed, 800 records screened. Of these, 600 records excluded at screening, leaving 200 reports sought for retrieval. 10 reports not retrieved. 190 reports retrieved and 190 reports assessed for eligibility. 40 records excluded with reasons. 150 studies included in the synthesis. -
prisma_fig_clean.md 273 B
1000 records identified 200 duplicates removed 800 records screened 600 records excluded at screening 200 reports sought for retrieval 10 reports not retrieved 190 reports retrieved 190 reports assessed for eligibility 40 records excluded with reasons 150 studies included -
prisma_fig_mismatch.md 273 B
1000 records identified 200 duplicates removed 800 records screened 600 records excluded at screening 200 reports sought for retrieval 10 reports not retrieved 190 reports retrieved 190 reports assessed for eligibility 40 records excluded with reasons 149 studies included
-
-
test_checklist_fail_fast.sh 3.9 KB
#!/usr/bin/env bash # Regression tests for check-reporting check_checklist_exists.py (fail-fast guard). set -uo pipefail REPO_ROOT="$(cd "$(dirname "$0")/../../.." && pwd)" SCRIPT="$REPO_ROOT/skills/check-reporting/scripts/check_checklist_exists.py" [[ -f "$SCRIPT" ]] || { echo "ENV-ERR: script missing" >&2; exit 2; } fail=0 ran=0 assert_exit() { local label="$1" expected="$2" actual="$3" ran=$((ran + 1)) if [[ "$expected" == "$actual" ]]; then printf ' PASS %-52s exit=%s\n' "$label" "$actual" else printf ' FAIL %-52s expected=%s actual=%s\n' "$label" "$expected" "$actual" fail=$((fail + 1)) fi } assert_contains() { local label="$1" needle="$2" haystack="$3" ran=$((ran + 1)) if [[ "$haystack" == *"$needle"* ]]; then printf ' PASS %-52s\n' "$label" else printf ' FAIL %-52s (missing: %s)\n' "$label" "$needle" fail=$((fail + 1)) fi } run() { python3 "$SCRIPT" "$@" >/dev/null 2>&1; echo $?; } # 1. Vendored checklists that exist on disk -> exit 0. assert_exit "STROBE present" 0 "$(run --guideline STROBE)" assert_exit "STARD-AI alias present" 0 "$(run --guideline STARD-AI)" assert_exit "TRIPOD+AI alias present" 0 "$(run --guideline 'TRIPOD+AI')" assert_exit "TRIPOD-LLM alias present" 0 "$(run --guideline 'TRIPOD-LLM')" assert_exit "AMSTAR 2 alias present" 0 "$(run --guideline 'AMSTAR 2')" assert_exit "RoB 2 alias present" 0 "$(run --guideline 'RoB 2')" assert_exit "QUADAS-C alias present" 0 "$(run --guideline QUADAS-C)" # 2. Now-vendored guidelines -> exit 0 (CONSORT 2025 / SPIRIT 2025 / CARE / CLAIM # 2024 were vendored in PR #43; they must resolve to real files). assert_exit "CONSORT now vendored" 0 "$(run --guideline 'CONSORT 2025')" assert_exit "CARE now vendored" 0 "$(run --guideline CARE)" assert_exit "SPIRIT now vendored" 0 "$(run --guideline 'SPIRIT 2025')" assert_exit "CLAIM 2024 now vendored" 0 "$(run --guideline 'CLAIM 2024')" assert_exit "DECIDE-AI now vendored" 0 "$(run --guideline 'DECIDE-AI')" assert_exit "REMARK now vendored" 0 "$(run --guideline REMARK)" assert_exit "TARGET now vendored" 0 "$(run --guideline TARGET)" assert_exit "target-trial-emulation alias -> TARGET" 0 "$(run --guideline 'target trial emulation')" # 2b. CONSORT-AI / SPIRIT-AI are now vendored -> exit 0 (routed AND a checklist file exists). # The advertised-but-unvendored contract path is exercised via --simulate-missing-checklist # (sections 4-6 below), since no routed guideline currently ships without a file. assert_exit "CONSORT-AI now vendored" 0 "$(run --guideline 'CONSORT-AI')" assert_exit "SPIRIT-AI now vendored" 0 "$(run --guideline 'SPIRIT-AI')" # 3. Unrecognised guideline -> exit 2. assert_exit "unknown guideline" 2 "$(run --guideline NOT-A-REAL-GUIDELINE)" # 4. Explicit opt-in downgrades missing/unknown to exit 0 but warns (never silent). assert_exit "opt-in missing -> 0" 0 "$(run --simulate-missing-checklist --allow-from-memory)" assert_exit "opt-in unknown -> 0" 0 "$(run --guideline NOPE --allow-from-memory)" optin_out="$(python3 "$SCRIPT" --simulate-missing-checklist --allow-from-memory 2>&1)" assert_contains "opt-in emits NON-AUTHORITATIVE warning" "NON-AUTHORITATIVE" "$optin_out" # 5. Contract-test simulation (codex Improvement B "Prove" step). sim_out="$(python3 "$SCRIPT" --simulate-missing-checklist 2>&1)" assert_exit "simulate-missing-checklist" 1 "$(run --simulate-missing-checklist)" assert_contains "simulate emits standard violation code" "MISSING_CHECKLIST_CONTRACT_VIOLATION" "$sim_out" # 6. Violation message carries the standardized machine-greppable code (via the simulated # missing-file path — no routed guideline currently ships without a vendored checklist). viol_out="$(python3 "$SCRIPT" --simulate-missing-checklist 2>&1)" assert_contains "missing emits standard violation code" "MISSING_CHECKLIST_CONTRACT_VIOLATION" "$viol_out" printf '\n%d/%d checks passed\n' "$((ran - fail))" "$ran" [[ "$fail" -eq 0 ]] || exit 1 -
test_checklist_fidelity.sh 6.5 KB
#!/usr/bin/env bash # Self-test for scripts/verify_checklist_fidelity.py — a bundled checklist must match the official # instrument it claims to be. # # Issue #352: the file labelled "TRIPOD+AI 2024" was actually TRIPOD 2015 + separately-numbered `-AI` # additions — wrong section sequence, non-canonical identifiers, no Open Science, no PPI. Nothing # caught it: check_checklist_exists only checks the file is present, check_framework_naming only # checks it names its base. # # The same fidelity audit (2026-07-21) found two more of the same class: # - CLEAR: invented a 7-topical-domain taxonomy (item 1 = "Study hypothesis"); official CLEAR is # numbered by manuscript section (item 1 = Title, item 44 = baseline demographics), non-essential # items are 53 and 58 (the file said 17 and 57). # - MI-CLEAR-LLM: labelled "Version 2025" but carried the 2024 SIX-item body; official 2025 has EIGHT # item categories (Access mode, Input data type, Adaptation strategy promoted to first-class items). # # This gate regresses all three defects and demands they fail, and holds the gate silent on the # corrected files. set -u REPO_ROOT="$(cd "$(dirname "$0")/../../.." && pwd)" G="$REPO_ROOT/skills/check-reporting/scripts/verify_checklist_fidelity.py" CK_DIR="skills/check-reporting/references/checklists" pass=0; fail=0 ck() { if [ "$2" = "$3" ]; then printf ' PASS %-56s exit=%s\n' "$1" "$3"; pass=$((pass+1)); else printf ' FAIL %-56s want=%s got=%s\n' "$1" "$2" "$3"; fail=$((fail+1)); fi; } # Seed a fixture tree with the LIVE (corrected) copies of every checklist the gate knows about, so a # regression can then overwrite exactly one file with a defect and isolate the failure to it. seed() { local dir="$1" mkdir -p "$dir/$CK_DIR" cp "$REPO_ROOT/$CK_DIR/TRIPOD_AI.md" "$dir/$CK_DIR/" cp "$REPO_ROOT/$CK_DIR/CLEAR.md" "$dir/$CK_DIR/" cp "$REPO_ROOT/$CK_DIR/MI_CLEAR_LLM.md" "$dir/$CK_DIR/" cp "$REPO_ROOT/$CK_DIR/GATHER.md" "$dir/$CK_DIR/" } # 1) the live repo's corrected files pass python3 "$G" --strict >/dev/null 2>&1 ck "live repo: all bundled checklists match their official inventory" 0 "$?" FIX="$(mktemp -d)"; trap 'rm -rf "$FIX"' EXIT # 2) REGRESSION — the #352 TRIPOD+AI defect (TRIPOD 2015 + -AI items, no Open Science / PPI, no DOI) seed "$FIX" cat > "$FIX/$CK_DIR/TRIPOD_AI.md" <<'MD' # TRIPOD+AI Checklist Version: TRIPOD+AI 2024 ## Checklist Items Items marked with **(AI)** are specific to the AI extension. All other items are from TRIPOD 2015. ### Title and Abstract | # | Item | Description | |---|------|-------------| | 1 | Title | Identify the study. | | 1-AI | Title (AI) | Identify AI/ML methods. | | 2 | Abstract | Provide a summary. | ### Methods | # | Item | Description | |---|------|-------------| | 10-AI-a | Model architecture (AI) | Describe the architecture. | ### Discussion | # | Item | Description | |---|------|-------------| | 18 | Limitations | Discuss limitations. | | 20 | Implications | Discuss clinical use. | ## MedSci supplemental MD python3 "$G" --root "$FIX" --strict >/dev/null 2>&1 ck "REGRESSION #352: TRIPOD 2015 + -AI (no 18/19) fails" 1 "$?" OUT="$(python3 "$G" --root "$FIX" 2>&1)" echo "$OUT" | grep -q "Open science" && ck " names the missing Open science section" 0 0 || ck " names the missing Open science section" 0 1 echo "$OUT" | grep -q "TRIPOD_AI.md.*forbidden" && ck " names the non-canonical -AI identifiers" 0 0 || ck " names the non-canonical -AI identifiers" 0 1 echo "$OUT" | grep -q "10.1136/bmj-2023-078378" && ck " names the missing source DOI marker" 0 0 || ck " names the missing source DOI marker" 0 1 # 3) REGRESSION — CLEAR regrouped into topical "domains" (item 1 = Study hypothesis), no manuscript # sections, no radiomic-title marker. The real bundled defect found on 2026-07-21. seed "$FIX" cat > "$FIX/$CK_DIR/CLEAR.md" <<'MD' # CLEAR Checklist Version: CLEAR 2023 Source: https://doi.org/10.1186/s13244-023-01415-8 ## Checklist Items (58 items) ### Domain 1: Study Design (Items 1-8) | # | Item | Description | |---|------|-------------| | 1 | Study hypothesis | State the study hypothesis. | | 2 | Study design | Describe the study design. | ### Domain 5: Modeling (Items 35-44) | # | Item | Description | |---|------|-------------| | 44 | Temporal validation | Report temporal validation. | ## Notes for assessors Items 17 and 57 are aspirational. MD python3 "$G" --root "$FIX" --strict >/dev/null 2>&1 ck "REGRESSION CLEAR: topical-domain regrouping fails" 1 "$?" OUT="$(python3 "$G" --root "$FIX" 2>&1)" echo "$OUT" | grep -q "CLEAR.md.*forbidden" && ck " names the invented Domain grouping" 0 0 || ck " names the invented Domain grouping" 0 1 echo "$OUT" | grep -q "### Title" && ck " names the missing Title section" 0 0 || ck " names the missing Title section" 0 1 echo "$OUT" | grep -q "specifying the radiomic" && ck " names the missing item-1 (Title) marker" 0 0 || ck " names the missing item-1 (Title) marker" 0 1 # 4) REGRESSION — MI-CLEAR-LLM labelled 2025 but carrying the 2024 six-item body (no numbered item # table, and none of the three items promoted to first-class in 2025). seed "$FIX" cat > "$FIX/$CK_DIR/MI_CLEAR_LLM.md" <<'MD' # MI-CLEAR-LLM Checklist **Version:** 2025 (expanded from 2024 original) **Source:** https://kjronline.org/DOIx.php?id=10.3348/kjr.2025.1522 ## Checklist Items ## Item 1 — LLM Identification and Specifications ## Item 2 — Stochasticity Handling ## Item 3 — Full Prompt Text ## Item 4 — Prompt Execution Details ## Item 5 — Prompt Testing and Optimization ## Item 6 — Test Data Independence ## Notes for assessors | Category | Items | |----------|-------| | **TOTAL** | **6** | MD python3 "$G" --root "$FIX" --strict >/dev/null 2>&1 ck "REGRESSION MI-CLEAR-LLM: 2024 body under a 2025 label fails" 1 "$?" OUT="$(python3 "$G" --root "$FIX" 2>&1)" echo "$OUT" | grep -q "expected 1..8" && ck " names the missing 8-item inventory" 0 0 || ck " names the missing 8-item inventory" 0 1 echo "$OUT" | grep -q "### 2. Access mode" && ck " names the missing 2025 Access-mode item" 0 0 || ck " names the missing 2025 Access-mode item" 0 1 # 5) NEGATIVE — a fully corrected fixture tree passes (proves the gate is not always-fail) seed "$FIX" python3 "$G" --root "$FIX" --strict >/dev/null 2>&1 ck "NEGATIVE: a corrected fixture tree passes" 0 "$?" echo echo " $pass passed, $fail failed" [ "$fail" -eq 0 ] || exit 1 -
test_checklist_version.sh 3.4 KB
#!/usr/bin/env bash # Test scripts/check_checklist_version.py — the A4b checklist-version staleness gate. # Synthetic, PII-free fixtures: a checklist targeting an older version / a changed # hash / a different file / no version metadata, vs an in-sync checklist. set -u HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" SCRIPT="$HERE/../scripts/check_checklist_version.py" PASS=0 FAIL=0 ok() { echo " PASS: $1"; PASS=$((PASS+1)); } bad() { echo " FAIL: $1"; FAIL=$((FAIL+1)); } WORK="$(mktemp -d)" trap 'rm -rf "$WORK"' EXIT # current manuscript (v8) + its sha256 (first 12) cat > "$WORK/manuscript_v8.md" <<'EOF' ## Title A cohort study, version 8. EOF SHA=$(python3 -c "import hashlib;print(hashlib.sha256(open('$WORK/manuscript_v8.md','rb').read()).hexdigest()[:12])") mk_json() { # file target_manuscript target_version source_sha256 cat > "$1" <<EOF {"check_reporting_version":"1.1","manuscript_title":"A cohort study", "target_manuscript":"$2","target_version":"$3","source_sha256":"$4", "guideline":"STROBE","total_items":22,"present":18} EOF } run() { python3 "$SCRIPT" "$@" 2>/dev/null; } # 1. in-sync checklist (same file/version/hash) -> exit 0 mk_json "$WORK/ck_ok.json" "manuscript_v8.md" "v8" "$SHA" run --checklist "$WORK/ck_ok.json" --manuscript "$WORK/manuscript_v8.md" --quiet [ $? -eq 0 ] && ok "in-sync checklist passes" || bad "in-sync should pass" # 2. older target_version -> stale (exit 1) mk_json "$WORK/ck_oldver.json" "manuscript_v6.md" "v6" "" run --checklist "$WORK/ck_oldver.json" --manuscript "$WORK/manuscript_v8.md" --out "$WORK/r.json" --quiet [ $? -eq 1 ] && ok "older target_version -> exit 1" || bad "older version should fail" python3 -c "import json,sys;d=json.load(open('$WORK/r.json'));sys.exit(0 if d['findings'][0]['type']=='checklist_version_stale' else 1)" \ && ok "reports checklist_version_stale" || bad "wrong finding type for version" # 3. changed content hash (same version) -> stale mk_json "$WORK/ck_hash.json" "manuscript_v8.md" "v8" "deadbeef0000" run --checklist "$WORK/ck_hash.json" --manuscript "$WORK/manuscript_v8.md" --out "$WORK/r2.json" --quiet python3 -c "import json,sys;d=json.load(open('$WORK/r2.json'));sys.exit(0 if d['findings'][0]['type']=='checklist_content_stale' else 1)" \ && ok "changed hash -> checklist_content_stale" || bad "hash drift not detected" # 4. no version metadata (pre-v1.1) -> unverifiable (exit 1) echo '{"check_reporting_version":"1.0","guideline":"STROBE","present":18}' > "$WORK/ck_old.json" run --checklist "$WORK/ck_old.json" --manuscript "$WORK/manuscript_v8.md" --out "$WORK/r3.json" --quiet [ $? -eq 1 ] && ok "no version metadata -> exit 1" || bad "missing metadata should fail" python3 -c "import json,sys;d=json.load(open('$WORK/r3.json'));sys.exit(0 if d['findings'][0]['type']=='checklist_no_version_metadata' else 1)" \ && ok "reports checklist_no_version_metadata" || bad "wrong finding for missing metadata" # 5. markdown report with header fields (no JSON) -> parsed + stale cat > "$WORK/report_v6.md" <<'EOF' ## Reporting Guideline Compliance Report Manuscript: A cohort study Target manuscript file: manuscript_v6.md Target version: v6 Guideline: STROBE 2007 EOF run --checklist "$WORK/report_v6.md" --manuscript "$WORK/manuscript_v8.md" --quiet [ $? -eq 1 ] && ok "markdown header (v6) flagged stale vs v8" || bad "markdown header not parsed" echo "" echo "test_checklist_version: $PASS passed, $FAIL failed" [ "$FAIL" -eq 0 ] -
test_framework_naming.sh 2.7 KB
#!/usr/bin/env bash # Regression test for the framework-naming audit (check-reporting Step 4e). # Synthetic, PII-free fixtures reproduce: (a) an AI extension used without naming # or citing its base instrument (BASE_MISSING), (b) mixed +AI / -AI hyphenation for # one family (HYPHEN_MIX), (c) vague "adapted per recent guidance" wording # (VAGUE_GUIDANCE), (d) a self-coined "12-AI" item label (SELF_COINED_LABEL). The # clean fixture names and cites both base instrument and extension. # Stdlib-only (python3). set -u HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" SCRIPT="$HERE/../scripts/check_framework_naming.py" BAD="$HERE/fixtures/framework_bad.md" CLEAN="$HERE/fixtures/framework_clean.md" OUT="$(mktemp -t fw_XXXX).json" trap 'rm -f "$OUT"' EXIT fail=0 check() { local label="$1"; shift if "$@" >/dev/null 2>&1; then printf ' PASS %s\n' "$label" else printf ' FAIL %s\n' "$label"; fail=$((fail+1)); fi } has_verdict() { python3 -c " import json,sys d=json.load(open('$OUT')) assert any(c['verdict']=='$1' for c in d['claims']), '$1 not found' "; } [[ -f "$SCRIPT" ]] || { echo "ENV-ERR: script missing" >&2; exit 2; } # (1) bad manuscript: BASE_MISSING is Major -> exit 1 under --strict python3 "$SCRIPT" --manuscript "$BAD" --out "$OUT" --strict --quiet >/dev/null 2>&1 check "exit 1 under --strict (Major present)" test "$?" -eq 1 check "JSON artifact written" test -s "$OUT" check "BASE_MISSING detected (extension without base)" has_verdict BASE_MISSING check "HYPHEN_MIX detected (PROBAST+AI vs PROBAST-AI)" has_verdict HYPHEN_MIX check "VAGUE_GUIDANCE detected (adapted per recent guidance)" has_verdict VAGUE_GUIDANCE check "SELF_COINED_LABEL detected (12-AI)" has_verdict SELF_COINED_LABEL # (2) clean manuscript: base named + cited, consistent hyphenation -> exit 0 python3 "$SCRIPT" --manuscript "$CLEAN" --strict --quiet >/dev/null 2>&1 check "exit 0 on clean manuscript (base named + cited)" test "$?" -eq 0 # (3) FP guard (less-defensive, F05): method-level "recent best-practice recommendations" # with NO reporting context must NOT trigger VAGUE_GUIDANCE. NOFP="$(mktemp -t fw_nofp_XXXX).md" trap 'rm -f "$OUT" "$NOFP"' EXIT cat > "$NOFP" <<'EOF' # Methods We performed external validation following recent best-practice recommendations. The imputation strategy was aligned with current guidance for handling missingness. EOF python3 "$SCRIPT" --manuscript "$NOFP" --out "$OUT" --quiet >/dev/null 2>&1 check "no VAGUE_GUIDANCE on method-level recommendations (no reporting cue)" python3 -c " import json d=json.load(open('$OUT')) raise SystemExit(0 if not any(c['verdict']=='VAGUE_GUIDANCE' for c in d['claims']) else 1) " echo "fail=$fail"; [[ "$fail" -eq 0 ]] && echo "ALL PASS" || echo "FAILURES: $fail" exit "$fail" -
test_prisma_cascade.sh 3.8 KB
#!/usr/bin/env bash # Regression tests for check-reporting prisma_cascade_check.py. set -uo pipefail REPO_ROOT="$(cd "$(dirname "$0")/../../.." && pwd)" SCRIPT="$REPO_ROOT/skills/check-reporting/scripts/prisma_cascade_check.py" TMP="$(mktemp -d -t prisma_cas.XXXXXX)" trap 'rm -rf "$TMP"' EXIT [[ -f "$SCRIPT" ]] || { echo "ENV-ERR: script missing" >&2; exit 2; } fail=0 ran=0 assert_exit() { local label="$1" expected="$2" actual="$3" ran=$((ran + 1)) if [[ "$expected" == "$actual" ]]; then printf ' PASS %-50s exit=%s\n' "$label" "$actual" else printf ' FAIL %-50s expected=%s actual=%s\n' "$label" "$expected" "$actual" fail=$((fail + 1)) fi } build_tsv() { local out="$1" col="$2" shift 2 { printf 'uid\t%s\n' "$col" local i=1 for label in "$@"; do printf 'UID_%03d\t%s\n' "$i" "$label" i=$((i + 1)) done } > "$out" } # -------------------------------------------------------------------------- # Case 1: matching counts (no manuscript) => PASS. # round 1: 5 INCLUDE, 5 EXCLUDE # round 2: 3 INCLUDE, 2 EXCLUDE # round 3: 2 INCLUDE, 1 EXCLUDE # -------------------------------------------------------------------------- build_tsv "$TMP/c1_r1.tsv" decision INCLUDE INCLUDE INCLUDE INCLUDE INCLUDE EXCLUDE EXCLUDE EXCLUDE EXCLUDE EXCLUDE build_tsv "$TMP/c1_r2.tsv" decision INCLUDE INCLUDE INCLUDE EXCLUDE EXCLUDE build_tsv "$TMP/c1_r3.tsv" round3_decision INCLUDE INCLUDE EXCLUDE python3 "$SCRIPT" --round1 "$TMP/c1_r1.tsv" --round2 "$TMP/c1_r2.tsv" \ --round3 "$TMP/c1_r3.tsv" --out "$TMP/c1.json" --quiet assert_exit "case 1: TSV only, no manuscript (PASS)" 0 $? python3 - "$TMP/c1.json" <<'PY' || fail=$((fail + 1)) import json, sys r = json.load(open(sys.argv[1])) sc = r["stage_counts"] assert sc["round1_total"] == 10, sc assert sc["round1_include"] == 5, sc assert sc["round3_include"] == 2, sc PY # -------------------------------------------------------------------------- # Case 2: manuscript claims correct count => PASS. # -------------------------------------------------------------------------- cat > "$TMP/c2_manuscript.md" <<'EOF' ## **METHODS** After title and abstract screening, 3 records were retrieved for full-text review. Finally included 2 studies in the meta-analysis. EOF python3 "$SCRIPT" --round1 "$TMP/c1_r1.tsv" --round2 "$TMP/c1_r2.tsv" \ --round3 "$TMP/c1_r3.tsv" \ --manuscript "$TMP/c2_manuscript.md" \ --out "$TMP/c2.json" --quiet assert_exit "case 2: manuscript matches (PASS)" 0 $? # -------------------------------------------------------------------------- # Case 3: manuscript prose off-by-one => FAIL. # Computed round3_include = 2; manuscript says "3 studies" # -------------------------------------------------------------------------- cat > "$TMP/c3_manuscript.md" <<'EOF' ## **RESULTS** Finally included 3 studies in the meta-analysis. EOF python3 "$SCRIPT" --round1 "$TMP/c1_r1.tsv" --round2 "$TMP/c1_r2.tsv" \ --round3 "$TMP/c1_r3.tsv" \ --manuscript "$TMP/c3_manuscript.md" \ --out "$TMP/c3.json" --quiet assert_exit "case 3: off-by-one prose drift (FAIL)" 1 $? python3 - "$TMP/c3.json" <<'PY' || fail=$((fail + 1)) import json, sys r = json.load(open(sys.argv[1])) drifts = r["manuscript_drift"] assert any(d["stage"] == "round3_include" and d["manuscript"] == 3 for d in drifts), drifts PY # -------------------------------------------------------------------------- # Case 4: bad column => exit 2. # -------------------------------------------------------------------------- build_tsv "$TMP/c4.tsv" notes INCLUDE EXCLUDE python3 "$SCRIPT" --round1 "$TMP/c4.tsv" --round2 "$TMP/c1_r2.tsv" \ --round3 "$TMP/c1_r3.tsv" --out "$TMP/c4.json" --quiet 2>/dev/null assert_exit "case 4: bad column (exit 2)" 2 $? echo "" echo "ran=$ran fail=$fail" [[ $fail -eq 0 ]] -
test_prisma_figure.sh 2.1 KB
#!/usr/bin/env bash # Regression test for the PRISMA Figure 1 arithmetic + cross-reference audit # (check-reporting Step 4d / check_prisma_figure.py). Synthetic, PII-free fixtures. # Stdlib-only (python3); no network, no pandoc. set -u HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" SCRIPT="$HERE/../scripts/check_prisma_figure.py" BODY="$HERE/fixtures/prisma_body.md" CLEAN="$HERE/fixtures/prisma_fig_clean.md" MM="$HERE/fixtures/prisma_fig_mismatch.md" OUT="$(mktemp -t prisma_fig_XXXX).json" trap 'rm -f "$OUT"' EXIT for f in "$SCRIPT" "$BODY" "$CLEAN" "$MM"; do [[ -f "$f" ]] || { echo "ENV-ERR: missing $f" >&2; exit 2; } done fail=0 pass() { printf ' PASS %s\n' "$1"; } bad() { printf ' FAIL %s\n' "$1"; fail=$((fail+1)); } echo "test_prisma_figure:" # 1. Clean figure (numbers match body, arithmetic consistent) -> audit_safe, exit 0. python3 "$SCRIPT" --md "$BODY" --figure "$CLEAN" --out "$OUT" >/dev/null 2>&1; rc=$? if [[ $rc -eq 0 ]] && python3 -c "import json,sys; sys.exit(0 if json.load(open('$OUT'))['audit_safe'] else 1)"; then pass "clean body/figure -> audit_safe, exit 0" else bad "clean case rc=$rc (expected 0 + audit_safe)" fi # 2. Mismatched figure (included 149 vs body 150) -> MISMATCH, exit 1, PRISMA-FIGURE flag. out="$(python3 "$SCRIPT" --md "$BODY" --figure "$MM" --out "$OUT" 2>&1)"; rc=$? if [[ $rc -eq 1 && "$out" == *"[PRISMA-FIGURE]"* ]] \ && python3 -c "import json,sys; d=json.load(open('$OUT')); sys.exit(0 if (not d['audit_safe'] and d['action_items']) else 1)"; then pass "mismatched figure -> MISMATCH flagged, exit 1" else bad "mismatch case rc=$rc (expected 1 + [PRISMA-FIGURE] + action_items)" fi # 3. Missing input -> clean error, exit 2 (no traceback). err="$(python3 "$SCRIPT" --md /nonexistent_prisma.md --figure "$CLEAN" --out "$OUT" 2>&1)"; rc=$? if [[ $rc -eq 2 && "$err" == *"not found"* && "$err" != *"Traceback"* ]]; then pass "missing manuscript -> clean error, exit 2" else bad "missing-input case rc=$rc: $err" fi if [[ $fail -eq 0 ]]; then echo " OK"; exit 0; else echo " $fail check(s) failed"; exit 1; fi
-
-
SKILL.md 41.9 KB
--- name: check-reporting description: Check manuscript compliance with medical research reporting guidelines. Supports 49 guidelines including STROBE, STROBE-MR, RECORD, REMARK (prognostic tumor-marker studies), TARGET (target trial emulation), GATHER (burden-of-disease / health-estimate modeling), CONSORT, CONSORT-AI, STARD, STARD-AI, TRIPOD, TRIPOD+AI, TRIPOD-LLM, PGS-RS, ARRIVE, PRISMA, PRISMA 2020 for Abstracts, PRISMA-DTA, PRISMA-P, PRISMA-ScR (scoping reviews), CARE, SPIRIT, SPIRIT-AI, CLAIM, DECIDE-AI, MI-CLEAR-LLM, SQUIRE 2.0, CLEAR, MOOSE, GRRAS, SWiM, AMSTAR 2, CHEERS 2022, CROSS (survey studies), SRQR and COREQ (qualitative research), and risk of bias tools (QUADAS-3, QUADAS-2, QUADAS-C, RoB 2, ROBINS-I, ROBINS-E, ROBIS, ROB-ME, PROBAST, PROBAST+AI, NOS, COSMIN, RoB NMA). Generates item-by-item assessment with PRESENT/MISSING/PARTIAL status. triggers: checklist, QUADAS-3, abstract checklist, structured abstract, reporting guideline, STROBE, STROBE-MR, Mendelian randomization, CONSORT, CONSORT-AI, STARD, STARD-AI, TRIPOD, TRIPOD-LLM, PGS-RS, PRS-RS, polygenic risk score, polygenic score, PRISMA, PRISMA-DTA, PRISMA-P, PRISMA-ScR, scoping review, scoping, evidence map, ARRIVE, CARE, CLAIM, DECIDE-AI, MI-CLEAR-LLM, SPIRIT, SPIRIT-AI, QUADAS, QUADAS-C, RoB, ROBINS, ROBINS-E, ROBIS, ROB-ME, PROBAST, NOS, COSMIN, AMSTAR, SWiM, CHEERS, economic evaluation, cost-effectiveness, cost-utility, QALY, ICER, RECORD, RECORD-PE, routinely-collected data, registry, claims, electronic health records, EHR, real-world data, CROSS, CHERRIES, survey, questionnaire, KAP, e-survey, response rate, SRQR, COREQ, qualitative research, interviews, focus groups, thematic analysis, grounded theory, reflexivity, REMARK, tumor marker, prognostic marker, prognostic biomarker, molecular residual disease, TARGET, target trial emulation, target trial, causal inference, estimand, immortal time bias, GATHER, burden of disease, global burden, GBD, health estimates, attributable burden, comparative risk assessment, population attributable fraction, disability-adjusted life years, DALY, forecasting, decomposition, risk of bias, compliance check, LLM accuracy, large language model, clinical deployment tools: Read, Write, Edit, Bash, Grep, Glob model: inherit --- # Check-Reporting Skill You are helping a medical researcher verify that their manuscript complies with the appropriate medical research reporting guideline. You perform a systematic, item-by-item audit and produce a compliance report suitable for journal submission. ## Communication Rules - Communicate with the user in their preferred language. - Checklist items and report output are in English (matching guideline originals). - Medical terminology is always in English. ## Reference Files - **Checklists (bundled, open license)**: `${CLAUDE_SKILL_DIR}/references/checklists/` - `STROBE.md` -- observational studies (CC BY) - `STROBE_MR.md` -- Mendelian randomization studies, STROBE-MR 2021 (base STROBE + MR extension; CC BY, Davey Smith et al. BMJ 2021) - `STARD.md` -- diagnostic accuracy studies (CC BY 4.0) - `STARD_AI.md` -- AI diagnostic accuracy studies (CC BY, Sounderajah et al. Nat Med 2025) - `TRIPOD.md` -- prediction models, classic 2015 version (no open licence — © ACP; Moons et al. Ann Intern Med 2015) - `TRIPOD_AI.md` -- prediction models with AI/ML (CC BY 4.0, Collins et al. BMJ 2024) - `TRIPOD_LLM.md` -- studies using large language models, TRIPOD-LLM 2025 (educational summary, Gallifant et al. Nat Med 2025) - `PGS_RS.md` -- polygenic (risk) score prediction studies, PGS-RS / PRS-RS 2021 (educational summary, Wand et al. Nature 2021) - `CHEERS_2022.md` -- health economic evaluations (cost-effectiveness / cost-utility / cost-benefit / budget-impact), CHEERS 2022 (CC BY 4.0, Husereau et al. BMJ 2022) - `RECORD.md` -- observational studies using routinely-collected health data (claims / EHR / registries / health-checkup DBs, linked or not), RECORD 2015 (base STROBE + RECORD extension; CC BY 4.0, Benchimol et al. PLoS Med 2015; RECORD-PE for drug studies) - `CROSS.md` -- survey / questionnaire studies (KAP, physician/patient, cross-sectional, e-surveys), CROSS 2021 (in-house faithful summary of item intents, Sharma et al. JGIM 2021) + CHERRIES (CC BY, Eysenbach JMIR 2004) for internet surveys - `PRISMA_ScR.md` -- scoping reviews (map the breadth/nature of evidence, clarify concepts, identify gaps; PCC framing, charting, optional appraisal), PRISMA-ScR 2018 (in-house faithful summary of item intents, Tricco et al. Ann Intern Med 2018; DOI 10.7326/M18-0850) - `SRQR.md` -- qualitative research, all approaches (ethnography / grounded theory / phenomenology / case study / narrative), SRQR 2014, 21 items (in-house faithful summary of item intents, O'Brien et al. Acad Med 2014; DOI 10.1097/ACM.0000000000000388) - `COREQ.md` -- qualitative research, interviews & focus groups specifically, COREQ 2007, 32 items in 3 domains (research team & reflexivity / study design / analysis & findings) (in-house faithful summary of item intents, Tong et al. Int J Qual Health Care 2007; DOI 10.1093/intqhc/mzm042) - `REMARK.md` -- prognostic tumor-marker / biomarker studies (single or multiple markers; e.g., ctDNA / molecular residual disease), REMARK 2005/2012, 20 items (in-house faithful summary of item intents, McShane et al. Br J Cancer 2005 + Altman et al. PLoS Med 2012) - `TARGET.md` -- observational studies emulating a target trial (causal / comparative-effectiveness questions on routinely-collected / registry / EHR data), TARGET 2025, 21 items (in-house faithful summary of item intents, Cashin/Hansford/Hernán et al. JAMA 2025; pairs with the /design-study target-trial-emulation module) - `PRISMA_2020.md` -- systematic reviews (CC BY) - `PRISMA_2020_Abstracts.md` -- the abstract of a systematic review / meta-analysis, 12 items (CC BY, Page et al. BMJ 2021). A separate instrument from the 27-item checklist, not a subset: item 2 of the main checklist defers to it. Score it with its own denominator. - `ARRIVE_2.md` -- animal studies (CC0) - `PRISMA_DTA.md` -- DTA systematic reviews (no open licence — © AMA; McInnes et al. JAMA 2018) - `QUADAS3.md` -- diagnostic accuracy risk of bias, **current recommended version** (no open licence -- (c) ACP; Whiting et al. Ann Intern Med 2026) - `QUADAS2.md` -- diagnostic accuracy risk of bias (no open licence — © ACP; Whiting et al. Ann Intern Med 2011) - `RoB2.md` -- RCT risk of bias (CC BY, Sterne et al. BMJ 2019) - `ROBINS_I.md` -- non-randomised studies risk of bias (CC BY-**NC** 3.0 — non-commercial; Sterne et al. BMJ 2016) - `PROBAST.md` -- prediction model risk of bias (no open licence — © ACP; Wolff et al. Ann Intern Med 2019) - `NOS.md` -- observational study quality (public domain, Ottawa Hospital) - `CONSORT.md` -- randomised controlled trials, CONSORT 2025 (CC BY 4.0, Hopewell et al. BMJ 2025) - `CONSORT_AI.md` -- AI clinical-trial reports, CONSORT-AI 2020 (CC BY 4.0, Liu et al. Nat Med 2020) - `CARE.md` -- case reports, CARE 2013 (no confirmed open licence — Elsevier TDM only; Gagnier et al. J Clin Epidemiol 2014) - `SPIRIT.md` -- clinical trial protocols, SPIRIT 2025 (CC BY 4.0, Chan et al. BMJ 2025) - `SPIRIT_AI.md` -- AI clinical-trial protocols, SPIRIT-AI 2020 (CC BY 4.0, Cruz Rivera et al. Nat Med 2020) - `CLAIM_2024.md` -- AI/ML in clinical imaging, CLAIM 2024 Update (RSNA open access, Tejani et al. Radiol Artif Intell 2024) - `DECIDE_AI.md` -- early-stage clinical evaluation of AI decision-support systems, DECIDE-AI 2022 (educational summary, CC BY-NC, Vasey et al. Nat Med 2022) - `MI_CLEAR_LLM.md` -- LLM accuracy studies in healthcare (CC BY-NC 4.0, Park et al. KJR 2024; 2025 update) - `SQUIRE_2.md` -- quality improvement in healthcare/education (no open licence — Crossref returns none; Ogrinc et al. BMJ Qual Saf 2016) - `CLEAR.md` -- radiomics studies (CC BY 4.0, Kocak et al. Insights Imaging 2023) - `MOOSE.md` -- meta-analysis of observational studies (Stroup et al. JAMA 2000) - `GRRAS.md` -- reliability and agreement studies (Kottner et al. J Clin Epidemiol 2011) - `QUADAS_C.md` -- comparative DTA risk of bias, extension to QUADAS-2 (no open licence — © ACP; Yang et al. Ann Intern Med 2021) - `ROBINS_E.md` -- non-randomised exposure studies risk of bias (CC BY-NC-ND 4.0, Higgins et al. Environ Int 2024) - `ROBIS.md` -- risk of bias in systematic reviews (Whiting et al. J Clin Epidemiol 2016) - `ROB_ME.md` -- risk of bias due to missing evidence in meta-analysis (no open licence — BMJ TDM policy only; Page et al. BMJ 2023) - `PROBAST_AI.md` -- prediction model risk of bias, updated for AI/ML (Moons et al. BMJ 2025) - `COSMIN_RoB.md` -- reliability/measurement error risk of bias (Mokkink et al. BMC Med Res Methodol 2020) - `RoB_NMA.md` -- risk of bias in network meta-analysis (Lunny et al. 2024) - `AMSTAR2.md` -- quality of systematic reviews (Shea et al. BMJ 2017) - `PRISMA_P.md` -- systematic review protocols (Shamseer et al. BMJ 2015) - `SWiM.md` -- synthesis without meta-analysis reporting (Campbell et al. BMJ 2020) - `GATHER.md` -- health-estimate / burden-of-disease modeling studies (GBD and GBD-satellite, comparative-risk / population-attributable-fraction, cause-of-death and prevalence/incidence estimation, with or without forecasts), GATHER 2016 (in-house faithful summary; CC BY, Stevens et al. Lancet 2016;388:e19-23 / PLoS Med 2016;13(6):e1002056). Pairs with `/analyze-stats` `references/analysis_guides/burden_decomposition_forecasting.md` for the analytic methods. - Fail-fast contract: if a routed guideline has no vendored checklist file, the skill does **not** silently construct items from memory. It halts with a `MISSING_CHECKLIST_CONTRACT_VIOLATION` and surfaces the gap. A from-memory assessment is allowed only with the explicit `--allow-from-memory` opt-in, and that report must be clearly labelled NON-AUTHORITATIVE. See Step 2 and `scripts/check_checklist_exists.py`. - **Critical-item floor**: `${CLAUDE_SKILL_DIR}/references/critical_item_floor.md` -- the small set of non-waivable items per study type (presence outranks the headline %), plus the AI/radiomics methodological-quality / risk-of-bias instruments (PROBAST+AI, METRICS/RQS, APPRAISE-AI) kept distinct from their reporting counterparts. Loaded in Step 4f. --- ## Workflow ### Step 0: Existing-checklist staleness pre-check If a checklist already exists for this project (`qc/reporting_checklist.json` or a prior `.md` report), verify it targets the **current** manuscript before reusing it — a checklist generated against an older version carries stale section/line references and a stale version label that a reviewer who cross-checks will catch: ```bash python3 "${CLAUDE_SKILL_DIR}/scripts/check_checklist_version.py" \ --checklist qc/reporting_checklist.json --manuscript manuscript_v8.md ``` A non-zero exit means the existing checklist is stale (older `target_version`, changed `source_sha256`, different `target_manuscript`) or pre-dates the version contract — regenerate it against the current manuscript (Steps 1–5) rather than reusing it. Every report you generate must carry the `target_manuscript` / `target_version` / `source_sha256` fields (Part A header + Part D JSON) so this check works next round. ### Step 1: Select Guideline Determine the appropriate reporting guideline. Auto-detect from the manuscript type or accept user specification. **Auto-detection mapping:** | Study Type | Primary Guideline | AI Extension | |------------|------------------|--------------| | Observational study | STROBE | -- | | Mendelian randomization study | STROBE-MR (base STROBE + MR extension) | -- | | Health economic evaluation (cost-effectiveness / cost-utility / cost-benefit / budget-impact) | CHEERS 2022 | -- | | Observational study using routinely-collected data (claims / EHR / registry / health-checkup DB) | RECORD (base STROBE + RECORD extension; RECORD-PE for drug studies) | -- | | Survey / questionnaire study (KAP, physician/patient, cross-sectional, e-survey) | CROSS (+ CHERRIES for internet surveys) | -- | | Scoping review (maps breadth/nature of evidence, clarifies concepts, identifies gaps — not a focused effectiveness/accuracy question) | PRISMA-ScR (base PRISMA + scoping-review extension) | -- | | Qualitative study (interviews, focus groups, ethnography, grounded theory, phenomenology, document analysis) | SRQR (all qualitative approaches); COREQ (interviews/focus groups specifically) | -- | | Randomized controlled trial | CONSORT 2025 | CONSORT-AI | | Diagnostic accuracy study | STARD 2015 | STARD-AI | | Prediction model (development/validation) | TRIPOD | TRIPOD+AI | | Polygenic (risk) score prediction study | PGS-RS (with TRIPOD / TRIPOD+AI) | -- | | Prognostic tumor-marker / biomarker study (single or multiple markers; e.g., ctDNA / molecular residual disease) | REMARK (pair with STROBE for the observational-design items; TRIPOD / TRIPOD+AI if a prognostic model is developed) | -- | | Causal / comparative-effectiveness question emulated on observational data (treatment vs treatment, screening vs none, drug A vs B on registry / EHR / claims data) | TARGET (pair with the /design-study target-trial-emulation module for design; RECORD / STROBE for the routinely-collected-data items) | -- | | Health-estimate / burden-of-disease modeling study (GBD or GBD-satellite, comparative-risk / population-attributable-fraction, cause-of-death or prevalence/incidence estimation, with or without forecasts) | GATHER (pair with `/analyze-stats` burden-decomposition-forecasting guide for the analytic layer) | -- | | Systematic review / meta-analysis | PRISMA 2020 | PRISMA 2020 for Abstracts (run on the abstract, scored separately) | | DTA systematic review / meta-analysis | PRISMA-DTA | PRISMA 2020 for Abstracts (run on the abstract, scored separately) | | Meta-analysis of observational studies | MOOSE | PRISMA 2020 (use both) | | Risk of bias (DTA studies) | **QUADAS-3** (current recommended version) | QUADAS-2 only when appraising or reproducing a review that used it | | Risk of bias (RCTs) | RoB 2 | -- | | Risk of bias (non-randomised intervention studies) | ROBINS-I | -- | | Risk of bias (non-randomised exposure studies) | ROBINS-E | -- | | Risk of bias (comparative DTA studies) | QUADAS-C | **QUADAS-3** (use both; apply the E&E's adaptation — see *Using QUADAS-C with QUADAS-3* in `QUADAS3.md`) | | Risk of bias (prediction models) | PROBAST | PROBAST+AI | | Risk of bias (systematic reviews) | ROBIS | AMSTAR 2 | | Risk of bias (missing evidence in MA) | ROB-ME | -- | | Risk of bias (network meta-analysis) | RoB NMA | -- | | Risk of bias (measurement properties) | COSMIN RoB | -- | | Quality assessment (observational) | NOS | -- | | Case report | CARE | -- | | Study protocol | SPIRIT 2025 | SPIRIT-AI | | Animal study | ARRIVE 2.0 | -- | | AI/ML study in clinical imaging | CLAIM 2024 | -- | | Study using a large language model (develop/fine-tune/prompt/evaluate an LLM) | TRIPOD-LLM | MI-CLEAR-LLM (use alongside when LLM accuracy is an outcome) | | Early-stage / live clinical evaluation of an AI decision-support system (human factors, workflow, safety) | DECIDE-AI | -- | | LLM accuracy evaluation in healthcare | MI-CLEAR-LLM | STARD-AI or CLAIM 2024 (use alongside) | | Reliability / agreement study | GRRAS | -- | | SR protocol | PRISMA-P | -- | | Synthesis without meta-analysis | SWiM | PRISMA 2020 (use both) | | Quality of systematic reviews | AMSTAR 2 | ROBIS | | Radiomics study | CLEAR | CLAIM 2024 (if deep learning component) | | Educational / QI study | SQUIRE 2.0 | -- | | Generative AI **images ARE the study object** (realism / real-vs-synthetic reader study / model-vs-model quality) | (no single guideline -- assemble) | see decision aid below | > **QUADAS-3 has two protocol-stage phases, and this skill usually runs too late for them.** > Phase 1 (state the synthesis question) and phase 2 (define the **ideal test accuracy trial** > each judgement is made against) are review-level and belong in the protocol, alongside the > review-specific guidance for answering each signalling question. Reaching them for the first > time during manuscript QC means writing the comparator after seeing the results. > If they are missing, say so as a limitation rather than reconstructing them — and route the > protocol work to `/meta-analysis` Phase 1. Phases 3–6 are what a QC pass can genuinely run. **Rules:** - If the study involves AI/ML, always apply the AI extension in addition to the base guideline. - **Exception — TRIPOD**: TRIPOD+AI 2024 (Collins et al., BMJ 2024) is a complete rewrite, not an addendum to TRIPOD 2015 (Moons et al., Ann Intern Med 2015). For non-AI prediction models, use TRIPOD 2015 only. For AI/ML prediction models, use TRIPOD+AI 2024 only. Do NOT apply both simultaneously. - **STARD-AI** (Sounderajah et al., Nat Med 2025) extends STARD 2015 with 14 new and 4 modified items (40 total). For AI diagnostic accuracy studies, use STARD-AI (which incorporates all STARD 2015 items). Do NOT apply both STARD 2015 and STARD-AI simultaneously — STARD-AI supersedes STARD 2015 for AI studies. - **TRIPOD-LLM** (Gallifant et al., Nat Med 2025) is the reporting guideline for studies that develop, fine-tune, prompt, or evaluate a large language model for a clinical/biomedical task. It extends the TRIPOD family (TRIPOD 2015 → TRIPOD+AI 2024 → TRIPOD-LLM 2025); name the base instrument and the extension and cite each. It is modular — task-specific items (Annotation, Prompting, Summarization, Instruction-tuning) are N/A when that component is absent. Use TRIPOD-LLM for LLM studies in place of TRIPOD+AI; pair with MI-CLEAR-LLM when LLM accuracy is an evaluated outcome. The vendored checklist is an educational summary (own-words paraphrase of item intent); complete the official instrument for a submission checklist. - **MI-CLEAR-LLM** is a supplementary checklist (8 item categories in the 2025 update; the 2024 original had 6), not a standalone reporting guideline. Always pair it with the study's primary guideline (e.g., STARD-AI for AI diagnostic accuracy, CLAIM for imaging AI). Apply MI-CLEAR-LLM whenever the study evaluates LLM accuracy as an outcome — do NOT apply it merely because the manuscript was written with LLM assistance. Its scope is **LLM accuracy** studies (including VLMs interpreting images); it does **not** apply at study level to studies where a generative model *produces* the images under study (see next bullet). - **Generative-AI images as the study object** (a generative model synthesizes images and the study evaluates their realism, controllability, real-vs-synthetic distinguishability, or model-vs-model quality) has **no single dominant checklist**. Assemble: CLAIM 2024 (imaging-AI umbrella; model-development items N/A when commercial models are used as-is) + FUTURE-AI traceability + MI-CLEAR-LLM **transparency items only** (prompt/model/version/params/runs — for generation provenance, not study-level compliance) on the generator side; STARD-AI (for real-vs-synthetic detection) + GRRAS (reader reliability) + MRMC reporting on the evaluation side. Map applicable items and cite base + extension; never claim wholesale compliance. Full decision aid: `${CLAUDE_SKILL_DIR}/references/genai_image_study_object_decision_aid.md`. - If multiple guidelines apply (e.g., a diagnostic accuracy study that is also an AI study), check against all relevant guidelines and merge into one report. - If the user requests a specific guideline, use that one regardless of auto-detection. ### Step 2: Load Checklist 1. **Run the fail-fast guard first** for every guideline you intend to apply: ```bash python "${CLAUDE_SKILL_DIR}/scripts/check_checklist_exists.py" --guideline "STARD-AI" ``` - Exit 0 → the vendored checklist exists; read it from `${CLAUDE_SKILL_DIR}/references/checklists/` and proceed. - Exit 1 (`MISSING_CHECKLIST_CONTRACT_VIOLATION`) → the guideline is routed but no checklist file is vendored. **Do not construct items from memory.** Halt, report the violation to the user, and stop unless they explicitly opt in (next bullet). - Exit 2 (`UNKNOWN_GUIDELINE`) → the name is not recognised; confirm the correct guideline with the user. 2. **No silent fallback.** A from-memory checklist is permitted only when the user explicitly accepts it — re-run the guard with `--allow-from-memory` (exit 0 + a NON-AUTHORITATIVE warning). In that case the output report MUST carry a prominent banner that the assessment was constructed from model knowledge and is not backed by a vendored checklist, and `submission_safe` must not be asserted on its basis. ### Step 3: Scan Manuscript Read all sections of the manuscript thoroughly: 1. Title and abstract 2. Introduction 3. Methods (all subsections) 4. Results (all subsections) 5. Discussion 6. Tables, figures, and their captions 7. Supplemental materials (if available) 8. References (for registration numbers, protocol references) Gather context from the full document before starting the item-by-item assessment. ### Step 4: Assess Each Item For every checklist item, determine: | Status | Criteria | |--------|----------| | **PRESENT** | The item is fully addressed with sufficient detail. | | **PARTIAL** | The item is mentioned or partially addressed but lacks required detail. | | **MISSING** | The item is not found anywhere in the manuscript. | | **N/A** | The item does not apply to this particular study (justify why). | For each item, record: - **Status**: PRESENT / PARTIAL / MISSING / N/A - **Location**: Section name and paragraph or approximate position (e.g., "Methods, paragraph 3") - **Notes**: What was found (if PRESENT/PARTIAL) or what should be added (if MISSING) **What is appraised is the source paper's reporting — never your convenience in using it.** This holds for every instrument here, reporting checklists and risk-of-bias / quality tools alike, and it is easiest to lose in a systematic review, where you read each paper *in order to extract from it*. An item asking "are the results clearly reported?" is not asking "were they reported in the unit my pool needs". A scorer working a case-series quality tool marks a paper down on the outcome-reporting item because its analysis unit does not match the pool's — treatment-level results against a patient-level denominator. The correction is one sentence, *that is a limit of our extraction, not a defect in their reporting*, and the score goes back up. Single-scorer appraisal is where this happens, because there is nobody to say it. So: if a downgrade's stated reason turns on a **denominator, an analysis unit, a subgroup you needed and they did not report separately, or a format you could not parse**, it is an extraction note, not a scoring reason. Record it in a separate **extraction-note** column and restore the score. Both belong in the table. An extraction limitation is a real constraint on *your* synthesis and often belongs in your limitations paragraph — it just is not evidence about the paper being appraised, and folding it into the score makes the appraisal unreproducible: another assessor with a different pool would score the same paper differently. ### Step 4b: Section Boundary Check In addition to checklist items, verify that: - **Results section** contains only factual findings: no interpretation, no "why" explanations, no prior literature comparisons, no evaluative adjectives without numbers. - **Discussion section** does not introduce new data not presented in Results. - Flag any boundary violation as a separate finding in Part C Action Items with the label `[BOUNDARY]`. ### Step 4c: Registration / Protocol Timing Consistency Check **Applies to:** systematic reviews, meta-analyses, and intervention studies with prospective registration (PRISMA 2020, PRISMA-DTA, PRISMA-P, MOOSE, CONSORT, SPIRIT). **Why this step exists:** the registration identifier is a single checklist item and can pass Step 4 even when the manuscript is internally inconsistent about *when* the registration or its amendments occurred relative to the analysis. An undisclosed post-hoc amendment is a common rejection trigger. **Five audit items (summary):** (1) registration identifier present in Methods, Abstract, and cover letter; (2) initial registration date precedes — or is explicitly disclosed as post-dating — the extraction milestone; (3) amendment dates appear in Methods, the described change is visible in Methods, analysis was re-run if amendment post-dates the lock, and no amendment post-dates submission; (4) cross-artifact agreement between Methods and the registry record (PROSPERO PDF, ClinicalTrials.gov export) — silent discrepancy is a finding; (5) retrospective-registration disclosure paragraph when evidence suggests post-extraction filing. **Registration-ID format gate:** a PROSPERO ID is `CRD42` + 9 digits = 14 characters (`^CRD42\d{9}$`, e.g. `CRD42024500001`). Run `grep -oE 'CRD42[0-9]+' manuscript.md` and assert each match is 14 characters long; a 15-character ID (a stray inserted digit) is a transcription error logged as `[REGISTRATION-TIMING]` (`fixable_by_ai: false` — verify against the live PROSPERO record, do not guess the correct digit). **Flagging:** any failure is logged in Part C Action Items with label `[REGISTRATION-TIMING]`. `fixable_by_ai: false` when reconciliation requires an external amendment filing; `true` only when the fix is a Methods-text insertion of a date already disclosed elsewhere. Part D JSON includes a `registration_timing` object (registry, id, initial_registration_date, amendments[], timing_consistency, findings[]). **Load-on-demand procedural detail** (exact item-by-item procedure, JSON schema, flagging edge cases): `${CLAUDE_SKILL_DIR}/references/step4c_registration_timing.md`. ### Step 4d: PRISMA Figure 1 Arithmetic & Cross-Reference Audit **Applies to:** systematic reviews and meta-analyses using PRISMA 2020 / PRISMA-DTA / PRISMA-P. Triggers when Item 16a (flow diagram) is PRESENT. **Why this step exists:** the flow diagram is a single checklist item and can pass Step 4 visually while still containing arithmetic errors (records screened ≠ identified − duplicates; sought-for-retrieval ≠ screened − excluded) or text↔figure number disagreements. Senior MA reviewers commonly require strict PRISMA 2020 diagram conformance and explicit body↔ figure number agreement; reviewers who detect these mismatches lose confidence in the study's data integrity immediately. **Four arithmetic checks:** 1. records screened = records identified − duplicates removed 2. records sought-for-retrieval = records screened − records excluded (screening) 3. reports retrieved = sought − reports not retrieved 4. studies included = reports assessed for eligibility − reports excluded (with reasons) **Two cross-reference checks:** - Body text PRISMA numbers (e.g., "315 records identified, 122 duplicates removed, 186 records screened") match Figure 1 box labels 1:1. - Reasons for exclusion (Methods + Figure legend) agree on counts and category names. **Procedure:** Run the deterministic implementation first — it performs steps 1, 4, 5, and 6 below automatically (same keyword regex, the four arithmetic equations, the body↔figure cross-reference) and writes `qc/prisma_figure_audit.json`: ```bash python3 ${CLAUDE_SKILL_DIR}/scripts/check_prisma_figure.py \ --md <manuscript.md> --figure <Figure 1 source: .md manifest / caption / text export> \ --out qc/prisma_figure_audit.json ``` Exit `1` = an arithmetic or cross-reference MISMATCH (log a Part C Action Item labelled `[PRISMA-FIGURE]`, `fixable_by_ai: false` — the author must reconcile the numbers); exit `2` = missing/unparsable input. The manual algorithm below documents exactly what the script checks and is the fallback when Figure 1 numbers live only in a PNG/SVG that must be transcribed by hand: 1. Extract numbers from manuscript Results / PRISMA flow paragraph (regex: integers near keywords `identified`, `duplicates`, `screened`, `excluded`, `sought`, `retrieved`, `assessed`, `included`). 2. Extract numbers from Figure 1 source — preferred order: (a) `analysis/figures/Figure1_PRISMA.md` markdown manifest, (b) caption text in `manuscript.md`, (c) PPTX text run if `.pptx` exists, (d) manual entry from PNG/SVG. 3. **Cross-check `analysis/figures/_figure_manifest.md`** (produced by `/make-figures`): verify that the row whose `Type = prisma` (or `Type = prisma-dta`) points at the same file path used as the audit source, and that the row's `Critic` field is `yes` or `partial` (not `no`). A missing manifest row, mismatched path, or `Critic = no` flag logs `[MANIFEST-XREF]` (advisory) — the arithmetic check still runs against the source identified in step 2. Skip this sub-step if `_figure_manifest.md` does not exist (older projects). 4. Run 4 arithmetic checks; emit PRESENT / MISSING / MISMATCH per equation. 5. Run 2 cross-reference checks; emit PRESENT / MISSING / MISMATCH per number. 6. Output `qc/prisma_figure_audit.json` and a short table. **Flagging:** any MISMATCH or arithmetic failure logs a Part C Action Item with label `[PRISMA-FIGURE]`. `fixable_by_ai: false` (numbers must be reconciled by the author). **Load-on-demand procedural detail** (exact regex set, JSON schema, edge cases — duplicates handled across databases, citation searching strand, dual-reviewer screening): `${CLAUDE_SKILL_DIR}/references/step4d_prisma_figure_audit.md`. **Cross-cutting**: integrates with `~/.claude/rules/numerical-safety.md` (PRISMA 5-way consistency: text ↔ Figure ↔ extraction CSV ↔ analysis script ↔ supplementary). ### Step 4e: Reporting-Framework Naming Audit **Applies to:** any manuscript that invokes an AI/extension reporting framework (PROBAST+AI, STARD-AI, TRIPOD+AI, TRIPOD-LLM, CONSORT-AI, SPIRIT-AI, PRISMA-DTA, QUADAS-C). **Why this step exists:** a base reporting tool and its extension are distinct instruments with separate citations (manuscript-style-classical §14). Step 1 routes to the right checklist but does not police how the framework is *named* in prose. The recurring failures are: invoking an extension without ever naming or citing the base instrument it extends; mixing `+AI` and `-AI` hyphenation for one family within a single document; coining item labels like "12-AI"; and waving at "recent guidance" instead of naming the framework. **Run the deterministic gate:** ```bash python3 "${CLAUDE_SKILL_DIR}/scripts/check_framework_naming.py" \ --manuscript manuscript.md --out qc/framework_naming.json --strict ``` **Verdicts:** `BASE_MISSING` (extension used, base instrument never named standalone) is a Major and logs `[FRAMEWORK-NAMING]` in Part C with `fixable_by_ai: true` (insert the base name + its citation). `HYPHEN_MIX`, `CITE_MISSING`, `SELF_COINED_LABEL`, and `VAGUE_GUIDANCE` are Minor (`fixable_by_ai: true`). Part D JSON includes a `framework_naming` object mirroring the script's `claims[]`. ### Step 4f: Critical-item floor cross-check **Applies to:** every guideline assessment **for which the floor defines a row** (load and check only those; do not invent a floor for an unlisted guideline). After the item-by-item table, load `${CLAUDE_SKILL_DIR}/references/critical_item_floor.md` and check the small set of **non-waivable** items for this study type. A MISSING critical item is surfaced as a **Critical gap** and becomes the report's headline regardless of the overall percentage — a high percentage with a missing critical item (undefined reference standard, no leakage-controlled partition, calibration absent for a prediction model, an unreconciled flow diagram) is not "broadly acceptable." For AI/ML and radiomics manuscripts, also confirm the chosen **methodological-quality / risk-of-bias** instrument (PROBAST+AI, METRICS/RQS, APPRAISE-AI) and its non-waivable concerns — a fully *reported* paper can still be at high risk of bias. For radiomics, the fuller METRICS breakdown (9 categories / 30 weighted items) is in `${CLAUDE_SKILL_DIR}/references/appraisal_tools/METRICS.md` (an appraisal reference, not a counted reporting checklist). Keep these distinct from the reporting counterparts (CLEAR, DECIDE-AI), which route through the normal checklist flow. Do not assert a numeric journal desk-reject threshold; the hard signals are a missing critical item and the journal's own required elements. ### Step 5: Generate Report Produce a structured compliance report in four parts. This report is an **internal working audit** — it carries auto-fix annotations, a machine-readable JSON block (`compliance_pct`, `fixable_by_ai`, …), and Action Items. It is **NOT** the official reporting checklist a journal expects (that is the blank guideline form with `Item | Recommendation | Reported in page/section`, which the authors fill in). **Never submit this report as the submission checklist.** So that the file is self-identifying and cannot be reused by filename into a later submission package, **the report MUST begin with this banner as its very first line**: ``` <!-- INTERNAL AUDIT — NOT FOR SUBMISSION. This is the /check-reporting working report, not the official journal checklist. Do not upload to a submission portal. --> ``` (`/sync-submission`'s `check_checklist_dump_leak` gate also catches this dump if it ever lands in a submission directory — but the banner is what makes it catchable.) **The four parts** — literal templates in `${CLAUDE_SKILL_DIR}/references/report_templates.md`: - **Part A — Summary.** Header (manuscript file, version token, guideline, date), the PRESENT/PARTIAL/MISSING/N-A count table, and overall compliance. The **headline is the critical items (Step 4f)**, not the percentage: report `{present}/{total}` and name every missing critical item with the section it belongs in. - **Part B — Item-by-item checklist.** One row per item: `# | Section | Item | Status | Location | Notes`. - **Part C — Action items** (MISSING and PARTIAL only), ordered by: items most journals enforce strictly (ethics approval, registration, sample size) → items in Methods (easiest to fix) → everything else. - **Part D — Machine-readable JSON**, appended as a fenced block. **MUST** be present under `--json` or when called from `/write-paper` Phase 7, which parses it. **JSON field contract** (the part other skills depend on — get these right): - `compliance_pct` — `present / (total_items - na) * 100`, one decimal. - `action_items` — MISSING and PARTIAL only; PRESENT and N/A are excluded. - `fixable_by_ai` — `true` when the fix inserts or expands text using information already in the manuscript or inferable from it; `false` when it needs external facts the author alone holds (registration number, IRB approval number, protocol details). - `suggested_fix` — concrete draft text, insertable as written. - `source_sha256` — first 12 hex chars of the SHA-256 of the manuscript bytes, so a stale report cannot be silently attributed to a newer manuscript. **Read on demand:** | File | Read it when | Cost if read blindly | |---|---|---| | `references/report_templates.md` | you have finished the audit and are writing the report | ~1,900 tokens of pure output format — it informs no part of the assessment itself | --- ## Assessment Standards ### Be Strict - PARTIAL means the item is mentioned but lacks specificity. For example: - "We used appropriate statistical tests" = PARTIAL (which tests?) - "We used the Mann-Whitney U test for continuous variables and Fisher's exact test for categorical variables" = PRESENT - A vague reference does not count as PRESENT. The detail level must match what the guideline expects. ### Be Specific in Suggestions - For MISSING items, provide a draft sentence the user can insert. - For PARTIAL items, point to the exact gap and suggest specific additions. - Reference the specific manuscript section where the addition should go. ### Common Gaps to Watch For These items are frequently missing in medical manuscripts: 1. **Study registration number** (CONSORT, PRISMA, STARD) 2. **Registration / amendment date consistency** (PRISMA 2020, PRISMA-DTA, CONSORT, SPIRIT) — run Step 4c whenever a registration identifier is present 3. **Sample size justification** (CONSORT, STROBE, STARD) 4. **Missing data handling** (all guidelines) 5. **Blinding details** (CONSORT, STARD) 6. **Funding and conflicts of interest** (all guidelines) 7. **Ethics approval with committee name and approval number** (all guidelines) 8. **Data availability statement** (increasingly required) 9. **AI-specific: training/validation/test split details** (TRIPOD+AI, CLAIM, STARD-AI) 10. **AI-specific: model architecture and hyperparameters** (TRIPOD+AI, CLAIM, STARD-AI) 11. **AI-specific: failure mode analysis** (CLAIM, STARD-AI) 12. **AI-specific: fairness/bias assessment** (STARD-AI) 13. **AI-specific: commercial interests and data/code availability** (STARD-AI) 14. **Power-aware framing of a null result** (STROBE 16a / 18 / 20) — for an observational study whose headline is a **non-significant** association, a flat "X was not associated with Y" overreads the data when the analysis is not powered to *exclude* a clinically meaningful effect. Mark item 18/20 PARTIAL unless the manuscript states the precision as an exclusion (e.g., "the 95% CI excluded an eGFR difference larger than ~1.7") or reports a minimum detectable effect — "no effect" vs "could not exclude an effect of size X" are different claims, and a negative conclusion needs the latter. 15. **Confounder-selection rationale, not "adjust for everything that differs"** (STROBE 16a explicitly asks *which confounders were adjusted for and why*) — flag a kitchen-sink adjustment set chosen because variables differ in Table 1. The Methods must give a causal rationale (DAG / prior literature) and must not adjust for a **mediator or consequence of the outcome** (over-adjustment, e.g. serum uric acid in an eGFR model); both an unjustified inclusion and an unjustified omission are item-16a gaps. --- ## PRISMA Cascade Arithmetic Auto-Verify PRISMA 2020 flow diagrams chain a cascade of subtractions (database records → after dedup → title/abstract screened → full-text reviewed → included in synthesis). Off-by-one errors in the prose cascade are a high-frequency reviewer red flag (e.g., `151 + 108 + 39 + 1 + 1 + 4 = 304` followed by a prose summary "305" four lines later). When PRISMA 2020 or PRISMA-DTA is selected and round-by-round screening TSV artifacts are available, run the cascade auto-verify: ```bash python "${CLAUDE_SKILL_DIR}/scripts/prisma_cascade_check.py" \ --round1 2_Screening/round1.tsv \ --round2 2_Screening/round2.tsv \ --round3 2_Screening/round3_adjudication.tsv \ --manuscript manuscript.md \ --out qc/prisma_cascade.json ``` The script: 1. Reads the round TSVs and counts `INCLUDE` / `EXCLUDE` / `MAYBE` decisions per round. 2. Computes the cascade arithmetic from raw decisions (no prose). 3. Optionally grep the manuscript for matching stage-count claims and emits per-stage drift when the prose disagrees. Treat any `manuscript_drift` entry as a P0 blocker — fix the prose to match the computed cascade and re-run. ## Submission Checklist Export Many journals require a filled reporting checklist to be submitted alongside the manuscript. When the user asks for a submission-ready checklist, format the output as: ``` {Guideline Name} Checklist Manuscript title: {title} Date: {YYYY-MM-DD} | Item # | Checklist Item | Reported on Page # | Reported in Section | |--------|---------------|-------------------|-------------------| | 1 | {item text} | {page or N/A} | {section} | | 2 | {item text} | {page or N/A} | {section} | | ... | ... | ... | ... | ``` Page numbers should be filled in by the user after final formatting. Use section names as placeholders. --- ## Skill Interactions | When | Call | Purpose | |------|------|---------| | During manuscript writing | `/write-paper` Phase 7 | Final compliance check | | Need to add Methods text | `/write-paper` Phase 3 | Draft missing Methods content | | Need statistical details | `/analyze-stats` | Generate missing statistical reporting | | Need flow diagram | `/make-figures` | Generate CONSORT/STARD/PRISMA diagram | --- ## Error Handling - If the manuscript file cannot be read, ask the user for the correct path. - If the study type is ambiguous, ask the user to confirm before selecting a guideline. - If a checklist item is genuinely unclear in its applicability, mark as N/A with justification. - This is a pre-screening tool. Always remind the user that final compliance should be verified by all co-authors and ideally by a methodologist. ## Language - Checklist content and compliance report: English - Communication with user: Match user's preferred language - Medical terms: English only ## Anti-Hallucination - **Never fabricate references.** All citations must be verified via `/search-lit` with confirmed DOI or PMID. Mark unverified references as `[UNVERIFIED - NEEDS MANUAL CHECK]`. - **Never invent clinical definitions, diagnostic criteria, or guideline recommendations.** If uncertain, flag with `[VERIFY]` and ask the user. --- ## Gates | Gate | Severity | Trigger | Action on fail | |---|---|---|---| | Mandatory items present | ENFORCED at submission | < 100% of guideline-mandatory items marked PRESENT | Auto-fix MISSING items where text exists; otherwise route to `/write-paper` Phase 7 for re-draft | | Step 4d PRISMA Figure 1 arithmetic & cross-reference audit (PRISMA / PRISMA-DTA only) | ENFORCED for SR/MA | flow numbers don't sum (e.g., screened ≠ included + excluded), or in-text counts mismatch flow diagram | HALT; reconcile against extraction artifacts | | Optional items (e.g., supplementary AI declarations) | ADVISORY | < 80% of optional items present | warn; user accepts | | Cross-reporting-guideline routing (study type → guideline) | ENFORCED | study type undeclared or guideline missing | Ask user; do not silently default | ## Global-rule references Some passages in this skill cite a path of the form `~/.claude/rules/<name>.md`. Those are the maintainer's personal global rules, kept outside this repository. They are **not shipped with this skill** and will not exist on your machine; they appear only as provenance for where a convention came from. If one of them looks like it is standing in for an instruction you actually need, that is a bug — please open an issue, because the instruction belongs here. -
skill.yml 1.9 KB
schema_version: 2 name: check-reporting layer: A owner_domain: reporting_compliance maturity: official when_to_use: - Pre-submission item-level checklist run for a target reporting guideline (STROBE / CONSORT / STARD / TRIPOD / PRISMA / ARRIVE / CARE / SPIRIT / CLAIM / SQUIRE 2.0 / etc.) - Risk-of-bias assessment using QUADAS-2 / QUADAS-C / RoB 2 / ROBINS-I / ROBINS-E / ROBIS / PROBAST / NOS / COSMIN / RoB NMA - AI-extension audit (STARD-AI, TRIPOD+AI, TRIPOD-LLM, CONSORT-AI, SPIRIT-AI, PROBAST+AI, CLAIM 2024, DECIDE-AI, MI-CLEAR-LLM, CLEAR) - Reviewer or editor asks the manuscript "to comply with [guideline]" — generate the gap list when_NOT_to_use: - Drafting manuscript prose to fill missing items (use /write-paper or /revise) - Inventing checklist items not in the published guideline (forbidden) - Reviewer-style critique of intellectual content (use /peer-review or /self-review) inputs: - manuscript/manuscript.md outputs: - qc/reporting_checklist.md - qc/reporting_checklist.json deterministic_scripts: - none_required side_effects: - writes_qc_artifacts downstream_consumers: - write-paper - self-review forbidden_actions: - invent_reporting_guideline_items - silently_mark_missing_items_present # v2.1 quality card purpose: "Audit a manuscript item-by-item against a chosen reporting guideline (36 supported) with PRESENT/MISSING/PARTIAL status." safety_boundaries: - "Checklist items are quoted from the guideline, never invented; missing items are not marked present." - "Fails fast if the requested checklist file is absent rather than generating a guessed checklist." known_limitations: - "Item judgements are advisory; a PRESENT mark is a locator, not a quality guarantee." - "Coverage is limited to the bundled checklists." validation_commands: - "python3 scripts/check_checklist_exists.py <guideline>" - "python3 scripts/prisma_cascade_check.py" evidence_surface: demo
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.