bensz-rmd-rules
规范 AI 开发 R Markdown 分析脚本的行为准则。当用户要求"写 Rmd 分析"、"开发 R 脚本"、"做数据分析"时触发。核心原则:遵循主业与副业分离架构(.R 保留完整数据,.Rmd 应用业务阈值),优先使用用户已有 R 包资源;图表默认按 Nature 级别可读性与出版质量生成;专家级解读兼顾弱背景读者,提供四层框架、指标导读与不常用指标首次解释协议;路径验证确保跨平台兼容性。前提:luckyBase 为硬依赖。
Install
npx skills add https://github.com/huangwb8/skills/tree/main/skills/beta/bensz-rmd-rules
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install huangwb8-skills@llmmart
git clone https://github.com/huangwb8/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole huangwb8/skills collection as a plugin from our marketplace. Git is the plain clone.
README
bensz-rmd-rules
面向 R 数据分析、R Markdown 报告、可复现 targets 流程、论文级图表与证据解读 的 Agent Skill。当前版本以 config.yaml 中的 skill_info.version 为准;本目录属于 beta 候选源。
什么时候使用
使用场景:从原始数据完成整理、统计/模型、图表、Rmd/HTML、科学产品和结果解释,或维护已有 R/Rmd 项目。
不要用于:主要交付物是跨项目复用的 R 函数、类、稳定 API 或 Package(改用 bensz-r-developer);仅渲染既有 Rmd(改用 knit-rmd-html);其它语言的数据分析。
最短用法
请用 bensz-rmd-rules 完成这个 R 分析。先只读判断项目是 new 还是 existing,再说明选择 simple 或 targets-first complex/pipeline 的理由。使用 renv,并在正式运行前从正式入口真实跑通 synthetic_fixture 或 project_subset 轻量测试。
Skill 会保留已有项目的现状:不因默认策略自动创建 targets/renv 或迁移目录;有 _targets.R 的项目按 complex 维护,无 targets(含历史编号脚本)的项目按 simple 语义维护。
工作流模式
| 模式 | 适用场景 | 固定要求 |
|---|---|---|
simple |
线性、低成本、整体重跑可接受 | renv、明确 R/Rmd 入口、真实轻量测试;不创建 targets |
complex(说明性别名 pipeline) |
非线性依赖、昂贵步骤、多下游复用、局部失效、恢复或并行 | _targets.R 唯一 DAG、R/ 计算函数、Rmd 消费 target、隔离 store 恢复验收 |
new/existing 是项目状态而不是第三种模式:已有 _targets.R 按 complex,否则按 simple。迁移必须由人类明确授权。
complex 的核心关系:
renv.lock → _targets.R → R/ functions → target results → Rmd → reports/
↘ optional products/
↘ _targets/ (machine state)
_targets/ 只存机器计算状态;products/ 只存需要审阅、复用或交付的科学对象,不实现第二套缓存、SUCCESS、identity hash 或恢复 runner。
推荐项目布局
项目根目录/
├── 00.Environment.R
├── R/ # target 调用的计算函数
├── _targets.R # complex 的唯一 DAG 入口
├── raw/ # 只读
├── products/ # 可选科学产物
├── reports/ # 图、表、HTML、补充材料
├── scripts/tests/ # 可版本化测试代码
├── tmp/tests/<run-id>/ # 隔离测试现场
├── renv.lock
└── renv/activate.R
模板 _targets.R 使用 tar_source("R");R_data_template.R 提供数据处理起点;Rmd_template.Rmd 通过 targets::tar_read() 消费 analysis_results;simple 可从 Rmd_simple_template.Rmd 开始。
统计推断完整性
论文级分析在计算前列出主要及影响结论的次要估计对象,确认目标人群、观察单位、对比、设计和推断目的。方法允许时,完整结果保留估计值与单位、有效 N/事件数、适当的默认 95% CI;只有零假设明确时才报告 p 值,批量检验同时说明校正方法及 q/调整后 p。预注册方案另定置信水平时按方案执行。报告引用计算结果并解释效应的实际意义和局限。固定常数或纯描述结果可标 not_applicable;数据或设计无法支持推断时标 not_estimable 并说明原因与替代分析;未运行标 not_run。不凭空补 CI/p,也不以文本关键词检查代替方法复核。详见 statistical_inference_protocol.md。
轻量测试与恢复
新建或实质修改的流程必须真实运行一种测试:
synthetic_fixture:无授权真实数据、数据敏感或数据过大时,保留 schema、类型、主键、分组、缺失和边缘条件,并固定随机种子。project_subset:有授权数据和现有代码时使用代表性子集,覆盖关键分组/结局/缺失/异常;不能只用head(n)。
测试代码放 scripts/tests/,每次使用唯一 tmp/tests/<run-id>/。simple 调用正式入口;complex 执行同一 DAG,但 store 放在 <run-id>/_targets,不能污染正式 _targets/、products/ 或 reports/。至少断言输入契约、关键类型/主键、重要数值不变量、产品/报告生成及 raw/ 无写入。未跑全量时记录 full_data_execution=NOT_RUN。
恢复验收:先让昂贵 target 成功,再中断后续 target;在输入、代码、参数和 renv 身份不变时再次 tar_make(),用 tar_meta()、outdated 集合和执行记录证明有效前序 target 被跳过。
检查与脚本入口
先检查项目状态和模式:
python3 <skill-root>/scripts/check_targets_renv.py <project> --project-state auto --workflow-mode auto
常用检查:
python3 <skill-root>/scripts/check_pipeline_contract.py <project>
python3 <skill-root>/scripts/check_interpretation_quality.py <project>/report.Rmd
python3 <skill-root>/scripts/check_figure_table_interpretation.py <project>/report.Rmd
python3 <skill-root>/scripts/check_htmlwidget_visibility.py <project>/report.Rmd
python3 <skill-root>/scripts/check_rmd_template_yaml.py <project>/report.Rmd
Rscript <project>/scripts/tests/smoke_test.R
图表可读性检查器为 check_plot_readability.R,路径安全检查器为 validate_paths.R。Liquid Glass 主题可用 bootstrap_liquid_glass.py 初始化;桌面动态目录会立即扩展可交互区域,避免鼠标移入时收回。HTML 渲染交给 knit-rmd-html。
图表与解读规范
默认图表语言为英文;中文期刊等场景通过 YAML params.plot_language 切换。每个可见图/表附近说明当前对象、方向/对比、可追溯数值、量级、不确定性和后续验证(方法 + 输入 + 判据)。参考:
four_tier_interpretation_framework.md:四层解读与 Fail Fast Gate。interpretation_templates.md:单因素、多因素和模型验证骨架。plot_quality_standards.md:Nature 级图表可读性。liquid_glass_theme_guide.md:HTML 主题与故障排查。
相关模板与参考
- 环境与依赖:
00.Environment.R、renv/activate.R。 - 测试:
templates/tests/、lightweight_testing.md。 - 模式与架构:
workflow_modes.md、hybrid_architecture_guide.md。 - 交付与审查:
delivery_verification.md、serial_review_protocol.md。
FAQ 与边界
可以把编号脚本自动迁移成 targets 吗? 不可以。已有无 targets 项目按 simple 维护;迁移前需人类明确授权、映射、结果校验和回退方式。
可以用自制 checkpoint 或 SUCCESS 文件吗? 不可以。complex 使用 targets 的 metadata、失效和增量重建;不引入第二套缓存/恢复协议。
轻量测试通过是否等于全量通过? 不是。必须分别报告 preflight、lightweight_execution 和 full_data_execution;未运行全量写 NOT_RUN。
需要通用 R 函数怎么办? 由 bensz-r-developer 实现,本 Skill 负责分析需求、接入和流程证据。
许可证与贡献
仓库许可证见项目根目录的 LICENSE。修改 Skill 时请同步 SKILL.md、config.yaml、必要 references 和变更记录,并运行与风险匹配的检查;不要在 raw/ 或 README 中写入凭据、隐私数据或私有提示词。
Skill manifest
bensz-rmd-rules
目标
把研究目标和数据组织为可审查、可复现的 R 分析流程。新项目必须用 renv,并在真实 R/Rmd 入口完成轻量测试:线性低成本流程用 simple;有非线性依赖、昂贵步骤、复用、局部失效、恢复、血缘或并行需求时用 targets-first complex/pipeline(_targets.R 唯一 DAG,计算函数在 R/,Rmd 只消费 target)。_targets/ 是机器状态,products//reports/ 是科学产物。已有项目不自动补齐或迁移:有 _targets.R 按 complex,无 targets(含历史编号脚本)按 simple;需要复杂能力时由人类显式迁移。本 Skill 不维护编号 runner。
主要交付物决定触发,而不是扩展名:
| 交付物 | 主导 Skill |
|---|---|
| 分析数据流、统计结果、可恢复产品、Rmd/HTML、图表与解读 | bensz-rmd-rules |
| 可复用、版本化或跨项目的函数、类、API、Package | bensz-r-developer |
| 分析需要可复用组件 | 本 Skill 定义需求/集成证据,bensz-r-developer 实现,本 Skill 接回验证 |
可创建只服务当前分析的 R/ 函数和 helper,但不扩展为公共 Package。luckyBase 是唯一固定 Bensz 依赖;其它包由项目写入 renv.lock。crew、autometric、集群插件只是项目级可选依赖;不自制调度器、日志协议、worker 状态机或资源采样器。本 Skill 不自动接入 State、Verifier、Pack 或 Gate。
流程
输入
收集:研究问题、用途、交付物、统计边界和完成判据;主要/次要估计对象、目标人群、观察单位、对比与效应尺度、研究设计、缺失/依赖结构、预注册方案及推断目的;原始输入/数据字典/隐私授权及 raw/ 只读边界;重算成本、依赖、外部请求、随机性、关键参数和恢复需求;现有 R/Rmd、00.Environment.R、产品路径、tmp/、_targets.R、renv.lock/renv/;可用测试数据与边缘条件;函数复用范围。缺失信息只有在改变行为或安全边界时才询问。
先读 config.yaml,再按任务读取最少 references:
- 模式/目录:
workflow_modes.md、hybrid_architecture_guide.md;测试:lightweight_testing.md。 - 缓存/示例:
analysis_workflow_cache.md、hybrid_architecture_examples.md;审查:serial_review_protocol.md。 - R 实现:
code_style_guide.md、no_overdefensive_code.md、cross_platform.md。 - 指标/解读:
metric_explanation_protocol.md、four_tier_interpretation_framework.md、interpretation_templates.md、interpretation_narrative_examples.md、expert_discussion_template.md。 - 论文级估计与检验:
statistical_inference_protocol.md;计算前逐项确定可估计性与方法,报告时逐项核对。 - 图表/HTML:
plot_quality_standards.md、plot_language.md、htmlwidget_visibility_rules.md、liquid_glass_theme_guide.md;ID/函数:gene_id_guidelines.md、candidate_r.md。
对 raw/ 仅做规划和选样所需的轻量只读盘点,不启动完整计算。
执行步骤
- 映射交付与推断:若只要公共函数/API/类/Package,转交
bensz-r-developer;否则建立 target—输入—科学产品—报告—完成判据映射,确保依赖无环。计算前列出影响论文结论的主要与次要估计对象,逐项判断是否需要且能够给出区间/检验,记录设计、方法理由和结果去向;预注册方案优先,不为显著性事后更换主终点或检验。纯描述、固定常数或不可估计项明确标状态与原因,不伪造 CI/p。新 pipeline 的计算函数放R/,由_targets.R的tar_source("R")发现;00.Environment.R默认唯一,只负责包、项目根、产品路径和全局配置。 - 判状态和模式:先运行
优先级为人类要求 → existing 边界 → 复杂度。新项目默认 simple;复杂信号(非线性、昂贵/高失败代价、多下游复用、局部失效、恢复、血缘、并行)选 complex。python3 <skill-root>/scripts/check_targets_renv.py <项目根> --project-state auto --workflow-mode autoexisting不是第三种模式:有_targets.R按 complex,否则按 simple;不创建缺失 targets/renv,不隐式迁移。详见workflow_modes.md。 - 固定路径:按需使用
00.Environment.R、R/、_targets.R、raw/、products/、reports/、templates/、scripts/tests/、tmp/tests/<run-id>/、tmp/scratch/、renv.lock和renv/activate.R,不预建空目录。products/只存需审阅/复用/交付的科学产物;路径优先级为人类指定 → 已有路径 → 项目统一设置 →products/,必须在项目内且不得与raw/、reports/、tmp/、_targets/、.bensz-api/重叠。 - 设计缓存边界:只导出昂贵、可复用、高风险/有损、需人工审查、依赖可变外部请求或中断代价高的结果。targets 管理
_targets/metadata、失效和增量重建;禁止SUCCESS、identity hash、force-step/resume runner 或 checkpoint helper。详见analysis_workflow_cache.md。 - 计算推断并实现报告:
00.Environment.R用luckyBase::Plus.library();分析代码优先pkg::fn(),基因 ID 用luckyBase::convert()。分析层对关键参数计算估计值、适当的默认 95% CI(方案另有水平时从其规定),并仅在有明确检验问题时给 p 值;多重检验记录校正方法与 q/调整后 p。保留未经显著性筛选的完整结果,含参数身份、尺度/单位、有效 N/事件数、方法、CI 界限/水平、适用的 p/q、来源及不适用/不可估计/未运行状态与理由。complex 的R/函数读取raw//上游 target;Rmd 用tar_read()/tar_load()消费真实结果,不在正文臆算。Top N/阈值/配色放 YAMLparams。target 语义化命名并按阶段注释;交付摘要附targets::tar_manifest()。静态图写reports/,HTML 留根目录;默认不在图内重复标题。只在 I/O 边界和硬前提检查,不用占位/静默降级掩盖失败。 - 真实轻量测试:新建或实质修改流程选择
synthetic_fixture(保 schema、类型、主键、分组、缺失和边缘条件,固定种子)或有授权时的project_subset(覆盖关键分组/结局/缺失/异常,不能仅head(n))。测试在scripts/tests/,每次使用唯一tmp/tests/<run-id>/;只改输入规模和路径,不加业务analysis_mode/test_mode。simple 调正式入口;complex 执行同一 DAG,store 放<run-id>/_targets。至少断言输入契约、类型/主键、数值不变量、产品/报告生成和raw/未写入。checker/语法/--dry-run仅为preflight;真实链为lightweight_execution;未跑全量记full_data_execution=NOT_RUN。按环境依赖、fixture/子集、路径隔离、分析代码、科学断言、渲染、外部服务分类失败。 - 串行审查:初稿和轻量测试后,按
serial_review_protocol.md完成结构/数据流、R/恢复、科学/统计三项只读审查;每项读取前项修正后的最新版本。影响执行、输出或断言的修正会使旧证据失效,须在最新 identity 上重跑受影响链;不能提供独立 Agent 时须记录降级。 - 正式运行与报告:simple 用明确 R/Rmd 入口;complex 先检查契约,再
targets::tar_make(),store 为_targets/。中断后用tar_meta()、outdated 集合和执行记录证明有效 target 被跳过,不手工写SUCCESS。图表默认英文,中文场景用params.plot_language;静态图保存 PDF,并在任务工作区生成 JPG 预览。每个关键结论给出与计算产品对应的估计值、CI、有明确假设时的 p/q、实际意义与局限;无法推断时写清状态和理由。图/表附近给对象、方向/对比与可追溯证据,不为装饰性数字逐个检验。
输出
交付需求—target—报告映射、唯一 00.Environment.R、R/ 函数、Rmd/HTML、renv.lock、renv/activate.R、scripts/tests/、reports/ 和验证摘要。complex 另交 _targets.R 与 tar_manifest() 快照,按需交分析计划和科学产品;simple 不创建 targets。测试记录绑定最新 identity,分别报告三层执行状态,不能以轻量通过暗示全量通过。
输出管理
正式代码、测试、Rmd/HTML、产品和报告留在用户项目约定位置;草稿、审查、JPG 和日志进入当前 .bensz-api/task-*;测试现场进入 tmp/tests/<run-id>/。不写 raw/;products/ 不放草稿;reports/ 不放源码、缓存或审查日志;tmp/scratch/ 的重要发现须晋升并补 renv/轻量测试证据。
校验
# 新 simple
python3 <skill-root>/scripts/check_targets_renv.py <项目根> --project-state new --workflow-mode simple
Rscript -e 'renv::status()'
Rscript <项目根>/scripts/tests/smoke_test.R
# 新 complex
python3 <skill-root>/scripts/check_targets_renv.py <项目根> --project-state new --workflow-mode complex
Rscript -e 'renv::status()'
Rscript <项目根>/scripts/tests/smoke_test.R
Rscript -e 'targets::tar_make()'
Rscript -e 'targets::tar_meta()'
# 已有项目
python3 <skill-root>/scripts/check_targets_renv.py <项目根> --project-state existing --workflow-mode auto
所有 Rmd 还要检查解读覆盖/质量、widget 可见性、图表可读性和数字追溯,并按实际情况做 R 语法/运行、Rmd 渲染;逐项复核主要估计对象的结果或不可估计理由、CI/p 方法与设计的一致性。现有文本检查器只作启发式预检,词语命中不证明推断已计算或方法正确。检查清单见 workflow_checklist.md,证据格式见 delivery_verification.md。最低验收:模式理由可复述、renv 可信、测试未污染正式路径、Rmd 未绕过 DAG、恢复可证明、已有项目未隐式迁移、跳过项不写成通过。
失败与恢复
- 缺
luckyBase或新项目无 renv:停止补齐前提,不更换包管理逻辑。 - simple 出现复杂信号:评估升级 complex,不扩展调度器或无理由降级。
- 测试失败:按失败类别最小修正并重跑;权限/网络/授权/依赖不足时记录阻塞,不伪造成功。
- complex 测试写正式
_targets/、products、reports 或 raw:停止、隔离并确认正式路径后重跑。 - targets 失效/失败按 outdated/error 重算必要节点;不伪造
SUCCESS。Rmd 渲染失败只从报告层恢复。 - existing 冲突不覆盖/迁移;沿现有 targets 或顺序入口做授权增量修改,需复杂能力时显式迁移。
约束
公共硬约束
本块由 docs/templates/skill-common-constraints.md 统一维护;每个 SKILL.md 的 ## 约束 必须逐字同步本块,不得在副本中改写公共规则。
- 任务需要落盘时,使用唯一的
./.bensz-api/task-{yyyymmdd-hhmm}-{简短描述}/根目录;共享材料放入shared/,Skill 专属材料放入该 Skill 的input/、output/、log/。 - 正式交付物、源代码和正式计划按项目约定保存,不写入任务工作区;未经授权不覆盖、删除、迁移或远程写入。
- 项目维护变更检查 BAC 可用性并记录需求、AI 产出、工具结果、文件改动和验证摘要;BAC 只做过程审计,不替代署名、责任或合规判断。
- 不记录 API Key、访问令牌、密码、Cookie、环境/凭据文件、私有 Prompt、身份信息、本地用户名、主机名或不必要的大体积原始数据。
- 文件路径必须规范化并限制在授权项目范围内;外部 URL、子进程和网络访问遵循最小权限,防止路径遍历、SSRF 和命令注入。
- Skill 版本唯一记录在自身
config.yaml:skill_info.version;公开 API、协议、目录或配置变更同步文档与CHANGELOG.md。 bensz-collect-bugs是一个 Agent Skill;仅将 Bensz Agent Skill 或 Bensz 基础设施本身的设计缺陷交给它。先脱敏写入~/.bensz-skills/bugs/,当前任务不中断,只有用户明确要求才公开上报,禁止直接修改用户已安装的 Skill 源码。
Skill 专属约束
- 只在授权范围内只读盘点
raw/;不得修改、覆盖或向其写入,也不在日志中记录绝对私有路径、凭据或不必要原始值。 - 不因模板存在就创建全部文件;模式、缓存边界和目录由真实需求决定。
- 新 pipeline 不引入 checkpoint helper、SUCCESS 或自定义 identity runner;该自制缓存机制已移除,不得在任何项目中新引入或修复。
Files (skills)
-
evals
-
evals.json 9.6 KB
{ "skill_name": "bensz-rmd-rules", "evals": [ { "id": 1, "prompt": "在空目录新建一个低成本描述性 Rmd 报告,只读取一张小表并生成一个 HTML。数据不能外传,也没有授权真实样本用于测试。", "expected_output": "选择 new + simple,使用 renv,不创建 _targets.R;选择 synthetic_fixture,并从正式 Rmd 入口真实 render 到唯一 tmp/tests/<run-id>/,记录 full_data_execution=NOT_RUN。", "assertions": [ "明确写出 project_state=new、workflow_mode=simple 及理由", "包含 renv.lock、renv/activate.R 与 scripts/tests/smoke_test.R", "不创建或要求 _targets.R", "fixture 覆盖 schema、类型、关键分组、缺失和边缘条件", "运行记录绑定 subject identity 且不把轻量测试冒充全量通过" ], "files": [] }, { "id": 2, "prompt": "新建一个线性小分析:清洗后立刻画图,整个流程几十秒可重跑。我明确要求 simple。请给出实现和验收入口。", "expected_output": "遵循人工 simple 覆盖,使用明确 R/Rmd 入口与 renv,不开发 targets 或自制 runner;真实 smoke test 通过后交付。", "assertions": [ "人工指定 simple 优先", "没有 targets 资产或自制调度框架", "checker/dry-run 未被当作端到端通过" ], "files": [] }, { "id": 3, "prompt": "空目录中建立清洗、外部注释、昂贵模型和两个下游报告。已有授权真实数据,模型结果会被两个报告复用,需要断点恢复。", "expected_output": "选择 new + complex;使用 targets + renv + products;选择 project_subset;同一 target 图在唯一 tmp/tests/<run-id>/_targets 真实运行,清理真实子集,正式 _targets/products/reports 不被测试污染。", "assertions": [ "complex 理由包含昂贵步骤、复用或恢复", "project_subset 有确定且代表性的抽样规则而非无理由 head(n)", "测试 store 明确位于唯一 tmp/tests/<run-id>/_targets", "真实子集默认清理且只保留不可还原证据", "区分 targets store 与正式 products" ], "files": [] }, { "id": 4, "prompt": "我坚持新项目用 simple,但任务含一次高成本 API 请求和三个下游报告。请按我的模式做。", "expected_output": "优先遵循人工 simple,同时明确高成本与多下游复用风险,不静默改为 complex;不通过自制 runner 伪装 complex,给出可审查的升级建议。", "assertions": [ "不静默覆盖人工模式", "披露 simple 的重跑/恢复风险", "不创建第二套轻量调度器" ], "files": [] }, { "id": 5, "prompt": "维护已有 R 项目:有 raw/、tmp/ 和旧 runner,但没有 targets 或 renv。只修复报告统计错误。", "expected_output": "判为 existing;无 _targets.R 的项目按 simple 语义维护(编号/旧脚本只是顺序执行的普通入口),renv 缺失只披露不自动初始化;修复后沿现有脚本入口做可隔离轻量试跑,无法隔离则明确阻塞;不创建 _targets.R、renv.lock 或迁移产品,也不提供或维护任何编号 runner。", "assertions": [ "existing 是项目状态,不再引入 preserved-existing 模式", "只报告复现性/恢复风险,不自动补建机制", "实质修改沿现有入口测试或记录 path_isolation 阻塞", "任务需要恢复/复用等复杂能力时给出显式迁移建议,而非自维护 runner" ], "files": [] }, { "id": 6, "prompt": "已有项目已经有 _targets.R 和 _targets/,但没有 renv。请增加一个 target 并运行检查。", "expected_output": "判为 existing + complex(有 _targets.R 自动解析为 complex),按 R/ + tar_source 契约维护现有 targets 并新增 target;说明 renv 缺失但不自动初始化;existing 检查不产生 renv.lock。", "assertions": [ "保留现有 target 图与正式 store", "不自动新增 renv", "测试如适用时使用隔离 store" ], "files": [] }, { "id": 7, "prompt": "新 complex 项目的正式派生产品必须写到 derived/stable,而不是 products。请确保所有计算和报告一致。", "expected_output": "通过单一项目设置覆盖 products 路径,所有 helper/R/Rmd/checker 使用同一路径;规范化并限制在项目内,不散落硬编码。", "assertions": [ "使用 BENSZ_PRODUCTS_DIR 或等价单一项目入口", "路径保持在项目内且不与 raw/reports/tmp/_targets 重叠", "targets store 仍与正式派生产品分离" ], "files": [] }, { "id": 8, "prompt": "请搭建一个 complex R 项目,并说明 CSS、测试代码、测试数据副本和正式结果各放哪里。", "expected_output": "CSS/渲染资产在 templates,测试代码在 scripts/tests,运行副本在 tmp/tests,正式产品在 products,正式图表在 reports;只创建实际使用目录。", "assertions": [ "templates 不包含自制缓存/完成标记文件", "scripts/tests 与 tmp/tests 职责分开", "tmp 发现成为正式依据时需要晋升并补测试" ], "files": [] }, { "id": 9, "prompt": "轻量测试第一次因为 fixture 缺少一个关键分组而失败,修正后报告又因缺包失败;审查随后还修改了执行代码。请给出交付记录。", "expected_output": "分别分类为 fixture_or_subset 与 environment_dependency;记录命令、退出状态、最小修正和真实链重跑;审查修改使旧 subject identity 证据过期,最新版本必须重跑;未跑全量时 full_data_execution=NOT_RUN。", "assertions": [ "两次失败类别准确且分开记录", "修正只作用于对应层", "包含最终断言、subject identity、三层执行状态与剩余风险" ], "files": [] }, { "id": 10, "prompt": "把多个研究项目重复复制的 normalize_counts() 抽成带稳定 API、roxygen2 文档和 testthat 的 R Package。", "expected_output": "主要交付物是可复用 Package,应转交 bensz-r-developer;若仍有分析流程,本 Skill 只定义集成需求并在组件完成后验证。", "assertions": [ "不因请求包含 R 文件而错误触发分析实现", "明确组件工程与分析集成边界" ], "files": [] }, { "id": 11, "prompt": "已有项目的旧入口硬编码绝对输出路径,轻量测试会覆盖正式报告;我只要求修一个分析 bug,没有授权迁移或重构路径。", "expected_output": "判为 existing 并按现有入口做最小修复;不强行运行有污染风险的测试,也不补建新版架构;将轻量执行记为 BLOCKED、失败类别记为 path_isolation,并请求人类决定是否授权最小可测试性改造或显式迁移。", "assertions": [ "没有把 checker 通过写成轻量执行通过", "没有擅自迁移、覆盖正式报告或改造旧目录", "记录 path_isolation 阻塞、未验证项和可选最小改造" ], "files": [] }, { "id": 12, "prompt": "新建一个需要数小时运行的 targets 分析:清洗、模型和报告都要可恢复。请按最新 targets-first 架构设计目录、函数、Rmd 消费关系和中断恢复验收。", "expected_output": "complex/pipeline 使用 _targets.R 作为唯一 DAG,计算函数在 R/,Rmd 用 tar_read()/tar_load(),_targets/ 只保存机器状态,products/reports 只保存科学产物;真实中断后 tar_make() 复用有效前序 target。", "assertions": [ "没有第二套 SUCCESS/identity/checkpoint runner", "明确 R/、_targets.R、Rmd 与 _targets/products/reports 职责", "恢复验收包含真实中断、再次 tar_make 和 tar_meta/outdated 证据" ], "files": [] }, { "id": 13, "prompt": "已有一个旧项目,里面有编号 R 脚本、checkpoint_helpers.R 和自定义 runner。我只想修复一个统计错误,不要迁移架构。", "expected_output": "判为 existing 并按 simple 语义维护:编号脚本只是顺序执行的普通入口,不动旧文件、不把 targets-first 规则隐式套用,也不创建 R/ 或 _targets.R;修复沿现有顺序脚本入口完成并做隔离试跑;如需恢复/复用等复杂能力,只提出显式迁移方案。", "assertions": [ "没有引入或维护任何编号 runner/checkpoint 机制", "没有隐式迁移或删除用户旧文件", "只在授权范围内沿现有入口修复并测试" ], "files": [] }, { "id": 14, "prompt": "一个 complex 项目有 100 个相互独立的 cohort 模型,想用 crew 加速;每个模型还会启动 8 个 BLAS 线程。请给出安全配置和观测方案。", "expected_output": "crew 作为可选增强,先评估任务成本与 worker 开销;限制 worker×内部线程不超过资源预算,并用 tar_poll/tar_watch、crew 日志/metrics、autometric 组合观测,不自制 dashboard。", "assertions": [ "明确防止 worker 与 BLAS/OpenMP/future 嵌套 oversubscription", "crew/autometric 被描述为项目级可选依赖", "观测优先 targets 状态且不新增自定义事件协议" ], "files": [] } ] }
-
-
qa
-
test_code_ligature_rendering.py 908 B
from pathlib import Path import re import unittest SKILL_ROOT = Path(__file__).resolve().parents[1] THEME_CSS = SKILL_ROOT / "templates" / "liquid_glass_theme.css" class CodeLigatureRenderingTests(unittest.TestCase): def test_code_surfaces_disable_programming_ligatures(self) -> None: css = THEME_CSS.read_text(encoding="utf-8") match = re.search( r"pre,\s*\ncode\s*\{(?P<body>.*?)\n\}", css, re.DOTALL, ) self.assertIsNotNone( match, "pre/code rule must preserve source operators such as R's <-", ) body = match.group("body") self.assertRegex(body, r"font-variant-ligatures:\s*none\s*;") self.assertRegex(body, r'font-feature-settings:[^;]*"liga"\s+0') self.assertRegex(body, r'font-feature-settings:[^;]*"calt"\s+0') if __name__ == "__main__": unittest.main() -
test_dual_mode_integration.py 13.8 KB
"""Real lightweight integration checks for bensz-rmd-rules simple/complex modes.""" from __future__ import annotations import hashlib import os import shutil import subprocess import sys import tempfile import textwrap import unittest from pathlib import Path SKILL_ROOT = Path(__file__).resolve().parents[1] REPO_ROOT = SKILL_ROOT.parents[2] CHECKER = SKILL_ROOT / "scripts" / "check_targets_renv.py" THEME_BOOTSTRAP = SKILL_ROOT / "scripts" / "bootstrap_liquid_glass.py" def digest(path: Path) -> str: return hashlib.sha256(path.read_bytes()).hexdigest() def r_has(packages: list[str], *, require_pandoc: bool = False) -> bool: rscript = shutil.which("Rscript") if not rscript: return False checks = " && ".join(f'requireNamespace("{name}", quietly=TRUE)' for name in packages) if require_pandoc: checks = f"({checks}) && rmarkdown::pandoc_available()" result = subprocess.run( [rscript, "-e", f"quit(status=if ({checks}) 0 else 1)"], text=True, capture_output=True, check=False, env=os.environ.copy(), ) return result.returncode == 0 class DualModeRealExecution(unittest.TestCase): @classmethod def setUpClass(cls) -> None: cls.rscript = shutil.which("Rscript") parent = Path(os.environ.get("BENSZ_TEST_TMP", REPO_ROOT / "tmp")) parent.mkdir(parents=True, exist_ok=True) cls.temp_parent = parent def add_renv_contract(self, root: Path, packages: list[str]) -> None: result = subprocess.run( [ self.rscript, "-e", "a <- commandArgs(TRUE); p <- a[[1]]; pkgs <- unique(c('renv', a[-1])); sources <- .libPaths(); renv::init(project=p, bare=TRUE, restart=FALSE); renv::hydrate(project=p, packages=pkgs, sources=sources, prompt=FALSE, report=FALSE); renv::snapshot(project=p, packages=pkgs, prompt=FALSE)", str(root), *packages, ], text=True, capture_output=True, check=False, env=os.environ.copy(), ) self.assertEqual(result.returncode, 0, result.stdout + result.stderr) self.assertTrue((root / "renv" / "activate.R").is_file()) self.assertTrue((root / "renv.lock").is_file()) def assert_renv_synchronized(self, root: Path) -> None: result = subprocess.run( [ self.rscript, "-e", "a <- commandArgs(TRUE); status <- renv::status(project=a[[1]], sources=FALSE); quit(status=if (isTRUE(status$synchronized)) 0 else 1)", str(root), ], text=True, capture_output=True, check=False, env=os.environ.copy(), ) self.assertEqual(result.returncode, 0, result.stdout + result.stderr) def run_checker(self, root: Path, workflow_mode: str) -> subprocess.CompletedProcess[str]: return subprocess.run( [ sys.executable, str(CHECKER), str(root), "--project-state", "new", "--workflow-mode", workflow_mode, ], text=True, capture_output=True, check=False, ) def test_bootstrap_places_pipeline_assets(self): with tempfile.TemporaryDirectory(prefix="rmd-bootstrap-", dir=self.temp_parent) as tmp: root = Path(tmp) result = subprocess.run( [sys.executable, str(THEME_BOOTSTRAP), "--project-root", str(root), "--with-extras", "--with-pipeline"], text=True, capture_output=True, check=False, ) self.assertEqual(result.returncode, 0, result.stdout + result.stderr) self.assertTrue((root / "_targets.R").is_file()) self.assertTrue((root / "R" / "analysis_functions.R").is_file()) self.assertTrue((root / "templates" / "liquid_glass_theme.css").is_file()) @unittest.skipUnless(r_has(["renv", "rmarkdown"], require_pandoc=True), "R renv + rmarkdown + pandoc required") def test_simple_synthetic_fixture_renders_real_rmd_without_targets(self): with tempfile.TemporaryDirectory(prefix="rmd-simple-", dir=self.temp_parent) as tmp: root = Path(tmp) self.add_renv_contract(root, ["rmarkdown"]) (root / "raw").mkdir() raw_path = root / "raw" / "protected.csv" raw_path.write_text("id,group,value\nR1,A,10\n", encoding="utf-8") raw_before = digest(raw_path) report = root / "01.00.00. 分析报告.Rmd" report.write_text( textwrap.dedent( """\ --- title: "Simple integration" output: html_document --- ```{r} input <- Sys.getenv("BENSZ_ANALYSIS_INPUT") data <- utils::read.csv(input, check.names = FALSE) stopifnot(nrow(data) == 6L, !anyDuplicated(data$id)) stopifnot(setequal(unique(data$group), c("A", "B"))) summary(data$value) ``` """ ), encoding="utf-8", ) smoke = root / "scripts" / "tests" / "smoke_test.R" smoke.parent.mkdir(parents=True) shutil.copy2( SKILL_ROOT / "templates" / "tests" / "test_harness.R", smoke.parent / "test_harness.R", ) smoke.write_text( textwrap.dedent( """\ source(file.path("renv", "activate.R")) source(file.path("scripts", "tests", "test_harness.R")) Sys.setenv(BENSZ_TEST_RUN_ID = "simple-integration") run_id <- Sys.getenv("BENSZ_TEST_RUN_ID") context <- bensz_test_begin("new", "simple", "synthetic_fixture", "01.00.00. 分析报告.Rmd") set.seed(20260920L) fixture <- data.frame( id = sprintf("S%02d", 1:6), group = rep(c("A", "B"), each = 3L), value = c(1, NA, 3, 4, 9, 6) ) fixture_path <- file.path(context$input_dir, "synthetic_fixture.csv") utils::write.csv(fixture, fixture_path, row.names = FALSE, na = "") Sys.setenv(BENSZ_ANALYSIS_INPUT = normalizePath(fixture_path, mustWork = TRUE)) dir.create(context$reports_dir, recursive = TRUE, showWarnings = FALSE) output <- rmarkdown::render( "01.00.00. 分析报告.Rmd", output_dir = context$reports_dir, quiet = TRUE, envir = new.env(parent = globalenv()) ) stopifnot(file.exists(output), file.info(output)$size > 0) unlink(fixture_path) bensz_test_finish(context, assertions = c("synthetic schema passed", "Rmd rendered")) """ ), encoding="utf-8", ) self.assert_renv_synchronized(root) checked = self.run_checker(root, "simple") self.assertEqual(checked.returncode, 0, checked.stdout + checked.stderr) run = subprocess.run( [self.rscript, str(smoke)], cwd=root, text=True, capture_output=True, check=False ) self.assertEqual(run.returncode, 0, run.stdout + run.stderr) self.assertEqual(digest(raw_path), raw_before) run_root = root / "tmp" / "tests" / "simple-integration" rendered = list((run_root / "reports").glob("*.html")) self.assertEqual(len(rendered), 1) self.assertGreater(rendered[0].stat().st_size, 0) self.assertTrue((run_root / "run-record.md").is_file()) record = (run_root / "run-record.md").read_text(encoding="utf-8") self.assertIn("full_data_execution: NOT_RUN", record) self.assertIn("subject_check: PASS", record) self.assertIn("test_style: synthetic_fixture", record) self.assertFalse((run_root / "input" / "synthetic_fixture.csv").exists()) self.assertFalse((root / "_targets.R").exists()) self.assertFalse((root / "products").exists()) @unittest.skipUnless(r_has(["renv", "targets"]), "R renv + targets packages required") def test_complex_project_subset_runs_real_graph_in_isolated_store(self): with tempfile.TemporaryDirectory(prefix="rmd-complex-", dir=self.temp_parent) as tmp: root = Path(tmp) self.add_renv_contract(root, ["targets"]) (root / "R").mkdir() (root / "raw").mkdir() raw_path = root / "raw" / "input.csv" raw_path.write_text( "id,group,value\nA1,A,1\nA2,A,\nA3,A,7\nB1,B,2\nB2,B,9\nB3,B,3\n", encoding="utf-8", ) raw_before = digest(raw_path) (root / "_targets.R").write_text( textwrap.dedent( """\ source(file.path("renv", "activate.R")) library(targets) tar_source("R") input_path <- Sys.getenv("BENSZ_ANALYSIS_INPUT", unset = file.path("raw", "input.csv")) output_root <- Sys.getenv("BENSZ_PRODUCTS_DIR", unset = "products") list( tar_target(data, utils::read.csv(input_path, check.names = FALSE)), tar_target( summary_file, { dir.create(output_root, recursive = TRUE, showWarnings = FALSE) path <- file.path(output_root, "summary.txt") writeLines(paste("rows", nrow(data)), path) path }, format = "file" ) ) """ ), encoding="utf-8", ) (root / "R" / "pipeline_functions.R").write_text( "identity_data <- function(input_path) utils::read.csv(input_path, check.names = FALSE)\n", encoding="utf-8", ) smoke = root / "scripts" / "tests" / "smoke_test.R" smoke.parent.mkdir(parents=True) shutil.copy2( SKILL_ROOT / "templates" / "tests" / "test_harness.R", smoke.parent / "test_harness.R", ) smoke.write_text( textwrap.dedent( """\ source(file.path("renv", "activate.R")) source(file.path("scripts", "tests", "test_harness.R")) Sys.setenv(BENSZ_TEST_RUN_ID = "complex-integration") run_id <- Sys.getenv("BENSZ_TEST_RUN_ID") context <- bensz_test_begin("new", "complex", "project_subset", "_targets.R") source_path <- file.path("raw", "input.csv") raw_before <- unname(tools::md5sum(source_path)) full <- utils::read.csv(source_path, check.names = FALSE) ordered <- full[order(full$group, full$id), , drop = FALSE] pick <- !duplicated(ordered$group) | is.na(ordered$value) | ordered$value == max(ordered$value, na.rm = TRUE) subset_data <- ordered[pick, , drop = FALSE] stopifnot(setequal(unique(subset_data$group), c("A", "B")), anyNA(subset_data$value)) test_input <- file.path(context$input_dir, "project_subset.csv") utils::write.csv(subset_data, test_input, row.names = FALSE, na = "") Sys.setenv( BENSZ_ANALYSIS_INPUT = normalizePath(test_input, mustWork = TRUE) ) targets::tar_make(store = context$targets_store) built <- targets::tar_read(data, store = context$targets_store) stopifnot(nrow(built) == nrow(subset_data)) stopifnot(file.exists(file.path(context$products_dir, "summary.txt"))) stopifnot(identical(raw_before, unname(tools::md5sum(source_path)))) unlink(test_input) bensz_test_finish(context, assertions = c("representative subset passed", "target graph executed")) """ ), encoding="utf-8", ) self.assert_renv_synchronized(root) checked = self.run_checker(root, "complex") self.assertEqual(checked.returncode, 0, checked.stdout + checked.stderr) run = subprocess.run( [self.rscript, str(smoke)], cwd=root, text=True, capture_output=True, check=False ) self.assertEqual(run.returncode, 0, run.stdout + run.stderr) self.assertEqual(digest(raw_path), raw_before) run_root = root / "tmp" / "tests" / "complex-integration" self.assertTrue((run_root / "_targets").is_dir()) self.assertTrue((run_root / "run-record.md").is_file()) record = (run_root / "run-record.md").read_text(encoding="utf-8") self.assertIn("full_data_execution: NOT_RUN", record) self.assertIn("mutation_check: PASS", record) self.assertIn("test_style: project_subset", record) self.assertFalse((run_root / "input" / "project_subset.csv").exists()) self.assertFalse((root / "_targets").exists()) self.assertFalse((root / "products").exists()) self.assertFalse((root / "reports").exists()) if __name__ == "__main__": unittest.main() -
test_metric_explanation_protocol.py 1.7 KB
from __future__ import annotations import importlib.util import sys import unittest from pathlib import Path SKILL_ROOT = Path(__file__).resolve().parents[1] SCRIPT_PATH = SKILL_ROOT / "scripts" / "check_interpretation_quality.py" def load_checker(): spec = importlib.util.spec_from_file_location("check_interpretation_quality", SCRIPT_PATH) if spec is None or spec.loader is None: raise RuntimeError(f"Unable to load {SCRIPT_PATH}") module = importlib.util.module_from_spec(spec) sys.modules[spec.name] = module spec.loader.exec_module(module) return module class MetricExplanationProtocolTests(unittest.TestCase): @classmethod def setUpClass(cls) -> None: cls.checker = load_checker() def test_required_metric_explanation_is_not_flagged_as_teaching_tone(self) -> None: text = ( "该指标用于评估预测概率与实际发生率的一致程度," "反映模型的校准关系;当数值接近参考点时可以认为偏差较小。" ) result = self.checker.scan_text(text) self.assertEqual(result["teaching_hits"], []) def test_template_teaching_prompt_is_still_flagged(self) -> None: result = self.checker.scan_text("提示:这个图用于快速判断模型是否可靠。") self.assertTrue(result["teaching_hits"]) def test_protocol_reference_contains_first_and_later_mention_rules(self) -> None: text = (SKILL_ROOT / "references" / "metric_explanation_protocol.md").read_text( encoding="utf-8" ) self.assertIn("## 首次出现协议", text) self.assertIn("## 后续出现协议", text) self.assertIn("## 指标导读表", text) if __name__ == "__main__": unittest.main() -
test_targets_renv.py 7.5 KB
import json import subprocess import sys import tempfile import unittest from pathlib import Path ROOT = Path(__file__).resolve().parents[1] SCRIPT = ROOT / "scripts" / "check_targets_renv.py" class TargetsRenvChecks(unittest.TestCase): def run_check(self, root, *args): out = subprocess.run( [sys.executable, str(SCRIPT), str(root), *args], text=True, capture_output=True, check=False, ) return out.returncode, json.loads(out.stdout) @staticmethod def add_renv(root: Path) -> None: (root / "renv.lock").write_text("{}\n", encoding="utf-8") (root / "renv").mkdir() (root / "renv" / "activate.R").write_text("# renv\n", encoding="utf-8") @staticmethod def add_smoke(root: Path, content: str = "run_id <- Sys.getenv('BENSZ_TEST_RUN_ID')\nsource('analysis.R')\n") -> None: entry = root / "scripts" / "tests" / "smoke_test.R" entry.parent.mkdir(parents=True) entry.write_text(content, encoding="utf-8") def test_new_simple_requires_renv_and_test_entry_but_not_targets(self): with tempfile.TemporaryDirectory() as tmp: code, result = self.run_check( Path(tmp), "--project-state", "new", "--workflow-mode", "simple" ) self.assertEqual(code, 1) self.assertEqual( result["missing"], ["renv.lock", "renv/activate.R", "scripts/tests/smoke_test.R"], ) self.assertNotIn("_targets.R", result["missing"]) def test_new_simple_passes_without_targets(self): with tempfile.TemporaryDirectory() as tmp: root = Path(tmp) self.add_renv(root) self.add_smoke(root) before = sorted(path.relative_to(root) for path in root.rglob("*")) code, result = self.run_check( root, "--project-state", "new", "--workflow-mode", "simple" ) after = sorted(path.relative_to(root) for path in root.rglob("*")) self.assertEqual(code, 0) self.assertEqual(result["workflow_mode"], "simple") self.assertEqual(before, after) self.assertFalse((root / "_targets.R").exists()) def test_new_complex_requires_targets_and_isolated_test_store(self): with tempfile.TemporaryDirectory() as tmp: root = Path(tmp) self.add_renv(root) self.add_smoke(root, "targets::tar_make(store = '_targets')\n") code, result = self.run_check( root, "--project-state", "new", "--workflow-mode", "complex" ) self.assertEqual(code, 1) self.assertIn("_targets.R", result["missing"]) codes = {item["code"] for item in result["issues"]} self.assertIn("complex-test-store-unbound", codes) self.assertIn("complex-test-uses-formal-store", codes) def test_new_complex_passes_with_isolated_store(self): with tempfile.TemporaryDirectory() as tmp: root = Path(tmp) self.add_renv(root) (root / "R").mkdir() (root / "_targets.R").write_text('targets::tar_source("R")\n# targets\n', encoding="utf-8") self.add_smoke( root, 'run_id <- Sys.getenv("BENSZ_TEST_RUN_ID", unset = "smoke")\n' 'test_store <- file.path("tmp", "tests", run_id, "_targets")\n' "targets::tar_make(store = test_store)\n", ) code, result = self.run_check( root, "--project-state", "new", "--workflow-mode", "complex" ) self.assertEqual(code, 0) self.assertEqual(result["test_store"], "tmp/tests/<run-id>/_targets") def test_packaged_smoke_templates_delegate_unique_run_root_to_harness(self): for mode in ("simple", "complex"): with self.subTest(mode=mode), tempfile.TemporaryDirectory() as tmp: root = Path(tmp) self.add_renv(root) if mode == "complex": (root / "R").mkdir() (root / "_targets.R").write_text('targets::tar_source("R")\n# targets\n', encoding="utf-8") entry = root / "scripts" / "tests" / "smoke_test.R" entry.parent.mkdir(parents=True) entry.write_text( (ROOT / "templates" / "tests" / f"{mode}_smoke_test.R").read_text(encoding="utf-8"), encoding="utf-8", ) code, result = self.run_check( root, "--project-state", "new", "--workflow-mode", mode ) self.assertEqual(code, 0, result) def test_existing_targets_project_auto_resolves_complex(self): with tempfile.TemporaryDirectory() as tmp: root = Path(tmp) (root / "analysis.R").write_text("# existing\n", encoding="utf-8") self.add_renv(root) (root / "R").mkdir() (root / "_targets.R").write_text('targets::tar_source("R")\n# targets\n', encoding="utf-8") self.add_smoke( root, 'run_id <- Sys.getenv("BENSZ_TEST_RUN_ID", unset = "smoke")\n' 'test_store <- file.path("tmp", "tests", run_id, "_targets")\n' "targets::tar_make(store = test_store)\n", ) code, result = self.run_check(root, "--project-state", "existing", "--workflow-mode", "auto") self.assertEqual(code, 0) self.assertEqual(result["project_state"], "existing") self.assertEqual(result["workflow_mode"], "complex") self.assertEqual(result["missing"], []) self.assertEqual(result["issues"], []) self.assertTrue(result["observed_mechanisms"]["targets"]) def test_existing_without_targets_reports_mechanisms_as_warnings(self): with tempfile.TemporaryDirectory() as tmp: root = Path(tmp) (root / "analysis.R").write_text("# existing\n", encoding="utf-8") before = sorted(path.relative_to(root) for path in root.rglob("*")) code, result = self.run_check(root, "--project-state", "auto") after = sorted(path.relative_to(root) for path in root.rglob("*")) self.assertEqual(code, 0) self.assertEqual(result["project_state"], "existing") self.assertEqual(result["workflow_mode"], "simple") self.assertEqual(result["missing"], []) self.assertEqual(result["issues"], []) self.assertGreaterEqual(len(result["warnings"]), 2) self.assertFalse(result["observed_mechanisms"]["renv"]) self.assertFalse(result["observed_mechanisms"]["targets"]) self.assertEqual(before, after) def test_legacy_mode_alias_remains_compatible(self): with tempfile.TemporaryDirectory() as tmp: root = Path(tmp) (root / "analysis.R").write_text("# existing\n", encoding="utf-8") code, result = self.run_check(root, "--mode", "existing") self.assertEqual(code, 0) self.assertEqual(result["mode"], "existing") self.assertEqual(result["workflow_mode"], "simple") def test_disagreeing_state_arguments_fail_stably(self): with tempfile.TemporaryDirectory() as tmp: code, result = self.run_check( Path(tmp), "--mode", "new", "--project-state", "existing" ) self.assertEqual(code, 2) self.assertEqual(result["status"], "error") if __name__ == "__main__": unittest.main() -
test_toc_hover_stability.py 2.6 KB
from pathlib import Path import re import unittest SKILL_ROOT = Path(__file__).resolve().parents[1] THEME_CSS = SKILL_ROOT / "templates" / "liquid_glass_theme.css" class TocHoverStabilityTests(unittest.TestCase): def test_dynamic_toc_does_not_animate_hit_test_geometry(self) -> None: css = THEME_CSS.read_text(encoding="utf-8") match = re.search( r"html\.lg-toc-mode-dynamic\.lg-toc-layout #TOC \{(?P<body>.*?)\n \}", css, re.DOTALL, ) self.assertIsNotNone(match, "dynamic TOC collapsed rule is missing") body = match.group("body") transition = re.search(r"transition:\s*(?P<value>.*?);", body, re.DOTALL) self.assertIsNotNone(transition, "dynamic TOC transition is missing") self.assertNotRegex( transition.group("value"), r"\b(?:all|width|max-height|height|padding|border-radius)\b", "animating TOC hit-test geometry can close the menu under the pointer", ) def test_fast_pointer_move_into_toc_keeps_it_open(self) -> None: try: from playwright.sync_api import Error, sync_playwright except ImportError: self.skipTest("Playwright is not installed") css = THEME_CSS.read_text(encoding="utf-8") page_html = ( '<!doctype html><html class="lg-toc-layout lg-toc-mode-dynamic">' f"<head><style>{css}</style><style>.tocify {{ margin-top: 25px; max-width: 260px; }}</style></head>" '<body><div id="TOC" class="tocify"><div class="lg-toc-header">目录</div>' '<ul><li><a href="#section">Section</a></li></ul></div></body></html>' ) with sync_playwright() as playwright: try: browser = playwright.chromium.launch(headless=True) except Error as exc: self.skipTest(f"Chromium is unavailable: {exc.__class__.__name__}") try: page = browser.new_page(viewport={"width": 1600, "height": 900}) page.set_content(page_html) page.mouse.move(600, 800) page.mouse.move(48, 73) page.wait_for_timeout(40) page.mouse.move(180, 75) page.wait_for_timeout(320) self.assertTrue( page.locator("#TOC").evaluate("element => element.matches(':hover')"), "moving into the TOC during expansion must not close it", ) finally: browser.close() if __name__ == "__main__": unittest.main() -
test_windows_cli_and_dt_helper.py 1.8 KB
from __future__ import annotations import json import subprocess import sys import tempfile import unittest from pathlib import Path SKILL_ROOT = Path(__file__).resolve().parents[1] FORBIDDEN_STATUS_GLYPHS = set("✓✗⚠✅❌📊🔍🚀⊙▶") class WindowsCliAndDtHelperTests(unittest.TestCase): def test_executable_sources_use_gbk_safe_status_prefixes(self) -> None: paths = list((SKILL_ROOT / "scripts").glob("*.py")) paths += list((SKILL_ROOT / "scripts").glob("*.R")) paths += list((SKILL_ROOT / "templates").glob("*.R")) violations = [] for path in paths: text = path.read_text(encoding="utf-8") found = sorted(FORBIDDEN_STATUS_GLYPHS.intersection(text)) if found: violations.append(f"{path.name}: {''.join(found)}") self.assertEqual(violations, []) def test_recommended_dt_helper_is_counted_as_table_output(self) -> None: with tempfile.TemporaryDirectory() as tmpdir: rmd = Path(tmpdir) / "helper.Rmd" rmd.write_text( "```{r table}\nrender_dt_output(data.frame(x = 1))\n```\n", encoding="utf-8", ) result = subprocess.run( [ sys.executable, str(SKILL_ROOT / "scripts" / "check_figure_table_interpretation.py"), str(rmd), "--json", ], check=False, capture_output=True, text=True, encoding="utf-8", ) report = json.loads(result.stdout) self.assertEqual(report["total_outputs"], 1) self.assertEqual(report["unmatched_outputs"], 1) if __name__ == "__main__": unittest.main()
-
-
references
-
analysis_workflow_cache.md 2.9 KB
# Pipeline 状态、科学产品与恢复契约 ## 1. 唯一计算事实来源 在 `complex/pipeline` 项目中,`_targets.R` 是唯一 DAG、依赖、失效和增量重建入口;`_targets/` 是 targets 自己管理的机器状态。`SUCCESS`、identity hash、metadata checkpoint 这套自制缓存/恢复体系已从本 Skill 移除,不得重新引入。 `analysis-plan.yaml` 只做需求映射、target/报告关系和验收记录,不是调度器,也没有任何编号 runner 承载它;迁移到 targets 必须显式授权。 ## 2. 科学产品与报告 `products/` 只保存需要人工审阅、下游复用或正式交付的科学对象,例如完整结果 RDS、矩形表、数据字典或结果摘要;并非每个 target 都要导出。`reports/` 保存 PDF、图、表、HTML 和补充材料。二者不承担 targets 的缓存命中、失效或恢复。 产品元数据可以记录科学说明、数据字典和交付来源,但不得成为第二个缓存判定协议,也不应写入凭据、绝对私有路径或大体积原始数据。 ## 3. R/ 与 Rmd 边界 计算函数放在项目 `R/`,由 `_targets.R` 的 `tar_source("R")` 发现并调用。target 产出完整、未按展示阈值截断的结果。Rmd 通过 `tar_read()`/`tar_load()` 消费结果;若使用 `tarchetypes::tar_render()`,报告 target 必须在 DAG 中明确声明且不能形成循环。 Top N、阈值、配色和版式是报告参数。只改报告参数时,应只重渲染报告或报告 target,不触发无关重型 target;Rmd 不得从 `raw/` 绕过 DAG 重做昂贵计算。 ## 4. 中断与恢复验收 恢复不是“文件存在”检查,而是真实运行证据: 1. 在隔离 `tmp/tests/<run-id>/_targets` store 中先让至少一个昂贵 target 成功; 2. 让后续 target 中断或失败并保留 store; 3. 在不改变输入、代码、参数和 renv 身份、且未显式强制重算的条件下再次执行 `targets::tar_make()`; 4. 从 `tar_meta()`、outdated target 集合和执行日志确认前序 target 被跳过,仅未完成或失效部分继续; 5. 删除 store、改变输入/代码/参数或显式请求重算时,确认 targets 重新计算必要节点。 恢复证据以 targets metadata 和实际执行记录为准,不手工补写标记,也不把 worker 日志当作 DAG 状态。 ## 5. 可选并行与观测 无 `crew` 时普通 `tar_make()` 必须正常运行。仅当独立昂贵 target 的收益覆盖 worker 启动和传输开销时,才在 `_targets.R` 中配置 crew controller;worker 内 BLAS/OpenMP/future/BiocParallel 线程数必须受资源预算约束。 进度优先使用 `tar_poll()` 或 `tar_watch()`;worker 日志/指标使用 crew 能力;资源诊断按需使用 `autometric::log_start()`、`log_read()` 和 `log_plot()`。Skill 不定义新的调度器、事件字段、心跳、资源采样器、dashboard 或告警平台;缺少可选依赖时记录观测降级。 -
candidate_r.md 2.4 KB
# candidate_r 函数库使用指南 ## 概述 `candidate_r` 是一个可选的独立分析脚本集合,包含特征选择、聚类、生存分析等函数。 ## 可选性说明 - `candidate_r` **不是必需的**,只是用户可能拥有的"自定义函数目录"的一种 - 如果用户没有 candidate_r,AI 应使用 lucky 系列、CRAN/Bioconductor 包或实现最小辅助函数 - candidate_r 典型场景:用户有一些常用的独立 R 脚本(如 `feature_selection.R`、`clustering.R`),但不方便打包成正式 R 包 ## 使用模式 ### 模式 1:已通过 00.Environment.R 加载(默认) 当用户已在 `00.Environment.R` 中加载 candidate_r 后,AI 可以直接使用这些函数: ```r # 先加载环境 source("00.Environment.R") # 轻量检查:函数是否已存在 fs <- get0("feature_selection", ifnotfound = NULL) if (is.function(fs)) { # 使用 candidate_r 中的函数 res <- fs(x, y, method = "lasso") } else { # 回退:使用 lucky/CRAN/or 实现最小 helper res <- glmnet::cv.glmnet(x, y, alpha = 1) } ``` **优点**: - 用户一次性配置(在 00.Environment.R),AI 无需猜测路径 - 轻量检查,不扫描磁盘,不假设文件结构 ### 模式 2:用户指定路径(可选) 用户可在环境设置代码块中定义 `candidate_r_path`,AI 按该路径加载: ```r # 在 00.Environment.R 或 Rmd 环境设置中 if (exists("candidate_r_path") && nzchar(candidate_r_path)) { f <- file.path(candidate_r_path, "feature_selection.R") if (!file.exists(f)) { stop("candidate_r script not found: ", f) } source(f) } ``` ## 安全提示 `source()` 只应指向**信任的本地代码**,不要对未知来源/下载来的脚本直接 `source()`。 ## 实践原则 1. **不猜路径**:AI 不应假设 candidate_r 的标准位置 2. **不扫描磁盘**:避免使用 `list.files()` 搜索 candidate_r 3. **优先现有资源**:lucky 系列不满足时,才考虑 candidate_r 4. **提供回退方案**:当 candidate_r 不可用时,使用 CRAN 包或最小实现 ## 资源使用优先级 1. **已安装的 R 包**(直接 `library()`) 2. **candidate_r 函数库**(已由 `00.Environment.R` 加载;或用户通过 `candidate_r_path` 指定路径后再 `source()`) 3. **CRAN/Bioconductor 包**(如确实需要新包) 4. **自定义函数**(最后选择,写入 `00.Environment.R`) -
code_block_explanations.md 3 KB
# 代码块前解释(分析决策叙述) 每个主要代码块前,建议写一段“分析决策叙述”(1 段、约 2-5 句),把“为什么这样做”说清楚,并为后续结果解读建立理解锚点。 硬性要求(内涵)必须覆盖: 1. 数据特征:本块使用的数据是什么?规模/分组/删失/缺失/分布等关键特征是什么? 2. 方法选择:为什么选这个方法/参数/阈值?它解决什么问题? 3. 输出解读:接下来会输出什么?关键指标怎么读?它将支持什么判断/决策? 形式要求(软约束):鼓励自然段落的专家叙述,避免机械分点(尤其避免“第一层/第二层/第三层”这样的标签化表述)。 ## 推荐写法模板(连贯叙述式) ````markdown ### 为什么选择 [方法/参数] 本部分使用 [方法名] 来评估/比较 [分析目标]。之所以这样做,是因为当前数据具有 [关键数据特征], 而 [方法/参数] 能在这些约束下给出 [你关心的输出/指标]。接下来我们将得到 [输出类型], 其中 [关键指标] 可以用来判断 [方向/效应/差异],并为后续的 [决策/验证/分层] 提供依据。 ```{r chunk-label} # code... ``` ```` ## 示例 1:生存分析(存在删失 + 分组比较) ````markdown ### 生存差异评估(为什么用 KM + Log-rank) 本部分用 Kaplan-Meier 曲线对比两组的总生存差异,因为当前结局存在删失观测且我们需要直观看到随时间变化的风险分离。 Log-rank 检验用于给出两组曲线差异的统计证据;同时我们会报告事件数与中位生存期,便于把“显著”翻译成可解释的效应量级。 接下来的输出将用于判断分组是否具有稳定的预后分层价值,并为后续 Cox 模型/分层验证提供起点。 ```{r survival-analysis} # ... ``` ```` ## 示例 2:差异分析(阈值/多重校正 + Top 信号) ````markdown ### 差异基因筛选(为什么用 FDR + 效应阈值) 本部分在全基因层面进行差异分析,面对大量同时检验,我们使用 FDR(q 值)控制多重比较带来的假阳性风险; 同时结合效应大小阈值,是为了避免“统计显著但效应极小”的结果占据解释篇幅。输出将给出每个基因的 log2FC 与 q 值, 我们会按 q 值排序提取 Top 信号并报告其方向、效应与不确定性,作为后续富集分析与机制解释的证据锚点。 ```{r de-analysis} # ... ``` ```` ## 示例 3:回归建模(协变量控制 + 系数可解释性) ````markdown ### 多因素模型(为什么需要协变量) 本部分使用多因素回归评估 X 与 Y 的关联,因为在当前数据里年龄/性别/批次等因素可能同时影响 X 与 Y,构成混杂。 通过将这些协变量纳入模型,我们希望得到在控制混杂后的主要效应,并用系数/OR/HR(含 CI)解释方向与量级。 输出将用于判断关联是否稳健,并为敏感性分析(替代协变量/分层)提供可复现的比较基线。 ```{r multivariable-model} # ... ``` ```` -
code_style_guide.md 5.8 KB
# R 代码风格指南 本文档规范 R Markdown 分析脚本中的代码风格,确保代码简洁、易读、易维护。 ## 核心原则 | 原则 | 说明 | |------|------| | **管道优先** | 使用 `%>%` 或 `|>` 构建可读的数据处理流程 | | **向量化思维** | 优先使用向量化操作,避免遍历 + if 判断 | | **函数式编程** | 使用 `purrr::map()` 等函数式工具替代循环 | | **数据驱动** | 用查找表、join 等数据驱动方式替代硬编码逻辑 | ## 基础代码风格 ### 管道操作 ```r # 好的风格:简洁、直观 result <- data %>% filter(condition) %>% group_by(category) %>% summarise(mean_value = mean(value)) # 避免的风格:过度嵌套、难以理解 result <- summarise(group_by(filter(data, condition), category), mean_value = mean(value)) ``` ### 命名规范 - **变量名**:小写字母 + 下划线,描述性强(如 `survival_data`, `gene_expression`) - **函数名**:动词开头,清晰表达功能(如 `calculate_survival()`, `plot_heatmap()`) - **中间函数**:以 `.` 开头,个性化命名(如 `.cf01_prepare_data()`) ## 减少 if 语句使用 **核心思想**:if 语句增加代码复杂度,优先使用向量化操作、内置函数和数据驱动方式替代。 ### 场景对照表 | 场景 | 推荐做法 | 避免 | |------|----------|------| | 条件赋值 | `dplyr::case_when()` / `ifelse()` / `dplyr::coalesce()` | 嵌套 `if-else` | | 分类映射 | `purrr::map()` / `dplyr::left_join()` | 手动 `if` 判断 | | 数值替换 | `dplyr::recode()` / `dplyr::na_if()` | `if (x == "a") x <- "b"` | | 存在性检查 | `any()` / `all()` / `%in%` | 遍历 + `if` 判断 | | 多分支逻辑 | 向量化函数 / 查找表 | 多层 `if-else if-else` | ### 条件赋值示例 ```r # 好的做法:向量化 + 数据驱动 data$category <- case_when( data$value > 90 ~ "high", data$value > 50 ~ "medium", TRUE ~ "low" ) # 避免:命令式 if 判断 data$category <- NA for (i in seq_len(nrow(data))) { if (data$value[i] > 90) { data$category[i] <- "high" } else if (data$value[i] > 50) { data$category[i] <- "medium" } else { data$category[i] <- "low" } } ``` ### 数值替换示例 ```r # 好的做法:向量化替换 data$status <- dplyr::recode(data$old_status, "a" = "active", "i" = "inactive", "u" = "unknown" ) # 避免:逐个 if 判断 data$status <- NA for (i in seq_len(nrow(data))) { if (data$old_status[i] == "a") { data$status[i] <- "active" } else if (data$old_status[i] == "i") { data$status[i] <- "inactive" } else if (data$old_status[i] == "u") { data$status[i] <- "unknown" } } ``` ### 分类映射示例 ```r # 好的做法:数据驱动(查找表) lookup <- tibble( code = c("A", "B", "C"), label = c("Type A", "Type B", "Type C") ) data <- left_join(data, lookup, by = c("category_code" = "code")) # 避免:手动 if 判断 data$category_label <- NA data$category_label[data$category_code == "A"] <- "Type A" data$category_label[data$category_code == "B"] <- "Type B" data$category_label[data$category_code == "C"] <- "Type C" ``` ## 边界规则 ### if 语句使用边界 **允许**: - 控制流程(如提前返回、参数验证) - 无法向量化的场景 - 函数内部的逻辑判断(非数据操作) **禁止**: - 遍历数据行做简单判断 - 用 if 做可用向量化替代的操作 ### 禁止防御性文件存在性检查 不要为"文件是否存在"写大量 if 语句。让代码在文件缺失时自然报错,由调用方处理。 ```r # 好的做法:直接读取,缺失时报错 data <- readRDS(file.path("tmp", "analysis_results.rds")) # 避免:防御性 if 检查 if (file.exists(file.path("tmp", "analysis_results.rds"))) { data <- readRDS(file.path("tmp", "analysis_results.rds")) } else { stop("文件不存在") # 与不写 if 效果相同,但更冗长 } ``` **理由**: - 文件缺失是异常情况,应通过报错暴露,而非静默吞掉 - 过多 if 检查增加维护负担,且容易掩盖真实问题 - R 的 `stop()` 本身就是处理异常的标准机制 ## 代码注释规范 ### 代码块头部 必须说明目的/输入/参数/输出,帮助人类快速理解上下文: ```r # 目的:一句话概括代码块目的 # 输入:关键输入来源或变量 # 参数:来自 YAML params / 全局变量的关键参数 # 输出:关键变量或输出文件 ``` ### 分步注释 用 `Step` 注释拆分流程: ```r # Step 1: 数据加载 # Step 2: 数据清洗 # Step 3: 统计分析 ``` ### 业务逻辑注释 说明筛选或判断的业务含义: ```r # 规则:说明筛选或判断的业务含义 # 回退策略:说明缺失数据时的处理方式 ``` ### 注释原则 - 代码块头部必须说明目的/输入/参数/输出 - 关键步骤用 `Step` 注释拆分流程 - 复杂逻辑、正则、魔法数、关键中间函数要补一句"为何如此" - 注释保持简洁,建议不超过代码行数的 20% ## 常见反模式 ### 反模式 1:遍历 + if ```r # 避免 for (i in seq_len(nrow(data))) { if (data$value[i] > threshold) { data$flag[i] <- "high" } else { data$flag[i] <- "low" } } # 推荐 data$flag <- ifelse(data$value > threshold, "high", "low") ``` ### 反模式 2:嵌套 if-else ```r # 避免 if (condition1) { if (condition2) { result <- "A" } else { result <- "B" } } else { if (condition3) { result <- "C" } else { result <- "D" } } # 推荐 result <- case_when( condition1 & condition2 ~ "A", condition1 & !condition2 ~ "B", !condition1 & condition3 ~ "C", TRUE ~ "D" ) ``` ### 反模式 3:防御性文件检查 ```r # 避免 if (file.exists(path)) { data <- readRDS(path) } else { stop("File not found") } # 推荐 data <- readRDS(path) # 缺失时自动报错 ``` -
cross_platform.md 2.6 KB
# 跨平台路径与文件 I/O ## 新项目边界 从项目根目录构造相对路径,并用 `file.path()` 适配 Windows、macOS 与 Linux: ```r input_path <- file.path("raw", "expression.tsv") product_path <- file.path(products_dir, "analysis_results.rds") figure_path <- file.path("reports", "figures", "02.00.00. 主要结果.pdf") ``` - `raw/` 只读;不得把清洗结果或缓存写回。 - 完整可恢复数据写 `products/`,正式图表/表格写 `reports/`。 - Rmd 与同名 HTML 位于项目根目录。 - AI 日志、预览和检查结果写当前 `.bensz-api/task-*`,由宿主传入任务根目录。 R/ 函数名、target 名称和报告文件名必须兼容 Windows:不得包含 `< > : " / \\ | ? *` 或控制字符,不得以点/空格结尾,也不得使用 `CON`、`PRN`、`AUX`、`NUL`、`COM1`–`COM9`、`LPT1`–`LPT9` 等设备保留名。 不要手写 `/` 或 `\\` 拼接路径,也不要硬编码用户名、盘符、`/tmp` 或本机绝对路径。 ## 读取与写入 ```r if (!file.exists(input_path)) stop("Missing raw input: ", input_path) data <- utils::read.delim(input_path, check.names = FALSE, fileEncoding = "UTF-8") output_dir <- file.path("reports", "tables") dir.create(output_dir, recursive = TRUE, showWarnings = FALSE) utils::write.table( data, file.path(output_dir, "02.00.00. 主要结果.tsv"), sep = "\t", row.names = FALSE, quote = TRUE, fileEncoding = "UTF-8" ) ``` 新 pipeline 不自行实现通用缓存/完成标记;由 targets 管理 `_targets/` 状态。自制 checkpoint/完成标记体系已从本 Skill 移除,不再提供、维护或校验。 ## 项目根与 Skill 根 用户分析代码以项目根为当前目录。Skill 自带脚本和资产必须从 Skill 文件自身位置解析,文档命令用 `<skill-root>/scripts/...` 表示,不假设用户当前目录是 Skill 根。 验证项目路径: ```bash Rscript <skill-root>/scripts/validate_paths.R /path/to/project ``` ## 检查清单 - [ ] 路径由 `file.path()` 构造,且从项目根或已解析的 Skill 根开始。 - [ ] `raw/` 只有读取,没有写入、移动、覆盖或删除。 - [ ] `products/`、`reports/`、`.bensz-api/task-*` 职责没有混用。 - [ ] 文件名大小写一致,不依赖 Windows 的大小写不敏感行为。 - [ ] 文本显式使用 UTF-8;Windows 控制台状态前缀使用 ASCII。 - [ ] 元数据和日志只记录相对路径,不泄露用户名、盘符或绝对私有路径。 - [ ] `tmp/{主脚本名}/` 产品路径模式已移除,科学产物只进 `products/`、正式材料只进 `reports/`。 -
delivery_verification.md 2.7 KB
# 交付验证记录 交付摘要只记录可复核证据,不复制敏感数据或绝对私有路径。 | 检查项 | 结论 | 证据 | | --- | --- | --- | | 项目状态与模式 | PASS/FAIL | `project_state`、`workflow_mode`、选择理由与人工覆盖 | | renv | PASS/FAIL/NA | lockfile、activation、`renv::status()`;existing 可为风险披露 | | 需求—target—报告映射 | PASS/FAIL/NA | targets 清单或 `analysis-plan.yaml` | | 目录与产品路径 | PASS/FAIL | raw/R/_targets/products/reports/scripts/templates/tmp 边界与单一路径设置 | | 测试风格 | PASS/FAIL | `synthetic_fixture` 或 `project_subset` 及选择理由 | | 真实轻量运行 | PASS/FAIL/BLOCKED | 正式入口、唯一 run root、命令、退出码;不得用 dry-run 替代 | | 测试隔离 | PASS/FAIL | raw/正式 products/reports/_targets 前后不变;test store 路径 | | 数据与科学断言 | PASS/FAIL | schema、主键/分组、范围/不变量、边缘条件 | | 主要估计对象推断 | PASS/FAIL/NA/NOT_RUN | 逐项列参数、估计值/CI/有效 N/方法及适用的 p/q,或明确 `not_applicable`、`not_estimable`、`not_run` 和理由 | | 方法与报告对应 | PASS/FAIL/NA/NOT_RUN | 计算代码/完整结果表/报告位置;配对、聚类、删失、缺失、模型假设和多重比较复核;文本检查不作方法证明 | | 预期交付生成 | PASS/FAIL | 产品、报告、图表可读性与数字追溯 | | targets 恢复 | PASS/FAIL/NA | `tar_meta()`、outdated 集合、store 损坏/失效/中断恢复用例 | | 串行审查 | PASS/FAIL/DEGRADED | 三项只读结果与修正摘要 | | 旧项目兼容 | PASS/FAIL/NA | 未自动补机制、迁移或覆盖旧路径 | ## 尝试记录 每次执行追加一行: | 尝试 | 命令 | 退出状态 | 失败类别 | 修正摘要 | 断言 | 剩余风险 | | --- | --- | --- | --- | --- | --- | --- | | 1 | | | environment_dependency / fixture_or_subset / path_isolation / analysis_code / scientific_assertion / report_rendering / external_service / none | | | | ## 运行摘要 - 正式入口与测试入口: - simple R/Rmd 或 complex target 实际执行证据: - test store / 临时输出: - 正式路径前后不变性: - subject identity 与环境来源: - preflight / lightweight_execution / full_data_execution(未跑全量为 NOT_RUN): - 逐参数推断状态(已计算并验证 / 不适用 / 受数据或设计限制 / 未运行)与报告数字来源: - project subset cleanup: - 未运行、阻塞或无法确认的项目: - 剩余风险与恢复入口: 宿主无独立子 Agent 时写 `DEGRADED: 三项职责分离顺序自检,非独立审查`;不得写 PASS 冒充。 -
expert_discussion_template.md 7.3 KB
# 专家级讨论模板 **来源**:基于 PanTCGA Mutation Diff-Variance 分析的真实讨论部分 **用途**:展示"有理有据、有启发性"的专家级讨论写作标准 --- ## 模板结构 ```markdown ## 讨论与分析 本次 [分析主题] 在"[宏观层面]"给出了相对清晰的总体图景:在 [样本量描述] 个可评估的 [分析对象] 中,有 [N] 个在阈值(q≤[阈值] 且 p≤[阈值])下达到统计学显著(约 [百分比])。从 [排序/分布] 看,[Top 信号的具体描述],且这些 top [对象] 的 [总体方向/特征] 以 [方向] 为主([具体统计描述:如 effect<0:X个,effect>0:Y个])。这意味着在当前数据与建模口径下,[关键发现的核心解释]。需要强调的是,[总体层面] effect 的量级并不大(多在 [量级范围] 左右),更像是一种"[定性描述:弱但稳定/偏单侧]"的总体倾向,而非单一强驱动因素。 然而,当把同一批 top [对象] 投影到 [细分维度:如分癌种/亚组] 后,一个更重要的信息是"[方向与强度的显著异质性/跨层级的差异模式]"。即便在 [总体层面] 显著的 [对象],在不同 [细分单元] 内的 [指标] 也可能出现 [方向翻转/效应放大或减弱/缺失],从而提示其更可能代表"[机制解释:某些组织背景下的选择压力/进化轨迹/亚群特异性]",而不是可直接泛化的普适标志物。以代表性 [对象] 为例,[对象A] 的跨 [细分单元] 方向一致性并不高(在部分 [单元] 出现与 [总体] 方向相反的 [指标];示例 [单元]:[具体列表]),[对象B] 也呈现类似的强异质性(示例 [单元]:[具体列表]);即使是 [对象C] 这样在 [总体] 层面最稳健的信号之一,在部分 [单元] 中也可观察到方向/强度的明显差异(示例 [单元]:[具体列表])。因此,本分析更稳妥的结论是:[总体] 的排序可用于"给出候选清单与总体方向",但真正可用于机制解释或临床分层的证据,仍应回到目标 [细分单元] 的 [数据文件](并结合 [关键维度:如效应大小、显著性与样本量])逐一核查。 在 [更高层级:如通路/模块/网络] 层面,本次结果呈现出"[总体有主题线索、但跨细分单元共性很弱/总体方向一致、但细分单元差异显著]"的格局。首先,[总体] 的 [分析方法:如 GSEA] 在当前阈值下仅有 [N] 条 [对象:如通路] 达到显著(Top:[Top 信号描述]),且从 [总体] 的 [可视化视角:如网络/热图] 来看,[可视化对象的统计描述:如 nodes=N、edges=M、modules=K],并呈现极强的 [方向特征:如负向主导/正向主导]([具体统计:NES<0:N个,NES>0:M个])。这与 [微观层面] 的"[方向特征]趋势"在方向上是一致的,提示与 [表型/状态] 相关的生物学过程在 [总体] 层面更容易被捕捉到;但与此同时,跨 [细分单元] 的复现性并不理想:在进入 [共性统计方法:如 commonality] 统计的 [N] 个 [细分单元] 中,共性 [对象] 仅 [N'] 条(≥2 [细分单元] 显著的仅 [N''] 条,且 [统计量] 最大仅 [值])。[模块/网络] 层面的跨 [细分单元]"出现"同样稀疏([统计描述]),并且 [可视化对象:如 heatmap] 里有相当比例的格子为 NA(约 [百分比]),反映出在多数 [细分单元] 中,当前阈值下"[模块内/网络中]没有任何显著 [成员]"。这一组现象更支持这样一种解释:[暴露/自变量] → [因变量/表型] 的关联在不同 [细分维度] 背景下驱动机制并不统一(或统计功效不足以在 [更高层级] 层面稳定复现),因此用"跨 [细分单元] 共性 [对象]"去讲强 [泛化结论] 机制故事并不稳健。 从机制推断角度,更可取的写作路径是把 [网络/模块/聚类] 的 [主题单元] 当作"主题锚点",而不是把 [共性统计方法:如 commonality] 当作"[泛化结论]"。本次 [总体] 网络模块的 [核心种子:seeds]([Top 5 描述])显示,显著 [对象:如通路] 主要围绕 [主题 1]、[主题 2] 以及 [主题 3] 等 [主题类别] 聚集(以 [分类方式:如 GO terms/Reactome] 为主,具备一定的"[广谱调控/特定功能]"特征)。这类主题本身容易与 [生物学过程:如肿瘤进化、分化状态、微环境互作] 等宏观表型相耦合,但在不同 [细分单元] 中的"触发方式"可能不同:有的 [单元] 由 [驱动因素 A] 的 [特征:如谱系差异] 驱动,有的 [单元] 可能更依赖 [驱动因素 B] 或 [因素 C] 的共同变化。因此,建议将本报告的结论框定为"[分析视角/限定条件] 下,[总体] 可观察到一组与 [表型/状态] 相关的 [信号特征],并在网络层面形成若干主题模块;但跨 [细分单元] 共性有限,解释与验证应以目标 [细分单元] 为单位进行"。 最后,从转化与可复现性角度,本次结果对后续分析给出两点直接启示:其一,若研究目标是"[泛化目标:如泛癌可复用/跨队列稳定]"的标志物或机制,应优先采用能提升跨 [细分单元] 稳定性的策略(例如在网络层面聚合主题、降低噪声、或对 [细分单元] 间异质性做分层/随机效应建模),而不是依赖单条 [对象:如通路/基因] 在多个 [细分单元] 同时达阈值;其二,若目标是具体 [细分单元] 的分层或靶点线索,则应把重点放在 [细分单元] 展开部分,优先筛选该 [单元] 内 [指标:如 effect 大且显著] 的 [信号源],并用该 [单元] 的 [高层级分析:如 GSEA/网络模块] 去组织机制叙事,再结合独立队列(或多组学证据)做验证与加固。 ``` --- ## 核心特征 | 特征 | 说明 | 示例 | |------|------|------| | **量化陈述** | 每个论断都配有具体数字 | "约 35.2% 的格子为 NA" | | **动态数值** | 使用 `` `r ...` `` 嵌入变量 | `r sprintf("%.1f%%", 100 * modules_prop_na)` | | **层次递进** | 从现象→机制→写作建议 | 先描述异质性,再解释机制,最后给出建议 | | **辩证思考** | 不回避局限,指出边界 | "更稳妥的结论是...而非..." | | **具体示例** | 用具体基因/通路举例 | BRAF、TTN、TP53 的跨癌种差异 | | **可操作建议** | 给出明确的后续方向 | "应优先采用能提升跨癌种稳定性的策略..." | | **启发性** | 不仅总结结果,还启发后续研究 | "从转化与可复现性角度,本次结果给出两点直接启示..." | --- ## 避免的浅层讨论 - ❌ "结果显示有显著差异,建议进一步研究"(空泛) - ❌ "发现了若干重要基因,值得深入探讨"(无具体指向) - ❌ "分析揭示了生物学机制,为临床应用提供依据"(套话) ## 推荐的深度讨论 - ✅ "Top 信号主要集中在经典 driver 基因附近(Top 8:[具体列表]),且方向以负向为主(effect<0:X个,effect>0:Y个)"(量化) - ✅ "同一基因在不同癌种的效应方向可翻转(如 BRAF 在 [具体癌种] 为正,在 [具体癌种] 为负),提示癌种特异性调控机制"(辩证) - ✅ "若研究目标是泛癌可复用标志物,应优先在网络层面聚合主题,而非依赖单条通路在多个癌种同时达阈值"(可操作) -
figure_interpretation_criteria.md 2.3 KB
# 图表/表格解读覆盖判定标准(硬编码) 本文件定义 `scripts/check_figure_table_interpretation.py` 的“覆盖检验”口径:**每个可见输出(图/表)都必须有对应的解读文本**。 ## 检验目标 - 解决“有图/表、无解读”的漏项(Coverage)。 - 不替代“解读质量检查”(Quality),后者由 `scripts/check_interpretation_quality.py` 提供启发式提示。 ## 什么算作“需要解读的输出” 脚本通过静态模式识别以下 R 代码块为“输出块”(可在 `config.yaml:figure_interpretation_check.check_patterns` 调整): - 图:`ggplot(...)` / `plot(...)` / `Heatmap(...)` / `pheatmap(...)` / `ggsave(...)` / `pdf(...)` / `png(...)` 等 - 表:`knitr::kable(...)` / `DT::datatable(...)` / `gt::gt(...)` / `flextable::flextable(...)` 等 以下情况默认**不纳入覆盖检验**(视作“不可见/不需要解读”): - chunk 设置了 `eval=FALSE` 或 `include=FALSE` - chunk 设置了 `results='hide'`(通常隐藏打印输出) - chunk 设置了 `interp_check=FALSE`(手动豁免:用于“命中输出模式但确实不产生可见输出”的少数情况) ## 什么算作“有解读” 对每个“输出块”,其后(默认 50 行内,见 `config.yaml:figure_interpretation_check.max_distance_lines`)必须出现 Markdown 正文(非代码块),并满足: 1. **长度门槛**(避免一句话套话) - 含中文:至少 100 个中文字符(默认,可用 CLI 参数覆盖) - 纯英文:至少 50 个英文单词(默认,可用 CLI 参数覆盖) 2. **解释性证据元素**(默认至少 2 个) 典型元素:展示/差异/显著性(p/FDR/q)/提示与推论等。 3. **四层解读标记(推荐)** 默认要求正文出现“数据描述/统计见解/领域见解/局限”等关键词之一(可在 `config.yaml:figure_interpretation_check.check_patterns.interpretation_markers` 调整)。 可选更严格模式:强制要求正文显式引用 “Figure/图/Table/表 N”(见脚本 `--require-reference` 或 `config.yaml:figure_interpretation_check.require_reference`)。 ## 运行方式(交付前强制) ```bash python3 bensz-rmd-rules/scripts/check_figure_table_interpretation.py your_report.Rmd --strict python3 bensz-rmd-rules/scripts/check_interpretation_quality.py your_report.Rmd ``` -
four_tier_interpretation_framework.md 4.9 KB
# 四层解读框架(证据锚定) 用于图、表和统计量的结果解读,目标是让每个判断都能回到当前结果,而非套话。弱背景读者先按 `metric_explanation_protocol.md` 解释自定义指标;首次完整说明定义、原理、选用理由和判读方法,后续只保留证据。 ## 四层与写法 | 层 | 必答问题 | 最低内容 | |---|---|---| | 数据描述 | 是什么? | N、阈值、输出结构、通过/未通过计数 | | 统计见解 | 统计上意味着什么? | 关键估计对象的方向、效应大小、适当 CI;有明确假设时的 p/q,或不适用/不可估计理由 | | 领域见解 | 为什么重要? | 绑定本次对象的机制/实践映射和可验证路径 | | 局限与后续 | 下一步怎么办? | 至少 2 条“方法 + 输入 + 判据”的验证动作 | 四层是内容门槛,不是固定四段。默认写成 1–2 段连贯叙述;多图、多组或需逐项审查时才分项。推荐骨架: ```markdown 本次 [分析] 在 [有效 N/条件/阈值] 下评估了 [估计对象与对比]。结果显示 [估计值、单位、CI 与水平;预定检验存在时给 p/q 和方法](用 `r ...` 引用完整结果)。这对 [本次变量/分层] 的领域意义是 [实际量级 + 怎么用/验证]。若推断不适用或不可估计,写明状态、原因与结论限制;后续用 [方法] 在 [输入] 上检验,判据为 [支持/不支持条件]。 ``` 每个 cohort/亚组可附“核心结论”:Top 1–3、证据强度(强/中/弱)、不确定性来源和可执行后续;阈值与判据来自 YAML `params`,不得在正文硬编码。 ## Fail Fast Gate - [ ] 每个关键推断判断能定位到对应估计值、CI/置信水平、有效 N、方法;有预定零假设时定位到 p/q,不能推断时定位到状态与原因。 - [ ] 优先用 `` `r ...` `` 动态嵌入关键数值,不用“较高/明显”等词代替量化证据。 - [ ] 点名本次最强 1–3 个信号,给对象和证据。 - [ ] 至少 3 句当前数据观察(对象 + 对比/方向 + 数值)。 - [ ] 其它描述性或探索性观察至少有一类适用的不确定性/稳定性证据;关键推断结果按上方逐项要求核对。 - [ ] 领域层回答“为什么重要 + 怎么用/怎么验证”,并落到本次变量/分层。 - [ ] 至少 2 条后续动作;每条包含方法、输入、判据。 - [ ] 不写无对象、无证据、无动作或无判据的套话。 ### 反套话替换 | 空泛句式 | 必须补齐 | |---|---| | “提示/可能与…相关” | 对象 + 指标/数值 + 对比 + 不确定性 | | “值得进一步研究/验证” | 方法 + 输入 + 判据 | | “具有临床意义/可解释” | 本次变量/分层 + 决策用途 + 验证方式 | | “可能存在混杂/非线性” | 受影响对象 + 检验方法 + 判据 | ## 数值与图表原则 AI 读 Rmd 图表时优先检查生成代码和原始数据,无需把渲染图当作定量来源。解读中的数值用内联 R 动态生成,例如: ```markdown 缺失比例为 `r sprintf("%.1f%%", 100 * prop_na)`;当前阈值下模块内没有显著通路成员。 ``` 每个数值都应能追溯到变量/函数:统计量来自模型摘要,N 来自维度或计数,比例来自计算逻辑,阈值来自 `params`。避免用代码拼接整句解释文本。 ## 常用句式 - **Top 信号**:`X` 是最强信号之一(指标 = 值,排名 = 位次/N),方向为 …;关键效应与 `Y` 的差异为 …(适当 CI = …;预定检验适用时 p/q = …),若不可估计则写原因。 - **阈值**:按 `params$...` 的阈值 `T`,该效应对应 … 的变化,适用边界为 …。 - **验证**:不确定性来自 …;用 … 在 … 上验证,支持条件为 …。 ## 按输出类型的最小模板 ### 表格 ```markdown 该表含 [N] 行 × [M] 列,展示 [内容],阈值为 [params 值],入围 [n] 个。关键 Top 结果为 [对象/方向/效应/适当 CI;预定检验适用时 p/q];若不可推断则列状态与原因。领域意义是 [映射]。局限为 [来源];后续:1) [方法 + 输入 + 判据];2) [方法 + 输入 + 判据]。 ``` ### 图表 ```markdown [图表类型]基于 [N] 个样本展示 [关系]。主要趋势/异常为 [带数值证据],效应和不确定性为 [值];替代解释是 [对象]。可通过 [方法] 在 [输入] 上检验,判据为 [条件]。 ``` ### 模型与验证 报告样本/事件数、模型/验证方案、Top 1–3 系数或性能(含 CI/校准/稳定性)、适用阈值和边界;后续至少包括假设诊断与外部/重采样验证,各写方法、输入和判据。具体骨架见 `interpretation_templates.md`。 ## 预检 ```bash python3 bensz-rmd-rules/scripts/check_interpretation_quality.py your_report.Rmd ``` 该启发式检查内联代码、当前观察句、Top 信号、不确定性和常见套话;不能替代语义审查。 -
gene_id_guidelines.md 4.8 KB
# 基因 ID 优先级与可读性指南 ## 适用场景 生物信息学、医学统计、临床数据分析等涉及基因标识符的 R Markdown 分析项目。 ## 核心原则 在解读、做表、画图时,**优先使用 SYMBOL 类的基因 ID**,而非 ENSEMBL ID 或 ENTREZID。 ## 原因 | ID 类型 | 可读性 | 典型读者认知 | |---------|--------|-------------| | **SYMBOL**(如 `TP53`、`EGFR`) | 高 | 大部分读者熟悉,直观易懂 | | ENSEMBL ID(如 `ENSG00000141510`) | 低 | 仅生信专业人员熟悉 | | ENTREZID(如 `7157`) | 低 | 仅生信专业人员熟悉 | **核心目标**:增强可读性,让更多读者(临床医生、生物学家、学生等)能够理解分析结果。 ## 实践规范 ### 1. 可视化与展示(SYMBOL 优先) 在热图、火山图、箱线图、网络图等可视化中: ```r # 推荐:使用 SYMBOL 作为标签,轴标题使用英文 ggplot(data, aes(x = SYMBOL, y = expression)) + geom_point() + labs(x = "Gene Symbol", y = "Expression Level") # 避免:使用 ENSEMBL ID 或 ENTREZID ggplot(data, aes(x = ensembl_id, y = expression)) + geom_point() # 也避免:使用中文轴标题 ggplot(data, aes(x = SYMBOL, y = expression)) + geom_point() + labs(x = "基因", y = "表达量") ``` ### 2. 表格展示(SYMBOL 为主) 在 DT::datatable、kable 等表格渲染中: ```r # 推荐:SYMBOL 作为主列,其他 ID 作为辅助列 DT::datatable(data[, .(SYMBOL, ensembl_id, entrezid, p_value, log2FC)], options = list(scrollX = TRUE)) # 列顺序:SYMBOL 在前,便于读者快速识别 ``` ### 3. 文本解读(使用 SYMBOL) 在 Rmd 的文本解读中: ````markdown ```r # 推荐:使用 SYMBOL 基因名 TP53 在肺癌中显著高表达(log2FC = 2.3, p < 0.001) # 避免:使用 ENSEMBL ID ENSG00000141510 在肺癌中显著高表达(log2FC = 2.3, p < 0.001) ``` ```` ### 4. 数据保存(多 ID 并存) **关键原则**:虽然展示时优先使用 SYMBOL,但**保存的数据中应包含多种 ID**,以确保数据准确性和可追溯性。 ```r # 推荐:保存包含所有 ID 类型的完整数据 final_data <- data[, .( SYMBOL, ensembl_id, entrezid, expression, p_value, log2FC )] # 保存为 CSV/RDS write.csv(final_data, "results/gene_expression.csv", row.names = FALSE) out_dir <- file.path("tmp", "analysis") dir.create(out_dir, recursive = TRUE, showWarnings = FALSE) saveRDS(final_data, file.path(out_dir, "gene_expression.rds")) ``` **好处**: - **准确性**:ENSEMBL ID 和 ENTREZID 是稳定的数据库标识符,避免 SYMBOL 歧义 - **可追溯性**:后续可通过其他 ID 类型回溯到原始数据库 - **灵活性**:可根据需要切换展示的 ID 类型 ## 基因 ID 转换 ### 常用转换函数 ```r # 强制:使用 luckyBase::convert()(本 skill 的统一接口) # ENSEMBL -> SYMBOL gene_symbols <- luckyBase::convert( ids = gene_ensembl_ids, from_type = "ENSEMBL", to_type = "SYMBOL", organism = "human" ) # ENTREZID -> SYMBOL gene_symbols2 <- luckyBase::convert( ids = gene_entrez_ids, from_type = "ENTREZID", to_type = "SYMBOL", organism = "human" ) ``` ### 在数据处理流程中添加 SYMBOL ```r # 在 .R 脚本的数据处理阶段,确保添加 SYMBOL 列(示例) .pp_add_symbol_column <- function(data, id_col = "ensembl_id", from_type = "ENSEMBL", organism = "human") { ids <- data[[id_col]] data$SYMBOL <- luckyBase::convert( ids = ids, from_type = from_type, to_type = "SYMBOL", organism = organism ) data } ``` ## 检查清单 在生成涉及基因的 R Markdown 分析时: - [ ] **可视化**:图表的轴标签、图例是否使用 SYMBOL? - [ ] **表格**:DT::datatable 的主列是否为 SYMBOL? - [ ] **文本解读**:解读中提到的基因名是否使用 SYMBOL? - [ ] **数据保存**:保存的数据是否包含 SYMBOL、ENSEMBL、ENTREZID 三列? - [ ] **ID 转换**:是否在数据处理阶段添加了 SYMBOL 列? - [ ] **可读性**:非生信专业背景的读者能否理解基因标识符? ## 常见问题 ### Q1:SYMBOL 有歧义怎么办? **A**:在数据保存时保留 ENSEMBL/ENTREZID 作为权威标识符,但在展示时使用 SYMBOL 并注明可能的歧义。 ### Q2:某些基因没有 SYMBOL 怎么办? **A**:使用 ENSEMBL ID 或 ENTREZID 作为备选,并在注释中说明。 ### Q3:需要展示大量基因时,SYMBOL 太长怎么办? **A**: - 使用缩写(如 `TP53` 而非 `tumor protein p53`) - 在图表中使用基因编号,在表格中提供完整 SYMBOL - 使用交互式表格(DT),支持搜索和滚动 ## 参考资源 - [HGNC 基因命名规范](https://www.genenames.org/) - [ENSEMBL 基因 ID](https://www.ensembl.org/info/genome/stable/index.html) - [NCBI Gene 数据库](https://www.ncbi.nlm.nih.gov/gene) -
htmlwidget_visibility_rules.md 2.4 KB
# HTML 可见性硬规则(htmlwidget/DT/plotly) 当 `.Rmd` 渲染为 HTML 时,`DT::datatable()` 等 **htmlwidget 必须作为 code chunk 的“可见结果”返回**,否则经常会出现“代码块在,但 HTML 不出表/不出图”的情况(尤其当 widget 被 `print()` / `invisible()` 包裹,或在 chunk 中间输出后又继续执行其它表达式)。 ## 硬规则 - 任何“需要在 HTML 中展示”的 htmlwidget(含 DT/plotly/leaflet 等),必须满足以下任意一种写法: - 方式 A(最推荐):让 widget 调用成为该 chunk 的最后一个表达式 - 方式 B(推荐):先赋值给变量,chunk 最后一行单独写该变量名(返回该对象) - 方式 C(多组件):用 `htmltools::tagList(...)` 作为 chunk 最后表达式,返回多个 widget - 禁止对“需要展示”的 htmlwidget 使用 `print()` / `invisible()` / `suppressMessages(print(...))` 等包裹写法。 - 如果一个 chunk 里要展示多个 widget,必须用 `htmltools::tagList(...)`(否则通常只会渲染最后一个)。 ## 常见错误示例 ```r print(DT::datatable(head(data, 100))) # 禁止:print() 包裹 widget ``` ```r DT::datatable(head(data, 100)) NULL # 禁止:widget 不是最后表达式 ``` ## 正确示例 ```r # 方式 A:最后表达式 DT::datatable(head(data, 100), options = list(scrollX = TRUE, pageLength = 10)) ``` ```r # 方式 B:先赋值,最后返回变量 tbl <- DT::datatable(head(data, 100), options = list(scrollX = TRUE, pageLength = 10)) tbl ``` ```r # 方式 C:多个 widget 一次性返回 tbl <- DT::datatable(head(data, 50)) fig <- plotly::ggplotly(ggplot2::qplot(1:10, 1:10)) htmltools::tagList(tbl, fig) ``` ## 标准化调用(推荐) 建议使用本 skill 的辅助函数统一渲染交互式表格(并确保“可见性规则”成立): ```r # 方式 1:直接使用 DT(推荐用于简单表格) DT::datatable(head(data, 100), options = list(scrollX = TRUE, pageLength = 10)) # 方式 2:使用辅助函数(推荐用于复杂场景) source(file.path("templates", "datatables_helper.R")) render_dt(data, n = 100) # 注意:必须作为 chunk 最后表达式(或先赋值,最后返回变量) # 如果你需要在同一 chunk 里继续写其它代码,但仍要确保表格可见, # 推荐改用 render_dt_output(...) 并把它放在 chunk 末尾(或用于 tagList 组合)。 render_dt_output(data, n = 100) ``` -
hybrid_architecture_examples.md 1.7 KB
# 双模式示例 ## simple:单报告 输入规模小、整体重跑可接受时,只创建 `00.Environment.R`、一个 Rmd、renv 和 smoke test。Rmd 直接读取授权的 `raw/`,不创建 `_targets.R`、products 缓存或自制 runner。 ## complex/pipeline:共享计算与多个报告 ```text R/analysis_functions.R _targets.R: analysis_input -> prepared_data -> model_results 01-report.Rmd: tar_read("model_results") -> reports/ 02-report.Rmd: tar_read("model_results") -> reports/ _targets/: targets metadata and computational state products/: only explicitly exported scientific objects ``` 报告参数变化只重渲染 Rmd;模型代码、输入或参数变化由 targets 失效必要节点。不要把报告层 Top N、配色或图表布局写进重型 target。 ## 中断恢复 在 `tmp/tests/<run-id>/_targets` 中让 `model_results` 上游成功、下游报告 target 失败,保留 store 后再次 `tar_make()`。使用 `tar_meta()` 与执行日志确认上游 target 被跳过、失败节点继续;不创建 SUCCESS 或人工 identity 文件。 ## 并行与观测 当 cohort/model/bootstrap target 足够独立且昂贵时,在 `_targets.R` 配置 crew controller;worker 数与内部线程数乘积必须在预算内。先用 `tar_poll()`/`tar_watch()` 观察 DAG,再用 crew 日志或 autometric 诊断资源。没有对应包或平台能力时明确记录降级。 ## existing 已有编号脚本或旧 runner 的项目按 simple 语义维护:编号脚本作为顺序执行的普通入口,不设专门兼容模式;自制 checkpoint 体系已移除,不再维护。迁移不是新默认;必须有人类授权、结果比对和回退安排。 -
hybrid_architecture_guide.md 2.7 KB
# R + Rmd 双模式架构 ## 新项目目录职责 ```text 项目根目录/ ├── 00.Environment.R ├── R/ # targets 调用的计算函数 ├── reports/ # 正式图、表、HTML 与补充材料 ├── products/ # 可选科学产品,不是缓存命中协议 ├── raw/ # 只读原始输入 ├── templates/ # 样式/渲染资产 ├── scripts/tests/ # 可版本化测试代码 ├── tmp/tests/<run-id>/ # 隔离测试现场与 test store ├── renv.lock ├── renv/activate.R ├── _targets.R # complex 的唯一 DAG 入口 └── _targets/ # complex 的正式 targets store ``` | 组件 | 负责 | 不负责 | | --- | --- | --- | | `00.Environment.R` | 包加载、项目根、路径与全局配置 | 具体统计逻辑 | | `R/` | 可被 target 调用的计算函数 | DAG 编排、报告展示 | | `_targets.R` | 依赖、失效、增量执行与可选报告 target | 科学解释与自定义缓存协议 | | `.Rmd` | `tar_read()`/`tar_load()`、展示参数、图表和解读 | 从 raw 重做昂贵计算 | | `_targets/` | targets 机器计算状态 | 正式科学产品 | | `products/` | 审阅、复用或交付的科学对象 | 缓存命中与恢复判断 | | `reports/` | 论文图、表、HTML、补充材料 | 源码、模型缓存、运行日志 | | `tmp/tests/<run-id>/` | 测试输入、隔离 store、日志与运行记录 | 正式结论唯一来源 | ## simple 线性、低成本单报告使用 renv 和明确的 R/Rmd 入口,smoke test 真实 render;不创建 targets、编号执行图或自制缓存。 ## complex / pipeline 复杂项目采用 `renv + _targets.R + R/ + Rmd`。普通 `tar_make()` 为默认执行;Rmd 只消费 target 结果。`products/` 按科学交付需要选择性导出,不能重新变成第二个 store。恢复验收必须真实证明中断后再次 `tar_make()` 会复用有效前序 target。 ## 并行和观测 `crew`、集群插件和 `autometric` 都是项目级可选依赖。先以 targets 状态判断 DAG,再用 `tar_poll()`/`tar_watch()` 看进度,最后用 worker 日志或资源图诊断。限制 worker 与内部线程乘积;没有资源预算时保持串行。 ## existing 与迁移 已有编号脚本或旧 runner 的项目按 simple 语义维护:编号脚本只是顺序执行的普通 Rscript 入口,不设专门兼容模式;自制 checkpoint、编号 runner 与 `tmp/{主脚本名}/` 产品路径模式已移除,不再维护。只有人类明确授权迁移时,才建立旧→新映射、结果比对、回退入口和清理计划;新默认不隐式迁移。 -
interpretation_narrative_examples.md 5.6 KB
# 专家解读思维示例(非模板) 本文件不是“替换变量名即可用”的写作模板,而是展示:如何从**本次结果的证据**出发,推导出对研究者有用的解读,并把“四层内涵”(数据描述/统计见解/领域映射/局限与后续)自然融入 1–2 段连贯叙述。 你应该学到的是**思考路径**,而不是句式本身。 ## 使用方式 - 先读每个示例的“思考过程”,用它检查自己是否真的想清楚了“结果意味着什么”。 - 再看“最终解读”,学习如何把证据、推断、边界和后续动作写成论文口吻。 - 写你自己的解读时:数值优先用 `` `r ...` `` 动态嵌入;阈值一律来自 YAML `params`(单一真相来源)。 ## 解读前四问(强制) 1. 如果我是研究者,看到这个结果,我最想确认/反驳的是什么? 2. 这个统计发现如何改变我对机制/分层/风险的理解(只写可证伪推断)? 3. 基于这个结果,我会优先做什么?不做什么?为什么? 4. 如果不复现,最可能的原因是什么(样本量/混杂/批次/模型假设)? --- ## 示例 1:差异分析(火山图 / DE 表) ### 场景(示范用,非真实结论) - 比较:高 normCCS vs 低 normCCS - 阈值:`r params$q_cutoff`(FDR/q) - 结果:Top 信号集中在 `TP53`、`KRAS`、`EGFR`(举例) ### 思考过程 **研究者最想知道什么?** - 这是不是“少数强信号 + 大量弱信号”的长尾?还是整体系统性偏移? - Top 信号是否足够稳定(分层同向?bootstrap/CV 入选率?)值得投入验证资源? **结果意味着什么?** - 统计层面:不仅要报 q,还要说明 log2FC/效应的量级是否“中等/较强”,以及差异主要由哪些对象驱动。 - 领域层面(最小可证伪写法):把 Top 1–3 变成可检验假设(如果复现,则支持某个机制/分型解释;如果不复现,优先排查混杂/批次/构成差异)。 **如何验证(必须写成方法 + 输入 + 判据)?** - 分层复核:在关键亚组内复算差异;判据:方向一致且 q 过阈值(或效应落在预设范围)。 - 外部复现:在独立队列按同口径复现;判据:方向一致 + 效应量级相近(CI 重叠或落在范围内)。 ### 最终解读(连贯叙述,示范写法) ```markdown 本次差异分析在 N=`r N` 个样本中比较了高/低 normCCS 两组,并以 q≤`r params$q_cutoff` 控制多重比较;在该阈值下共有 `r n_sig` 个基因入围。**最强信号集中在 `TP53`、`KRAS`、`EGFR`**,其 log2FC 分别为 `r lfc_tp53` / `r lfc_kras` / `r lfc_egfr`,对应 q 值为 `r q_tp53` / `r q_kras` / `r q_egfr`,提示差异主要由少数头部基因驱动而非整体系统性漂移。 对研究者而言,这一结果更适合先作为“候选机制锚点”而非直接下结论:若上述 Top 信号能在关键亚组内保持同向(判据:方向一致且 q≤`r params$q_cutoff`),并在独立队列复现(判据:效应方向一致且 log2FC 落在 `r lfc_floor`–`r lfc_ceil`),则支持其作为 normCCS 分层差异的稳定组成部分;若不复现,优先排查构成混杂与批次差异(例如组间关键临床变量分布是否不均衡),再决定是否继续投入后续功能验证。 ``` --- ## 示例 2:模型性能(AUC/C-index、校准、阈值) ### 场景(示范用,非真实结论) - 任务:二分类风险模型 - 指标:测试集 AUC=0.72(95% CI: 0.65–0.79),基线 AUC=0.62 - 目标:把“性能数字”翻译为“可执行的决策方式” ### 思考过程 **研究者最想知道什么?** - 0.72 是不是“足够可用”?在什么使用方式下可用(粗筛/分层/辅助决策)? - 置信区间下限意味着什么(最坏情况下仍可接受吗)? **如何把性能变成决策动作?** - 选择一个阈值并给出对应的敏感性/特异性(或 PPV/NPV),说明该阈值下的收益与代价。 - 明确适用边界:单中心回顾/样本外泛化/校准是否可靠。 ### 最终解读(连贯叙述,示范写法) ```markdown 本模型在测试集(N=`r n_test`)上的 AUC 为 `r auc`(bootstrap 95% CI: `r auc_ci`),较临床基线模型提升 `r delta_auc`,提示在当前数据分布下具有中等区分度。若以 `r thr` 作为决策阈值,对应敏感性/特异性为 `r sens`/`r spec`,可用于将人群粗分为“高风险需强化随访”与“常规随访”两类;但需要注意 CI 下限 `r auc_ci_low` 意味着在部分人群中性能可能下降。 下一步建议:(1)先做校准评估并在必要时校准概率输出(输入:校准曲线/校准 slope + Brier;判据:校准 slope 处于 `r slope_low`–`r slope_high` 且 Brier 下降),避免风险概率误导阈值决策;(2)在外部队列复现(输入:外部测试集;判据:AUC≥`r auc_floor` 且校准 slope 落在同一范围),通过后再讨论是否将该阈值固化为工作流的一部分。 ``` --- ## 当缺少“领域知识库”时怎么写(避免装懂) 如果你无法可靠地解释某基因/通路的具体机理,不要编故事。推荐用“**最小领域映射**”保持专业性与可验证性: - 只陈述:本次数据里**谁**最突出、**方向/对比**是什么、**量级/不确定性**是什么。 - 把“领域见解”写成**可证伪假设**(如果...那么...),并把后续验证写成“方法 + 输入 + 判据”。 - 明确边界:把“可能的原因/风险”落到数据与方法层面(混杂、批次、样本量、模型假设),而不是泛泛一句“需要进一步研究”。 -
interpretation_templates.md 5.5 KB
# 深度解读写作骨架(证据锚定) 用于把四层内容(数据描述、统计见解、领域映射、局限与后续)写成当前结果的连贯叙述。占位符必须替换为本次对象和数值,数值优先使用 `` `r ...` ``;阈值统一来自 YAML `params`。任何“可能/提示/建议验证”都要补对象、动作和判据。禁止用代码拼接整句解释文本。 ## 写作前四问与硬门槛 1. 负责人最想确认/反驳什么? 2. 发现如何改变机制、分层或风险理解(只写可证伪推断)? 3. 基于结果优先做/不做什么,为什么? 4. 不复现时最可能的样本量、混杂、批次或模型原因是什么? 不要机械输出“数据描述:…统计见解:…”四段,也不要写“这张图/表用于……”而不陈述当前观察。最终正文通常为 1–2 段自然叙述,关键观点可加粗。 ## 默认连贯叙述 ```markdown 本次分析在有效 N=`r N` 个样本/对象上评估了 `r target`(`r comparison`),阈值为 `r threshold`(来自 `params`)。 **关键估计对象 `r parameter` 的效应为 `r estimate` `r unit`**,`r confidence_level` CI [`r ci_lower`, `r ci_upper`](`r method`);若有预定零假设,再给原始 p=`r p_value` 及适用的 `r adjustment_method` 后 q=`r q_value`。 这对 `r domain_context` 的实际意义是 `r domain_mapping`;主要局限为 `r limitation`,当前证据等级为 `r evidence_level`。 后续:1) 用 `r method_1` 在 `r input_1` 上验证,判据 `r criterion_1`;2) 用 `r method_2` 在 `r input_2` 上验证,判据 `r criterion_2`。 ``` 上述 p/q 句仅在预定检验及多重校正确实适用时保留,数字须来自计算层完整结果。若状态为 `not_applicable` 或 `not_estimable`,改写为“本结果的 [CI/检验] [不适用/不可估计],原因是 [具体设计或数据限制];可采用 [替代分析],当前只能得出 [有限结论]”。`not_run` 应说明未执行,不能作阴性结论。 可为每个 cohort/亚组附核心结论:Top 1–3、证据强度、主要不确定性和后续动作。小表应只汇总当前对象、数值、排名/对比、证据等级和下一步,不重复整段解释。 ## 反模式与替代 - 只报 p/q,不报效应大小、方向和不确定性。 - 只报点估计,未判断关键参数能否给 CI 或预定假设下的 p 值。 - 只写“具有临床意义/可解释/可行动”,不落到变量、决策动作和验证。 - 写“建议进一步研究”,但没有方法、输入、判据。 - 写“可能混杂/非线性”,但不指出对象和检验方式。 替换为:“`X` 的效应为 …(CI …),与 `Y` 相比 …;不确定性来自 …。用 … 在 … 上验证,支持条件为 …。” ## 模板 0:核心结论(cohort/亚组) ```markdown 本 cohort(`r cohort_label`,有效 N=`r n_all`)中,`r event_set` 与目标的关键效应为 `r estimate`(`r confidence_level` CI [`r ci_lower`, `r ci_upper`],方法 `r method`);仅有预定零假设时报告 p/q。主导维度为 `r dominant_dimension`。 **最强信号为 `r top_1`、`r top_2`、`r top_3`**(当前值 + 阈值 + 稳健性证据),与宏观方向 `r consistency`。因此证据等级为 `r evidence_level`,主要不确定性是 `r uncertainty`。 后续:1) `r stratified_method` 于 `r input`,判据 `r criterion_1`;2) `r resampling_method`,判据 `r criterion_2`;3) 必要时在 `r external_cohort` 复现,判据 `r criterion_3`。 ``` ## 模板 1:单因素筛选 / 批量回归 报告候选特征数、方法/结局、q/p 阈值和入围数;保留完整未筛选结果,Top 1–3 关键效应给对象、方向、beta/OR/HR/差值、适当 CI,以及预定检验适用时的 p/q;说明整体效应量级与模式。领域段须把一个 Top 信号映射到可检验假设(方法 + 输入 + 判据)。后续至少包括: ```markdown 1. 多因素校正:`r model` 加入 `r covariates`;判据为 Top 信号方向一致且 `r stability_rule`。 2. 稳定性:`r resampling_method`(`r n_resamples` 次);判据为入选频率/CI 达到 `r threshold`。 ``` ## 模板 2:多因素模型(logistic/Cox/线性) 报告 N、事件数、参数数、EPV(适用时)、模型/结局、特征集和缺失处理。Top 1–3 给系数或 OR/HR、CI、方向和 p/q;说明校准、PH、共线性或 bootstrap 稳健性。领域段明确决策动作和适用边界。后续包括: ```markdown 1. 假设诊断:`r diagnostic_method` 检查 `r assumption`;判据 `r diagnostic_criterion`。 2. 泛化:`r validation_method` 评估性能/系数稳定性;判据为性能下降 ≤ `r drop_threshold` 且方向一致率 ≥ `r stability_threshold`。 ``` ## 模板 3:模型性能与验证 报告验证方案、训练/验证/测试规模和事件数;给 AUC/C-index、Brier、校准 slope、net benefit 等及 CI;比较基线提升、训练—测试差距与稳定性。将性能映射到具体阈值、收益、代价和人群边界。后续至少包括: ```markdown 1. 外部验证:在 `r external_dataset` 复现;判据性能 ≥ `r floor` 且 slope ∈ `r slope_range`。 2. 阈值/决策:用 `r decision_method`;判据 net benefit 在 `r range` 内为正或优于基线。 ``` ## 数据溯源 统计量引用模型摘要或计算变量,N 引用维度/计数,比例引用计算逻辑,阈值引用 `params`。写完逐项核对解读中的数字能在代码/数据中定位;图表证据应符合 `plot_quality_standards.md`。 -
lightweight_testing.md 5.6 KB
# 轻量测试协议 ## 适用范围与通过定义 每个新建或实质修改的分析流程必须选择一种测试风格,并从正式入口真实运行适用的计算与报告。checker、静态检查、R 语法检查、targets graph 预览和 `--dry-run` 只算预检;只有 R/Rmd/target 图实际执行且断言通过,才能写 PASS。 测试代码进入 `scripts/tests/` 并版本化;每次测试使用唯一 `tmp/tests/<run-id>/`,其中包含 fixture/子集、隔离 products/reports/store、日志与渲染结果,默认加入 `.gitignore`。运行前后比较正式 `raw/`、products、reports、`_targets/` 以及相关代码/配置,必须保持不变。 ## 测试风格 ### synthetic_fixture 在没有授权真实数据、数据敏感或原始数据过大时使用。fixture 至少保留: - 与正式输入一致的列名、类型、主键和关键分组; - 业务允许的缺失模式; - 至少一个会触发重要分支或边界的案例; - 固定随机种子及生成规则。 不要为了“容易通过”删除正式流程必须处理的复杂性,也不要把真实敏感值复制进 fixture。 ### project_subset 存在授权数据和已有流程时优先。抽样规则必须确定、可复述且有代表性,覆盖关键分组、结局、缺失和异常边界。除非顺序本身有业务意义且已说明,不能只用 `head(n)`。子集采用最小字段和最少样本的真实副本,不链接回原始文件;断言后默认删除,只保留抽样规则、schema、行列数、非敏感摘要与不可还原证据。只有人类明确要求且项目具备合适访问控制时才保留。 ## 隔离与同入口 - simple:测试 wrapper 准备隔离输入后,调用与正式运行相同的 Rscript 或 `rmarkdown::render()` 入口;输出重定向到本次 run root。 - complex/pipeline:同一 `_targets.R` 和 target 图接收隔离输入,`targets::tar_make(store = file.path(run_root, "_targets"))`;不得写正式 `_targets/`。至少一次在前序 target 成功后中断下游,再次 `tar_make()`,用 `tar_meta()`/outdated 证明前序复用;不得用 SUCCESS 或自定义 checkpoint 替代。 - 测试专用 helper 只存在于 `scripts/tests/` 或隔离副本。不得向业务逻辑添加长期 `analysis_mode`、`test_mode` 或两套科学行为。 - harness 先设置 `BENSZ_TEST_RUN_ROOT`,再用正式入口已经识别的 `BENSZ_ANALYSIS_INPUT`、`BENSZ_PRODUCTS_DIR` 和 `BENSZ_REPORTS_DIR` 把输入、产品和报告绑定到同一 run root;`00.Environment.R` 只接受与该 run root 精确匹配的 `tmp/tests/<run-id>/{products,reports}`,不能借测试变量放宽任意 `tmp/` 写入。 - 正式与测试必须调用同一计算函数和同一报告入口。若正式代码无法在不加入测试分支的情况下隔离路径,应先修正路径边界。 existing 项目实质修改后仍应沿其现有 targets 或顺序脚本入口做可行的轻量运行,但不得为测试补建 renv、targets 或持久测试框架。可以使用一次性隔离副本;若旧入口硬编码绝对路径或会覆盖正式结果,将其记录为 `path_isolation` 阻塞/风险,请人类决定是否授权最小可测试性改造或显式迁移。 ## 最低断言 按任务补充科学断言,通用最低集包括: 1. 输入 schema、列类型和必需字段; 2. 行列数、主键唯一性与关键分组覆盖; 3. 重要数值范围或统计不变量; 4. 预期产品、表格、图和报告确实生成且可读; 5. `raw/` 内容摘要在运行前后不变; 6. 测试未修改正式 products/reports/targets store 或代码/配置; 7. 对 complex,隔离 store 中存在实际构建结果,且中断恢复复用仍有效 target,而非仅图解析成功。 ## 执行—诊断—修正—重跑 一次尝试至少记录:命令、退出状态、失败类别、修正摘要、断言结果和剩余风险。失败类别固定为: - `environment_dependency` - `fixture_or_subset` - `path_isolation` - `analysis_code` - `scientific_assertion` - `report_rendering` - `external_service` 失败后只修正对应层并重跑真实受影响路径。不得无限重试;遇到权限、网络、缺失依赖或数据授权等外部阻塞时,停止并明确未验证项。不得把跳过、预检通过或旧缓存命中写成端到端通过。 运行记录至少包含:`project_state`、`workflow_mode`、已有项目观察机制、`test_style`、`formal_entrypoint`、相对 `run_root` 与路径覆盖、R/renv 来源、绑定代码/配置/lockfile 的 `subject_identity`、`preflight`、`lightweight_execution`、`full_data_execution`、断言、正式路径前后对照、尝试和 cleanup。三种执行状态取 `PASS`、`FAIL`、`BLOCKED` 或 `NOT_RUN`;未实际运行全量数据时,`full_data_execution` 必须为 `NOT_RUN`。 轻量测试通过只证明该代码路径和指定不变量在轻量输入上成立,不证明全量资源消耗、罕见值、外部服务稳定性或最终科学结论。审查修正若影响执行、输出或断言,旧 subject identity 的证据立即过期,必须对最新版本重跑。 ## 模板入口 - `templates/tests/synthetic_fixture.R`:模拟数据骨架; - `templates/tests/project_subset.R`:授权子集骨架; - `templates/tests/simple_smoke_test.R`:simple 真实入口; - `templates/tests/complex_smoke_test.R`:complex 隔离 target store 入口。 - `templates/tests/test_harness.R`:唯一 run root、subject identity、正式路径不变性与运行记录 helper。 复制到项目后按真实 schema、文件名和科学不变量修改,不能保留占位断言后声称通过。 -
liquid_glass_theme_guide.md 3.9 KB
# Liquid Glass Theme 使用指南 本主题为 R Markdown HTML 提供 glassmorphism 风格:半透明/模糊卡片、渐变与多层阴影、自动深色模式、响应式目录、图片 Lightbox、代码复制按钮和默认居中图片。默认无无限循环动画,避免复制/选中文本抖动;目录支持桌面动态/静态浮动和窄屏折叠。 ## 快速开始 将 `templates/liquid_glass_theme.css` 与 `templates/liquid_glass_lightbox.html` 放入项目 `templates/`,或运行: ```bash python3 /path/to/bensz-rmd-rules/scripts/bootstrap_liquid_glass.py --with-env ``` 在 Rmd YAML 中使用: ```yaml output: html_document: toc: true toc_float: true theme: default highlight: tango code_folding: show includes: after_body: "templates/liquid_glass_lightbox.html" css: "templates/liquid_glass_theme.css" ``` `theme: default` 兼容 `toc_float`;`css` 加载主题;`after_body` 启用图片放大、代码复制和视图状态保持。不要用 `includes.in_header` 直接 include CSS,否则 Pandoc 可能把 CSS 当正文。 渲染: ```bash python3 skills/knit-rmd-html/scripts/knit_rmd_html.py your_report.Rmd ``` ## 目录与交互 - 桌面(>1024px):动态模式以左上角小圆点展开;目录的可交互区域立即展开,鼠标可直接移入正文,可切换静态常驻,偏好写入浏览器存储。 - 窄屏(≤1024px):目录回到顶部并折叠为 sticky 卡片;点击展开,跳转后收起并处理标题偏移。 - 代码块:启用 `after_body` 后在右上角显示 Copy;现代浏览器用 Clipboard API,旧浏览器或 `file://` 降级为 `execCommand`。主题会关闭代码字体连字,确保 R 的 `<-` 等多字符运算符保持源码字面形态,不被合成箭头字形。 - Lightbox:点击图片放大,Esc/空白处关闭;`Ctrl/Cmd + (+/-/0)` 可用页面缩放,刷新后尽量恢复位置。 ## 视觉与自定义 主题使用 `backdrop-filter`、半透明背景和 CSS custom properties;默认支持系统深色模式。常用变量: ```css :root { --lg-accent-blue: #007aff; --lg-accent-purple: #5856d6; --lg-accent-pink: #ff2d55; --lg-gradient-primary: linear-gradient(135deg, #667eea 0%, #764ba2 100%); } ``` 可用工具类: ```html <div class="lg-glass">玻璃容器</div> <h2 class="lg-gradient-text">渐变标题</h2> <div class="lg-card">卡片内容</div> ``` 在 Rmd 末尾用 CSS chunk 覆盖样式;动画应克制,并尊重 `prefers-reduced-motion`。主题默认最大容器宽度约 1200px,带明显焦点轮廓与 WCAG AA 对比目标。 ## 兼容性与故障排查 需要支持 `backdrop-filter` 和 CSS 自定义属性;已验证目标为 Chrome 90+、Safari 14+、Firefox 88+、Edge 90+(Safari 需要 `-webkit-` 前缀)。主题无运行时 JS 依赖,`after_body` 只提供交互增强。 - 样式无效:确认两个模板存在,YAML 使用正确相对路径,且 `toc_float` 时保留 `theme: default`;必要时强制刷新。 - 目录不显示:检查文档有 H2/H3,桌面窗口宽度是否 >1024px。 - 目录悬停时闪烁或刚展开就收回:更新项目中的 `liquid_glass_theme.css` 并重新渲染 HTML;已有项目的模板不会随 Skill 源码自动更新,若使用 `bootstrap_liquid_glass.py --force`,请先检查项目里是否有自定义主题改动。 - 深色模式不切换:检查系统偏好和浏览器 `prefers-color-scheme` 支持。 - `<-` 显示成 `←`:确认项目中的 `liquid_glass_theme.css` 已更新;旧项目需重新运行初始化脚本并显式使用 `--force`,或手动同步主题文件。该现象只影响字体显示,复制出的源码通常仍是 `<-`。 - 不需要缩放/滚动保持:移除 `includes.after_body` 或换用仅含 Lightbox 的文件。 主题文件:`templates/liquid_glass_theme.css`;交互文件:`templates/liquid_glass_lightbox.html`。问题和建议请在 bensz-rmd-rules issue 反馈。 -
metric_explanation_protocol.md 5.8 KB
# 指标解释协议 本协议用于让专业 R Markdown 报告同时服务相关背景较弱的读者与资深研究者。默认假定读者不了解任务特有指标,但具备阅读基础医学统计结果的意愿。写作顺序是先建立最低限度的理解支架,再给出精确、可追溯的结果判断;不得用通俗化牺牲定义、方向、参考点或不确定性。 ## 判定规则 先按 `config.yaml:metric_explanation.common_metrics` 判断。只有含义稳定、跨医学研究广泛使用且通常无需任务专属定义的指标,才按常用指标处理;未命中白名单、无法确定或存在多种主流定义时,一律按不常用指标处理。 默认常用指标包括: - 描述统计:样本量、频数、比例、均值、中位数、标准差、四分位距、发生率。 - 推断与不确定性:p 值、q 值/FDR、95% 置信区间。 - 常见效应与关联:均值差、回归系数、OR、RR、HR、Pearson `r`、Spearman `rho`。 - 经典生存分析:Kaplan–Meier 生存概率、中位生存期、log-rank 检验、Cox HR、C-index。 - 经典诊断与二分类性能:ROC AUC、灵敏度、特异度、阳性/阴性预测值、准确率。 以下情况即使名称看似熟悉,也按不常用指标处理: - 用户或项目自行构建的评分、指数、组合终点或归一化分数。 - 只服务特定模型或窄领域的指标,如 Brier score、校准斜率/截距、net benefit、NRI、IDI、SHAP 汇总量、silhouette、modularity、kME、伪时间等。 - 采用非标准公式、转换、截断、方向、量纲或参考点的常用指标。 - 同一名称存在多个定义,而当前报告未明确采用哪一种。 分类只决定解释密度,不代表指标重要性。常用指标若存在反直觉方向或关键判读门槛,首次出现时仍应补一条简短定向说明。 ## 指标导读表 只要报告使用至少一个不常用指标,就在“关键函数、参数与源代码位置”之后、首个分析模块之前增加 `## 指标导读`。无不常用指标时删除该章节,避免为形式完整而制造冗余。 表格至少包含: | 指标 | 它回答什么问题 | 尺度、参考点与趋势含义 | 不确定性与置信区间怎么读 | 本报告为何使用 | 判读边界 | |---|---|---|---|---|---| | `{metric}` | 用一句普通语言说明 | 写清范围、单位、零点/基线;升高或降低分别意味着什么 | 说明 CI/SE/bootstrap/IQR 等来自哪里、跨过哪个参考点意味着什么;无 CI 时如实说明 | 绑定本次分析目标,不写通用宣传语 | 写清适用前提、不可推出的结论及非标准定义位置 | 要求: - 表格解释判读规则,不提前替代具体结果解读。 - 不得假设每个指标都有置信区间;没有 CI 时报告实际使用的不确定性证据,不得编造。 - 自定义指标必须给出公式或算法位置,例如 `{name}_functions.R:85`,并解释高低方向。 - 名称、缩写、单位、方向和参考点必须与代码及结果表一致。 ## 首次出现协议 不常用指标在结果解读正文第一次出现时,按以下逻辑自然写成一个短段落,然后继续当前结果;不要机械输出固定标签: 1. 用一句普通语言说明它衡量的对象或问题。 2. 用一至两句说明计算原理:输入是什么、如何汇总或比较、输出尺度是什么。自定义指标需指出公式或源代码位置。 3. 说明为什么本次分析选择它,以及它补充了哪个常用指标无法回答的问题。 4. 交代方向、参考点和不确定性:高/低意味着什么,CI 是否跨越关键参考点,或实际采用何种稳定性证据。 5. 立即回到本次结果,用可追溯数字解释当前对象的方向、量级、边界和价值。 推荐采用“双层表达”:先给不失真的普通语言结论,再给精确术语和数值证据。例如,先说明“模型给出的概率与真实发生率仍有偏差”,再报告校准斜率、截距及其不确定性。类比只用于辅助直觉,不能替代正式定义。 ## 后续出现协议 同一指标从第二次结果解读开始,不再重复定义、原理和选用理由,只保留: - 当前对象、对比或排序。 - 指标值、方向、参考点和不确定性。 - 对当前研究问题的具体意义。 “同一指标”要求名称、公式、量纲、方向和参考点均一致。任一项变化,或切换到不同时间窗、不同标准化方式、不同版本算法时,视为新指标,重新执行首次出现协议。 ## 写作边界 - 深入浅出不等于教学口号。避免“大家都知道”“简单来说就是”“越高一定越好”等失真表达。 - 不把统计关联写成因果,不把区分度写成校准度,不把模型性能写成临床获益。 - 常用指标不做教科书式长篇复述;必要时只给一句与当前结果相关的方向或参考点提醒。 - 指标解释属于必要的读者支架,不受“避免教学口吻”规则中的模板化提示语限制;但第二次起必须收敛,防止报告反复上课。 - 所有具体数字继续遵循 `` `r ...` `` 动态嵌入和解读—代码一致性要求。 ## 交付自检 - [ ] 是否先完成常用/不常用指标盘点,无法确定的指标是否按不常用处理? - [ ] 存在不常用指标时,是否在分析前提供指标导读表? - [ ] 表格是否覆盖尺度、参考点、趋势、不确定性、选用理由和判读边界? - [ ] 每个不常用指标的首次正文出现是否解释了定义、原理、理由、价值和本次结果? - [ ] 第二次及以后是否避免重复教学,只保留证据锚定的结果解读? - [ ] 同名但定义变化的指标是否重新解释? - [ ] 通俗表达是否仍与公式、代码、数值和统计边界一致? -
no_overdefensive_code.md 4.6 KB
# 不过度保护原则(避免“防御性过载”) 核心理念:信任 `00.Environment.R` 的环境配置,避免对已加载的包和函数进行冗余检查;让错误显式暴露,避免静默失败。 ## 绝对禁止的反模式 ```r # 包加载检查(绝对禁止) if (!requireNamespace("DT", quietly = TRUE)) { # 降级方案 } else { DT::datatable(data) } # 正确做法:直接使用 DT::datatable(data),因为 00.Environment.R 已加载 # 包加载多重检查(绝对禁止) if (requireNamespace("luckyBase", quietly = TRUE)) { luckyBase::Plus.library("DT") } else { suppressPackageStartupMessages(library(DT)) } # 正确做法:00.Environment.R 已执行加载,主脚本无需重复 # 数据多重验证(绝对禁止) if(!is.null(data) && nrow(data) > 0 && !all(is.na(data$value))) { if(class(data$value) %in% c("numeric", "integer")) { result <- mean(data$value, na.rm = TRUE) } else { result <- NA warning("数据类型不正确") } } else { result <- NA warning("数据为空或无效") } # 简洁直接(推荐) result <- mean(data$value, na.rm = TRUE) # 表格渲染过度检查(绝对禁止) if (!is.null(data) && requireNamespace("DT", quietly = TRUE)) { DT::datatable(utils::head(data, n), options = list(scrollX = TRUE, pageLength = 10)) } else if (!is.null(data)) { utils::head(data, n) } # 正确做法:DT::datatable(utils::head(data, n), options = list(scrollX = TRUE, pageLength = 10)) ``` ## 检查边界规则 必须检查的场景(白名单): - 文件 I/O 前检查路径有效性 - 用户输入的关键参数(如分组变量是否存在于数据中) - 外部数据源连接(如数据库、API) - **00.Environment.R 存在性检查**:作为整个分析的硬性前提,显式 `if (!file.exists("00.Environment.R")) stop(...)` 提供比自然报错更清晰的错误信息,属于合理异常(见 `templates/Rmd_template.Rmd` 示例) 不应检查的场景(黑名单): - R 包是否已加载(`00.Environment.R` 已通过 `luckyBase::Plus.library()` 统一管理) - 函数是否存在(对于通过 `00.Environment.R` 加载的函数) - 内部中间变量的类型检查 - 标准 R 包函数的返回值验证 - 已由用户代码处理过的数据 硬性规定: - 主脚本中禁止使用 `requireNamespace()` 检查包可用性 - 主脚本中禁止使用 `if (!is.null(...) && requireNamespace(...))` 模式 - 主脚本中禁止出现包加载的“降级方案”或“回退逻辑” ## 禁止占位性代码(Fail Loud) 核心问题:在科研分析场景中,用占位性代码保证“不报错”会隐藏功能失败,导致分析结果不完整且难以察觉。 占位性代码定义:使用 `try()` 捕获错误后,将结果赋值为 `NULL`、`NA`、空数据框,或直接跳过/打印警告继续,从而保证代码表面运行成功的模式。 绝对禁止的占位性模式: | 模式类型 | 示例代码 | 危害 | |---------|---------|------| | try-catch 后赋值 NULL | `if (inherits(fit, "try-error")) { fit <- NULL }` | 静默失败,后续代码可能崩溃或产生误导性结果 | | try-catch 后打印警告继续 | `if (inherits(fit, "try-error")) { cat("Failed, skipping...\n"); next }` | 用户不易察觉,分析结果不完整 | | 降级到占位符 | `if (!requireNamespace("pkg")) { result <- data.frame() }` | 空结果导致误导性输出 | | 条件分支后无有效逻辑 | `if (error) { return(NA) } else { ... }` | 功能形同虚设,分析目标未达成 | 切实落地要求: - 必须确保功能在正常流程中切实落地,不得使用占位符保证“不报错” - 如功能失败,应 `stop()` 报错而非静默处理 - 如需容错,必须提供有效降级方案(替代算法/简化模型),且降级方案也要切实落地 正确示例对比: ```r # ❌ 危险:占位性代码 fit_lda <- try(MASS::lda(f_main, data = train), silent = TRUE) if (inherits(fit_lda, "try-error")) { cat("LDA model fitting failed, skipping...\n") fit_lda <- NULL } # ✅ 正确:切实落地或报错 fit_lda <- MASS::lda(f_main, data = train) # ✅ 正确:需要容错时,提供有效降级方案 fit_lda <- try(MASS::lda(f_main, data = train), silent = TRUE) if (inherits(fit_lda, "try-error")) { message("LDA failed, falling back to QDA...") fit_lda <- MASS::qda(f_main, data = train) } if (inherits(fit_lda, "try-error")) { stop("Both LDA and QDA failed. Please check data or use other methods.") } ``` 权衡原则: - 科研分析 > 代码美观(宁可报错,也要结果可信) - 显式失败 > 静默成功(失败要暴露) - 有效降级 > 占位符(降级也要实现目标) -
numeric_accuracy_verification.md 1.3 KB
# 数字准确性验证(末尾检验) 每个 `.Rmd` 的末尾应包含“数字准确性验证”章节,确保解读中的关键数字均可追溯到具体变量或计算逻辑。 > 核心目标:避免出现“写了一个比例/效应量/样本数,但无法解释它从哪里来的”。 ## 章节模板 ```markdown ## 数字准确性验证 本报告中的关键数字均可追溯到以下变量或计算逻辑(必要时附代码位置/章节名): | 数字描述 | 来源变量/计算 | 位置 | |---|---|---| | 样本量 | `nrow(data)` | 数据加载与处理 | | 事件数 | `sum(data$event == 1)` | 生存分析 | | HR | `hr` | 多因素 Cox 回归 | | ... | ... | ... | 验证结论: - [ ] 所有数字均可追溯到变量或计算逻辑 - [ ] 关键数字均使用 `` `r ...` `` 动态嵌入(或明确说明计算来源) - [ ] 未出现无法解释来源的数字或“凭空捏造”的比例/效应 ``` ## 禁止行为 - 在解读中写出无法追溯的数字(如“约 35%”但无变量来源) - 用“可能/大约/左右”替代可计算的关键数字 ## 允许行为 - 使用 `` `r sprintf('%.2f', value)` `` 或 `` `r round(value, 1)` `` 做格式化 - 在明确给出来源变量的前提下写相对描述(如“约 `r sprintf('%.1f%%', 100*prop)`”) -
plot_language.md 2.7 KB
# 图表语言规范(plot_language) ## 英文优先原则 **核心规则**:所有可视化图表(ggplot2、survminer、lattice 等)的文本元素默认必须使用英文。 适用范围: | 文本元素 | 要求 | 示例 | |---------|------|------| | **轴标题** | 必须英文 | `x = "Gene Expression"` (非 `x = "基因表达"`) | | **图例标签** | 必须英文 | `color = "Cancer Type"` (非 `color = "癌种"`) | | **图表标题** | 默认不使用;如使用必须英文 | `title = "Survival Analysis"` (非 `title = "生存分析"`) | | **副标题/注释** | 默认不使用;如使用必须英文 | `subtitle = "Log-rank test"` (非 `subtitle = "Log-rank 检验"`) | | **刻度标签** | 推荐英文 | 分类变量使用英文标签(如 "LUAD"、"BRCA") | 实践示例: ```r # 正确做法(默认;强制):不在图内使用 title/subtitle,把语义放在文件名与 .Rmd 正文结构里 ggplot(data, aes(x = SYMBOL, y = expression, color = cancer_type)) + geom_point() + labs( x = "Gene Symbol", y = "Expression Level (log2 TPM)", color = "Cancer Type" ) + theme_minimal() # 错误做法(禁止) ggplot(data, aes(x = SYMBOL, y = expression, color = cancer_type)) + geom_point() + labs( title = "不同癌种的基因表达", # 禁止中文标题 x = "基因", # 禁止中文轴标签 y = "表达量", # 禁止中文轴标签 color = "癌种" # 禁止中文图例 ) ``` ## 例外场景 | 场景 | 处理方式 | 示例 | |------|----------|------| | **中文期刊投稿** | 在 YAML params 中添加 `plot_language: "zh"`,由 AI 根据参数动态调整 | `params$plot_language == "zh"` 时使用中文 | | **特定临床术语** | 优先使用国际通用英文术语 | 使用 "Overall Survival" 而非"总生存期" | | **基因/蛋白名称** | 始终使用国际标准符号(如 TP53、EGFR) | 无需翻译 | 实现建议: ```r # 方式 1:在 00.Environment.R 中定义全局标签函数 .set_plot_labels <- function(language = "en") { if (language == "en") { list( x_title = "Gene Expression", y_title = "Expression Level", legend_title = "Cancer Type" ) } else if (language == "zh") { list( x_title = "基因表达", y_title = "表达水平", legend_title = "癌种" ) } } # 方式 2:在 .Rmd 的 YAML params 中配置 params: plot_language: "en" # 或 "zh" ``` 检查清单: - [ ] 图表轴标题是否使用英文? - [ ] 图例标签是否使用英文? - [ ] 若显式使用了图表标题/副标题(默认不建议),是否使用英文? - [ ] 是否避免了中文字符在图表中出现(除非有特殊例外)? -
plot_quality_standards.md 7.4 KB
# 图表质量规范(Nature 级别) 本文件定义“默认即出版级别”的图表质量标准,用于指导 AI 在生成绘图代码时自动应用统一规范。 唯一目标:最大程度保证人类可读性(屏幕阅读 + 打印均清晰)。 说明:当前仅强制 Nature 级别默认规范;不提供多期刊风格切换(保持 KISS)。 ## AI 自主规划原则(必须) - 可读性优先:任何参数调整都以“读者能否清晰阅读”为第一准则。 - 场景自适应:根据数据量/图表类型/展示载体(论文/幻灯/交互)动态调整字体、线宽、点大小、图例位置等。 - 默认值 + 允许偏离:下方提供推荐默认值;若偏离需有明确理由(例如“点数>1e4,避免遮挡需减小点大小”)。 - 避免视觉缺陷:主动检查并修复标签重叠、过密、字体过小、对比不足、中文乱码、图例遮挡等问题。 ## 可读性自检清单(生成代码时必须自检) - [ ] 字体大小:轴标签/刻度/图例/标题是否清晰?是否存在过大/过小? - [ ] 线条与点:线宽/点大小是否与图密度匹配?是否过粗/过细? - [ ] 标签重叠:轴刻度/注释/数据标签是否重叠?是否需要旋转/缩写/分面? - [ ] 字符编码:中文/希腊字母/上下标是否正常?必要时指定字体或使用表达式。 - [ ] 图例布局:图例是否遮挡数据?是否应改到底部/右侧/外置? - [ ] 配色对比:颜色是否足够区分?是否色盲友好?避免彩虹色。 - [ ] 尺寸比例:宽高比是否与内容匹配?是否过扁或过窄导致信息拥挤? ## 通用技术规范(跨包) | 维度 | 推荐默认值 | 说明 | |------|------------|------| | 输出格式(静态) | PDF(矢量,优先) | 矢量优先;位图仅在特殊场景使用并说明原因 | | 字体 | Arial / Helvetica(无衬线) | 系统缺失时回退到 `sans`,不要为了字体引入复杂依赖 | | 背景与网格 | 白底 + 减少网格 | 避免灰底;网格默认去除或弱化,强调数据本身 | | 配色 | Nature 调色板 / viridis | 必须色盲友好;避免彩虹色、低对比度组合 | ## ggplot2 推荐默认值 | 维度 | 推荐默认值 | 说明 | |------|------------|------| | 字体大小 | 轴刻度 10pt,图例 9pt,标题 12pt | 图密度高时可整体下调 2–4pt | | 线宽 | 数据线 0.8–1.0,边框/轴线 0.5 | 复杂图可降至 0.5;强调可升至 1.5 | | 图例位置 | 右侧或底部 | 避免遮挡数据;优先外置 | | 保存 | `ggsave(..., device = "pdf")` | PDF 无需 dpi;若输出位图需指定 dpi 并说明 | 推荐实现(详见模板): - `templates/nature_colors.R`:`nature_colors` - `templates/nature_theme.R`:`theme_nature()` / `theme_nature_readable()` ## ComplexHeatmap 推荐默认值 | 维度 | 推荐默认值 | 说明 | |------|------------|------| | 字体 | 行/列名 8–10pt,图例标题 9pt | 大矩阵需动态减小字体并可隐藏部分行/列名 | | 配色 | Nature / viridis | 使用 `circlize::colorRamp2()` 构造连续色带 | | 尺寸 | 随矩阵维度动态调整 | 避免挤压导致文字不可读 | 参考模板:`templates/complexheatmap_template.R`(含 `make_heatmap_nature_safe()` 等可读性安全入口)。 ## plotly 推荐默认值 | 维度 | 推荐默认值 | 说明 | |------|------------|------| | 字体 | 轴标题 14px,刻度 12px,图例 11px | 随图大小调整 ±2–3px | | 线宽 | 2–3px | 点多时可减小点大小与透明度 | | 配色 | Nature / viridis | 保证色盲友好与高对比度 | | 输出 | 交互优先 HTML | 静态图优先 ggplot2/ComplexHeatmap 的 PDF | 参考模板:`templates/plotly_template.R`。 ## 可读性问题诊断流程 当图表出现可读性问题时,按以下流程诊断与修复(优先用“最小改动”解决核心问题): 1. **识别问题类型** - 文字溢出:长标签超出画布边界(常见于分类轴) - 文字遮挡:标签/注释与数据点、图例、标题重叠 - 字体过小:打印或投影不可读 - 标签过密:刻度过多、点太多、注释太多 - 中文/特殊字符异常:乱码、方块、缺字 2. **选择处理策略(按优先级)** - 调整布局:增大画布尺寸、增加边距、调整宽高比、移动图例 - 调整内容:旋转/换行/缩写、减少 breaks、仅标注关键点 - 调整展示方式:分面(facet)、拆成多图;必要时改用交互(plotly) 3. **验证修复效果** - 文字是否完全可见、无裁剪 - 文字是否清晰可读(屏幕 + 打印) - 修复是否引入新问题(如旋转后边距不足、图例挤压数据) 确定性 PDF/JPG 入口: ```bash Rscript <skill-root>/scripts/check_plot_readability.R path/to/figure.pdf \ --render-jpg --out-dir <任务根>/bensz-rmd-rules/output/plot-check ``` `<任务根>` 必须是当前 `.bensz-api/task-*`;未提供 `--out-dir` 时脚本读取 `BENSZ_TASK_ROOT`,不会回退到共享缓存或系统临时目录。 ## 常见场景处理示例(最小可复用) 说明:以下 ggplot2 示例默认已 `source("templates/nature_theme.R")`(提供 `theme_nature()` / `theme_nature_readable()`)。 ### 场景 1:长分类名(基因名/样本名/分组名很长) 推荐:旋转 + 增加边距(或换行/缩写)。 ```r p <- ggplot2::ggplot(df, ggplot2::aes(x = category, y = value)) + ggplot2::geom_col(width = 0.8) + theme_nature_readable(base_size = 10, x_text_angle = 45) + ggplot2::theme( axis.text.x = ggplot2::element_text(hjust = 1, vjust = 1), plot.margin = ggplot2::margin(10, 10, 18, 10) ) ``` ### 场景 2:数据点很多导致标签过密(散点/火山图/条形图 topN) 推荐:只标注关键点(极值/topN);其余用交互(plotly)或表格补充。 ```r # 不引入额外依赖的最小做法:仅保留 topN,再用 geom_text(check_overlap=TRUE) df_label <- df[order(df$score, decreasing = TRUE), ][seq_len(min(10, nrow(df))), , drop = FALSE] p <- ggplot2::ggplot(df, ggplot2::aes(x = x, y = y)) + ggplot2::geom_point(size = 1, alpha = 0.7) + ggplot2::geom_text( data = df_label, ggplot2::aes(label = label), size = 3, check_overlap = TRUE ) + theme_nature_readable(base_size = 10) ``` ### 场景 3:图例遮挡数据(尤其面板小、分组多) 推荐:优先外置到底部,并横向排布(降低遮挡概率)。 ```r p <- p + ggplot2::theme(legend.position = "bottom") + ggplot2::guides(color = ggplot2::guide_legend(nrow = 1)) ``` ### 场景 4:中文显示问题(方块/乱码/缺字) 推荐:按“是否包含中文字符”选择字体族,并允许用户在 `.Rmd` 的 `params` 中覆盖。 ```r has_han <- grepl("\\p{Han}", paste(df$label, collapse = ""), perl = TRUE) base_family <- if (has_han) "sans" else "Arial" p <- p + theme_nature_readable(base_size = 10, base_family = base_family) ``` ## 常见反模式 | 反模式 | 问题 | 更好的做法 | |--------|------|------------| | 硬编码图例位置 | 不同数据分布下可能遮挡 | 默认外置(底部/侧边),必要时再微调 | | 固定字体大小 | 图尺寸变化时不可读 | 基于输出场景调整 `base_size` 与图尺寸 | | 显示所有标签 | 数据多时重叠、不可读 | 采样/减少 breaks/只标注关键点 | | 忽略中文字体 | 方块/乱码/缺字 | 显式指定支持中文的字体族(跨平台候选) | -
serial_review_protocol.md 1.9 KB
# 真实运行前的三轮串行审查 ## 固定顺序 主 Agent 完成全部初稿后、首次昂贵或不可逆运行前,依次创建三个新的独立审查者。每轮只读,主 Agent 修正后下一轮再读取最新版本。 ### 1. 结构与数据流 核对需求—target—报告覆盖、DAG 无环、`_targets.R` 与 `R/` 边界、`raw/` 只读、products/reports/_targets 职责、是否误建第二套缓存、旧项目兼容。 ### 2. R 实现与恢复 核对对象生命周期、路径、随机种子、参数归属、输入/代码/上游身份、完成标记、损坏识别、强制重算和从失败边界恢复。 ### 3. 科学与统计 逐项核对主要估计对象、分析单位、目标人群、对比和效应尺度;方法是否匹配配对/聚类/重复测量/删失/权重/缺失设计;CI 水平和界限、检验零假设与尾数、p/q 和多重比较族是否对应同一参数与数据;检查偏差与混杂、数据泄漏、模型诊断、不可估计状态、外推边界及解读是否超过证据。文本检查通过不能替代计算产品与代码复核。 ## 权限与输出 审查者不得修改文件、执行真实昂贵计算、写远程状态或代替主 Agent 做决定。输出只包含:结论(通过/需修正)、问题、证据位置、影响、建议与未能检查项。无问题时明确记录通过,不制造形式性意见。 主 Agent 是唯一修改主体。每轮修正摘要与审查输入版本保存在当前 `.bensz-api` 任务目录,确保下一轮针对最新版本。 ## 能力降级 宿主不能创建独立子 Agent 时: 1. 先告知用户降级原因; 2. 主 Agent 按三个职责分离、串行完成自检; 3. 记录为“顺序自检”,不得写成“独立审查”; 4. 用户明确要求真实多 Agent,或任务包含高风险临床/安全/不可逆运行时,在运行前停止并请求具备能力的环境。 -
statistical_inference_protocol.md 4 KB
# 统计推断决策与结果对应 适用于论文级分析中需要由样本推断目标人群、处理效应或预测性能的关键结果。先确定估计对象与研究设计,再选择方法、计算并解释;不能从一个点估计自动推出 CI 或 p 值。预注册方案、主要终点和预先约定的置信水平优先,不能看过结果后为显著性更换假设或分析人群。 ## 计算前的估计对象清单 对每个主要和影响结论的次要结果,记录:目标人群与分析样本、观察单位、结局/参数、比较对象或参考点、效应尺度与单位、设计及抽样机制、配对/聚类/重复测量/删失/权重、缺失处理、目的(描述/估计/检验/预测)、有效 N/事件数、预定 CI 水平、检验零假设与单/双侧(仅在确有检验问题时)、多重比较族与调整策略、可估计性、方法理由、计算入口和报告去向。simple 可用短表或方法段;complex 可用分析计划中的可选 `inference` 节点。未解决的设计问题先标为待确认,不能默认为独立同分布。 | 对象示例 | 推断决定 | | --- | --- | | 两组独立样本的主要均值差 | 估计差值及单位、各组有效 N;依据设计与分布选相应区间。预先定义零差假设时才给匹配的检验与 p 值。 | | 同一受试者前后差 | 以受试者差值为分析单位,区间和检验保留配对;不能将两列观测当作独立样本。 | | 单组比例或外部验证 AUC | 给适合分母/验证设计的区间;没有明确零假设时不附 p 值。 | | 固定常数、全体已知的描述量 | 标 `not_applicable`,写明无抽样推断目的;不造 CI/p。 | ## 方法与可估计性 优先使用与设计、参数尺度和假设相符的解析区间/检验。若条件不满足,再评估有效的 bootstrap、稳健或聚类方差、置换/精确方法;重抽样单位必须与独立抽样单位一致。配对、重复测量、删失、权重、缺失机制和模型诊断都会改变可用方法。检查小样本、稀疏事件、完全分离、模型不可识别、非独立性、选择偏差和数据泄漏;不能给所有结果套同一种 Wald CI 或 t 检验。预测任务须保持训练与验证分离,区间覆盖整个验证流程需要考虑的重抽样层级。 默认展示 95% CI;方案另有预定水平时按方案执行并标明。检验问题明确时,写出零假设、方向、所用检验和原始 p 值;批量检验同时写清比较族、校正方法及 q 值或调整后 p 值。CI、p 值和显著性不替代效应大小、实际意义或因果识别。CI 与 p 值来自不同模型、样本、尾数或调整策略时不得并列成同一个推断结论。 ## 计算产品与报告 计算层保留未按显著性阈值筛掉的完整推断结果。每个关键参数至少可追溯到:参数身份及比较、点估计、尺度/单位、有效 N/事件数、方法、CI 下上限和置信水平;适用时还包括检验问题、原始 p、校正方法、q/调整后 p、模型或重抽样设置、数据和代码来源。可用表或等价结构,不规定全项目统一 schema。 对每项标明 `estimated`、`not_applicable`、`not_estimable` 或 `not_run`。后三者保留原因;`not_estimable` 写可做的替代分析与结论限制,`not_run` 不写成计算失败或阴性结果。Rmd 只消费计算层的真实结果,用动态引用或可核对来源的值给出“估计值 + CI + 有明确假设时 p/q + 实际意义/局限”。图表附近按关键结论给证据,不为装饰性数字逐个做检验。描述性目标若不需抽样推断,报告定义、分母和边界即可。 ## 复核 逐项核对估计对象清单与完整结果表、报告数字及代码:分析单位和方法是否相符;CI 覆盖的是哪个参数、水平和人群;p 是否回答预定零假设;多重比较是否处理;缺失/删失/权重与诊断是否影响结论;无区间或检验的状态与理由是否诚实。文本检查器只能提示措辞或数字缺口,不能证明这些问题已通过。 -
workflow_checklist.md 4.2 KB
# 工作流检查清单 ## 状态、模式与环境 - [ ] 已分开记录 `project_state` 与 `workflow_mode`,并说明选择理由。 - [ ] 人类显式模式优先;existing 未被按新默认重构。 - [ ] 新 simple 有 renv 和测试入口,无 `_targets.R`;新 complex 额外有 `_targets.R`、`R/`,且 Rmd 消费 target。 - [ ] 已有项目未因检查新增 targets、renv、目录或 runner。 - [ ] 目标、输入、授权、数据字典、统计边界、报告用途和重跑成本已确认。 - [ ] 论文级主要/关键次要估计对象在计算前逐项记录目标人群、分析单位、对比、效应尺度、设计、推断目的、CI/检验可用性及方法理由;预注册方案优先。 ## 目录与产品 - [ ] `raw/` 只读;正式派生产品与 targets store 没有混淆。 - [ ] `templates/` 只有样式/渲染资产;新 pipeline 不复制 checkpoint helper。 - [ ] 测试代码在 `scripts/tests/`;每次运行现场在唯一 `tmp/tests/<run-id>/` 并默认忽略。 - [ ] products 路径来自一个项目设置,留在授权项目范围内,无散落硬编码。 - [ ] `tmp/scratch/` 的正式发现已晋升到代码、产品和报告并补测试。 - [ ] 只创建实际使用的目录;旧 `tmp/` 项目未自动迁移或覆盖。 ## 轻量测试 - [ ] 明确选择 `synthetic_fixture` 或 `project_subset`,并记录理由。 - [ ] fixture/子集覆盖 schema、类型、主键、关键分组、缺失和边缘条件。 - [ ] simple 真实执行正式 R/Rmd 入口;complex 真实执行同一 target 图。 - [ ] complex test store 是 `tmp/tests/<run-id>/_targets`,未接触正式 `_targets/`。 - [ ] 未向业务逻辑加入 `analysis_mode`/`test_mode` 分支。 - [ ] 输入、主键/分组、统计不变量、产品、报告/图表与 raw 不变性断言通过。 - [ ] 每次尝试有命令、退出状态、失败分类、修正、断言和剩余风险。 - [ ] 运行记录绑定最新 subject identity,并区分 preflight/lightweight/full-data;未跑全量时明确 `NOT_RUN`。 - [ ] project subset 已默认删除,或有明确保留授权与访问控制。 - [ ] 审查后若代码、配置或 lockfile 改变,受影响轻量真实链已重跑。 ## targets、审查与报告 - [ ] 完整结果未按显著性筛掉参数;关键项可追溯到估计值、单位/尺度、有效 N/事件数、方法、CI 界限/水平、适用的 p/q、数据与代码来源。 - [ ] `not_applicable`、`not_estimable`、`not_run` 分别写明理由;受设计或数据限制时给替代分析和结论边界,不造区间或显著性结论。 - [ ] 逐项复核 CI/p 是否来自一致的设计、样本与参数;配对/聚类/删失/权重/缺失、模型诊断和多重比较已按实际情况处理。 - [ ] 报告的关键结论与计算结果逐项相符;静态文本检查只作启发式预检。 - [ ] `_targets/` 由 targets 管理;`products/` 只保存需要审阅、复用或交付的科学对象。 - [ ] 中断后再次 `tar_make()` 跳过仍有效前序 target;store/输入/代码/参数变化会正确失效。 - [ ] Rmd 通过 `tar_read()`/`tar_load()` 消费结果,不从 raw/ 重做昂贵计算。 - [ ] 三项只读审查按序完成,或明确记录非独立降级。 - [ ] PDF/JPG、HTML widget、图表/表格解读覆盖、数字追溯均通过。 ## 交付命令 ```bash # simple python3 <skill-root>/scripts/check_targets_renv.py <项目根> --project-state new --workflow-mode simple Rscript -e 'renv::status()' Rscript <项目根>/scripts/tests/smoke_test.R # complex/pipeline:smoke_test 内部使用 tmp/tests/<run-id>/_targets python3 <skill-root>/scripts/check_targets_renv.py <项目根> --project-state new --workflow-mode complex python3 <skill-root>/scripts/check_pipeline_contract.py <项目根> Rscript -e 'renv::status()' Rscript <项目根>/scripts/tests/smoke_test.R Rscript -e 'targets::tar_make()' # existing(有 _targets.R 时 auto 解析为 complex,否则 simple) python3 <skill-root>/scripts/check_targets_renv.py <项目根> --project-state existing --workflow-mode auto ``` 所有 Rmd 继续执行解读覆盖、解读质量、widget 可见性和图表可读性检查。checker/`--dry-run` 不替代上面的真实轻量运行。 -
workflow_modes.md 3.7 KB
# 项目状态与工作流模式 ## 两个判断维度 先判 `project_state`,再判 `workflow_mode`。工作流模式只有 `simple` 与 `complex`;`complex` 在文档中也称 **pipeline**,不是另一套模式。`existing` 是项目状态,不是第三种工作流模式。 | 维度 | 取值 | 含义 | | --- | --- | --- | | 项目状态 | `new` / `existing` | 决定能否采用新布局;existing 不自动迁移或补齐 | | 工作流模式 | `simple` / `complex` | complex = targets-first pipeline | 决策优先级固定为:人类显式指定 → existing 不自动迁移边界 → AI 根据依赖、重算成本、复用、恢复与并行需求选择。选择结果和理由进入分析计划或交付摘要。 ## simple 适合线性、低成本、整体重跑可接受的小分析或单报告: - 必须有 `renv.lock` 与 `renv/activate.R`; - 不创建 `_targets.R` 或 `_targets/`; - 正式入口是明确的 Rscript、`rmarkdown::render()` 或项目操作脚本; - smoke test 调用同一正式入口并真实执行; - 不开发自制 runner、checkpoint 或缓存命中协议。 ## complex / pipeline 出现非线性依赖、昂贵步骤、多下游复用、局部失效、断点恢复、血缘或并行需求时选择 complex。其唯一计算组织方式是: ```text renv.lock + renv/activate.R _targets.R # 唯一 DAG、依赖与增量重建入口 R/ # target 调用的计算函数 Rmd # tar_read()/tar_load() 消费结果并负责科学沟通 _targets/ # targets 自己管理的机器计算状态 products/ # 需要审阅、复用或交付的科学对象(可选) reports/ # 图、表、HTML 与补充材料 ``` 不得再为 pipeline 生成编号执行图、`SUCCESS`、identity hash、force-step/resume runner 或自定义 checkpoint helper。`products/` 不是第二个缓存系统;只有科学上需要持久化的对象才导出。 `_targets.R` 的人类可读性契约:target 按阶段用注释块分组(输入/计算/交付),target 一律语义化命名;交付摘要附 `targets::tar_manifest()` 快照,作为人类可读的流程地图。 普通 `tar_make()` 是默认执行方式。只有存在足够多相互独立且昂贵的 target 时才增加 `crew` controller,并同时限制 worker 与内部线程,防止 CPU/内存超卖。 ## existing 项目 已有任一 R/Rmd、raw、products、reports、`_targets.R`、renv 或历史产品的项目判为 `existing`;只读盘点、披露观察到的机制,不自动补建、迁移或删除: - 已有 `_targets.R`:按 complex 契约维护现有 DAG 与正式 store。 - 无 targets(含历史编号脚本):按 simple 语义维护。编号脚本只是顺序执行的普通 Rscript 入口,本 Skill 不提供、不维护任何编号 runner、编号单元检查器或第二套执行编排。 - 任务出现非线性依赖、昂贵步骤、恢复或复用等复杂信号时,先给出旧→新映射、结果比对和回退入口的显式迁移方案,经人类授权后再迁移到 complex;缺失 renv/targets 机制只披露风险,不隐式初始化。 ## 检查入口 ```bash python3 <skill-root>/scripts/check_targets_renv.py <project> \ --project-state new --workflow-mode simple python3 <skill-root>/scripts/check_targets_renv.py <project> \ --project-state new --workflow-mode complex python3 <skill-root>/scripts/check_targets_renv.py <project> \ --project-state existing --workflow-mode auto python3 <skill-root>/scripts/check_pipeline_contract.py <project> ``` 检查器只读:simple 拒绝 targets 入口;complex 要求 `_targets.R`、`R/` 与隔离 test store,并检查 Rmd 是否消费 target;existing 的缺失机制降级为 warning,不因缺失新机制失败,模式契约冲突仍报 error。
-
-
scripts
-
bootstrap_liquid_glass.py 3.8 KB
#!/usr/bin/env python3 """ Bootstrap Liquid Glass assets into a new project. Why: - The default Rmd template assumes the project root contains: - templates/liquid_glass_theme.css - templates/liquid_glass_lightbox.html (lightbox + view state persistence) - Optionally, it also benefits from: - 00.Environment.R (project root) - templates/datatables_helper.R, templates/nature_theme.R, etc. - opt-in _targets.R and R/ starter via --with-pipeline This script copies those files from the installed skill directory into the current working directory (or --project-root). Design: - Deterministic, no network. - Path-aware: locate skill root via __file__ so it works after installation. - Safe-by-default: never overwrite unless --force. """ from __future__ import annotations import argparse import shutil import sys from pathlib import Path def _skill_root() -> Path: # scripts/bootstrap_liquid_glass.py -> skill_root return Path(__file__).resolve().parents[1] def _copy(src: Path, dst: Path, *, force: bool) -> tuple[bool, str]: if not src.exists(): return False, f"missing source: {src}" dst.parent.mkdir(parents=True, exist_ok=True) if dst.exists() and not force: return False, f"skip exists: {dst}" shutil.copy2(src, dst) return True, f"copied: {dst}" def main(argv: list[str]) -> int: parser = argparse.ArgumentParser( description="Copy Liquid Glass theme assets into a project root.", ) parser.add_argument( "--project-root", default=".", help="Target project root directory (default: current directory).", ) parser.add_argument( "--with-env", action="store_true", help='Also copy "00.Environment.R" into project root (if missing).', ) parser.add_argument( "--with-extras", action="store_true", help="Also copy common helper templates (DT/theme/etc.).", ) parser.add_argument( "--with-pipeline", action="store_true", help="Also copy the opt-in targets-first _targets.R and R/ function starter.", ) parser.add_argument( "--force", action="store_true", help="Overwrite existing files.", ) args = parser.parse_args(argv) skill_root = _skill_root() src_templates = skill_root / "templates" project_root = Path(args.project_root).expanduser().resolve() tasks: list[tuple[Path, Path]] = [ (src_templates / "liquid_glass_theme.css", project_root / "templates" / "liquid_glass_theme.css"), (src_templates / "liquid_glass_lightbox.html", project_root / "templates" / "liquid_glass_lightbox.html"), ] if args.with_env: tasks.append((src_templates / "00.Environment.R", project_root / "00.Environment.R")) if args.with_extras: for name in [ "datatables_helper.R", "nature_colors.R", "nature_theme.R", "complexheatmap_template.R", "plotly_template.R", ]: tasks.append((src_templates / name, project_root / "templates" / name)) if args.with_pipeline: # New projects use targets-managed state; the legacy checkpoint helper # has been removed and is never copied. tasks.extend( [ (src_templates / "_targets.R", project_root / "_targets.R"), (src_templates / "R_data_template.R", project_root / "R" / "analysis_functions.R"), ] ) tasks.append((src_templates / "analysis_plan_template.yaml", project_root / "analysis-plan.yaml")) ok = 0 for src, dst in tasks: success, msg = _copy(src, dst, force=args.force) print(msg) ok += int(success) print(f"done: {ok}/{len(tasks)} copied") return 0 if __name__ == "__main__": raise SystemExit(main(sys.argv[1:])) -
check_figure_table_interpretation.py 19.8 KB
#!/usr/bin/env python3 """ Hard-coded coverage check: ensure every *visible* figure/table output chunk in a .Rmd has a nearby, evidence-anchored interpretation in Markdown prose. Why this exists: - Existing checks focus on interpretation *quality* for already-written prose. - This script focuses on interpretation *coverage*: "there is output, there is interpretation". Design goals (KISS): - No knitr execution, purely static. - Heuristic detection (best-effort), configurable via config.yaml. - Fail fast in --strict mode to block delivery when coverage is incomplete. """ from __future__ import annotations import argparse import json import re import sys from dataclasses import dataclass from pathlib import Path from typing import Any, Iterable FENCE_START_RE = re.compile(r"^\s*```{r\b(?P<header>[^}]*)}\s*$") FENCE_END_RE = re.compile(r"^\s*```\s*$") NEXT_FENCE_ANY_RE = re.compile(r"^\s*```") HEADING_RE = re.compile(r"^\s*#{1,6}\s+") FIG_REF_RE = re.compile(r"(?:Figure|Fig|图)\s*(\d+[a-zA-Z]?)", re.IGNORECASE) TAB_REF_RE = re.compile(r"(?:Table|表)\s*(\d+[a-zA-Z]?)", re.IGNORECASE) HAS_CJK_RE = re.compile(r"[\u4e00-\u9fff]") WORD_RE = re.compile(r"[A-Za-z]+") _CONFIG_WARNED = False def _skill_root() -> Path: # scripts/xxx.py -> {skill_root}/scripts/xxx.py return Path(__file__).resolve().parents[1] def _warn_once(msg: str) -> None: global _CONFIG_WARNED if _CONFIG_WARNED: return _CONFIG_WARNED = True print(msg, file=sys.stderr) def _load_yaml_config() -> dict[str, Any]: config_path = _skill_root() / "config.yaml" if not config_path.exists(): return {} try: import yaml # type: ignore except Exception: _warn_once( "[WARN] PyYAML not available; config.yaml will be ignored and defaults will be used. " "Install PyYAML to enable configurable checks." ) return {} try: cfg = yaml.safe_load(config_path.read_text(encoding="utf-8")) or {} if not isinstance(cfg, dict): _warn_once("[WARN] config.yaml parsed but is not a mapping; ignoring config.") return {} return cfg except Exception as e: _warn_once(f"[WARN] Failed to parse config.yaml; ignoring config and using defaults. Error: {e}") return {} def _as_list(v: Any) -> list[str]: if v is None: return [] if isinstance(v, list): return [str(x) for x in v] return [str(v)] @dataclass(frozen=True) class Chunk: start_line: int end_line: int header: str label: str | None options_raw: str code: str @dataclass(frozen=True) class OutputItem: kind: str # "figure" | "table" | "mixed" start_line: int end_line: int chunk_label: str | None reason: str @dataclass(frozen=True) class MatchResult: output: OutputItem ok: bool details: str interpretation_preview: str | None def _split_header_parts(header: str) -> list[str]: # Best-effort split by comma, respecting simple quotes. parts: list[str] = [] buf: list[str] = [] in_squote = False in_dquote = False for ch in header: if ch == "'" and not in_dquote: in_squote = not in_squote elif ch == '"' and not in_squote: in_dquote = not in_dquote if ch == "," and not in_squote and not in_dquote: part = "".join(buf).strip() if part: parts.append(part) buf = [] continue buf.append(ch) tail = "".join(buf).strip() if tail: parts.append(tail) return parts def _parse_chunk_header(header: str) -> tuple[str | None, dict[str, str]]: parts = _split_header_parts(header.strip()) label: str | None = None options: dict[str, str] = {} for idx, p in enumerate(parts): if "=" in p: k, v = p.split("=", 1) options[k.strip()] = v.strip() continue # First non k=v token is commonly the chunk label. if idx == 0 and p and p.lower() != "r": label = p.strip() return label, options def _is_hidden_chunk(options: dict[str, str]) -> bool: # Coverage check is for *visible* outputs in the rendered report. def _truthy(v: str) -> bool: return v.strip().lower() in {"true", "t", "1", "yes"} def _norm(v: str) -> str: return v.strip().strip("'\"").lower() # Escape hatch: for rare cases where a chunk matches output patterns # but is intentionally non-output (e.g. constructing a ggplot object only). if "interp_check" in options and not _truthy(options["interp_check"]): return True if "eval" in options and not _truthy(options["eval"]): return True if "include" in options and not _truthy(options["include"]): return True # Plots can be explicitly hidden without disabling execution. if "fig.show" in options and "hide" in _norm(options["fig.show"]): return True if "fig.keep" in options and "none" in _norm(options["fig.keep"]): return True return False def parse_rmd_chunks(text: str) -> list[Chunk]: lines = text.splitlines() chunks: list[Chunk] = [] in_chunk = False header = "" start_line = 0 code_lines: list[str] = [] for i, line in enumerate(lines, start=1): if not in_chunk: m = FENCE_START_RE.match(line) if m: in_chunk = True header = m.group("header") or "" start_line = i code_lines = [] continue # in_chunk if FENCE_END_RE.match(line): in_chunk = False label, options = _parse_chunk_header(header) chunks.append( Chunk( start_line=start_line, end_line=i, header=header, label=label, options_raw=header, code="\n".join(code_lines).strip("\n"), ) ) header = "" start_line = 0 code_lines = [] continue code_lines.append(line) return chunks def _compile_any(patterns: Iterable[str]) -> list[re.Pattern[str]]: out: list[re.Pattern[str]] = [] for p in patterns: try: out.append(re.compile(p, re.IGNORECASE)) except re.error: # Ignore broken patterns in config rather than crash. continue return out def detect_outputs(chunks: list[Chunk], cfg: dict[str, Any]) -> list[OutputItem]: section = (cfg.get("figure_interpretation_check") or {}) if isinstance(cfg, dict) else {} fig_pats = _as_list(((section.get("check_patterns") or {}).get("figure_generation"))) tab_pats = _as_list(((section.get("check_patterns") or {}).get("table_generation"))) # Conservative defaults to reduce false negatives while staying readable. if not fig_pats: fig_pats = [ r"\bggplot\s*\(", r"\bgeom_[a-zA-Z0-9_]+\s*\(", r"\bplot\s*\(", r"\bheatmap\b", r"\bComplexHeatmap\b", r"\bHeatmap\s*\(", r"\bpheatmap\s*\(", r"\bggsave\s*\(", r"\bpdf\s*\(", r"\bpng\s*\(", r"\btiff\s*\(", r"\bjpeg\s*\(", ] if not tab_pats: tab_pats = [ r"\bknitr::kable\s*\(", r"\bkable\s*\(", r"\bDT::datatable\s*\(", r"\bgt::gt\s*\(", r"\bflextable::flextable\s*\(", r"\breactable::reactable\s*\(", ] fig_re = _compile_any(fig_pats) tab_re = _compile_any(tab_pats) outputs: list[OutputItem] = [] for c in chunks: _label, options = _parse_chunk_header(c.header) if _is_hidden_chunk(options): continue code = c.code fig_hit = next((r.pattern for r in fig_re if r.search(code)), None) tab_hit = next((r.pattern for r in tab_re if r.search(code)), None) # Also treat chunks with fig.cap/fig.caption as figure outputs. if not fig_hit: for k in ("fig.cap", "fig.caption"): if k in options: fig_hit = f"chunk_option:{k}" break if not fig_hit and not tab_hit: continue kind = "mixed" if (fig_hit and tab_hit) else ("figure" if fig_hit else "table") reason = fig_hit or tab_hit or "pattern" outputs.append( OutputItem( kind=kind, start_line=c.start_line, end_line=c.end_line, chunk_label=c.label, reason=reason, ) ) return outputs def _extract_prose_window(lines: list[str], start_line: int, max_lines: int) -> tuple[str, int]: """ Return (prose_text, end_line) starting *after* start_line, stopping at next fence or max_lines, whichever comes first. """ start_idx = min(len(lines), start_line) # 1-based line -> 0-based idx after it end_idx = min(len(lines), start_idx + max_lines) buf: list[str] = [] last_line_num = start_line for idx in range(start_idx, end_idx): line = lines[idx] line_num = idx + 1 if NEXT_FENCE_ANY_RE.match(line): break buf.append(line) last_line_num = line_num return "\n".join(buf).strip(), last_line_num def _is_interpretation_valid( prose: str, *, required_markers: list[re.Pattern[str]], require_ref: bool, min_cjk_chars: int, min_en_words: int, min_content_elements: int, ) -> tuple[bool, str]: # Remove headings-only / whitespace-only. raw_lines = [ln.rstrip() for ln in prose.splitlines()] lines = [ln for ln in raw_lines if ln.strip()] if not lines: return False, "no prose after output chunk" non_heading = [ln for ln in lines if not HEADING_RE.match(ln)] text = "\n".join(non_heading).strip() if not text: return False, "only headings found after output chunk" # Length check (language-aware-ish). has_cjk = bool(HAS_CJK_RE.search(text)) cjk_chars = sum(1 for ch in text if HAS_CJK_RE.match(ch)) en_words = len(WORD_RE.findall(text)) if has_cjk and cjk_chars < min_cjk_chars: return False, f"interpretation too short (CJK chars {cjk_chars} < {min_cjk_chars})" if (not has_cjk) and en_words < min_en_words: return False, f"interpretation too short (EN words {en_words} < {min_en_words})" # Reference check (optional, but recommended). if require_ref: if not (FIG_REF_RE.search(text) or TAB_REF_RE.search(text)): return False, "missing explicit figure/table reference (e.g., 'Figure 1'/'图 1')" # Marker check (e.g. '解读/结论/观察' markers). if required_markers: if not any(r.search(text) for r in required_markers): return False, "missing interpretation markers (e.g., '解读/结论/小结/观察')" # Content elements: require at least N evidence-anchoring elements. elements = 0 if re.search(r"展示|显示|show|display", text, re.IGNORECASE): elements += 1 if re.search(r"高于|低于|差异|difference|higher|lower|increase|decrease", text, re.IGNORECASE): elements += 1 if re.search(r"p\\s*[<>=]\\s*0\\.\\d+|显著|significant|FDR|q\\s*[<>=]", text, re.IGNORECASE): elements += 1 if re.search(r"提示|表明|suggest|indicat", text, re.IGNORECASE): elements += 1 if elements < min_content_elements: return False, f"content elements insufficient ({elements} < {min_content_elements})" return True, "ok" def check_coverage( rmd_path: Path, *, max_distance_lines: int, strict_reference: bool, interpretation_marker_patterns: list[str], min_cjk_chars: int, min_en_words: int, min_content_elements: int, ) -> dict[str, Any]: text = rmd_path.read_text(encoding="utf-8", errors="replace") lines = text.splitlines() chunks = parse_rmd_chunks(text) cfg = _load_yaml_config() outputs = detect_outputs(chunks, cfg) marker_res = _compile_any(interpretation_marker_patterns) results: list[MatchResult] = [] for out in outputs: prose, _ = _extract_prose_window(lines, out.end_line, max_distance_lines) ok, details = _is_interpretation_valid( prose, required_markers=marker_res, require_ref=strict_reference, min_cjk_chars=min_cjk_chars, min_en_words=min_en_words, min_content_elements=min_content_elements, ) preview = None if prose: preview = "\n".join([ln for ln in prose.splitlines() if ln.strip()][:8]).strip() results.append( MatchResult( output=out, ok=ok, details=details, interpretation_preview=preview, ) ) unmatched = [r for r in results if not r.ok] return { "file": str(rmd_path), "total_outputs": len(outputs), "unmatched_outputs": len(unmatched), "pass": len(unmatched) == 0, "unmatched_details": [ { "kind": r.output.kind, "start_line": r.output.start_line, "end_line": r.output.end_line, "chunk_label": r.output.chunk_label, "reason": r.output.reason, "details": r.details, "interpretation_preview": r.interpretation_preview, } for r in unmatched ], } def main(argv: list[str]) -> int: p = argparse.ArgumentParser( description="Coverage check for figure/table interpretations in .Rmd files." ) p.add_argument("rmd_file", type=Path, help="Path to a .Rmd file") p.add_argument("--strict", action="store_true", help="Exit non-zero if check fails") p.add_argument("--json", action="store_true", help="Print JSON report") p.add_argument( "--max-distance-lines", type=int, default=None, help="Max allowed line distance from output chunk end to interpretation prose (default: from config.yaml or 50)", ) p.add_argument( "--require-reference", action="store_true", help="Require explicit 'Figure/Table/图/表 N' reference in interpretation prose (stricter, fewer false passes).", ) g = p.add_mutually_exclusive_group() g.add_argument( "--require-markers", action="store_true", default=None, help="Require interpretation markers (default: follow config.yaml).", ) g.add_argument( "--no-require-markers", action="store_true", default=None, help="Disable interpretation marker requirement (default: follow config.yaml).", ) p.add_argument( "--min-cjk-chars", type=int, default=None, help="Minimum CJK char count for interpretation prose when CJK is detected (default: from config.yaml or 100)", ) p.add_argument( "--min-en-words", type=int, default=None, help="Minimum English word count for interpretation prose when no CJK is detected (default: from config.yaml or 50)", ) p.add_argument( "--min-content-elements", type=int, default=None, help="Require at least N evidence-anchoring elements in interpretation prose (default: from config.yaml or 2)", ) args = p.parse_args(argv) if not args.rmd_file.exists(): print(f"[FAIL] file not found: {args.rmd_file}", file=sys.stderr) return 1 # Allow config.yaml to set defaults while keeping CLI explicit override. cfg = _load_yaml_config() sec = cfg.get("figure_interpretation_check") if isinstance(cfg, dict) else None sec = sec if isinstance(sec, dict) else {} enabled = bool(sec.get("enabled", True)) if not enabled: # Explicitly disabled -> always pass (no-op). report = { "file": str(args.rmd_file), "total_outputs": 0, "unmatched_outputs": 0, "pass": True, "note": "figure_interpretation_check.disabled", } if args.json: print(json.dumps(report, ensure_ascii=False, indent=2)) else: print("[PASS] figure/table interpretation coverage check is disabled by config.yaml") return 0 check_patterns = sec.get("check_patterns") if isinstance(sec.get("check_patterns"), dict) else {} has_marker_key = "interpretation_markers" in check_patterns marker_patterns = _as_list(check_patterns.get("interpretation_markers")) if (not marker_patterns) and (not has_marker_key): # Neutral fallback markers (avoid forcing tier labels). marker_patterns = [ r"解读", r"结论", r"小结", r"观察", r"核心结论", r"结果(?:显示|表明|提示|可见)", r"\binterpretation\b", r"\btakeaway\b", r"(?:图|表|Figure|Table)\s*\d+.*?(?:解读|分析|结果|interpretation|analysis)", ] # Config defaults (CLI still overrides via flags). max_distance = int(sec.get("max_distance_lines", 50)) if args.max_distance_lines is None else int(args.max_distance_lines) min_layers = int(sec.get("required_layers", 4)) # informational; not enforced directly _ = min_layers strict_mode = bool(sec.get("strict_mode", False)) require_ref_default = bool(sec.get("require_reference", False)) if args.require_markers is True: require_markers = True elif args.no_require_markers is True: require_markers = False else: require_markers = bool(sec.get("require_markers", True)) # If require_markers=false OR marker list empty -> skip marker check entirely. if (not require_markers) or (not marker_patterns): marker_patterns = [] min_cjk_chars = ( int(sec.get("min_cjk_chars", 100)) if args.min_cjk_chars is None else int(args.min_cjk_chars) ) min_en_words = int(sec.get("min_en_words", 50)) if args.min_en_words is None else int(args.min_en_words) min_content_elements = ( int(sec.get("min_content_elements", 2)) if args.min_content_elements is None else int(args.min_content_elements) ) report = check_coverage( args.rmd_file, max_distance_lines=max_distance, strict_reference=bool(args.require_reference or require_ref_default), interpretation_marker_patterns=marker_patterns, min_cjk_chars=min_cjk_chars, min_en_words=min_en_words, min_content_elements=min_content_elements, ) if args.json: print(json.dumps(report, ensure_ascii=False, indent=2)) else: print("图表/表格-解读覆盖检验报告") print(f"文件: {report['file']}") print("-" * 60) print(f"检测到的输出块数: {report['total_outputs']}") print(f"未通过块数: {report['unmatched_outputs']}") if report["unmatched_outputs"]: print("\n未通过明细(需要补充/加强解读):") for d in report["unmatched_details"]: lab = f" chunk={d['chunk_label']}" if d.get("chunk_label") else "" print( f"- {d['kind']} L{d['start_line']}-L{d['end_line']}{lab}: {d['details']}" ) if d.get("interpretation_preview"): prev = str(d["interpretation_preview"]).strip().replace("\t", " ") if prev: print(" 现有解读片段(预览):") for ln in prev.splitlines()[:6]: print(f" {ln}") print("\n" + ("[PASS] 通过" if report["pass"] else "[FAIL] 未通过")) should_fail = (args.strict or strict_mode) and (not report["pass"]) return 2 if should_fail else 0 if __name__ == "__main__": raise SystemExit(main(sys.argv[1:])) -
check_htmlwidget_visibility.py 10.2 KB
#!/usr/bin/env python3 """ Static checker for R Markdown (.Rmd) htmlwidget visibility issues. Goal: prevent the common failure mode "code chunk shows, but HTML doesn't render the widget" when DT::datatable()/plotly/etc. are wrapped in print()/invisible() or are not returned as the chunk's visible result. This is intentionally heuristic (fast, no R execution). It focuses on high-signal patterns. """ from __future__ import annotations import argparse import re import sys from dataclasses import dataclass from pathlib import Path _CHUNK_START_RE = re.compile(r"^\s*```+\s*\{r([^}]*)\}\s*$") _CHUNK_END_RE = re.compile(r"^\s*```+\s*$") # Common htmlwidget constructors / wrappers we want to keep visible in HTML. _WIDGET_CALL_RE = re.compile( r"(?P<call>" r"DT::datatable\s*\(" r"|render_dt_output\s*\(" r"|render_dt\s*\(" r"|plotly::" r"|leaflet::" r"|reactable::" r"|highcharter::" r"|visNetwork::" r"|dygraphs::" r"|htmlwidgets::" r")" ) _BAD_WRAPPER_RE = re.compile( r"\b(?P<wrapper>print|invisible|suppressMessages|suppressWarnings)\s*\(" ) _ASSIGN_WIDGET_RE = re.compile( r"^\s*(?P<var>[.A-Za-z][\w.]*)\s*<-\s*.*(" r"DT::datatable\s*\(" r"|render_dt_output\s*\(" r"|render_dt\s*\(" r"|plotly::" r"|leaflet::" r"|reactable::" r"|highcharter::" r"|visNetwork::" r"|dygraphs::" r"|htmlwidgets::" r")" ) _TAGLIST_RE = re.compile(r"\b(htmltools::)?tagList\s*\(") @dataclass(frozen=True) class Finding: severity: str # "ERROR" | "WARN" path: Path line: int chunk: str message: str def _chunk_label(chunk_header: str) -> str: # header like: " data-preview, message=FALSE" # We consider the first token as the label if it's not an option assignment. tokens = [t.strip() for t in chunk_header.strip().lstrip().split(",") if t.strip()] if not tokens: return "<unnamed>" first = tokens[0] if "=" in first: return "<unnamed>" return first def _is_comment_or_blank(line: str) -> bool: s = line.strip() return (not s) or s.startswith("#") def _last_meaningful_line(lines: list[str]) -> tuple[int, str] | None: for i in range(len(lines) - 1, -1, -1): if _is_comment_or_blank(lines[i]): continue return i, lines[i].strip() return None def _has_bad_wrapper(line: str) -> bool: # High-signal: print(invisible(...)) etc around a widget call on the same line. if not _BAD_WRAPPER_RE.search(line): return False return _WIDGET_CALL_RE.search(line) is not None def _is_assigned_widget_line(line: str) -> bool: return _ASSIGN_WIDGET_RE.search(line) is not None _R_STRING_RE = re.compile( r"(" # Replace strings to avoid counting parentheses inside them. r"\"(?:[^\"\\\n]|\\.)*\"" # "..." r"|'(?:[^'\\\n]|\\.)*'" # '...' r"|`[^`\n]*`" # `...` r")" ) def _strip_strings_and_comments(line: str) -> str: # Remove R comments (heuristic: everything after #) after stripping string literals. # This is intentionally lightweight; we only need stable parenthesis counting. no_str = _R_STRING_RE.sub('""', line) return no_str.split("#", 1)[0] def _find_expr_end_idx0(lines: list[str], start_idx0: int) -> int | None: """ Find the last (meaningful) line index of the expression starting at start_idx0. We use parenthesis depth counting across lines. This is heuristic but works well for common multi-line function calls like: DT::datatable( ..., options = list(...) ) """ depth = 0 seen = False last_meaningful_idx0: int | None = None for j in range(start_idx0, len(lines)): raw = lines[j] if _is_comment_or_blank(raw): continue last_meaningful_idx0 = j s = _strip_strings_and_comments(raw) opens = s.count("(") closes = s.count(")") if opens or closes: seen = True depth += opens - closes # When we've returned to depth 0, the expression is closed. if seen and depth <= 0: return last_meaningful_idx0 return None def check_rmd(path: Path) -> list[Finding]: findings: list[Finding] = [] text = path.read_text(encoding="utf-8") lines = text.splitlines() in_chunk = False chunk_header = "" chunk_name = "<outside>" chunk_start_line = 0 # 1-based chunk_lines: list[str] = [] def flush_chunk() -> None: nonlocal chunk_lines if not in_chunk: chunk_lines = [] return assigned_widget_vars: set[str] = set() naked_widget_calls: list[tuple[int, str]] = [] # (0-based chunk line idx, line) for idx0, l in enumerate(chunk_lines): if _has_bad_wrapper(l): findings.append( Finding( severity="ERROR", path=path, line=chunk_start_line + idx0, chunk=chunk_name, message=( "htmlwidget 被 print()/invisible()/suppress*() 包裹,HTML 可能不渲染;" "请让 widget 作为 chunk 的可见返回值输出(不要包裹)。" ), ) ) m_assign = _ASSIGN_WIDGET_RE.search(l) if m_assign: assigned_widget_vars.add(m_assign.group("var")) if _WIDGET_CALL_RE.search(l): if _is_assigned_widget_line(l): continue # Heuristic: if '<-' appears before the match, treat it as assigned (avoid FP). m_call = _WIDGET_CALL_RE.search(l) assert m_call is not None before = l[: m_call.start()] if "<-" in before: continue naked_widget_calls.append((idx0, l)) last = _last_meaningful_line(chunk_lines) if last is None: chunk_lines = [] return last_idx0, last_line = last # Rule: if we have naked widget calls, the chunk must end by returning a widget / tagList / widget var. if naked_widget_calls: ok = False if _WIDGET_CALL_RE.search(last_line): ok = True elif _TAGLIST_RE.search(last_line): ok = True elif last_line in assigned_widget_vars: ok = True else: # Multi-line widget call: the last meaningful line might be just ")". In that case, # check whether the last expression *started* with a naked widget call and ends at last_idx0. for idx0, _ in reversed(naked_widget_calls): end_idx0 = _find_expr_end_idx0(chunk_lines, idx0) if end_idx0 is not None and end_idx0 == last_idx0: ok = True break if not ok: first_idx0, _ = naked_widget_calls[0] findings.append( Finding( severity="ERROR", path=path, line=chunk_start_line + first_idx0, chunk=chunk_name, message=( "检测到 htmlwidget 调用,但该 chunk 的最后表达式不是 widget/tagList/已赋值的 widget 变量;" "这通常会导致 HTML 不出表/不出图。" ), ) ) # Rule: multiple naked widgets must be returned via tagList. if len(naked_widget_calls) > 1 and not _TAGLIST_RE.search(last_line): first_idx0, _ = naked_widget_calls[0] findings.append( Finding( severity="ERROR", path=path, line=chunk_start_line + first_idx0, chunk=chunk_name, message=( "同一 chunk 内检测到多个 htmlwidget 需要展示,但末尾未使用 htmltools::tagList(...) 统一返回;" "通常只会渲染最后一个。" ), ) ) chunk_lines = [] for i, line in enumerate(lines, start=1): if not in_chunk: m = _CHUNK_START_RE.match(line) if m: in_chunk = True chunk_header = m.group(1) or "" chunk_name = _chunk_label(chunk_header) chunk_start_line = i + 1 # code starts next line chunk_lines = [] continue # in chunk if _CHUNK_END_RE.match(line): flush_chunk() in_chunk = False chunk_header = "" chunk_name = "<outside>" chunk_start_line = 0 chunk_lines = [] continue chunk_lines.append(line) # Unclosed chunk (still analyze). if in_chunk: flush_chunk() return findings def main(argv: list[str]) -> int: parser = argparse.ArgumentParser( description="Check Rmd htmlwidget visibility pitfalls (DT/plotly/etc.)", ) parser.add_argument("paths", nargs="+", help="Rmd file(s) to check") args = parser.parse_args(argv) all_findings: list[Finding] = [] for p in args.paths: path = Path(p) if not path.exists(): print(f"ERROR: not found: {path}", file=sys.stderr) return 2 if path.is_dir(): # Keep it simple: only *.Rmd directly under the directory (no recursion by default). for f in sorted(path.glob("*.Rmd")): all_findings.extend(check_rmd(f)) else: all_findings.extend(check_rmd(path)) if not all_findings: return 0 # Stable output order. all_findings.sort(key=lambda f: (str(f.path), f.line, f.severity)) for f in all_findings: loc = f"{f.path}:{f.line}" chunk = f"[chunk: {f.chunk}]" print(f"{f.severity}: {loc} {chunk} {f.message}") has_errors = any(f.severity == "ERROR" for f in all_findings) return 2 if has_errors else 1 if __name__ == "__main__": raise SystemExit(main(sys.argv[1:])) -
check_interpretation_quality.py 29.2 KB
#!/usr/bin/env python3 """ Lightweight static checks for Rmd interpretation text. Goal: detect obvious "non-evidence-anchored" patterns before rendering/review. This is intentionally heuristic (not 100% accurate). """ from __future__ import annotations import argparse import re import sys from dataclasses import dataclass from pathlib import Path FENCE_RE = re.compile(r"^\s*```") YAML_DELIM_RE = re.compile(r"^\s*---\s*$") HEADING_RE = re.compile(r"^\s*(#{1,6})\s+(.+?)\s*$") # Prefer multi-word phrases to reduce false positives. DEFAULT_BLACKLIST = [ r"建议进一步研究", r"值得进一步研究", r"建议进一步验证", r"需要进一步验证", r"值得深入探讨", r"提示可能", r"可能提示", r"为.*提供依据", ] INLINE_R_RE = re.compile(r"`r\s+[^`]+`") TOP_RE = re.compile(r"\bTop\s*\d+\b|\btop\s*\d+\b|前\s*\d+\b|最强信号") UNCERTAINTY_RE = re.compile( r"\bCI\b|置信区间|\bSE\b|标准误|bootstrap|Bootstr|交叉验证|\bCV\b|稳定性|敏感性分析", re.IGNORECASE, ) BOLD_SPAN_RE = re.compile(r"\*\*[^*\n]+\*\*") NUMBER_RE = re.compile(r"(?<![A-Za-z_])\d+(?:\.\d+)?%?") ENTITY_LIKE_RE = re.compile(r"`[^`\n]+`|\b[A-Za-z][A-Za-z0-9_.-]{2,}\b") INTERP_SECTION_RE = re.compile(r"(结果解读|讨论与分析|讨论|interpretation|discussion)", re.IGNORECASE) INTERP_BOLD_MARKER_RE = re.compile(r"\*\*\s*结果解读\s*\*\*|结果解读\s*[::]") TIER_LABEL_LINE_RE = re.compile( r"^\s*(?:\*\*)?\s*(数据描述|统计见解|领域见解|局限与后续)\s*(?:\*\*)?\s*[::]\s*", re.IGNORECASE, ) ACTION_SENTENCE_RE = re.compile(r"(后续|下一步|建议|可通过|可用|推荐|验证|复核)", re.IGNORECASE) ACTION_METHOD_RE = re.compile( r"(用|采用|通过|进行|复算|复现|验证|检验|调整|校正|分层|多因素|bootstrap|CV)", re.IGNORECASE, ) ACTION_INPUT_RE = re.compile(r"(在|使用|基于|输入|数据|表|列|图|队列|样本|模型|结果)", re.IGNORECASE) ACTION_CRITERION_RE = re.compile( r"(判据|阈值|满足|支持|不支持|若|当|>=|<=|≥|≤|<|>|p\\s*[<=>]|q\\s*[<=>])", re.IGNORECASE, ) STRONG_CLAIM_RE = re.compile(r"(更支持|支持真实|证明|证实|确定|无疑)", re.IGNORECASE) # Mechanical anti-pattern: "fill-the-structure" four-part template. DEFAULT_MECHANICAL_TEMPLATE_PATTERNS = [ ( r"这张(?:图|表|图/表).{0,20}?(?:在)?本次.{0,20}?直接观察是[::].*?" r"统计.{0,20}?含义是[::].*?" r"研究者.{0,20}?意义是[::].*?" r"(?:你可以)?(?:立即)?(?:执行)?的?下一步" ) ] # Vague phrases are not forbidden *per se*; they become failures when they lack # nearby evidence (entity + number/inline-R) and/or executable next steps. DEFAULT_VAGUE_PHRASE_PATTERNS = [ r"提示可能", r"可能提示", r"值得深入探讨", r"需要进一步研究", r"建议进一步研究", r"建议进一步验证", r"需要进一步验证", r"为.*提供依据", ] # Broad markers for "this time / current data" statements. # We combine this with NUMBER/INLINE_R evidence to reduce false positives. CURRENT_OBS_MARKERS_RE = re.compile( r"(本次|当前数据|该结果|结果显示|数据显示|我们观察到|在本次分析中|在当前队列中|\bN\s*=\s*\d+\b|样本量|在\s*\d+\s*(?:个)?\s*(?:样本|患者))", re.IGNORECASE, ) # Expression-style checks (optional, can be noisy). TITLE_PAREN_RE = re.compile(r"[((].+[))]") TITLE_SUBTITLE_RE = re.compile(r"\s-\s") TEACHING_MARKERS = [ r"提示:", r"注意:", r"需要注意的是", r"即:", r"即:", r"用于快速识别", r"用于快速判断", r"帮助你判断", r"用于把方向落到", ] def _iter_non_fenced_non_yaml_lines(text: str) -> list[tuple[int, str]]: """Return (lineno, line) for non-fenced lines, skipping YAML frontmatter if present.""" out: list[tuple[int, str]] = [] in_fence = False in_yaml = False yaml_delims_seen = 0 for i, raw in enumerate(text.splitlines(), start=1): line = raw.rstrip("\n") if i <= 200 and YAML_DELIM_RE.match(line): # Only treat YAML frontmatter when it starts at the beginning of the file. if i == 1 and yaml_delims_seen == 0: in_yaml = True yaml_delims_seen = 1 continue if in_yaml and yaml_delims_seen == 1: in_yaml = False yaml_delims_seen = 2 continue if in_yaml: continue if FENCE_RE.match(line): in_fence = not in_fence continue if in_fence: continue out.append((i, line)) return out def _compile_many(patterns: object) -> list[re.Pattern[str]]: res: list[re.Pattern[str]] = [] if not isinstance(patterns, list): return res for p in patterns: if not isinstance(p, str) or not p.strip(): continue try: res.append(re.compile(p)) except re.error: continue return res def _untracked_number_issues(text: str, *, exempt_res: list[re.Pattern[str]]) -> list[str]: """ Heuristic: report literal numbers in non-code prose lines that do not have inline `r ...`. Returns issue strings with line numbers for actionable fixes. """ issues: list[str] = [] for ln, line in _iter_non_fenced_non_yaml_lines(text): if not line.strip(): continue if INLINE_R_RE.search(line): # Treat inline-R presence on this line as "tracked enough" to avoid over-blocking. continue bad = False for m in NUMBER_RE.finditer(line): window = line[max(0, m.start() - 24) : min(len(line), m.end() + 24)] if any(r.search(window) for r in exempt_res): continue bad = True break if bad: issues.append(f"L{ln}: literal numbers without inline `r ...`: {line.strip()[:200]}") return issues def strip_fenced_blocks(text: str) -> str: """Remove fenced code blocks (``` ... ```), keeping only prose.""" lines: list[str] = [] in_fence = False for line in text.splitlines(): if FENCE_RE.match(line): in_fence = not in_fence continue if not in_fence: lines.append(line) return "\n".join(lines) def split_sentences(prose: str) -> list[str]: parts = re.split(r"[。!?!?]+|\n{2,}", prose) out: list[str] = [] for p in parts: p = p.strip() if not p: continue sub = re.split(r"\n(?=(?:\s*[-*]\s+|\s*\d+[.)]\s+))", p) out.extend([s.strip() for s in sub if s.strip()]) return out def count_current_observations(prose: str) -> int: """ Count sentences that look like "current-data observations": - has a marker like '本次/当前数据/结果显示/...' - AND has numeric evidence (inline `r ...` OR a literal number) """ count = 0 for sent in split_sentences(prose): if not CURRENT_OBS_MARKERS_RE.search(sent): continue if INLINE_R_RE.search(sent) or NUMBER_RE.search(sent): count += 1 return count def _skill_root() -> Path: # scripts/xxx.py -> {skill_root}/scripts/xxx.py return Path(__file__).resolve().parents[1] def _load_yaml_config() -> dict[str, object]: config_path = _skill_root() / "config.yaml" if not config_path.exists(): return {} try: import yaml # type: ignore except Exception: return {} try: cfg = yaml.safe_load(config_path.read_text(encoding="utf-8")) or {} if not isinstance(cfg, dict): return {} return cfg except Exception: return {} def _get_cfg_section(cfg: dict[str, object], key: str) -> dict[str, object]: sec = cfg.get(key, {}) if isinstance(sec, dict): return sec return {} @dataclass(frozen=True) class Block: title: str start_line: int end_line: int prose: str def _non_fenced_lines(text: str) -> list[tuple[int, str]]: out: list[tuple[int, str]] = [] in_fence = False for i, line in enumerate(text.splitlines(), start=1): if FENCE_RE.match(line): in_fence = not in_fence continue if in_fence: continue out.append((i, line)) return out def _extract_heading_blocks(text: str) -> list[Block]: lines = _non_fenced_lines(text) headings: list[tuple[int, int, str]] = [] for line_no, line in lines: m = HEADING_RE.match(line) if not m: continue headings.append((line_no, len(m.group(1)), m.group(2).strip())) if not headings: return [] blocks: list[Block] = [] total_lines = len(text.splitlines()) for idx, (h_line, h_level, h_title) in enumerate(headings): if not INTERP_SECTION_RE.search(h_title): continue start_line = h_line + 1 end_line = total_lines for j in range(idx + 1, len(headings)): n_line, n_level, _ = headings[j] if n_line <= h_line: continue if n_level <= h_level: end_line = n_line - 1 break buf: list[str] = [] for line_no, line in lines: if start_line <= line_no <= end_line: buf.append(line) blocks.append(Block(title=h_title, start_line=start_line, end_line=end_line, prose="\n".join(buf).strip())) return blocks def _extract_bold_marker_blocks(text: str) -> list[Block]: lines = _non_fenced_lines(text) markers = [line_no for line_no, line in lines if INTERP_BOLD_MARKER_RE.search(line)] if not markers: return [] total_lines = len(text.splitlines()) blocks: list[Block] = [] for idx, marker_line in enumerate(markers): start_line = marker_line + 1 end_line = (markers[idx + 1] - 1) if idx + 1 < len(markers) else total_lines buf: list[str] = [] for line_no, line in lines: if start_line <= line_no <= end_line: buf.append(line) blocks.append(Block(title="**结果解读** marker", start_line=start_line, end_line=end_line, prose="\n".join(buf).strip())) return blocks def extract_interpretation_blocks(text: str) -> list[Block]: blocks = _extract_heading_blocks(text) if blocks: return blocks return _extract_bold_marker_blocks(text) def _top_grounded(prose: str) -> bool: for m in TOP_RE.finditer(prose): window = prose[m.start() : min(len(prose), m.end() + 120)] if INLINE_R_RE.search(window) or ENTITY_LIKE_RE.search(window): return True return False def _actionability_issues(prose: str) -> list[str]: issues: list[str] = [] for sent in split_sentences(prose): if not ACTION_SENTENCE_RE.search(sent): continue missing: list[str] = [] if not ACTION_METHOD_RE.search(sent): missing.append("method") if not ACTION_INPUT_RE.search(sent): missing.append("input") if not ACTION_CRITERION_RE.search(sent): missing.append("criterion") if missing: issues.append(f"missing {','.join(missing)}: {sent[:160]}") return issues def _strong_claim_issues(prose: str) -> list[str]: issues: list[str] = [] for sent in split_sentences(prose): if not STRONG_CLAIM_RE.search(sent): continue if (INLINE_R_RE.search(sent) or NUMBER_RE.search(sent)) and UNCERTAINTY_RE.search(sent): continue issues.append(f"strong-claim without evidence+stability guard: {sent[:160]}") return issues def _uncertainty_value_issues(prose: str) -> list[str]: issues: list[str] = [] for sent in split_sentences(prose): if not UNCERTAINTY_RE.search(sent): continue if INLINE_R_RE.search(sent) or NUMBER_RE.search(sent): continue issues.append(sent[:160]) return issues def _mechanical_template_hits(prose: str, patterns: list[str]) -> list[str]: hits: list[str] = [] for pat in patterns: try: if re.search(pat, prose, flags=re.DOTALL): hits.append(pat) except re.error: continue return hits def _vague_phrase_issues( prose: str, patterns: list[str], window_chars: int = 160, require_entity: bool = True ) -> list[str]: issues: list[str] = [] for pat in patterns: try: for m in re.finditer(pat, prose): start = max(0, m.start() - window_chars) end = min(len(prose), m.end() + window_chars) window = prose[start:end] has_number = bool(INLINE_R_RE.search(window) or NUMBER_RE.search(window)) has_entity = bool(ENTITY_LIKE_RE.search(window)) if require_entity: if has_entity and has_number: continue else: if has_number or has_entity: continue snippet = prose[m.start() : min(len(prose), m.end() + 40)].replace("\n", " ").strip() issues.append(f"{pat}: {snippet[:160]}") except re.error: continue return issues def scan_headings(text: str) -> list[dict[str, object]]: """Return heading issues outside fenced code blocks.""" in_fence = False bad: list[dict[str, object]] = [] for i, line in enumerate(text.splitlines(), start=1): if FENCE_RE.match(line): in_fence = not in_fence continue if in_fence: continue m = HEADING_RE.match(line) if not m: continue level = len(m.group(1)) title = m.group(2).strip() issues: list[str] = [] if TITLE_PAREN_RE.search(title): issues.append("parenthetical_note") if TITLE_SUBTITLE_RE.search(title): issues.append("main_subtitle_dash") if issues: bad.append({"line": i, "level": level, "title": title, "issues": issues}) return bad def scan_text(text: str) -> dict[str, object]: blocks = extract_interpretation_blocks(text) if blocks: prose = "\n\n".join(b.prose for b in blocks if b.prose.strip()) mode = "interpretation_blocks" else: prose = strip_fenced_blocks(text) mode = "all_prose" inline_r = INLINE_R_RE.findall(prose) top_hits = TOP_RE.findall(prose) uncertainty_hits = UNCERTAINTY_RE.findall(prose) bold_spans = BOLD_SPAN_RE.findall(prose) current_obs_count = count_current_observations(prose) tier_label_line_count = sum(1 for ln in prose.splitlines() if TIER_LABEL_LINE_RE.search(ln)) action_sentence_count = sum(1 for sent in split_sentences(prose) if ACTION_SENTENCE_RE.search(sent)) blacklist_hits: list[str] = [] for pat in DEFAULT_BLACKLIST: if re.search(pat, prose): blacklist_hits.append(pat) teaching_hits: list[str] = [] for pat in TEACHING_MARKERS: if re.search(pat, prose): teaching_hits.append(pat) bad_headings = scan_headings(text) return { "mode": mode, "_prose": prose, "inline_r_count": len(inline_r), "current_observation_count": current_obs_count, "top_hits": top_hits, "uncertainty_hits": uncertainty_hits, "blacklist_hits": blacklist_hits, "bold_span_count": len(bold_spans), "teaching_hits": teaching_hits, "tier_label_line_count": tier_label_line_count, "action_sentence_count": action_sentence_count, "blocks": blocks, "top_grounded": _top_grounded(prose), "actionability_issues": _actionability_issues(prose), "strong_claim_issues": _strong_claim_issues(prose), "uncertainty_value_issues": _uncertainty_value_issues(prose), "bad_headings": bad_headings, } def main(argv: list[str]) -> int: p = argparse.ArgumentParser( description="Heuristic checks for evidence-anchored interpretation in .Rmd files." ) p.add_argument("files", nargs="+", help="One or more .Rmd files to check.") p.add_argument( "--min-inline-r", type=int, default=None, help="Minimum count of inline R snippets (`r ...`) in prose (default: 3).", ) p.add_argument( "--min-current-observation", type=int, default=None, help=( "Minimum count of 'current data observation' sentences in prose (default: 3). " "A sentence counts if it includes a marker like '本次/当前数据/结果显示' and also includes numeric evidence " "(inline `r ...` or a literal number). Set to 0 to disable." ), ) p.add_argument( "--strict", action="store_true", help=( "Enable stricter anti-template checks: require grounded Top list in interpretation sections, " "flag strong claims without evidence+stability guard, and require stability mentions to include current values." ), ) p.add_argument( "--warn-only", action="store_true", help="Print findings but always exit 0 (useful for CI pre-check).", ) p.add_argument( "--check-title-style", action="store_true", default=None, help="Check headings for '(...)'/'(...)' and ' - ' main-subtitle style (off by default).", ) p.add_argument( "--check-tone", action="store_true", default=None, help="Check for teaching-style markers like '提示:/注意:/需要注意的是' (off by default).", ) p.add_argument( "--max-teaching-hits", type=int, default=None, help="Max allowed teaching-tone marker hits when check-tone is enabled (default: 0).", ) p.add_argument( "--check-actionability", action="store_true", default=None, help="Check that next-step suggestions include method+input+criterion (off by default).", ) p.add_argument( "--max-actionability-issues", type=int, default=None, help="Max allowed actionability issues when check-actionability is enabled (default: 0).", ) p.add_argument( "--check-template-labels", action="store_true", default=None, help="Check for repeated '数据描述/统计见解/...' label-lines (off by default).", ) p.add_argument( "--max-tier-label-lines", type=int, default=None, help="Max allowed tier-label lines when check-template-labels is enabled (default: 8).", ) p.add_argument( "--check-bold", action="store_true", default=None, help="Check for excessive bold spans '**...**' in prose (off by default).", ) p.add_argument( "--max-bold-per-1000", type=float, default=None, help="Max bold spans per 1000 chars when --check-bold is enabled (default: 10).", ) args = p.parse_args(argv) global DEFAULT_BLACKLIST, TEACHING_MARKERS cfg = _load_yaml_config() cfg_sec = _get_cfg_section(cfg, "interpretation_quality_check") if cfg_sec.get("enabled") is False: print("[PASS] interpretation quality check disabled by config.yaml") return 0 min_inline_r = args.min_inline_r if min_inline_r is None: min_inline_r = int(cfg_sec.get("min_inline_r", 3)) min_current_observation = args.min_current_observation if min_current_observation is None: min_current_observation = int(cfg_sec.get("min_current_observation", 3)) # Recommended structure gates (default: off; strict: on). require_top_hits = bool(cfg_sec.get("require_top_hits", False)) require_uncertainty_hits = bool(cfg_sec.get("require_uncertainty_hits", False)) require_actionability = bool(cfg_sec.get("require_actionability", False)) if args.strict: require_top_hits = True require_uncertainty_hits = True require_actionability = True # Traceable numbers gate (heuristic). check_untracked_numbers = bool(cfg_sec.get("check_untracked_numbers", True)) max_untracked_numbers = int(cfg_sec.get("max_untracked_numbers", 1)) if args.strict: max_untracked_numbers = 0 untracked_exempt_res = _compile_many(cfg_sec.get("untracked_number_exempt_patterns", [])) blacklist_patterns = cfg_sec.get("blacklist_patterns", None) if blacklist_patterns is None: blacklist_patterns = DEFAULT_BLACKLIST if not isinstance(blacklist_patterns, list): blacklist_patterns = DEFAULT_BLACKLIST check_title_style = args.check_title_style if check_title_style is None: check_title_style = bool(cfg_sec.get("check_title_style", False)) check_tone = args.check_tone if check_tone is None: check_tone = bool(cfg_sec.get("check_tone", False)) max_teaching_hits = args.max_teaching_hits if max_teaching_hits is None: max_teaching_hits = int(cfg_sec.get("max_teaching_hits", 0)) teaching_markers = cfg_sec.get("teaching_markers", None) if teaching_markers is None: teaching_markers = TEACHING_MARKERS if not isinstance(teaching_markers, list): teaching_markers = TEACHING_MARKERS check_actionability = args.check_actionability if check_actionability is None: check_actionability = bool(cfg_sec.get("check_actionability", False)) max_actionability_issues = args.max_actionability_issues if max_actionability_issues is None: max_actionability_issues = int(cfg_sec.get("max_actionability_issues", 0)) check_template_labels = args.check_template_labels if check_template_labels is None: check_template_labels = bool(cfg_sec.get("check_template_labels", False)) max_tier_label_lines = args.max_tier_label_lines if max_tier_label_lines is None: max_tier_label_lines = int(cfg_sec.get("max_tier_label_lines", 8)) check_bold = args.check_bold if check_bold is None: check_bold = bool(cfg_sec.get("check_bold", False)) max_bold_per_1000 = args.max_bold_per_1000 if max_bold_per_1000 is None: max_bold_per_1000 = float(cfg_sec.get("max_bold_per_1000", 10.0)) mechanical_template_patterns = cfg_sec.get("mechanical_template_patterns", None) if mechanical_template_patterns is None: mechanical_template_patterns = DEFAULT_MECHANICAL_TEMPLATE_PATTERNS if not isinstance(mechanical_template_patterns, list): mechanical_template_patterns = DEFAULT_MECHANICAL_TEMPLATE_PATTERNS check_mechanical_templates = bool(cfg_sec.get("check_mechanical_templates", True)) max_mechanical_template_hits = int(cfg_sec.get("max_mechanical_template_hits", 0)) vague_phrase_patterns = cfg_sec.get("vague_phrase_patterns", None) if vague_phrase_patterns is None: vague_phrase_patterns = DEFAULT_VAGUE_PHRASE_PATTERNS if not isinstance(vague_phrase_patterns, list): vague_phrase_patterns = DEFAULT_VAGUE_PHRASE_PATTERNS check_vague_phrases = bool(cfg_sec.get("check_vague_phrases", True)) vague_phrase_window_chars = int(cfg_sec.get("vague_phrase_window_chars", 160)) max_vague_phrase_issues = int(cfg_sec.get("max_vague_phrase_issues", 0)) DEFAULT_BLACKLIST = blacklist_patterns TEACHING_MARKERS = teaching_markers any_fail = False for fp in args.files: path = Path(fp) if not path.exists(): print(f"[FAIL] {fp}: file not found") any_fail = True continue text = path.read_text(encoding="utf-8", errors="replace") r = scan_text(text) inline_r_count = int(r["inline_r_count"]) current_obs_count = int(r["current_observation_count"]) top_hits = r["top_hits"] uncertainty_hits = r["uncertainty_hits"] blacklist_hits = r["blacklist_hits"] bold_span_count = int(r["bold_span_count"]) teaching_hits = r["teaching_hits"] tier_label_line_count = int(r.get("tier_label_line_count", 0)) actionability_issues = list(r.get("actionability_issues", [])) action_sentence_count = int(r.get("action_sentence_count", 0)) top_grounded = bool(r.get("top_grounded", False)) strong_claim_issues = list(r.get("strong_claim_issues", [])) uncertainty_value_issues = list(r.get("uncertainty_value_issues", [])) bad_headings = r["bad_headings"] prose_for_checks = str(r.get("_prose") or "") failures: list[str] = [] if inline_r_count < min_inline_r: failures.append(f"inline `r ...` too few: {inline_r_count} < {min_inline_r}") if min_current_observation > 0 and current_obs_count < min_current_observation: failures.append( "current-data observations too few: " f"{current_obs_count} < {min_current_observation} " "(need sentences with markers like '本次/当前数据/结果显示' + numeric evidence; " "each observation must describe THIS data, not generic rules)" ) if require_top_hits and (not top_hits): failures.append("missing Top signal hint (e.g. 'Top 3' / '前3' / '最强信号')") if require_uncertainty_hits and (not uncertainty_hits): failures.append("missing uncertainty/stability hint (e.g. CI/SE/bootstrap/CV)") if require_actionability and action_sentence_count <= 0: failures.append("missing actionable next-step sentence (need at least 1 sentence with method+input+criterion)") if blacklist_hits: failures.append(f"blacklist phrases found: {', '.join(blacklist_hits)}") if check_untracked_numbers: issues = _untracked_number_issues(text, exempt_res=untracked_exempt_res) if len(issues) > max_untracked_numbers: failures.append( f"untracked literal numbers: {len(issues)} lines (limit={max_untracked_numbers}); e.g. {issues[0]}" ) if check_mechanical_templates: hits = _mechanical_template_hits(prose=prose_for_checks, patterns=mechanical_template_patterns) if len(hits) > max_mechanical_template_hits: failures.append( f"mechanical template detected: {len(hits)} (limit={max_mechanical_template_hits})" ) if check_vague_phrases: issues = _vague_phrase_issues( prose=prose_for_checks, patterns=vague_phrase_patterns, window_chars=vague_phrase_window_chars, require_entity=True, ) if len(issues) > max_vague_phrase_issues: failures.append( f"vague phrases without local evidence: {len(issues)} issues (limit={max_vague_phrase_issues}); " f"e.g. {issues[0]}" ) if check_title_style and bad_headings: examples = "; ".join( f"L{h['line']}:{h['title']}({','.join(h['issues'])})" for h in bad_headings[:5] ) failures.append(f"non-natural headings found: {len(bad_headings)} (e.g. {examples})") if check_tone and teaching_hits and len(teaching_hits) > max_teaching_hits: failures.append( f"teaching-style markers found (limit={max_teaching_hits}): {', '.join(teaching_hits)}" ) if check_template_labels and tier_label_line_count > max_tier_label_lines: failures.append( f"tier-label lines too many: {tier_label_line_count} > {max_tier_label_lines} " "(suggest narrative style and hide '数据描述/统计见解/...' labels)" ) if check_actionability and len(actionability_issues) > max_actionability_issues: failures.append( f"non-actionable next steps: {len(actionability_issues)} issues (limit={max_actionability_issues}); " f"e.g. {'; '.join(actionability_issues[:3])}" ) if args.strict: if top_hits and not top_grounded: failures.append( "Top mention looks ungrounded: found 'Top 3/前3/最强信号' but no nearby entity-like token or inline `r ...` " "in interpretation sections" ) if strong_claim_issues: failures.append("strong claims need evidence+stability guard; e.g. " + "; ".join(strong_claim_issues[:3])) if uncertainty_value_issues: failures.append( "uncertainty/stability mentioned without current value evidence; e.g. " + "; ".join(uncertainty_value_issues[:3]) ) if check_bold: per_1000 = bold_span_count / max(1.0, len(strip_fenced_blocks(text)) / 1000.0) if per_1000 > max_bold_per_1000: failures.append( f"bold spans too dense: {bold_span_count} spans (~{per_1000:.1f}/1000 chars) > {max_bold_per_1000}" ) if failures: any_fail = True print(f"[FAIL] {fp}") for f in failures: print(f" - {f}") else: mode = r.get("mode", "scan") print(f"[PASS] {fp} ({mode}; inline_r={inline_r_count}, current_obs={current_obs_count})") if args.warn_only: return 0 return 2 if any_fail else 0 if __name__ == "__main__": raise SystemExit(main(sys.argv[1:])) -
check_pipeline_contract.py 3.1 KB
#!/usr/bin/env python3 """Read-only checks for the targets-first R analysis pipeline contract.""" from __future__ import annotations import argparse import json import re from pathlib import Path def finding(code: str, message: str, severity: str = "error") -> dict[str, str]: return {"code": code, "message": message, "severity": severity} def main() -> int: parser = argparse.ArgumentParser(description=__doc__) parser.add_argument("project_root", type=Path) parser.add_argument( "--workflow-mode", choices=("auto", "simple", "complex"), default="auto", ) args = parser.parse_args() root = args.project_root.expanduser().resolve() if not root.is_dir(): print(json.dumps({"status": "error", "message": "project root is not a directory"}, ensure_ascii=False)) return 2 targets_entry = root / "_targets.R" mode = args.workflow_mode if mode == "auto": mode = "complex" if targets_entry.is_file() else "simple" if mode == "simple": result = {"status": "pass", "workflow_mode": mode, "findings": [], "skipped": "no targets pipeline (simple mode)"} print(json.dumps(result, ensure_ascii=False, indent=2)) return 0 findings: list[dict[str, str]] = [] if not targets_entry.is_file(): findings.append(finding("missing-targets-entry", "complex/pipeline requires _targets.R")) r_dir = root / "R" if not r_dir.is_dir(): findings.append(finding("missing-r-directory", "targets-first pipeline requires project R/ for target functions")) targets_text = targets_entry.read_text(encoding="utf-8", errors="replace") if targets_entry.is_file() else "" if targets_text and not re.search(r"tar_source\s*\(\s*[\"']R[\"']", targets_text): findings.append(finding("r-source-not-declared", "_targets.R should discover R/ through tar_source(\"R\")")) rmd_files = sorted( path for path in root.rglob("*.Rmd") if "templates" not in path.parts and ".bensz-api" not in path.parts ) for rmd in rmd_files: text = rmd.read_text(encoding="utf-8", errors="replace") if not re.search(r"(?:tar_read|tar_load|tar_render)\s*\(", text): findings.append(finding("rmd-bypasses-targets", f"{rmd.relative_to(root)} must consume target results with tar_read/tar_load/tar_render")) if re.search(r"(?:read\.(?:csv|delim)|readRDS)\s*\(\s*[\"']raw[/\\]", text): findings.append(finding("rmd-reads-raw-directly", f"{rmd.relative_to(root)} reads raw/ directly instead of consuming targets")) errors = [item for item in findings if item["severity"] == "error"] result = { "status": "pass" if not errors else "fail", "workflow_mode": mode, "targets_entry": str(targets_entry.relative_to(root)) if targets_entry.exists() else None, "r_directory": str(r_dir.relative_to(root)) if r_dir.exists() else None, "reports_checked": [str(path.relative_to(root)) for path in rmd_files], "findings": findings, } print(json.dumps(result, ensure_ascii=False, indent=2)) return 0 if not errors else 1 if __name__ == "__main__": raise SystemExit(main()) -
check_plot_readability.R 10.9 KB · in bundle
-
check_rmd_template_yaml.py 3.4 KB
#!/usr/bin/env python3 """ Validate that templates/Rmd_template.Rmd YAML header matches config.yaml:rmd_template.yaml_header. Why: - config.yaml is the single source of truth for the YAML header structure. - The template keeps a convenience copy for "copy-and-start", which can drift over time. This script is for maintainers (CI/local checks), not for end-users. """ from __future__ import annotations import sys from pathlib import Path from typing import Any try: import yaml # type: ignore except Exception as e: # pragma: no cover print(f"[FAIL] PyYAML is required for this check. Error: {e}", file=sys.stderr) raise SystemExit(2) def _skill_root() -> Path: # scripts/check_rmd_template_yaml.py -> {skill_root}/scripts/... return Path(__file__).resolve().parents[1] def _load_yaml(path: Path) -> Any: return yaml.safe_load(path.read_text(encoding="utf-8")) def _extract_frontmatter(text: str) -> str | None: # Extract the first YAML frontmatter block between --- and ---. lines = text.splitlines() if not lines or lines[0].strip() != "---": return None for i in range(1, len(lines)): if lines[i].strip() == "---": return "\n".join(lines[1:i]) + "\n" return None def main() -> int: root = _skill_root() config_path = root / "config.yaml" tmpl_paths = [ root / "templates" / "Rmd_template.Rmd", root / "templates" / "Rmd_simple_template.Rmd", ] if not config_path.exists(): print(f"[FAIL] missing: {config_path}", file=sys.stderr) return 2 for tmpl_path in tmpl_paths: if not tmpl_path.exists(): print(f"[FAIL] missing: {tmpl_path}", file=sys.stderr) return 2 cfg = _load_yaml(config_path) or {} if not isinstance(cfg, dict): print("[FAIL] config.yaml is not a mapping.", file=sys.stderr) return 2 sec = cfg.get("rmd_template") if not isinstance(sec, dict): print("[FAIL] missing config.yaml:rmd_template section.", file=sys.stderr) return 2 cfg_header = sec.get("yaml_header") if not isinstance(cfg_header, dict): print("[FAIL] missing config.yaml:rmd_template.yaml_header mapping.", file=sys.stderr) return 2 for tmpl_path in tmpl_paths: fm = _extract_frontmatter(tmpl_path.read_text(encoding="utf-8", errors="replace")) if fm is None: print(f"[FAIL] {tmpl_path.name} does not start with YAML frontmatter.", file=sys.stderr) return 1 try: tmpl_header = yaml.safe_load(fm) or {} except Exception as e: print(f"[FAIL] failed to parse {tmpl_path.name} frontmatter as YAML: {e}", file=sys.stderr) return 1 if tmpl_header != cfg_header: dumped_cfg = yaml.safe_dump(cfg_header, sort_keys=True, allow_unicode=False) dumped_tmpl = yaml.safe_dump(tmpl_header, sort_keys=True, allow_unicode=False) print(f"[FAIL] YAML header mismatch: {tmpl_path.name}", file=sys.stderr) print("--- config.yaml:rmd_template.yaml_header", file=sys.stderr) print(dumped_cfg, file=sys.stderr) print(f"--- {tmpl_path.name} frontmatter", file=sys.stderr) print(dumped_tmpl, file=sys.stderr) return 1 print("[PASS] Rmd template YAML headers match config.yaml:rmd_template.yaml_header") return 0 if __name__ == "__main__": raise SystemExit(main()) -
check_targets_renv.py 9 KB
#!/usr/bin/env python3 """Inspect project state and validate simple/complex R workflow prerequisites.""" from __future__ import annotations import argparse import json import re from pathlib import Path DEFAULT_TEST_ENTRY = "scripts/tests/smoke_test.R" DEFAULT_TEST_STORE = "tmp/tests/<run-id>/_targets" def existing_signals(root: Path) -> list[str]: signals = [] for rel in ("_targets.R", "renv.lock", "renv/activate.R", "analysis-plan.yaml"): if (root / rel).exists(): signals.append(rel) for directory in ("raw", "products", "reports", "renv", "R"): if (root / directory).is_dir(): signals.append(directory + "/") if list(root.glob("*.R")) or list(root.glob("*.Rmd")): signals.append("root-R-files") return sorted(set(signals)) def observed_mechanisms(root: Path) -> dict[str, object]: return { "renv": (root / "renv.lock").is_file() or (root / "renv" / "activate.R").is_file(), "targets": (root / "_targets.R").is_file(), "product_roots": [name for name in ("products", "tmp", "reports") if (root / name).is_dir()], } def project_path(root: Path, value: str, label: str) -> tuple[Path | None, str | None]: candidate = Path(value) if candidate.is_absolute() or any(part in {"", ".", ".."} for part in candidate.parts): return None, f"{label} must be a normalized relative path inside the project" resolved = (root / candidate).resolve() try: resolved.relative_to(root) except ValueError: return None, f"{label} must stay inside the project" return resolved, None def validate_test_contract( root: Path, workflow_mode: str, test_entry: str, test_store: str, ) -> tuple[list[str], list[dict[str, str]]]: missing: list[str] = [] issues: list[dict[str, str]] = [] entry_path, entry_error = project_path(root, test_entry, "test entry") store_path, store_error = project_path(root, test_store, "test store") if entry_error: issues.append({"code": "unsafe-test-entry", "message": entry_error}) elif entry_path is not None and not entry_path.is_file(): missing.append(test_entry) if store_error: issues.append({"code": "unsafe-test-store", "message": store_error}) elif store_path is not None: forbidden = [root / name for name in ("_targets", "raw", "products", "reports")] if any(store_path == path.resolve() or path.resolve() in store_path.parents for path in forbidden): issues.append( { "code": "unsafe-test-store", "message": "test store must stay under an isolated tmp/ test directory", } ) expected_parent = (root / "tmp" / "tests").resolve() if store_path != expected_parent and expected_parent not in store_path.parents: issues.append( { "code": "test-store-not-isolated", "message": "test store must be tmp/tests or one of its descendants", } ) if entry_path is None or not entry_path.is_file(): return missing, issues text = entry_path.read_text(encoding="utf-8", errors="replace") has_tar_make = bool(re.search(r"\btar_make\s*\(", text)) delegates_run_root = "test_harness.R" in text if "BENSZ_TEST_RUN_ID" not in text and "run_id" not in text and not delegates_run_root: issues.append( { "code": "test-run-root-not-unique", "message": "smoke tests must create a unique tmp/tests/<run-id>/ run root", } ) if workflow_mode == "simple" and has_tar_make: issues.append( { "code": "simple-test-uses-targets", "message": "simple smoke tests must execute the R/Rmd entry directly, without targets", } ) if workflow_mode == "complex": if not has_tar_make: issues.append( { "code": "complex-test-missing-tar-make", "message": "complex smoke tests must execute the real target graph with tar_make()", } ) harness_bound = "test_harness.R" in text and "targets_store" in text literal_bound = "store" in text and all(part in text for part in ("tmp", "tests", "_targets")) if not (harness_bound or literal_bound): issues.append( { "code": "complex-test-store-unbound", "message": "complex smoke tests must bind tar_make() to an isolated tmp/tests/<run-id>/_targets store", } ) if re.search(r"store\s*=\s*['\"]_targets/?['\"]", text): issues.append( { "code": "complex-test-uses-formal-store", "message": "complex smoke tests cannot use the formal _targets/ store", } ) return missing, issues def main() -> int: parser = argparse.ArgumentParser(description=__doc__) parser.add_argument("project_root", type=Path) parser.add_argument( "--mode", choices=("auto", "new", "existing"), default=None, help="Deprecated compatibility alias for --project-state.", ) parser.add_argument("--project-state", choices=("auto", "new", "existing"), default=None) parser.add_argument( "--workflow-mode", choices=("auto", "simple", "complex"), default="auto", ) parser.add_argument("--test-entry", default=DEFAULT_TEST_ENTRY) parser.add_argument("--test-store", default=DEFAULT_TEST_STORE) args = parser.parse_args() root = args.project_root.expanduser().resolve() if not root.is_dir(): print(json.dumps({"status": "error", "message": "project root is not a directory"}, ensure_ascii=False)) return 2 project_state = args.project_state or args.mode or "auto" if args.project_state and args.mode and args.project_state != args.mode: print( json.dumps( {"status": "error", "message": "--mode and --project-state disagree"}, ensure_ascii=False, ) ) return 2 signals = existing_signals(root) if project_state == "auto": project_state = "existing" if signals else "new" workflow_mode = args.workflow_mode if workflow_mode == "auto": workflow_mode = "complex" if (root / "_targets.R").is_file() else "simple" missing: list[str] = [] issues: list[dict[str, str]] = [] warnings: list[str] = [] def report_absent(rel: str, display: str, action: str) -> None: if (root / rel).exists(): return if project_state == "new": missing.append(display) else: warnings.append(f"existing project has no {display}; reported, {action}") for rel, display in (("renv.lock", "renv.lock"), ("renv/activate.R", "renv/activate.R")): report_absent(rel, display, "not implicitly initialized") test_missing, test_issues = validate_test_contract( root, workflow_mode, args.test_entry, args.test_store, ) if test_missing: if project_state == "new": missing.extend(test_missing) else: warnings.append(f"existing project has no {args.test_entry}; not implicitly created") issues.extend(test_issues) if workflow_mode == "complex": report_absent("_targets.R", "_targets.R", "not implicitly migrated") report_absent("R", "R/", "not implicitly created") if (root / "_targets.R").is_file(): targets_text = (root / "_targets.R").read_text(encoding="utf-8", errors="replace") if not re.search(r"tar_source\s*\(\s*[\"']R[\"']", targets_text): issues.append( { "code": "targets-not-r-first", "message": "complex pipeline _targets.R must discover computation functions from R/ via tar_source(\"R\")", } ) if workflow_mode == "simple" and (root / "_targets.R").exists(): issues.append( { "code": "simple-has-targets-entry", "message": "simple mode must not create _targets.R; choose complex or remove the unintended entry", } ) missing = sorted(set(missing)) status = "pass" if not missing and not issues else "fail" result = { "status": status, "project_state": project_state, "workflow_mode": workflow_mode, "mode": project_state, "signals": signals, "missing": missing, "issues": issues, "warnings": warnings, "observed_mechanisms": observed_mechanisms(root) if project_state == "existing" else None, "test_entry": args.test_entry, "test_store": args.test_store if workflow_mode == "complex" else None, } print(json.dumps(result, ensure_ascii=False, indent=2)) return 0 if status == "pass" else 1 if __name__ == "__main__": raise SystemExit(main()) -
validate_paths.R 7.3 KB · in bundle
-
-
templates
-
renv
-
activate.R 384 B · in bundle
-
-
tests
-
complex_smoke_test.R 1.8 KB · in bundle
-
project_subset.R 1.4 KB · in bundle
-
simple_smoke_test.R 1.7 KB · in bundle
-
synthetic_fixture.R 837 B · in bundle
-
test_harness.R 7.8 KB · in bundle
-
-
00.Environment.R 7.4 KB · in bundle
-
analysis_plan_template.yaml 2.2 KB
workflow: "pipeline" project_state: "new" requirements: - id: "REQ-01" description: "说明用户目标或交付要求" targets: ["prepared_data", "analysis_results"] targets: - name: "prepared_data" purpose: "从 raw 输入形成完整、可复用的分析数据" depends_on: ["analysis_input"] function: "R/analysis_functions.R::prepare_data" inputs: ["raw/input.tsv"] scientific_products: [] completion: "target 成功且 schema/主键/缺失断言通过" - name: "analysis_results" purpose: "执行模型或统计计算,输出未按报告阈值截断的结果" depends_on: ["prepared_data"] function: "R/analysis_functions.R::analyze_data" inputs: ["prepared_data"] scientific_products: ["products/analysis_results.rds"] completion: "target 成功,关键数值与不变量断言通过" # 可选:仅在关键结果涉及样本推断时填写;simple 可用等价短表或方法段,不需此文件。 inference: - parameter: "主要结局的组间均值差(示例,需替换)" estimand: "目标人群中 A 组减 B 组的均值差,单位同结局" design: "独立两组;确认抽样单位、缺失和组内相关后选方法" method: "按实际设计选择有效区间与检验;此处不预设 t/Wald 方法" confidence_level: 0.95 # 仅示例默认值;预注册方案另有规定时以方案为准 null_hypothesis: "两组均值差为 0;若只做估计则填 not_applicable" test_sidedness: "预先声明单侧或双侧;不检验时填 not_applicable" multiplicity: "说明比较族、校正方法;无多重检验时填 not_applicable" not_estimable_when: "有效样本不足、设计/模型不支持有效推断时记录原因与替代分析" result_destination: "analysis_results 的完整结果及 Rmd 主要结果段" reports: - path: "reports/" consumer: "Rmd 使用 targets::tar_read('analysis_results')" presentation_parameters: ["top_n", "q_cutoff", "plot_language"] recovery: store: "_targets/" test_store: "tmp/tests/<run-id>/_targets/" acceptance: "中断后再次 tar_make() 跳过仍有效的前序 target" observability: progress: ["targets::tar_poll", "targets::tar_watch"] optional_workers: "crew" -
complexheatmap_template.R 2.8 KB · in bundle
-
datatables_helper.R 1.5 KB · in bundle
-
functions_template.R 558 B · in bundle
-
liquid_glass_lightbox.html 22.6 KB · in bundle
-
liquid_glass_theme.css 31.7 KB · in bundle
-
nature_colors.R 356 B · in bundle
-
nature_theme.R 2.6 KB · in bundle
-
plotly_template.R 1.7 KB · in bundle
-
plot_delivery_helpers.R 7.3 KB · in bundle
-
Rmd_simple_template.Rmd 2.9 KB · in bundle
-
Rmd_template.Rmd 12.2 KB · in bundle
-
R_data_template.R 928 B · in bundle
-
_targets.R 1.7 KB · in bundle
-
-
CHANGELOG.md 67.2 KB
# Changelog All notable changes to bensz-rmd-rules will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/zh-CN/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/lang/zh-CN/). ## [Unreleased] ### Added - 版本更新至 `0.29.0`:论文级关键估计对象在计算前逐项决定推断目的、方法、CI/检验可用性;新增 `statistical_inference_protocol.md`,规定完整计算结果、不可估计状态与报告对应。分析计划提供可选推断节点,simple/complex Rmd 模板、解读指南、审查和交付清单同步要求估计值、适当 CI、适用的 p/q 及来源。保留文本检查器的启发式定位,不将关键词命中当作推断完整性证明。 ### Fixed - 版本更新至 `0.28.2`:Liquid Glass 桌面动态 TOC 的宽度、最大高度和内边距立即切换,避免鼠标在展开动画期间移入目录正文时触发收回;保留阴影和内容透明度过渡,并新增 Chromium 交互回归测试。 - 版本更新至 `0.28.1`:Liquid Glass 在 `pre`/`code` 区域关闭标准与上下文连字,避免 Cascadia Code 等编程字体把 R 的 ASCII 赋值运算符 `<-` 显示成单个 `←` 字形;新增静态回归测试并补充旧项目主题同步说明。 ### Changed - 版本更新至 `0.28.0`:全面转向 targets-first,移除历史编号脚本执行体系的全部兼容——`preserved-existing` 模式删除,工作流模式收敛为 `simple`/`complex`;`existing` 降级为项目状态:有 `_targets.R` 按 complex 维护,无 targets(含历史编号脚本)按 simple 语义维护,需要复杂能力时显式迁移。新增 `_targets.R` 人类可读性契约:阶段注释分组、语义化 target 命名、交付摘要附 `tar_manifest()` 快照。 ### Removed - 删除 `scripts/run_analysis_workflow.py`(编号 runner)、`scripts/check_analysis_workflow.py`(编号单元检查器)与 `qa/test_analysis_workflow.py`;`config.yaml` 删除 `naming_convention.existing_unit_pattern` 与 `existing_project.workflow_mode: preserved-existing`。 - `check_targets_renv.py` 与 `check_pipeline_contract.py` 移除 `preserved-existing` 工作流模式与全部兼容分支:`--workflow-mode` 只接受 `auto|simple|complex`,auto 按 `_targets.R` 存在性解析;existing 项目的缺失机制降级为 warning,模式契约冲突(simple 携带 `_targets.R`、complex 缺 `tar_source("R")`)仍报 error。 - `templates/tests/test_harness.R` 解除 `preserved-existing` 模式耦合,`workflow_mode` 只接受 `simple`/`complex`;同步更新 `SKILL.md`、`README.md`、`workflow_modes.md`、`analysis_workflow_cache.md`、`hybrid_architecture_guide.md`、`hybrid_architecture_examples.md`、`lightweight_testing.md`、`workflow_checklist.md` 与 `evals/evals.json` 的 existing 语义。 ### Changed - 版本更新至 `0.27.0`:移除全部 legacy 兼容功能(自制 checkpoint 体系、旧 `tmp/{主脚本名}/` 产品路径模式与 legacy 信号检测/检查器分支),`preserved-existing` 只保留编号 runner 与"不自动迁移"边界。 ### Removed - 移除旧 checkpoint/自制缓存体系:删除 `templates/checkpoint_helpers.R` 与 `qa/test_checkpoint_helpers.R`;`check_analysis_workflow.py` 删除 checkpoint/SUCCESS/metadata 完整性校验、plan 单元 `cache` 字段契约与 `--allow-stale-checkpoints`;`run_analysis_workflow.py` 删除 `--force-step`/`--resume-from` 与 `BENSZ_FORCE_STEP`/`BENSZ_RESUME_FROM` 环境传递。 - 移除旧 `tmp/{主脚本名}/` 产品路径模式:删除 `config.yaml` 的 `legacy_temp_pattern`、`check_analysis_workflow.py` 的 `--legacy` 参数与 legacy tmp 项目自动识别、`cross_platform.md` 的"旧 `tmp/` 项目"专节;科学产物只进 `products/`、正式材料只进 `reports/`。 - 移除 legacy 信号检测与检查器分支:`check_targets_renv.py` 的 `observed_mechanisms` 不再报告 `legacy_runner` 信号、删除 `complex-uses-legacy-cache` token 检查;`check_pipeline_contract.py` 删除 SUCCESS/checkpoint/force-step/resume token 黑名单与 `legacy-checkpoint-in-new-pipeline` 文件检查;`config.yaml` 删除 `legacy_runner_fallback`、`preserve_existing_runner`、`legacy_runner` 块与 `legacy_checkpoint_allowed`。 - 同步清理 `SKILL.md`、`hybrid_architecture_guide.md`、`hybrid_architecture_examples.md`、`analysis_workflow_cache.md`、`evals/evals.json` 中"历史项目可继续维护原 helper/tmp 路径"的兼容描述。 ### Fixed - 修复 `qa/test_analysis_workflow.py` 模板断言过时:`Rmd_template.Rmd` 已切换为 `targets::tar_read()` 消费,断言仍检查旧 `file.path(products_dir, ...)` 写法导致失败;同步为断言 `targets::tar_read(`。 ### Added - 新增 `scripts/check_pipeline_contract.py`,只读检查 targets-first pipeline 的 `_targets.R`、`R/`、Rmd target 消费、legacy checkpoint 禁用与 raw 绕过风险。 - 新增 targets 恢复验收、`tar_poll()`/`tar_watch()`、crew/autometric 可选观测和嵌套并行防护契约;模板新增 `tar_source("R")`、target 函数与可重跑 smoke 证据。 - 新增 `workflow_modes.md` 与 `lightweight_testing.md`,把项目状态、simple/complex 模式、人工覆盖、两种测试风格、隔离与失败分类固化为独立契约。 - 新增 `templates/tests/` 的 synthetic fixture、project subset、simple smoke 与 complex smoke 骨架;`evals/evals.json` 扩展到模式选择、人工覆盖、已有项目兼容、测试风格、自定义产品路径、目录职责和失败闭环。 - 新增 Skill `qa/` 双模式 R 集成回归,真实执行 simple Rmd render 和 complex target 图,验证唯一 run root、测试 store、真实子集清理、subject identity、raw 与正式输出不被污染。 ### Changed - 版本更新至 `0.26.0`:complex 统一解释为 targets-first pipeline,计算函数转入 `R/`,Rmd 通过 `tar_read()`/`tar_load()` 消费结果;新项目不再生成第二套 SUCCESS/identity/checkpoint runner,existing 旧机制保留兼容。 - `products/` 与 `_targets/` 职责重新分离:前者只保存需审阅、复用或交付的科学产物,后者由 targets 管理机器状态、失效与恢复。 - `SKILL.md` 以主要交付物与验收标准明确和 `bensz-r-developer` 的兄弟边界:分析流程、结果、报告及分析本地 helper 由本 Skill 主导;可独立复用的 API/类/Package 交给 `bensz-r-developer`。 - 补充混合任务的“分析定义组件契约 → 组件工程实现 → 分析集成验证”协作顺序、README 分流说明和近邻触发用例,版本更新至 `0.24.1`。 - 新项目从“全部 targets”改为 simple/complex 双模式:两者都强制 renv 与真实轻量测试,只有 complex 使用 targets;已有项目仍不自动迁移或补齐。 - `check_targets_renv.py` 分离 `--project-state` 与 `--workflow-mode`,保留旧 `--mode` 兼容别名;existing 固定为 `preserved-existing` 并报告观察机制,同时校验 simple 无 targets、complex 唯一 run root 下的隔离 test store 和测试入口。 - 新项目目录职责调整为样式在 `templates/`、共享 helper 在 `scripts/lib/`、测试代码在 `scripts/tests/`、运行现场在 `tmp/tests/`;版本更新至 `0.25.0`。 - products 根目录支持通过单一 `BENSZ_PRODUCTS_DIR` 项目设置安全覆盖;checkpoint helper、工作流 checker/runner 和 Rmd 模板同步使用统一路径并拒绝越界与保留目录。 ### Fixed - 不再把 checker、`--dry-run` 或 target 图解析误作端到端测试通过;交付证据要求实际 R/Rmd/target 执行、断言和重跑记录。 - 不再把 `products/` 描述为普通技术缓存,也不再把 checkpoint helper 生成到用户项目的样式目录。 ## [0.24.0] - 2026-09-19 ### Added - 新项目生命周期契约:统一使用 `targets` 编排(含单步骤项目)与 `renv` 锁定环境。 - `templates/_targets.R`、`templates/renv/activate.R` 和 `scripts/check_targets_renv.py`,支持最小初始化与只读状态检查。 ### Changed - `SKILL.md`、`README.md`、架构指南与检查清单同步“新项目强制、已有项目独立兼容、不自动迁移、不建立旧 runner 回退入口”。 - `config.yaml` 的固定 R 依赖收缩为 `luckyBase`;具体分析包改由项目 `renv.lock` 管理。 ## [0.23.0] - 2026-09-17 ### Added - `templates/analysis_plan_template.yaml`:新增需求、依赖、产品、缓存与报告关系的机器可读分析图模板。 - `templates/checkpoint_helpers.R`:新增项目内相对路径、输入/参数/代码/上游身份、完整性复验、最后写入 `SUCCESS`、强制重算与恢复控制。 - `scripts/check_analysis_workflow.py`、`scripts/run_analysis_workflow.py` 与 `qa/test_analysis_workflow.py`:新增编号、依赖、目录边界、checkpoint 元数据、输出摘要、旧 `tmp/` 兼容和 `raw/` 只读检查,并提供按编号运行、强制重算与恢复入口。 - `references/analysis_workflow_cache.md` 与 `references/serial_review_protocol.md`:新增缓存选择/失效/恢复及三轮串行只读审查契约。 - `references/metric_explanation_protocol.md`:新增面向弱背景读者的指标解释协议,定义常用指标白名单、不常用指标兜底判定、指标导读表及首次出现完整解释规则。 - `plans/TOC移动支持-v202603011129.md`:新增 Liquid Glass 的 TOC 移动端支持优化计划(sticky 折叠目录条、跳转后自动收起、锚点偏移与触控可用性等)。 - `tests/TOC移动支持-v202603011129/`:新增对应轻量测试会话(PLAN/REPORT + demo.Rmd/render.sh + 静态断言脚本 + 3 组 viewport 截图证据,含“存储不可用”模拟)。 - `templates/plot_delivery_helpers.R`:新增“PDF 交付 + JPG 预览”最小 helper(`bensz_run_dir()` / `bensz_pdf_to_jpg()`),优先使用 poppler/ImageMagick,R 依赖可选。 - `tests/图片可读性-v202603011109/`:新增轻量测试会话(PLAN/REPORT + 3 个自生成 fixture),覆盖“PDF 产出 + 预览生成 + checker 渲染”闭环。 ### Changed - `SKILL.md`:重组为“理解—规划—实现—审查—运行/恢复—交付”主链;新项目默认使用根目录编号分析单元和 `raw/products/reports/.bensz-api` 边界,简单任务保留最小路径。 - `config.yaml`:以三段式编号、多分析单元、数据产品、缓存身份、报告目录和串行审查替代单主脚本默认;删除未落地的 `analysis_mode`,版本更新至 `0.23.0`。 - `templates/00.Environment.R`、`R_data_template.R`、`functions_template.R`、`Rmd_template.Rmd`:改为单编号分析单元模板,Rmd 从完整 products 读取并将正式材料写入 reports。 - `README.md`、混合架构指南、示例、检查清单和交付验证:同步新项目默认与旧 `tmp/` 非破坏兼容,移除未实现的 `dv_mut_mode` 和 README 版本字面值。 - `SKILL.md` / `README.md` / `templates/Rmd_template.Rmd` / `references/four_tier_interpretation_framework.md` / `references/workflow_checklist.md` / `references/delivery_verification.md`:专家级解读默认假定读者背景较弱;存在不常用指标时,要求先给“指标导读”表,首次出现解释定义、原理、选用理由、判读方向与不确定性,后续出现改为简洁结果解读。 - `scripts/check_interpretation_quality.py` / `config.yaml`:教学口吻检测移除“用于评估/反映关系/当……时……”等中性指标解释表达,只保留模板化提示语;新增 `qa/test_metric_explanation_protocol.py` 防止必要的首次解释被误拦截。 - `config.yaml`:新增 `metric_explanation` 默认分类与解释策略,并将版本号 `0.21.3 → 0.22.0`。 - `config.yaml`:版本号 `0.21.1 → 0.21.2`;JPG 预览 run 目录默认从 `/tmp/bensz-rmd-rules/` 迁移到当前项目下 `.bensz-api/skills/bensz-rmd-rules/`。 - `SKILL.md` / `templates/Rmd_template.Rmd` / `templates/plot_delivery_helpers.R` / `scripts/check_plot_readability.R`:同步更新 PDF→JPG 预览中间目录口径,避免默认写入系统 `/tmp`。 - `SKILL.md`:新增“参数面向用户 + PDF 交付 + JPG 视觉自检”硬约束口径,并补齐推荐落地方式与 YAML params(`plot_run_dir/plot_preview_dpi`)。 - `SKILL.md`:补充“HTML 表格默认使用 DT(DT::datatable)渲染”的规范口径(用户另有指定时以用户要求为准)。 - `SKILL.md` / `templates/Rmd_template.Rmd`:补充“HTML 展示图默认使用由 PDF 渲染得到的 JPG 预览图,以保证 HTML 与 PDF 比例一致;展示可用 out.width 缩小但避免 out.height 失真”的规范口径。 - `SKILL.md` / `templates/*.R*` / `plans/Nature级别图-v202601270654.md` / `references/plot_language.md`:图表默认不在图内使用 `title/subtitle`(例如 `ggplot2::labs(title=..., subtitle=...)`、`ggtitle()`、`plotly::layout(title=...)`),改用文件名 + `.Rmd` 标题/图注承载语义;仅在用户明确要求或确有必要时才添加。 - `templates/Rmd_template.Rmd`:在模板中加入“集中参数块 + PDF 交付 + JPG 预览”的示例口径,并在 `plot-style` chunk 约定 `run_dir` 以支持预览生成。 - `scripts/check_plot_readability.R`:新增 `--render-jpg/--out-dir/--dpi/--page`(轻依赖渲染预览 + proxy 指标),保留原有仅检查 PDF 的用法不变。 - `templates/liquid_glass_theme.css`:TOC 窄屏(≤1024px)升级为“顶部 sticky 折叠目录条”,并补齐触控可用性(按钮/链接 ≥44px hitbox)与移动端锚点跳转偏移口径,降低跳转遮挡与滚动后不可达的问题。 - `templates/liquid_glass_lightbox.html`:窄屏下点击任意 TOC 条目后自动收起(并持久化折叠态),避免跳转后遮挡;同时补充 `resize/orientationchange` 下的按钮文案同步兜底。 - `references/liquid_glass_theme_guide.md` / `README.md`:同步更新“移动端目录(窄屏)”口径(sticky + 点击后自动收起 + 跳转偏移)。 - `config.yaml`:新增 `params.plot_run_dir/plot_preview_dpi`(与模板同步),并更新版本号 `0.20.0 → 0.21.1`(Single Source of Truth)。 - `SKILL.md`:对齐“不过度防御”边界口径(允许 I/O 边界与硬前提做显式检查+stop,避免冗余包/函数检查),并澄清 Liquid Glass YAML 字段的单一真相来源。 - `config.yaml`:澄清 `references/liquid_glass_theme_guide.md` 为说明文档,具体字段以 `rmd_template.yaml_header` 为准。 - `templates/Rmd_template.Rmd`:补齐 `templates/datatables_helper.R` 缺失时的 fail-fast 提示(指向 `bootstrap_liquid_glass.py --with-extras`)。 ### Fixed - `templates/liquid_glass_theme.css`:修复桌面端动态 TOC 悬停展开时偶发抖动的问题。移除圆角半径过渡,避免展开动画改变鼠标命中边界并触发 `mouseenter`/`mouseleave` 振荡;版本号 `0.22.0 → 0.22.1`。 - `scripts/*.py`、`scripts/*.R` 与可执行 R 模板统一使用 `[PASS]` / `[FAIL]` / `[WARN]` ASCII 状态前缀,修复 Windows GBK 控制台因 emoji/符号不可编码而在成功或失败分支崩溃的问题;版本号 `0.21.2 → 0.21.3`。 - `figure_interpretation_check.check_patterns.table_generation` 补齐推荐 helper `render_dt_output()` / `render_dt()`,使图表解读覆盖检查与 htmlwidget 可见性检查采用一致的 DT helper 边界。 - `templates/liquid_glass_theme.css`:修复 Liquid Glass 主题对 Plotly 使用过宽的 `.plotly` 选择器会误伤 Plotly.js 内部节点,导致“双层背景/额外阴影与内边距/图像像浮在背景上”的观感问题;改为仅作用于 `.plotly.html-widget` 外层容器并采用更保守的默认样式。 - `scripts/check_figure_table_interpretation.py`:修复将 `results='hide'` 误判为“隐藏输出”导致的漏检;新增 `fig.show='hide'` 与 `fig.keep='none'` 的隐藏判定。 - `scripts/validate_paths.R`:修复仅扫描 `.Rmd` 与“行首绝对路径”的检测盲区;降低“路径拼接”误报;重复运行时重置结果避免累积旧告警。 ## [0.20.0] - 2026-02-13 ### Added - `plans/结果解读-流于表面-v202602130728.md`:新增优化计划,将“结果解读”硬约束从“写作枷锁”调整为“交付前质检门禁”,并明确保留/放松项的取舍口径。 - `tests/结果解读-流于表面-v202602130728/`:新增轻量测试会话(PLAN/REPORT + 4 个最小 Rmd fixture),覆盖探索期通过、覆盖缺失失败、不可追溯数字失败、严格模式叙述式通过。 ### Changed - `SKILL.md`:将“结果后解读”从“写法强制”调整为“两阶段门禁”口径;明确硬门禁(证据锚定/数字可追溯/覆盖不漏项/禁止代码生成解释)与推荐结构(四层/Top/不确定性/可执行后续);补充探索期/交付期自检命令。 - `config.yaml`: - `figure_interpretation_check`:默认 `strict_mode=false`(仅报告),交付期用 `--strict` 显式 Fail Fast;marker 默认改为更中性词(不再隐性要求四层标签);并将 `min_cjk_chars/min_en_words/min_content_elements/require_markers` 作为默认参数来源。 - `interpretation_quality_check`:放松零容忍阈值(teaching/vague/actionability);默认不强制 Top/不确定性/可执行后续;新增“不可追溯字面数字”启发式门禁与白名单。 - 版本号 `0.19.0 → 0.20.0`(Single Source of Truth)。 - `scripts/check_interpretation_quality.py`:默认不再强制 Top/不确定性/行动建议;`--strict` 下强制;新增“不可追溯字面数字”检测并输出行号示例,降低解释性表达误杀与模板化填空。 - `scripts/check_figure_table_interpretation.py`:marker/长度/要素门槛从 `config.yaml` 读取为默认值(CLI 可覆盖);支持 `require_markers` 配置;默认 marker 口径改为更中性词。 ### Fixed - **Liquid Glass TOC 点击定位精确度**:修复目录(TOC)链接点击后标题定位贴边或不够精确的问题 - `templates/liquid_glass_theme.css`:为所有标题(h1-h6)添加 `scroll-margin-top: 2rem`,确保锚点跳转时预留顶部空间 - 测试验证:`tests/liquid_glass_demo.html` 已包含修复后的样式 - **Liquid Glass 代码块复制工具条样式**:将“复制/Hide”工具条改为代码块右上角浮动胶囊样式,减少空白占用并统一按钮视觉 - `templates/liquid_glass_theme.css`:工具条改为 `position: absolute` 浮动展示;为工具条按钮提供一致的玻璃拟态样式;为代码块增加顶部内边距以避免遮挡首行 - `templates/liquid_glass_lightbox.html`:复制按钮不再依赖 Bootstrap 的 `.btn*` class,避免被主题/外部 CSS 干扰 - 测试验证:`tests/代码复制样式优化-v202602102004/`(PLAN/断言脚本/REPORT) ## [0.19.0] - 2026-02-10 ### Added - Liquid Glass:代码块右上角新增“复制/Copy”按钮(支持 `code_folding` 场景,将按钮放在 Hide 左侧;并兼容 Clipboard API 与 `execCommand('copy')` 降级)。 - `tests/代码复制功能-新增-v202602082340/`:轻量测试会话(PLAN/脚本断言/产物证据)。 ### Changed - `templates/liquid_glass_theme.css`:新增代码块工具条与复制按钮样式(玻璃拟态 + 交互反馈)。 - `templates/liquid_glass_lightbox.html`:新增代码块复制按钮注入逻辑(after_body)。 - `references/liquid_glass_theme_guide.md`:补充复制按钮使用说明与兼容性说明。 - `config.yaml`:版本号 `0.18.0 → 0.19.0`(Single Source of Truth)。 ## [0.18.0] - 2026-02-09 ### Added - `tests/结果解读-流于表面-v202602091747/`:轻量测试会话,覆盖“机械四段式模板”与“空泛句式缺少本地证据”两类反例,并提供正例 Rmd。 - `config.yaml:interpretation_quality_check`:新增机械模板检测与空泛句式缺少证据检测(可配置开关、阈值与正则)。 ### Changed - `SKILL.md`:将“结果后解读”从“结构门槛”进一步收敛为“思考路径 + 可证伪推理 + 可执行后续”,并显式禁止机械模板化四段式。 - `references/interpretation_narrative_examples.md`:从“替换变量名的示例库”重写为“专家解读思维示例(非模板)”,强调推理链条与风险边界。 - `references/interpretation_templates.md`:强化“写作前列要点、交付时合并为叙述”的使用方式,增加解读前四问与反机械模板说明。 - `scripts/check_interpretation_quality.py`:新增机械四段式模板识别与空泛句式缺少本地证据检测;相关阈值默认从 `config.yaml` 读取。 - `config.yaml`:将结果解读检查从“硬黑名单”迁移到“反模板 + 证据邻域”口径,并更新版本号 `0.17.9 → 0.18.0`(Single Source of Truth)。 ### Fixed - `scripts/check_interpretation_quality.py`:修复 `config.yaml` 中空列表(如 `blacklist_patterns: []`)会被误回退到内置默认值的问题。 ## [0.17.9] - 2026-02-09 ### Added - `tests/结果解读-流于表面-v202602091308/`:轻量测试会话与证据(PLAN/REPORT + 正反例 Rmd)。 - `references/interpretation_templates.md`:新增“核心结论(证据链收敛)”模板(面向 cohort/亚组收敛叙事)。 - `references/interpretation_narrative_examples.md`:新增“核心结论(证据链收敛)”连贯叙述式示例段。 ### Changed - `scripts/check_interpretation_quality.py`:强化“教学口吻 / 不可落地后续 / 模板化分层标签”静态阻断,并支持 `--strict` 做更强反模板化检查;默认从 `config.yaml:interpretation_quality_check` 读阈值与开关。 - `SKILL.md`:明确“连贯叙述式”为默认交付形态,并补充 `check_interpretation_quality.py --strict` 用法。 - `README.md`:补充“避免结果解读流于表面”的推荐默认策略与自检命令。 - `config.yaml`:补齐 `rmd_template.yaml_header.params`(与模板一致)并扩展 `interpretation_quality_check` 配置;版本号 `0.17.8 → 0.17.9`。 - `templates/Rmd_template.Rmd`:补齐解读相关 `params` 与“结果解读”占位段,减少“有结果无收敛解读”的漏项风险。 ## [0.17.8] - 2026-02-09 ### Changed - **B 轮 P1/P2 遗留问题修复** - `SKILL.md`:从 508 行精简到 486 行(≤500 行),将"禁止通用套话"详细黑白名单表格下沉到 references,仅保留核心口径和引用链接 - `config.yaml`:新增 `interpretation_quality_check` 配置节,包含 blacklist_patterns(单一真相来源) - `scripts/check_figure_table_interpretation.py`:更新注释说明脚本默认值为降级方案,config.yaml 为真相来源 - `config.yaml`:版本号 `0.17.7 → 0.17.8` - references 目录审计:所有 20 个文件均被引用,建议后续合并 `hybrid_architecture_guide.md` + `hybrid_architecture_examples.md` ## [0.17.7] - 2026-02-09 ### Changed - **A 轮一致性修复(v202602090010)** - `references/four_tier_interpretation_framework.md:58`:将"每句必须包含至少一个 `` `r ...` ``"改为"每句必须包含数值证据(优先 `` `r ...` `` 动态嵌入,或明确可追溯的字面数字)",对齐 check_interpretation_quality.py 的实际检测行为 - `references/no_overdefensive_code.md:56`:增加 00.Environment.R 存在性检查的白名单例外说明,解决与 Rmd_template.Rmd 的矛盾 - `references/workflow_checklist.md:40`:新增 YAML 头一致性检查项,引用 `scripts/check_rmd_template_yaml.py` - `config.yaml`:版本号 `0.17.6 → 0.17.7` - **B 轮质量检查(v202602090010)** - 完成 8 维度质量评估,得分 101/115(88%) - 确认 A 轮修复生效,无新增 P0 问题 - 遗留 P1:SKILL.md 508 行超限、配置集中化;遗留 P2:references 文件数可精简 ### Added - **强化结果解读当前数据观察门槛** - `bensz-rmd-rules/SKILL.md`:在"结果后解读"章节新增"禁止通用套话"硬门槛(黑名单示例 + 强制要求 + 白名单示例) - `bensz-rmd-rules/references/four_tier_interpretation_framework.md`:将示例从"骨架版"改为"血肉版",每句话都绑定 `` `r ...` `` 动态数值;新增"半成品解读"反例说明;将"当前数据观察"门槛从 2 句提升到 3 句 - `bensz-rmd-rules/references/interpretation_templates.md`:增强"反模式"章节,新增典型模板化句式黑名单、专家级写作公式、示例对比表格 - `bensz-rmd-rules/scripts/check_interpretation_quality.py`:新增 8 个教学口吻检测标记(如"用于判断"、"展示...的...形态"、"当...时...更支持"等);将 `--min-current-observation` 默认值从 2 提高到 3 - `bensz-rmd-rules/tests/结果解读-流于表面-v202602082340/`:问题修复的轻量测试会话(PLAN + REPORT) - **TOC 动画问题记录** - `bensz-rmd-rules/plans/TOC动画优化-v202602052237.md`:记录动态浮动 TOC 的动画不自然问题与位置 - **Liquid Glass 目录(TOC)两种浮动模式:静态/动态** - `bensz-rmd-rules/templates/liquid_glass_theme.css`:新增动态浮动目录的收起/展开样式;静态模式下正文为目录预留空间 - `bensz-rmd-rules/templates/liquid_glass_lightbox.html`:新增目录模式切换逻辑与目录头部(“静态/动态”按钮),并在桌面宽屏下将 TOC 从 bootstrap 栅格中移出以释放正文空间 - `bensz-rmd-rules/references/liquid_glass_theme_guide.md` / `bensz-rmd-rules/README.md`:补充用户侧切换方式与行为说明 - **HTML 视图状态保持(刷新后缩放 + 阅读位置尽量维持)** - `bensz-rmd-rules/templates/liquid_glass_lightbox.html`:新增内部缩放快捷键(`Ctrl/Cmd + (+/-/0)`)与刷新后 scroll 恢复逻辑(基于 `sessionStorage`) - `bensz-rmd-rules/references/liquid_glass_theme_guide.md` / `bensz-rmd-rules/README.md` / `bensz-rmd-rules/SKILL.md`:补充使用说明与故障排除口径 - **图表/表格解读覆盖检验(Fail Fast)** - `bensz-rmd-rules/scripts/check_figure_table_interpretation.py`:静态检测“有输出但无解读”的漏项(支持 `--strict` 阻断交付) - `bensz-rmd-rules/references/figure_interpretation_criteria.md`:覆盖检验口径与参数说明 - **YAML header 一致性校验(维护者用)** - `bensz-rmd-rules/scripts/check_rmd_template_yaml.py`:校验 `templates/Rmd_template.Rmd` 的 YAML header 是否与 `config.yaml:rmd_template.yaml_header` 一致 - **SKILL.md 渐进披露下沉文档** - `bensz-rmd-rules/references/htmlwidget_visibility_rules.md`:htmlwidget/DT/plotly 的 HTML 可见性硬规则与示例(从 SKILL.md 下沉) - `bensz-rmd-rules/references/numeric_accuracy_verification.md`:数字准确性验证章节模板与规范(从 SKILL.md 下沉) - **轻量测试会话** - `bensz-rmd-rules/tests/v202602041511/`:脚本级静态检查(PLAN/REPORT + 样例 Rmd + 产物) - `bensz-rmd-rules/tests/放大或当位位置维持-v202602042153/`:视图状态保持(缩放/scroll)相关静态校验与 bootstrap 复现(PLAN/REPORT + 断言输出) ### Changed - **Liquid Glass 图片/代码块居中观感修复** - `bensz-rmd-rules/templates/liquid_glass_theme.css`:补充 `div.figure` 包裹场景的居中样式,修复图片偏左 - `bensz-rmd-rules/config.yaml`:版本号 `0.17.3 → 0.17.4` - `bensz-rmd-rules/tests/v202602051015/`:figure 包裹图片居中的轻量测试计划与报告 - **Liquid Glass 图片默认居中** - `bensz-rmd-rules/templates/liquid_glass_theme.css`:独立图片自动居中,避免移动端/桌面端左对齐观感 - `bensz-rmd-rules/references/liquid_glass_theme_guide.md`:补充图片居中说明 - `bensz-rmd-rules/config.yaml`:版本号 `0.17.2 → 0.17.3` - `bensz-rmd-rules/tests/v202602050940/`:图片居中的轻量测试计划与报告 - **Liquid Glass 目录移动端可折叠** - `bensz-rmd-rules/templates/liquid_glass_theme.css`:移动端目录折叠卡片样式、触控友好间距与安全区适配 - `bensz-rmd-rules/templates/liquid_glass_lightbox.html`:窄屏下按钮改为“展开/收起”,并记忆折叠状态 - `bensz-rmd-rules/references/liquid_glass_theme_guide.md` / `bensz-rmd-rules/README.md` / `bensz-rmd-rules/SKILL.md`:补充移动端行为说明 - `bensz-rmd-rules/config.yaml`:版本号 `0.17.1 → 0.17.2` - `bensz-rmd-rules/tests/v202602050900/`:移动端目录折叠的轻量测试计划与报告 - **Liquid Glass 目录(TOC)字体层级更协调** - `bensz-rmd-rules/templates/liquid_glass_theme.css`:覆盖 tocify 默认的 subheader 12px,缩小 # / ## 的字号落差,使目录更易读 - **Liquid Glass 动态 TOC 动画节奏对齐** - `bensz-rmd-rules/templates/liquid_glass_theme.css`:统一容器过渡节奏并为内容显隐增加轻微延迟 - `bensz-rmd-rules/config.yaml`:版本号 `0.17.4 → 0.17.5` - `bensz-rmd-rules/tests/v202602052300/`:轻量测试计划与报告 - **结果解读避免“教学口吻/通用套话”**:强化“当前数据观察”门槛,减少“流于表面”的解读 - `bensz-rmd-rules/SKILL.md`:在“四层硬门槛”处新增“必须陈述本次结果具体观察”的硬门槛 - `bensz-rmd-rules/references/interpretation_templates.md`:新增“反模式:模板化写作 vs 专家级写作” - `bensz-rmd-rules/references/four_tier_interpretation_framework.md`:Fail Fast Gate 增加“当前数据观察”项,并收敛示例的通用推断写法 - `bensz-rmd-rules/references/interpretation_narrative_examples.md`:示例收敛教学式“用于判断”句式 - `bensz-rmd-rules/scripts/check_interpretation_quality.py`:新增 `--min-current-observation`(默认 2)并在 PASS 输出 `current_obs` - `bensz-rmd-rules/tests/v202602082340/`:新增轻量测试(bad/good Rmd + 输出证据) - `bensz-rmd-rules/config.yaml`:版本号 `0.17.5 → 0.17.6` - **交付前强制检查链路**:新增“图表/表格解读覆盖检验”作为交付前必跑步骤 - `bensz-rmd-rules/SKILL.md`:新增强制命令与失败处理口径 - `bensz-rmd-rules/README.md`:同步用户侧运行方式 - `bensz-rmd-rules/config.yaml`:新增 `figure_interpretation_check` 配置段 - `bensz-rmd-rules/references/workflow_checklist.md` / `bensz-rmd-rules/references/delivery_verification.md`:同步检查项 - **配置/模板口径对齐** - `bensz-rmd-rules/config.yaml`:明确 YAML header 的 SSOT 口径,并移除易误导的 `plot_quality.enforce_nature_level`;版本号 `0.15.1 → 0.16.0` - `bensz-rmd-rules/templates/Rmd_template.Rmd`:去除时间/品牌绑定叙事(iOS 26),并明确“模板 YAML 为便捷拷贝,需与 config 同步” - `bensz-rmd-rules/templates/liquid_glass_theme.css` / `bensz-rmd-rules/references/liquid_glass_theme_guide.md`:将主题描述改为时间无关的 glassmorphism 叙事 - **最小惊讶原则(模板副作用收敛)** - `bensz-rmd-rules/templates/00.Environment.R`:默认不自动启用 `showtext::showtext_auto(TRUE)`;改为显式开关/环境变量启用 - **配置集中化更安全** - `bensz-rmd-rules/scripts/check_figure_table_interpretation.py`:缺少 PyYAML 或 config 解析失败时输出清晰警告,避免“以为配置生效但实际静默回退” - **SKILL.md 瘦身** - `bensz-rmd-rules/SKILL.md`:将长示例/模板块下沉到 references,并将主体收敛到 500 行以内 ### Changed - **Rmd 报告可追溯性**:在“数据概览”后新增 `## 关键函数、参数与源代码位置` 章节规范与模板 - `bensz-rmd-rules/SKILL.md`:新增“章节结构(硬门槛)”与该章节的填写要求(关键过程/关键函数/关键参数与理由/源代码文件+行号) - `bensz-rmd-rules/templates/Rmd_template.Rmd`:新增对应章节模板与示例表格 - `bensz-rmd-rules/README.md`:同步用户侧说明 - **templates/liquid_glass_theme.css**:增大解释类文字的字体大小 - `p`(段落):字体从继承 16px 调整为 1.05rem(约 16.8px) - `li`(列表项):字体从继承 16px 调整为 1.05rem(约 16.8px) - `figcaption`(图片说明):字体从 0.9rem 调整为 0.95rem - 所有解释类文字行高从 1.6 调整为 1.7,提升可读性 ### Added - **references/code_style_guide.md**:新增 R 代码风格指南(从 SKILL.md 下沉) - 核心原则:管道优先、向量化思维、函数式编程、数据驱动 - 减少 if 语句使用:场景对照表(条件赋值/分类映射/数值替换/存在性检查/多分支逻辑) - 禁止防御性文件存在性检查:让代码在文件缺失时自然报错 - 代码注释规范:代码块头部、分步注释、业务逻辑注释 - 常见反模式:遍历+if、嵌套 if-else、防御性文件检查 ### Changed - **SKILL.md**:"人类可读原则"章节重构 - 移除详细的代码风格要求内容(基础代码风格、减少 if 语句使用、命名规范、代码注释规范) - 改为引用 `references/code_style_guide.md`,保留核心要点摘要 - 遵循渐进披露原则,降低 SKILL.md 维护负担 - **templates/liquid_glass_theme.css**:新增 Liquid Glass HTML 主题(iOS 26 风格) - Glassmorphism 效果(背景模糊与半透明层次) - 流体渐变动画与有机色彩过渡 - 弹性动效系统(类似 iOS 交互反馈) - 多层深度阴影营造空间感 - 自动深色模式支持(根据系统偏好切换) - 保留完整的浮动目录功能 - **references/liquid_glass_theme_guide.md**:新增 Liquid Glass 主题使用指南 - 快速开始教程 - 设计特性详解 - 组件样式展示 - 自定义工具类说明 - 故障排除与进阶定制 ### Changed - **config.yaml**:更新 `rmd_template.yaml_header` 配置 - `theme: default`(满足 `toc_float` 的 rmarkdown 约束;主要视觉由 Liquid Glass CSS 覆盖) - 仅使用 `css` 引用 Liquid Glass 样式(避免把原始 CSS 作为“头部文本”插入导致页面顶部出现代码墙) - **templates/Rmd_template.Rmd**:更新模板 YAML 头部 - 同步 config.yaml 中的主题配置 - 添加 Liquid Glass 主题说明注释 - **SKILL.md**:新增"HTML 主题(Liquid Glass,iOS 26 风格)"章节 - 核心特性说明 - YAML 配置示例 - 使用前提与文档链接 ### Fixed - **代码选中高亮“漂移/闪动”**:默认关闭 Liquid Glass 主题中的无限动画(尤其是代码块光泽划过效果),避免复制/选中时视觉干扰 - **Liquid Glass HTML 样式异常**:修复“页面顶部出现一大段 CSS 代码墙”的问题 - 移除 `includes.in_header` 直接 include `.css` 的写法,统一通过 `css:` 引入 - **浮动目录位置**:桌面端 TOC 默认固定到左侧(响应式逻辑保持不变) - **测试演示文档**:Liquid Glass demo 表格改为 `DT::datatable()`(不再使用 knitr 表格渲染) - **DataTables 表头/表体列宽错位**:将 Liquid Glass 的“卡片化 table”样式限制为 `table:not(.dataTable):not(.display)`(兼容 DT 初始化前的 `class="display"`),并对 `.dataTables_wrapper table.dataTable/.display` 禁用影响列宽计算的装饰样式,避免滚动模式下列对齐异常 - `bensz-rmd-rules/tests/v202602041426/`:轻量回归测试(PLAN/REPORT + 渲染产物) - **中文字体**:在 `templates/00.Environment.R` 增加绘图字体自动选择/注册逻辑,降低中文字符乱码风险 - **静态检查脚本**:修复 `scripts/check_htmlwidget_visibility.py` 对多行 `DT::datatable(...)` 等 widget 调用的误报 - **代码块滚动条**:移除代码折叠块底部的横向滚动条(通过强制换行与禁用横向 overflow) - **图片查看体验**:新增点击图片放大预览(Lightbox),并在模板 YAML 中默认启用(`includes.after_body`) - **图片可拷贝**:修复 Liquid Glass 主题 HTML 中右键菜单无法“拷贝图像”的问题 - `templates/liquid_glass_theme.css`:将 `img` 的 `transition: all` 收敛为仅动画视觉属性,并显式启用交互相关属性 - `templates/liquid_glass_lightbox.html`:Lightbox 仅响应左键点击,避免影响右键上下文菜单 ### Added - **新项目初始化脚本**:新增 `scripts/bootstrap_liquid_glass.py`,一键将 Liquid Glass 主题资源复制到新项目根目录 - **scripts/check_plot_readability.R**:新增 PDF 图表可读性基础检查脚本(文件存在/大小/可选文本提取) - **templates/nature_theme.R**:新增 `theme_nature_readable()`,为旋转标签/边距/图例外置提供更友好默认值 - **templates/complexheatmap_template.R**:新增 `make_heatmap_nature_safe()` 与 `truncate_labels()`,默认处理行/列名过长 - **references/plot_quality_standards.md**:新增 Nature 级别图表质量规范(跨包) - **templates/nature_colors.R**:新增 Nature 调色板(色盲友好) - **templates/nature_theme.R**:新增 ggplot2 Nature 级别主题 `theme_nature()` - **templates/complexheatmap_template.R**:新增 ComplexHeatmap 可读性模板(动态字体/尺寸) - **templates/plotly_template.R**:新增 plotly 交互图模板(统一字体/布局/配色) - **references/interpretation_narrative_examples.md**:新增“连贯叙述式”专家解读示例库 - 覆盖差异分析/生存分析/PCA/模型性能/缺失模式等常见输出类型 - 用自然段落内化四层内涵(数据描述 + 统计见解 + 领域映射 + 局限与后续),避免机械分点 - **references/expert_discussion_template.md**:新增专家级讨论模板文档 - 完整的讨论章节模板结构(基于 PanTCGA Mutation 分析) - 七大核心特征:量化陈述、动态数值、层次递进、辩证思考、具体示例、可操作建议、启发性 - 避免的浅层讨论 vs 推荐的深度讨论示例 - **references/four_tier_interpretation_framework.md**:新增四层解读框架文档 - AI 解读图表的独特优势说明(代码+数据 vs 视觉感知) - 最佳实践:动态数值嵌入(使用 Rmd 内联代码) - 数据溯源要求和解读-代码一致性验证 - 按分析类型的解读模板(表格类、图表类) - 解读质量自检清单 - **references/interpretation_templates.md**:新增“证据锚定版”深度解读模板 - 覆盖 3 类高频输出:单因素筛选、多因素模型、模型性能与验证 - 强制模板包含:Top 信号 + 方向/效应 + 不确定性 + 领域映射 + 可落地后续(方法+输入+判据) - **scripts/check_interpretation_quality.py**:新增解读静态预检脚本(可选) - 启发式检测:内联 `r ...` 数量、Top 信号提示、不确定性/稳定性提示、常见套话黑名单 - **references/hybrid_architecture_examples.md**:新增混合架构示例(从 SKILL.md 下沉) - 展示 `.R` 生成全量结果、`.Rmd` 基于 `params` 做阈值筛选的典型写法 - **references/hybrid_architecture_guide.md**:新增混合架构指南(从 SKILL.md 下沉) - 汇总 `.R/.Rmd/_functions/00.Environment/tmp` 的职责边界、模板入口与最小口径 - **references/plot_language.md**:新增图表语言规范文档(从 SKILL.md 下沉) - 英文优先 + `params.plot_language` 例外切换 + 示例与检查清单 - **references/no_overdefensive_code.md**:新增“不过度保护/禁止占位性代码”反模式说明(从 SKILL.md 下沉) - **references/code_block_explanations.md**:新增“代码块前解释(分析决策叙述)”模板与示例(从 SKILL.md 下沉) - **references/delivery_verification.md**:新增交付自检报告模板(从 SKILL.md 下沉) - **references/workflow_checklist.md**:新增工作流检查清单(从 SKILL.md 下沉) ### Changed - **htmlwidget/DT 输出可见性加固(避免“代码在但 HTML 不出表/不出图”)**: - `bensz-rmd-rules/SKILL.md`:新增“HTML 可见性硬规则(htmlwidget/DT)”与正误示例,并将“交付前强制检查”写成硬门槛 - `bensz-rmd-rules/README.md`:新增 FAQ,解释根因与推荐写法/检查命令 - `bensz-rmd-rules/templates/datatables_helper.R`:新增 `render_dt_output()`(用于稳定输出为可见结果) - `bensz-rmd-rules/templates/Rmd_template.Rmd`:表格示例改为 `render_dt_output(...)` 并确保为 chunk 最后表达式 - `bensz-rmd-rules/scripts/check_htmlwidget_visibility.py`:新增 Rmd 静态检查脚本,拦截 print()/invisible 包裹与“widget 非最后表达式”等高风险模式;版本号 `0.14.1 → 0.14.2` - **图表可读性硬检查升级**: - `bensz-rmd-rules/SKILL.md`:将“图表可读性自检”升级为“生成后必做硬检查”,并补充高频问题场景的最小可复用示例 - `bensz-rmd-rules/config.yaml`:新增 `plot_readability` 阈值配置(max_ticks/min_font_pt/long_label_chars/heatmap_max_label_chars);版本号 `0.14.0 → 0.14.1` - `bensz-rmd-rules/references/plot_quality_standards.md`:新增“可读性问题诊断流程”与“常见反模式”章节 - `bensz-rmd-rules/templates/nature_theme.R`:新增 `theme_nature_readable()`,减少长标签/旋转标签场景的重复样板代码 - `bensz-rmd-rules/templates/complexheatmap_template.R`:新增 `make_heatmap_nature_safe()`,默认处理行/列名过长 - `bensz-rmd-rules/README.md`:补齐 `plot_readability` 配置说明与新增模板/脚本的可发现性 - **图表默认质量升级(Nature 级别)**: - `bensz-rmd-rules/SKILL.md`:新增“图表质量规范(Nature 级别)”章节与必检清单 - `bensz-rmd-rules/config.yaml`:新增 `plot_quality` 配置;版本号 `0.13.4 → 0.14.0` - `bensz-rmd-rules/templates/Rmd_template.Rmd`:新增 plot-style 代码块与三包示例骨架 - `bensz-rmd-rules/references/delivery_verification.md` / `bensz-rmd-rules/references/workflow_checklist.md`:交付自检新增“图表质量(Nature)”检查项 - `bensz-rmd-rules/references/hybrid_architecture_examples.md` / `bensz-rmd-rules/references/interpretation_templates.md`:补齐图表质量与解读联动说明 - 兼容性与鲁棒性修复: - `bensz-rmd-rules/templates/nature_theme.R`:移除 `linewidth` 用法,兼容旧版 ggplot2 - `bensz-rmd-rules/templates/plotly_template.R`:修复 hover 默认值(避免空提示);保存前创建父目录 - `bensz-rmd-rules/templates/complexheatmap_template.R`:补齐数值/有限值校验与 `min==max` 处理;保存前创建父目录并校验尺寸 - `bensz-rmd-rules/templates/R_data_template.R` / `bensz-rmd-rules/templates/Rmd_template.Rmd` / `bensz-rmd-rules/references/cross_platform.md`:统一 `file.path()`,避免硬编码路径写法 - `bensz-rmd-rules/config.yaml`:`plot_quality.allowed_levels` 收敛为仅 `nature`(避免暗示未实现的多期刊切换) - **基因 ID 指南示例与口径修复**: - `bensz-rmd-rules/references/gene_id_guidelines.md`:移除 `biomaRt`/`useMart` 示例,统一为 `luckyBase::convert()`;示例路径改为 `file.path()` 并补齐目录创建 - **代码注释规范落地**: - `bensz-rmd-rules/SKILL.md`:新增代码注释规范(头部说明 + Step 注释 + 复杂逻辑说明) - `bensz-rmd-rules/templates/R_data_template.R`:补齐头部注释模板与 Step 注释示例 - `bensz-rmd-rules/templates/Rmd_template.Rmd`:补齐代码块头部注释模板与 Step 注释示例 - `bensz-rmd-rules/config.yaml`:版本号 `0.13.3 → 0.13.4` - **解读表达规范升级(标题/语气/数字/加粗)**: - `bensz-rmd-rules/SKILL.md`:新增“Rmd 标题规范”“解读语气与风格”“文本强调规范”“数字准确性验证(末尾检验)”,并在“结果后解读”补充“禁止代码生成解释”条款 - `bensz-rmd-rules/config.yaml`:版本号 `0.13.2 → 0.13.3`,description 同步表达规范口径 - `bensz-rmd-rules/templates/Rmd_template.Rmd`:补齐 `## 数字准确性验证` 章节骨架 - `bensz-rmd-rules/references/four_tier_interpretation_framework.md` / `bensz-rmd-rules/references/interpretation_narrative_examples.md` / `bensz-rmd-rules/references/interpretation_templates.md`:示例统一为论文口吻与自然标题,并加入关键观点适度加粗示例 - `bensz-rmd-rules/scripts/check_interpretation_quality.py`:新增可选表达规范检查 flags(标题括号/主-副标题/教学式标记/加粗密度),默认关闭以保持兼容 - `bensz-rmd-rules/tests/v202601211445/`:新增轻量测试会话,覆盖脚本默认检查与新增 flags 行为 - **luckyBase 硬前提口径统一**: - `bensz-rmd-rules/SKILL.md`:资源优先级调整为“包加载集中在 `00.Environment.R`(`luckyBase::Plus.library()`)+ 分析脚本优先 `pkg::fn()`” - `bensz-rmd-rules/templates/00.Environment.R`:补齐 luckyBase 的最小依赖边界与集中包加载骨架(不提供降级) - **YAML 头单一真相来源**: - `bensz-rmd-rules/SKILL.md`:移除 YAML 头的硬编码示例,改为引用 `config.yaml:rmd_template.yaml_header` - `bensz-rmd-rules/templates/Rmd_template.Rmd`:移除 `author` 行(避免与 config 重复口径) - `bensz-rmd-rules/README.md`:配置项表格改为引用 `rmd_template.yaml_header`,不再单独展示 `author` - **SKILL.md 瘦身(渐进披露)**: - 将“实践示例/模板代码/长篇示例”下沉到 `references/` 与 `templates/`,SKILL.md 仅保留硬规则与入口指引 - **“完整因果链”实践形式升级(从模板化到内化式叙述)**: - `bensz-rmd-rules/SKILL.md`:将“代码块前解释”升级为分析决策叙述(数据特征 → 方法选择 → 输出解读),并补充 3 个不同分析类型示例 - `bensz-rmd-rules/SKILL.md`:在“结果后解读”中新增“连贯叙述式 vs 分项式”的双模式指南,并提供段落骨架与写完后自检清单 - `bensz-rmd-rules/SKILL.md`:在“交付验证”新增“解读-代码一致性自检”与报告检查项(关键数值/Top 信号/不确定性可追溯) - `bensz-rmd-rules/references/four_tier_interpretation_framework.md`:新增连贯叙述式框架、正向写作指南(不确定性/Top 信号句式)与分项式 vs 连贯式对比示例 - `bensz-rmd-rules/README.md`:补充“四层是内容门槛而非固定格式”的说明,并增加示例库入口 - **SKILL.md 瘦身优化**(遵循渐进披露原则): - **R 包资源清单**:删除镜像展示,改为引用 `config.yaml:r_packages`,消除维护冗余 - **结果后解读章节**:从 ~180 行精简至 ~20 行,详细内容移至 `references/four_tier_interpretation_framework.md` - **末尾讨论与分析章节**:从 ~80 行精简至 ~40 行,详细模板移至 `references/expert_discussion_template.md` - 总行数:从 1161 行精简至 ~969 行(减少约 16%) - **自动化验证增强**: - 路径验证脚本说明从"可选"改为"推荐" - 新增两种运行方式说明(R 中运行、命令行运行) - 新增检测项说明(硬编码路径、绝对路径、路径拼接、路径分隔符) - 新增引用 `references/cross_platform.md` - **四层解读框架通用化**: - 第三层从"生物学/临床见解"改为"领域见解"(domain insights) - 适用于所有数据分析领域,不限定生信/医学场景 - **config.yaml**: - `skill_info.version`:0.10.0 → 0.11.0 - `skill_info.description`:移除领域限定,强调通用性 - **结果解读质量门槛(反套话 / 证据锚定)**: - `references/four_tier_interpretation_framework.md`:新增 Fail Fast Gate、结果锚点句式库、黑名单套话→替代写法对照表,并补充可选静态预检说明 - `SKILL.md`:在“结果后解读”章节新增硬性最小内容(Minimum Content Requirements)与反套话机制,并在工作流检查清单中新增硬门槛检查项 - `README.md`:术语统一为“领域见解”,并补充证据锚定要求与参考文档入口 - **config.yaml**: - `skill_info.version`:0.11.0 → 0.12.0 - `skill_info.version`:0.12.0 → 0.13.0 - `skill_info.description`:补充“连贯叙述式写法”和“解读-代码一致性自检” - **config.yaml**: - `skill_info.version`:0.13.0 → 0.13.1 - `skill_info.description`:补充 luckyBase 为硬依赖的前提说明 - **config.yaml**: - `skill_info.version`:0.13.1 → 0.13.2 - **SKILL.md 瘦身(继续下沉到 references)**: - 将“图表语言规范/混合架构细节/反模式/代码块前解释/交付验证/检查清单”等内容进一步下沉到 `references/` - SKILL.md 保留硬规则与入口链接,降低维护冗余与口径漂移风险 ### Removed - 移除 SKILL.md 中的包清单镜像展示(仅保留简要表格,完整清单引用 config.yaml) - 移除 SKILL.md 中专家级讨论模板的长篇示例(已移至 references/expert_discussion_template.md) - 移除 SKILL.md 中四层解读框架的详细说明(已移至 references/four_tier_interpretation_framework.md) - **禁止占位性代码模式**:在"不过度保护原则"章节中新增占位性代码的禁止规则 - **占位性代码定义**:使用 `try()` 捕获错误后,将结果赋值为 `NULL`、`NA`、空数据框,或直接跳过/打印警告继续,从而保证代码表面运行成功的模式 - **绝对禁止的占位性模式表格**:列出四种危险模式(try-catch 后赋值 NULL、打印警告继续、降级到占位符、条件分支后无有效逻辑)及其危害 - **切实落地要求**:功能必须确保在正常流程中切实落地,不得使用占位符保证"不报错";如功能失败应 `stop()` 报错;如需容错必须有有效降级方案 - **正确示例对比**:提供占位性代码 vs 正确处理的对比示例(包含 LDA 建模和包依赖场景) - **权衡原则**:科研分析 > 代码美观、显式失败 > 静默成功、有效降级 > 占位符 - **交付验证更新**:自检报告新增"无占位性代码"检查项,自检报告示例同步更新 - **工作流检查清单更新**:新增两项占位性代码相关检查项 - **图表语言规范**:新增可视化图表英文优先原则 - 在 SKILL.md"人类可读原则"章节后新增"图表语言规范"小节 - 核心规则:所有可视化图表的文本元素(轴标题、图例、标题、副标题)必须使用英文 - 提供正确/错误实践示例对比 - 特殊场景例外说明:中文期刊投稿可通过 YAML params 设置 `plot_language: "zh"` - 实现建议:提供全局标签函数模板和 YAML params 配置方式 - 检查清单:四项自检标准确保图表语言符合规范 - **图表语言配置**:在 config.yaml 中新增 `plot_language` 配置项 - 默认语言:`"en"`(英文) - `force_english: true` 强制英文规则 - 支持通过 YAML params 覆盖配置(用于中文期刊投稿等特殊场景) - **gene_id_guidelines.md 示例更新**:将图表示例代码中的中文轴标题改为英文 - 原示例:`labs(x = "基因", y = "表达量")` - 更新为:`labs(x = "Gene Symbol", y = "Expression Level")` - 新增反面示例展示应避免的中文轴标题用法 - **基因 ID 优先级指南**:新增生物医学分析专用的基因标识符使用规范 - 新增 `references/gene_id_guidelines.md`:完整的基因 ID 优先级与可读性指南 - 展示优先级:可视化、表格、文本解读中使用 SYMBOL(如 `TP53`),而非 ENSEMBL ID 或 ENTREZID - 数据完整性:保存的数据应包含 SYMBOL、ENSEMBL、ENTREZID 多种 ID,确保准确性和可追溯性 - 核心理由:增强可读性,让临床医生、生物学家等非生信专业读者也能理解 - 在"完整因果链原则"章节中添加引用和简要说明 - **数据筛选分离实例化说明**:新增基于真实项目的代码示例 - .R 脚本示例:展示如何保存全量表(不筛选 p/q 值) - .Rmd 脚本示例:展示如何动态应用阈值筛选 - 阈值参数配置原则:明确阈值仅在 .Rmd 的 YAML params 中定义 - YAML params 配置示例:提供完整的阈值参数配置模板 - **postprocess-only 快速模式**:新增针对耗时分析的优化模式 - 支持通过环境变量控制运行模式(`DV_MUT_MODE`) - 首次运行后,调整可视化时无需重跑耗时计算 - 在 Rmd 中通过 params 控制运行模式 - 使用方式和好处说明 - **文件写入最佳实践**:新增跨平台兼容的文件操作规范 - 安全写入 CSV 函数示例(`.dvmut_safe_write_csv`) - 安全读取 CSV 函数示例(`.dvmut_safe_read_csv`) - 统一路径分隔符、自动创建目录、统一 UTF-8 编码 - 错误处理和降级方案 - **专家级讨论模板**:新增基于真实项目的深度讨论写作指南 - 完整的讨论章节模板(约900字级别) - 专家级讨论的七大核心特征(量化陈述、动态数值、层次递进、辩证思考、具体示例、可操作建议、启发性) - 避免浅层讨论的反面示例 - 推荐的深度讨论示例 - **自检报告真实示例**:新增基于 PanTCGA 项目的完整自检报告 - 展示如何填写每个检查项的证据/说明 - 包含 15 个检查项的详细证据 - 新增"专家级讨论"检查项 - **DT::datatable() 标准化调用**:新增 `templates/datatables_helper.R` 辅助函数 - 提供 `render_dt(data, n, scrollX, pageLength)` 标准化接口 - 减少重复代码,统一表格渲染方式 - DT 包已通过 `00.Environment.R` 加载,直接使用即可 - **references/candidate_r.md**:新增 candidate_r 详细使用指南 - 明确 candidate_r 的可选性和使用场景 - 提供两种使用模式(已加载 vs 指定路径) - 强调安全提示和实践原则 - **references/cross_platform.md**:新增跨平台兼容性最佳实践指南 - 路径拼接、文件 I/O、平台差异处理详细说明 - 常见陷阱和跨平台测试清单 ### Changed - **config.yaml**: - `rmd_template.yaml_header.author`:从占位符改为“技能作者名”(仅在 config.yaml 托管) - `rmd_template.yaml_header.output.html_document.number_sections`:新增 `true` - 新增 `rmd_template.datatables_helper` 配置,指向 DT 辅助函数 - `skill_info.version`:0.7.0 → 0.10.0 - `skill_info.description`:更新为包含实例化说明、postprocess-only 模式、安全写入函数、专家级讨论模板、自检报告示例、禁止占位性代码模式 - **SKILL.md**(功能增强 + 瘦身优化): - **数据筛选分离原则**:新增优秀实践示例(.R 和 .Rmd 代码对比) - **阈值参数配置原则**:新增配置原则说明和 YAML params 示例 - **数据脚本规范**:新增"支持快速后处理模式"要求和完整实现示例 - **跨平台兼容性原则**:新增"文件写入最佳实践"章节(安全写入/读取函数) - **末尾讨论与分析**:新增专家级讨论模板(基于 PanTCGA Mutation 分析) - **专家级讨论的核心特征**:新增七大特征对比表和正反示例 - **交付验证**:新增自检报告的真实示例(15 个检查项) - **candidate_r 章节**:简化为核心原则,详细用法移至 `references/candidate_r.md` - **跨平台兼容性章节**:从 ~160 行精简至 ~30 行,详细内容移至 `references/cross_platform.md` - **YAML header 章节**:改为引用 `config.yaml:rmd_template.yaml_header`,确保单一真相来源 - **表格渲染章节**:新增 DT 辅助函数的使用方式,提供两种调用方法 - **交付验证章节**:YAML 规范检查项改为"见 config.yaml" - **工作流检查清单**:最后一项更新为"见 config.yaml" - 总行数:约 1000+ 行(包含新增的实例化内容) - **README.md**: - 新增"快速迭代可视化(postprocess-only 模式)"使用场景 - 新增"阈值参数灵活配置"和"快速后处理模式"核心特性 - 新增"专家级讨论模板"核心特性说明 - 更新核心特性表格(从 6 项扩展到 9 项) - **metadata.version**:0.8.0 → 0.9.0 ### Removed - 移除 SKILL.md 中 candidate_r 的重复代码示例(已在 references/candidate_r.md 详细说明) - 移除 SKILL.md 中跨平台兼容性的详细示例代码(已在 references/cross_platform.md 详细说明) ## [0.6.0] - 2026-01-20 ### Added - **专家级结果解读框架**:新增四层解读指导原则,避免仅有"描述"而无"见解"的浅层解读 - **第一层:数据描述**:清晰陈述输出的基本结构和内容 - **第二层:统计见解**:解释效应大小和方向的实际意义 - **第三层:生物学/临床见解**:将统计结果与生物学机制/临床实践关联 - **第四层:局限与后续**:指出局限性和不确定因素,提出验证建议 - **AI 图表解读的独特优势**:明确 AI 应基于**代码+原始数据**解读图表,而非"看图说话" - 核心原则:AI 无需查看渲染后的图像,直接分析生成图表的代码和原始数据 - 实践方式:读取代码理解图表类型和参数,分析数据提取精确数值,生成量化解读 - 优势对比:精确量化、无视觉偏差、可复现 - **最佳实践:动态数值嵌入**:解读中的关键数值应使用 `` `r ...` `` 动态生成,确保精确且可追溯 - **数据溯源要求**:解读中提及的数值必须明确其来源(变量名、计算逻辑) - **解读-代码一致性验证**:自检每个数值是否都能在代码/数据中找到对应来源 - 检查清单:六项标准确保基于代码和数据的精准解读(含动态嵌入和数据溯源) - **示例对比**:提供浅层解读(仅有描述)与专家级解读(描述 + 见解)的具体对比示例 - **解读模板**:新增表格类和图表类输出的标准化解读模板 - **解读质量自检清单**:新增六项自检标准(数据层、统计层、科学层、局限层、后续层、避免套话) - **交付验证更新**:自检报告新增"专家级解读"、"避免浅层描述"、"AI 图表解读"检查项 - **工作流检查清单更新**:新增四项解读质量相关检查项(含 AI 图表解读) ### Changed - **完整因果链原则章节重构**:从简单示例升级为系统化的解读框架 - **末尾讨论与分析章节增强**: - 新增"与现有研究的关联"小节 - 局限性细化为数据/方法/解释三个层面 - 进一步分析建议细化为短期/中期/长期三个方向 - **config.yaml**: - `skill_info.version`:0.5.0 → 0.6.0 - `skill_info.description`:新增专家级结果解读框架说明 - **SKILL.md**: - `metadata.version`:0.5.0 → 0.6.0 - `metadata.short-description`:更新为包含专家级结果解读框架 - YAML frontmatter `description`:新增专家级解读说明 ## [0.5.0] - 2026-01-20 ### Added - **数据筛选分离原则**:明确 .R 和 .Rmd 关于数据筛选的职责分工 - .R 脚本:生成未经阈值筛选的完整数据集,保存所有样本和特征,方便用户审查数据最原始的状态 - .Rmd 分析:根据具体业务需求和统计要求,对 .R 提供的完整数据应用阈值筛选(如表达量阈值、p值阈值、样本量要求等),从而保证� -
config.yaml 14.9 KB
# bensz-rmd-rules 配置文件 # R Markdown 开发规范与最佳实践 # # 注意: # - Rmd 的 YAML 头(含 author)以 `rmd_template.yaml_header` 为准(单一真相来源)。 # 为了便于“复制即用”,两个 Rmd 模板会保留 YAML header 的拷贝; # 维护者修改本节后,应同步更新模板(用 `scripts/check_rmd_template_yaml.py` 校验)。 # - luckyBase 是本 skill 的硬前提:包加载统一通过 `luckyBase::Plus.library()` 完成。 skill_info: name: bensz-rmd-rules version: "0.29.0" description: "规划并实现可复现的 R/R Markdown 分析流水线及论文级统计推断:先逐项判断关键估计对象的区间、检验与可估计性,再将完整计算结果对应到报告。新项目按复杂度选择 simple 或 targets-first complex/pipeline,两种模式均由 renv 锁定环境并通过真实轻量测试;complex 以 _targets.R + R/ 为唯一计算编排,Rmd 只消费 targets 结果。已有项目不自动迁移或补齐机制:有 _targets.R 的按 complex 维护,无 targets 的(含历史编号脚本)按 simple 语义维护,需要复杂能力时显式迁移;不维护任何历史编号 runner。前提:luckyBase 为唯一固定 Bensz 依赖" author: "Bensz Conan" category: "开发规范" # R 包资源配置 # 固定依赖仅保留 luckyBase;其它包由项目声明并由 renv.lock 锁定。 r_packages: fixed: - name: luckyBase description: "Bensz 基础包:包加载、项目初始化与稳定通用辅助能力" project_managed: description: "具体分析所需的 CRAN、Bioconductor 或用户自有包;必须写入项目 renv.lock,不在 Skill 中预设" # 项目状态与生命周期契约 project_strategy: state_detection: new_project: "空目录或用户明确声明新建" existing_project: "存在 R/Rmd、raw、products、reports、runner、_targets.R 或 renv/renv.lock" ambiguous: "只读盘点后请求澄清,不猜测迁移意图" new_project: require_renv: true default_workflow_mode: "simple" workflow_modes: - "simple" - "complex" complex_alias: "pipeline" pipeline_first: true human_override_precedence: true existing_project: workflow_modes: ["simple", "complex"] mode_resolution: "已有 _targets.R 判 complex;否则(含历史编号脚本)按 simple 语义维护,编号脚本只是顺序执行的普通 Rscript 入口" create_missing_targets: false create_missing_renv: false migrate_layout: false maintain_numbered_runner: false targets: entrypoint: "_targets.R" store_dir: "_targets" required_for_workflow_modes: - "complex" test_store_dir: "tmp/tests/{run_id}/_targets" output_contract: "targets 管理计算状态、依赖、失效与增量重建;products/ 与 reports/ 只保存需要审阅、复用或交付的科学产物" function_dir: "R" source_function_call: 'targets::tar_source("R")' report_consumer: "Rmd 通过 tar_read()/tar_load() 或受控 tarchetypes::tar_render() 消费 target" recovery_acceptance: "中断后再次 tar_make() 必须复用仍有效的前序 target" renv: lockfile: "renv.lock" activation: "renv/activate.R" required_for_new_project: true initialize_once: true silent_install_or_snapshot_during_run: false excluded_paths: - "renv/library" - "renv/staging" - "renv/cache" # 新建或实质修改的分析流程都必须真实执行一种轻量测试。 lightweight_testing: required_for_new_or_materially_changed_flow: true styles: - "synthetic_fixture" - "project_subset" default_without_authorized_data: "synthetic_fixture" preferred_with_authorized_existing_data: "project_subset" entrypoint: "scripts/tests/smoke_test.R" source_dir: "scripts/tests" run_dir: "tmp/tests/{run_id}" path_overrides: input: "BENSZ_ANALYSIS_INPUT" products: "BENSZ_PRODUCTS_DIR" reports: "BENSZ_REPORTS_DIR" run_root_guard: "BENSZ_TEST_RUN_ROOT" require_same_formal_entrypoint: true forbid_business_logic_test_flags: true required_attempt_fields: - "command" - "exit_status" - "failure_category" - "fix_summary" - "assertions" - "remaining_risks" failure_categories: - "environment_dependency" - "fixture_or_subset" - "path_isolation" - "analysis_code" - "scientific_assertion" - "report_rendering" - "external_service" execution_statuses: - "preflight" - "lightweight_execution" - "full_data_execution" default_full_data_execution: "NOT_RUN" cleanup_project_subset_by_default: true # Rmd 模板配置 rmd_template: yaml_header: title: "AA.BB.CC. 报告标题" author: "Bensz Conan" date: "`r Sys.Date()`" params: plot_language: "en" top_n: 3 q_cutoff: 0.05 min_sample_n: 30 interpret_top_n: 3 interpret_min_stratum_n: 30 interpret_min_consistent_prop: 0.70 interpret_min_boot_prop: 0.70 # Plot delivery helpers (optional): # - plot_run_dir: current .bensz-api/task-* root; default reads BENSZ_TASK_ROOT # - plot_preview_dpi: raster preview DPI (higher => clearer but larger files) plot_run_dir: ~ plot_preview_dpi: 200 output: html_document: toc: true toc_float: true toc_depth: 3 number_sections: true theme: default highlight: tango code_folding: show includes: after_body: "templates/liquid_glass_lightbox.html" css: "templates/liquid_glass_theme.css" # DT::datatable 标准化调用 datatables_helper: "templates/datatables_helper.R" # 注意:完整的 YAML 头配置说明、Liquid Glass 主题特性、自定义选项详见: # references/liquid_glass_theme_guide.md(说明文档;具体字段以本文件 rmd_template.yaml_header 为准) # 图表语言配置 plot_language: # 默认语言:en(英文)或 zh(中文) default: "en" # 强制英文:图表轴标题、图例、标题等必须使用英文 # 例外场景:中文期刊投稿可在 YAML params 中设置 plot_language: "zh" force_english: true # 图表质量配置(默认即 Nature 级别) plot_quality: default_level: "nature" supported_packages: - ggplot2 - ComplexHeatmap - plotly allowed_levels: - nature # 图表可读性硬检查阈值(经验默认值,可按项目需求调整) # 用途:指导 AI 在生成绘图代码后“必须显式检查并修复”的常见问题。 plot_readability: # 轴刻度过密阈值(> max_ticks 时,优先减少 breaks 或采样显示) max_ticks: 20 # 最小可接受字体(pt);低于该值通常打印/投影不可读 min_font_pt: 8 # 分类标签过长的参考阈值(字符数);超过后优先旋转/换行/缩写 long_label_chars: 20 # Heatmap 行/列名建议最大长度;超过后优先截断并保留前缀 heatmap_max_label_chars: 15 # 文件结构配置(新项目默认) file_structure: function_dir: "R" pipeline_entrypoint: "_targets.R" report_script: "{report}.Rmd" rendered_report: "{report}.html" environment_script: "00.Environment.R" r_dir: "R" raw_dir: "raw" products_dir: "products" reports_dir: "reports" styles_dir: "templates" operations_dir: "scripts/operations" library_dir: "scripts/lib" tests_dir: "scripts/tests" test_runs_dir: "tmp/tests/{run_id}" scratch_dir: "tmp/scratch" product_storage: default_dir: "products" environment_override: "BENSZ_PRODUCTS_DIR" require_single_project_setting: true stay_inside_project: true forbidden_roots: - "raw" - "reports" - "tmp" - "_targets" - ".bensz-api" # 文件命名约定(pipeline 使用语义化 target 命名,不使用编号脚本) naming_convention: target_name_pattern: '^[A-Za-z][A-Za-z0-9_.-]*$' report_name_pattern: '^[A-Za-z0-9_. -]+$' # 输出配置 output: html_skill: "knit-rmd-html" html_location: "project_root" figures_dir: "reports/figures" tables_dir: "reports/tables" supplementary_dir: "reports/supplementary" # 多分析单元、缓存和恢复的稳定默认值 analysis_workflow: plan_file: "analysis-plan.yaml" default_workflow: "pipeline" product_path: "{products_dir}/<scientific-product>" contract_version: 2 cache_identity: "targets metadata and store only; no SUCCESS/identity runner" report_dependency: "reports consume target values and must not read raw/ to repeat expensive computation" recovery: required: true command: "targets::tar_make()" evidence: ["tar_meta", "outdated target set", "execution log"] controls: force_rebuild: "targets::tar_make(callr_function = NULL) or explicit targets invalidation" validation: default_mode: "report" delivery_mode: "strict" serial_reviews: - "structure-dataflow" - "r-implementation-recovery" - "scientific-statistical" observability: default: "targets" progress: ["targets::tar_poll", "targets::tar_watch"] worker_metrics: ["crew logging", "crew_options_metrics"] resource_diagnostics: ["autometric::log_start", "autometric::log_read", "autometric::log_plot"] policy: "Skill 只组合成熟组件,不定义新的事件字段、调度器、worker 状态机或 dashboard" parallel: optional: true controller: "crew" enable_when: "存在足够多相互独立且昂贵的 target,收益覆盖 worker 启动与传输开销" default_workers: 1 oversubscription_guard: "限制 worker 数与 BLAS/OpenMP/future/BiocParallel 内部线程数乘积不超过资源预算" random_seed: "在 target 图中显式声明可复现随机数策略" # 指标解释协议:常用指标使用白名单;未命中、无法确定或非标准定义时按不常用指标处理。 metric_explanation: audience_assumption: "相关背景较弱,但需要获得专业、准确且可复核的解释" classification: "common_allowlist_else_uncommon" require_guide_table_for_uncommon: true require_full_first_mention: true concise_after_first_mention: true reset_when_definition_changes: true common_metrics: descriptive: - "样本量/频数/比例" - "均值/中位数/标准差/四分位距" - "发生率" inference: - "p 值" - "q 值/FDR" - "95% 置信区间" effect_and_association: - "均值差/回归系数" - "OR/RR/HR" - "Pearson r/Spearman rho" survival: - "Kaplan-Meier 生存概率/中位生存期" - "log-rank 检验" - "Cox HR" - "C-index" diagnostic_and_binary_prediction: - "ROC AUC" - "灵敏度/特异度" - "阳性预测值/阴性预测值" - "准确率" # 图表/表格解读覆盖检验(硬编码静态检查) # 目的:彻底解决“有图表/表格但缺少解读”的漏项。 figure_interpretation_check: enabled: true # 默认仅报告(不阻断),交付前请在 CLI 显式加 --strict 做“Fail Fast”门禁 strict_mode: false required_layers: 4 # 期望四层解读框架(脚本主要做覆盖检查,层数作为口径) max_distance_lines: 50 # 允许解读距离图表/表格输出块的最大行距 require_reference: false # 是否强制要求显式引用“Figure/图/Table/表 N”(更严格但更易误杀) # 是否要求解读段落包含某些“解读/结论/观察”等标记词(默认 true;marker 更中性,不再强迫四层标签) # - 设为 false:仅做“有 prose + 最低长度/要素”门禁 # - 设为 true 且 markers 为空:自动跳过 marker 检查(避免误杀) require_markers: true # 解读最小长度门槛(脚本默认值从此读取,CLI 仍可覆盖) min_cjk_chars: 80 min_en_words: 40 # 证据锚定要素数量(启发式 proxy,避免“空话解读”) min_content_elements: 2 check_patterns: figure_generation: - "\\bggplot\\s*\\(" - "\\bgeom_[a-zA-Z0-9_]+\\s*\\(" - "\\bplot\\s*\\(" - "\\bheatmap\\b" - "\\bComplexHeatmap\\b" - "\\bHeatmap\\s*\\(" - "\\bpheatmap\\s*\\(" - "\\bggsave\\s*\\(" - "\\bpdf\\s*\\(" - "\\bpng\\s*\\(" - "\\btiff\\s*\\(" - "\\bjpeg\\s*\\(" table_generation: - "\\bknitr::kable\\s*\\(" - "\\bkable\\s*\\(" - "\\bDT::datatable\\s*\\(" - "\\brender_dt_output\\s*\\(" - "\\brender_dt\\s*\\(" - "\\bgt::gt\\s*\\(" - "\\bflextable::flextable\\s*\\(" - "\\breactable::reactable\\s*\\(" interpretation_markers: # 更中性的 marker:用于识别“这是在写解读/结论/观察”,而非强制四层标签写法 - "解读" - "结论" - "小结" - "观察" - "核心结论" - "结果(?:显示|表明|提示|可见)" - "\\binterpretation\\b" - "\\btakeaway\\b" - "(?:图|表|Figure|Table)\\s*\\d+.*?(?:解读|分析|结果|interpretation|analysis)" # 解读质量静态预检(启发式检查) # 脚本:scripts/check_interpretation_quality.py # 说明:脚本优先使用此配置;若未配置则使用脚本内置默认值(降级) interpretation_quality_check: enabled: true min_inline_r: 3 min_current_observation: 3 # 推荐结构(默认不强迫;交付/收敛阶段可用 --strict 或改此配置启用) require_top_hits: false require_uncertainty_hits: false require_actionability: false # 数字可追溯:识别 prose 中“字面数字”但缺少附近 `r ...` 的情况(启发式) check_untracked_numbers: true max_untracked_numbers: 1 # 允许的“字面数字”白名单(避免误杀:年份/图表编号等) untracked_number_exempt_patterns: - "\\b20\\d{2}\\b" - "(?:Figure|Table|图|表)\\s*\\d+" check_tone: true max_teaching_hits: 2 check_title_style: false check_bold: false max_bold_per_1000: 10.0 check_actionability: true max_actionability_issues: 2 check_template_labels: true max_tier_label_lines: 8 teaching_markers: - "提示:" - "注意:" - "需要注意的是" - "即:" - "即:" - "用于快速识别" - "用于快速判断" - "帮助你判断" - "用于把方向落到" # 硬黑名单:命中即阻断交付(尽量少用;优先用“反模板/可证伪/可执行”检查替代) blacklist_patterns: [] # 机械模板检测:命中即判定为“结构填空”,通常不看数据也能写出来 check_mechanical_templates: true max_mechanical_template_hits: 0 mechanical_template_patterns: - "这张(?:图|表|图/表).{0,20}?(?:在)?本次.{0,20}?直接观察是[::].*?统计.{0,20}?含义是[::].*?研究者.{0,20}?意义是[::].*?(?:你可以)?(?:立即)?(?:执行)?的?下一步" # 空泛句式检测:这些短语不一定永远错误,但若附近缺少“对象+数值证据”,通常意味着在说官话 check_vague_phrases: true vague_phrase_window_chars: 160 max_vague_phrase_issues: 2 vague_phrase_patterns: - "提示可能" - "可能提示" - "值得深入探讨" - "需要进一步研究" - "建议进一步研究" - "建议进一步验证" - "需要进一步验证" - "为.*提供依据" -
README.md 8.1 KB
# bensz-rmd-rules 面向 **R 数据分析、R Markdown 报告、可复现 targets 流程、论文级图表与证据解读** 的 Agent Skill。当前版本以 [`config.yaml`](config.yaml) 中的 `skill_info.version` 为准;本目录属于 beta 候选源。 ## 什么时候使用 使用场景:从原始数据完成整理、统计/模型、图表、Rmd/HTML、科学产品和结果解释,或维护已有 R/Rmd 项目。 不要用于:主要交付物是跨项目复用的 R 函数、类、稳定 API 或 Package(改用 `bensz-r-developer`);仅渲染既有 Rmd(改用 `knit-rmd-html`);其它语言的数据分析。 ## 最短用法 ```text 请用 bensz-rmd-rules 完成这个 R 分析。先只读判断项目是 new 还是 existing,再说明选择 simple 或 targets-first complex/pipeline 的理由。使用 renv,并在正式运行前从正式入口真实跑通 synthetic_fixture 或 project_subset 轻量测试。 ``` Skill 会保留已有项目的现状:不因默认策略自动创建 targets/renv 或迁移目录;有 `_targets.R` 的项目按 complex 维护,无 targets(含历史编号脚本)的项目按 simple 语义维护。 ## 工作流模式 | 模式 | 适用场景 | 固定要求 | | --- | --- | --- | | `simple` | 线性、低成本、整体重跑可接受 | `renv`、明确 R/Rmd 入口、真实轻量测试;不创建 targets | | `complex`(说明性别名 `pipeline`) | 非线性依赖、昂贵步骤、多下游复用、局部失效、恢复或并行 | `_targets.R` 唯一 DAG、`R/` 计算函数、Rmd 消费 target、隔离 store 恢复验收 | `new`/`existing` 是项目状态而不是第三种模式:已有 `_targets.R` 按 complex,否则按 simple。迁移必须由人类明确授权。 complex 的核心关系: ```text renv.lock → _targets.R → R/ functions → target results → Rmd → reports/ ↘ optional products/ ↘ _targets/ (machine state) ``` `_targets/` 只存机器计算状态;`products/` 只存需要审阅、复用或交付的科学对象,不实现第二套缓存、`SUCCESS`、identity hash 或恢复 runner。 ## 推荐项目布局 ```text 项目根目录/ ├── 00.Environment.R ├── R/ # target 调用的计算函数 ├── _targets.R # complex 的唯一 DAG 入口 ├── raw/ # 只读 ├── products/ # 可选科学产物 ├── reports/ # 图、表、HTML、补充材料 ├── scripts/tests/ # 可版本化测试代码 ├── tmp/tests/<run-id>/ # 隔离测试现场 ├── renv.lock └── renv/activate.R ``` 模板 [`_targets.R`](templates/_targets.R) 使用 `tar_source("R")`;[`R_data_template.R`](templates/R_data_template.R) 提供数据处理起点;[`Rmd_template.Rmd`](templates/Rmd_template.Rmd) 通过 `targets::tar_read()` 消费 `analysis_results`;simple 可从 [`Rmd_simple_template.Rmd`](templates/Rmd_simple_template.Rmd) 开始。 ## 统计推断完整性 论文级分析在计算前列出主要及影响结论的次要估计对象,确认目标人群、观察单位、对比、设计和推断目的。方法允许时,完整结果保留估计值与单位、有效 N/事件数、适当的默认 95% CI;只有零假设明确时才报告 p 值,批量检验同时说明校正方法及 q/调整后 p。预注册方案另定置信水平时按方案执行。报告引用计算结果并解释效应的实际意义和局限。固定常数或纯描述结果可标 `not_applicable`;数据或设计无法支持推断时标 `not_estimable` 并说明原因与替代分析;未运行标 `not_run`。不凭空补 CI/p,也不以文本关键词检查代替方法复核。详见 [`statistical_inference_protocol.md`](references/statistical_inference_protocol.md)。 ## 轻量测试与恢复 新建或实质修改的流程必须真实运行一种测试: - `synthetic_fixture`:无授权真实数据、数据敏感或数据过大时,保留 schema、类型、主键、分组、缺失和边缘条件,并固定随机种子。 - `project_subset`:有授权数据和现有代码时使用代表性子集,覆盖关键分组/结局/缺失/异常;不能只用 `head(n)`。 测试代码放 `scripts/tests/`,每次使用唯一 `tmp/tests/<run-id>/`。simple 调用正式入口;complex 执行同一 DAG,但 store 放在 `<run-id>/_targets`,不能污染正式 `_targets/`、`products/` 或 `reports/`。至少断言输入契约、关键类型/主键、重要数值不变量、产品/报告生成及 `raw/` 无写入。未跑全量时记录 `full_data_execution=NOT_RUN`。 恢复验收:先让昂贵 target 成功,再中断后续 target;在输入、代码、参数和 renv 身份不变时再次 `tar_make()`,用 `tar_meta()`、outdated 集合和执行记录证明有效前序 target 被跳过。 ## 检查与脚本入口 先检查项目状态和模式: ```bash python3 <skill-root>/scripts/check_targets_renv.py <project> --project-state auto --workflow-mode auto ``` 常用检查: ```bash python3 <skill-root>/scripts/check_pipeline_contract.py <project> python3 <skill-root>/scripts/check_interpretation_quality.py <project>/report.Rmd python3 <skill-root>/scripts/check_figure_table_interpretation.py <project>/report.Rmd python3 <skill-root>/scripts/check_htmlwidget_visibility.py <project>/report.Rmd python3 <skill-root>/scripts/check_rmd_template_yaml.py <project>/report.Rmd Rscript <project>/scripts/tests/smoke_test.R ``` 图表可读性检查器为 [`check_plot_readability.R`](scripts/check_plot_readability.R),路径安全检查器为 [`validate_paths.R`](scripts/validate_paths.R)。Liquid Glass 主题可用 [`bootstrap_liquid_glass.py`](scripts/bootstrap_liquid_glass.py) 初始化;桌面动态目录会立即扩展可交互区域,避免鼠标移入时收回。HTML 渲染交给 `knit-rmd-html`。 ## 图表与解读规范 默认图表语言为英文;中文期刊等场景通过 YAML `params.plot_language` 切换。每个可见图/表附近说明当前对象、方向/对比、可追溯数值、量级、不确定性和后续验证(方法 + 输入 + 判据)。参考: - [`four_tier_interpretation_framework.md`](references/four_tier_interpretation_framework.md):四层解读与 Fail Fast Gate。 - [`interpretation_templates.md`](references/interpretation_templates.md):单因素、多因素和模型验证骨架。 - [`plot_quality_standards.md`](references/plot_quality_standards.md):Nature 级图表可读性。 - [`liquid_glass_theme_guide.md`](references/liquid_glass_theme_guide.md):HTML 主题与故障排查。 ## 相关模板与参考 - 环境与依赖:[`00.Environment.R`](templates/00.Environment.R)、[`renv/activate.R`](templates/renv/activate.R)。 - 测试:[`templates/tests/`](templates/tests/)、[`lightweight_testing.md`](references/lightweight_testing.md)。 - 模式与架构:[`workflow_modes.md`](references/workflow_modes.md)、[`hybrid_architecture_guide.md`](references/hybrid_architecture_guide.md)。 - 交付与审查:[`delivery_verification.md`](references/delivery_verification.md)、[`serial_review_protocol.md`](references/serial_review_protocol.md)。 ## FAQ 与边界 **可以把编号脚本自动迁移成 targets 吗?** 不可以。已有无 targets 项目按 simple 维护;迁移前需人类明确授权、映射、结果校验和回退方式。 **可以用自制 checkpoint 或 `SUCCESS` 文件吗?** 不可以。complex 使用 targets 的 metadata、失效和增量重建;不引入第二套缓存/恢复协议。 **轻量测试通过是否等于全量通过?** 不是。必须分别报告 `preflight`、`lightweight_execution` 和 `full_data_execution`;未运行全量写 `NOT_RUN`。 **需要通用 R 函数怎么办?** 由 `bensz-r-developer` 实现,本 Skill 负责分析需求、接入和流程证据。 ## 许可证与贡献 仓库许可证见项目根目录的 `LICENSE`。修改 Skill 时请同步 `SKILL.md`、`config.yaml`、必要 references 和变更记录,并运行与风险匹配的检查;不要在 `raw/` 或 README 中写入凭据、隐私数据或私有提示词。 -
README_EN.md 8.9 KB
# bensz-rmd-rules An Agent Skill for **R data analysis, R Markdown reports, reproducible targets workflows, publication-quality plots, and evidence-anchored interpretation**. The current version is `skill_info.version` in [`config.yaml`](config.yaml); this directory is a beta candidate source. ## When to use it Use it to turn raw data into preprocessing, statistics/models, plots, Rmd/HTML reports, scientific products, and result interpretation, or to maintain an existing R/Rmd project. Do not use it when the primary deliverable is a reusable cross-project R function, class, stable API, or Package (use `bensz-r-developer`); when only an existing Rmd must be rendered (use `knit-rmd-html`); or for analysis in another language. ## Minimal prompt ```text Please use bensz-rmd-rules for this R analysis. First inspect read-only whether the project is new or existing, then explain the choice between simple and targets-first complex/pipeline. Use renv and run a real lightweight synthetic_fixture or project_subset test from the formal entry point before the production run. ``` The Skill preserves existing projects: it does not automatically create targets/renv assets or migrate layouts; a project with `_targets.R` is maintained as complex, while an existing project without targets (including numbered scripts) is maintained with simple semantics. ## Workflow modes | Mode | Use when | Required contract | | --- | --- | --- | | `simple` | Linear, inexpensive work where a full rerun is acceptable | `renv`, an explicit R/Rmd entry point, and a real lightweight test; no targets | | `complex` (descriptive alias `pipeline`) | Non-linear dependencies, expensive steps, reuse, partial invalidation, recovery, or parallelism | `_targets.R` as the only DAG, compute functions in `R/`, Rmd consumes targets, and isolated-store recovery evidence | `new`/`existing` describe project state, not a third mode: an existing `_targets.R` means complex; otherwise simple. Migration requires explicit human authorization. The complex relationship is: ```text renv.lock → _targets.R → R/ functions → target results → Rmd → reports/ ↘ optional products/ ↘ _targets/ (machine state) ``` `_targets/` stores machine computation state only. `products/` stores scientific objects that need review, reuse, or delivery; it is not a second cache, `SUCCESS` marker, identity-hash, or recovery runner. ## Recommended project layout ```text project-root/ ├── 00.Environment.R ├── R/ # compute functions called by targets ├── _targets.R # the only complex DAG entry point ├── raw/ # read-only ├── products/ # optional scientific products ├── reports/ # plots, tables, HTML, supplements ├── scripts/tests/ # versioned test code ├── tmp/tests/<run-id>/ # isolated test run ├── renv.lock └── renv/activate.R ``` The [`_targets.R`](templates/_targets.R) template uses `tar_source("R")`; [`R_data_template.R`](templates/R_data_template.R) provides a data-preparation starting point; [`Rmd_template.Rmd`](templates/Rmd_template.Rmd) consumes `analysis_results` with `targets::tar_read()`; simple projects can start from [`Rmd_simple_template.Rmd`](templates/Rmd_simple_template.Rmd). ## Statistical inference completeness Before a paper-level analysis, list the primary and conclusion-relevant secondary estimands, including the target population, analysis unit, comparison, design, and inferential purpose. Where valid, retain the estimate and unit, effective N/event count, and an appropriate 95% CI by default; follow a preregistered confidence level when it differs. Report a p-value only for a defined hypothesis, and identify the adjustment method plus q-value or adjusted p-value for multiple tests. Reports must cite the computed results and explain practical magnitude and limitations. Mark fixed constants or descriptive quantities `not_applicable`, unsupported inference `not_estimable` with reasons and alternatives, and unexecuted work `not_run`. Do not fabricate CI/p-values or treat keyword checks as proof of statistical validity. See [`statistical_inference_protocol.md`](references/statistical_inference_protocol.md). ## Lightweight testing and recovery Every new or materially changed flow must run one real test style: - `synthetic_fixture`: when real data are unavailable, sensitive, or too large; preserve schema, types, keys, groups, missingness, and an edge case with a fixed seed. - `project_subset`: when authorized data and existing code are available; use a representative subset covering key groups/outcomes/missingness/anomalies, not only `head(n)`. Put test code in `scripts/tests/` and use a unique `tmp/tests/<run-id>/` per run. Simple invokes the formal entry point; complex runs the same DAG with a store under `<run-id>/_targets`, never the formal `_targets/`, `products/`, or `reports/`. Assert input contracts, key types/keys, important numeric invariants, product/report creation, and no writes to `raw/`. Record `full_data_execution=NOT_RUN` when the full dataset was not run. For recovery evidence, first complete an expensive target, interrupt a later target, then rerun `tar_make()` with unchanged input, code, parameters, and renv identity. Use `tar_meta()`, the outdated set, and execution records to show valid upstream targets were skipped. ## Checks and script entry points Inspect project state and mode first: ```bash python3 <skill-root>/scripts/check_targets_renv.py <project> --project-state auto --workflow-mode auto ``` Common checks: ```bash python3 <skill-root>/scripts/check_pipeline_contract.py <project> python3 <skill-root>/scripts/check_interpretation_quality.py <project>/report.Rmd python3 <skill-root>/scripts/check_figure_table_interpretation.py <project>/report.Rmd python3 <skill-root>/scripts/check_htmlwidget_visibility.py <project>/report.Rmd python3 <skill-root>/scripts/check_rmd_template_yaml.py <project>/report.Rmd Rscript <project>/scripts/tests/smoke_test.R ``` Use [`check_plot_readability.R`](scripts/check_plot_readability.R) for plot readability and [`validate_paths.R`](scripts/validate_paths.R) for path safety. Initialize the Liquid Glass theme with [`bootstrap_liquid_glass.py`](scripts/bootstrap_liquid_glass.py); its desktop dynamic TOC expands the interactive area immediately so the pointer can enter the menu. Delegate HTML rendering to `knit-rmd-html`. ## Plot and interpretation rules Plots default to English; switch with YAML `params.plot_language` for Chinese-journal contexts. Place an interpretation near every visible figure/table: current object, direction/comparison, traceable values, magnitude, uncertainty, and follow-up validation (method + input + criterion). See: - [`four_tier_interpretation_framework.md`](references/four_tier_interpretation_framework.md): four-layer interpretation and the Fail Fast Gate. - [`interpretation_templates.md`](references/interpretation_templates.md): univariate, multivariable, and model-validation skeletons. - [`plot_quality_standards.md`](references/plot_quality_standards.md): publication-quality readability rules. - [`liquid_glass_theme_guide.md`](references/liquid_glass_theme_guide.md): HTML theme and troubleshooting. ## Related templates and references - Environment and dependencies: [`00.Environment.R`](templates/00.Environment.R), [`renv/activate.R`](templates/renv/activate.R). - Testing: [`templates/tests/`](templates/tests/), [`lightweight_testing.md`](references/lightweight_testing.md). - Modes and architecture: [`workflow_modes.md`](references/workflow_modes.md), [`hybrid_architecture_guide.md`](references/hybrid_architecture_guide.md). - Delivery and review: [`delivery_verification.md`](references/delivery_verification.md), [`serial_review_protocol.md`](references/serial_review_protocol.md). ## FAQ and boundaries **Can numbered scripts be migrated to targets automatically?** No. Existing projects without targets stay simple; migration needs explicit human authorization, mapping, result checks, and rollback. **Can I add a custom checkpoint or `SUCCESS` file?** No. Complex uses targets metadata, invalidation, and incremental rebuilds; do not introduce a second cache/recovery protocol. **Does a passing lightweight test mean the full dataset passed?** No. Report `preflight`, `lightweight_execution`, and `full_data_execution` separately; write `NOT_RUN` when the full run was skipped. **What if I need a reusable R function?** Have `bensz-r-developer` implement it; this Skill owns analysis requirements, integration, and workflow evidence. ## License and contribution See the repository-root `LICENSE` for licensing terms. When changing this Skill, keep `SKILL.md`, `config.yaml`, required references, and change records aligned, and run checks appropriate to the risk. Never put credentials, private data, or private prompts in `raw/` or the README. -
SKILL.md 14.5 KB
--- name: bensz-rmd-rules description: 当主要交付物是基于 R 的数据分析流程、科学结果、可复现数据产品、R Markdown/HTML 报告或论文级图表与解读时使用;即使流程中包含 `.R` 脚本和仅服务当前分析的辅助函数,也由本 Skill 主导。⚠️ 不适用:主要交付物是可独立复用的 R 函数、稳定公共 API、类或 R Package;这些任务使用 bensz-r-developer。也不用于仅渲染既有 Rmd 或其它语言的数据分析。 metadata: author: Bensz Conan short-description: R/Rmd 多阶段分析、可恢复数据产品与论文级报告规范 keywords: - bensz-rmd-rules - R Markdown - Rmd - 数据分析 - 断点续算 - 可复现报告 --- # bensz-rmd-rules ## 目标 把研究目标和数据组织为可审查、可复现的 R 分析流程。新项目必须用 `renv`,并在真实 R/Rmd 入口完成轻量测试:线性低成本流程用 `simple`;有非线性依赖、昂贵步骤、复用、局部失效、恢复、血缘或并行需求时用 **targets-first complex/pipeline**(`_targets.R` 唯一 DAG,计算函数在 `R/`,Rmd 只消费 target)。`_targets/` 是机器状态,`products/`/`reports/` 是科学产物。已有项目不自动补齐或迁移:有 `_targets.R` 按 complex,无 targets(含历史编号脚本)按 simple;需要复杂能力时由人类显式迁移。本 Skill 不维护编号 runner。 主要交付物决定触发,而不是扩展名: | 交付物 | 主导 Skill | | --- | --- | | 分析数据流、统计结果、可恢复产品、Rmd/HTML、图表与解读 | `bensz-rmd-rules` | | 可复用、版本化或跨项目的函数、类、API、Package | `bensz-r-developer` | | 分析需要可复用组件 | 本 Skill 定义需求/集成证据,`bensz-r-developer` 实现,本 Skill 接回验证 | 可创建只服务当前分析的 `R/` 函数和 helper,但不扩展为公共 Package。`luckyBase` 是唯一固定 Bensz 依赖;其它包由项目写入 `renv.lock`。`crew`、`autometric`、集群插件只是项目级可选依赖;不自制调度器、日志协议、worker 状态机或资源采样器。本 Skill 不自动接入 State、Verifier、Pack 或 Gate。 ## 流程 ### 输入 收集:研究问题、用途、交付物、统计边界和完成判据;主要/次要估计对象、目标人群、观察单位、对比与效应尺度、研究设计、缺失/依赖结构、预注册方案及推断目的;原始输入/数据字典/隐私授权及 `raw/` 只读边界;重算成本、依赖、外部请求、随机性、关键参数和恢复需求;现有 R/Rmd、`00.Environment.R`、产品路径、`tmp/`、`_targets.R`、`renv.lock`/`renv/`;可用测试数据与边缘条件;函数复用范围。缺失信息只有在改变行为或安全边界时才询问。 先读 `config.yaml`,再按任务读取最少 references: - 模式/目录:[`workflow_modes.md`](references/workflow_modes.md)、[`hybrid_architecture_guide.md`](references/hybrid_architecture_guide.md);测试:[`lightweight_testing.md`](references/lightweight_testing.md)。 - 缓存/示例:[`analysis_workflow_cache.md`](references/analysis_workflow_cache.md)、[`hybrid_architecture_examples.md`](references/hybrid_architecture_examples.md);审查:[`serial_review_protocol.md`](references/serial_review_protocol.md)。 - R 实现:[`code_style_guide.md`](references/code_style_guide.md)、[`no_overdefensive_code.md`](references/no_overdefensive_code.md)、[`cross_platform.md`](references/cross_platform.md)。 - 指标/解读:[`metric_explanation_protocol.md`](references/metric_explanation_protocol.md)、[`four_tier_interpretation_framework.md`](references/four_tier_interpretation_framework.md)、[`interpretation_templates.md`](references/interpretation_templates.md)、[`interpretation_narrative_examples.md`](references/interpretation_narrative_examples.md)、[`expert_discussion_template.md`](references/expert_discussion_template.md)。 - 论文级估计与检验:[`statistical_inference_protocol.md`](references/statistical_inference_protocol.md);计算前逐项确定可估计性与方法,报告时逐项核对。 - 图表/HTML:[`plot_quality_standards.md`](references/plot_quality_standards.md)、[`plot_language.md`](references/plot_language.md)、[`htmlwidget_visibility_rules.md`](references/htmlwidget_visibility_rules.md)、[`liquid_glass_theme_guide.md`](references/liquid_glass_theme_guide.md);ID/函数:[`gene_id_guidelines.md`](references/gene_id_guidelines.md)、[`candidate_r.md`](references/candidate_r.md)。 对 `raw/` 仅做规划和选样所需的轻量只读盘点,不启动完整计算。 ### 执行步骤 1. **映射交付与推断**:若只要公共函数/API/类/Package,转交 `bensz-r-developer`;否则建立 target—输入—科学产品—报告—完成判据映射,确保依赖无环。计算前列出影响论文结论的主要与次要估计对象,逐项判断是否需要且能够给出区间/检验,记录设计、方法理由和结果去向;预注册方案优先,不为显著性事后更换主终点或检验。纯描述、固定常数或不可估计项明确标状态与原因,不伪造 CI/p。新 pipeline 的计算函数放 `R/`,由 `_targets.R` 的 `tar_source("R")` 发现;`00.Environment.R` 默认唯一,只负责包、项目根、产品路径和全局配置。 2. **判状态和模式**:先运行 ```bash python3 <skill-root>/scripts/check_targets_renv.py <项目根> --project-state auto --workflow-mode auto ``` 优先级为人类要求 → existing 边界 → 复杂度。新项目默认 simple;复杂信号(非线性、昂贵/高失败代价、多下游复用、局部失效、恢复、血缘、并行)选 complex。`existing` 不是第三种模式:有 `_targets.R` 按 complex,否则按 simple;不创建缺失 targets/renv,不隐式迁移。详见 [`workflow_modes.md`](references/workflow_modes.md)。 3. **固定路径**:按需使用 `00.Environment.R`、`R/`、`_targets.R`、`raw/`、`products/`、`reports/`、`templates/`、`scripts/tests/`、`tmp/tests/<run-id>/`、`tmp/scratch/`、`renv.lock` 和 `renv/activate.R`,不预建空目录。`products/` 只存需审阅/复用/交付的科学产物;路径优先级为人类指定 → 已有路径 → 项目统一设置 → `products/`,必须在项目内且不得与 `raw/`、`reports/`、`tmp/`、`_targets/`、`.bensz-api/` 重叠。 4. **设计缓存边界**:只导出昂贵、可复用、高风险/有损、需人工审查、依赖可变外部请求或中断代价高的结果。targets 管理 `_targets/` metadata、失效和增量重建;禁止 `SUCCESS`、identity hash、force-step/resume runner 或 checkpoint helper。详见 [`analysis_workflow_cache.md`](references/analysis_workflow_cache.md)。 5. **计算推断并实现报告**:`00.Environment.R` 用 `luckyBase::Plus.library()`;分析代码优先 `pkg::fn()`,基因 ID 用 `luckyBase::convert()`。分析层对关键参数计算估计值、适当的默认 95% CI(方案另有水平时从其规定),并仅在有明确检验问题时给 p 值;多重检验记录校正方法与 q/调整后 p。保留未经显著性筛选的完整结果,含参数身份、尺度/单位、有效 N/事件数、方法、CI 界限/水平、适用的 p/q、来源及不适用/不可估计/未运行状态与理由。complex 的 `R/` 函数读取 `raw/`/上游 target;Rmd 用 `tar_read()`/`tar_load()` 消费真实结果,不在正文臆算。Top N/阈值/配色放 YAML `params`。target 语义化命名并按阶段注释;交付摘要附 `targets::tar_manifest()`。静态图写 `reports/`,HTML 留根目录;默认不在图内重复标题。只在 I/O 边界和硬前提检查,不用占位/静默降级掩盖失败。 6. **真实轻量测试**:新建或实质修改流程选择 `synthetic_fixture`(保 schema、类型、主键、分组、缺失和边缘条件,固定种子)或有授权时的 `project_subset`(覆盖关键分组/结局/缺失/异常,不能仅 `head(n)`)。测试在 `scripts/tests/`,每次使用唯一 `tmp/tests/<run-id>/`;只改输入规模和路径,不加业务 `analysis_mode`/`test_mode`。simple 调正式入口;complex 执行同一 DAG,store 放 `<run-id>/_targets`。至少断言输入契约、类型/主键、数值不变量、产品/报告生成和 `raw/` 未写入。checker/语法/`--dry-run` 仅为 `preflight`;真实链为 `lightweight_execution`;未跑全量记 `full_data_execution=NOT_RUN`。按环境依赖、fixture/子集、路径隔离、分析代码、科学断言、渲染、外部服务分类失败。 7. **串行审查**:初稿和轻量测试后,按 [`serial_review_protocol.md`](references/serial_review_protocol.md) 完成结构/数据流、R/恢复、科学/统计三项只读审查;每项读取前项修正后的最新版本。影响执行、输出或断言的修正会使旧证据失效,须在最新 identity 上重跑受影响链;不能提供独立 Agent 时须记录降级。 8. **正式运行与报告**:simple 用明确 R/Rmd 入口;complex 先检查契约,再 `targets::tar_make()`,store 为 `_targets/`。中断后用 `tar_meta()`、outdated 集合和执行记录证明有效 target 被跳过,不手工写 `SUCCESS`。图表默认英文,中文场景用 `params.plot_language`;静态图保存 PDF,并在任务工作区生成 JPG 预览。每个关键结论给出与计算产品对应的估计值、CI、有明确假设时的 p/q、实际意义与局限;无法推断时写清状态和理由。图/表附近给对象、方向/对比与可追溯证据,不为装饰性数字逐个检验。 ### 输出 交付需求—target—报告映射、唯一 `00.Environment.R`、`R/` 函数、Rmd/HTML、`renv.lock`、`renv/activate.R`、`scripts/tests/`、`reports/` 和验证摘要。complex 另交 `_targets.R` 与 `tar_manifest()` 快照,按需交分析计划和科学产品;simple 不创建 targets。测试记录绑定最新 identity,分别报告三层执行状态,不能以轻量通过暗示全量通过。 ### 输出管理 正式代码、测试、Rmd/HTML、产品和报告留在用户项目约定位置;草稿、审查、JPG 和日志进入当前 `.bensz-api/task-*`;测试现场进入 `tmp/tests/<run-id>/`。不写 `raw/`;`products/` 不放草稿;`reports/` 不放源码、缓存或审查日志;`tmp/scratch/` 的重要发现须晋升并补 renv/轻量测试证据。 ### 校验 ```bash # 新 simple python3 <skill-root>/scripts/check_targets_renv.py <项目根> --project-state new --workflow-mode simple Rscript -e 'renv::status()' Rscript <项目根>/scripts/tests/smoke_test.R # 新 complex python3 <skill-root>/scripts/check_targets_renv.py <项目根> --project-state new --workflow-mode complex Rscript -e 'renv::status()' Rscript <项目根>/scripts/tests/smoke_test.R Rscript -e 'targets::tar_make()' Rscript -e 'targets::tar_meta()' # 已有项目 python3 <skill-root>/scripts/check_targets_renv.py <项目根> --project-state existing --workflow-mode auto ``` 所有 Rmd 还要检查解读覆盖/质量、widget 可见性、图表可读性和数字追溯,并按实际情况做 R 语法/运行、Rmd 渲染;逐项复核主要估计对象的结果或不可估计理由、CI/p 方法与设计的一致性。现有文本检查器只作启发式预检,词语命中不证明推断已计算或方法正确。检查清单见 [`workflow_checklist.md`](references/workflow_checklist.md),证据格式见 [`delivery_verification.md`](references/delivery_verification.md)。最低验收:模式理由可复述、renv 可信、测试未污染正式路径、Rmd 未绕过 DAG、恢复可证明、已有项目未隐式迁移、跳过项不写成通过。 ### 失败与恢复 - 缺 `luckyBase` 或新项目无 renv:停止补齐前提,不更换包管理逻辑。 - simple 出现复杂信号:评估升级 complex,不扩展调度器或无理由降级。 - 测试失败:按失败类别最小修正并重跑;权限/网络/授权/依赖不足时记录阻塞,不伪造成功。 - complex 测试写正式 `_targets/`、products、reports 或 raw:停止、隔离并确认正式路径后重跑。 - targets 失效/失败按 outdated/error 重算必要节点;不伪造 `SUCCESS`。Rmd 渲染失败只从报告层恢复。 - existing 冲突不覆盖/迁移;沿现有 targets 或顺序入口做授权增量修改,需复杂能力时显式迁移。 ## 约束 <!-- BEGIN COMMON CONSTRAINTS --> <!-- Source-Hash: sha256:15120201e9e0c7569517261d57ecefb63ac279c26ed13876f8e95b6dc35854d3 --> <!-- Template-ID: skill-common-constraints; Template-Version: 1; Sync-Policy: exact-block --> ### 公共硬约束 本块由 `docs/templates/skill-common-constraints.md` 统一维护;每个 `SKILL.md` 的 `## 约束` 必须逐字同步本块,不得在副本中改写公共规则。 - 任务需要落盘时,使用唯一的 `./.bensz-api/task-{yyyymmdd-hhmm}-{简短描述}/` 根目录;共享材料放入 `shared/`,Skill 专属材料放入该 Skill 的 `input/`、`output/`、`log/`。 - 正式交付物、源代码和正式计划按项目约定保存,不写入任务工作区;未经授权不覆盖、删除、迁移或远程写入。 - 项目维护变更检查 BAC 可用性并记录需求、AI 产出、工具结果、文件改动和验证摘要;BAC 只做过程审计,不替代署名、责任或合规判断。 - 不记录 API Key、访问令牌、密码、Cookie、环境/凭据文件、私有 Prompt、身份信息、本地用户名、主机名或不必要的大体积原始数据。 - 文件路径必须规范化并限制在授权项目范围内;外部 URL、子进程和网络访问遵循最小权限,防止路径遍历、SSRF 和命令注入。 - Skill 版本唯一记录在自身 `config.yaml:skill_info.version`;公开 API、协议、目录或配置变更同步文档与 `CHANGELOG.md`。 - `bensz-collect-bugs` 是一个 Agent Skill;仅将 Bensz Agent Skill 或 Bensz 基础设施本身的设计缺陷交给它。先脱敏写入 `~/.bensz-skills/bugs/`,当前任务不中断,只有用户明确要求才公开上报,禁止直接修改用户已安装的 Skill 源码。 <!-- End of canonical common constraints. --> <!-- END COMMON CONSTRAINTS --> ### Skill 专属约束 - 只在授权范围内只读盘点 `raw/`;不得修改、覆盖或向其写入,也不在日志中记录绝对私有路径、凭据或不必要原始值。 - 不因模板存在就创建全部文件;模式、缓存边界和目录由真实需求决定。 - 新 pipeline 不引入 checkpoint helper、SUCCESS 或自定义 identity runner;该自制缓存机制已移除,不得在任何项目中新引入或修复。
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.