Claude Skill

3d-deep-research

用证据链和 X/Y/Z 立体分析法研究产品、公司、技术、概念、人物、行业、市场或复杂事件,交付可追溯的深度研究报告。用户要求 deep research、系统调研、竞品或市场研究、尽职调查、来龙去脉分析、证据链或正式研究报告时使用。简单名词解释、新闻摘要、短篇观点、仿写,以及 3D 建模、渲染、CAD 或图形设计不使用。

LLM Mart · 0 points · 0 views 18 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download jeffy-peng-jeffy-skills-3d-deep-research-c29fd9b.zip · 46 KB
Part of jeffy-peng/jeffy-skills — 2 skills

Install

skills CLI npx skills add https://github.com/jeffy-Peng/jeffy-skills/tree/main/3d-deep-research
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install jeffy-peng-jeffy-skills@llmmart
Git git clone https://github.com/jeffy-Peng/jeffy-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole jeffy-peng/jeffy-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

3D Deep Research

先建立研究底图,确认研究对象的边界、运转方式和当前状态;再用 X/Y/Z 逐层解释:X 识别显性与非显性的关键变化及其后续影响,Y 沿具体变化切开并比较成因,Z 拆开尚不清楚的作用连接。分析结果持续返回发展路径和研究底图,用证据校正解释并检查重要遗漏。

本 Skill 为主研究流程时,默认交付完整的 Markdown、HTML 和 PDF 三份报告;用户明确缩小交付范围时按其要求减少。用户点名其他研究 Skill 或选择其他工作模式时,以其作为主流程;仅在需要本方法时辅助使用,不叠加独立默认产物。不要仅因通用的 deep research 字样同时运行两套流程,也不要把普通问答扩写成长报告。

必读资源

执行完整研究时:

  1. 读取 references/evidence-protocol.md,建立来源与 Claim 账本并执行事实审计。
  2. 读取 references/xyz-method.md,建立研究底图并执行 X/Y/Z、路径校正和覆盖复核。
  3. 根据研究对象读取 references/object-adapters.md 的对应部分。
  4. 读取 references/visual-guidelines.md,用于分析阶段的视觉规划、HTML/PDF 渲染和图表检查。
  5. 写作时复制 assets/report-template.md,保留元数据与证据附录;正文可按问题重组。

报告措辞

以下规则适用于正文、标题和图注;证据附录以准确、可追溯为先。

  • 先直接回答研究问题,再解释事实、原因、限制和行动含义。
  • 先具体,后抽象。 优先说清谁做了什么、发生什么变化,再解释背后的概念。能直接说明的,不用抽象名词代替。
  • 先判断,后解释。 一句话先说清主要意思,原因和必要限制随后展开,避免把多层意思挤进一句话。
  • 用自然、准确的书面语解释给非专业读者听。 写完后检查:读者是否需要把这句话“翻译成白话”才能理解?如果需要,就改写,同时保留因果关系与证据边界。
  • 每段只推进一个判断。删除不提供事实、解释或决策价值的句子。
  • 明确区分事实、解释和预测。证据只能支持“可能”时,不写成“证明”或“必然”。
  • 不把 Claim 类型、证据门槛、X/Y/Z 等内部方法术语写进正文;它们只用于组织研究和附录说明。
  • 数字首次出现时说明时间、单位、统计口径和比较对象。
  • 不使用宣传式形容词、无依据的最高级,以及“值得注意的是”“不容忽视”等空洞过渡语。
  • 只在不确定性会影响结论时说明限制,并紧邻相关判断书写。

执行流程

0. 明确研究设定

记录研究对象、对象类型、用户要做的决策、特别关注点、时间基准、范围边界和交付要求。无法从上下文解决且会显著改变目标或正确性的问题才追问;其余情况直接开始。根据对象适配器建立研究底图,先确认主体边界、价值或作用结构、关键参与者及当前状态,再选择解释主线。

复杂或持续时间较长的任务可以把上述设定记录为工作笔记;不要为简单任务强制创建独立过程文件。

研究范围由用户问题和证据决定,不追求来源数、Claim 数、字数或图表数量。每个进入 A2 的关键判断和数字都必须通过适用的证据门槛并完成审计;资料不足时缩小可确认结论的范围,保留未解决问题,不静默改变用户的研究目标。

涉及“最新、现在、最近”时联网核实,并记录发布日期和访问日期。输出到用户指定位置;未指定时使用当前项目的 output/。

1. 规划检索

先列出:

  1. 建立研究底图必须确认的事实和口径;
  2. 关键变化及其前后状态;
  3. 需要比较的成因和需要打开的作用连接;
  4. 需要主动寻找的反向证据、替代解释和失败案例。

检索过程中根据证据调整问题。每个影响结论的问题最终都要有来源支持,或在 A2/A3 明确标为未解决。复杂任务可以维护检索笔记,但不强制生成独立表格。来源选择遵循 evidence protocol,优先使用适合该对象的原始材料、官方记录和独立来源。

2. 维护证据账本

在 report.md 附录 A1 和 A2 分别维护来源与 Claim 账本。来源使用稳定 ID(S01、S02),关键判断使用 Claim ID(C01、C02)。把“来源出处”和“证据作用”分开记录。

研究过程中持续更新账本。每条关键判断的支持证据、替代解释、反向材料、独立性、置信度、资料缺口和修订条件,按 references/evidence-protocol.md 记录和判断;因果和机制判断分别记录过程证据与归因边界。过程已发生不等于它足以解释总体结果。

证据不足时交付“已确认部分 + 资料缺口 + 下一步验证路径”。因果或机制证据不完整时降级表述,不补写猜测,也不因局部缺口停止整份交付。

3. 建立底图并执行 X/Y/Z 分析

先用研究底图横向看清对象,避免只研究容易形成主线的材料。再建立暂定的关键变化路径,把可观察变化、状态改变和后续影响分开;围绕影响核心判断的节点,比较成因并拆解尚不清楚的作用过程。根据证据返回修正节点、原因和路径,不要求一次完成 X 后再单向完成 Y、Z。

复杂研究可在已有工作笔记中维护节点、待解释问题、因素、作用连接和证据状态的对应关系,不要求生成独立过程文件,也不把内部编号写入正文。X 只呈现原因摘要,Y 承担原因比较,Z 必须补出中间环节、成立条件或可观察后果,避免三部分重复叙述。详细方法和未来表达见 references/xyz-method.md。

正文中的重要判断使用 [S01] 形式引用来源。最终解释需要说明哪些因素通过什么作用过程造成关键变化、变化如何影响后续发展,以及判断最可能错在哪个证据缺口或前提。深度由解释缺口决定,不按固定节点数或拆解层数执行。

在分析阶段识别支撑核心结论的关键比较、演化路径、主体关系和机制,并规划相应的视觉表达,随证据更新调整。不要等正文完成后才考虑配图;可在已有研究笔记中简记,无需新增独立账本。

4. 写作与视觉表达

使用 assets/report-template.md 的元数据与必要附录。正文先回答问题并建立对象底图,再按问题组织解释。模板的演化式目录是默认示例;横向比较或稳定机制研究可以围绕具体问题连续展开观察、原因和机制,不必补写无关历史。重组目录不免除证据审计、解释追踪和双重回返;删除空章节及没有解释增量的重复。

关键关系用图能显著减少读者对照、记忆或推演负担时,应制作相应分析图;数量由需要解释的关系决定。不能以正文已有描述或缺少量化数据为由省略有价值的图。图、表与文字各取所长,避免重复和装饰。量化图必须有可靠且可比较的数据;具体触发条件、证据要求和图表契约见 references/visual-guidelines.md。

5. 执行事实审计

结构校验不能证明事实。验证前按 evidence protocol 执行归属审计和数字复核:

  1. 回到来源原文,确认每个关键判断能够由所引材料推出;
  2. 复算正文数字,检查单位、币种、时间和统计口径;
  3. 确认反向材料与正文中的限制一致,没有被删除或弱化;
  4. 检查 X 中的可观察变化没有被候选解释替代,Y 中的因素都指向明确节点,Z 中的机制没有停在术语或同义扩写;
  5. 返回研究底图,检查商业模式、经济性、竞争、治理、制度责任或其他对象适配器提示的重要侧面是否因主线过强而被静默删除;
  6. 把审计结果反映到正文、A2 和 A3。

审计失败时改写正文,或降级、删除 Claim;不修改账本去迁就结论。

6. 验证与交付

运行唯一的报告校验入口:

python [skill目录]/scripts/validate_report.py report.md

validate_report.py 只检查机器可验证的结构和证据一致性,不证明外部事实真实,也不判断是否遗漏了有价值的图。交付前按 visual guidelines 复核核心关系的视觉表达;全文无图时必须执行零图复核。

校验通过后默认生成 HTML 和 PDF,并校验 PDF:

python [skill目录]/scripts/render_report.py report.md output.html
python [skill目录]/scripts/render_report.py report.md output.pdf
python [skill目录]/scripts/validate_report.py report.md --html output.html --pdf output.pdf

渲染前在 report.md 同级的 fonts/ 放置 NotoSansCJKsc-Regular.otf 和 NotoSansCJKsc-Bold.otf,并安装 WeasyPrint 69.0。渲染器会检查并嵌入这两种字体。

新报告使用模板中的 证据契约:3。校验历史报告时显式加 --legacy-schema,其通过只表示旧契约通过,不证明现行证据要求或产物一致性已满足。新交付不得用兼容模式绕过失败。渲染器为每个产物生成同名 .manifest.json,核对 Markdown、产物与本地资源摘要;保留这些构建凭据以防误发旧版本,正文或资源变更后重新渲染。摘要检查用于版本一致性,不证明事实或版面正确。

PDF 只通过本 Skill 的 render_report.py 生成,不另写 ReportLab 或其他 Markdown-to-PDF 实现,也不绕过 assets/report.css。渲染依赖或字体不可用时说明阻塞,不切换到其他引擎或排版实现降级交付。完成可行且已授权的依赖修复;PDF 仍阻塞时,提供已完成并通过适用检查的 Markdown,以及能够合规生成的 HTML,明确 PDF 尚未完成,不把部分交付称为完整交付。

标题取报告第一行 H1。渲染器会把 [Sxx] 引用转为可点击锚点。交付前按 visual guidelines 检查输出,并通读正文。本 Skill 为主流程时,完整交付包含 report.md、output.html 和 output.pdf;用户明确缩小范围时按其要求,PDF 阻塞时按上一段提供部分成果并标明未完成项。

质量红线

  • 不把新闻排序当作因果链,也不把力量分类表当作 Y 轴。
  • 不把待验证的解释当作节点事实,不用心理揣测代替机制。
  • 不把合理机制直接当作本案例中已证实的原因,不用术语或同义扩写冒充机制拆解。
  • 不让单一解释主线挤掉会改变结论的重要业务、经济性、竞争、治理或制度侧面。
  • 不用同一原始材料的转载数量冒充独立证据。
  • 不隐藏冲突、样本偏差、访问失败、资料缺口或反向材料。
  • 不生成没有数据口径的数字图,不交付未经复核的数字。
  • 不留下模板占位符、未渲染图表或无法追溯的来源。
  • 不把过期证据当作当前状态。
Files (jeffy-skills)
  • agents
    • openai.yaml 275 B
      interface:
        display_name: "3D Deep Research"
        short_description: "用证据链和 X/Y/Z 分析复杂问题,输出可追溯的深度研究报告"
        default_prompt: "使用 $3d-deep-research 研究[对象],回答[核心问题],并输出可追溯的 Markdown 报告。"
      
  • assets
    • report-template.md 3.4 KB
      # [研究对象]深度研究报告
      
      > 研究问题:[一句话问题] | 资料截止:[YYYY-MM-DD] | 完成日期:[YYYY-MM-DD]
      
      > 证据契约:3
      
      <!-- 以下正文目录适合演化研究,可按问题重组;保留核心回答、对象底图、解释与证据边界,删除空章节。 -->
      <!-- 正文遵循 SKILL.md 的“报告措辞”:先说清具体事实与主要判断,再解释机制和必要限制。 -->
      
      ## 一、核心结论
      
      ### 一句话判断
      
      [直接回答研究问题。在支撑判断的关键事实后标注 Source ID,例如 [S01]。]
      
      ### 对象与当前状态
      
      [用最少但足够的内容建立研究底图:说明对象边界、组成部分、价值或作用结构、关键参与者、当前状态,以及必须分开的统计或概念口径。只保留会影响核心判断的内容。]
      
      ### 关键发现
      
      [写清最重要的发现、成立原因和实质不确定性。说明哪些过程得到观察、证据支持到什么结果;不要把内部字段标签机械写进正文。]
      
      ## 二、哪些变化塑造了今天
      
      ### [自然语言节点标题]
      
      [区分可观察变化、由证据支持的状态改变和对后续路径的影响。只给原因摘要,完整因素比较留在第三章;按研究需要复制或删除本小节,不按年份罗列事件。]
      
      ## 三、为什么发生这些变化
      
      ### [针对具体变化的原因问题]
      
      [说明要解释的是哪个变化的发生、时点、方向、规模或形式;比较实际可能发挥作用的因素、相互关系、替代解释和证据边界。避免重述第二章。按研究需要复制或删除本小节。]
      
      ## 四、这些因素如何产生结果
      
      ### [自然语言机制标题]
      
      [说明本节承接哪个关键变化、当前状态或成因解释;补出尚未说明的中间环节、成立条件、可观察后果和证据状态。既可以解释历史节点,也可以解释当前系统如何产生价值、成本、风险或其他核心结果。只保留具有解释增量的机制,按研究需要复制或删除本小节。]
      
      ## 五、对当前问题意味着什么
      
      [仅在研究涉及决策或未来变化时保留。说明行动含义、需要观察的信号,以及什么变化会推翻或下调判断;否则删除整章。]
      
      ## 附录:来源与证据边界
      
      ### A1 来源账本
      
      | Source ID | 来源与日期 | 证据作用 | 限制 |
      |---|---|---|---|
      | S01 | [标题](URL);发布者:[名称];发布:[YYYY-MM-DD];访问:[YYYY-MM-DD] | [支持或质疑什么判断;是否独立;关键依据页码/表名/章节] | [访问、样本或统计口径限制] |
      
      ### A2 关键判断与证据
      
      | Claim ID | 可核查判断 | 类型/证据范围 | 支持证据与独立性 | 替代解释/反向证据 | 置信度、缺口与修订条件 |
      |---|---|---|---|---|---|
      | C01 | [对象、结果、范围明确的判断] | [fact / causal / mechanism / market / forecast;因果和机制另填过程:具体连接的证据;归因:支持范围与未区分解释] | [支持 S01;independent / shared-origin / mixed / unknown;必要的来源定位] | [反向 S02;有争议时说明区分证据,否则记录反向核查方向] | 置信度:[high / medium / low];缺口:[具体缺失或无重大已知缺口];修订条件:[改变本条主张的证据] |
      
      ### A3 资料边界
      
      [说明缺失数据、冲突口径、样本偏差、访问失败,以及哪些重要侧面因证据不足无法纳入。确认研究底图中的关键方面没有被主解释线静默删除。]
      
    • report.css 4.5 KB · in bundle
  • references
    • evidence-protocol.md 6.4 KB
      # 证据协议
      
      ## 核心原则
      
      把来源身份、证据作用和判断置信度分开记录。只有影响核心结论的关键判断需要 Claim ID;普通事实在正文引用 Source ID,不进入 Claim 账本。
      
      ## 来源记录
      
      在附录 A1 为每个来源分配稳定 ID(`S01`、`S02`),并记录:
      
      - 原始标题、发布者和具体链接或文件;
      - 发布日期;涉及当前状态时同时记录访问日期;
      - 它支持或质疑什么判断;
      - 是否属于原始材料,是否独立于其他来源;
      - 利益关系、样本、统计口径、访问或时间限制。
      
      为便于区分日期角色,A1 来源栏使用 `发布者:名称;发布:YYYY-MM-DD;访问:YYYY-MM-DD`。发布日期无法确定写 `发布:未知`,确实无需访问日期的离线固定记录写 `访问:不适用`,不编造日期。统计/观察期间与发布日不同且影响判断时另外说明。关键判断引用长材料时记录页码、表名、章节或字段;派生数字保留公式与输入口径,便于再次定位和复算。
      
      原始材料不等于独立评价。公司公告可以证明公司发布了什么,却不能独立证明其影响。同一新闻稿的转载属于同一个来源。搜索摘要、聚合页和无法核实出处的内容只能帮助发现线索,不能单独支撑关键判断。
      
      社区帖子、用户评论和单个 issue 可以说明某种现象存在,但不能直接代表总体。使用时必须说明样本边界。
      
      ## 关键判断记录
      
      在附录 A2 为每个影响核心结论的判断分配稳定 ID(`C01`、`C02`),并记录:
      
      - 一句话、可核查或可证伪的判断,明确对象、结果、范围及必要时间边界;不同部分需不同证据强度时才拆分;
      - 类型:`fact` / `causal` / `mechanism` / `market` / `forecast`;
      - 因果和机制判断分别填写 `过程:关键连接的已观察/理论支持/未知情况` 与 `归因:支持的结果范围及未区分的解释`;
      - 支持来源及其独立性;
      - 可信的替代解释、反向证据或未解决的反向检索方向;
      - 置信度、尚缺少的证据和反证条件。
      
      事实、市场和预测判断不强行填写过程与归因。过程已发生和贡献未知可以同时成立;不能用一个状态替整条链或不同结果背书。A2 最后一栏使用 `置信度:high/medium/low;缺口:具体缺失或无重大已知缺口;修订条件:会改变此主张的具体证据`,三个标签分别填写,不用置信度代替缺口。
      
      如果问题无法解决,在 A2 或 A3 明确说明,不为填满结构而补写结论。
      
      ## 证据门槛
      
      | 判断类型 | 最低要求 |
      |---|---|
      | `fact` | 可核验的原始记录;没有原始记录时,使用相互独立的可靠来源 |
      | `causal` | 事实节点、行为证据、作用过程和能够比较主要替代解释的材料;不能排除时降低状态与结论强度 |
      | `mechanism` | 说明中间环节和成立条件;案例运行须有相关行为或过程证据,理论机制须有假设、推导或可重复实验支持;存在实质竞争解释时比较区分证据,不强造反命题 |
      | `market` | 根据主张使用可观察指标;涉及需求或采用时检查用户/客户样本及代表性;公布规模可由相关原始统计支持,不从个别样本外推总体 |
      | `forecast` | 关键变量、基准路径、触发信号、替代解释和反证条件 |
      
      证据未达到对应门槛时,不把判断写成确定性结论。将其改为暂定解释、降低结论强度,或列入资料缺口。资料不足时缩小可确认结论的范围,保留用户问题与未解决项,不静默缩换研究目标。
      
      ## 置信度
      
      置信度评价的是判断,不是来源本身:
      
      | 等级 | 判断标准 |
      |---|---|
      | `high` | 证据直接、可核验且覆盖主张,没有足以改变结论的未解释反向证据;单一权威原始记录可支持“记录内容是什么”,记录外的影响与归因须另有适用证据 |
      | `medium` | 证据可靠但仍有缺口、独立性不足,或存在尚不能完全排除的替代解释 |
      | `low` | 主要依据线索、同源材料、过时信息或未解决推断,只能作为暂定解释或背景 |
      
      不要根据来源数量机械定级。给不出判断依据时,降低结论强度或删除判断。
      
      独立性取决于证据来源关系,不等于没有利益关系,也不是所有事实都需要两个来源。多篇独立评论不能替代因果识别所需的证据。
      
      涉及“当前、最新、最近”的判断必须联网核实,并在正文说明资料截止时间。历史事实不设置任意失效日期;重新使用或更新报告时,再核实其中时间敏感的判断。
      
      ## 引用与反向证据
      
      正文关键事实和判断使用 `[S01]` 形式引用。引用紧邻它支持的句子;一个段落包含多个不同事实时,不把所有 Source ID 统一堆在段尾。
      
      A1 必须让每个正文 Source ID 都能解析到来源、发布者、日期和链接。A2 使用 Claim ID 和 Source ID,把支持证据与替代解释分开记录,不写“官方资料”“媒体报道”等无法追溯的描述。
      
      对每个关键判断主动寻找最有力的反向材料或替代解释,例如失败与回滚、竞争路径、监管争议、时间线冲突、口径差异和样本偏差。没有找到时,在 A2 记录检索方向;“未发现”不等于“不存在”。
      
      对影响结论的解释分歧,说明不同解释各自预期的观察、共同可解释的结果和真正能区分它们的证据。共同作用或中间环节不自动构成替代解释。某个结果可触发经营判断更新,却未必能反驳机制判断;修订条件须对应本条主张。
      
      ## 交付前核查
      
      结构校验只能检查编号和格式,不能证明结论来自所引来源。交付前必须:
      
      1. 回到来源原文,确认每个关键句都能由所引材料推出;
      2. 复算正文数字,检查单位、币种、时间范围和统计口径;
      3. 确认反向材料和限制没有在正文中被删除或弱化;
      4. 确认过程证据、归因边界与正文措辞一致,没有把理论可能或局部运行写成总体结果的主要原因;
      5. 覆盖全部关键判断和数字,不抽样。
      
      核查失败时修改或删除结论,不修改证据记录去迁就结论。核查结果反映在正文、A2 和 A3;不强制独立摘录存档或审计日志,可在已有笔记保留必要的来源定位与计算依据。
      
    • object-adapters.md 4 KB
      # 对象适配
      
      根据研究问题选择最相关的对象类型。对象跨类型时可以组合使用,不强行归入单一类别。
      
      下表用于建立研究底图、发现候选节点和检查常见遗漏,不预先决定原因,也不要求逐项写成报告章节。所有对象共用“研究底图—关键变化—成因解释—作用过程—双重回返”的方法。
      
      | 类型 | 研究底图要看什么 | 可以从哪些变化寻找节点 | Y/Z 常见解释重点 | 常见遗漏 |
      |---|---|---|---|---|
      | 产品 | 用户任务、价值主张、功能与服务边界、定价、分发和采用 | 定位、采用、留存、定价、体验、渠道依赖和关键产品决策 | 用户摩擦、替代选择、分发条件、团队与平台约束如何影响采用和留存 | 把发布当采用;把功能当价值;忽略渠道、迁移成本和失败用户 |
      | 公司 | 业务组合、客户与供应方、价值链、收入成本利润、资本治理、制度责任 | 战略转向、经营拐点、能力积累或衰退、组织与依赖关系变化 | 商业模式、单位经济性、资源配置、组织能力、竞争、资本和制度如何传导 | 只讲战略不讲现金;只看公司不看交易对手;只讲增长不讲价值分配和责任 |
      | 技术/概念 | 定义与边界、能力条件、实现路径、成本、可靠性、采用环境 | 能力跃迁、瓶颈变化、路线分化、标准、采用转折与争议 | 理论条件、工程实现、基础设施、成本曲线、生态和使用环境如何作用 | 把演示当生产能力;忽略可靠性、集成、维护、失败条件和替代路线 |
      | 人物 | 经历、公开行为、组织位置、关系、资源和权力边界 | 关键选择、公开立场、行为边界、权力位置和关系变化 | 经历与环境如何改变选择空间,公开行为如何通过组织和关系产生结果 | 心理揣测;用事后结果倒推动机;忽略组织、时代与他人行动 |
      | 事件 | 参与者、时间边界、当时可知信息、传播与制度环境 | 前因、触发、扩散、响应、结果以及重要的未发生变化 | 信息、激励、协调、制度响应和反馈如何改变事件方向与规模 | 后见之明;单一归因;把同期变化当因果;忽略未响应和反事实边界 |
      | 行业/生态 | 边界、价值链、利润池、供需、地域口径、规则和主要参与者 | 技术与商业阶段、周期、规则、价值分配和依赖关系变化 | 多主体竞争、上下游关系、规模经济、替代路径、政策和宏观条件如何传导 | 混用统计口径;只看头部企业;忽略利润池迁移、非消费方和区域差异 |
      
      ## 公司研究的覆盖复核
      
      用户提出宽泛的“深度拆解一家公司”时,完成主因果链后检查以下方面是否会改变结论:
      
      1. **业务结构**:卖什么、服务谁,各业务如何关联,哪些口径不能直接比较;
      2. **价值与经济性**:收入如何产生,成本、利润、现金和资本投入在哪里形成;
      3. **供需与竞争**:客户、供应方、渠道、替代品和合作方如何影响价值分配;
      4. **组织与治理**:决策权、组织能力、激励和资本配置怎样改变选择;
      5. **制度与责任**:监管、安全、数据及社会责任如何进入经营过程;
      6. **扩张与未来选项**:新地域、新业务或新技术需要补齐哪些能力和约束。
      
      这六项是遗漏检查,不是固定目录。与研究问题无关的内容可以省略;证据不足但可能改变判断的内容写入 A3,不能静默删除。
      
      ## 边界提醒
      
      - 人物研究避免心理揣测,优先使用公开行为、利益结构、组织位置和同期约束。
      - 事件研究区分当时可知信息与事后信息,避免后见之明。
      - 技术研究优先原始论文、标准和官方实现;预印本不自动高于同行评审版本。
      - 行业研究先定义边界、价值链和地域范围,不直接比较不同口径的数据。
      - 公司研究区分业务规模、会计收入、利润和现金,不用单一指标替代完整经济性。
      
    • visual-guidelines.md 4.1 KB
      # 视觉表达规范
      
      ## 何时使用图表
      
      在分析阶段主动识别需要视觉表达的核心关系,随证据更新决定图形与位置。每张图必须回答一个真实的分析问题,并且比正文或表格更容易理解。不要为了装饰、填充版面或展示方法框架而配图。
      
      当核心判断依赖一组可比较的数据,并且图形能更清楚地呈现趋势、差异或构成时,必须制作数据图。满足该条件时,不能用结构图或定性矩阵代替。
      
      只有数据口径可比、能够绑定 Source ID,并能说明时间、单位和统计口径时才制作数据图。数据不足时不要编造分值、补齐数据或混用不可比口径。
      
      核心判断涉及多方相互作用、层级与依赖、因果机制、关键转折或行动路径,且用图能显著减少读者跨段对照、记忆或推演负担时,应制作相应分析图。不能以正文已经描述为由省略,也不能把缺少量化数据当作省略结构图的理由。
      
      不设统一图数或每章配额;同一张图可以解释多个紧密关联的问题,不同问题需要分别解释时可以增加图。简单顺序可以用短列表,精确的逐项对照可以用表格;涉及路径、网络或反馈时,检查表格是否掩盖了关键关系。
      
      ## 交付前复核
      
      检查支撑核心结论的比较、演化与机制分析:是否仍有需要读者反复对照段落才能理解的关键关系?有则补图,或调整已有图使其表达清楚;删除重复、装饰性或证据不足的图。
      
      全文无图时必须执行上述复核。确实没有适合作图的核心关系,或证据不足以支持关系表达时,允许无图;在已有工作记录中简要指出所检查的核心问题及适合文字或表格的具体原因,不用“文字已经足够清晰”等泛泛理由代替判断,也无需新增报告附录或申请豁免。
      
      校验器负责检查图注、来源标识、文件引用与 SVG 属性等客观条件,不按图数判定报告质量。脚本通过后仍需完成内容复核与渲染检查。
      
      ## 选择图形
      
      | 分析问题 | 优先图形 |
      |---|---|
      | 数量、构成或对象比较 | 柱状图、条形图、散点图、构成图 |
      | 指标如何变化 | 折线图、阶段趋势图 |
      | 路径如何演化 | 时间轴、阶段演进图、关键转折链 |
      | 多方力量如何作用 | 力场图、利益相关方网络、压力—响应图 |
      | 机制如何运转 | 因果链、反馈回路、约束结构图 |
      | 未来变量如何组合 | 驱动变量地图、情景矩阵、领先指标表 |
      | 方案如何取舍、行动如何衔接 | 决策矩阵、决策树、分阶段路线图 |
      
      没有量化数据但有证据支持的关系可以使用结构图,不能伪装成量化结果;推断关系须明确标注,并解释证据边界。
      
      ## 图表要求
      
      每张图必须包含:
      
      1. 写明图号和问题的 `figcaption`;
      2. 支撑数值、节点或关系的 Source ID;
      3. 正文对主要发现和证据边界的解释;
      4. SVG 的 `viewBox`、`role="img"` 和 `aria-label`。
      
      数据图还要标明单位、时间范围和统计口径。结构图的节点和箭头必须有证据支持;没有反馈闭环时不要称为“飞轮”,没有方向证据时不要强行添加箭头。
      
      所有 inline SVG 和已渲染图片都必须放在 `figure` 中,并用 `figcaption` 标明图号和问题。最终报告不能保留 Mermaid 源码。
      
      ## HTML/PDF
      
      使用 [../assets/report.css](../assets/report.css)。SVG 文本和数据标签在 A4 页面上应保持可读。
      
      尽量避免图形和单行表格跨页;超长表格允许自然分页。
      
      生成 PDF 后检查封面、正文、图表页和末页,确认中文字体、数值标签、裁切、重叠、空白图、页眉页脚和编号正确。
      
      渲染器将直接引用的本地图片嵌入 HTML,避免发送后依赖本机绝对路径。外部 SVG 的嵌套依赖与远程媒体仍需人工检查;正式交付优先使用已核实的本地或自包含资源。HTML/PDF 同名 `.manifest.json` 用于核对输入、产物和直接本地依赖的摘要,不能替代实际视觉检查,也不能保证远程资源内容不变。
      
    • xyz-method.md 11.2 KB
      # 研究底图与 X/Y/Z 解释方法
      
      这套方法先横向看清研究对象,再沿关键变化向下追问原因和作用过程,最后返回全局检查解释是否成立、是否完整。
      
      研究底图回答“它是什么、如何运转、当前处于什么状态”;X 识别显性与非显性的关键变化及其后续影响;Y 针对具体变化提出和比较成因解释;Z 打开解释中尚不清楚的连接,说明因素如何产生影响。
      
      研究的追问方向是:
      
      > 当前状态与关键变化 → 成因解释 → 作用过程。
      
      最终需要说明的因果关系是:
      
      > 在特定条件下,哪些因素通过什么过程造成变化;变化又如何改变主体的处境,并影响后续发展。
      
      分析结果持续返回 X 和研究底图。深入的标准是减少关键解释中的未知和未经检验的假设,同时保留会改变结论的重要侧面,而不是增加术语、层级、篇幅或固定数量的节点。
      
      ## 研究底图:先横向看清对象
      
      进入因果主线前,先确认:
      
      - 主体边界以及需要区分的组成部分;
      - 它如何产生价值、收入、成本、风险或其他核心结果;
      - 影响结果的主要参与者、关系和约束;
      - 当前最重要的状态、矛盾和口径差异;
      - 用户真正需要判断的问题。
      
      根据 [object-adapters.md](object-adapters.md) 选择适合对象的覆盖问题。覆盖问题用于发现材料缺口,不要求逐项写成正文,也不能替代 X/Y/Z。某个侧面缺失后会改变核心判断时,应纳入正文、补充证据或在 A3 说明缺口;无关侧面可以省略。
      
      先有研究底图,再选择解释主线。不能因为某条叙事容易成立,就只检索支持该叙事的部分。
      
      ## X:识别关键变化,建立待解释的路径
      
      回答:主体经历了哪些重要变化,它们改变了什么,又如何影响后续发展?
      
      关键节点是解释路径上的定位单位,可以是事件、选择、转折,也可以是一段累积过程或重要的持续状态。依据其对研究问题的解释价值选择节点,不按新闻热度或公开可见度选择。
      
      同时关注:
      
      - **显性节点**:有明确事件或公开记录的变化;
      - **非显性节点**:没有单一宣布时点,但有证据表明主体的能力、资源、约束、关系或选择空间发生了实质变化。
      
      非显性不等于隐秘,也不等于可以推测。渐进变化使用时间区间或阶段表述,不强行指定发生日期。
      
      ### 节点契约
      
      先把节点拆成三个层次:
      
      1. **可观察变化**:相较此前发生了什么,有什么证据;
      2. **状态改变**:该变化使主体的能力、资源、约束或位置发生了什么变化;
      3. **后续影响**:它改变了哪些选择,并如何成为后续变化的条件。
      
      “交付周期持续延长”可以是可观察变化;“组织能力衰退”首先是候选解释。不能把待验证的解释包装成节点事实。
      
      每个进入正文的关键节点说明:
      
      - 为什么它对研究问题重要;
      - 经 Y 和 Z 支持的原因摘要,哪些仍未解决;
      - 影响在什么条件下持续、减弱或逆转。
      
      X 只呈现足够理解路径的原因摘要。完整的因素比较放在 Y,具体作用过程放在 Z,避免同一段论证写三遍。
      
      筛选节点时,同时检查不符合初步解释的转折、重要的持续状态,以及有合理预期却没有发生的变化。未发生的变化只能在存在预期依据和可观察结果时使用,不能凭空构造反事实。
      
      影响主体的外部变化也可进入 X,但应说明它如何改变主体。仅用于解释其他节点的背景通常留在 Y。先建立暂定路径,再随证据修正;时间先后本身不能证明因果,也不把发展过程写成必然通向今天的直线。
      
      ## Y:沿具体节点切开,比较成因解释
      
      回答:为什么这一变化会在当时、以这种方式发生?
      
      每次分析先说明要解释的是节点的发生、时点、方向、规模,还是具体形式。一个节点可以包含多个待解释问题,它们可能对应不同因素。
      
      识别实际可能发挥作用的因素,并说明它们分别解释结果的哪一部分。因素可以来自主体内部、外部环境或双方关系;不按资本、技术、市场、监管等类别机械填满清单。
      
      根据实际问题辨认:
      
      - 长期积累的背景条件;
      - 使变化在当时发生的触发因素;
      - 使变化能够实现的能力或机会;
      - 限制、阻碍或改变变化形式的因素。
      
      这些角色是思考提示,不要求每个节点四项齐全。
      
      对进入解释的因素检查:
      
      - 它当时处于什么状态,或发生了什么变化;
      - 为什么将它纳入解释,有什么行为或结果证据;
      - 它与其他因素如何增强、依赖、抵消或相互替代;
      - 可信的替代解释是什么,现有材料能否区分。
      
      “如果该因素不存在或更弱,结果会不会不同”可用于提出比较问题和寻找证据,不能把想象中的答案当作事实。
      
      Y 应形成有主次、有关系的成因解释。无法判断相对重要性时明确保留,不随意排序或分配贡献比例;同时出现的变化不能自动视为原因。
      
      对影响结论的争议,说明候选解释是相互竞争、共同作用,还是同一链条的中间环节。比较它们分别预期出现什么、哪些观察两者都能解释,以及什么证据能够改变相对判断。共同预期的结果不是区分性反证;没有可区分材料时保留多种解释,不强选赢家,也不为无实质争议的问题强造反命题。
      
      ## Z:打开尚不清楚的连接,解释影响如何产生
      
      回答:成因解释中的影响究竟如何传递,为什么能够形成观察到的结果?
      
      Z 拆解的是解释链中尚未说明的连接,而不只是继续给因素分类。可以拆单个因素如何作用,也可以拆多个因素怎样共同产生结果。
      
      根据实际对象追踪:
      
      > 因素或条件变化 → 哪些参与者或系统状态受到影响 → 什么能力、成本、收益、信息或约束改变 → 引发什么行为或系统响应 → 如何形成观察结果。
      
      这是一组追问,不是所有机制必须经过的固定步骤。机制可以发生在主体内部、外部关系、市场、技术或制度中,也不必全部归结为人的动机。
      
      机制拆解必须补出此前缺失的中间环节、成立条件或可观察后果。改述因素名称、扩写原句,或使用“网络效应”“规模经济”“激励错配”等抽象术语,不算解释增量。使用机制名称后,继续说明谁或什么发生了变化、经过哪些环节、对结果造成什么影响。
      
      ### 过程证据与归因边界
      
      先明确主张的对象、结果、范围及必要的时间边界;只有不同部分需要不同证据强度时才拆成多个 Claim。对核心因果和机制判断分别说明:
      
      - **过程证据**:哪些关键连接在本案例中已观察,哪些只有理论或类比支持,哪些未知;
      - **归因边界**:对所声称结果,证据支持到作用存在、方向、重要性或贡献幅度的哪一范围;哪些替代解释仍无法区分。
      
      这两个问题不组成统一等级。“过程已经运行”和“对结果的贡献未知”可以同时成立。只需标明影响核心判断的薄弱连接,不要求每条箭头建立账本。结论不能跨越未经支持的必要连接,也不机械计算全链条的最低置信度。定量材料不足时不强求贡献估算;项目交付、实际运营、经济效果等不同结果必须分别按主张核验。
      
      存在共同作用、反馈或阈值时,说明条件和时序。结果可能反过来改变原有因素,但不能用循环叙述代替解释。证据不足的连接保留为缺口,不用心理揣测或更抽象的概念补齐。
      
      Z 也用于解释当前系统如何产生价值、成本、风险或其他核心结果。一个重要经营机制即使不对应戏剧性的历史事件,只要解释研究底图中的关键状态,也应纳入分析。
      
      ## 在工作笔记中保持对应关系
      
      复杂研究可在已有笔记中维护以下关系,不要求生成独立过程文件,也不作为最终报告模板:
      
      | 节点 | 可观察变化 | 状态改变 | 待解释问题 | 候选因素 | 待拆连接 | 证据状态 | 后续影响 |
      |---|---|---|---|---|---|---|---|
      
      这张内部图用于保持节点数量和命名一致,确认 Y 指向明确节点、Z 承接真实解释缺口。正文仍使用自然语言,不暴露内部编号或强制表格结构。
      
      ## 第一次回返:校正路径与解释
      
      将因素和机制放回对应节点,检查:
      
      - 能否解释变化的发生、时点、方向、规模和具体形式;
      - 是否只能解释方向,无法解释影响程度;
      - 是否存在不符合解释的观察或更有力的替代解释;
      - 节点如何改变后续条件,又影响了哪些发展可能;
      - 哪个连接、证据缺口或前提最可能使判断失效。
      
      分析结果改变了节点定义时,返回修改 X。不要为了保持原有叙事而保留前后不一致的节点。
      
      同一机制可能影响多个节点,同一节点也可能由多个因素共同造成。保持对应关系,避免重复拆解,也避免用一个宏大原因解释全部变化。
      
      ## 第二次回返:检查研究底图是否被主线压缩
      
      完成核心解释后,返回对象适配器的覆盖问题,检查:
      
      - 是否遗漏会改变结论的重要组成部分、参与者或口径;
      - 是否只解释增长或变化,没有解释价值、成本、风险和约束;
      - 是否只分析主体,没有分析用户、供应方、渠道、合作方或制度环境;
      - 是否因一条主线叙述顺畅,删掉了重要反向材料或另一套机制;
      - 未纳入正文的侧面是否应在 A3 说明范围和证据缺口。
      
      覆盖复核不要求把所有问题写进报告。只有缺失后会改变核心判断的侧面才需要补入;避免把全面研究变成按类别堆积资料。
      
      ## 停止条件
      
      优先深入最影响核心判断、解释分歧最大或证据最薄弱的连接,不要求所有节点使用相同篇幅或拆解层数。
      
      当关键作用过程、成立条件和证据边界已经说明,研究底图中的重要侧面没有被静默遗漏,且继续拆解预计不会改变原因排序、适用边界或后续判断时,可以停止。仍有重要连接无法解释时,将其保留为未解决问题。
      
      最终可以确认既有解释、修正其边界,或提出有证据支持的新解释,不强求新颖性。关键因果与机制判断写入 A2,沿用证据协议记录支持、反向材料、证据状态、缺口和反证条件。
      
      ## 未来表达
      
      仅在研究问题涉及决策或未来变化时展开。从已识别的因素、机制及成立条件出发,说明什么变化可能改变后续路径,以及可以观察到什么信号。
      
      - 某组条件明显占优:说明基准路径、触发信号和反证条件;
      - 两个高影响、高不确定变量适合共同分析:可使用 2×2 情景矩阵;
      - 证据不足:列出待验证问题和观察信号,不给出确定路径。
      
      未来判断必须承接已有解释,不机械生成多套情景,也不用“增长、停滞、衰退”代替机制分析。
      
  • scripts
    • linkify_sources.py 3 KB
      """Add source links to HTML generated by ``render_report.py``."""
      
      from __future__ import annotations
      
      import re
      
      
      SOURCE_ID = r"S\d{2,}"
      SKIP_ELEMENTS = {"a", "code", "pre", "script", "style"}
      SOURCE_ROW = re.compile(
          rf"<tr\b(?=[^>]*\bid\s*=\s*['\"]src-{SOURCE_ID}['\"])[^>]*>",
          flags=re.IGNORECASE,
      )
      
      
      def linkify_html(html: str) -> tuple[str, int, int]:
          """Link ``[Sxx]`` citations and return HTML, links added, and ledger rows."""
      
          row_pattern = re.compile(
              rf"<tr(?P<attrs>[^>]*)>\s*"
              rf"(?P<cells><td(?:\s[^>]*)?>\s*(?P<sid>{SOURCE_ID})\s*</td>.*?</tr>)",
              flags=re.IGNORECASE | re.DOTALL,
          )
      
          def add_anchor(match: re.Match[str]) -> str:
              attrs = match.group("attrs")
              if re.search(r"\bid\s*=", attrs, flags=re.IGNORECASE):
                  return match.group(0)
              return f'<tr{attrs} id="src-{match.group("sid")}">{match.group("cells")}'
      
          html = row_pattern.sub(add_anchor, html)
          url_map: dict[str, str] = {}
          for row in re.finditer(r"<tr\b([^>]*)>(.*?)</tr>", html, re.IGNORECASE | re.DOTALL):
              anchor = re.search(
                  rf"\bid\s*=\s*['\"]src-({SOURCE_ID})['\"]",
                  row.group(1),
                  flags=re.IGNORECASE,
              )
              if not anchor:
                  continue
              href = re.search(
                  r"href\s*=\s*(['\"])(https?://.*?)\1",
                  row.group(2),
                  flags=re.IGNORECASE | re.DOTALL,
              )
              bare = re.search(r'(?<!["\'=])(https?://[^\s<)"\']+)', row.group(2))
              url = href.group(2) if href else bare.group(1) if bare else ""
              if url:
                  url_map[anchor.group(1)] = url.rstrip(".,;。")
      
          def citation(match: re.Match[str]) -> str:
              sid = match.group(1)
              if sid in url_map:
                  return (
                      f'<a href="{url_map[sid]}" target="_blank" '
                      f'rel="noopener noreferrer" class="src-link">[{sid}]</a>'
                  )
              return f'<a href="#src-{sid}" class="src-link">[{sid}]</a>'
      
          output: list[str] = []
          in_ledger_row = False
          skip_depth = 0
          n_links = 0
          for part in re.split(r"(<[^>]+>)", html):
              if not part.startswith("<"):
                  if in_ledger_row or skip_depth:
                      output.append(part)
                  else:
                      linked, count = re.subn(rf"\[({SOURCE_ID})\]", citation, part)
                      output.append(linked)
                      n_links += count
                  continue
      
              tag = re.match(r"<\s*(/?)\s*([a-z0-9]+)", part, flags=re.IGNORECASE)
              if tag:
                  closing, name = bool(tag.group(1)), tag.group(2).lower()
                  if name == "tr":
                      in_ledger_row = False if closing else bool(SOURCE_ROW.match(part))
                  if name in SKIP_ELEMENTS:
                      if closing:
                          skip_depth = max(0, skip_depth - 1)
                      elif not part.rstrip().endswith("/>"):
                          skip_depth += 1
              output.append(part)
      
          return "".join(output), n_links, len(SOURCE_ROW.findall(html))
      
    • render_report.py 12 KB
      #!/usr/bin/env python3
      """Render a 3d-deep-research Markdown report to HTML or PDF."""
      
      from __future__ import annotations
      
      import argparse
      import base64
      import mimetypes
      import html
      import os
      import re
      import shutil
      import subprocess
      import sys
      from pathlib import Path
      from urllib.parse import unquote, urlsplit, urlunsplit
      
      from linkify_sources import linkify_html
      from report_artifacts import MediaReferences, local_path, sha256, write_manifest
      
      
      SKILL_DIR = Path(__file__).resolve().parent.parent
      DEFAULT_CSS = SKILL_DIR / "assets" / "report.css"
      REQUIRED_WEASYPRINT_VERSION = "69.0"
      FONT_FILES = {
          400: "NotoSansCJKsc-Regular.otf",
          700: "NotoSansCJKsc-Bold.otf",
      }
      
      
      def _configure_utf8_console() -> None:
          for stream in (sys.stdout, sys.stderr):
              reconfigure = getattr(stream, "reconfigure", None)
              if reconfigure:
                  reconfigure(encoding="utf-8", errors="replace")
      
      
      def _marked_command() -> list[str] | None:
          node = os.environ.get("CODEX_PRIMARY_RUNTIME_NODE")
          modules = os.environ.get("CODEX_PRIMARY_RUNTIME_NODE_MODULES")
          if node and modules:
              cli = Path(modules) / "marked" / "bin" / "marked.js"
              if Path(node).is_file() and cli.is_file():
                  return [node, str(cli)]
      
          marked = shutil.which("marked")
          return [marked] if marked else None
      
      
      def markdown_to_html(md_text: str) -> tuple[str, str]:
          """Return rendered HTML and the converter name."""
          try:
              import markdown  # type: ignore
      
              rendered = markdown.markdown(
                  md_text,
                  extensions=["tables", "fenced_code", "sane_lists"],
                  output_format="html5",
              )
              version = getattr(markdown, "__version__", "")
              return rendered, f"python-markdown {version}".strip()
          except ModuleNotFoundError:
              pass
      
          command = _marked_command()
          if command:
              result = subprocess.run(
                  [*command, "--gfm"],
                  input=md_text,
                  text=True,
                  encoding="utf-8",
                  capture_output=True,
                  check=False,
                  timeout=60,
              )
              if result.returncode == 0 and result.stdout.strip():
                  return result.stdout, "marked"
              detail = result.stderr.strip() or f"exit code {result.returncode}"
              raise RuntimeError(f"marked failed: {detail}")
      
          raise RuntimeError(
              "No Markdown converter found. Install Python-Markdown or expose "
              "marked on PATH."
          )
      
      
      def _split_report(md_text: str) -> tuple[str, str, str]:
          lines = md_text.splitlines()
          first_content = next(
              (index for index, line in enumerate(lines) if line.strip()),
              None,
          )
          if first_content is None:
              raise RuntimeError("The report is empty.")
      
          title_match = re.fullmatch(r"#\s+(.+?)\s*", lines[first_content])
          if not title_match:
              raise RuntimeError("The report must start with an H1 title.")
          title = title_match.group(1)
          lines[first_content] = ""
      
          meta_line = ""
          for index in range(first_content + 1, min(len(lines), first_content + 12)):
              line = lines[index].lstrip()
              if line.startswith(">"):
                  meta_line = line[1:].strip()
                  lines[index] = ""
                  break
      
          body = "\n".join(lines)
          body = re.sub(r"^>\s*证据契约:3\s*$", "", body, flags=re.MULTILINE)
          return title, meta_line, body
      
      
      def _resolve_local_media(html_body: str, base_dir: Path | None, embed: bool = False) -> str:
          """Resolve relative image references against the Markdown directory."""
          if base_dir is None:
              return html_body
      
          attribute = re.compile(
              r'(<(?:img|image|use)\b[^>]*?\b(?:src|href|xlink:href)\s*=\s*)(["\'])(.*?)\2',
              flags=re.IGNORECASE | re.DOTALL,
          )
      
          def resolve(match: re.Match[str]) -> str:
              value = match.group(3).strip()
              path = local_path(html.unescape(value), base_dir)
              if embed and path is not None:
                  mime = mimetypes.guess_type(path.name)[0] or "application/octet-stream"
                  encoded = base64.b64encode(path.read_bytes()).decode("ascii")
                  fragment = urlsplit(value).fragment
                  uri = f"data:{mime};base64,{encoded}" + (f"#{fragment}" if fragment else "")
                  return f"{match.group(1)}{match.group(2)}{uri}{match.group(2)}"
              parsed = urlsplit(value)
              if (
                  not value
                  or parsed.scheme
                  or parsed.netloc
                  or value.startswith(("//", "#", "/"))
              ):
                  return match.group(0)
      
              local_uri = (base_dir / unquote(parsed.path)).resolve().as_uri()
              resolved = urlsplit(local_uri)
              rewritten = urlunsplit(
                  (resolved.scheme, resolved.netloc, resolved.path, parsed.query, parsed.fragment)
              )
              return f"{match.group(1)}{match.group(2)}{rewritten}{match.group(2)}"
      
          return attribute.sub(resolve, html_body)
      
      
      def _font_face_css(font_dir: Path) -> str:
          missing = [name for name in FONT_FILES.values() if not (font_dir / name).is_file()]
          if missing:
              raise RuntimeError(
                  f"Missing required Noto CJK font(s) in {font_dir}: "
                  + ", ".join(missing)
              )
          rules = []
          for weight, filename in FONT_FILES.items():
              font_uri = (font_dir / filename).resolve().as_uri()
              rules.append(
                  '@font-face {\n'
                  '  font-family: "Noto Sans CJK SC";\n'
                  f'  src: url("{font_uri}") format("opentype");\n'
                  '  font-style: normal;\n'
                  f'  font-weight: {weight};\n'
                  '}'
              )
          return "\n".join(rules)
      
      
      def build_html(
          md_text: str,
          asset_base: Path | None = None,
          font_dir: Path | None = None,
      ) -> tuple[str, str, str]:
          report_title, meta_line, body_md = _split_report(md_text)
          html_body, converter = markdown_to_html(body_md)
          html_body = _resolve_local_media(html_body, asset_base, embed=True)
      
          css = DEFAULT_CSS.read_text(encoding="utf-8")
          if "HEADER_TEXT" not in css:
              raise RuntimeError(f"Missing HEADER_TEXT placeholder in {DEFAULT_CSS}")
          css_header = report_title.replace("\\", "\\\\").replace('"', '\\"')
          css = css.replace("HEADER_TEXT", css_header)
          if font_dir is not None:
              css = _font_face_css(font_dir) + "\n" + css
      
          safe_title = html.escape(report_title, quote=True)
          meta_html = (
              f'<div class="meta">{html.escape(meta_line, quote=True)}</div>'
              if meta_line
              else ""
          )
          cover = f"""
      <div class="cover">
        <h1 style="page-break-before: avoid; border: none;">{safe_title}</h1>
        {meta_html}
        <hr class="divider">
      </div>
      """.strip()
      
          document = f"""<!DOCTYPE html>
      <html lang="zh-CN">
      <head>
        <meta charset="UTF-8">
        <meta name="viewport" content="width=device-width, initial-scale=1">
        <title>{safe_title}</title>
        <style>{css}</style>
      </head>
      <body>
      {cover}
      {html_body}
      </body>
      </html>
      """
          return document, report_title, converter
      
      
      def render_with_weasyprint(html_text: str, input_path: Path, output_path: Path) -> str:
          try:
              import weasyprint  # type: ignore
          except Exception as exc:
              raise RuntimeError(f"WeasyPrint is unavailable: {exc}") from exc
      
          version = getattr(weasyprint, "__version__", "unknown")
          if version != REQUIRED_WEASYPRINT_VERSION:
              raise RuntimeError(
                  f"WeasyPrint {REQUIRED_WEASYPRINT_VERSION} is required; found {version}. "
                  f"Install with `python -m pip install weasyprint=={REQUIRED_WEASYPRINT_VERSION}`."
              )
      
          weasyprint.HTML(
              string=html_text,
              base_url=str(input_path.parent),
          ).write_pdf(str(output_path))
          if not output_path.is_file() or output_path.stat().st_size < 1024:
              raise RuntimeError("WeasyPrint did not create a valid PDF.")
          return f"weasyprint {version}"
      
      
      def parse_args() -> argparse.Namespace:
          parser = argparse.ArgumentParser(
              description="Render 3d-deep-research Markdown to HTML or PDF."
          )
          parser.add_argument("input", help="Input Markdown path")
          parser.add_argument("output", help="Output .html or .pdf path")
          parser.add_argument(
              "--font-dir",
              default="fonts",
              help="Directory containing NotoSansCJKsc-Regular.otf and Bold.otf; "
              "relative paths resolve from the Markdown directory",
          )
          return parser.parse_args()
      
      
      def main() -> None:
          _configure_utf8_console()
          args = parse_args()
          input_path = Path(args.input).expanduser().resolve()
          output_path = Path(args.output).expanduser().resolve()
      
          if not input_path.is_file():
              raise SystemExit(f"Input Markdown not found: {input_path}")
          if output_path.suffix.lower() not in {".html", ".pdf"}:
              raise SystemExit("Output path must end in .html or .pdf.")
      
          font_dir: Path | None = None
          if output_path.suffix.lower() == ".pdf":
              font_dir = Path(args.font_dir).expanduser()
              if not font_dir.is_absolute():
                  font_dir = input_path.parent / font_dir
              font_dir = font_dir.resolve()
      
          output_path.parent.mkdir(parents=True, exist_ok=True)
          md_text = input_path.read_text(encoding="utf-8")
          media = MediaReferences()
          media.feed(md_text)
          refs = media.refs + re.findall(r"!\[[^\]]*\]\(([^)\s]+)", md_text)
          dependencies = [DEFAULT_CSS]
          remote = []
          for ref in refs:
              path = local_path(ref, input_path.parent)
              if path is not None:
                  dependencies.append(path)
              elif ref.startswith(("https://", "http://", "//")):
                  remote.append(ref)
          if font_dir is not None:
              dependencies.extend(font_dir / name for name in FONT_FILES.values())
          try:
              expected_inputs = {path: sha256(path) for path in {input_path, *dependencies}}
              # Catch an edit between reading Markdown and taking the resource snapshot.
              if input_path.read_text(encoding="utf-8") != md_text:
                  raise ValueError("Markdown changed while preparing render; render again.")
              rendered_html, report_title, converter = build_html(
                  md_text,
                  asset_base=input_path.parent,
                  font_dir=font_dir,
              )
              rendered_html, n_links, n_rows = linkify_html(rendered_html)
          except (OSError, RuntimeError, ValueError) as exc:
              raise SystemExit(f"HTML rendering failed: {exc}") from exc
      
          if output_path.suffix.lower() == ".html":
              output_path.write_text(rendered_html, encoding="utf-8")
              try:
                  write_manifest(input_path, output_path, dependencies, remote, expected_inputs)
              except (OSError, ValueError) as exc:
                  raise SystemExit(f"Artifact manifest failed; render again: {exc}") from exc
              print(f"[OK] HTML: {output_path}")
              print(f"[OK] Markdown converter: {converter}")
              print(f"[OK] Source links: {n_links}; ledger anchors: {n_rows}")
              return
      
          html_path = output_path.with_name(
              f".{output_path.name}.{os.getpid()}.rendering.html"
          )
          temporary_pdf = output_path.with_name(
              f".{output_path.name}.{os.getpid()}.rendering.pdf"
          )
          html_path.write_text(rendered_html, encoding="utf-8")
      
          temporary_pdf.unlink(missing_ok=True)
          try:
              selected_engine = render_with_weasyprint(
                  rendered_html,
                  input_path,
                  temporary_pdf,
              )
          except Exception as exc:
              temporary_pdf.unlink(missing_ok=True)
              raise SystemExit(
                  "PDF rendering failed. Intermediate HTML was kept for debugging:\n"
                  f"{html_path}\n- {exc}"
              ) from exc
      
          os.replace(temporary_pdf, output_path)
          try:
              write_manifest(input_path, output_path, dependencies, remote, expected_inputs)
          except (OSError, ValueError) as exc:
              raise SystemExit(f"Artifact manifest failed; render again: {exc}") from exc
          html_path.unlink(missing_ok=True)
          size_kb = output_path.stat().st_size / 1024
          print(f"[OK] PDF: {output_path} ({size_kb:.1f} KB)")
          print(f"[OK] PDF engine: {selected_engine}")
          print(f"[OK] Markdown converter: {converter}")
          print(f"[OK] Source links: {n_links}; ledger anchors: {n_rows}")
          print(f"[OK] Title: {report_title}")
      
      
      if __name__ == "__main__":
          main()
      
    • report_artifacts.py 3.7 KB
      """Track local report inputs and output identity (not a security signature)."""
      from __future__ import annotations
      
      import hashlib
      import json
      from html.parser import HTMLParser
      from pathlib import Path
      from urllib.parse import unquote, urlsplit
      from urllib.request import url2pathname
      
      
      class MediaReferences(HTMLParser):
          def __init__(self) -> None:
              super().__init__()
              self.refs: list[str] = []
      
          def handle_starttag(self, tag: str, attrs: list[tuple[str, str | None]]) -> None:
              if tag in {"img", "image", "use"}:
                  values = dict(attrs)
                  ref = values.get("src") if tag == "img" else values.get("href") or values.get("xlink:href")
                  if ref:
                      self.refs.append(ref)
      
      
      def local_path(ref: str, base: Path) -> Path | None:
          if len(ref) > 2 and ref[1] == ":" and ref[2] in "\\/":
              return Path(ref).resolve()
          parsed = urlsplit(ref)
          if ref.startswith("#") or parsed.scheme in {"http", "https", "data"} or parsed.netloc:
              return None
          if parsed.scheme == "file":
              return Path(url2pathname(parsed.path)).resolve()
          if parsed.scheme:
              return None
          return (base / unquote(parsed.path)).resolve()
      
      
      def sha256(path: Path) -> str:
          return hashlib.sha256(path.read_bytes()).hexdigest()
      
      
      def manifest_path(output: Path) -> Path:
          return output.with_name(output.name + ".manifest.json")
      
      
      def write_manifest(
          source: Path, output: Path, dependencies: list[Path], remote: list[str],
          expected_inputs: dict[Path, str] | None = None,
      ) -> None:
          import os
          def dependency_name(path: Path) -> str:
              try:
                  return os.path.relpath(path, source.parent)
              except ValueError:  # Different Windows drives have no relative path.
                  return str(path.resolve())
      
          input_hashes = {path: sha256(path) for path in {source, *dependencies}}
          if expected_inputs is not None and input_hashes != expected_inputs:
              raise ValueError("Report inputs changed during rendering; render again.")
          data = {
              "version": 1,
              "source_sha256": input_hashes[source],
              "output_sha256": sha256(output),
              "dependencies": [
                  {"path": dependency_name(path), "sha256": input_hashes[path]}
                  for path in sorted(set(dependencies))
              ],
              "remote_resources": sorted(set(remote)),
          }
          destination = manifest_path(output)
          temporary = destination.with_name(destination.name + ".tmp")
          temporary.write_text(json.dumps(data, ensure_ascii=False, indent=2), encoding="utf-8")
          temporary.replace(destination)
      
      
      def validate_manifest(source: Path, output: Path) -> tuple[list[str], list[str]]:
          errors: list[str] = []
          warnings: list[str] = []
          try:
              data = json.loads(manifest_path(output).read_text(encoding="utf-8"))
              if data["version"] != 1:
                  raise ValueError("unsupported manifest version")
              if data["source_sha256"] != sha256(source):
                  errors.append("Artifact was built from a different Markdown revision.")
              if data["output_sha256"] != sha256(output):
                  errors.append("Artifact does not match its build manifest.")
              for dependency in data["dependencies"]:
                  path = source.parent / dependency["path"]
                  if not path.is_file() or sha256(path) != dependency["sha256"]:
                      errors.append(f"Build dependency is missing or changed: {dependency['path']}")
              if data.get("remote_resources"):
                  warnings.append("Remote media are not content-pinned; verify availability and rendered content.")
          except (OSError, ValueError, KeyError, TypeError) as exc:
              errors.append(f"Missing or invalid artifact manifest; render again: {exc}")
          return errors, warnings
      
    • validate_report.py 25.4 KB
      #!/usr/bin/env python3
      """Validate machine-checkable parts of a 3d-deep-research report."""
      
      from __future__ import annotations
      
      import argparse
      import re
      import sys
      from datetime import date
      from html.parser import HTMLParser
      from math import isfinite
      from pathlib import Path
      
      from report_artifacts import MediaReferences, local_path, validate_manifest
      
      
      REQUIRED_MAIN_SECTIONS = ["一", "二", "三", "四"]
      CLAIM_TYPES = {
          "fact",
          "causal",
          "mechanism",
          "market",
          "forecast",
          "事实",
          "因果",
          "机制",
          "市场",
          "预测",
      }
      MECHANISM_STATUSES = {
          "机制可行",
          "案例中运行",
          "解释力已确认",
          "解释力未确认",
          "possible",
          "operating",
          "confirmed",
          "unconfirmed",
      }
      CONFIDENCE_LEVELS = {"high", "medium", "low", "高", "中", "低"}
      INDEPENDENCE_MARKERS = {
          "independent",
          "shared-origin",
          "unknown",
          "mixed",
          "独立",
          "非独立",
          "同源",
          "未知",
      }
      PLACEHOLDER_PATTERNS = (
          r"\[研究对象\]",
          r"\[YYYY(?:-MM-DD)?\]",
          r"\[一句话(?:问题|判断)?\]",
          r"\{\{[^}]+\}\}",
          r"\bTODO\b",
          r"\bTBD\b",
          r"\bURL\b",
          r"^\s*(?:#{1,6}\s*)?\[[^\]\n]+\]\s*$",
          r"\|\s*\[[^|\n]+\]\s*(?=\|)",
      )
      A4_WIDTH_POINTS = 595.28
      A4_HEIGHT_POINTS = 841.89
      PAGE_SIZE_TOLERANCE_POINTS = 5.0
      REQUIRED_PDF_PRODUCER = "WeasyPrint 69.0"
      REQUIRED_PDF_FONTS = {
          "Noto-Sans-CJK-SC",
          "Noto-Sans-CJK-SC-Bold",
      }
      LITERAL_HTML_PATTERN = re.compile(
          r"</?(?:br|figure|figcaption|svg|div|span|table|thead|tbody|tr|td|th)\b[^>]*>",
          flags=re.IGNORECASE,
      )
      
      
      def _configure_utf8_console() -> None:
          for stream in (sys.stdout, sys.stderr):
              reconfigure = getattr(stream, "reconfigure", None)
              if reconfigure:
                  reconfigure(encoding="utf-8", errors="replace")
      
      
      def _read_pdf(
          pdf_path: Path,
      ) -> tuple[str, list[tuple[float, float]], str, set[str]]:
          try:
              from pypdf import PdfReader  # type: ignore
          except ModuleNotFoundError:
              try:
                  from PyPDF2 import PdfReader  # type: ignore
              except ModuleNotFoundError as exc:
                  raise RuntimeError("Install pypdf or PyPDF2 to validate PDF output.") from exc
      
          reader = PdfReader(str(pdf_path))
          text_parts: list[str] = []
          page_sizes: list[tuple[float, float]] = []
          font_names: set[str] = set()
          for page in reader.pages:
              text_parts.append(page.extract_text() or "")
              width = float(page.mediabox.width)
              height = float(page.mediabox.height)
              rotation = int(page.get("/Rotate", 0) or 0) % 360
              if rotation in (90, 270):
                  width, height = height, width
              page_sizes.append((width, height))
              resources = page.get("/Resources")
              resources = resources.get_object() if resources else None
              fonts = resources.get("/Font") if resources else None
              fonts = fonts.get_object() if fonts else {}
              for font_ref in fonts.values():
                  font = font_ref.get_object()
                  base_font = str(font.get("/BaseFont", ""))
                  descendants = font.get("/DescendantFonts", [])
                  candidates = [font, *(item.get_object() for item in descendants)]
                  embedded = False
                  for candidate in candidates:
                      descriptor = candidate.get("/FontDescriptor")
                      descriptor = descriptor.get_object() if descriptor else {}
                      for key in ("/FontFile", "/FontFile2", "/FontFile3"):
                          stream = descriptor.get(key)
                          if stream and stream.get_object().get_data():
                              embedded = True
                  if base_font and embedded:
                      font_names.add(base_font.lstrip("/").split("+")[-1])
          metadata = reader.metadata or {}
          producer = str(metadata.get("/Producer", ""))
          return "\n".join(text_parts), page_sizes, producer, font_names
      
      
      def _validate_pdf_data(
          pdf_text: str,
          page_sizes: list[tuple[float, float]],
          producer: str,
          font_names: set[str],
      ) -> tuple[list[str], list[str]]:
          errors: list[str] = []
          warnings: list[str] = []
          if not page_sizes:
              errors.append("PDF has no pages.")
          if producer != REQUIRED_PDF_PRODUCER:
              errors.append(
                  f"PDF producer is {producer or '<missing>'}; expected {REQUIRED_PDF_PRODUCER}."
              )
          missing_fonts = REQUIRED_PDF_FONTS - font_names
          if missing_fonts:
              errors.append(
                  "PDF does not embed required Noto CJK fonts: "
                  + ", ".join(sorted(missing_fonts))
                  + "."
              )
          for page_number, (width, height) in enumerate(page_sizes, start=1):
              if width > height:
                  errors.append(f"PDF page {page_number} is landscape; expected A4 portrait.")
                  continue
              if (
                  abs(width - A4_WIDTH_POINTS) > PAGE_SIZE_TOLERANCE_POINTS
                  or abs(height - A4_HEIGHT_POINTS) > PAGE_SIZE_TOLERANCE_POINTS
              ):
                  errors.append(
                      f"PDF page {page_number} is {width:.1f} x {height:.1f} pt; "
                      "expected A4 portrait."
                  )
          literal_tag = LITERAL_HTML_PATTERN.search(pdf_text)
          if literal_tag:
              errors.append(f"PDF contains a literal HTML tag: {literal_tag.group(0)!r}.")
          if len(pdf_text.strip()) < 100:
              warnings.append("PDF text extraction returned very little text.")
          if "\ufffd" in pdf_text:
              warnings.append("PDF text contains Unicode replacement characters.")
          return errors, warnings
      
      
      def validate_pdf(pdf_path: Path) -> tuple[list[str], list[str]]:
          if not pdf_path.is_file() or pdf_path.stat().st_size < 1024:
              return [f"PDF is missing or too small: {pdf_path}"], []
          try:
              pdf_text, page_sizes, producer, font_names = _read_pdf(pdf_path)
          except Exception as exc:
              return [f"PDF validation failed: {exc}"], []
          return _validate_pdf_data(pdf_text, page_sizes, producer, font_names)
      
      
      def _split_body_and_appendix(md_text: str) -> tuple[str, str]:
          match = re.search(r"^##\s+附录\b", md_text, flags=re.MULTILINE)
          if not match:
              return md_text, ""
          return md_text[: match.start()], md_text[match.start() :]
      
      
      def _section_text(md_text: str, section_id: str) -> str | None:
          match = re.search(
              rf"^###\s+{re.escape(section_id)}\b[^\n]*\n(.*?)(?=^###\s+|^##\s+|\Z)",
              md_text,
              flags=re.MULTILINE | re.DOTALL,
          )
          return match.group(1) if match else None
      
      
      def _split_row(line: str) -> list[str]:
          value = line.strip()
          if not value.startswith("|"):
              return []
          value = value[1:-1] if value.endswith("|") else value[1:]
          return [cell.replace(r"\|", "|").strip() for cell in re.split(r"(?<!\\)\|", value)]
      
      
      def _is_separator(cells: list[str]) -> bool:
          return bool(cells) and all(re.fullmatch(r":?-{3,}:?", cell) for cell in cells)
      
      
      def _extract_table(
          md_text: str,
          section_id: str,
      ) -> tuple[list[str], list[list[str]], int]:
          section = _section_text(md_text, section_id)
          if section is None:
              return [], [], 0
      
          lines = section.splitlines()
          start = next(
              (index for index, line in enumerate(lines) if line.strip().startswith("|")),
              None,
          )
          if start is None:
              return [], [], 0
      
          table_lines: list[str] = []
          for line in lines[start:]:
              if not line.strip().startswith("|"):
                  if table_lines:
                      break
                  continue
              table_lines.append(line)
          if len(table_lines) < 2:
              return [], [], 0
      
          header = _split_row(table_lines[0])
          separator = _split_row(table_lines[1])
          if len(header) != len(separator) or not _is_separator(separator):
              return [], [], 0
      
          rows: list[list[str]] = []
          malformed = 0
          for line in table_lines[2:]:
              cells = _split_row(line)
              if len(cells) == len(header):
                  rows.append(cells)
              else:
                  malformed += 1
          return header, rows, malformed
      
      
      def _plain_text(value: str) -> str:
          value = re.sub(r"<[^>]+>", " ", value)
          value = re.sub(r"!?\[([^]]*)\]\([^)]+\)", r"\1", value)
          value = re.sub(r"[`*_~]", "", value)
          return re.sub(r"\s+", " ", value).strip()
      
      
      def _contains_enum(value: str, markers: set[str]) -> bool:
          for marker in markers:
              boundary = r"A-Za-z0-9_-" if marker.isascii() else r"\w-"
              if re.search(rf"(?<![{boundary}]){re.escape(marker)}(?![{boundary}])", value, re.I):
                  return True
          return False
      
      
      def _validate_metadata(md_text: str) -> list[str]:
          lines = md_text.splitlines()
          first_content = next(
              (index for index, line in enumerate(lines) if line.strip()),
              None,
          )
          if first_content is None:
              return ["The report metadata line is missing or malformed."]
      
          next_content = next(
              (line.strip() for line in lines[first_content + 1 :] if line.strip()),
              "",
          )
          match = re.fullmatch(
              r">\s*研究问题\s*[::]\s*(?P<question>.+?)\s*\|\s*"
              r"资料截止\s*[::]\s*(?P<cutoff>\d{4}-\d{2}-\d{2})\s*\|\s*"
              r"完成日期\s*[::]\s*(?P<completed>\d{4}-\d{2}-\d{2})\s*",
              next_content,
          )
          if not match or not _is_filled(match.group("question") if match else ""):
              return ["The report metadata line is missing or malformed."]
      
          try:
              cutoff = date.fromisoformat(match.group("cutoff"))
              completed = date.fromisoformat(match.group("completed"))
          except ValueError:
              return ["The report metadata contains an invalid ISO date."]
          if completed < cutoff:
              return ["The report completion date cannot precede the source cutoff date."]
          return []
      
      
      def _is_filled(value: str) -> bool:
          return bool(_plain_text(value).strip("-—[]()()。.;;:: "))
      
      
      def _validate_sources(
          rows: list[list[str]],
          legacy_schema: bool = False,
      ) -> tuple[list[str], set[str]]:
          errors: list[str] = []
          source_ids: set[str] = set()
          location_pattern = re.compile(
              r"https?://|(?:^|[\s(])[^\s)]+\."
              r"(?:pdf|html?|md|txt|csv|xlsx?|docx?)\b",
              flags=re.IGNORECASE,
          )
      
          for row_no, row in enumerate(rows, start=1):
              if len(row) != 4:
                  errors.append(f"A1 row {row_no} must have four columns.")
                  continue
              source_id = row[0].strip()
              if not re.fullmatch(r"S\d{2,}", source_id):
                  errors.append(f"A1 row {row_no} has an invalid Source ID: {source_id or '<empty>'}.")
                  continue
              if source_id in source_ids:
                  errors.append(f"A1 has duplicate Source ID {source_id}.")
              source_ids.add(source_id)
      
              details, role, limits = row[1:4]
              if not _is_filled(details):
                  errors.append(f"Source {source_id} has no source details.")
              if not location_pattern.search(details + " " + limits):
                  errors.append(f"Source {source_id} has no URL or file location.")
              if legacy_schema and not re.search(r"\b\d{4}-\d{2}-\d{2}\b", details):
                  errors.append(f"Source {source_id} has no ISO date.")
              for value in re.findall(r"\b\d{4}-\d{2}-\d{2}\b", details):
                  try:
                      date.fromisoformat(value)
                  except ValueError:
                      errors.append(f"Source {source_id} has an invalid ISO date: {value}.")
              if not legacy_schema:
                  fields = _labeled_fields(details)
                  for label in ("发布者", "发布", "访问"):
                      if not _is_filled(fields.get(label, "")):
                          errors.append(f"Source {source_id} is missing {label}.")
                  for label, unknown in (("发布", "未知"), ("访问", "不适用")):
                      value = fields.get(label, "")
                      if value != unknown and not re.fullmatch(r"\d{4}-\d{2}-\d{2}", value):
                          errors.append(f"Source {source_id} {label} must be an ISO date or {unknown}.")
              if not _is_filled(role):
                  errors.append(f"Source {source_id} has no evidence role.")
              elif not _contains_enum(role, INDEPENDENCE_MARKERS):
                  errors.append(f"Source {source_id} does not state source independence.")
              if not _is_filled(limits):
                  errors.append(f"Source {source_id} has no limitations.")
      
          return errors, source_ids
      
      
      def _labeled_fields(value: str) -> dict[str, str]:
          fields = {}
          for part in re.split(r"[;;]|<br\s*/?>", value):
              match = re.match(r"\s*([^::]+)[::]\s*(.*?)\s*$", part)
              if match:
                  fields[match.group(1)] = match.group(2)
          return fields
      
      
      def _validate_claims(
          rows: list[list[str]],
          source_ids: set[str],
          header: list[str] | None = None,
          legacy_schema: bool = False,
      ) -> tuple[list[str], set[str]]:
          errors: list[str] = []
          claim_ids: set[str] = set()
          used_sources: set[str] = set()
      
          new_schema = bool(
              header
              and len(header) == 6
              and ("机制状态" in header[2] or "证据范围" in header[2] or "mechanism status" in header[2].lower())
          )
      
          for row_no, row in enumerate(rows, start=1):
              if len(row) != 6:
                  errors.append(f"A2 row {row_no} must have six columns.")
                  continue
              claim_id = row[0].strip()
              if not re.fullmatch(r"C\d{2,}", claim_id):
                  errors.append(f"A2 row {row_no} has an invalid Claim ID: {claim_id or '<empty>'}.")
                  continue
              if claim_id in claim_ids:
                  errors.append(f"A2 has duplicate Claim ID {claim_id}.")
              claim_ids.add(claim_id)
      
              statement = row[1]
              claim_type = row[2]
              if not _is_filled(statement):
                  errors.append(f"Claim {claim_id} has no statement.")
              if not _contains_enum(claim_type, CLAIM_TYPES):
                  errors.append(f"Claim {claim_id} has no recognized type.")
              elif not legacy_schema and re.split(r"[;;]", claim_type)[0].strip().lower() not in CLAIM_TYPES:
                  errors.append(f"Claim {claim_id} must have exactly one type before its evidence fields.")
      
              if new_schema:
                  support_part, counterevidence, confidence_gap = row[3:6]
                  evidence = f"{support_part} {counterevidence}"
                  confidence = confidence_gap
                  gap = confidence_gap
                  if not _is_filled(counterevidence):
                      errors.append(
                          f"Claim {claim_id} has no counterevidence, alternative explanation, "
                          "or reverse-search note."
                      )
                  causal_or_mechanism = _contains_enum(
                      claim_type, {"causal", "mechanism", "因果", "机制"}
                  )
                  if not legacy_schema:
                      fields = _labeled_fields(confidence_gap)
                      for label in ("置信度", "缺口", "修订条件"):
                          if not _is_filled(fields.get(label, "")):
                              errors.append(f"Claim {claim_id} is missing {label}.")
                      confidence = fields.get("置信度", "")
                      gap = fields.get("缺口", "")
                      if causal_or_mechanism:
                          scope = _labeled_fields(claim_type)
                          for label in ("过程", "归因"):
                              if not _is_filled(scope.get(label, "")):
                                  errors.append(f"Claim {claim_id} is missing mechanism evidence scope: {label}.")
                  elif causal_or_mechanism and not _contains_enum(
                      claim_type, MECHANISM_STATUSES
                  ):
                      errors.append(f"Claim {claim_id} has no mechanism evidence status.")
              else:
                  if not legacy_schema:
                      errors.append("Legacy A2 schema requires --legacy-schema; migrate new reports to contract 3.")
                  evidence, confidence, gap = row[3:6]
                  reverse_marker = re.search(
                      r"反向|替代|未解决|counter",
                      evidence,
                      flags=re.IGNORECASE,
                  )
                  support_part = evidence[: reverse_marker.start()] if reverse_marker else evidence
                  if not reverse_marker:
                      errors.append(
                          f"Claim {claim_id} has no counterevidence or reverse-search note."
                      )
      
              supporting_sources = set(re.findall(r"S\d{2,}", support_part))
              claim_sources = set(re.findall(r"S\d{2,}", evidence))
              used_sources.update(claim_sources)
              if not supporting_sources:
                  errors.append(f"Claim {claim_id} has no supporting Source ID.")
              undefined = claim_sources - source_ids
              if undefined:
                  errors.append(
                      f"Claim {claim_id} references undefined Source IDs: "
                      + ", ".join(sorted(undefined))
                      + "."
                  )
      
              confidence_valid = (
                  _contains_enum(confidence, CONFIDENCE_LEVELS)
                  if legacy_schema else confidence.strip().lower() in CONFIDENCE_LEVELS
              )
              if not confidence_valid:
                  errors.append(f"Claim {claim_id} has no confidence level.")
              independence_value = support_part if new_schema else confidence
              if not _contains_enum(independence_value, INDEPENDENCE_MARKERS):
                  errors.append(f"Claim {claim_id} does not state evidence independence.")
              if not _is_filled(gap):
                  errors.append(f"Claim {claim_id} has no evidence gap or disconfirmation condition.")
      
          return errors, used_sources
      
      
      def _inside(position: int, ranges: list[tuple[int, int]]) -> bool:
          return any(start <= position < end for start, end in ranges)
      
      
      class _SVGAttributes(HTMLParser):
          def __init__(self) -> None:
              super().__init__()
              self.items: list[dict[str, str | None]] = []
      
          def handle_starttag(self, tag: str, attrs: list[tuple[str, str | None]]) -> None:
              if tag == "svg":
                  self.items.append(dict(attrs))
      
      
      def _validate_figures(md_text: str, base_dir: Path | None) -> list[str]:
          errors: list[str] = []
          figures = list(
              re.finditer(r"<figure\b.*?</figure>", md_text, flags=re.IGNORECASE | re.DOTALL)
          )
          ranges = [(figure.start(), figure.end()) for figure in figures]
      
          for index, match in enumerate(figures, start=1):
              figure = match.group(0)
              if not re.search(r"<figcaption\b[^>]*>\s*.+?</figcaption>", figure, re.I | re.S):
                  errors.append(f"Figure {index} has no non-empty figcaption.")
              if not re.search(r"\[S\d{2,}\]", figure):
                  errors.append(f"Figure {index} has no Source ID.")
              parser = _SVGAttributes()
              parser.feed(figure)
              for attributes in parser.items:
                  try:
                      box = [float(x) for x in re.split(r"[\s,]+", attributes.get("viewbox", "").strip())]
                      valid_box = len(box) == 4 and all(isfinite(x) for x in box) and box[2] > 0 and box[3] > 0
                  except (ValueError, AttributeError):
                      valid_box = False
                  if not valid_box:
                      errors.append(f"Figure {index} SVG has missing or invalid viewbox.")
                  if attributes.get("role") != "img":
                      errors.append(f"Figure {index} SVG requires role=img.")
                  if not (attributes.get("aria-label") or "").strip():
                      errors.append(f"Figure {index} SVG requires a non-empty aria-label= value.")
      
          media = list(re.finditer(r"<(?:img|image)\b[^>]*>", md_text, flags=re.IGNORECASE))
          media += list(re.finditer(r"!\[[^\]]*\]\([^)]+\)", md_text))
          media += list(re.finditer(r"<svg\b[^>]*>", md_text, flags=re.IGNORECASE))
          for match in media:
              if not _inside(match.start(), ranges):
                  errors.append("An image or SVG appears outside a <figure> element.")
      
          if base_dir is not None:
              media_parser = MediaReferences()
              media_parser.feed(md_text)
              refs = media_parser.refs
              refs += re.findall(r"!\[[^\]]*\]\(([^)\s]+)", md_text)
              for ref in sorted(set(refs)):
                  candidate = local_path(ref, base_dir)
                  if candidate is not None and not candidate.is_file():
                      errors.append(f"Image file not found: {ref}")
      
          return errors
      
      
      def validate_markdown(
          md_text: str,
          base_dir: Path | None = None,
          legacy_schema: bool = False,
      ) -> tuple[list[str], list[str]]:
          errors: list[str] = []
          warnings: list[str] = []
      
          first_content = next((line for line in md_text.splitlines() if line.strip()), "")
          h1 = re.findall(r"^#\s+.+$", md_text, flags=re.MULTILINE)
          if len(h1) != 1 or first_content != h1[0]:
              errors.append("The report must start with exactly one H1 title.")
          errors.extend(_validate_metadata(md_text))
          if legacy_schema:
              warnings.append("Legacy schema mode: evidence contract 3 is not enforced.")
          elif not re.search(r"^>\s*证据契约:3\s*$", md_text, re.M):
              errors.append("New reports require > 证据契约:3; use --legacy-schema only for historical reports.")
      
          main_sections = re.findall(
              r"^##\s+([一二三四五六七八九十]+)、.+$",
              md_text,
              flags=re.MULTILINE,
          )
          valid_sequences = [REQUIRED_MAIN_SECTIONS, [*REQUIRED_MAIN_SECTIONS, "五"]]
          if legacy_schema and main_sections not in valid_sequences:
              errors.append(
                  "Main sections must be 一 through 四, with 五 optional. "
                  f"Found: {main_sections or 'none'}."
              )
      
          placeholders: list[str] = []
          for pattern in PLACEHOLDER_PATTERNS:
              placeholders.extend(re.findall(pattern, md_text, flags=re.IGNORECASE | re.MULTILINE))
          if placeholders:
              errors.append("Unresolved template placeholders were found.")
          if re.search(r"```\s*mermaid\b", md_text, flags=re.IGNORECASE):
              errors.append("Unrendered Mermaid source found.")
      
          body_text, appendix_text = _split_body_and_appendix(md_text)
          if not appendix_text:
              errors.append("The report has no appendix.")
      
          a1_header, a1_rows, a1_malformed = _extract_table(appendix_text, "A1")
          a2_header, a2_rows, a2_malformed = _extract_table(appendix_text, "A2")
          if len(a1_header) != 4 or not a1_rows:
              errors.append("A1 must contain a four-column source ledger with data rows.")
          if a1_malformed:
              errors.append(f"A1 has {a1_malformed} malformed data row(s).")
          if len(a2_header) != 6 or not a2_rows:
              errors.append("A2 must contain a six-column Claim ledger with data rows.")
          if a2_malformed:
              errors.append(f"A2 has {a2_malformed} malformed data row(s).")
      
          source_errors, source_ids = _validate_sources(a1_rows, legacy_schema)
          claim_errors, claim_source_ids = _validate_claims(a2_rows, source_ids, a2_header, legacy_schema)
          errors.extend(source_errors)
          errors.extend(claim_errors)
      
          body_source_ids = set(re.findall(r"\[(S\d{2,})\]", body_text))
          all_sources = set(re.findall(r"\[(S\d{2,})\]", md_text))
          if all_sources - source_ids:
              errors.append("Report references undefined Source IDs: " + ", ".join(sorted(all_sources - source_ids)))
          for heading, content in re.findall(r"^(##\s+[^\n]+)\n(.*?)(?=^##\s+|\Z)", body_text, re.M | re.S):
              content = re.sub(r"^#{3,6}\s+.*$", "", content, flags=re.M)
              if not _is_filled(content):
                  errors.append(f"Empty analysis section: {heading}")
          if not body_source_ids:
              errors.append("The report body has no [Sxx] citations.")
          undefined_body_sources = body_source_ids - source_ids
          if undefined_body_sources:
              errors.append(
                  "Report body references undefined Source IDs: "
                  + ", ".join(sorted(undefined_body_sources))
                  + "."
              )
          unused_sources = source_ids - body_source_ids - claim_source_ids
          if unused_sources:
              warnings.append(
                  "Source ledger entries are unused: " + ", ".join(sorted(unused_sources)) + "."
              )
      
          a3_text = _section_text(appendix_text, "A3")
          if a3_text is None or not _is_filled(a3_text):
              errors.append("A3 must describe evidence boundaries.")
      
          errors.extend(_validate_figures(md_text, base_dir))
          return errors, warnings
      
      
      def parse_args() -> argparse.Namespace:
          parser = argparse.ArgumentParser(description="Validate a 3d-deep-research report.")
          parser.add_argument("markdown", help="Markdown report path")
          parser.add_argument("--pdf", help="Optional rendered PDF path")
          parser.add_argument("--html", help="Optional rendered HTML path")
          parser.add_argument("--legacy-schema", action="store_true", help="Read historical evidence ledgers; does not certify current contract")
          return parser.parse_args()
      
      
      def main() -> None:
          _configure_utf8_console()
          args = parse_args()
          markdown_path = Path(args.markdown).expanduser().resolve()
          if not markdown_path.is_file():
              raise SystemExit(f"Markdown report not found: {markdown_path}")
      
          md_text = markdown_path.read_text(encoding="utf-8")
          errors, warnings = validate_markdown(md_text, base_dir=markdown_path.parent, legacy_schema=args.legacy_schema)
          if args.pdf:
              pdf_errors, pdf_warnings = validate_pdf(Path(args.pdf).expanduser().resolve())
              errors.extend(pdf_errors)
              warnings.extend(pdf_warnings)
          for artifact in (args.pdf, args.html):
              if artifact:
                  if args.legacy_schema:
                      warnings.append("Legacy artifact identity is not verified; render again for current-contract delivery.")
                  else:
                      identity_errors, identity_warnings = validate_manifest(markdown_path, Path(artifact).expanduser().resolve())
                      errors.extend(identity_errors)
                      warnings.extend(identity_warnings)
          for error in errors:
              print(f"[ERROR] {error}")
          for warning in warnings:
              print(f"[WARN] {warning}")
          passed = not errors
          print("[OK] Report validation passed." if passed else "[FAIL] Report validation failed.")
          raise SystemExit(0 if passed else 1)
      
      
      if __name__ == "__main__":
          main()
      
  • tests
    • test_evidence_contract.py 8.6 KB
      from __future__ import annotations
      
      import base64
      import sys
      import tempfile
      import unittest
      from pathlib import Path
      
      sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "scripts"))
      from report_artifacts import manifest_path, sha256, validate_manifest, write_manifest
      from render_report import _resolve_local_media
      from test_validate_report import VALID_REPORT, validate_report
      
      
      class EvidenceContractTests(unittest.TestCase):
          def check_rejected(self, report: str, expected: str) -> None:
              errors, _ = validate_report.validate_markdown(report)
              self.assertTrue(any(expected in error for error in errors), errors)
      
          def test_confidence_does_not_substitute_for_gap(self):
              row = next(line for line in VALID_REPORT.splitlines() if line.startswith("| C01"))
              cells = row.split("|")
              cells[-2] = " high "
              self.check_rejected(VALID_REPORT.replace(row, "|".join(cells)), "缺口")
      
          def test_explicit_no_known_gap_is_valid(self):
              report = VALID_REPORT.replace("缺口:缺少原位储量测量", "缺口:无重大已知缺口")
              self.assertEqual(validate_report.validate_markdown(report)[0], [])
      
          def test_confidence_and_type_are_single_values(self):
              self.check_rejected(VALID_REPORT.replace("置信度:high", "置信度:high / low"), "confidence level")
              self.check_rejected(VALID_REPORT.replace("| fact |", "| fact / causal |"), "exactly one type")
      
          def test_english_independence_can_touch_chinese(self):
              report = VALID_REPORT.replace("independent", "来源为shared-origin且非独立复验")
              self.assertEqual(validate_report.validate_markdown(report)[0], [])
              self.assertFalse(validate_report._contains_enum("non-independent", {"independent"}))
      
          def test_invalid_source_date_is_rejected(self):
              self.check_rejected(VALID_REPORT.replace("2026-08-01", "2026-99-99"), "invalid ISO date")
      
          def test_missing_publisher_is_rejected(self):
              self.check_rejected(VALID_REPORT.replace("发布者:NASA;", ""), "发布者")
      
          def test_unknown_publication_is_explicitly_allowed(self):
              self.assertEqual(validate_report.validate_markdown(VALID_REPORT.replace("发布:2026-08-01", "发布:未知"))[0], [])
      
          def test_offline_unknown_date_does_not_need_invented_date(self):
              report = VALID_REPORT.replace("发布:2026-08-01;访问:2026-08-29", "发布:未知;访问:不适用")
              self.assertEqual(validate_report.validate_markdown(report)[0], [])
      
          def test_missing_contract_requires_explicit_legacy_mode(self):
              report = VALID_REPORT.replace("> 证据契约:3\n", "")
              self.check_rejected(report, "证据契约")
              errors, warnings = validate_report.validate_markdown(report, legacy_schema=True)
              self.assertEqual(errors, [])
              self.assertTrue(any("Legacy" in warning for warning in warnings))
      
          def test_process_and_attribution_are_independent(self):
              report = VALID_REPORT.replace("| fact |", "| mechanism;过程:部分连接已观察;归因:结果贡献未知 |")
              self.assertEqual(validate_report.validate_markdown(report)[0], [])
              self.check_rejected(report.replace(";归因:结果贡献未知", ""), "归因")
      
          def test_undefined_appendix_reference_is_rejected(self):
              self.check_rejected(VALID_REPORT + "\n附加限制:[S99]\n", "undefined Source IDs")
      
          def test_empty_chapter_is_rejected(self):
              self.check_rejected(VALID_REPORT.replace("早期观测推动了后续直接测量。[S01]", ""), "Empty analysis section")
      
          def test_question_led_sections_need_no_chronology(self):
              appendix = VALID_REPORT[VALID_REPORT.index("## 附录"):]
              metadata = VALID_REPORT[:VALID_REPORT.index("## 一")]
              report = metadata + "## 哪种观测更可靠\n\n比较口径与测量。[S01]\n\n## 误差如何产生\n\n误差边界。[S01]\n\n" + appendix
              self.assertEqual(validate_report.validate_markdown(report)[0], [])
      
          def test_missing_svg_image_is_rejected(self):
              figure = '<figure><svg viewBox="0 0 10 10" role="img" aria-label="test"><image href="absent.png"/></svg><figcaption>图 1 [S01]</figcaption></figure>'
              with tempfile.TemporaryDirectory() as directory:
                  errors, _ = validate_report.validate_markdown(VALID_REPORT + figure, Path(directory))
              self.assertTrue(any("Image file not found" in error for error in errors), errors)
      
          def test_substring_attributes_do_not_satisfy_svg_contract(self):
              report = VALID_REPORT + '<figure><svg data-viewbox="0 0 10 10" role="presentation" aria-label=""></svg><figcaption>图 1 [S01]</figcaption></figure>'
              for expected in ("viewbox", "role=", "aria-label="):
                  self.check_rejected(report, expected)
      
          def test_portable_media_preserves_local_image_bytes(self):
              with tempfile.TemporaryDirectory() as directory:
                  root = Path(directory)
                  data = b"test-image-bytes"
                  (root / "chart.png").write_bytes(data)
                  output = _resolve_local_media('<img src="chart.png">', root, embed=True)
                  self.assertNotIn("file:///", output)
                  self.assertIn(base64.b64encode(data).decode(), output)
                  (root / "chart.png").unlink()
                  self.assertIn("data:image/png;base64,", output)
      
          def test_manifest_binds_source_output_and_dependencies(self):
              with tempfile.TemporaryDirectory() as directory:
                  root = Path(directory)
                  source, output, resource = [root / name for name in ("report.md", "output.pdf", "figure.svg")]
                  source.write_text("version A")
                  output.write_bytes(b"artifact A")
                  resource.write_text("figure A")
                  write_manifest(source, output, [resource], [])
                  self.assertEqual(validate_manifest(source, output), ([], []))
                  source.write_text("version B")
                  self.assertIn("different Markdown", " ".join(validate_manifest(source, output)[0]))
                  source.write_text("version A")
                  output.write_bytes(b"artifact B")
                  self.assertIn("does not match", " ".join(validate_manifest(source, output)[0]))
                  output.write_bytes(b"artifact A")
                  resource.write_text("figure B")
                  self.assertIn("dependency", " ".join(validate_manifest(source, output)[0]))
      
          def test_missing_manifest_and_remote_resource_boundary(self):
              with tempfile.TemporaryDirectory() as directory:
                  root = Path(directory)
                  source, output = root / "report.md", root / "output.html"
                  source.write_text("source")
                  output.write_text("output")
                  self.assertTrue(validate_manifest(source, output)[0])
                  write_manifest(source, output, [], ["https://example.com/image.png"])
                  errors, warnings = validate_manifest(source, output)
                  self.assertEqual(errors, [])
                  self.assertTrue(warnings)
                  manifest_path(output).write_text("{}")
                  self.assertTrue(validate_manifest(source, output)[0])
      
          def test_edit_during_render_cannot_certify_wrong_input(self):
              with tempfile.TemporaryDirectory() as directory:
                  root = Path(directory)
                  source, output = root / "report.md", root / "output.pdf"
                  source.write_text("input A")
                  expected = {source: sha256(source)}
                  source.write_text("input B")
                  output.write_bytes(b"render of input A")
                  with self.assertRaisesRegex(ValueError, "changed during rendering"):
                      write_manifest(source, output, [], [], expected)
      
          def test_pdf_font_names_alone_do_not_prove_embedding(self):
              try:
                  from pypdf import PdfWriter
                  from pypdf.generic import DictionaryObject, NameObject
              except ImportError:
                  self.skipTest("pypdf is needed for the PDF resource regression")
              with tempfile.TemporaryDirectory() as directory:
                  writer = PdfWriter()
                  page = writer.add_blank_page(width=595.28, height=841.89)
                  fonts = DictionaryObject()
                  for index, name in enumerate(validate_report.REQUIRED_PDF_FONTS):
                      fonts[NameObject(f"/F{index}")] = DictionaryObject({
                          NameObject("/Type"): NameObject("/Font"),
                          NameObject("/Subtype"): NameObject("/Type1"),
                          NameObject("/BaseFont"): NameObject("/" + name),
                      })
                  page[NameObject("/Resources")] = DictionaryObject({NameObject("/Font"): fonts})
                  path = Path(directory) / "names-only.pdf"
                  writer.write(path)
                  self.assertEqual(validate_report._read_pdf(path)[3], set())
      
      
      if __name__ == "__main__":
          unittest.main()
      
    • test_linkify_sources.py 2.1 KB
      from __future__ import annotations
      
      import importlib.util
      import unittest
      from pathlib import Path
      
      
      SCRIPT = Path(__file__).resolve().parents[1] / "scripts" / "linkify_sources.py"
      SPEC = importlib.util.spec_from_file_location("linkify_sources", SCRIPT)
      assert SPEC and SPEC.loader
      linkify_sources = importlib.util.module_from_spec(SPEC)
      SPEC.loader.exec_module(linkify_sources)
      
      
      class LinkifySourcesTests(unittest.TestCase):
          def test_ledger_url_links_citations_and_adds_anchor(self) -> None:
              value = """<p>结论 [S01]</p>
      <table><tr><td>S01</td><td><a href="https://example.com/source">来源</a></td></tr></table>"""
      
              output, links, rows = linkify_sources.linkify_html(value)
      
              self.assertIn('id="src-S01"', output)
              self.assertIn('href="https://example.com/source"', output)
              self.assertEqual(links, 1)
              self.assertEqual(rows, 1)
      
          def test_protected_contexts_and_ledger_are_not_relinked(self) -> None:
              value = """<p>[S01]</p><a href="#existing">[S01]</a><code>[S01]</code>
      <pre>[S01]</pre><table><tr><td>S01</td><td>[S01]</td></tr></table>"""
      
              output, links, _ = linkify_sources.linkify_html(value)
      
              self.assertEqual(links, 1)
              self.assertIn('<a href="#existing">[S01]</a>', output)
              self.assertIn('<code>[S01]</code>', output)
              self.assertIn('<pre>[S01]</pre>', output)
      
          def test_missing_source_url_uses_ledger_anchor(self) -> None:
              value = "<p>结论 [S01]</p><table><tr><td>S01</td><td>离线材料</td></tr></table>"
      
              output, links, rows = linkify_sources.linkify_html(value)
      
              self.assertIn('href="#src-S01"', output)
              self.assertEqual((links, rows), (1, 1))
      
          def test_second_pass_is_idempotent(self) -> None:
              value = "<p>结论 [S01]</p><table><tr><td>S01</td><td>离线材料</td></tr></table>"
              first, _, _ = linkify_sources.linkify_html(value)
      
              second, links, rows = linkify_sources.linkify_html(first)
      
              self.assertEqual(second, first)
              self.assertEqual(links, 0)
              self.assertEqual(rows, 1)
      
      
      if __name__ == "__main__":
          unittest.main()
      
    • test_render_report.py 3 KB
      from __future__ import annotations
      
      import importlib.util
      import sys
      import tempfile
      import unittest
      from pathlib import Path
      from unittest import mock
      
      
      SCRIPT_DIR = Path(__file__).resolve().parents[1] / "scripts"
      sys.path.insert(0, str(SCRIPT_DIR))
      SPEC = importlib.util.spec_from_file_location("render_report", SCRIPT_DIR / "render_report.py")
      assert SPEC and SPEC.loader
      render_report = importlib.util.module_from_spec(SPEC)
      SPEC.loader.exec_module(render_report)
      
      
      class RenderReportTests(unittest.TestCase):
          def test_build_html_uses_report_title_and_metadata(self) -> None:
              markdown = """# 自定义报告
      
      > 研究问题:测试 | 资料截止:2026-08-28 | 完成日期:2026-08-29
      
      正文。[S01]
      """
              with mock.patch.object(
                  render_report,
                  "markdown_to_html",
                  return_value=("<p>正文。[S01]</p>", "test-converter"),
              ) as converter:
                  document, title, name = render_report.build_html(markdown)
      
              body_md = converter.call_args.args[0]
              self.assertEqual(title, "自定义报告")
              self.assertEqual(name, "test-converter")
              self.assertIn("研究问题:测试", document)
              self.assertNotIn("研究问题:测试", body_md)
              self.assertNotIn("立体分析法深度研究报告", document)
              self.assertNotIn("作者:jeffy", document)
      
          def test_relative_media_resolves_against_markdown_directory(self) -> None:
              with tempfile.TemporaryDirectory() as directory:
                  base = Path(directory)
                  value = """<img src="images/chart.png">
      <svg><image href='assets/map.svg'></image></svg>
      <img src="https://example.com/remote.png"><img src="data:image/png;base64,abc">
      <svg><image href="#symbol"></image></svg>"""
      
                  output = render_report._resolve_local_media(value, base)
      
              self.assertIn((base / "images/chart.png").resolve().as_uri(), output)
              self.assertIn((base / "assets/map.svg").resolve().as_uri(), output)
              self.assertIn('src="https://example.com/remote.png"', output)
              self.assertIn('src="data:image/png;base64,abc"', output)
              self.assertIn('href="#symbol"', output)
      
          def test_pdf_fonts_are_explicitly_embedded(self) -> None:
              with tempfile.TemporaryDirectory() as directory:
                  font_dir = Path(directory)
                  for filename in render_report.FONT_FILES.values():
                      (font_dir / filename).write_bytes(b"font")
      
                  css = render_report._font_face_css(font_dir)
      
              self.assertIn("NotoSansCJKsc-Regular.otf", css)
              self.assertIn("NotoSansCJKsc-Bold.otf", css)
              self.assertIn("font-weight: 400", css)
              self.assertIn("font-weight: 700", css)
      
          def test_missing_pdf_font_fails(self) -> None:
              with tempfile.TemporaryDirectory() as directory:
                  with self.assertRaisesRegex(RuntimeError, "Missing required Noto CJK font"):
                      render_report._font_face_css(Path(directory))
      
      
      if __name__ == "__main__":
          unittest.main()
      
    • test_validate_report.py 13.2 KB
      from __future__ import annotations
      
      import importlib.util
      import re
      import sys
      import unittest
      from pathlib import Path
      
      
      SCRIPT = Path(__file__).resolve().parents[1] / "scripts" / "validate_report.py"
      sys.path.insert(0, str(SCRIPT.parent))
      SPEC = importlib.util.spec_from_file_location("validate_report", SCRIPT)
      assert SPEC and SPEC.loader
      validate_report = importlib.util.module_from_spec(SPEC)
      SPEC.loader.exec_module(validate_report)
      
      
      VALID_REPORT = """# 月球研究报告
      
      > 研究问题:月球是否存在水冰 | 资料截止:2026-08-28 | 完成日期:2026-08-29
      
      > 证据契约:3
      
      ## 一、核心结论
      
      公开测量支持月球极区存在水冰。[S01]
      
      ## 二、为什么会走到今天
      
      早期观测推动了后续直接测量。[S01]
      
      ## 三、哪些力量改变了路径
      
      探测能力和任务目标共同改变了证据质量。[S01]
      
      ## 四、它为什么这样运转
      
      永久阴影区允许挥发物长期保存。[S01]
      
      ## 五、对当前问题意味着什么
      
      后续任务需要直接测量储量和分布。[S01]
      
      ## 附录:来源与证据边界
      
      ### A1 来源账本
      
      | Source ID | 来源与日期 | 证据作用 | 限制 |
      |---|---|---|---|
      | S01 | [NASA fact sheet](https://example.com/moon);发布者:NASA;发布:2026-08-01;访问:2026-08-29 | 支持 C01;原始材料;independent | 仅覆盖公开测量 |
      
      ### A2 关键判断与证据
      
      | Claim ID | 可证伪判断 | 类型/证据范围 | 支持证据与独立性 | 替代解释/反向证据 | 置信度、缺口与反证条件 |
      |---|---|---|---|---|---|
      | C01 | 月球极区存在水冰 | fact | 支持 S01;independent | 反向检索:检查任务更正和相反测量,未发现 | 置信度:high;缺口:缺少原位储量测量;修订条件:若后续原位测量不支持则修改判断 |
      
      ### A3 资料边界
      
      现有证据不能确定水冰的完整储量和开采难度。
      """
      
      
      class ValidateReportTests(unittest.TestCase):
          def test_current_contract_passes(self) -> None:
              errors, warnings = validate_report.validate_markdown(VALID_REPORT)
      
              self.assertEqual(errors, [])
              self.assertEqual(warnings, [])
      
          def test_optional_fifth_section_can_be_removed(self) -> None:
              report = VALID_REPORT.replace(
                  "## 五、对当前问题意味着什么\n\n后续任务需要直接测量储量和分布。[S01]\n\n",
                  "",
              )
      
              errors, _ = validate_report.validate_markdown(report)
      
              self.assertEqual(errors, [])
      
          def test_required_structure_rejects_missing_parts(self) -> None:
              cases = {
                  "H1": (
                      VALID_REPORT.replace("# 月球研究报告", "月球研究报告", 1),
                      "exactly one H1",
                  ),
                  "A1 shape": (
                      VALID_REPORT.replace(
                          "| Source ID | 来源与日期 | 证据作用 | 限制 |",
                          "| Source ID | 来源与日期 | 证据作用 |",
                          1,
                      ),
                      "A1 must",
                  ),
                  "A2 shape": (
                      VALID_REPORT.replace(
                          "| Claim ID | 可证伪判断 | 类型/证据范围 | 支持证据与独立性 | 替代解释/反向证据 | 置信度、缺口与反证条件 |",
                          "| Claim ID | 可证伪判断 | 支持证据与独立性 | 替代解释/反向证据 | 置信度、缺口与反证条件 |",
                          1,
                      ),
                      "A2 must",
                  ),
                  "A3": (
                      VALID_REPORT.replace(
                          "### A3 资料边界\n\n现有证据不能确定水冰的完整储量和开采难度。\n",
                          "",
                          1,
                      ),
                      "A3 must",
                  ),
              }
      
              for name, (report, expected) in cases.items():
                  with self.subTest(name=name):
                      errors, _ = validate_report.validate_markdown(report)
                      self.assertTrue(any(expected in error for error in errors), errors)
      
          def test_metadata_is_required_and_valid(self) -> None:
              cases = {
                  "missing": VALID_REPORT.replace(
                      "> 研究问题:月球是否存在水冰 | 资料截止:2026-08-28 | 完成日期:2026-08-29\n\n",
                      "",
                      1,
                  ),
                  "malformed": VALID_REPORT.replace("研究问题:", "问题:", 1),
                  "invalid date": VALID_REPORT.replace("2026-08-28", "2026-02-30", 1),
                  "reversed dates": VALID_REPORT.replace("完成日期:2026-08-29", "完成日期:2026-08-27", 1),
              }
      
              for name, report in cases.items():
                  with self.subTest(name=name):
                      errors, _ = validate_report.validate_markdown(report)
                      self.assertTrue(any("metadata" in error.lower() or "date" in error.lower() for error in errors), errors)
      
          def test_enum_values_require_exact_tokens(self) -> None:
              cases = {
                  "english substring": VALID_REPORT.replace("置信度:high", "置信度:follow-up", 1),
                  "chinese substring": VALID_REPORT.replace("置信度:high", "置信度:中立", 1),
              }
      
              for name, report in cases.items():
                  with self.subTest(name=name):
                      errors, _ = validate_report.validate_markdown(report)
                      self.assertTrue(any("confidence level" in error for error in errors), errors)
      
          def test_non_independent_marker_is_valid(self) -> None:
              report = VALID_REPORT.replace("independent", "非独立")
      
              errors, _ = validate_report.validate_markdown(report)
      
              self.assertEqual(errors, [])
      
          def test_legacy_claim_ledger_is_still_valid(self) -> None:
              report = VALID_REPORT.replace(
                  "| Claim ID | 可证伪判断 | 类型/证据范围 | 支持证据与独立性 | 替代解释/反向证据 | 置信度、缺口与反证条件 |",
                  "| Claim ID | 关键判断 | 判断类型 | 支持与反向证据 | 置信度与独立性 | 缺口与反证条件 |",
                  1,
              ).replace(
                  "| C01 | 月球极区存在水冰 | fact | 支持 S01;independent | 反向检索:检查任务更正和相反测量,未发现 | 置信度:high;缺口:缺少原位储量测量;修订条件:若后续原位测量不支持则修改判断 |",
                  "| C01 | 月球极区存在水冰 | fact | 支持 S01;反向检索:检查任务更正和相反测量,未发现 | high;independent | 缺少原位储量测量;若后续原位测量不支持则修改判断 |",
                  1,
              )
      
              errors, _ = validate_report.validate_markdown(report, legacy_schema=True)
      
              self.assertEqual(errors, [])
      
          def test_new_schema_requires_process_and_attribution(self) -> None:
              report = VALID_REPORT.replace("| fact |", "| mechanism |", 1)
      
              errors, _ = validate_report.validate_markdown(report)
      
              self.assertTrue(any("mechanism evidence scope" in error for error in errors), errors)
      
              report_with_status = VALID_REPORT.replace(
                  "| fact |", "| mechanism;过程:已有原位观测;归因:无法判断全部储量 |", 1
              )
              errors, _ = validate_report.validate_markdown(report_with_status)
      
              self.assertEqual(errors, [])
      
          def test_extra_numbered_section_is_allowed(self) -> None:
              report = VALID_REPORT.replace(
                  "## 附录:来源与证据边界",
                  "## 六、额外章节\n\n问题需要的额外分析。[S01]\n\n## 附录:来源与证据边界",
                  1,
              )
      
              errors, _ = validate_report.validate_markdown(report)
      
              self.assertEqual(errors, [])
      
          def test_ledgers_must_be_inside_appendix(self) -> None:
              before_appendix, appendix = VALID_REPORT.split("## 附录:来源与证据边界", 1)
              report = before_appendix + appendix + "\n\n## 附录:来源与证据边界\n"
      
              errors, _ = validate_report.validate_markdown(report)
      
              self.assertTrue(any("A1 must" in error for error in errors), errors)
              self.assertTrue(any("A2 must" in error for error in errors), errors)
      
          def test_unresolved_template_placeholder_fails(self) -> None:
              report = VALID_REPORT.replace("# 月球研究报告", "# [研究对象]深度研究报告")
      
              errors, _ = validate_report.validate_markdown(report)
      
              self.assertTrue(any("placeholder" in error.lower() for error in errors), errors)
      
          def test_duplicate_source_id_fails(self) -> None:
              original = (
                  "| S01 | [NASA fact sheet](https://example.com/moon);发布者:NASA;发布:2026-08-01;访问:2026-08-29 | "
                  "支持 C01;原始材料;independent | 仅覆盖公开测量 |"
              )
              duplicate = (
                  "| S01 | [Duplicate](https://example.com/duplicate);发布者:NASA;发布:2026-08-02;访问:不适用 | "
                  "补充 C01;independent | 仅覆盖摘要 |"
              )
              report = VALID_REPORT.replace(original, original + "\n" + duplicate)
      
              errors, _ = validate_report.validate_markdown(report)
      
              self.assertTrue(any("duplicate Source ID S01" in error for error in errors), errors)
      
          def test_dangling_claim_source_fails(self) -> None:
              report = VALID_REPORT.replace("支持 S01;independent", "支持 S99;independent", 1)
      
              errors, _ = validate_report.validate_markdown(report)
      
              self.assertTrue(any("undefined Source IDs: S99" in error for error in errors), errors)
      
          def test_claim_requires_reverse_search_note(self) -> None:
              report = VALID_REPORT.replace(
                  "反向检索:检查任务更正和相反测量,未发现",
                  " ",
              )
      
              errors, _ = validate_report.validate_markdown(report)
      
              self.assertTrue(any("counterevidence" in error for error in errors), errors)
      
          def test_body_citation_must_resolve(self) -> None:
              report = VALID_REPORT.replace("公开测量支持月球极区存在水冰。[S01]", "公开测量支持月球极区存在水冰。[S99]")
      
              errors, _ = validate_report.validate_markdown(report)
      
              self.assertTrue(any("Report body references undefined" in error for error in errors), errors)
      
          def test_media_must_follow_figure_contract(self) -> None:
              report = VALID_REPORT.replace(
                  "## 二、为什么会走到今天",
                  "![无标题图片](missing.png)\n\n## 二、为什么会走到今天",
              )
      
              errors, _ = validate_report.validate_markdown(report, base_dir=Path("/tmp"))
      
              self.assertTrue(any("outside a <figure>" in error for error in errors), errors)
              self.assertTrue(any("Image file not found" in error for error in errors), errors)
      
          def test_figure_contract(self) -> None:
              figure = """<figure>
      <svg viewBox="0 0 10 10" role="img" aria-label="月球水冰示意">
        <circle cx="5" cy="5" r="4"></circle>
      </svg>
      <figcaption>图 1:月球极区水冰证据 [S01]</figcaption>
      </figure>"""
      
              def add_figure(value: str) -> str:
                  return VALID_REPORT.replace("## 二、为什么会走到今天", value + "\n\n## 二、为什么会走到今天", 1)
      
              errors, _ = validate_report.validate_markdown(add_figure(figure))
              self.assertEqual(errors, [])
      
              cases = {
                  "figcaption": (re.sub(r"<figcaption>.*?</figcaption>", "", figure), "figcaption"),
                  "Source ID": (figure.replace(" [S01]", "", 1), "no Source ID"),
                  "viewBox": (figure.replace(' viewBox="0 0 10 10"', "", 1), "viewbox"),
                  "role": (figure.replace(' role="img"', "", 1), "role="),
                  "aria-label": (figure.replace(' aria-label="月球水冰示意"', "", 1), "aria-label="),
              }
              for name, (broken_figure, expected) in cases.items():
                  with self.subTest(name=name):
                      errors, _ = validate_report.validate_markdown(add_figure(broken_figure))
                      self.assertTrue(any(expected in error for error in errors), errors)
      
          def test_pdf_layout_accepts_a4_portrait(self) -> None:
              errors, warnings = validate_report._validate_pdf_data(
                  "可提取的正文" * 20,
                  [(595.28, 841.89), (595.28, 841.89)],
                  "WeasyPrint 69.0",
                  {"Noto-Sans-CJK-SC", "Noto-Sans-CJK-SC-Bold"},
              )
      
              self.assertEqual(errors, [])
              self.assertEqual(warnings, [])
      
          def test_pdf_layout_rejects_landscape_and_literal_html(self) -> None:
              errors, _ = validate_report._validate_pdf_data(
                  "研究时间<br/>资料截止" + "正文" * 50,
                  [(595.28, 841.89), (841.89, 595.28)],
                  "WeasyPrint 69.0",
                  {"Noto-Sans-CJK-SC", "Noto-Sans-CJK-SC-Bold"},
              )
      
              self.assertTrue(any("page 2 is landscape" in error for error in errors), errors)
              self.assertTrue(any("literal HTML tag" in error for error in errors), errors)
      
          def test_pdf_layout_rejects_wrong_engine_and_fonts(self) -> None:
              errors, _ = validate_report._validate_pdf_data(
                  "可提取的正文" * 20,
                  [(595.28, 841.89)],
                  "ReportLab PDF Library",
                  {"Helvetica"},
              )
      
              self.assertTrue(any("expected WeasyPrint 69.0" in error for error in errors), errors)
              self.assertTrue(any("required Noto CJK fonts" in error for error in errors), errors)
      
      
      if __name__ == "__main__":
          unittest.main()
      
  • SKILL.md 11.6 KB
    ---
    name: 3d-deep-research
    description: |
      用证据链和 X/Y/Z 立体分析法研究产品、公司、技术、概念、人物、行业、市场或复杂事件,交付可追溯的深度研究报告。用户要求 deep research、系统调研、竞品或市场研究、尽职调查、来龙去脉分析、证据链或正式研究报告时使用。简单名词解释、新闻摘要、短篇观点、仿写,以及 3D 建模、渲染、CAD 或图形设计不使用。
    ---
    
    # 3D Deep Research
    
    先建立研究底图,确认研究对象的边界、运转方式和当前状态;再用 X/Y/Z 逐层解释:X 识别显性与非显性的关键变化及其后续影响,Y 沿具体变化切开并比较成因,Z 拆开尚不清楚的作用连接。分析结果持续返回发展路径和研究底图,用证据校正解释并检查重要遗漏。
    
    本 Skill 为主研究流程时,默认交付完整的 Markdown、HTML 和 PDF 三份报告;用户明确缩小交付范围时按其要求减少。用户点名其他研究 Skill 或选择其他工作模式时,以其作为主流程;仅在需要本方法时辅助使用,不叠加独立默认产物。不要仅因通用的 deep research 字样同时运行两套流程,也不要把普通问答扩写成长报告。
    
    ## 必读资源
    
    执行完整研究时:
    
    1. 读取 [references/evidence-protocol.md](references/evidence-protocol.md),建立来源与 Claim 账本并执行事实审计。
    2. 读取 [references/xyz-method.md](references/xyz-method.md),建立研究底图并执行 X/Y/Z、路径校正和覆盖复核。
    3. 根据研究对象读取 [references/object-adapters.md](references/object-adapters.md) 的对应部分。
    4. 读取 [references/visual-guidelines.md](references/visual-guidelines.md),用于分析阶段的视觉规划、HTML/PDF 渲染和图表检查。
    5. 写作时复制 [assets/report-template.md](assets/report-template.md),保留元数据与证据附录;正文可按问题重组。
    
    ## 报告措辞
    
    以下规则适用于正文、标题和图注;证据附录以准确、可追溯为先。
    
    - 先直接回答研究问题,再解释事实、原因、限制和行动含义。
    - **先具体,后抽象。** 优先说清谁做了什么、发生什么变化,再解释背后的概念。能直接说明的,不用抽象名词代替。
    - **先判断,后解释。** 一句话先说清主要意思,原因和必要限制随后展开,避免把多层意思挤进一句话。
    - **用自然、准确的书面语解释给非专业读者听。** 写完后检查:读者是否需要把这句话“翻译成白话”才能理解?如果需要,就改写,同时保留因果关系与证据边界。
    - 每段只推进一个判断。删除不提供事实、解释或决策价值的句子。
    - 明确区分事实、解释和预测。证据只能支持“可能”时,不写成“证明”或“必然”。
    - 不把 Claim 类型、证据门槛、X/Y/Z 等内部方法术语写进正文;它们只用于组织研究和附录说明。
    - 数字首次出现时说明时间、单位、统计口径和比较对象。
    - 不使用宣传式形容词、无依据的最高级,以及“值得注意的是”“不容忽视”等空洞过渡语。
    - 只在不确定性会影响结论时说明限制,并紧邻相关判断书写。
    
    ## 执行流程
    
    ### 0. 明确研究设定
    
    记录研究对象、对象类型、用户要做的决策、特别关注点、时间基准、范围边界和交付要求。无法从上下文解决且会显著改变目标或正确性的问题才追问;其余情况直接开始。根据对象适配器建立研究底图,先确认主体边界、价值或作用结构、关键参与者及当前状态,再选择解释主线。
    
    复杂或持续时间较长的任务可以把上述设定记录为工作笔记;不要为简单任务强制创建独立过程文件。
    
    研究范围由用户问题和证据决定,不追求来源数、Claim 数、字数或图表数量。每个进入 A2 的关键判断和数字都必须通过适用的证据门槛并完成审计;资料不足时缩小可确认结论的范围,保留未解决问题,不静默改变用户的研究目标。
    
    涉及“最新、现在、最近”时联网核实,并记录发布日期和访问日期。输出到用户指定位置;未指定时使用当前项目的 `output/`。
    
    ### 1. 规划检索
    
    先列出:
    
    1. 建立研究底图必须确认的事实和口径;
    2. 关键变化及其前后状态;
    3. 需要比较的成因和需要打开的作用连接;
    4. 需要主动寻找的反向证据、替代解释和失败案例。
    
    检索过程中根据证据调整问题。每个影响结论的问题最终都要有来源支持,或在 A2/A3 明确标为未解决。复杂任务可以维护检索笔记,但不强制生成独立表格。来源选择遵循 evidence protocol,优先使用适合该对象的原始材料、官方记录和独立来源。
    
    ### 2. 维护证据账本
    
    在 `report.md` 附录 A1 和 A2 分别维护来源与 Claim 账本。来源使用稳定 ID(`S01`、`S02`),关键判断使用 Claim ID(`C01`、`C02`)。把“来源出处”和“证据作用”分开记录。
    
    研究过程中持续更新账本。每条关键判断的支持证据、替代解释、反向材料、独立性、置信度、资料缺口和修订条件,按 [references/evidence-protocol.md](references/evidence-protocol.md) 记录和判断;因果和机制判断分别记录过程证据与归因边界。过程已发生不等于它足以解释总体结果。
    
    证据不足时交付“已确认部分 + 资料缺口 + 下一步验证路径”。因果或机制证据不完整时降级表述,不补写猜测,也不因局部缺口停止整份交付。
    
    ### 3. 建立底图并执行 X/Y/Z 分析
    
    先用研究底图横向看清对象,避免只研究容易形成主线的材料。再建立暂定的关键变化路径,把可观察变化、状态改变和后续影响分开;围绕影响核心判断的节点,比较成因并拆解尚不清楚的作用过程。根据证据返回修正节点、原因和路径,不要求一次完成 X 后再单向完成 Y、Z。
    
    复杂研究可在已有工作笔记中维护节点、待解释问题、因素、作用连接和证据状态的对应关系,不要求生成独立过程文件,也不把内部编号写入正文。X 只呈现原因摘要,Y 承担原因比较,Z 必须补出中间环节、成立条件或可观察后果,避免三部分重复叙述。详细方法和未来表达见 [references/xyz-method.md](references/xyz-method.md)。
    
    正文中的重要判断使用 `[S01]` 形式引用来源。最终解释需要说明哪些因素通过什么作用过程造成关键变化、变化如何影响后续发展,以及判断最可能错在哪个证据缺口或前提。深度由解释缺口决定,不按固定节点数或拆解层数执行。
    
    在分析阶段识别支撑核心结论的关键比较、演化路径、主体关系和机制,并规划相应的视觉表达,随证据更新调整。不要等正文完成后才考虑配图;可在已有研究笔记中简记,无需新增独立账本。
    
    ### 4. 写作与视觉表达
    
    使用 [assets/report-template.md](assets/report-template.md) 的元数据与必要附录。正文先回答问题并建立对象底图,再按问题组织解释。模板的演化式目录是默认示例;横向比较或稳定机制研究可以围绕具体问题连续展开观察、原因和机制,不必补写无关历史。重组目录不免除证据审计、解释追踪和双重回返;删除空章节及没有解释增量的重复。
    
    关键关系用图能显著减少读者对照、记忆或推演负担时,应制作相应分析图;数量由需要解释的关系决定。不能以正文已有描述或缺少量化数据为由省略有价值的图。图、表与文字各取所长,避免重复和装饰。量化图必须有可靠且可比较的数据;具体触发条件、证据要求和图表契约见 [references/visual-guidelines.md](references/visual-guidelines.md)。
    
    ### 5. 执行事实审计
    
    结构校验不能证明事实。验证前按 evidence protocol 执行归属审计和数字复核:
    
    1. 回到来源原文,确认每个关键判断能够由所引材料推出;
    2. 复算正文数字,检查单位、币种、时间和统计口径;
    3. 确认反向材料与正文中的限制一致,没有被删除或弱化;
    4. 检查 X 中的可观察变化没有被候选解释替代,Y 中的因素都指向明确节点,Z 中的机制没有停在术语或同义扩写;
    5. 返回研究底图,检查商业模式、经济性、竞争、治理、制度责任或其他对象适配器提示的重要侧面是否因主线过强而被静默删除;
    6. 把审计结果反映到正文、A2 和 A3。
    
    审计失败时改写正文,或降级、删除 Claim;不修改账本去迁就结论。
    
    ### 6. 验证与交付
    
    运行唯一的报告校验入口:
    
    ```bash
    python [skill目录]/scripts/validate_report.py report.md
    ```
    
    `validate_report.py` 只检查机器可验证的结构和证据一致性,不证明外部事实真实,也不判断是否遗漏了有价值的图。交付前按 visual guidelines 复核核心关系的视觉表达;全文无图时必须执行零图复核。
    
    校验通过后默认生成 HTML 和 PDF,并校验 PDF:
    
    ```bash
    python [skill目录]/scripts/render_report.py report.md output.html
    python [skill目录]/scripts/render_report.py report.md output.pdf
    python [skill目录]/scripts/validate_report.py report.md --html output.html --pdf output.pdf
    ```
    
    渲染前在 `report.md` 同级的 `fonts/` 放置 `NotoSansCJKsc-Regular.otf` 和 `NotoSansCJKsc-Bold.otf`,并安装 `WeasyPrint 69.0`。渲染器会检查并嵌入这两种字体。
    
    新报告使用模板中的 `证据契约:3`。校验历史报告时显式加 `--legacy-schema`,其通过只表示旧契约通过,不证明现行证据要求或产物一致性已满足。新交付不得用兼容模式绕过失败。渲染器为每个产物生成同名 `.manifest.json`,核对 Markdown、产物与本地资源摘要;保留这些构建凭据以防误发旧版本,正文或资源变更后重新渲染。摘要检查用于版本一致性,不证明事实或版面正确。
    
    PDF 只通过本 Skill 的 `render_report.py` 生成,不另写 ReportLab 或其他 Markdown-to-PDF 实现,也不绕过 [assets/report.css](assets/report.css)。渲染依赖或字体不可用时说明阻塞,不切换到其他引擎或排版实现降级交付。完成可行且已授权的依赖修复;PDF 仍阻塞时,提供已完成并通过适用检查的 Markdown,以及能够合规生成的 HTML,明确 PDF 尚未完成,不把部分交付称为完整交付。
    
    标题取报告第一行 H1。渲染器会把 `[Sxx]` 引用转为可点击锚点。交付前按 visual guidelines 检查输出,并通读正文。本 Skill 为主流程时,完整交付包含 `report.md`、`output.html` 和 `output.pdf`;用户明确缩小范围时按其要求,PDF 阻塞时按上一段提供部分成果并标明未完成项。
    
    ## 质量红线
    
    - 不把新闻排序当作因果链,也不把力量分类表当作 Y 轴。
    - 不把待验证的解释当作节点事实,不用心理揣测代替机制。
    - 不把合理机制直接当作本案例中已证实的原因,不用术语或同义扩写冒充机制拆解。
    - 不让单一解释主线挤掉会改变结论的重要业务、经济性、竞争、治理或制度侧面。
    - 不用同一原始材料的转载数量冒充独立证据。
    - 不隐藏冲突、样本偏差、访问失败、资料缺口或反向材料。
    - 不生成没有数据口径的数字图,不交付未经复核的数字。
    - 不留下模板占位符、未渲染图表或无法追溯的来源。
    - 不把过期证据当作当前状态。
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related