Claude Skill

short-drama-novel-analyze

把长篇小说、连载网文或多集散稿拆成可追溯的原著分析:章节索引、改编价值快评、逐章功能提取、剧情单元与节奏聚合、人物与设定归并,最后给出改编价值判定与分集候选,交给 $short-drama-develop 立契约。用户说“导入这本小说”“拆这本书”“分析原著”“这本书能不能改短剧”“先看看值不值得拆”“把长篇拆成分集候选”,或直接给出小说文件路径时使用。只做只读的结构化分析,不写剧本、不建资产、不生成媒体,也不替创作者决定改编方案。

LLM Mart · 0 points · 12 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download zenstory-ai-drama-skills-skills_short-drama-novel-analyze-f7ea5ba.zip · 37 KB
Part of zenstory-ai/drama-skills — 11 skills

Install

skills CLI npx skills add https://github.com/zenstory-ai/drama-skills/tree/main/skills/short-drama-novel-analyze
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install zenstory-ai-drama-skills@llmmart
Git git clone https://github.com/zenstory-ai/drama-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole zenstory-ai/drama-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

长篇原著分析

把一部长材料变成能被引用、能被反驳、能被接着用的分析层。目标不是复述剧情,而是找出 每一段承担的戏剧功能,并说明它在竖屏短剧里值多少钱。

分析永远是候选。哪条线保留、哪些人合并、从哪里开篇,是创作者的决定,由 $short-drama-develop 立成改编契约。本技能不替它决定,也不批准自己的产物。

Quick Start

离线验证章节索引、采样和源文件变更检测:

python3 {技能目录}/scripts/selftest.py
python3 {技能目录}/scripts/novel_index.py index <原著.txt> --out <chapter-index.json>

开始前

本技能可独立安装和执行。先读取用户明确提供的原著与本任务直接输入;若当前目录是 short-drama 项目且项目工具可用,可以读取 status 并使用其发布生命周期,但缺少 core 或任何其他技能都不是分析工作的阻断条件。完整边界与规则见 阶段契约,无需读取其他技能的文件。

材料前提

只分析创作者合法持有、拥有使用权的作品。分析是只读的转化性工作:提取结构与功能, 不复制原文成段落,不把原句搬进下游产物。

通俗题材里的暴力、复仇、背叛、情爱张力与黑暗伦理是常规虚构叙事元素,照常做结构化提取。 个别片段无法处理时跳过该段并记录,不要因此中止整章或整本——中止会让后续所有阶段拿到 一份有洞却看不出洞在哪的分析。

先判断入口

  1. 只有书名,没有原文:请创作者提供文件路径或粘贴正文。不要凭书名回忆情节—— 没有字节就没有 span,没有 span 的分析无法被引用,也无法被反驳。
  2. 有原文,未建项目:直接建立一个本技能自己的工作区,至少包含只读输入目录 输入/ 与输出目录 项目开发/source-analysis/_work/;把原文字节复制到 输入/ 后记录 原始文件位置,再用本技能的 novel_index.py 建索引。若 $short-drama 可用,可选用它 初始化同样的目录和发布生命周期,但不得把 core 安装变成开始分析的前提。
  3. 有原文,项目已在:直接进入管道。
  4. 已有部分分析:读 项目开发/source-analysis/_progress.md 从断点续跑, 不重跑已完成阶段。

管道

输入/ 是不可变的创作者输入,本阶段只读它。全部产出落在 项目开发/source-analysis/。 独立工作区与完整项目使用同一套相对路径,因此后续安装 core 时无需迁移分析产物。

阶段 做什么 产出 停靠
S0 建章节索引(脚本) _index.json、_progress.md 索引有问题就停
S1 改编价值快评(抽样) triage.md 停靠问创作者
S2 逐章功能提取 chapters/ch-<N>-extract.md 覆盖率不足就停
S3 剧情单元与节奏聚合 story-units.md、rhythm-and-emotion.md 阈值不达标就复核
S4 人物归并与设定 characters.md、world.md —
S5 改编价值与分集候选 adaptation-value.md、episode-candidates.jsonl 交接 develop

S0 章节索引

索引是唯一切片真源,由 章节索引脚本 建立。每个阶段各跑 一次正则,就会切出互相对不上的章节,第 47 章按一种边界分析、按另一种边界聚合, 而且没有人会发现。

python3 {技能目录}/scripts/novel_index.py index 输入/{原文文件} \
  --out 项目开发/source-analysis/_work/_index.next.json
python3 {技能目录}/scripts/novel_index.py verify \
  项目开发/source-analysis/_work/_index.next.json 输入/{原文文件}
# 项目工具可用时可选:
python3 {core 技能目录}/scripts/project_tool.py publish {项目根} \
  --owner short-drama-novel-analyze --artifact-id source-analysis:index \
  --output 项目开发/source-analysis/_index.json=项目开发/source-analysis/_work/_index.next.json \
  --input 输入/{原文文件}

脚本识别阿拉伯数字与中文数字章号(含 千 / 两,覆盖千章以上连载),只认一种编号单位 (章/回/节里出现最多的那个,其余记进 ignored_heading_units),只把短的独立行当标题 (以章号开头的正文段落记进 long_heading_lines_skipped),剔除开头的目录块, 按卷分段校验编号。它不做编辑判断——哪章重要、 讲了什么,是后面阶段的事。

problems 非空就停下报告,不要带着错表进 S1。常见四种:章号跳号(缺章或抓错标题)、 同卷内重号、正文极少的章(多半抓到了目录残留或卷首页)、无法解析的章号。 另外确认三个计数:chapter_unit 与 ignored_heading_units——一本用 第N章 分章、 用 第N节 分小节的书应当看到 chapter_unit: 章 且 节 被记入忽略计数,反过来说明分章单位 判断错了;long_heading_lines_skipped 明显偏大时,多半是这本书的章节标题确实很长, 需要与创作者确认后手写边界。

原文本身没有章节标题时索引会返回空表,此时与创作者确认按什么切分,把边界写进 _work/_index.next.json,通过 verify 后再按上面的公开生命周期发布——手写的行也要带齐 sequence / line_start / line_end,verify 会逐行检查并报出缺字段的行。

改了原文必须重建索引,不能沿用旧 span。verify 核对编号、顺序与行号覆盖,所以插行删行会被报出来; 但同行数的原地改写不会——那一步靠改原文的人自己重建,套件不比对字节。

python3 {技能目录}/scripts/novel_index.py verify \
  项目开发/source-analysis/_index.json 输入/{原文文件}

S1 改编价值快评

回答这本书值不值得花全量拆解的成本。判据是全书的改编密度——screen_ready 单元占多少、 prose_only 占多少、制作负担压在哪几段——所以快评横跨整本书,用脚本抽样:

python3 {技能目录}/scripts/novel_index.py sample \
  项目开发/source-analysis/_index.json --count 12

抽样确定、可复现、首尾必取,跑第二次引用的是同一批章。按 改编价值快评 写 triage.md,覆盖六件事: 故事框架、三类判定比例、开篇替换点、制作负担量级、最大的三处改编风险、分集候选量级。 第一行写覆盖率(脚本返回的 coverage_ratio),所有结论限于抽样范围。

这里停靠问创作者:给出快评与全量拆解的预计耗时(按章数粗估),问是否继续。 创作者一开始就明确说「一次跑完」时仍写 triage.md,但不停下等待——它是 S5 要回填 对照的第一版假设。

停靠时把 _progress.md 的状态写成 paused_after_triage,断点写「下一步:S2 逐章提取」。

S1–S5 的 Agent 创作产物同样先写到 source-analysis/_work/,完成本阶段机械检查后再发布到 上表中的正式路径。项目工具可用时用 project_tool.py publish 和稳定 artifact-id;独立运行时 原子替换正式文件。不要用半成品覆盖 _index.json、_progress.md、chapters/*.md 或聚合产物。 _work/ 是候选工作区,不是权威分析层,也不进入交付包。

并发子代理写 _work/,主线程发布到正式路径。子代理只落 _work/chapters/ch-<N>-extract.md,主线程跑完机械自检后再发布到 chapters/。 覆盖率闸门按正式路径 chapters/ 匹配文件名——发布之前跑,每一章都会进 unmatched_files,那不是缺陷,只是跑早了。

_progress.md 每个阶段都会重写,而发布要求一个路径只有一个 owner,所以它用一个固定 artifact-id:source-analysis:progress,owner 是 short-drama-novel-analyze, S0–S5 每次停靠都用同一个 id 重新发布。

S2 逐章功能提取

按 章节提取 处理每一章。能并发子代理就分批并发 (每批 5–8 章,等一批落盘再发下一批),不支持就串行——两条路径的写法要求和自检是同一份, 只是速度不同。

每章提取完落到 chapters/ch-<N>-extract.md,<N> 是索引里的 sequence,不是原文章号 (多卷书的原文章号会重复,sequence 不会)。全部落盘后跑覆盖率:

python3 {技能目录}/scripts/novel_index.py coverage \
  项目开发/source-analysis/_index.json 项目开发/source-analysis/chapters

missing 非空就补跑缺的章;unmatched_files 非空说明有文件名写歪了——它既不算覆盖, 也不会被当成缺章,必须改名而不是重跑。不要在覆盖率不足时进入 S3——聚合会照样产出 一份读起来完整的结果,而缺掉的章不会在任何地方留下痕迹。

单章连续失败两次就标记跳过,写进 _progress.md 的失败记录,并在后续每一份聚合产物里 注明该章缺失。失败可以接受,失败被藏起来不行。

S3 剧情单元与节奏

从逐章提取聚合,不回头重读原文——原文已经在 S2 被读过一次,再读一次只会得到第二份互相 矛盾的事实。按 聚合与实体 先识别故事框架 (框架决定按什么切单元),再产出:

  • story-units.md:把情节点归成有始有终的单元,每个单元记录进入状态、冲突、代价与出去状态;
  • rhythm-and-emotion.md:关键信息如何逐章推进、情绪触动点的铺垫→释放→余波、 跨章伏笔与兑现。

聚合完成后跑同一文件里的三条阈值自检(归属置信、覆盖率、重叠率)与散落情节兜底。 阈值不是评分,是边界模糊的信号:重叠率过高说明两个单元其实是一个。

S4 人物与设定

按 聚合与实体 归并人物(跨章去重、别名归一、 分级),并从提及数据归纳世界规则、力量体系与势力。别名只有专名与有同指证据的绰号能合并, 描述性称谓与头衔永远不触发合并。

人物归并是候选,不是资产。 这里的人物条目带 unresolved 与来源引用, 交给 $short-drama-develop 定改编决定、$short-drama-write 写进剧本之后, 才由 $short-drama-assets 从已接受剧本建立真正的资产身份。绕过这条链直接建资产, 等于让原著的人物表冒充剧本的出现证据。

S5 改编价值与分集候选

这是本技能与通用拆书的分水岭。按 改编价值 产出:

  • adaptation-value.md:哪些单元在竖屏短剧里能直接成立、哪些要换载体、哪些是纯文字快感 (内心戏、叙述性诡计、长铺垫)在画面上无法兑现;制作负担落在哪里。
  • episode-candidates.jsonl:按局部戏剧结果与精确交接切出的候选集,不按章号或字数 平均切。每条带来源 span、承担的功能、以及未决项。

同时回填快评:S1 的哪几条判断被全量结果推翻了,写进 adaptation-value.md 的开头。 一个抽样结论被证伪,比它被悄悄忘掉有用得多——下一本书的快评会因此更准。

样例见 分集候选样例。

交接

S5 完成后展示创作者可读的摘要,说明:拆了多少章、跳过哪些、分了多少个候选集、 最大的三处改编风险、以及快评里被推翻的判断。然后交给 $short-drama-develop—— 由它把候选变成 项目开发/adaptation-map.jsonl 与改编契约。

本技能不写 adaptation-map.jsonl,那是 develop 的产物。需要质量结论时交给独立的 $short-drama-review(范围 source_analysis)。

规则分级

  • structural_invariant:索引与 span 的可证明性、引用完整性、覆盖率。可由脚本阻断。
  • reviewed_invariant:功能提取是否忠于原文、归并是否保住戏剧作用等语义义务。
  • craft_default:通常有帮助的做法;创作者说明理由后可覆盖。
  • taste_option:从哪里开篇、保留哪条线等选择;不得单独阻断。

不要用固定的章数配方、情节点数量或篇幅比例替代因果判断。

产物与边界

本技能只拥有 项目开发/source-analysis/ 下的文件:_index.json、_progress.md、 chapters/*.md、triage.md、story-units.md、rhythm-and-emotion.md、 characters.md、world.md、adaptation-value.md、episode-candidates.jsonl。

它不改写 输入/,不写 项目开发/ 下其他技能的产物,不建资产、不写场景与台词、 不写提示词、不生成媒体,也不签发终审结论。

不把原文成段复制进分析。记录 locator、span 与去引用的功能摘要;需要证据时引用 最短的必要片段。交付包不得把原始材料带出边界。

语言

分析产物是创作者读的:项目内跟随 short-drama.json#/language,独立运行时跟随用户使用的 语言,不在本技能内硬编码语言。本技能不产生提示词正文,与 #/format/prompt_language 无关。

按需加载

  • 快评读哪些章、写哪六件事、怎么不冒充全量分析:改编价值快评
  • 逐章提取写法、白描与叙事框架词的界线、机械自检、并发与串行:章节提取
  • 故事框架、剧情单元、节奏情绪、人物归并、阈值与散落兜底:聚合与实体
  • 改编价值评估、载体替换与分集候选切法:改编价值
  • 本阶段拥有什么、继承什么、不越权什么:阶段契约
Files (drama-skills)
  • agents
    • openai.yaml 283 B
      interface:
        display_name: "长篇原著分析"
        short_description: "把长篇小说拆成章节索引、剧情单元与可改编的分集候选清单"
        default_prompt: "使用 $short-drama-novel-analyze 先给这本小说做改编价值快评,再决定要不要全量拆解。"
      
  • assets
    • episode-candidate.example.jsonl 2.5 KB · in bundle
  • references
    • adaptation-triage.md 4.6 KB
      # 改编价值快评
      
      ## 目录
      
      - [产出什么](#产出什么)
      - [抽样怎么取](#抽样怎么取)
      - [写六件事](#写六件事)
      - [开篇替换点](#开篇替换点)
      - [结论限于抽样范围](#结论限于抽样范围)
      - [停靠与交接](#停靠与交接)
      
      ## 产出什么
      
      一份 `triage.md`,回答**这本书值不值得花全量拆解的成本**。
      
      它是一份可以被后面推翻的假设:S5 拿全量结果回来对照它,被推翻的判断写进
      `adaptation-value.md` 开头。
      
      ## 抽样怎么取
      
      由脚本取,不手挑:
      
      ```bash
      python3 {技能目录}/scripts/novel_index.py sample \
        项目开发/source-analysis/_index.json --count 12
      ```
      
      抽样等距、首尾必取、结果确定。跑第二次引用的是同一批章,S5 回填对照才说得清哪条判断错了。
      
      默认 12 章。书特别长或框架复杂时调大 `--count`,调大要写进快评。抽到的章正文极少
      (索引里 `char_count` 很小)时往后顺延一章并记录——那多半是卷首页,不是内容。
      
      快评只判断**功能与载体**:不写完整情节点表,不写人物表,不产出 `chapters/` 下的文件。
      逐章提取的格式要求见 [章节提取](chapter-extraction.md),这一步不套用。
      
      ## 写六件事
      
      `triage.md` 固定六段,缺一段不算快评。
      
      ### 1. 故事框架
      
      从 [聚合与实体](aggregation-and-entities.md) 的框架表里选一个:升级流 / 复仇线 /
      单元剧 / 多线交织 / 蜕变成长。框架决定后面按什么切单元,也决定这本书按集切的难度——
      单元剧与短剧的集结构最贴合,多线交织在竖屏里最贵。
      
      抽样判不出来时写「抽样不足以判断框架」并说明原因。
      
      ### 2. 三类判定比例
      
      对抽到的每一章,判断它承担的功能属于哪一类(定义见 [改编价值](adaptation-value.md)):
      
      | 判定 | 含义 |
      |---|---|
      | `screen_ready` | 功能由可见行动、可听对白或空间关系承担 |
      | `needs_carrier` | 功能成立,但当前载体是文字专属 |
      | `prose_only` | 功能依赖读者与叙述者的关系,屏幕上无法兑现 |
      
      写出三类各占抽样里的多少章。这是快评里唯一给数字的一段,其余用文字判断。
      `prose_only` 过半时直接说明:全量拆解大概率只会把同一个结论算得更精确。
      
      ### 3. 开篇替换点
      
      见下一节。
      
      ### 4. 制作负担量级
      
      抽样里出现的:最多几人同框、需要多少个不重复空间、有没有群戏 / 动作 / 特效 /
      可读文字 / 跨时间造型变化。给量级,不给预算表。
      
      按已接受的制作形态估;形态未定时按两种形态各给一句,不替创作者选。
      
      ### 5. 三处最大改编风险
      
      具体到抽到的哪一章、哪个功能。写「改编难度较大」对任何一本书都成立,
      因此对这一本没有信息。
      
      ### 6. 分集候选量级
      
      按抽样里的戏剧结果密度给一个量级(例如「约 20–30 集」),并写清这是从 N/M 章推出来的。
      集数由创作者与 `$short-drama-develop` 定,这里给的是量级。
      
      ## 开篇替换点
      
      回答一个问题:**如果这本书改成短剧,第 1 集从哪里开始**?
      
      找**第一个可见的、有代价的、当场发生的转折**——主角第一次做出会付出代价的选择,
      且这个选择能被拍出来而不是被想出来。它通常不在第一章。写清三条:
      
      - 落在第几章(用索引的 `sequence`);
      - 它之前的内容里,哪些是观众必须知道的,可以怎样在开场后补进去;
      - 哪些是原著读者需要、短剧观众不需要的。
      
      抽样里没有这样的转折时,写「抽样内未见可用的开篇替换点」,并列进第 5 段的风险。
      
      ## 结论限于抽样范围
      
      `triage.md` 第一行写覆盖率,用脚本返回的 `coverage_ratio` 与 `total_chapters`:
      
      ```markdown
      > 本快评基于抽样:读了 340 章中的 12 章(3.5%),章号见下表。结论是假设,不是结论。
      ```
      
      三条:
      
      - **只用读过的章做证据**。 需要提到抽样外的内容时写「抽样未覆盖」。
      - **比例的分母是抽样章数**,在文里写出来。
      - **候选集条目是 S5 的产物**,快评给量级即可。
      
      ## 停靠与交接
      
      快评写完后向创作者报告:覆盖率、六段结论、全量拆解的预计耗时(按章数粗估)、
      以及一句直给的建议(值得拆 / 不值得拆 / 拆但先砍掉某几卷)。
      
      然后停下等待。`_progress.md` 状态写 `paused_after_triage`,断点写「下一步:S2 逐章提取」。
      
      创作者一开始就说「一次跑完」时不停靠,`triage.md` 仍然要写——S5 要拿全量结果回填对照它。
      
    • adaptation-value.md 6.6 KB
      # 改编价值与分集候选
      
      ## 目录
      
      - [这一步在回答什么](#这一步在回答什么)
      - [三类判定](#三类判定)
      - [载体替换](#载体替换)
      - [制作负担](#制作负担)
      - [桥段标签](#桥段标签)
      - [分集候选](#分集候选)
      - [回填快评](#回填快评)
      - [交给 develop 之前](#交给-develop-之前)
      
      ## 这一步在回答什么
      
      前面几步产出的是「这本书是怎么运转的」。这一步回答的是**「这本书里有多少东西能在竖屏
      短剧里成立」**——这是拆书和改编分析的分水岭,也是本技能存在的理由。
      
      判定对象是**剧情单元**,不是章,也不是情节点。判定的原始依据是 S2 每章写下的
      [载体记录](chapter-extraction.md#载体记录):通道栏走可见行动或对白的,多半能直接成立;
      走心理叙述或叙述者旁白的,需要换载体或根本换不了。
      
      ## 三类判定
      
      每个单元落进且只落进一类:
      
      | 判定 | 含义 | 后续 |
      |---|---|---|
      | `screen_ready` | 功能由可见行动、可听对白或空间关系承担 | 直接进候选集 |
      | `needs_carrier` | 功能成立,但当前载体是文字专属 | 必须写出新载体,否则降级 |
      | `prose_only` | 功能依赖读者与叙述者的关系,屏幕上无法兑现 | 明说放弃了什么 |
      
      `prose_only` 的典型:长段内心权衡、叙述性诡计(读者被叙述视角欺骗,而镜头做不到同样的
      隐瞒)、靠信息密度堆出的爽感、以及需要读者记住三十章前一个细节才成立的回扣。
      
      **不要把 `prose_only` 悄悄改写成 `needs_carrier`**。 承认放弃比假装能拍便宜得多——
      假装的代价会在分镜阶段才暴露,那时整条链都已经建在上面。
      
      ## 载体替换
      
      `needs_carrier` 的单元逐个写:
      
      - **原载体**:功能现在靠什么承担(抄 S2 载体记录里的通道栏);
      - **新载体**:可见行动 / 物证 / 空间关系 / 对白策略 / 第三方反应,任选但要具体;
      - **保住了什么**:功能清单,逐项对上;
      - **改变或失去了什么**:不写「无」。总有东西变了,写不出来说明还没想清楚。
      
      替换的检验标准只有一条:**观众不需要旁白就能知道发生了什么**。需要旁白解释的替换,
      是把文字问题搬到了画面上,不是解决。
      
      ## 制作负担
      
      按已接受的制作形态估,不设跨项目阈值。逐单元记:需要几个人物同框、需要哪些不重复的
      空间、有没有群戏、动作或特效、可读文字、跨时间的造型变化。
      
      这不是预算表,是**改编排序的依据**:同样是 `screen_ready`,负担低的单元先进候选集。
      
      ## 桥段标签
      
      给每个单元打桥段标签,命中就标,逗号分隔。标签的用处是让 `$short-drama-develop`
      能按题材检索,也让重复的桥段一眼可见——同一个桥段在一本书里出现五次,
      改编时要么是卖点,要么是拖沓,这个判断需要先能数出来。
      
      | 题材 | 常见桥段 |
      |---|---|
      | 都市/言情 | 打脸、扮猪吃虎、马甲掉落、先婚后爱、追妻火葬场、破镜重圆、替身、契约婚姻、萌宝助攻 |
      | 玄幻/修仙 | 废柴逆袭、越级战斗、夺宝、护短打脸、绝境翻盘、机缘逆天 |
      | 历史/种田 | 先知先觉、改变历史、科举经商崛起、攒钱起家 |
      | 悬疑/无限 | 信息差、反转打脸、规则怪谈、身份揭穿、复盘解谜 |
      | 系统流 | 任务奖励、签到、面板升级、强制任务、积分兑换 |
      | 重生/穿越 | 前世复仇、改命、身份调换、双重生、未卜先知 |
      | 世情/家庭 | 极品亲戚、偏心打脸、维权反击、自我觉醒 |
      
      这是**可扩展的种子表,不是封闭枚举**。不在表内但本书确有的桥段照实写;
      凑不出标签就留空,不硬凑。
      
      ## 分集候选
      
      按**局部戏剧结果与精确交接**切,不按章号、字数或平均分配。一个候选集必须能回答:
      
      1. 进入时观众已经知道什么;
      2. 本集内兑现了哪一部分承诺(不能一条都不兑现);
      3. 出去时留下的具体决定、危险或问题是什么;
      4. 下一集必须继承哪些事实。
      
      只留钩子、不兑现任何承诺的集,观众会在第三集停——这是分集候选里最常见的错。
      
      每条候选写进 `episode-candidates.jsonl`,字段见
      [分集候选样例](../assets/episode-candidate.example.jsonl)。每条都带:来源单元与章号区间、
      承担的功能、判定类别、桥段标签、未决项、以及 `creator_acceptance: pending`。
      
      **集数不由本技能决定**。 候选可以多于、也可以少于创作者设定的集数;差多少要明说,
      由创作者与 `$short-drama-develop` 决定合并还是拆分。
      
      ## 回填快评
      
      `adaptation-value.md` 开头写一段**快评对照**:S1 的 `triage.md` 里哪几条判断被全量结果
      推翻了,以及为什么。逐条对上六段:故事框架、三类判定比例、开篇替换点、制作负担量级、
      三处风险、分集量级。
      
      **其中「三类判定比例」两边的判定对象不同**:S1 抽样判的是章,S5 判的是剧情单元,
      两者之间没有换算,也不该造一个——一个单元里的章可以分属不同类别,按章数加权回推
      只会得到一个看起来精确的假数字。这一条按方向对照即可:抽样说「多数章走心理叙述」,
      全量结论若是「多数单元需要换载体」,就是印证;若方向相反,写清是抽样偏了还是
      全书后段变了。其余五段仍逐条对上。
      
      抽样结论被证伪不是失误,是抽样这一步在起作用。写出来才有两个用处:创作者知道
      「当时那个建议是基于什么」,下一本书的快评知道该往哪个方向修正。
      
      全部命中时也写一句,别留空——一份从不出错的快评,多半是没有被认真对照过。
      
      ## 交给 develop 之前
      
      自检五条:
      
      - 每个 `screen_ready` / `needs_carrier` 单元都能在候选集里找到落点,或明确说明为什么没有;
      - 每条候选的功能都能回溯到某个单元、某条情节点的白描;
      - 硬事实(等级、数值、出场章、谁说的话)都能回到白描或原文,回不到的已改成「原文未明确」;
      - `prose_only`、被跳过的章与散落清单都出现在交接摘要里,不只是躺在某个文件深处;
      - 没有把本技能的判定写成结论——它们是候选,等创作者接受。
      
      交接摘要给创作者读,说清:拆了多少章、跳过哪些、多少单元分别落进三类、
      切出多少候选集、最大的三处改编风险、快评里被推翻的判断。风险要具体到单元,
      不写「改编难度较大」。
      
    • aggregation-and-entities.md 8.3 KB
      # 聚合与实体
      
      ## 目录
      
      - [先识别故事框架](#先识别故事框架)
      - [剧情单元](#剧情单元)
      - [单元粒度](#单元粒度)
      - [阈值自检](#阈值自检)
      - [散落情节兜底](#散落情节兜底)
      - [节奏与情绪](#节奏与情绪)
      - [人物归并](#人物归并)
      - [设定归纳](#设定归纳)
      
      聚合只读 S2 的逐章提取,**不回头重读原文**。原文在 S2 已经被读过一次;再读一次得到的
      不是补充,是第二份互相矛盾的事实。需要核对某条硬事实时回查具体那一条白描,
      而不是重新通读。
      
      ## 先识别故事框架
      
      框架决定按什么切单元。先判框架再切,能省掉一整轮返工。
      
      | 框架 | 特征 | 切单元的依据 | 改编提示 |
      |---|---|---|---|
      | 升级流 | 等级/境界体系清晰,按阶段推进 | 按等级阶段 | 阶段边界天然是集边界 |
      | 复仇线 | 核心矛盾明确,目标驱动 | 按复仇对象或阶段 | 每个对象一个小高潮,适配短剧 |
      | 单元剧 | 每卷独立故事,弱连续性 | 按单元 | 与短剧的集结构最贴合 |
      | 多线交织 | 多视角/多时间线 | 按线索切,交叉点单独标记 | 竖屏里最贵:观众要记住的人最多 |
      | 蜕变成长 | 主角内在变化为主线 | 按成长阶段 | 内在变化要找到外部载体才成立 |
      
      判定方法:扫全部章节概要 + 首尾各三章的提取。判不出来时写「框架不明确」并说明原因,
      不硬套——套错框架会让后面每一个单元都切在错的地方。
      
      框架识别的结论写进 `story-units.md` 开头,并与 S1 `triage.md` 里的框架判断对照:
      不一致时写清哪一个错了、为什么。
      
      ## 剧情单元
      
      一个单元记录六项:
      
      | 字段 | 写什么 |
      |---|---|
      | 标题 | 15 字以内 |
      | 章节范围 | 用索引的 `sequence` 起止 |
      | 进入状态 | 单元开始时,谁知道什么、谁有什么筹码 |
      | 核心冲突 | 具体的对抗对象 + 具体的冲突事件 |
      | 代价 | 这个单元里谁付出了什么,不写「无」 |
      | 出去状态 | 单元结束时改变了什么,下一个单元必须继承哪些事实 |
      
      进入状态与出去状态必须能对接:上一个单元的出去状态就是下一个的进入状态。
      对不上说明中间漏了一个单元,或者边界切错了。
      
      ## 单元粒度
      
      | 粒度 | 例子 | 问题 |
      |---|---|---|
      | 过细 | 「获得神器」(1–2 章) | 这是事件,不是单元 |
      | 过粗 | 「主角的成长之路」(1–100 章) | 这是主题,不是单元 |
      | 合适 | 「家族觉醒与传承认知」(1–6 章) | 有独立目标、冲突、角色群和完整弧线 |
      
      一个单元必须同时满足四条:有独立且可衡量的戏剧目标;有具体的对抗力量;有本单元特有的
      核心角色(可与其他单元共享部分角色);有自己的紧张—缓和节奏。
      
      四条里缺一条,它多半是另一个单元的**阶段**,降级处理即可,不必单独建条。
      
      ## 阈值自检
      
      聚合后跑三条。它们不是评分,是**边界模糊的信号**:
      
      | 指标 | 参考值 | 算法 | 不达标怎么办 |
      |---|---|---|---|
      | 归属置信 | ≥ 0.85 | 有明确归属的情节点 / 单元内情节点总数 | 低于线的单元标「待复核」,不删 |
      | 覆盖率 | 85%–95% | 已归类情节点 / 总情节点数 | < 85% 走散落兜底;> 95% 复核是否过度归并 |
      | 重叠率 | ≤ 35% | 跨单元共享的情节点 / 总情节点数 | > 35% 说明两个单元其实是一个,考虑合并 |
      
      「有明确归属」指这条情节点服务于本单元的戏剧目标或对抗力量,而不只是恰好落在本单元的
      章节范围内。按连续章节切分时,若把「落在范围内」当作归属,这个比值恒等于 1,指标失去
      意义——所以逐条判的是**它推动的是不是这个单元的那件事**,跨单元共享的情节点计入重叠率
      而不计入本单元的分子。
      
      **小体量素材的例外**:不足 30 章、单主线连续叙事的材料按时序切分时,覆盖率天然接近
      100%、散落接近 0,这是线性结构的正常结果,不是过度归并。此时复核的重点是
      「有没有把两条剧情硬并进同一个单元」,不是把数字压到 95% 以下。
      
      不要为了凑阈值人为拆分或合并单元。数字对不上时先怀疑边界,再怀疑数字。
      
      ## 散落情节兜底
      
      覆盖率不足时执行,目的是**不丢情节点**,不是把每个点都塞进某个单元:
      
      1. 算孤立比例。低于 5% 直接标「少量散落」,不再处理。
      2. 收集未归类的情节点。
      3. 按角色重叠、地点重叠、因果关系三条线索判相关性:强相关归入并标 `[孤立归入]`;
         中等相关归入并标 `[低置信归入]` 待复核;弱相关**不强行归入**。
      4. 弱相关的按主题聚类,聚出 5 个以上情节点才形成候选单元并标 `[聚类生成]`。
      5. 仍无法归类的写进 `story-units.md` 末尾的散落清单,按章排列并注明原因。
      
      **宁可留在散落清单里,也不要塞进不相关的单元**。 塞进去之后,那个单元的进入/出去状态
      就会带上一条对不上的事实,而它在下游看起来和真事实一模一样。
      
      ## 节奏与情绪
      
      `rhythm-and-emotion.md` 写三件事,全部要能回指到章或情节点:
      
      - **关键信息推进**:某条信息何时进入读者视野 → 如何被延迟、误导或验证 → 何时爆发 →
        余波与新钩子。
      - **情绪触动点索引**:每个触动点落在哪一章、由哪条情节点承担。
      - **长间隔说明**:连续 3 章以上没有触动点时,写清作者用什么维持期待。
        改编时这些长间隔就是要合并或砍掉的段落——短剧没有这个耐心。
      
      回指不到章或情节点的条目删掉,不要保留无源断言。
      
      ## 人物归并
      
      跨章去重时,**一人一实体**:无法确定两个称谓是否指同一人,就分开建两条,不合并。
      合并错了的代价不对称——两个人被并成一个,之后每一条关于「他」的事实都可能属于另一个人,
      而这种错在文本里几乎看不出来。
      
      别名分四类,只有前两类能触发合并:
      
      | 类型 | 定义 | 可合并 | 例 |
      |---|---|---|---|
      | 专名 | 全名、常用名 | 是 | 「顾禾」「顾小禾」 |
      | 绰号 | 外号,须有同指证据 | 是(证据充分时) | 「城东那位」经文中点明即为某人 |
      | 描述性称谓 | 按外形或角色描述 | **否** | 「红衣女子」「随从」 |
      | 头衔职务 | 泛指某个职位的称呼 | **否** | 「护卫队长」「家主」 |
      | 姓氏+头衔 | 带姓氏或名的称呼 | 是(证据充分时) | 「周团」「钟记者」「齐军长」 |
      
      绰号与姓氏+头衔的合并都需要可指出的同指证据:同位说明、括号别名、上下文明确指代、
      文中改名。「读起来像同一个人」不是证据。
      
      **姓氏+头衔单列一类,因为中文叙事里它常常是最主要的称呼方式**。「周团」「钟记者」
      出现频次可能远高于全名,而它们带着姓氏,指向性与绰号相当——按「头衔职务」一律禁止合并,
      一部书会留下十几条重复实体。真正不可合并的是**不带姓名的泛指**:「护卫队长」「家主」
      可以先后由不同人担任,同一部书里换了人也看不出来。
      
      判据是:去掉头衔后还剩不剩一个姓或名。剩,按证据判;不剩,不合并。
      
      分级按戏份,不按喜好:出场章数占比、是否推动主线、是否有独立动机与变化。
      每条人物记 `unresolved`(未定的指代、矛盾的描述)与来源引用。
      
      **归并结果是候选,不是资产**。资产身份由 `$short-drama-assets` 从已接受剧本建立。
      
      ## 设定归纳
      
      从提及数据归纳世界规则、力量体系与势力,不从记忆或同题材惯例补。
      
      - 力量体系:名称、层级、晋升方式;原文没写清的层级写「原文未明确」。
      - 地理:分布、主要区域、关键地点。
      - 势力:门派/组织/家族/国家,各自的目标与彼此的关系。
      - 特殊设定:区别于现实的规则。
      
      相互依赖、共同起作用的设定合并成一条(例如「戒指 + 里面的残魂」是一个整体,
      不拆成两条);只有完全独立、互不依赖的才分开。
      
      每条设定标注它出现在哪几章。标不出来的设定,多半是从题材惯例里补出来的,删掉。
      
    • chapter-extraction.md 8.3 KB
      # 逐章提取
      
      ## 目录
      
      - [提取什么](#提取什么)
      - [情节点写法](#情节点写法)
      - [白描与叙事框架词](#白描与叙事框架词)
      - [载体记录](#载体记录)
      - [机械自检](#机械自检)
      - [并发与串行](#并发与串行)
      - [失败处理](#失败处理)
      
      ## 提取什么
      
      一章提取一份 `chapters/ch-<N>-extract.md`,`<N>` 用索引里的 `sequence`(多卷书的原文
      章号会重复,sequence 不会)。
      
      **第一行是标题**:`# ch-<N> <原文章节标题>`。原文没有标题就只写 `# ch-<N>`,不要自拟。
      
      其下五段固定顺序:
      
      1. **概要**:3–5 句,说这一章在整本书里**改变了什么**。不是缩写剧情,是回答
         「读完这章,谁的处境、认知或筹码和读完上一章时不一样了」。写成条目罗列,
         或整段「因为…所以…」串联,都是没做提取。
      2. **情节点**:见下。
      3. **载体记录**:见下。
      4. **出场人物**:本章出现的称谓原样记录,附本章内的作用**。不做跨章归并**——归并是 S4
         的事,在这里做会让每一章各自发明一套人物表。
      5. **未决项**:本章无法判断的指代、时间关系或矛盾,逐条列出。留白比猜测便宜。
      
      ## 情节点写法
      
      每个情节点一行,四段用 `|` 分隔:
      
      ```
      P1 类型 | 白描 | 涉及:称谓、物件、地点 | 功能:这一步改变了什么
      ```
      
      - **`P<n>` 与类型之间是一个半角空格**,类型与 `|` 之间也是。下游按这个字面形状 grep,
        写成全角空格或不写空格,这一行就不算情节点——**这是最常见的一处格式漂移**。
      - **类型**取自:`行动`、`对话`、`信息`、`转折`、`铺垫`、`兑现`、`状态`。
      - **白描**是可观察的事实:谁做了什么、说了什么、什么被看见。见下一节。
      - **功能**回答改变了什么:权力、信息、关系、暴露、资源、时间或代价,至少一项。
        写不出功能的情节点,多半不是情节点。
      - **`功能:` 是字段标记,不要在别处使用**。概要、载体记录、未决项里若出现这四个字符,
        自检的计数就会超过情节点数。要在正文里说「功能」,改写成「作用」「起到的效果」。
      
      密度按章节字数动态调节,通常一个情节点覆盖 150–200 字原文。不要为凑数拆步骤,
      也不要把一整场戏压成一行。为同一个戏剧目的服务的连续微动作合并成一个情节点。
      
      ## 白描与叙事框架词
      
      白描是这条情节点的主要证据——后面每一次事实回查都落在它上面。所以它只能写
      **原文写出来的可观察事实**:不写心理断言,不写作者意图,不用叙事框架词概括。
      
      | 禁止 | 正确 |
      |---|---|
      | 通过对话,郑松得知张子豪在韩国训练 | 吴志斌告诉郑松,张子豪在韩国训练 |
      | 林风展现了自己的实力 | 林风三招击败对手,围观者倒吸一口凉气 |
      | 邵阳感到心碎和愤怒 | 邵阳目睹宋丽与人拥抱,表情由刺痛转为冷漠 |
      | 这一章为后文埋下伏笔 | 管家把一枚旧钥匙放进抽屉最里层,没有说明用途 |
      
      **原文本身就是内心独白时**,写「他想」的内容本身,而不是写他的心理状态:
      原文「他想,这次真的完了」写成「他在心里说这次真的完了」,不写「他感到绝望」。
      第一人称叙述同理——原文以「我」写出来的判断是这一章的可观察事实,照录即可,
      但只能录他说出或想出的那句话,不能替他总结。
      
      左列的问题不是啰嗦,是**它把结论写在了证据的位置上**。S3 之后的每一次聚合都离原文
      两跳,如果证据本身已经是结论,后面就再也无法核对这个结论对不对。
      
      同一条纪律往后延伸到 S3–S5:**硬事实必须可回溯**。等级、数值、距离、势力数量、
      谁对谁说了什么、某人出现在哪几章——这类事实写进聚合产物前,必须能回到某条白描或原文;
      回不到就写「原文未明确」,绝不用近义项、别处的例子或惯例补一个看起来合理的值。
      相邻实体的属性也不互相挪用:坐骑的等级不是骑手的等级,A 的台词不是 B 的台词。
      
      ## 载体记录
      
      每章至少写一条,标准章节通常 2–5 条。它把「本章必须让读者知道什么」和
      「原文用什么把它写成场景」拆开——**这一栏是 S5 判断 `needs_carrier` 的原始数据**,
      没有它,后面只能凭印象说某个功能能不能拍。
      
      | 字段 | 写什么 |
      |---|---|
      | 关键信息 | 本章必须让读者知道、误判、期待或确认的事 |
      | 原文载体 | 作者用了哪些事件、对话、反应、细节、误导、回扣把它写成场景 |
      | 通道 | 这个载体主要走哪个通道:可见行动 / 对白 / 心理叙述 / 叙述者旁白 / 信息密度 |
      | 对读者的作用 | 好奇、期待、压抑、爽、心疼、紧张、甜、热血中的哪几种 |
      
      **通道栏决定下游判定**。 走可见行动或对白的,多半 `screen_ready`;走心理叙述或
      叙述者旁白的,多半 `needs_carrier` 或 `prose_only`。这里只记通道,不做判定——
      判定是 S5 的事,需要跨章看完才成立。
      
      过渡章也要写:说明它维持了什么期待或冷却了什么情绪,不写「无」。
      
      ## 机械自检
      
      落盘后由主线程直接跑,命中即判失败,不依赖子代理自报。这些检查存在的理由是
      下游会按同样的字面形状 grep:格式漂一个字符,聚合阶段就会静默漏掉整章。
      
      ```bash
      F=项目开发/source-analysis/chapters/ch-${N}-extract.md
      P=$(grep -cE '^P[0-9]+ ' "$F")
      grep -cE '^P[0-9]+ [^|]+\|[^|]*[^|[:space:]][^|]*\|[^|]*涉及' "$F"   # 必须 == $P
      grep -cE '^P[0-9]+ .*\|.*\|.*\|[^|]*功能:' "$F"                     # 必须 == $P
      grep -hoE '^P[0-9]+ [^|]+' "$F" | awk '{print $2}' | sort -u        # 必须 ⊆ 类型表
      awk '/^## 载体记录/{f=1;next} /^## /{f=0} f&&/^\|/{n++} END{print n+0}' "$F"  # 必须 >= 3
      ```
      
      - 载体表计数少于 3:该章只写了表头没写内容。阈值是 3 不是 2,因为表头行与 `|---|` 分隔行
        本身就占 2 行——按 2 计的话,一张空表照样过闸。**这一条也必须只数「载体记录」小节内的
        表格行**:数全文的 `|` 行会把「出场人物」表算进来,同样让空载体表蒙混过关。
      - 白描段计数少于 `$P`:有情节点缺白描,或字段顺序、分隔符不对。
      - 功能段计数少于 `$P`:有情节点没回答改变了什么。注意全角冒号——半角会漏。
        这一条只数情节点行第四段里的 `功能:`:早先的写法数的是全文任何位置的这四个字符,
        于是正文里提一句「功能」就会让计数超过 `$P`,把一份合规的提取判成失败。
      - 类型超出枚举:自造类型会让 S3 的归类整段落空。
      - 载体表行数少于 2:表头加至少一行数据。只核对表头存在,是数不出「表里一条记录都
        没有」的——而 S5 判定 `needs_carrier` 要的正是那些数据行。
      
      文件名由 S2 的覆盖率闸门按字面匹配,写歪一个字符会进 `unmatched_files`——
      那既不算覆盖也不算缺章,改名即可,不必重跑。
      
      ## 并发与串行
      
      **两条路径的写法要求和自检是同一份**,区别只有速度。
      
      - **并发**:每批 5–8 章,一批全部落盘再发下一批,每批结束更新 `_progress.md`。
        每条子任务的提示必须自带材料授权前缀,说明这是创作者合法持有的虚构作品、
        任务是只读的转化性分析、通俗题材的黑暗元素照常提取——否则子代理会把正常的戏剧化
        情节误判为有害内容并拒绝整章,而主线程只会看到一个空结果。
      - **串行**:主线程逐章处理,同样跑机械自检。
      
      运行环境不支持子代理、或本技能已在子代理上下文里运行时,直接走串行,并在汇报里说明。
      
      ## 失败处理
      
      | 情况 | 处理 |
      |---|---|
      | 子代理崩溃、超时、空输出 | 同条件重试 1 次 |
      | 机械自检不过 | 带上失败项重写本章 1 次 |
      | 重试仍失败 | 标记 `⚠️ 跳过`,写入 `_progress.md` 失败记录 |
      
      单章失败不阻断管道。但**跳过的章必须出现在每一份继承它的聚合产物里**——
      一份看起来完整、实际缺了第 47 章的分析,比一份明说缺章的分析危险得多。
      
    • stage-contract.md 5.2 KB
      # 原著分析阶段契约
      
      ## 目录
      
      - [独立运行与项目集成](#独立运行与项目集成)
      - [所有权边界](#所有权边界)
      - [材料授权与只读转化](#材料授权与只读转化)
      - [本阶段规则](#本阶段规则)
      
      本文件是本技能的自包含契约:预检、所有权、材料边界与规则表都在这里,
      不需要读取其他技能的文件。
      
      ## 独立运行与项目集成
      
      本技能可以单独安装并运行;下面几条说明项目工具存在时怎样集成。
      
      1. **读取直接输入**:只读取用户明确提供或当前任务实际需要的文件,不批量加载整个项目。
      2. **可选项目集成**:若存在 `short-drama.json` 且 core 项目工具可用,可以运行
         `python3 <core>/scripts/project_tool.py status <project>`,使用返回的目录布局和语言设置;
         core 不可用时,直接基于已提供输入产出本阶段文件。
      3. **可选发布生命周期**:项目工具可用时,用 `publish` 原子发布并用 `--input <path>`
         声明直接输入;`accept`、`review` 与 `package` 继续承担确认、复核与交付。
      4. **保持职责分离**:创作者确认、内容修订和复核是不同动作;reviewer 提修改要求,负责人改文件。
      
      ## 所有权边界
      
      - **本阶段拥有**:`项目开发/source-analysis/` 下的章节索引、改编价值快评、逐章提取、
        剧情单元、节奏情绪、人物与设定候选、改编价值评估与分集候选。
      - **本阶段继承**:创作者提供的原始材料与授权说明、项目语言、已接受的形式约束与制作形态。
      - **本阶段不越权**:不改写 `输入/`;不写改编契约、创作简报、故事引擎或
        `adaptation-map.jsonl`(`$short-drama-develop` 拥有);不建资产身份;不写场景、台词、
        分镜与提示词;不生成媒体;不批准自己的产物。
      
      分析层与决策层分开的理由很直接:**分析可以被推翻,契约不能**。把两者写进同一份文件,
      创作者就再也无法只推翻分析而保留已确认的改编承诺。
      
      ## 材料授权与只读转化
      
      - 只处理创作者声明**合法持有、拥有使用权**的作品;授权不清时保留问题,不替创作者定案。
      - 分析是转化性的:提取结构与功能,不复制原文成段落,不模仿原作文风生成新文本。
      - 引用采用最短必要片段,并绑定 span。功能摘要必须是**去引用**的重述。
      - 通俗题材的暴力、复仇、背叛、情爱张力与黑暗伦理照常提取;个别片段无法处理时跳过并记录,
        不中止整章或整本。
      - `输入/` 与其中的原始材料不进入交付包。
      
      ## 本阶段规则
      
      ### `NVA`
      
      | ID | Class | Knowledge |
      |---|---|---|
      | NVA-01 | structural_invariant | All stages slice the source through one chapter index built from that source; `verify` blocks on a chapter set whose numbering, ordering or line spans do not cover that source. An edited source still invalidates the index and everything derived from it, but only a changed line count or numbering is reported — a same-line-count rewrite in place is not, so whoever edits the source re-indexes it. |
      | NVA-02 | structural_invariant | Every extracted claim carries a source locator and span; a claim with no span cannot be cited downstream. |
      | NVA-03 | structural_invariant | Aggregation may not start while chapter coverage is incomplete; missing chapters are named in every aggregate that inherits the gap. |
      | NVA-04 | structural_invariant | Analysis records de-quoted function summaries, never copied source paragraphs. |
      | NVA-05 | structural_invariant | A sampled stage states its own coverage and confines every claim to the chapters it read. |
      | NVA-06 | reviewed_invariant | A function summary states what a passage does to character choice, information, power or relationship, not what happens in it. |
      | NVA-07 | reviewed_invariant | Hard facts — levels, counts, distances, who said what, which chapters a character appears in — trace back to a description line or the source; an unsupported one is written as unstated, never filled in by plausibility. |
      | NVA-08 | reviewed_invariant | Character merges preserve dramatic role, knowledge scope, relational position and causal bridge; only proper names and evidenced nicknames may merge, never descriptors or titles. |
      | NVA-09 | reviewed_invariant | Adaptation value distinguishes what the screen can show from what only prose can deliver, and names the new carrier for each function it keeps. |
      | NVA-10 | reviewed_invariant | Episode candidates are cut on local dramatic result and precise handoff, not on chapter count or word budget. |
      | NVA-11 | craft_default | Triage a deterministic spread of chapters across the whole book and stop for the creator before committing to a full pass. |
      | NVA-12 | taste_option | Where to open, which line to keep and which ending to promise remain creator choices; analysis may argue but never blocks. |
      
      规则分级由高到低:`structural_invariant`(结构缺陷,阻断)、
      `reviewed_invariant`(需证据判断)、`craft_default`(常用做法,可覆盖)、
      `taste_option`(创作者选择,不作缺陷)。创作者已接受的事实优先于本表。
      
  • scripts
    • novel_index.py 25.5 KB
      #!/usr/bin/env python3
      """Build, verify and sample one stable chapter index for a long source novel.
      
      Every stage of source analysis slices the same file. If each stage ran its own
      regular expression the slices would disagree, and a chapter analysed under one
      boundary would be aggregated under another. This script owns the only slicing
      truth: it detects chapter boundaries once, binds each span to a content hash,
      and refuses to stay quiet about numbering it cannot explain.
      
      It performs no editorial judgement. Which chapters matter, what they mean and
      how they adapt are decisions for the skill workflow, not for this file.
      """
      
      from __future__ import annotations
      
      import argparse
      import json
      import os
      import re
      import sys
      import unicodedata
      from pathlib import Path
      from typing import Any
      
      # Creators run these scripts on whatever interpreter their machine provides, so
      # an unsupported version must say so instead of failing inside an import.
      MINIMUM_PYTHON = (3, 9)
      if sys.version_info < MINIMUM_PYTHON:
          raise SystemExit(
              "short-drama needs Python {}.{} or newer; this interpreter is {}.{}".format(
                  *MINIMUM_PYTHON, sys.version_info.major, sys.version_info.minor
              )
          )
      
      SCHEMA_VERSION = "1.0.0-draft"
      
      # Chapter numbers appear as Arabic digits or Chinese numerals. 千 and 两 are
      # included because serialized fiction routinely passes 1000 chapters, and
      # 第两百章 is as common as 第二百章 in that range.
      CHINESE_NUMERALS = "零一二三四五六七八九十百千两"
      # Books number their chapters with 章, 回 or 节 — but a book that uses 章 for
      # chapters may also use 节 for subsections inside them. Matching every unit at
      # once turns those subsections into chapters, so the unit is captured here and
      # a single dominant unit is chosen below.
      CHAPTER_UNITS = ("章", "回", "节")
      CHAPTER_RE = re.compile(
          r"^[ \t ]*第\s*([0-9]+|[" + CHINESE_NUMERALS + r"]+)\s*([章回节])"
          r"[ \t ]*(.*?)[ \t ]*$"
      )
      CHINESE_DIGITS = {
          "零": 0, "一": 1, "二": 2, "两": 2, "三": 3, "四": 4,
          "五": 5, "六": 6, "七": 7, "八": 8, "九": 9,
      }
      CHINESE_UNITS = {"十": 10, "百": 100, "千": 1000}
      # A chapter heading is a short standalone line. A body paragraph that happens to
      # open with 第三章 — a character recalling what happened back then — otherwise
      # becomes a chapter boundary, and the index silently gains a chapter whose span
      # starts mid-scene. Generous enough for the long titles serialized fiction uses.
      MAX_HEADING_LINE_WIDTH = 50
      
      
      def chinese_to_int(text: str) -> int | None:
          """Convert 一 / 十五 / 两百零三 / 一千零一 to an int, or None if malformed.
      
          Returning None rather than raising matters: a heading whose number cannot be
          read must be reported as an unreadable heading, not crash the whole index
          and leave the creator with no boundary table at all.
          """
      
          if not text:
              return None
          if text.isdigit():
              return int(text)
          total = 0
          section = 0
          digit_seen = False
          for character in text:
              if character in CHINESE_DIGITS:
                  section = CHINESE_DIGITS[character]
                  digit_seen = True
              elif character in CHINESE_UNITS:
                  # 十五 means 15, so a leading 十 carries an implicit one.
                  total += (section or 1) * CHINESE_UNITS[character]
                  section = 0
                  digit_seen = True
              else:
                  return None
          if not digit_seen:
              return None
          return total + section
      
      
      def _atomic_write_text(path: Path, content: str) -> None:
          path.parent.mkdir(parents=True, exist_ok=True)
          temporary = path.with_name(f".{path.name}.{os.getpid()}.tmp")
          try:
              with temporary.open("x", encoding="utf-8") as handle:
                  handle.write(content)
                  handle.flush()
                  os.fsync(handle.fileno())
              os.replace(temporary, path)
          finally:
              if temporary.exists():
                  temporary.unlink()
      
      
      def _visible_width(text: str) -> int:
          """Count characters as a creator counts them, not as bytes."""
      
          return sum(1 for character in text if not unicodedata.combining(character))
      
      
      def find_heading_lines(lines: list[str]) -> tuple[list[dict[str, Any]], int]:
          headings: list[dict[str, Any]] = []
          long_lines = 0
          for offset, line in enumerate(lines):
              candidate = line.removeprefix("\ufeff") if offset == 0 else line
              match = CHAPTER_RE.match(candidate)
              if match is None:
                  continue
              if _visible_width(candidate.strip()) > MAX_HEADING_LINE_WIDTH:
                  long_lines += 1
                  continue
              headings.append(
                  {
                      "line_index": offset,
                      "raw": candidate.strip(),
                      "number": chinese_to_int(match.group(1)),
                      "unit": match.group(2),
                      "title": match.group(3).strip(),
                  }
              )
          return headings, long_lines
      
      
      # Two headings are the least that can be a numbering run. One 第三回 inside a
      # chapter of dialogue is a character speaking, not a book changing its unit.
      MINIMUM_NUMBERING_RUN = 2
      
      
      def select_chapter_unit(
          headings: list[dict[str, Any]], forced: str | None = None
      ) -> tuple[str | None, dict[str, int]]:
          """Keep one numbering unit and report what the other units cost.
      
          A book that writes chapters as 第N章 and subsections as 第N节 has two
          interleaved numbering runs. Accepting both produces duplicate numbers, and a
          duplicate used to switch off every numbering check at once — so the very
          file that needed reporting came back clean.
      
          The unit is chosen by depth, not by frequency. Choosing the most frequent
          unit inverted every such book, because subsections always outnumber the
          chapters that contain them: a book of 4 章 with 3 节 each was indexed as 12
          chapters of 节, `problems` empty, and every later stage then cut episodes on
          subsection boundaries. `CHAPTER_UNITS` is ordered coarse to fine, so the
          coarsest unit that forms a real numbering run wins, and a book that numbers
          its chapters 第N节 with nothing coarser still works.
          """
      
          counts = {unit: 0 for unit in CHAPTER_UNITS}
          for heading in headings:
              counts[heading["unit"]] += 1
          if not headings:
              return None, counts
          if forced is not None:
              return forced, {
                  unit: count for unit, count in counts.items() if unit != forced and count
              }
          chosen = next(
              (unit for unit in CHAPTER_UNITS if counts[unit] >= MINIMUM_NUMBERING_RUN),
              None,
          )
          if chosen is None:
              # No unit forms a run. Fall back to the most frequent, resolving ties by
              # the declared order so the choice never depends on which heading came
              # first.
              chosen = max(
                  CHAPTER_UNITS, key=lambda unit: (counts[unit], -CHAPTER_UNITS.index(unit))
              )
          ignored = {unit: count for unit, count in counts.items() if unit != chosen and count}
          return chosen, ignored
      
      
      # A contents entry is one line; the next follows immediately or after a blank.
      # Chapter prose never packs headings this tightly, so the gap is an absolute
      # signal and does not need calibrating against the rest of the file.
      CONTENTS_MAX_GAP = 3
      CONTENTS_MIN_ENTRIES = 3
      # A volume shorter than this is not a volume. Treating every restart as a legal
      # multi-volume structure is what let a single stray duplicated heading disable
      # numbering validation for an entire book.
      MIN_VOLUME_CHAPTERS = 3
      
      
      def drop_leading_table_of_contents(
          headings: list[dict[str, Any]]
      ) -> tuple[list[dict[str, Any]], int]:
          """Discard a leading contents block, which repeats every chapter heading.
      
          Density alone is not enough to decide. A median-based threshold breaks
          exactly when the contents block is a large share of all headings, because
          the block then drags the median into its own range and survives the test.
      
          So the rule is density plus corroboration: a leading run of headings packed
          within `CONTENTS_MAX_GAP` lines is dropped only if the chapter numbers it
          lists reappear among the headings that follow. A contents block always
          duplicates the chapters it indexes; a genuinely dense opening does not.
          Without this the index carries two "第一章" rows and every later stage
          slices at the contents entry, which holds no prose at all.
          """
      
          if len(headings) < CONTENTS_MIN_ENTRIES * 2:
              return headings, 0
          numbers = [heading["number"] for heading in headings]
          if numbers[0] is None:
              return headings, 0
          # The cut is where numbering restarts: a contents block is followed by the
          # same chapters again. Anchoring on the gap instead is off by one, because
          # the step from the last contents entry to chapter one is itself small.
          cut = 0
          for index in range(1, len(headings)):
              if numbers[index] is not None and numbers[index] <= numbers[0]:
                  cut = index
                  break
          if cut < CONTENTS_MIN_ENTRIES or cut > len(headings) - cut:
              return headings, 0
          internal_gaps = [
              headings[index + 1]["line_index"] - headings[index]["line_index"]
              for index in range(cut - 1)
          ]
          if any(gap > CONTENTS_MAX_GAP for gap in internal_gaps):
              # Numbering restarts, but the prefix holds real prose between headings.
              # That is a multi-volume book, not a contents block.
              return headings, 0
          dropped = {number for number in numbers[:cut] if number is not None}
          remaining = {number for number in numbers[cut:] if number is not None}
          if not dropped or not dropped <= remaining:
              # The packed prefix indexes chapters that never appear. It is part of
              # the story, not a contents list, so keep it.
              return headings, 0
          return headings[cut:], cut
      
      
      def build_chapters(
          lines: list[str], headings: list[dict[str, Any]], text: str
      ) -> list[dict[str, Any]]:
          line_offsets = _line_offsets(lines)
          chapters: list[dict[str, Any]] = []
          for position, heading in enumerate(headings):
              start_line = heading["line_index"]
              end_line = (
                  headings[position + 1]["line_index"]
                  if position + 1 < len(headings)
                  else len(lines)
              )
              body = text[
                  line_offsets[start_line] : (
                      line_offsets[end_line] if end_line < len(line_offsets) else len(text)
                  )
              ]
              chapters.append(
                  {
                      "sequence": position + 1,
                      "source_number": heading["number"],
                      "unit": heading["unit"],
                      "title": heading["title"],
                      "heading": heading["raw"],
                      "line_start": start_line + 1,
                      "line_end": end_line,
                      "char_count": _visible_width(body),
                  }
              )
          return chapters
      
      
      def _line_offsets(lines: list[str]) -> list[int]:
          offsets: list[int] = []
          cursor = 0
          for line in lines:
              offsets.append(cursor)
              cursor += len(line) + 1
          return offsets
      
      
      def segment_by_restart(chapters: list[dict[str, Any]]) -> list[list[dict[str, Any]]]:
          """Split the chapter list wherever source numbering stops increasing.
      
          Each segment is a candidate volume. Validation then runs *inside* every
          segment, so a book whose volumes restart at 第一章 stays legal while a
          missing chapter inside any one volume is still reported.
          """
      
          segments: list[list[dict[str, Any]]] = []
          current: list[dict[str, Any]] = []
          previous: int | None = None
          for chapter in chapters:
              number = chapter["source_number"]
              if number is not None and previous is not None and number <= previous:
                  segments.append(current)
                  current = []
              if number is not None:
                  previous = number
              current.append(chapter)
          if current:
              segments.append(current)
          return segments
      
      
      def validate_chapters(
          chapters: list[dict[str, Any]], segments: list[list[dict[str, Any]]]
      ) -> list[str]:
          problems: list[str] = []
          if not chapters:
              problems.append("no chapter heading matched; the source may need manual spans")
              return problems
          unreadable = [
              chapter["heading"] for chapter in chapters if chapter["source_number"] is None
          ]
          if unreadable:
              problems.append("unreadable chapter numbers: " + ", ".join(unreadable[:5]))
          for position, segment in enumerate(segments, start=1):
              label = "" if len(segments) == 1 else f" in volume {position}"
              if len(segments) > 1 and len(segment) < MIN_VOLUME_CHAPTERS:
                  # Too short to be a volume, so the restart is far more likely to be
                  # a duplicated or stray heading than a legal structure.
                  problems.append(
                      f"numbering restarts after only {len(segment)} chapter(s) at "
                      f"sequence {segment[0]['sequence']}; expected a volume boundary, "
                      "check for a stray or duplicated heading"
                  )
              numbers = [
                  chapter["source_number"]
                  for chapter in segment
                  if chapter["source_number"] is not None
              ]
              for previous, current in zip(numbers, numbers[1:]):
                  if current == previous:
                      problems.append(f"duplicate chapter number {current}{label}")
                      break
                  if current != previous + 1:
                      problems.append(
                          f"chapter numbering jumps from {previous} to {current}{label}"
                      )
                      break
          empty = [chapter["sequence"] for chapter in chapters if chapter["char_count"] < 50]
          if empty:
              problems.append(
                  "chapters with almost no body text at sequence "
                  + ", ".join(str(item) for item in empty[:5])
              )
          return problems
      
      
      def build_index(source: Path, unit_override: str | None = None) -> dict[str, Any]:
          raw = source.read_bytes()
          text = raw.decode("utf-8")
          lines = text.split("\n")
          headings, long_heading_lines = find_heading_lines(lines)
          unit, ignored_units = select_chapter_unit(headings, unit_override)
          headings = [heading for heading in headings if heading["unit"] == unit]
          headings, dropped = drop_leading_table_of_contents(headings)
          chapters = build_chapters(lines, headings, text)
          segments = segment_by_restart(chapters)
          problems = validate_chapters(chapters, segments)
          if unit is not None:
              coarser = [
                  other
                  for other in CHAPTER_UNITS[: CHAPTER_UNITS.index(unit)]
                  if ignored_units.get(other)
              ]
              if coarser:
                  # SKILL.md names this exact shape as the sign that the unit judgment
                  # is wrong. Reporting it is what lets a creator act on it; `problems`
                  # coming back empty here is how a subsection index passed for a
                  # chapter index all the way to episode cutting.
                  problems.append(
                      "chapters were indexed as 第N{} while coarser headings were skipped: {}"
                      "; pass --unit to choose the unit yourself".format(
                          unit,
                          ", ".join(
                              "第N{} x{}".format(other, ignored_units[other])
                              for other in coarser
                          ),
                      )
                  )
          return {
              "schema_version": SCHEMA_VERSION,
              "record_type": "novel_chapter_index",
              "source": {
                  "artifact": source.name,
                  "char_count": _visible_width(text),
                  "line_count": len(lines),
              },
              "chapter_unit": unit,
              # Kept in the record rather than only warned about: a creator whose book
              # numbers chapters as 第N节 needs to see that 章 headings were skipped.
              "ignored_heading_units": ignored_units,
              # Body paragraphs that merely open with a chapter number. Reported rather
              # than silently dropped: if a book genuinely titles its chapters this
              # long, the count is how a creator finds out why chapters went missing.
              "long_heading_lines_skipped": long_heading_lines,
              "contents_entries_dropped": dropped,
              "volume_numbering_restarts": len(segments) > 1,
              "volume_count": len(segments),
              "chapter_count": len(chapters),
              "chapters": chapters,
              "problems": problems,
          }
      
      
      REQUIRED_CHAPTER_FIELDS = ("sequence", "line_start", "line_end")
      
      
      def verify_index(index_path: Path, source: Path) -> dict[str, Any]:
          """Re-check an index against its source without trusting either.
      
          The workflow tells creators to write spans by hand when a source carries no
          headings at all, so this must report a malformed row rather than raise on
          it: an exception here reads as a broken tool, not as a fixable index.
          """
      
          index = json.loads(index_path.read_text(encoding="utf-8"))
          raw = source.read_bytes()
          problems: list[str] = []
          if not isinstance(index, dict):
              return {"verified": False, "problems": ["index must be a JSON object"]}
          source_record = index.get("source")
          if not isinstance(source_record, dict):
              return {
                  "verified": False,
                  "problems": [
                      "index is missing its source record"
                  ],
              }
          text = raw.decode("utf-8")
          lines = text.split("\n")
          chapters = index.get("chapters", [])
          if not isinstance(chapters, list) or not chapters:
              return {"verified": False, "problems": ["index carries no chapters"]}
          chapter_count = index.get("chapter_count")
          if (
              isinstance(chapter_count, bool)
              or not isinstance(chapter_count, int)
              or chapter_count != len(chapters)
          ):
              problems.append("chapter_count does not match the chapter rows")
          previous_end: int | None = None
          for position, chapter in enumerate(chapters, start=1):
              if not isinstance(chapter, dict):
                  problems.append(f"chapter row {position} is not an object")
                  continue
              missing = [field for field in REQUIRED_CHAPTER_FIELDS if field not in chapter]
              if missing:
                  problems.append(
                      f"chapter row {position} is missing {', '.join(missing)}; "
                      "a hand-written span must carry the same fields the script writes"
                  )
                  continue
              sequence = chapter["sequence"]
              start_line, end_line = chapter["line_start"], chapter["line_end"]
              if (
                  isinstance(sequence, bool)
                  or not isinstance(sequence, int)
                  or sequence != position
              ):
                  problems.append(
                      f"chapter row {position} sequence must be the contiguous value {position}"
                  )
                  continue
              if any(
                  isinstance(value, bool) or not isinstance(value, int)
                  for value in (start_line, end_line)
              ):
                  problems.append(f"chapter {sequence} line span must use integer line numbers")
                  continue
              if not (1 <= start_line <= len(lines)) or not (
                  start_line <= end_line <= len(lines)
              ):
                  problems.append(
                      f"chapter {chapter['sequence']} spans lines "
                      f"{start_line}-{end_line}, outside the source's {len(lines)} lines"
                  )
                  continue
              if previous_end is not None and start_line != previous_end + 1:
                  problems.append(
                      f"chapter {sequence} must start at line {previous_end + 1}; "
                      f"found {start_line}"
                  )
                  continue
              previous_end = end_line
          if previous_end is not None and previous_end != len(lines):
              problems.append(
                  f"chapter spans end at line {previous_end}; source ends at line {len(lines)}"
              )
          return {"verified": not problems, "problems": problems}
      
      
      # The triage stage reads a fraction of the book. A fixed, explainable count beats
      # a formula nobody can predict, and the reported ratio is what stops a triage
      # from being mistaken for a full pass.
      DEFAULT_SAMPLE_COUNT = 12
      # One literal for the stage marker, shared by the function default and the CLI
      # flag. Spelling it twice is how the two drift apart, and a drifted default
      # reports every chapter as missing while the files sit right there.
      DEFAULT_COVERAGE_STAGE = "extract"
      
      
      def sample_chapters(index_path: Path, count: int = DEFAULT_SAMPLE_COUNT) -> dict[str, Any]:
          """Choose a reproducible spread of chapters for the adaptation triage.
      
          Evenly spaced and always including the first and last chapter: the triage
          asks whether a book is worth adapting, and that question is answered by how
          the whole spine behaves, not by how well the opening performs. The selection
          is deterministic so a second run cites the same chapters as the first.
          """
      
          index = json.loads(index_path.read_text(encoding="utf-8"))
          chapters = index.get("chapters", [])
          total = len(chapters)
          if total == 0:
              return {"total_chapters": 0, "sampled": [], "coverage_ratio": 0.0}
          count = max(1, min(count, total))
          if count == 1:
              positions = [0]
          else:
              positions = sorted(
                  {round(step * (total - 1) / (count - 1)) for step in range(count)}
              )
          sampled = [
              {
                  "sequence": chapters[position]["sequence"],
                  "source_number": chapters[position].get("source_number"),
                  "title": chapters[position].get("title", ""),
                  "line_start": chapters[position]["line_start"],
                  "line_end": chapters[position]["line_end"],
              }
              for position in positions
          ]
          return {
              "total_chapters": total,
              "sampled_count": len(sampled),
              "coverage_ratio": round(len(sampled) / total, 4),
              "sampled": sampled,
          }
      
      
      def coverage(
          index_path: Path, analysis_dir: Path, stage: str = DEFAULT_COVERAGE_STAGE
      ) -> dict[str, Any]:
          """Report which chapters have no analysis file for one named stage.
      
          The stage marker is required, not cosmetic. Chapter files for different
          stages live in the same directory, so matching any file that merely contains
          a chapter number lets one stage's output satisfy another stage's gate — and
          the gate that exists to stop a silently skipped chapter then reports a book
          as complete.
          """
      
          index = json.loads(index_path.read_text(encoding="utf-8"))
          expected = {chapter["sequence"] for chapter in index.get("chapters", [])}
          pattern = re.compile(rf"^ch-(\d+)-{re.escape(stage)}$")
          found: set[int] = set()
          unmatched: list[str] = []
          for path in sorted(analysis_dir.glob("*.md")):
              match = pattern.fullmatch(path.stem)
              if match is None:
                  unmatched.append(path.name)
              else:
                  found.add(int(match.group(1)))
          missing = sorted(expected - found)
          return {
              "stage": stage,
              "expected": len(expected),
              "present": len(expected & found),
              "missing": missing,
              "unexpected": sorted(found - expected),
              # Named rather than counted: a file whose name drifted by one character
              # is invisible to the gate, and looks identical to a missing chapter.
              "unmatched_files": unmatched,
              "complete": not missing,
          }
      
      
      def main(argv: list[str] | None = None) -> int:
          parser = argparse.ArgumentParser(description=__doc__)
          sub = parser.add_subparsers(dest="command", required=True)
      
          build = sub.add_parser("index", help="build the chapter index for one source")
          build.add_argument("source", type=Path)
          build.add_argument("--out", type=Path, default=None)
          build.add_argument(
              "--unit",
              choices=CHAPTER_UNITS,
              default=None,
              help="force the chapter unit instead of detecting it (章 / 回 / 节)",
          )
      
          check = sub.add_parser("verify", help="re-check an index against its source")
          check.add_argument("index", type=Path)
          check.add_argument("source", type=Path)
      
          pick = sub.add_parser("sample", help="choose triage chapters spread across the book")
          pick.add_argument("index", type=Path)
          pick.add_argument("--count", type=int, default=DEFAULT_SAMPLE_COUNT)
      
          cover = sub.add_parser("coverage", help="report chapters with no analysis file")
          cover.add_argument("index", type=Path)
          cover.add_argument("analysis_dir", type=Path)
          cover.add_argument("--stage", default=DEFAULT_COVERAGE_STAGE)
      
          args = parser.parse_args(argv)
      
          if args.command == "index":
              document = build_index(args.source, args.unit)
              payload = json.dumps(document, ensure_ascii=True, indent=2, sort_keys=True)
              if args.out is not None:
                  _atomic_write_text(args.out, payload + "\n")
                  print(
                      json.dumps(
                          {
                              "chapter_count": document["chapter_count"],
                              "chapter_unit": document["chapter_unit"],
                              "ignored_heading_units": document["ignored_heading_units"],
                              "problems": document["problems"],
                              "out": str(args.out),
                          },
                          ensure_ascii=True,
                          indent=2,
                      )
                  )
              else:
                  print(payload)
              return 1 if document["problems"] else 0
      
          if args.command == "verify":
              result = verify_index(args.index, args.source)
              print(json.dumps(result, ensure_ascii=True, indent=2))
              return 0 if result["verified"] else 1
      
          if args.command == "sample":
              result = sample_chapters(args.index, args.count)
              print(json.dumps(result, ensure_ascii=True, indent=2))
              return 0 if result["sampled"] else 1
      
          result = coverage(args.index, args.analysis_dir, args.stage)
          print(json.dumps(result, ensure_ascii=True, indent=2))
          return 0 if result["complete"] else 1
      
      
      if __name__ == "__main__":
          raise SystemExit(main())
      
    • selftest.py 1.7 KB
      #!/usr/bin/env python3
      """Offline self-test for the standalone novel index."""
      
      from __future__ import annotations
      
      import json
      import sys
      import tempfile
      from pathlib import Path
      
      from novel_index import build_index, sample_chapters, verify_index
      
      MINIMUM_PYTHON = (3, 9)
      if sys.version_info < MINIMUM_PYTHON:
          raise SystemExit("selftest.py requires Python 3.9 or newer")
      
      
      def require(condition: bool, message: str) -> None:
          if not condition:
              raise AssertionError(message)
      
      
      def main() -> int:
          with tempfile.TemporaryDirectory() as directory:
              root = Path(directory)
              source = root / "novel.txt"
              index_path = root / "index.json"
              source.write_text(
                  "第一章 起势\n"
                  + "人物做出选择,代价随之发生,旧关系因此失衡,新的目标也被迫提前。" * 4
                  + "\n第二章 反转\n"
                  + "旧承诺被新的证据推翻,人物必须在名誉与家人之间付出不可逆的代价。" * 4
                  + "\n",
                  encoding="utf-8",
              )
              index = build_index(source)
              require(index["chapter_count"] == 2, "chapter count")
              require(index["problems"] == [], "valid chapter index")
              index_path.write_text(json.dumps(index, ensure_ascii=False), encoding="utf-8")
              require(verify_index(index_path, source)["verified"] is True, "fresh index")
              require(sample_chapters(index_path, 2)["sampled_count"] == 2, "sample count")
      
              source.write_text(source.read_text(encoding="utf-8") + "变化\n", encoding="utf-8")
              require(verify_index(index_path, source)["verified"] is False, "source drift")
      
          print("5 self-tests passed")
          return 0
      
      
      if __name__ == "__main__":
          raise SystemExit(main())
      
  • SKILL.md 14.5 KB
    ---
    name: short-drama-novel-analyze
    description: 把长篇小说、连载网文或多集散稿拆成可追溯的原著分析:章节索引、改编价值快评、逐章功能提取、剧情单元与节奏聚合、人物与设定归并,最后给出改编价值判定与分集候选,交给 $short-drama-develop 立契约。用户说“导入这本小说”“拆这本书”“分析原著”“这本书能不能改短剧”“先看看值不值得拆”“把长篇拆成分集候选”,或直接给出小说文件路径时使用。只做只读的结构化分析,不写剧本、不建资产、不生成媒体,也不替创作者决定改编方案。
    license: MIT
    ---
    
    # 长篇原著分析
    
    把一部长材料变成**能被引用、能被反驳、能被接着用**的分析层。目标不是复述剧情,而是找出
    每一段承担的戏剧功能,并说明它在竖屏短剧里值多少钱。
    
    分析永远是候选。哪条线保留、哪些人合并、从哪里开篇,是创作者的决定,由
    `$short-drama-develop` 立成改编契约。本技能不替它决定,也不批准自己的产物。
    
    ## Quick Start
    
    离线验证章节索引、采样和源文件变更检测:
    
    ```bash
    python3 {技能目录}/scripts/selftest.py
    python3 {技能目录}/scripts/novel_index.py index <原著.txt> --out <chapter-index.json>
    ```
    
    ## 开始前
    
    本技能可独立安装和执行。先读取用户明确提供的原著与本任务直接输入;若当前目录是
    `short-drama` 项目且项目工具可用,可以读取 `status` 并使用其发布生命周期,但缺少 core
    或任何其他技能都不是分析工作的阻断条件。完整边界与规则见
    [阶段契约](references/stage-contract.md),无需读取其他技能的文件。
    
    ## 材料前提
    
    只分析创作者**合法持有、拥有使用权**的作品。分析是只读的转化性工作:提取结构与功能,
    不复制原文成段落,不把原句搬进下游产物。
    
    通俗题材里的暴力、复仇、背叛、情爱张力与黑暗伦理是常规虚构叙事元素,照常做结构化提取。
    个别片段无法处理时跳过该段并记录,不要因此中止整章或整本——中止会让后续所有阶段拿到
    一份有洞却看不出洞在哪的分析。
    
    ## 先判断入口
    
    1. **只有书名,没有原文**:请创作者提供文件路径或粘贴正文。不要凭书名回忆情节——
       没有字节就没有 span,没有 span 的分析无法被引用,也无法被反驳。
    2. **有原文,未建项目**:直接建立一个本技能自己的工作区,至少包含只读输入目录
       `输入/` 与输出目录 `项目开发/source-analysis/_work/`;把原文字节复制到 `输入/` 后记录
       原始文件位置,再用本技能的 `novel_index.py` 建索引。若 `$short-drama` 可用,可选用它
       初始化同样的目录和发布生命周期,但不得把 core 安装变成开始分析的前提。
    3. **有原文,项目已在**:直接进入管道。
    4. **已有部分分析**:读 `项目开发/source-analysis/_progress.md` 从断点续跑,
       不重跑已完成阶段。
    
    ## 管道
    
    `输入/` 是不可变的创作者输入,本阶段只读它。全部产出落在 `项目开发/source-analysis/`。
    独立工作区与完整项目使用同一套相对路径,因此后续安装 core 时无需迁移分析产物。
    
    | 阶段 | 做什么 | 产出 | 停靠 |
    |---|---|---|---|
    | S0 | 建章节索引(脚本) | `_index.json`、`_progress.md` | 索引有问题就停 |
    | S1 | 改编价值快评(抽样) | `triage.md` | **停靠问创作者** |
    | S2 | 逐章功能提取 | `chapters/ch-<N>-extract.md` | 覆盖率不足就停 |
    | S3 | 剧情单元与节奏聚合 | `story-units.md`、`rhythm-and-emotion.md` | 阈值不达标就复核 |
    | S4 | 人物归并与设定 | `characters.md`、`world.md` | — |
    | S5 | 改编价值与分集候选 | `adaptation-value.md`、`episode-candidates.jsonl` | 交接 develop |
    
    ### S0 章节索引
    
    索引是**唯一切片真源**,由 [章节索引脚本](scripts/novel_index.py) 建立。每个阶段各跑
    一次正则,就会切出互相对不上的章节,第 47 章按一种边界分析、按另一种边界聚合,
    而且没有人会发现。
    
    ```bash
    python3 {技能目录}/scripts/novel_index.py index 输入/{原文文件} \
      --out 项目开发/source-analysis/_work/_index.next.json
    python3 {技能目录}/scripts/novel_index.py verify \
      项目开发/source-analysis/_work/_index.next.json 输入/{原文文件}
    # 项目工具可用时可选:
    python3 {core 技能目录}/scripts/project_tool.py publish {项目根} \
      --owner short-drama-novel-analyze --artifact-id source-analysis:index \
      --output 项目开发/source-analysis/_index.json=项目开发/source-analysis/_work/_index.next.json \
      --input 输入/{原文文件}
    ```
    
    脚本识别阿拉伯数字与中文数字章号(含 千 / 两,覆盖千章以上连载),**只认一种编号单位**
    (章/回/节里出现最多的那个,其余记进 `ignored_heading_units`),只把**短的独立行**当标题
    (以章号开头的正文段落记进 `long_heading_lines_skipped`),剔除开头的目录块,
    按卷分段校验编号。它**不做编辑判断**——哪章重要、
    讲了什么,是后面阶段的事。
    
    `problems` 非空就停下报告,不要带着错表进 S1。常见四种:章号跳号(缺章或抓错标题)、
    同卷内重号、正文极少的章(多半抓到了目录残留或卷首页)、无法解析的章号。
    另外确认三个计数:`chapter_unit` 与 `ignored_heading_units`——一本用 `第N章` 分章、
    用 `第N节` 分小节的书应当看到 `chapter_unit: 章` 且 `节` 被记入忽略计数,反过来说明分章单位
    判断错了;`long_heading_lines_skipped` 明显偏大时,多半是这本书的章节标题确实很长,
    需要与创作者确认后手写边界。
    
    原文本身没有章节标题时索引会返回空表,此时与创作者确认按什么切分,把边界写进
    `_work/_index.next.json`,通过 `verify` 后再按上面的公开生命周期发布——手写的行也要带齐
    `sequence` / `line_start` / `line_end`,`verify` 会逐行检查并报出缺字段的行。
    
    改了原文必须重建索引,不能沿用旧 span。`verify` 核对编号、顺序与行号覆盖,所以插行删行会被报出来;
    但同行数的原地改写不会——那一步靠改原文的人自己重建,套件不比对字节。
    
    ```bash
    python3 {技能目录}/scripts/novel_index.py verify \
      项目开发/source-analysis/_index.json 输入/{原文文件}
    ```
    
    ### S1 改编价值快评
    
    回答**这本书值不值得花全量拆解的成本**。判据是全书的改编密度——`screen_ready` 单元占多少、
    `prose_only` 占多少、制作负担压在哪几段——所以快评横跨整本书,用脚本抽样:
    
    ```bash
    python3 {技能目录}/scripts/novel_index.py sample \
      项目开发/source-analysis/_index.json --count 12
    ```
    
    抽样确定、可复现、首尾必取,跑第二次引用的是同一批章。按
    [改编价值快评](references/adaptation-triage.md) 写 `triage.md`,覆盖六件事:
    故事框架、三类判定比例、开篇替换点、制作负担量级、最大的三处改编风险、分集候选量级。
    第一行写覆盖率(脚本返回的 `coverage_ratio`),所有结论限于抽样范围。
    
    **这里停靠问创作者**:给出快评与全量拆解的预计耗时(按章数粗估),问是否继续。
    创作者一开始就明确说「一次跑完」时仍写 `triage.md`,但不停下等待——它是 S5 要回填
    对照的第一版假设。
    
    停靠时把 `_progress.md` 的状态写成 `paused_after_triage`,断点写「下一步:S2 逐章提取」。
    
    S1–S5 的 Agent 创作产物同样先写到 `source-analysis/_work/`,完成本阶段机械检查后再发布到
    上表中的正式路径。项目工具可用时用 `project_tool.py publish` 和稳定 artifact-id;独立运行时
    原子替换正式文件。不要用半成品覆盖 `_index.json`、`_progress.md`、`chapters/*.md` 或聚合产物。
    `_work/` 是候选工作区,不是权威分析层,也不进入交付包。
    
    **并发子代理写 `_work/`,主线程发布到正式路径**。子代理只落
    `_work/chapters/ch-<N>-extract.md`,主线程跑完机械自检后再发布到 `chapters/`。
    覆盖率闸门按正式路径 `chapters/` 匹配文件名——发布之前跑,每一章都会进
    `unmatched_files`,那不是缺陷,只是跑早了。
    
    `_progress.md` 每个阶段都会重写,而发布要求一个路径只有一个 owner,所以它用一个固定
    artifact-id:`source-analysis:progress`,owner 是 `short-drama-novel-analyze`,
    S0–S5 每次停靠都用同一个 id 重新发布。
    
    ### S2 逐章功能提取
    
    按 [章节提取](references/chapter-extraction.md) 处理每一章。能并发子代理就分批并发
    (每批 5–8 章,等一批落盘再发下一批),不支持就串行——两条路径的写法要求和自检是同一份,
    只是速度不同。
    
    每章提取完落到 `chapters/ch-<N>-extract.md`,`<N>` 是索引里的 `sequence`,不是原文章号
    (多卷书的原文章号会重复,sequence 不会)。全部落盘后跑覆盖率:
    
    ```bash
    python3 {技能目录}/scripts/novel_index.py coverage \
      项目开发/source-analysis/_index.json 项目开发/source-analysis/chapters
    ```
    
    `missing` 非空就补跑缺的章;`unmatched_files` 非空说明有文件名写歪了——它既不算覆盖,
    也不会被当成缺章,必须改名而不是重跑。**不要在覆盖率不足时进入 S3**——聚合会照样产出
    一份读起来完整的结果,而缺掉的章不会在任何地方留下痕迹。
    
    单章连续失败两次就标记跳过,写进 `_progress.md` 的失败记录,并在后续每一份聚合产物里
    注明该章缺失。失败可以接受,失败被藏起来不行。
    
    ### S3 剧情单元与节奏
    
    从逐章提取聚合,不回头重读原文——原文已经在 S2 被读过一次,再读一次只会得到第二份互相
    矛盾的事实。按 [聚合与实体](references/aggregation-and-entities.md) 先识别故事框架
    (框架决定按什么切单元),再产出:
    
    - `story-units.md`:把情节点归成有始有终的单元,每个单元记录进入状态、冲突、代价与出去状态;
    - `rhythm-and-emotion.md`:关键信息如何逐章推进、情绪触动点的铺垫→释放→余波、
      跨章伏笔与兑现。
    
    聚合完成后跑同一文件里的三条阈值自检(归属置信、覆盖率、重叠率)与散落情节兜底。
    阈值不是评分,是**边界模糊的信号**:重叠率过高说明两个单元其实是一个。
    
    ### S4 人物与设定
    
    按 [聚合与实体](references/aggregation-and-entities.md) 归并人物(跨章去重、别名归一、
    分级),并从提及数据归纳世界规则、力量体系与势力。别名只有专名与有同指证据的绰号能合并,
    描述性称谓与头衔**永远不触发合并**。
    
    **人物归并是候选,不是资产**。 这里的人物条目带 `unresolved` 与来源引用,
    交给 `$short-drama-develop` 定改编决定、`$short-drama-write` 写进剧本之后,
    才由 `$short-drama-assets` 从已接受剧本建立真正的资产身份。绕过这条链直接建资产,
    等于让原著的人物表冒充剧本的出现证据。
    
    ### S5 改编价值与分集候选
    
    这是本技能与通用拆书的分水岭。按 [改编价值](references/adaptation-value.md) 产出:
    
    - `adaptation-value.md`:哪些单元在竖屏短剧里能直接成立、哪些要换载体、哪些是纯文字快感
      (内心戏、叙述性诡计、长铺垫)在画面上无法兑现;制作负担落在哪里。
    - `episode-candidates.jsonl`:按**局部戏剧结果与精确交接**切出的候选集,不按章号或字数
      平均切。每条带来源 span、承担的功能、以及未决项。
    
    同时**回填快评**:S1 的哪几条判断被全量结果推翻了,写进 `adaptation-value.md` 的开头。
    一个抽样结论被证伪,比它被悄悄忘掉有用得多——下一本书的快评会因此更准。
    
    样例见 [分集候选样例](assets/episode-candidate.example.jsonl)。
    
    ### 交接
    
    S5 完成后展示创作者可读的摘要,说明:拆了多少章、跳过哪些、分了多少个候选集、
    最大的三处改编风险、以及快评里被推翻的判断。然后交给 `$short-drama-develop`——
    由它把候选变成 `项目开发/adaptation-map.jsonl` 与改编契约。
    
    **本技能不写 `adaptation-map.jsonl`**,那是 develop 的产物。需要质量结论时交给独立的
    `$short-drama-review`(范围 `source_analysis`)。
    
    ## 规则分级
    
    - **`structural_invariant`**:索引与 span 的可证明性、引用完整性、覆盖率。可由脚本阻断。
    - **`reviewed_invariant`**:功能提取是否忠于原文、归并是否保住戏剧作用等语义义务。
    - **`craft_default`**:通常有帮助的做法;创作者说明理由后可覆盖。
    - **`taste_option`**:从哪里开篇、保留哪条线等选择;不得单独阻断。
    
    不要用固定的章数配方、情节点数量或篇幅比例替代因果判断。
    
    ## 产物与边界
    
    本技能只拥有 `项目开发/source-analysis/` 下的文件:`_index.json`、`_progress.md`、
    `chapters/*.md`、`triage.md`、`story-units.md`、`rhythm-and-emotion.md`、
    `characters.md`、`world.md`、`adaptation-value.md`、`episode-candidates.jsonl`。
    
    它不改写 `输入/`,不写 `项目开发/` 下其他技能的产物,不建资产、不写场景与台词、
    不写提示词、不生成媒体,也不签发终审结论。
    
    **不把原文成段复制进分析**。记录 locator、span 与去引用的功能摘要;需要证据时引用
    最短的必要片段。交付包不得把原始材料带出边界。
    
    ## 语言
    
    分析产物是创作者读的:项目内跟随 `short-drama.json#/language`,独立运行时跟随用户使用的
    语言,不在本技能内硬编码语言。本技能不产生提示词正文,与
    `#/format/prompt_language` 无关。
    
    ## 按需加载
    
    - **快评读哪些章、写哪六件事、怎么不冒充全量分析**:[改编价值快评](references/adaptation-triage.md)
    - **逐章提取写法、白描与叙事框架词的界线、机械自检、并发与串行**:[章节提取](references/chapter-extraction.md)
    - **故事框架、剧情单元、节奏情绪、人物归并、阈值与散落兜底**:[聚合与实体](references/aggregation-and-entities.md)
    - **改编价值评估、载体替换与分集候选切法**:[改编价值](references/adaptation-value.md)
    - **本阶段拥有什么、继承什么、不越权什么**:[阶段契约](references/stage-contract.md)
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related