Claude Skill

video-recap

从输入视频生成中文解说成片或原声剧情短片。用户提供 .mp4 / .mov / .mkv / .webm,并要求剪辑、添加旁白、 配音、总结、短剧/电视剧/电影/纪录片/科普解说时使用。负责编排 video-* 技能链:视频理解 → Agent 制定故事与视听方案 → 剪辑 → 配音 → 合成。触发词:视频解说、视频旁白、生成解说、 视频 recap、video recap、voiceover、narration、auto-dub、recap。

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download zenstory-ai-oh-story-dsh-packages_knowledge_video-recap_skills_video-recap-d734089.zip · 125 KB
Part of zenstory-ai/oh-story-dsh — 31 skills

Install

skills CLI npx skills add https://github.com/zenstory-ai/oh-story-dsh/tree/main/packages/knowledge/video-recap/skills/video-recap
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install zenstory-ai-oh-story-dsh@llmmart
Git git clone https://github.com/zenstory-ai/oh-story-dsh.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole zenstory-ai/oh-story-dsh collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

1. 定位与流程

本技能是五个独立技能的轻量编排器。各技能只通过 work_dir 中的 JSON / MP4 产物通信,不共享代码:

video-understanding ─▶ Agent 按 video-script 制定方案并写稿 ─▶ [video-cut] ─▶ video-voiceover ─▶ video-assemble

流程支持断点续跑:写好 narration.json 后重复同一条命令即可继续。第二阶段会比对 recap_run_manifest.json 记录的源视频路径、文件大小/修改时间与运行参数,拒绝复用来自其他源视频或其他参数的旧工作目录;视频理解产物也只在来源一致时复用。

画面流程 --edit-mode full|cut|dub 与声音策略 --audio-mode 是两个独立开关; narration 保留上述解说流程,source-mix 不做配音,adopted-packet-copy 冻结当前输入的已采用 AAC 音轨。 组合只有下面几条路径,其余组合在启动时直接报错:

输入 --edit-mode --audio-mode Agent 暂停点 流程 详见
单视频 full narration 1:narration.json 理解 → 写稿 → 校验 → 配音 → 合成 §4
单视频 cut narration 2:clip_plan.json,再对着成片写 narration.json 理解 → 剪辑 → 重建输出时间 brief → 写稿 → 配音 → 合成 §4
多视频 cut narration 2:同上,clip 必须带 source_id 逐源理解 → 剪辑 → 写稿 → 配音 → 合成 §4.3
单视频 full source-mix / adopted-packet-copy 无 直接合成当前整段 §4.7
单 / 多视频 cut source-mix / adopted-packet-copy 1:clip_plan.json 理解 → 剪辑 → 合成,不写稿 §4.7
单视频 dub narration 1:dub_script.json 英文转写 → 译稿 → 克隆音色整轨替换 §5
已剪好的母版 full narration + 三个采用 JSON 无 只做严格合成 下文

所有 full/cut 路径共用同一段收尾:(有旁白时)评审 → TTS → 合成 → 成片 QC。使用原声模式时读 references/audio-routing.md。

已有预制画面和本地采用的完整声音三件套时,可走严格 assembly-only 路径:

python3 scripts/recap.py picture.mp4 --edit-mode full --work-dir NEW_WORK \
  --output-dir DELIVERY \
  --tts-meta tts_meta.json \
  --narration-adoption narration_adoption.json \
  --audio-mix-adoption audio_mix_adoption.json

三个 JSON 参数必须同时出现。该入口只接受单视频、full、narration、音轨 0、新工作目录和未存在的 交付文件;不运行理解、写稿、解说评审、TTS、cut、MiMo QC 或剪映导出。语义与媒体形状仍由 video-assemble 严格验证,recap 只核对子技能绑定记录引用的是同一批采用文件与母版路径,不把调用方 采用的声音或混音声明成自动创作或发布批准。详见 references/audio-routing.md。

这里的单视频是已经剪好的母版。重剪后可以复用未改动的 WAV 与 tts_meta.json,但必须按新母版 重新写混音采用文件里的落点与准备好的音床;衔接步骤见 references/audio-routing.md 的 “Keep adopted voice after a cut”。

2. 创作职责

这不是单纯的 JSON / 渲染流水线。Agent 是本次内容的创作负责人。先判断本轮的创作控制模式(CREATE / DIRECTED / REVISION,与 --edit-mode 无关),再在进入昂贵的下游处理前完成五次判断:

  1. 导演判断:观众承诺、POV、戏剧问题、情绪终点与揭示节奏。
  2. 故事编辑:beat 定义为“发生了什么变化”,不是场景摘要。
  3. 画面剪辑:选择真正值得保留的具体时刻、反应、入点与出点。
  4. 声音/旁白:先分配画面、原声、沉默和旁白的任务,再写解说词。
  5. 观众复核:分别检查无旁白、只听声音和第一次观看的体验。

三种控制模式的定义、REVISION 的修改/冻结规则、创作方法以及 recap_story_plan.json / visual_audio_board.json / style_card.json 的写法,全部按 video-script 执行;它会要求先读创作手册。这些文件只记录可审计的当前决定,不增加服务或渲染依赖。

3. 环境与脚本路径

# ffmpeg: brew install ffmpeg | apt install ffmpeg | choco install ffmpeg
export MIMO_API_KEY=***

同一个 MiMo key 驱动:

  • ASR:mimo-v2.5-asr
  • VLM:mimo-v2.5
  • TTS:mimo-v2.5-tts

TTS 供应商由 --tts-provider mimo-tts|fish-audio|index-tts(或 TTS_PROVIDER)透传给配音技能;Fish Audio 与自托管 index-tts 各自的环境变量、默认音色和能力限制见该技能。ASR/VLM 始终使用 MiMo。--doctor 只做离线配置检查。

tp-* Token Plan 密钥默认使用中国区集群,可用 MIMO_TOKEN_PLAN_CLUSTER 覆盖。

可选能力:

  • --mimo-video-overview:按场景块补充 MiMo 视频理解。
  • --mimo-qc pre-assemble|post-render|both:在合成前、成片后或两个阶段给出建议型复核。

MiMo QC 默认关闭;每个选定阶段最多请求一次,写入 mimo_qc.json。任何凭证缺失、限流、超时、格式错误或采样失败都只记录状态,不阻断流程。可覆盖配置见 references/config-playbook.md,QC 报告的最小契约见 references/shift-left-qc-schema.md。

下面的 scripts/... 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。脚本启动后会自行定位兄弟技能和资源。

4. 标准解说流程(audio-mode narration)

4.1 背景调研

若能识别影片、剧集或主题,先按 video-understanding 技能的调研指南 research-guide.md 调研并写入 work_dir/background_research.json。视频理解会把人物名和剧情背景折入 VLM 上下文,避免只得到“黑衣男子”一类模糊描述。无法识别来源时可跳过。

4.2 分析并暂停创作

python3 scripts/recap.py <video> --work-dir <work_dir> --context "背景"

命令完成视频理解、写出 agent_narration_brief.md,然后暂停。此时按以下顺序执行 video-script:

  1. 查看创作 brief 与原片故事板。
  2. 写 recap_story_plan.json 和 visual_audio_board.json。
  3. full 模式写 narration.json;cut 模式第一阶段只写 clip_plan.json。
  4. cut 模式第二阶段查看剪后故事板,补充输出时间与声音分工,再写 narration.json。

不要从标题或旁白句子开始;先锁定故事体验和素材选择。

时间线有两条不可降级的硬约束:原声只能在可靠句末/静音边界被切入、切出或恢复;旁白必须使用 完整逐段音频,任何 clip 映射裁段、TTS 裁尾或剪映引用更长的加速前素材都阻断。Agent 收到 interrupts_source_sentence / unsafe_clip_sentence_boundary / no_safe_fit / timeline_audio_mismatch 时,应移动边界、缩短整句或删除该块,而不是增加抢断 override。

4.3 多视频与素材库

多视频只支持 cut 模式。项目 brief 会列出稳定的 source_id,clip_plan.json 中每个片段都必须填写来源:

python3 scripts/recap.py ep1.mp4 ep2.mp4 --edit-mode cut --target-duration 10m --work-dir work_dir_multi_ep

可选文件系统素材库:

python3 scripts/recap.py ep1.mp4 --material-library-dir .video-materials --save-materials
python3 scripts/recap.py ep1.mp4 ep2.mp4 --edit-mode cut --material-library-dir .video-materials --use-materials

素材检索只是对 JSON / MD / JSONL 做 grep,例如 grep -R "keyword" .video-materials。当前版本不复制原始媒体,也不提供数据库、向量或语义搜索。

同一根目录还可以登记可复用的资源(BGM、音效、音色、字体、图片)、带版本与采用记录的模板(字幕样式、包装图层)和样片。 格式与 scripts/library.py check|list|show 只读工具见 references/resource-library.md。用 --project recap_project.json 把已采用的字幕样式、音色与 BGM 绑定到这次运行; 每次 full / cut 合成后 work_dir/resource_lock.json 记下实际用到的资源与授权状态。

4.4 继续生成成片

写好所需产物后,重复同一条命令:

python3 scripts/recap.py <video> --work-dir <work_dir>  # 可追加 --edit-mode cut / --no-burn-subtitles

流程会校验当前阶段的硬输入(clip_plan.json / narration.json);两份创作计划仍是 Agent 与建议型评审使用的工作记录,不是渲染门禁。cut 模式随后生成 edited_source.mp4,再合成旁白并输出 recap_<name>.mp4。

若需要建议型 MiMo 复核:

python3 scripts/recap.py <video> --work-dir <work_dir> --mimo-qc both

合成前复核会读取脚本、计划和 TTS 元数据;成片后还会读取最多六张临时 JPEG。输入文件的大小/修改时间、模型与提示都未变时直接复用上次报告,--mimo-qc-refresh 可强制刷新。帧的 base64 与凭证不会写入磁盘。

已有批准解说稿时加 --preserve-approved-text:校验与 TTS 原样保留批准稿(只更新 overlaps_speech),装不下时间窗即失败,不缩稿、不降级为部分成功。

4.5 字幕与克隆旁白

若要把旁白字幕固定在原片字幕区域,先在仓库根目录运行:

python3 tools/measure_subtitle.py <video>

再传入测得的 --subtitle-y-top/--subtitle-y-bot。坐标基于 ffmpeg 自动旋转后的显示画布,区间为半开 [top, bot),并要求底对齐 ASS 样式;显式设置后,该区域默认使用 60% 透明度的旁白窗口遮罩。

解说模式如需克隆参考声音,使用 --voice-ref <audio>;它与 dub 模式不同。

4.6 最终观看与交付复核

脚本、接点检测、样帧和 QC 报告都不能替代观看。每轮准备交付前,必须检查本轮实际要交付的最终文件,而不是旧别名、无字幕母版或中间代理:

  1. 正常速度完整播放一次短片,不边看边改;先记录真实观看问题。
  2. 播放每个拼接点前后约 0.5–1 秒,检查闪帧、原片叠化被截断、动作跳变和半句原声。
  3. 完整只听声音一次,检查旁白是否碎成一句一停、场景间声音是否接得上、关键原声是否完整。
  4. 单独复看开头、核心情绪/表演点和结尾,确认进入时机、回报停留和收束都成立。
  5. REVISION 分别验证本轮修改项已经改变、冻结项没有意外变化;然后再做解码、时长、音画规格等机械检查。

scene score、亮度统计、contact sheet 与自动 QC 只负责定位候选问题;最终判断以真实播放为准。密集切点的来源判断与处理规则按剪辑技能执行。修复失败时回到剪点、声音或文案层,不用更多包装掩盖。

full/cut 交付如需让确定性的最终检查影响命令退出状态,显式传 --require-final-qc。只有 final_qc.json 与 golden_eval.json 的摘要均为 ok: true 且整数 blocker_count: 0 才打印完成并返回成功;缺失、畸形或 blocker 会保留报告和已渲染诊断媒体,但命令非零退出且不打印完成。默认仍是仅报告、不阻断。 该参数不支持 --edit-mode dub;dub 未传该参数时的准备和渲染行为不变。

4.7 不需要解说的片子

# 对当前整段输入直接合成;不隐式跑理解/ASR/TTS
python3 scripts/recap.py locked_picture.mp4 --work-dir source_work --audio-mode source-mix
# 剪辑计划仍按 cut 流程产生,剪完不再暂停等待 narration.json
python3 scripts/recap.py ep1.mp4 ep2.mp4 --edit-mode cut --work-dir cut_work --audio-mode source-mix
# 只换包装时冻结当前整片 AAC;不允许同时加 BGM/TTS
python3 scripts/recap.py adopted.mp4 --work-dir packaging_work --audio-mode adopted-packet-copy

source-mix 仍会混音和重编码;cut + adopted-packet-copy 冻结的是剪后中间片的声音,不是原片的 AAC 包。 当前严格字幕轨只支持 adopted 模式;其他字幕来源没有因此变成精确对齐。切换声音模式须新工作目录,不得把旧 TTS、QC 或自动生成的解说花字混入本轮原声生产。细节见 references/audio-routing.md。

5. 英译中原声复刻模式

--edit-mode dub 把英文视频翻译为中文,并用原说话者的克隆音色替换人声;它不是在压低原声上叠加解说。

python3 scripts/recap.py <video> --edit-mode dub --work-dir <work_dir>

准备阶段会转写英文、提取一段参考音频,并写出 dub_brief.md 与 dub_transcript.json。Agent 随后写:

[{"start": 0.0, "end": 2.0, "zh": "中文译文"}]

要求:

  • 逐句忠实翻译,不删钩子、不合并、不擅自压缩;原文重复,译文也按时间重复。
  • 每句沿用原声 [start, end],相邻句不重叠。
  • 译文尽量控制在约 5 字/秒,使其能在原时间窗内说完。

重复同一命令后输出 dub_<name>.mp4。每句单独克隆并贴回原时间线;只有即将覆盖下一句时才局部加速。当前版本只支持单说话者、整轨替换,不分离背景音乐。

6. 自检与只读 dashboard

python3 scripts/recap.py --doctor

6.1 只读 dashboard

python3 scripts/dashboard_server.py --root <目录> [--port 0] [--open]

前台运行并打印本机地址(只绑定 127.0.0.1,Agent 启动时放到后台)。它在 --root 下按 library.json、recap_project.json、 recap_run_manifest.json 发现资源库、项目与运行,按阶段显示剪辑节奏、旁白、成片与时间线、QC 和 resource_lock.json。 严格只读:只接受 GET / HEAD,不写任何文件;页面上的「复制给助手」只复制一句请求,改动回到对话里做。

7. 输出与参数

主要输出:

  • recap_<video>.mp4:最终成片。
  • subtitles.srt / subtitles.ass:字幕。
  • work_dir/:全部中间产物,契约见 references/data-schema.md。
  • work_dir/recap_story_plan.json / visual_audio_board.json:Agent 创作意图与剪辑决定。
  • work_dir/mimo_qc.json:可选的建议型复核,不作为发布门禁。

完整参数列表以 python3 scripts/recap.py --help 为准。--style 是原样传给 Agent 的自由文本指导,不是 preset、枚举、开关或有限风格分类。

8. 能力边界

  • 语义评审默认建议型、失败开放;只有调用方显式启用严格解说评审时,事实矛盾、残句或评审不可用才会在 TTS 前阻断。确定性校验阶段始终负责硬校验。
  • MiMo QC 不能阻断、自动修复或改变退出状态,只提供定位建议。
  • 宣发标题、花字或外部文案回填见 video-script 的 references/promotional-copy.md。
Files (oh-story-dsh)
  • references
    • audio-routing.md 7 KB
      # Recap audio routing
      
      `--edit-mode` selects the picture workflow; `--audio-mode` independently
      selects the final audio workflow:
      
      | Audio mode | Narration validation/review | TTS | Assemble input |
      | --- | --- | --- | --- |
      | `narration` (default) | Runs as before | Runs as before | source or rendered cut |
      | `source-mix` | Not applicable | Not run | current source or rendered cut |
      | `adopted-packet-copy` | Not applicable | Not run | current source or rendered cut |
      
      `full` with either non-narration mode goes directly to assemble and does not
      implicitly run understanding. `cut` still requires the existing `clip_plan`:
      the first run analyzes and pauses for that plan, then later runs the established
      single- or multi-source cut renderer. It does not pause for `narration.json`.
      
      ## Local adopted full-sound assembly
      
      The strict local-adoption route accepts one prebuilt picture in `full` mode and
      exactly this all-or-none bundle:
      
      ```text
      --tts-meta PATH --narration-adoption PATH --audio-mix-adoption PATH
      ```
      
      It requires narration mode, stream 0, a new explicit `--work-dir`, and a delivery path
      that does not already exist. `--output-dir` is optional; when omitted, delivery uses the
      new work directory's parent. The route calls only the video-assemble CLI; it does
      not run understanding, script validation, narration review, voiceover, MiMo QC, cut,
      continuation, editor export, or material-cache reuse. Ambient TTS provider and voice
      configuration are inert. Explicit TTS/voice/review/MiMo/editor flags are rejected rather
      than silently ignored.
      
      `recap_run_manifest.json` records the resolved path of all three local artifacts under
      `audio.local_adoption`, and the picture's path plus `{size, mtime_ns}` as
      `source_video_identity`. Analysis settings remain separate and unchanged. Recap validates
      routing and, after the child run, checks that the assembler's binding records reference
      the same adoption files and picture; video-assemble remains authoritative for the adoption
      schemas, PCM shapes, frame clock, mix, bindings, and final media transaction. This path
      executes caller-adopted assets and does not claim that recap authored or approved their
      story, voice, or mix.
      
      ## Keep adopted voice after a cut
      
      The single input above is the **finished picture**, which may contain one or many source
      episodes. Use two existing stages rather than sending an old mix binding into a new cut:
      
      1. Finish the selected cut with `video-cut`: either run `recap.py ... --edit-mode cut`,
         or call that skill's own `scripts/cut.py` from its installed directory, with an
         existing clip plan and source manifest:
      
         ```bash
         python3 <video-cut>/scripts/cut.py ep1.mp4 --work-dir CUT_WORK \
           --sources-manifest SOURCES_JSON
         ```
      
         A single-source cut omits `--sources-manifest`. For already locked frame decisions,
         use the picture-plan path instead. Read `clip_plan_validated.json` (or the locked path's
         `picture_map.json`) before placing sound.
      2. Keep the selected WAVs, `tts_meta.json`, and `narration_adoption.json` when the words,
         WAV bytes, and selected voice are unchanged. Do not run voiceover just to move a line.
         In this explicit mix, new `output_start_sample` values own placement; old generation
         windows in `tts_meta.json` are not the revised output timeline.
      3. Match source dialogue, ambience and music to the new picture. Reuse the prepared bed
         only if that sound and its complete sample clock remain appropriate; otherwise rebuild
         it with the existing `video-assemble` source/score producer. Write a **new** mix adoption
         with the prepared receipt path, total samples and complete WAV placements for the new
         picture. Reordering shots can require moving dialogue too; a mix adoption written for
         the old picture cannot simply be pointed at the new one.
      4. Run the local adopted `full` command in `SKILL.md` with `CUT_WORK/edited_source.mp4` and a new
         assembly work/delivery directory. Inspect the emitted narration placements and subtitles
         against that mixed output. Rebind any precise subtitle track to the new finished audio.
      
      The same steps handle a later revision: keep unchanged picture/voice assets, rebuild only
      the affected bed and placement decisions, then assemble into a new directory. The local
      assembly command still does not resume an old work directory or decide new placements.
      
      ## Identity and resume
      
      `recap_run_manifest.json` binds `{mode, selected_stream_index}` separately from
      understanding/material settings. Changing audio mode or stream requires a new
      work directory; it never reuses narrated TTS accidentally. Legacy manifests
      without the audio field are interpreted as `narration`, stream 0, preserving
      old in-flight narration runs. Continuation commands retain the audio selection.
      
      An unbound work directory containing `narration.json` is ambiguous and cannot
      be adopted by a source mode. Existing run-local QC stages are reset after the
      work-directory policy is validated. Source runs record TTS, narration
      validation, and narration review as `not_applicable`; old `tts_meta.json` and
      narration review files are not read as evidence for the current run.
      
      ## Fail-closed combinations
      
      - `dub` cannot be combined with a non-narration audio mode.
      - Explicit TTS/text-policy options are rejected in source modes rather than
        ignored. Ambient TTS provider/voice configuration is inert because TTS does
        not run, and continuation commands do not promote it into explicit flags.
      - Cut audio modes support selected stream 0 only. Full mode may pass
        another stream to assemble, subject to assemble/export support.
      - Source modes do not support advisory MiMo QC; use `off`.
      - Local adopted full-sound assembly consumes a prebuilt picture in `full` mode; use the
        two-stage workflow above for single- or multi-source cuts. That assembly invocation
        does not resume an old work directory or export an editor draft.
      - `source-mix` rejects an explicit `subtitle_track.json` before analysis or cut,
        because the precise-track contract only binds adopted packet audio.
      
      `--preserve-approved-text` remains opt-in and is forwarded unchanged to
      voiceover and continuation only in narration mode.
      
      Narration may explicitly select `--tts-provider index-tts`; `auto` does not
      select Index TTS. Index TTS cannot be combined with MiMo voice selection,
      `--voice-ref`, or `dub`. Its endpoint and voice remain backend configuration,
      not recap defaults.
      
      ## Packaging and frozen-audio boundary
      
      Source modes never regenerate `visual_overlays.json`. A work directory already
      bound to narration is rejected, preventing narration-derived overlays from
      silently becoming source-mode author input. Otherwise existing overlays are
      treated as caller-authored packaging and preserved.
      
      `adopted-packet-copy` freezes only the selected audio stream on the actual file
      passed to assemble. In `full`, that is the input video. In `cut`, the cut stage
      has already rendered `edited_source.mp4`; packet-copy verification therefore
      applies to that file's audio, **not** to AAC packets from any original source.
      This routing does not prove listening quality or publication approval.
      
    • config-playbook.md 11.6 KB
      # Config playbook (override-only)
      
      The bundle runs **zero-config** with sensible defaults. To change behavior, set the
      environment variables below (or pass the noted CLI flags) — they **override** the defaults.
      Nothing here is required; this is documentation only. The only config file any tool reads is an
      optional `recap_project.json` passed with `--project` (see `resource-library.md`): it resolves to the same
      variables below, and a variable you set that disagrees with the project stops the run. The bundle ships no
      root `CLAUDE.md` (so it never collides with your project/global instructions).
      Defaults below are bundle-level defaults unless a note scopes them to a specific stage.
      
      | Concern | Env var / flag | Default | Notes |
      |---|---|---|---|
      | MiMo API key | `MIMO_API_KEY` | — | **required for the default pipeline**; one key drives ASR + VLM + default MiMo TTS. `tp-*` Token-Plan keys auto-route to the cluster base URL; `sk-*` keys support pay-as-you-go without a subscription |
      | Token-Plan cluster | `MIMO_TOKEN_PLAN_CLUSTER` | `cn` | `cn` / `sgp` / `ams` (only for `tp-*` keys) |
      | VLM / chat model | `MIMO_MODEL` | `mimo-v2.5` | frame VLM + reviewer + consolidate |
      | ASR model | `MIMO_ASR_MODEL` | `mimo-v2.5-asr` | speech-to-text |
      | ASR language | `MIMO_ASR_LANGUAGE` | `auto` | `auto` / `zh` / `en` |
      | ASR window | `ASR_SEGMENT_SECONDS` | `15` | smaller → finer dialogue timestamps (stays under MiMo's 10MB base64 cap) |
      | TTS provider | `TTS_PROVIDER` / `--tts-provider {auto,mimo-tts,fish-audio,index-tts}` | `auto` | `auto` prefers configured MiMo, then Fish Audio, and never picks index-tts; explicit selection is recommended for repeatable runs |
      | MiMo TTS model | `MIMO_TTS_MODEL` | `mimo-v2.5-tts` | MiMo provider only |
      | MiMo voice | `MIMO_TTS_VOICE` / `--mimo-tts-voice` | `冰糖` | |
      | Cloned narration voice | `VOICE_REF` / `--voice-ref` | off | MiMo full/cut only; lazily normalize once, then use `mimo-v2.5-tts-voiceclone`; mutually exclusive with `--mimo-tts-voice`; requires authorization and sends the reference to MiMo |
      | Fish Audio key | `FISH_API_KEY` | — | required when Fish Audio is selected; never written to artifacts or cache metadata |
      | Fish Audio model | `FISH_TTS_MODEL` | `s2.1-pro-free` | free model under Fair Use, no SLA; check Fish Audio's current policy |
      | Fish Audio voice | `FISH_TTS_REFERENCE_ID` | `5653cea4ac83480aaf2bf45406556185`(娱乐扒妹) | optional override for the built-in narration voice; part of the per-segment cache settings |
      | Fish Audio endpoint | `FISH_TTS_API_URL` | `https://api.fish.audio/v1/tts` | returns WAV directly to the existing voiceover pipeline |
      | Self-hosted TTS | `INDEX_TTS_ENDPOINT` / `INDEX_TTS_VOICE` | — | explicit `--tts-provider index-tts` only; the endpoint is never written to disk, and segment `emotion`/style is rejected |
      | TTS transport | `TTS_TIMEOUT` / `TTS_WORKERS` / `TTS_RETRIES` | `300` / `4` / `3` | request timeout, parallel segments, and per-segment retries for all providers |
      | Advisory MiMo QC | `MIMO_QC` / `--mimo-qc {off,pre-assemble,post-render,both}` | `off` | optional subjective review at the selected stage(s), one request per stage. Always fail-open: results only point the agent/user to `mimo_qc.json`, never block or auto-repair |
      | MiMo QC refresh/model | `MIMO_QC_REFRESH` / `--mimo-qc-refresh`; `MIMO_QC_MODEL` | cache on / VLM fallback | the reuse check compares evidence file sizes/mtimes, model and prompt, never absolute paths. Post-render temporarily samples at most 6 JPEGs (≤768px); base64 is never persisted. The standalone adapter requires explicit `mimo_qc.py --live` for network access |
      | Narration block coverage | `NARRATION_COVERAGE_TARGET` / `NARRATION_BLOCK_SECONDS` | `0.7` / `9.0` | current block-recap density controls |
      | Narration speed | `NARRATION_SPEED` | `1.15` | global atempo on the voiceover; set `1.0` for long-form/documentary |
      | Narration authored start | `NARRATION_DELAY_SECONDS` | `0` | the renderer uses the Agent-authored `start` exactly. Set a non-zero value only for legacy drafts; hidden delay can move a validated sentence-boundary entry back into source speech |
      | Source sentence boundary detector | `SOURCE_BOUNDARY_NOISE_THRESHOLD` / `SOURCE_BOUNDARY_MIN_PAUSE` / `SOURCE_BOUNDARY_MAX_ALIGNMENT_ERROR` | `-18dB` / `0.12` / `2.1` | aligns ASR terminal punctuation to short acoustic pauses and writes `speech_boundary_anchors.json`; unsafe narration entries and cut in/out points are blocked. There is no intentional-interrupt override |
      | Cut sentence snapping | `SNAP_CLIP_LINE_END` / `CLIP_START_SNAP_MAX_PREPEND` / `CLIP_START_SNAP_MAX_TRIM` / `CLIP_SNAP_MAX_EXTEND` | on / `1.8` / `0.35` / `2.0` | scene-cut cleanup runs first; reliable sentence/quiet snapping runs last. An ASR-owned edge still inside speech produces `unsafe_clip_sentence_boundary` instead of a half sentence |
      | Mask source subs | `MASK_SOURCE_SUBTITLES` / `SOURCE_SUBTITLE_MASK_POLICY` / `SOURCE_SUBTITLE_MASK_RATIO` | off / `off` / `0.14` | masking requires an explicit policy (`opt_in`, `safe`, or `forced`) and burned recap subtitles. Passing measured `--subtitle-y-top/--subtitle-y-bot` opts in automatically. With `--no-burn-subtitles`, the MP4 stays unmasked |
      | Measured subtitle band | `SUBTITLE_Y_TOP` / `SUBTITLE_Y_BOT`; `--subtitle-y-top/--subtitle-y-bot` | off | half-open `[top, bot)` ffmpeg auto-rotated display-frame pixel coordinates; square-pixel video and bottom ASS alignment (1/2/3) only; use `tools/measure_subtitle.py` to generate source-identified previews and suggestions |
      | Subtitle mask look | `SUBTITLE_MASK_OPACITY` / `SOURCE_SUBTITLE_MASK_TIMING` / `SUBTITLE_MASK_PADDING` | `0.6` / `narration` / `4` | generic ratio masks stay translucent by default; an explicitly measured `--subtitle-y-top/--subtitle-y-bot` band defaults to opaque so source glyphs cannot show through. Set `SUBTITLE_MASK_OPACITY` to override; timing `all` restores full-time masking |
      | Original ducking | `IDLE_ORIG_VOLUME` / `SPEECH_DUCKING_VOLUME` | `1.0` / `0.2` | the original returns to full-volume `IDLE` in deliberate gaps/original blocks, and ducks to `SPEECH` under narration. Inter-beat gaps shorter than `DUCK_BRIDGE_SECONDS` stay ducked so a single narration block does not swell between sentences. `DUCKING_ORIG_VOLUME` (`0.3`) is only the fallback when beats carry no placement info |
      | Foreign source audio | `FOREIGN_SOURCE_AUDIO` | off | set when the original audio is in a language the narration is **not** (e.g. a Japanese drama recapped in Chinese). The under-narration original (`SPEECH_DUCKING_VOLUME` / `ZONE_DUCKING_VOLUME`) drops from `0.2`/`0.12` to `0.05` so the foreign speech doesn't bleed under the narration as 怪音; original-audio gap blocks still play full-volume (`IDLE_ORIG_VOLUME`). Explicit `SPEECH_DUCKING_VOLUME`/`ZONE_DUCKING_VOLUME` still override. Pairs with bring-your-own `user_subtitles.*` for the foreign dialogue |
      | Duck fade | `DUCK_FADE_SECONDS` | `0.3` | attack ramp target. Sentence-safe source restore is shortened to fit wholly inside the measured acoustic pause (`pause_start`→anchor), so it never reveals the previous source sentence tail |
      | Duck bridge | `DUCK_BRIDGE_SECONDS` | `1.5` | inter-beat gaps shorter than this stay ducked inside one narration block; gaps >= this are treated as intentional original-audio blocks and return to `IDLE_ORIG_VOLUME` |
      | Background music | `BGM_PATH` / `BGM_VOLUME` / `BGM_DUCKING_VOLUME` | off / `0.18` / `0.10` | optional looped music bed mixed as its own track; point `BGM_PATH` at any audio file. It ducks to `BGM_DUCKING_VOLUME` under narration |
      | Final loudness | `FINAL_LOUDNORM` / `TARGET_LUFS` | `true` / `-14` | end-of-pipeline normalize |
      | Output compression | `OUTPUT_CRF` / `OUTPUT_PRESET` / `OUTPUT_MAX_HEIGHT` | `18` / `veryfast` / `0` | x264 re-encode controls, applied whenever the final mux re-encodes (burning subtitles / masking / scaling / `FORCE_VIDEO_REENCODE`). Higher `OUTPUT_CRF` = smaller file/lower quality (18≈visually lossless, 23–26 much smaller); `slow`/`slower` preset shrinks more at the same CRF; `OUTPUT_MAX_HEIGHT>0` downscales the final height (keeps aspect, even width), e.g. `720` to halve 1080p pixels. Subtitles/mask render at native res then downscale, so they stay crisp |
      | Style | `--style` | `纪录片` | freeform verbatim guidance passed into the agent brief. The agent synthesizes voice/pacing from this exact text plus evidence; it is not a fixed option list, preset, switch, or finite style taxonomy |
      | Edit mode | `EDIT_MODE` / `--edit-mode` | `full` | `full`, `cut`, or `dub`; multi-source input supports `cut` only |
      | Cut target | `TARGET_DURATION` / `--target-duration` | — | e.g. `10m` (cut mode) |
      | Scene threshold | `--scene-threshold` | `0.1` | scene-cut sensitivity |
      | Shot-change-aware cut | `SCENE_CUT_SNAP` / `SCENE_CUT_SNAP_MARGIN` / `SCENE_CUT_DETECT_THRESHOLD` | on / `0.5` / `0.4` | cut mode: nudge each clip boundary off the original footage's hard cuts so the edit point doesn't flash a sliver of the adjacent shot (闪烁). source_start moves forward onto / source_end back onto any shot-change within the margin; boundaries already on a cut, or that would shrink a clip below ~0.5s, are left as-is. Set `SCENE_CUT_SNAP=0` to disable |
      | VLM workers | `VLM_WORKERS` | `8` | lower to 1 if a proxy/WAF rate-limits |
      | Subtitle size | `SUBTITLE_FONT_SIZE` / `SUBTITLE_MARGIN_V` | `42` / `48` | look & placement |
      | Subtitle font | `SUBTITLE_FONT_NAME` / `SUBTITLE_FONT_FILE` | `Arial` / — | with a font file, burned ASS subtitles load fonts from its directory (`fontsdir`) and on-screen text overlays use it (`fontfile`); `SUBTITLE_FONT_NAME` must be that file's family name. A bound `subtitle_style` template sets both |
      | 整理 / index | `--no-consolidate` / `--consolidate-asr` | on | build the understanding index (and optionally clean ASR); use `--no-consolidate` to skip |
      | Advisory / strict narration review | `REVIEW_NARRATION` / `--review-narration` / `--no-review-narration`; strict: `REQUIRE_NARRATION_REVIEW` / `--require-narration-review` | advisory on, strict off | runs the narration review stage after validation and before TTS. Default advisory mode is fail-open; strict mode blocks TTS on review failure, parse error, or error-severity findings. In cut mode the reviewer uses `clip_plan_validated.json` to remap VLM/ASR grounding onto the output timeline |
      | 剪映 export (optional) | `--export-jianying` / `EXPORT_JIANYING` | off | after rendering, also write a 剪映/JianYing draft from `timeline.json`. Decoupled — the core render never needs it |
      | 剪映 draft dir | `JIANYING_DRAFT_DIR` | work_dir | parent folder for the exported draft (point it at 剪映's drafts root to open in-app) |
      | 剪映 bundle media | `JIANYING_BUNDLE_MEDIA` / `--jianying-no-bundle-media` | **on** | copies video/audio/photo into `Resources/local/*`, uses draft-placeholder paths, and writes the type-0 material index in `draft_meta_info.json`. This makes a cloned/moved draft self-contained and is required on sandboxed macOS. Use `--jianying-no-bundle-media` only if 剪映 can reach the original paths |
      | Source video | `--source-video` | — | original video (cut mode) so `timeline.json` / 剪映 export reference the real source clips instead of the concatenated `edited_source.mp4`; direct `video-assemble` runs intentionally ignore ambient `SOURCE_VIDEO` unless `--source-video` is passed |
      
      `video-assemble` always writes `timeline.json` — a backend-neutral multi-track model
      (video / original-audio / narration / BGM / subtitle, with ducking automation). The
      canonical renderer is ffmpeg; the 剪映 exporter is an optional consumer of the same file. Subtitle text in `timeline.json` is display-ready and follows the same terminal-punctuation policy as SRT/ASS.
      
      See each stage skill's SKILL.md for the full per-stage option list.
      
    • data-schema.md 27.7 KB
      # 数据格式(中间 JSON)
      
      所有中间文件均在 pipeline 工作目录(work_dir/)下。
      
      ## vlm_analysis.json
      
      每场景的 VLM 分析结果,数组格式:
      
      ```json
      [
        {
          "scene_id": 1,
          "start": 5.0,
          "end": 15.0,
          "description": "男子闯入房间",
          "depth_analysis": "角色情绪分析...",
          "frame_facts": {
            "5.0": ["男子闯入房间, 头发蓬乱表情紧张"],
            "10.0": ["男子俯身盯着床上男孩, 男孩睁眼惊醒"]
          }
        }
      ]
      ```
      
      | 字段 | 类型 | 说明 |
      |------|------|------|
      | `scene_id` | int | 场景编号 |
      | `start` | float | 开始时间(秒) |
      | `end` | float | 结束时间(秒) |
      | `description` | string | 画面简述(≤80字) |
      | `depth_analysis` | string | 深层分析(情绪/关系/潜台词) |
      | `frame_facts` | object | 帧级事实,key 为时间戳字符串 |
      
      ## asr_result.json
      
      语音转文字结果:
      
      ```json
      [
        {"start": 0.0, "end": 3.5, "text": "What are you doing here?"}
      ]
      ```
      
      ## asr_writing_chunks.json
      
      由 CLI 在生成 `agent_narration_brief.md` 时自动写出。它把长 ASR 按句子边界拆成适合 Agent 消化的语义块;中文按字符计数,非 CJK 文本按词数计数,并尽量保留 scene 对齐。
      
      ```json
      [
        {
          "chunk_id": 0,
          "start": 0.0,
          "end": 28.5,
          "scene_ids": [0, 1],
          "char_count": 642,
          "text": "第一段对白……",
          "segments": [
            {"start": 0.0, "end": 3.5, "text": "第一句。", "char_count": 4}
          ]
        }
      ]
      ```
      
      ## silence_periods.json
      
      静音窗口列表(适合放解说):
      
      ```json
      [
        {"start": 2.0, "end": 8.5, "duration": 6.5, "has_speech": false}
      ]
      ```
      
      `has_speech` 标记该窗口是否与检测到的 ASR 语音重叠;下游(pipeline / narration)只把 `has_speech=false` 的窗口当作可放解说的安静窗口。
      
      ## speech_boundary_anchors.json
      
      原声句末安全切入点。脚本把 ASR 终止标点的估算时间吸附到短声学停顿;它与
      `silence_periods.json` 分开,因为约 0.1–0.5 秒的停顿适合作为旁白入口,却不足以容纳整段旁白。
      
      ```json
      {
        "schema_version": 1,
        "sentence_anchors": [
          {
            "time": 5.809,
            "confidence": "high",
            "text_tail": "带你重走詹姆斯的二十一年。",
            "pause_start": 5.222,
            "pause_end": 5.809
          }
        ]
      }
      ```
      
      当 `overlaps_speech=true` 且旁白不是从 0 秒冷开场时,`narration` lint 要求 `start`
      贴近 `high`/`medium` 锚点。否则在 TTS 前用 `interrupts_source_sentence` 阻断,并返回
      `suggested_start` 与 `source_text_tail` 给 Agent 调整。常规块使用
      `source_entry_policy: "sentence_boundary"`;原声语句完整性没有抢断 override。最后一个可靠
      锚点之后又进入已声明的原声讲话区时,`suggested_start` 可为 `null`,Agent 必须移动、缩短或删除该旁白块。
      
      剪辑模式第二阶段会另外生成 `speech_boundary_anchors_output.json`,把锚点、ASR 语音区间和
      安静窗口映射到剪后 OUTPUT 时间轴,避免拿原片时间检查剪后旁白:
      
      ```json
      {
        "schema_version": 2,
        "timeline": "cut_output",
        "sentence_anchors": [{"time": 4.0, "pause_start": 3.8, "confidence": "high"}],
        "speech_spans": [{"start": 0.0, "end": 3.8}],
        "quiet_windows": [{"start": 3.8, "end": 4.1}]
      }
      ```
      
      该文件必须不早于当前 `clip_plan_validated.json`(按修改时间判断)。缺失、过期或畸形的
      output 证据一律 fail closed,不能回退到原片时钟或信任 Agent 写入的
      `overlaps_speech=false`。多来源剪辑的每条映射记录还保留 `source_id` 和原片起止时间。
      
      ## timeline_fusion.json
      
      由 CLI 在生成 brief 时自动写出。它把 VLM 场景、ASR 对白和静音窗口按时间轴 overlap 合并,减少写稿时手工推断“这一幕有没有对白/能不能插解说”的成本。
      
      ```json
      [
        {
          "scene_id": 0,
          "time_range": [0.0, 10.0],
          "visual_description": "两人在门口对峙",
          "depth_analysis": "关系紧张",
          "frame_facts": {"1.0": ["女子回头"]},
          "dialogue_segments": [
            {"start": 2.0, "end": 4.0, "overlap_seconds": 2.0, "text": "你到底是谁"}
          ],
          "dialogue_overlap_seconds": 2.0,
          "narration_slots": [
            {"start": 5.0, "end": 7.0, "duration": 2.0, "char_budget": 5}
          ],
          "recommended_mode": "ducked-bed"
        }
      ]
      ```
      
      ## narration.json
      
      Agent 撰写的解说词。full 模式下使用原视频时间;**orchestrated cut 模式(`video-recap --edit-mode cut`)下,第二次暂停时已经先剪出 `edited_source.mp4`,因此 `narration.json` 必须直接使用剪后成片的 OUTPUT 时间轴(0..成片总时长);不存在原视频时间→输出时间的旁白映射产物:
      
      ```json
      [
        {"start": 2.5, "end": 7.0, "narration": "解说文本", "pause_after_ms": 250, "overlaps_speech": true}
      ]
      ```
      
      ## narration_lint.json
      
      `--step script` 或续跑验证 `narration.json` 时生成的预检结果。它检查写稿、时间安全和解说覆盖。`metrics` 为 full 模式下的诊断指标(cut 模式为空对象),不是要求命中某个旁白比例的创作配额;低覆盖 warning 应回到 `visual_audio_board.json` 检查是否为有意的原声/沉默选择。
      
      ```json
      {
        "ok": false,
        "error_count": 1,
        "warning_count": 1,
        "metrics": {
          "segment_count": 12,
          "narration_coverage": 0.68,
          "narration_seconds": 61.2,
          "timeline_seconds": 90.0,
          "avg_block_chars": 48,
          "original_block_count": 4
        },
        "errors": [
          {"level": "error", "index": 2, "code": "time_overlap", "message": "Segment overlaps the previous narration segment"}
        ],
        "warnings": [
          {"level": "warning", "index": 0, "code": "over_budget", "budget_chars": 28, "actual_chars": 42}
        ]
      }
      ```
      
      常见 code:`invalid_time`、`empty_narration`、`time_overlap`、`outside_clip_plan`、`over_budget`、`incomplete_sentence`、`slot_too_short`、`under_narrated`、`over_narrated`、`fragmented_beats`、`no_original_blocks`。
      
      ## recap_story_plan.json / visual_audio_board.json(Agent 创作工作产物)
      
      这两个 JSON 是 skill 层的创作决策记录:前者保存导演意图、备选剪辑假设、POV/主线和 change-based beats;后者保存每拍的画面/表演选择、入点/出点、原声锚点、`audio_owner` 与 `narration_job`。完整字段与工作流见 video-script 技能的 `creative-editing-playbook.md`。
      
      CLI 不以它们作为渲染硬门禁,也不新增解析服务;建议型解说评审在文件存在时读取它们,Agent 则用它们保证 cut、旁白和声音选择没有偏离同一个创作意图。
      
      ## style_card.json(Agent 撰写,可选/按 brief 要求)
      
      `style_card.json` 是表达层契约:由 Agent 根据 `--style`、`--context`、素材证据、ASR 和用户偏好信号综合撰写。`--style` 是 freeform verbatim guidance(原样自由文本指导),不是枚举、preset、switch,也不是一组可穷举风格名;不要把它翻译成固定档位。它是当前版本的活动契约:用户对声音、节奏、字幕阅读或禁忌提出新反馈后,更新原文件并移除过期偏好,不要只改 `narration.json`。
      
      这个文件记录声音、节奏、回收意图和证据支撑的表达判断;字段可以随项目增减,下游只把它当 JSON object 读取,不要求固定键名。它不负责标题、封面、首句承诺或卖点包装。
      
      ```json
      {
        "voice": "冷静但有压迫感,少讲大道理,多用人物动作和台词里的证据推进",
        "pacing": "前 15 秒紧凑建立冲突,每个 beat 连续说完一个思路;中段留原声喘息,结尾回收开头疑问",
        "payoff_intent": "让观众先看到误会,再看到人物选择的代价",
        "subtitle_read_posture": "TTS 保持连续口语,字幕按阅读宽度拆 cue,不用字幕换行切碎朗读",
        "evidence_intent": ["优先引用画面动作", "关键转折保留原声"]
      }
      ```
      
      ## packaging_plan.json(Agent 撰写,可选)
      
      `packaging_plan.json` 是内容锁定后的可选包装层契约:标题、封面帧/视觉钩子、首句、观众承诺、卖点和发布包装信息。它帮助 review 判断“包装承诺”和正文前 15 秒是否对齐;不应反过来驱动故事取舍。
      
      它不是文风策略,不覆盖 `style_card.json` 的声音、节奏或表达规则;如果包装需要某个承诺,正文仍要用素材证据兑现。
      
      ```json
      {
        "title": "一句能对外展示的标题",
        "cover_frame": {"time": 12.4, "reason": "人物第一次正面做出关键选择"},
        "first_line": "开场第一句解说",
        "viewer_promise": "观众看完会明白的冲突/反转/信息增量",
        "selling_points": ["强冲突", "原声高光"],
        "packaging_notes": "发布侧备注,不写文风规则"
      }
      ```
      
      ## deslop_qc_requirements.json(工具/brief 生成的运行契约)
      
      `deslop_qc_requirements.json` 是 tool/brief generated run contract:工具或 brief 生成本次运行的 QC 要求,供 `deslop_qc` 读取,不由 Agent 手写。字段为 `schema_version` 与 `style_card_required`。
      
      `style_card_required` 默认 `false`(advisory):缺少 `style_card.json` 只是 warning,不阻断出片。将来的 opt-in 运行可把它设为 `true`,让 `style_card.json` 成为硬性要求——`deslop_qc` 只读这个字段判断缺少 `style_card.json` 是否是 blocker,不扫描 `agent_narration_brief.md` 的 prompt wording 来推断。如果 requirements 文件缺失或损坏,按 legacy/migration advisory 处理,不作为 hard failure。
      
      该契约不改变 `--style`:`--style` 仍是 freeform verbatim guidance,不增加固定风格档位。它也不改变 `deslop_qc` 边界:仍然是 report-only,不是 AIGC detector,不自动改写。
      
      最小示例:
      
      ```json
      {
        "schema_version": 1,
        "style_card_required": false
      }
      ```
      
      ## deslop_qc.json(CLI 生成,报告型 QC)
      
      `deslop_qc.json` 由本地 deterministic scanner 生成,Agent 不手写。它只是 report-only QC:不是 AIGC detector,不判断文本是不是 AI 写的,不会自动改写。修改仍由 Agent/人工根据报告回到 `narration.json`、`style_card.json` 或字幕源里处理。
      
      报告分两层:
      
      - `blockers`:客观阻断项,会并入 `narration_lint.json` 的 error,例如 requirements 要求但缺少/损坏 `style_card.json`、破折号、占位符泄漏。
      - `advisories`:建议项,只提示可读性/口语化风险,例如模板化“不是……而是……”转折、套话密度、抽象总结词、解释链、比喻标记、过长段落;它们不自动阻断,也不自动改写。
      
      ```json
      {
        "ok": false,
        "contract": "Local readability/QC report only: this is not an AIGC detector, does not claim AI-generation accuracy, and never rewrites text. Corrections remain human/agent rewrite work.",
        "scanner": "deslop_qc.py",
        "style_card_required": true,
        "blocker_count": 1,
        "advisory_count": 1,
        "blockers": [
          {"severity": "blocker", "code": "missing_style_card", "source": "style_card", "index": null, "message": "style_card.json is required by this expression/packaging run but is missing"}
        ],
        "advisories": [
          {"severity": "advisory", "code": "cliche_density", "source": "narration", "index": null, "message": "套话/高频抽象词偏密,建议换成具体行动、选择和后果"}
        ],
        "metrics": {"segments_scanned": 12, "text_units": 860, "sentence_count": 38}
      }
      ```
      
      ## multi_source_manifest.json(多视频 cut)
      
      多视频剪辑模式下,项目级 `work_dir/multi_source_manifest.json` 是编排、剪辑与合成阶段共用的来源契约。`source_id` 由源文件名主干与文件大小派生为 `src_<stem>_<size>`;同一项目里得到相同 id 的后续来源按输入顺序追加 `_2`、`_3` 后缀。`source_video_identity` 记录该文件的 `{size, mtime_ns}`,续跑时与当前输入逐项比对。
      
      ```json
      {
        "schema_version": 1,
        "sources": [
          {
            "source_id": "src_episode1_734003200",
            "source_path": "/abs/episode1.mp4",
            "source_name": "episode1.mp4",
            "source_video_identity": {"size": 734003200, "mtime_ns": 1758326400000000000},
            "source_work_dir": "sources/src_episode1_734003200",
            "material_id": "episode1-734003200"
          }
        ]
      }
      ```
      
      ## clip_plan.json
      
      cut 模式下 Agent 选择要保留的原片片段,数组或 `{ "clips": [...] }` 都可接受。默认片段不能重叠,避免同一原片时间映射到多个输出位置:
      
      ```json
      {
        "target_duration": "10m",
        "clips": [
          {"start": 12.0, "end": 38.0, "reason": "b01 | hook | knowledge: unknown→threat | POV=主角 | 保留倾听反应 | 入点=问题已问出 | 出点=沉默落地"}
        ]
      }
      ```
      
      | 字段 | 类型 | 说明 |
      |------|------|------|
      | `start` | float | 原视频片段开始秒数 |
      | `end` | float | 原视频片段结束秒数 |
      | `reason` | string | 选择该片段的剧情/信息原因 |
      
      多视频 cut 的 `clip_plan.json` 必须给每个片段加 `source_id`;重叠检测按 `source_id` 分开计算,不同源视频的相同时间范围不互相冲突:
      
      ```json
      {
        "target_duration": "10m",
        "clips": [
          {"source_id": "src_episode1_734003200", "start": 12.0, "end": 38.0, "reason": "b01 | setup | knowledge: unknown→clue | POV=主角 | 保留迟疑反应 | 入点=线索出现 | 出点=疑问成立"},
          {"source_id": "src_episode2_689110016", "start": 4.0, "end": 22.0, "reason": "b02 | payoff | power: suspect→hero | POV=主角 | 保留最终选择 | 入点=证据落下 | 出点=代价显现"}
        ]
      }
      ```
      
      ## clip_plan_validated.json
      
      CLI 校验 `clip_plan.json` 后写出,额外包含输出时间轴:
      
      ```json
      {
        "clips": [
          {
            "clip_id": 0,
            "source_start": 12.0,
            "source_end": 38.0,
            "output_start": 0.0,
            "output_end": 26.0,
            "duration": 26.0,
            "reason": "b01 | hook | knowledge: unknown→threat | POV=主角 | 保留倾听反应 | 入点=问题已问出 | 出点=沉默落地"
          }
        ],
        "total_duration": 26.0,
        "target_duration": 600.0
      }
      ```
      
      `qc.boundary_status.sentence_checks` 逐项记录每个片段 start/end 是 `safe`、`unchecked`
      还是 `blocking`。理解阶段已有 ASR 讲话时间时,任何未落到源头/源尾、可靠句末/静音窗,且
      不是同源无损连续连接的边界都会写入 `qc.blocking[].code=unsafe_clip_sentence_boundary`。
      切镜吸附先执行,句末吸附最后执行,保证视觉边界不会覆盖声音安全边界。
      
      多视频 validated clip 会额外保留来源字段,供 pass2 brief、timeline 和剪映导出追溯原素材:
      
      ```json
      {
        "clips": [
          {
            "clip_id": 0,
            "source_id": "src_episode1_734003200",
            "source_path": "/abs/episode1.mp4",
            "source_start": 12.0,
            "source_end": 38.0,
            "output_start": 0.0,
            "output_end": 26.0,
            "duration": 26.0,
            "reason": "b01 | hook | knowledge: unknown→threat | POV=主角 | 保留倾听反应 | 入点=问题已问出 | 出点=沉默落地"
          }
        ]
      }
      ```
      
      ## material library(可选,grep 复用)
      
      `--material-library-dir <dir> --save-materials` 会把每个源视频的已分析小文件复制到 `<dir>/materials/<material_id>/`,不复制原始媒体。`--use-materials` 会在源文件路径、`source_video_identity`(`{size, mtime_ns}`)和分析 `settings` 都相等时把这些 JSON/MD 产物恢复到当前 per-source `work_dir`。
      
      ```text
      .video-materials/
        materials_index.jsonl          # 追加式 grep journal;旧行可保留为历史
        materials/<material_id>/
          material.json                # 当前权威 metadata
          material.md                  # grep 友好摘要
          artifacts/scenes.json
          artifacts/asr_result.json
          artifacts/asr_clean.json       # 启用 --consolidate-asr 时保留,恢复后 brief/review 优先使用
          artifacts/vlm_analysis.json
          artifacts/understanding_index.json
      ```
      
      `materials_index.jsonl` 每次保存追加一行,字段包括 `schema_version`, `event`, `material_id`, `source_name`, `source_path`, `source_video_identity`, `summary`, `tags`, `material_dir`, `updated_at`;`material.json` 另外记录分析 `settings` 字典与每个产物的 `bytes`。当前权威状态始终以 `materials/<material_id>/material.json` 为准。MVP 只承诺 `grep -R "关键词" <library>` 这类文件检索;没有 DB、embedding 或语义搜索。
      
      保存时会对凭证形态(`tp-`/`sk-`/`gh*_`/`AKIA`/JWT 与 `KEY=VALUE` 赋值)和凭证命名的 JSON key 做脱敏,但这只是**尽力而为**的兜底,不是保证:陌生格式的密钥仍可能漏过。请从源头避免把密钥写进分析产物——key 从环境变量/`.env` 读取,不需要落进 scenes/ASR/VLM/summary 等 JSON。
      
      ## original_subtitles.json / user_subtitles.{json,srt,ass}(可选,原声留白字幕)
      
      解说块之间的原声留白会把【原声台词】烧成字幕(assemble 阶段用 `「」` 包裹以区分解说)。来源优先级(高→低):
      
      1. **`work_dir/user_subtitles.json`**(用户自带,最准)— 数组 `[{start,end,text}]` 默认按**成片 OUTPUT** 时间轴直接使用;也可写成 `{"timeline": "source"|"output", "lines": [...]}`,`source` 表示按**原片**时间轴给出,由 assemble 依 `clip_plan_validated.json` 映射到成片。
      2. **`work_dir/user_subtitles.srt` / `.ass`**(用户自带)— 默认按**原片**时间轴解析后映射到成片。
      3. **`work_dir/original_subtitles.json`**(Agent 校对,cut pass2 写)— OUTPUT 时间轴 `[{start,end,text}]`,订正 ASR 错字/人名、只写留白里真正出声的句子。
      4. **ASR 兜底** — 无上述文件时,用 `asr_result.json` 按留白粗略映射(中点估时,可能偏多偏乱)。
      
      来源 1–3 为「精确来源」:每条按句**区间裁剪**落到所覆盖的留白边界(跨边界会拆分),不走 ASR 兜底的中点估时;over-dense 行截断显示而非丢弃。
      
      ```json
      [
        {"start": 2.0, "end": 5.0, "text": "原声台词一句"}
      ]
      ```
      
      ## background_research.json
      
      可选的背景调研结果(由 Agent 使用任意可用搜索/浏览方式整理):
      
      ```json
      {
        "synopsis": "剧情概要",
        "characters": {"角色名": "角色简介"},
        "worldbuilding": "世界观设定",
        "episode_context": "集数上下文",
        "character_details": {
          "角色名": {
            "aliases": ["别名/昵称"],
            "role": "主角|配角|反派|次要角色",
            "relationships": ["与XX是夫妻", "与YY是师徒"]
          }
        },
        "plot_arcs": [
          {"name": "线索名称", "description": "简要描述", "status": "进行中|已解决|伏笔"}
        ],
        "cultural_notes": [
          {"item": "文化梗/典故/时代背景", "explanation": "解释"}
        ]
      }
      ```
      
      > `character_details`、`plot_arcs`、`cultural_notes` 为(可选,新增)字段。仅含 `synopsis`、`characters`、`worldbuilding`、`episode_context` 四个原始字段的旧 JSON 仍然有效。
      
      ## narration_review.json
      
      解说评审阶段输出 LLM-as-judge 结果。旧字段 `verdict/summary/findings` 仍然有效;新增一份 **advisory** 的内容效果 scorecard 与改稿清单。**scorecard 不改变 verdict、也不作硬门禁**——硬门禁仍是 `findings` 里的 error(事实矛盾/残句)经 `--require-narration-review` 严格模式拦截。`verdict` 词表为 `PASS|REVISE|FAIL`,`OK` 作为旧值的兼容别名。
      
      ```json
      {
        "verdict": "PASS|REVISE|FAIL|OK",
        "summary": "总体判断",
        "scorecard": {
          "promise_match": 4, "hook_3s": 4, "first_15s_delivery": 4, "spine_clarity": 4,
          "stakes_escalation": 4, "information_gain": 4, "spoken_language": 4, "sentence_brevity": 4,
          "tts_pacing": 4, "grounding": 4, "original_audio_use": 4, "subtitle_readability": 4
        },
        "hook_candidates_review": [{"candidate": "首句", "type": "suspense", "score": 4, "keep": true}],
        "retention_risk_points": [{"time": "00:28", "risk": "信息重复可能掉人", "fix": "删掉复述画面的句子"}],
        "highest_return_edits": ["最值得先改的一件事"],
        "information_gain_notes": [{"segment": 0, "label": "motive|...|visual_restatement", "note": "证据/改法"}],
        "spoken_language_rewrites": [{"segment": 0, "original": "原句", "rewrite": "口语改写", "why": "为什么更适合听"}],
        "grounding_assertions": [{"segment": 0, "assertion": "人物/关系/因果断言", "source": "visual|asr|research|user_context|unsupported", "risk": "谨慎说明"}],
        "findings": [{"segment": 0, "severity": "warning", "category": "weak_hook", "issue": "问题", "fix": "改法"}]
      }
      ```
      
      > 若 `work_dir` 提供了 `packaging_plan.json` / `recap_story_plan.json` / `visual_audio_board.json`(可选、Agent 撰写),review 会把它们并入评估上下文;缺失时 review 仅基于解说与画面/对白证据评分,行为不受影响。
      
      ## tts_meta.json(partial 失败可见性)
      
      `video-voiceover` 正常输出 `{segments, engine, narration}`。当显式允许 partial TTS(`--allow-partial-tts` / `ALLOW_PARTIAL_TTS=1`)让运行在部分段失败后继续时,失败段不会只埋在日志里,而会写入 `partial` 与 `failures[]`:
      
      ```json
      {
        "segments": [{"index": 0, "start": 0.0, "end": 1.0, "narration": "第一段。", "audio_path": "tts_segments/narr_000.wav"}],
        "engine": "mimo-tts",
        "narration": "narration.json",
        "partial": true,
        "failures": [{"index": 1, "start": 1.0, "end": 2.5, "text": "第二段。", "error": "network timeout"}]
      }
      ```
      
      > 正常(无失败)运行也会带 `"partial": false, "failures": []`。partial 成片只适合预览,不建议直接发布。
      
      `voice` 记录本次实际使用的音色:`{"provider": "mimo-tts", "model": "mimo-v2.5-tts", "voice_id": "冰糖", "reference": null}`。
      用参考音频克隆时 `voice_id` 为 `null`,`reference` 为 `{path, size, mtime_ns}`;Fish Audio 的 `voice_id` 是 reference id,
      index-tts 的是 `INDEX_TTS_VOICE`。旧 `tts_meta.json` 没有这个字段时按只知道 `engine` 处理。
      
      ## resource_lock.json(本次运行用到的资源)
      
      full / cut 流程(含本地采用路径)合成完成后,video-recap 在 `work_dir` 写出 `resource_lock.json`,汇总运行清单、
      `tts_meta.json.voice` 与 `assembly_manifest.json`(BGM 路径、字幕字体)里已有的事实;配置了资源库时,按解析后的路径
      (音色按 provider + voice_id)把每项对上已登记的资源,带出授权与声音授权状态。dub 模式不写。
      
      ```json
      {
        "schema": "video-recap.resource-lock.v1",
        "work_dir": "/abs/work_dir",
        "library": "/abs/library",
        "project": null,
        "templates": [],
        "resources": [
          {"role": "bgm", "path": "/abs/library/resources/bgm/pulse-demo/pulse-demo.wav", "size": 8044, "mtime_ns": 1,
           "detail": {}, "library": {"id": "pulse-demo", "kind": "bgm", "license": "owned", "consent": null}}
        ],
        "attention": [{"code": "license_unknown", "role": "voice", "message": "…"}]
      }
      ```
      
      `role` ∈ `source_video` / `voice` / `bgm` / `subtitle_font`(项目绑定后还会有模板引入的资源)。`attention` 列出需要人确认的项:
      `license_unknown` / `license_restricted`、参考音频的 `consent_unknown` / `consent_denied`,以及配置了资源库但没有登记的
      `unregistered`。它只提示、不阻断;`final_qc.json` 只承载阻断项,不包含这些提示。
      
      组装后,`assembly_manifest.json.audio_segments[]` 另外记录 `fit_status`、`truncated`、
      `truncate_reason`、`placed_audio_duration`、`placed_audio_path`、`source_duck_end`、
      `source_restore_at` 与 `source_handoff_status`。组装阶段从不按时间裁旁白尾音:放不下时用
      `no_safe_fit` 阻断。`placed_audio_path` 是实际写入 canonical `narration.wav` 的完整逐段 PCM;
      `timeline.json`/剪映必须引用它而不是更长的加速前文件。素材时长与序列化后的时间线段长不一致时,
      `assembly_qc.json` 用 `timeline_audio_mismatch` 阻断。原声恢复在 `pause_start` 前保持压低,
      只在实测停顿内渐强,并于 `source_restore_at` 完成,避免渐强提前泄露上一句尾音。
      
      ## dub_lint.json / dub_review.json
      
      Dub 模式下,`dub_script.json` 在 voiceclone **之前**先经过 deterministic lint,把明显不可发布的脚本挡在昂贵的克隆 TTS 之前。空译文、相邻行重叠、时间越界、`room < 0.4s` 等 **error** 会 `verdict=FAIL` 并阻断 render;`fast_speech`、`trim_risk` 等是 warning,不阻断。
      
      每行 voiceclone 原始 WAV 会把中文台词、模型/提示等合成设置和参考音频信息写入相邻的 `*.wav.meta.json`。台词与设置完全相等且 WAV 可读取时,dub render 直接复用并在 `dub_manifest.json.lines[].tts_cache` 记录 `hit`;台词、参考音频、模型或提示变化都会重新合成。
      
      ```json
      {
        "schema_version": 1,
        "verdict": "PASS|FAIL",
        "blocking": false,
        "errors": [],
        "issues": [{"severity": "warning", "code": "fast_speech", "line": 2, "message": "translation is dense (8.4 chars/s)", "start": 1.0, "end": 2.0}],
        "summary": {"lines": 12, "errors": 0, "warnings": 1, "max_chars_per_second": 8.4, "trim_risk_lines": []}
      }
      ```
      
      `dub_review.json` 是脚本级 review scaffold(确定性派生自 lint,语义忠实/语气仍需 agent/人工判断):
      
      ```json
      {
        "schema_version": 1,
        "verdict": "PASS|REVISE|FAIL",
        "checks": {"faithful_to_source": "needs_agent_review", "spoken_chinese": "PASS", "speaker_tone": "needs_agent_review", "timing_fit": "PASS", "platform_fit": "needs_agent_review"},
        "highest_return_edits": [],
        "coverage": {"transcript_windows": 8, "script_lines": 12}
      }
      ```
      
      > CLI:`dub.py --stage lint|review`(无需 video)、`dub.py --print-schema` 打印以上全部 dub artifact 契约;`--stage render` 会在克隆前自动写 `dub_lint.json` / `dub_review.json`,lint 非 PASS 即中止。
      > 最终 `dub_<name>.mp4` 显式输出 48 kHz AAC;不能沿用 `loudnorm` 内部的 96 kHz 分析采样率。
      ## shift-left QC artifacts
      
      `preflight_qc.json`、`final_qc.json`、`golden_eval.json`、`mimo_qc.json` 共用最小 QC 契约;stage 仅允许 `pre_cut` / `post_cut` / `pre_tts` / `post_tts` / `pre_assemble` / `post_render` / `golden`,其中 `mimo_qc.json` 是 artifact 而不是 stage。详见 `shift-left-qc-schema.md`。
      
      `recap.py` 可通过 `--mimo-qc pre-assemble|post-render|both`(默认 `off`)在组装前和/或成片后写 `mimo_qc.json`。每个 stage 最多一次 live request;报告的 `metadata.cache_input`(证据文件的 kind/bytes/mtime_ns、模型、提示与抽帧元数据)与本次完全相等时复用上次结果,`--mimo-qc-refresh` 可刷新。`post_render` 最多临时抽取 6 张、最长边 768px 的 JPEG;base64 只进入请求,不写进 artifact。多 stage 报告聚合在 `metadata.stages`,状态为 `completed` / `cached` / `unavailable` / `failed`,任何状态都不阻断、也不自动修复。关闭功能会清理旧 `mimo_qc.json`,避免陈旧建议被误认为本轮结果。
      
      QC 证据把 `source_asr` 与 `generated_subtitles` 分开:前者只用于源事实/原声时序,后者是本轮旁白派生字幕,不能反过来充当事实证据。多视频项目从 `multi_source_manifest.json` 指向的逐源 work dir 汇集 ASR;`cut_output` 解说评审同样按 `source_id` 映射逐源 VLM/ASR,避免项目根目录没有单一 ASR 文件时产生空证据。
      
      渲染后,`recap.py` 先更新 `preflight_qc.json` 的 `post_render` stage,再运行可选 MiMo 提示,最后写 `final_qc.json` 和 `golden_eval.json`。`final_qc.json` 汇总最终 mp4、`assembly_manifest.json`、`assembly_qc.json`、`visual_qc.json`、`preflight_qc.json`、`mimo_qc.json` 的本地元数据;缺失/空成片、ffprobe 不可用或失败、以及 assembly/visual QC 的客观 blocker 会进入 deterministic blockers。MiMo 和其他 non-deterministic finding 永远不能成为 blocker;客观佐证必须由 deterministic producer 另发 finding。`golden_eval.json` 默认要求 `final_qc.json.ok=true`,也可用 golden fixture 做简单的时长、codec 和必需 artifact 断言。所有 QC metadata/evidence 写入前都经过 `qc_contract.redact_secrets`:secret-looking key/value 会被替换,URL userinfo/query/fragment 会被移除,仅保留必要 host/path 诊断信息。
      
    • resource-library.md 8.3 KB
      # 资源库、模板与样片
      
      资源库是你自己的目录,与素材库共用同一个根目录(`--material-library-dir` / `VIDEO_RECAP_MATERIAL_LIBRARY_DIR`)。
      仓库不内置任何真实 BGM、音效、音色、字体或包装;合成示例见仓库的 `examples/resource-library/`。
      
      ```text
      <library>/
        library.json                                     # {"schema": "video-recap.library.v1", "name": "…"}
        materials/…                                      # 理解分析库(见 data-schema.md),不变
        resources/<kind>/<id>/resource.json (+ 文件)
        templates/<kind>/<id>/v<version>/template.json
        samples/<id>/sample.json (+ 样片文件,或指向库外的绝对路径)
      ```
      
      只读工具(本技能目录下):
      
      ```bash
      python3 scripts/library.py --library-dir <library> check          # 有错误时退出码 1;--json 输出机器可读报告
      python3 scripts/library.py --library-dir <library> list [--kind bgm]
      python3 scripts/library.py --library-dir <library> show <id|id@vN>
      ```
      
      浏览资源、模板与样片(含图片、音频与样片预览)可用只读 dashboard:`python3 scripts/dashboard_server.py --root <library 或其上层目录>`。
      
      工具从不写库。新增或修改记录时直接编辑 JSON,再跑 `check`。**错误**表示该条目不能用;**警告**表示能用但需要人看一眼
      (授权未确认、声音授权未确认、样片不在本机、模板采用后资源文件已变化)。
      
      ## 通用规则
      
      - `id` 与所在目录名一致,只用小写字母、数字和 `. _ -`;模板目录名为 `v<version>`。
      - 顶层字段严格:出现未列出的字段即报错,避免拼错的字段被静默忽略。
      - 文件路径相对记录所在目录,解析后必须仍在库根目录内(`..` 越界、绝对路径和指向库外的符号链接都报错)。
        唯一例外是样片可以写绝对路径,因为成片常放在别的盘上。
      - 文件身份是 `{size, mtime_ns}`,不计算内容哈希。
      - 派生文件(归一化 WAV、转码片段)只出现在 work_dir,不回写为资源。
      
      ## 资源 `resource.json`
      
      ```json
      {
        "schema": "video-recap.resource.v1",
        "id": "pulse-demo",
        "kind": "bgm",
        "title": "合成脉冲底噪(演示)",
        "files": [{"role": "main", "path": "pulse-demo.wav"}],
        "origin": {"creator": "…", "url": "…"},
        "license": {"status": "owned", "terms": "…", "evidence": "…"},
        "tags": ["低频"],
        "notes": ""
      }
      ```
      
      | `kind` | 文件 | 额外字段 |
      |---|---|---|
      | `bgm` / `sfx` | 至少一个音频文件(wav / mp3 / m4a / aac / flac / ogg) | — |
      | `voice` | 可选参考音频 | `voice.provider` ∈ `mimo-tts` / `fish-audio` / `index-tts`,以及 `voice.voice_id` 或一个参考音频;有参考音频时必须写 `consent.status` ∈ `unknown` / `granted` / `denied` |
      | `font` | 至少一个字体文件(ttf / otf / ttc) | 可选 `font.family`、`font.index` |
      | `image` | 至少一个图片(png / jpg / webp),用于 logo、包框、片尾卡 | — |
      
      `license.status` ∈ `unknown` / `owned` / `licensed` / `restricted`,只能由人填写;工具不会因为文件放在库里就认为有授权。
      `unknown` 与 `restricted` 在 `check` 中是警告。
      
      ## 模板 `template.json`
      
      ```json
      {
        "schema": "video-recap.template.v1",
        "id": "clean-white",
        "version": 1,
        "kind": "subtitle_style",
        "title": "白字细描边",
        "canvas": {"width": 900, "height": 1600},
        "params": {
          "font": {"family": "Arial"},
          "size_px": {"value": 52, "provenance": "specified"},
          "max_chars": {"value": 15, "provenance": "measured"},
          "band": {"value": {"y_top": 1280, "y_bot": 1440}, "provenance": "measured"}
        },
        "samples": ["demo-sample"],
        "status": "adopted",
        "adoption": {"date": "2026-09-27", "by": "user", "statement": "用户原话", "scope": "适用范围"}
      }
      ```
      
      - 模板只对 `canvas` 声明的画布有效;换画幅就是新模板,不自动缩放套用。
      - 任何带 `provenance` 的参数都要有 `value`,`provenance` ∈ `measured`(从成片实测)/ `fitted`(反复调出来的)/
        `specified`(人直接给的数值)/ `unknown`。编辑器面板上的读数不是像素,按 `specified` 或 `unknown` 记。
      - 引用资源写 `{"resource": "<id>"}`:`params.font` 必须指向 `font` 资源,图层的 `image` 必须指向 `image` 资源。
      - `subtitle_style` 需要 `params.font`(`resource` 或 `family`)、`params.size_px` 与 `params.max_chars`(每行最多字数,用最长一行在这块画布上校准);
        `size_px × max_chars` 超过画布宽减去两侧各 40px 默认边距时报错。可选 `outline_px`、`shadow_px`、`primary_color`、`outline_color`、`max_lines`、`band`,其他参数名报错;`params.band.value` 若给出,需满足
        `0 <= y_top < y_bot <= canvas.height`。
      - `packaging` 需要非空 `params.layers`,每层有唯一 `name`、`image` 引用与画布内的整数 `rect {x, y, width, height}`;
        可选 `params.safe_rect`。
      - `status` ∈ `draft` / `adopted` / `retired`。`adopted` 必须有 `adoption.date`(YYYY-MM-DD)、`by`、`statement`(用户原话)
        与 `scope`。可选 `adoption.resources` 记录采用时所用资源文件的身份:
      
        ```json
        "resources": {"frame-demo": [{"path": "resources/image/frame-demo/frame-demo.png", "size": 6063, "mtime_ns": 0}]}
        ```
      
        之后文件的大小或修改时间变化,`check` 给出 `changed_since_adoption` 警告:重新采用,或出一个新版本。
        修改时间是本机事实,把库拷到另一台机器会让所有快照都显示为已变化。
      
      ## 样片 `sample.json`
      
      ```json
      {
        "schema": "video-recap.sample.v1",
        "id": "demo-sample",
        "title": "纯色演示片",
        "file": {"path": "demo-sample.mp4"},
        "canvas": {"width": 900, "height": 1600},
        "demonstrates": ["字幕带位置", "画布尺寸"],
        "not_reusable": "这条样片里哪些东西不能照搬",
        "templates": ["clean-white@v1"]
      }
      ```
      
      样片是证据,不是模板:`demonstrates` 写它示范了什么,`not_reusable` 写不能照搬什么(人物、字幕内容、时间码……)。
      `templates` 用 `id@vN` 指回它所示范的模板版本。
      
      ## 项目绑定 `recap_project.json`
      
      ```json
      {
        "schema": "video-recap.project.v1",
        "name": "某系列竖屏解说",
        "library": "../video-library",
        "bindings": {"subtitle_style": "clean-white@v1", "voice": "narrator-demo", "bgm": "pulse-demo"}
      }
      ```
      
      `python3 scripts/recap.py <video> --project <recap_project.json 或所在目录> …` 在任何阶段开始前解析绑定:`library` 相对项目文件;
      模板必须是 `adopted`;资源与模板必须通过 `check`。解析结果只以各阶段已有的设置下发,阶段技能不读资源库:
      
      | 绑定 | 下发为 |
      |---|---|
      | `subtitle_style` | `SUBTITLE_PLAY_RES_X/Y` = 模板画布,`SUBTITLE_FONT_SIZE` ← `size_px`,`SUBTITLE_OUTLINE` ← `outline_px`,`SUBTITLE_SHADOW` ← `shadow_px`,`SUBTITLE_PRIMARY_COLOR` / `SUBTITLE_OUTLINE_COLOR`(ASS `&HAABBGGRR`),`SUBTITLE_MAX_CHARS` / `SUBTITLE_MAX_LINES`;`band` → 底对齐 `SUBTITLE_ALIGNMENT=2` 且 `SUBTITLE_MARGIN_V` = 画布高 − `y_bot`;`font.family` → `SUBTITLE_FONT_NAME`,字体资源 → 再加 `SUBTITLE_FONT_FILE` |
      | `voice` | provider → `--tts-provider`;MiMo 预置音色 → `--mimo-tts-voice`,参考音频 → `--voice-ref`;Fish Audio → `FISH_TTS_REFERENCE_ID`;index-tts → `INDEX_TTS_VOICE` |
      | `bgm` | `BGM_PATH` ← 该资源的第一个文件 |
      | `packaging` | 合成前写出 `work_dir/packaging_layers.json`:每个图层的图片资源第一个文件 + `rect`,由 video-assemble 叠加到成片并写进 `timeline.json` 的 image 轨 |
      
      - 你已经显式设置的参数或环境变量与绑定不一致时,运行在开始前停止并指出是哪一项,不会静默覆盖。
      - 合成前核对模板画布与实际成片画布;不一致即停止——换画幅要用另一个模板。
      - 参考音频的 `consent.status` 为 `denied` 时拒绝绑定;dub 模式与本地采用三件套不接受 `--project`。
      - 续跑命令只带 `--project`,不重复写出由绑定得到的值,改了绑定后续跑会按新绑定解析。
      
      ## 运行记录
      
      每次 full / cut 合成后,`work_dir/resource_lock.json` 记下这次用到的原片、音色、BGM 与字幕字体,并对上资源库里的登记和授权状态;
      格式见 `data-schema.md` 的 resource_lock.json 一节。
      
    • shift-left-qc-schema.md 2.2 KB
      # 前置 QC 数据契约
      
      `final_qc.json`、`golden_eval.json`、`mimo_qc.json` 与 `preflight_qc.json` 共用一套最小结构,由本技能的 `scripts/qc_contract.py` 实现。
      
      ## 字段契约
      
      - `schema_version`:整数 `1`。
      - `artifact`:只能是 `final_qc.json`、`golden_eval.json`、`mimo_qc.json` 或 `preflight_qc.json`。
      - `stage`:只能是 `pre_cut`、`post_cut`、`pre_tts`、`post_tts`、`pre_assemble`、`post_render`、`golden`。
        - 不得使用 `pre_voiceover`、`post_voiceover`、`pre_export`、`post_export`、`final`、`golden_eval` 或 `mimo_qc` 作为阶段值。
        - `mimo_qc.json` 是产物名;MiMo finding 必须挂在 `post_tts`、`pre_assemble` 等真实阶段上。
      - `findings[]`:每项必须包含 `finding_id`、`stage`、`severity`、`blocking`、`deterministic`、`confidence`、`rule_id`、`decision_reason`、`location`、`evidence`、`sample_policy`、`model_used` 与 `next_action`。
      - `sample_policy`:至少包含 `type`;其值只能是 `all`、`deterministic`、`sampled`、`semantic` 或 `aesthetic`。
      - `location.timecode` 与 `location.source_span` 必须存在;剪辑前阶段可以把任一字段设为 `null`。
      - 为兼容与诊断,当前辅助函数还可能写出 `id`、`category`、`code`、`message`、`source` 与 `objective_corroboration`。
      
      ## 阻断语义
      
      只有确定性的客观规则可以阻断,例如产物缺失或过期、时长或媒体流错误、字幕与 TTS 放置错误、数据结构无效。
      
      MiMo 语义/审美 finding 以及其他所有非确定性 finding **始终只给建议,不能阻断**。若存在客观佐证,必须由确定性生产者单独写成确定性 finding;它不能把主观模型观察升级成 blocker。运行时不存在 allow-list 或规则表逃生口。
      
      `qc_contract.py` 只负责数据结构:不调用 MiMo、不连接流水线,也不自动修复。独立的 `mimo_qc.py` 只有在显式开启时,才会在 `pre_assemble` 和/或 `post_render` 各发起一次实时请求;它保持 fail-open,且不落盘凭证或抽帧 base64。所有报告构建器都通过 `qc_contract.redact_secrets` 处理元数据与证据:脱敏疑似密钥的键值,并从 URL 移除 userinfo、query 与 fragment,同时保留有用的 host/path 上下文。
      
  • scripts
    • qc
      • mimo_client.py 410 B
        """Load this skill's MiMo client without depending on the ambient ``lib`` module."""
        
        import importlib.util
        from pathlib import Path
        
        _SPEC = importlib.util.spec_from_file_location(
            "video_recap_mimo_qc_lib", Path(__file__).resolve().parents[1] / "lib.py"
        )
        _LIB = importlib.util.module_from_spec(_SPEC)
        _SPEC.loader.exec_module(_LIB)
        
        DEFAULT_CONFIG = _LIB.CONFIG
        mimo_qc_api_call = _LIB.mimo_qc_api_call
        
      • mimo_contract.py 239 B
        """Shared constants for the self-contained MiMo QC adapter."""
        
        ARTIFACT_NAME = "mimo_qc.json"
        DEFAULT_STAGE = "pre_assemble"
        MAX_OBSERVATIONS = 12
        MAX_MESSAGE_CHARS = 800
        MAX_FRAMES = 6
        MAX_FRAME_DIMENSION = 768
        FRAME_SAMPLER_VERSION = 2
        
      • mimo_evidence.py 9.5 KB
        """Collect and redact local evidence for MiMo QC."""
        
        from __future__ import annotations
        
        import json
        import os
        from pathlib import Path
        from typing import Any
        from collections.abc import Mapping, Sequence
        
        import qc_contract
        from qc.mimo_client import DEFAULT_CONFIG
        
        _JSON_ARTIFACTS = (
            "narration.json",
            "visual_overlays.json",
            "clip_plan_validated.json",
            "clip_plan.json",
            "assembly_manifest.json",
            "tts_meta.json",
        )
        
        _SOURCE_ASR_ARTIFACTS = (
            "asr_clean.json",
            "asr.json",
            "asr_result.json",
            "asr_segments.json",
        )
        
        _GENERATED_SUBTITLE_ARTIFACTS = (
            "subtitles.json",
            "subtitle.json",
            "subtitles.srt",
            "subtitle.srt",
            "subtitles.vtt",
            "subtitle.vtt",
            "output.srt",
            "output.vtt",
        )
        
        _OPTIONAL_VISUAL_METADATA = (
            "sampled_frames.json",
            "frame_samples.json",
            "storyboard.json",
            "storyboard_meta.json",
            "frames_manifest.json",
        )
        
        
        def _load_json(path: Path) -> Any:
            return json.loads(path.read_text(encoding="utf-8"))
        
        
        def _read_text_sample(path: Path, *, max_chars: int = 4000) -> dict[str, Any]:
            text = path.read_text(encoding="utf-8", errors="replace")
            return {
                "kind": "text",
                "bytes": path.stat().st_size,
                "truncated": len(text) > max_chars,
                "sample": text[:max_chars],
            }
        
        
        def _summarize(
            value: Any, *, max_items: int = 8, max_string: int = 700, depth: int = 0
        ) -> Any:
            """Keep request/report evidence bounded while retaining useful structure."""
            if depth >= 4:
                return {"type": type(value).__name__, "omitted": True}
            if isinstance(value, Mapping):
                out = {
                    str(key): _summarize(
                        item, max_items=max_items, max_string=max_string, depth=depth + 1
                    )
                    for key, item in list(value.items())[:max_items]
                }
                if len(value) > max_items:
                    out["_omitted_keys"] = len(value) - max_items
                return out
            if isinstance(value, list):
                return {
                    "count": len(value),
                    "items": [
                        _summarize(
                            item, max_items=max_items, max_string=max_string, depth=depth + 1
                        )
                        for item in value[:max_items]
                    ],
                    "omitted": max(0, len(value) - max_items),
                }
            if isinstance(value, str) and len(value) > max_string:
                return {"text": value[:max_string], "truncated": True, "chars": len(value)}
            return value
        
        
        def _collect_file(work_dir: Path, name: str) -> dict[str, Any] | None:
            path = work_dir / name
            if not path.is_file():
                return None
            if path.suffix.lower() == ".json":
                kind, summary = "json", _summarize(_load_json(path))
            else:
                kind, summary = "text", _summarize(_read_text_sample(path))
            stat = path.stat()
            return {
                "path": name,
                "kind": kind,
                "bytes": stat.st_size,
                "mtime_ns": stat.st_mtime_ns,
                "summary": summary,
            }
        
        
        def _first_existing(work_dir: Path, names: Sequence[str]) -> str | None:
            return next((name for name in names if (work_dir / name).is_file()), None)
        
        
        def _collect_multi_source_asr(work_dir: Path) -> dict[str, Any]:
            """Collect per-source ASR when a multi-source project has no root transcript."""
            manifest_path = work_dir / "multi_source_manifest.json"
            if not manifest_path.is_file():
                return {}
            collected = {}
            for source in _load_json(manifest_path)["sources"]:
                relative_dir = source["source_work_dir"]  # recap writes sources/<source_id>
                name = _first_existing(work_dir / relative_dir, _SOURCE_ASR_ARTIFACTS)
                if name is not None:
                    collected[source["source_id"]] = _collect_file(work_dir, f"{relative_dir}/{name}")
            return collected
        
        
        def _resolve_candidate(work_dir: Path, candidate: str | Path) -> Path:
            path = Path(candidate)
            return path if path.is_absolute() else work_dir / path
        
        
        def _final_output_path(
            work_dir: Path, final_output: str | Path | None
        ) -> tuple[Path, str] | None:
            """(path, display) of the final output: the caller's, else assembly_manifest.final_output
            (video-assemble always writes it); None before the assembler has run."""
            if final_output is None:
                manifest_path = work_dir / "assembly_manifest.json"
                if not manifest_path.is_file():
                    return None
                final_output = _load_json(manifest_path)["final_output"]
            return _resolve_candidate(work_dir, final_output), str(final_output)
        
        
        def _final_output_metadata(
            work_dir: Path, final_output: str | Path | None = None
        ) -> dict[str, Any]:
            resolved = _final_output_path(work_dir, final_output)
            if resolved is None:
                return {"candidates": []}
            path, display = resolved
            item: dict[str, Any] = {"path": display, "exists": path.is_file()}
            if item["exists"]:
                stat = path.stat()
                item.update({"bytes": stat.st_size, "mtime_ns": stat.st_mtime_ns})
            return {"candidates": [item]}
        
        
        def _existing_final_output(
            work_dir: Path, final_output: str | Path | None
        ) -> Path | None:
            resolved = _final_output_path(work_dir, final_output)
            return resolved[0] if resolved is not None and resolved[0].is_file() else None
        
        
        def collect_evidence(
            work_dir: str | Path, *, final_output: str | Path | None = None
        ) -> dict[str, Any]:
            """Collect lightweight, secret-scrubbed evidence from one work directory."""
            root = Path(work_dir)
            artifacts: dict[str, Any] = {}
            preferred_plan = _first_existing(
                root, ("clip_plan_validated.json", "clip_plan.json")
            )
            for name in _JSON_ARTIFACTS:
                if (
                    name in {"clip_plan_validated.json", "clip_plan.json"}
                    and name != preferred_plan
                ):
                    continue
                item = _collect_file(root, name)
                if item is not None:
                    artifacts[name] = item
            evidence = {
                # Display only; excluded from cache_input so moving the work directory is a cache hit.
                "work_dir": str(root),
                "artifacts": artifacts,
                "source_asr": {
                    name: item
                    for name in _SOURCE_ASR_ARTIFACTS
                    if (item := _collect_file(root, name)) is not None
                },
                "generated_subtitles": {
                    name: item
                    for name in _GENERATED_SUBTITLE_ARTIFACTS
                    if (item := _collect_file(root, name)) is not None
                },
                "visual_metadata": {
                    name: item
                    for name in _OPTIONAL_VISUAL_METADATA
                    if (item := _collect_file(root, name)) is not None
                },
                "final_output": _final_output_metadata(root, final_output),
            }
            if not evidence["source_asr"]:
                evidence["source_asr"] = _collect_multi_source_asr(root)
            # The evidence goes into the MiMo request as well as the persisted report, so it is
            # redacted here, before either.
            return qc_contract.redact_secrets(evidence)
        
        
        def _cache_file_group(group: Mapping[str, Any]) -> dict[str, Any]:
            return {
                name: {key: item[key] for key in ("kind", "bytes", "mtime_ns")}
                for name, item in sorted(group.items())
            }
        
        
        def _cache_evidence(evidence: Mapping[str, Any]) -> dict[str, Any]:
            """The per-file identities ({kind, bytes, mtime_ns}) a cached stage report was built
            from; work_dir is excluded so moving the directory is still a cache hit."""
            return {
                "artifacts": _cache_file_group(evidence["artifacts"]),
                "source_asr": _cache_file_group(evidence["source_asr"]),
                "generated_subtitles": _cache_file_group(evidence["generated_subtitles"]),
                "visual_metadata": _cache_file_group(evidence["visual_metadata"]),
                "final_output": [
                    {key: item[key] for key in ("exists", "bytes", "mtime_ns") if key in item}
                    for item in evidence["final_output"]["candidates"]
                ],
            }
        
        
        def _effective_config(config: Mapping[str, Any] | None = None) -> dict[str, Any]:
            source = dict(DEFAULT_CONFIG)
            if config:
                source.update(dict(config))
                if not config.get("mimo_qc_model"):
                    source["mimo_qc_model"] = (
                        config.get("mimo_video_model")
                        or config.get("mimo_model")
                        or source["mimo_qc_model"]
                    )
            # MIMO_QC_MODEL is intentionally read at call time for embedded/CLI tests and
            # long-running agent processes whose environment may be adjusted between runs.
            if os.environ.get("MIMO_QC_MODEL") and not (config and config.get("mimo_qc_model")):
                source["mimo_qc_model"] = os.environ["MIMO_QC_MODEL"]
            return source
        
        
        def safe_mimo_config(config: Mapping[str, Any] | None = None) -> dict[str, Any]:
            """Return non-secret settings suitable for persisted report metadata."""
            source = _effective_config(config)
            keep = (
                "api_provider",
                "mimo_api_url",
                "mimo_api_url_source",
                "mimo_video_api_url",
                "mimo_video_api_url_source",
                "mimo_qc_model",
                "mimo_qc_model_source",
                "mimo_model",
                "mimo_model_source",
                "mimo_video_model",
                "mimo_video_model_source",
                "mimo_disable_thinking",
                "mimo_disable_thinking_source",
                "mimo_media_resolution",
                "mimo_media_resolution_source",
            )
            safe = {key: source[key] for key in keep}
            safe["model"] = (
                source["mimo_qc_model"] or source["mimo_video_model"] or source["mimo_model"]
            )
            key_present = bool(
                source.get("mimo_video_api_key")
                or source.get("mimo_api_key")
                or source.get("api_key")
            )
            safe = qc_contract.redact_secrets(safe)
            safe["key_present"] = key_present
            return safe
        
      • mimo_observations.py 6.3 KB
        """Normalize MiMo QC observations into the local contract."""
        
        from __future__ import annotations
        
        import json
        import re
        from typing import Any
        from collections.abc import Mapping
        
        import qc_contract
        from qc.mimo_evidence import _summarize, safe_mimo_config
        from qc.mimo_payload import _strip_json_fence
        from qc.mimo_contract import (
            ARTIFACT_NAME,
            DEFAULT_STAGE,
            MAX_MESSAGE_CHARS,
            MAX_OBSERVATIONS,
        )
        
        
        def _extract_observations(model_output: Any) -> list[Mapping[str, Any]]:
            """Bound whatever shape the model (or an offline fixture) returned to observation dicts."""
            if isinstance(model_output, str):
                try:
                    model_output = json.loads(_strip_json_fence(model_output))
                except ValueError:
                    return [
                        {
                            "code": "freeform_observation",
                            "message": model_output,
                            "confidence": "low",
                        }
                    ]
            if isinstance(model_output, list):
                return [item for item in model_output if isinstance(item, Mapping)][
                    :MAX_OBSERVATIONS
                ]
            if not isinstance(model_output, Mapping):
                return []
            for key in ("observations", "findings"):
                if isinstance(model_output.get(key), list):
                    return [item for item in model_output[key] if isinstance(item, Mapping)][
                        :MAX_OBSERVATIONS
                    ]
            if "choices" in model_output:  # a raw chat-completion response saved as a fixture
                return _extract_observations(model_output["choices"][0]["message"]["content"])
            return [model_output] if model_output else []
        
        
        def _norm_choice(raw: Any, allowed: set[str], default: str) -> str:
            value = str(raw or "").strip().lower()
            return value if value in allowed else default
        
        
        def _caption_match_text(value: Any) -> str:
            return re.sub(r"[^0-9A-Za-z㐀-鿿]+", "", str(value or "")).lower()
        
        
        def _generated_subtitle_corpus(payload: Mapping[str, Any]) -> str:
            strings = []
        
            def visit(value: Any) -> None:
                if isinstance(value, str):
                    strings.append(value)
                elif isinstance(value, Mapping):
                    for child in value.values():
                        visit(child)
                elif isinstance(value, list):
                    for child in value:
                        visit(child)
        
            visit(payload["evidence"]["generated_subtitles"])
            return _caption_match_text(" ".join(strings))
        
        
        def _misclassified_generated_caption(
            observation: Mapping[str, Any], payload: Mapping[str, Any], stage: str
        ) -> bool:
            """Drop an objectively contradicted 'source subtitle visible' observation.
        
            MiMo occasionally calls the intended recap cue on an opaque mask band a leftover source
            subtitle even when its own ``visible_text`` exactly matches generated_subtitles.  This is not
            a subjective disagreement: the artifact provides direct provenance for that text.
            """
            if stage != "post_render":
                return False
            code_and_message = (
                f"{observation.get('code', '')} {observation.get('message', '')}".lower()
            )
            if not any(
                token in code_and_message
                for token in ("source_subtitle", "source subtitle", "源字幕")
            ):
                return False
            if any(
                token in code_and_message for token in ("overlap", "double", "两层", "重叠")
            ):
                return False
            evidence = observation.get("evidence")
            visible = evidence.get("visible_text") if isinstance(evidence, Mapping) else None
            visible = _caption_match_text(visible)
            corpus = _generated_subtitle_corpus(payload)
            return len(visible) >= 4 and bool(corpus) and visible in corpus
        
        
        def normalize_observations(
            model_output: Any,
            *,
            payload: Mapping[str, Any],
            stage: str = DEFAULT_STAGE,
            model_config: Mapping[str, Any] | None = None,
        ) -> list[dict[str, Any]]:
            """Normalize bounded model output into permanently non-blocking findings."""
            cfg = safe_mimo_config(model_config)
            model_used = cfg["model"]
            findings = []
            for index, observation in enumerate(_extract_observations(model_output), start=1):
                if _misclassified_generated_caption(observation, payload, stage):
                    continue
                obs = _summarize(observation, max_items=10, max_string=MAX_MESSAGE_CHARS)
                category_hint = str(
                    obs.get("category") or obs.get("type") or "semantic"
                ).lower()
                category = "mimo_aesthetic" if "aesthetic" in category_hint else "mimo_semantic"
                code = str(
                    obs.get("code") or obs.get("rule_id") or f"mimo_observation_{index}"
                ).strip()
                code = (code or f"mimo_observation_{index}")[:96]
                message = str(
                    obs.get("message")
                    or obs.get("summary")
                    or obs.get("text")
                    or "MiMo QC observation"
                )
                message = message.strip()[:MAX_MESSAGE_CHARS] or "MiMo QC observation"
                confidence = _norm_choice(
                    obs.get("confidence"), {"low", "medium", "high"}, "low"
                )
                sample_type = _norm_choice(
                    obs.get("sample_policy") or obs.get("sample_policy_type"),
                    {"semantic", "aesthetic", "sampled"},
                    "aesthetic" if category == "mimo_aesthetic" else "semantic",
                )
                raw_evidence = obs.get("evidence")
                evidence = {
                    **(raw_evidence if isinstance(raw_evidence, Mapping) else {}),
                    "model": model_used,
                    "config": cfg,
                }
                location = obs.get("location")
                findings.append(
                    qc_contract.build_finding(
                        finding_id=str(
                            obs.get("finding_id")
                            or obs.get("id")
                            or f"mimo-{stage}-{index:03d}"
                        )[:160],
                        stage=stage,
                        severity="advisory",
                        confidence=confidence,
                        sample_policy={"type": sample_type},
                        category=category,
                        code=code,
                        message=message,
                        deterministic=False,
                        blocking=False,
                        source={"artifact": ARTIFACT_NAME, "adapter": "mimo_qc.py"},
                        location=location if isinstance(location, Mapping) else {},
                        evidence=evidence,
                        model_used=model_used,
                        next_action="human_review",
                        decision_reason=message,
                    )
                )
            return findings
        
      • mimo_payload.py 10.4 KB
        """Build bounded semantic and multimodal MiMo QC payloads."""
        
        from __future__ import annotations
        
        import json
        from typing import Any
        from collections.abc import Mapping, Sequence
        
        from qc.mimo_evidence import safe_mimo_config
        from qc.mimo_contract import ARTIFACT_NAME, DEFAULT_STAGE
        
        
        def _semantic_evidence(evidence: Mapping[str, Any], *, stage: str) -> dict[str, Any]:
            """Keep real artifact values visible to MiMo without sending irrelevant paths.
        
            ``collect_evidence`` already bounds every artifact independently.  The additional
            summarization here therefore needs enough depth to retain the narration/ASR/TTS
            scalars nested inside those summaries.  A shallow second pass used to replace the
            values with opaque placeholders, which made live QC invent missing-script and failed-TTS
            findings from evidence it could no longer read.
            """
            semantic = dict(evidence)
            semantic.pop("work_dir", None)
            semantic["evidence_roles"] = {
                "narration.json": (
                    "Planned recap voiceover on the OUTPUT timeline; judge its factual and temporal "
                    "fit rather than expecting verbatim source dialogue."
                ),
                "tts_meta.json": (
                    "Synthesis and placement metadata for narration.json; an empty failures list means "
                    "all requested narration segments synthesized successfully."
                ),
                "source_asr": (
                    "Transcript of the SOURCE media audio, not a transcript of generated TTS. "
                    "Wording differences from narration are expected; use this evidence for factual "
                    "support and original-dialogue timing, not verbatim equality."
                ),
                "generated_subtitles": (
                    "Recap subtitles derived from generated narration/TTS during assembly. They are not "
                    "source ASR and must never be used as independent factual support. Visible text that "
                    "matches one of these cues is the intended generated caption, including when it sits "
                    "on the black source-subtitle mask band."
                ),
                "visual_metadata": (
                    "Diagnostic storyboard/frame metadata. labels_burned refers only to diagnostic "
                    "timestamp labels and does not mean labels were burned into the final video."
                ),
                "final_output": (
                    "Only present for post_render and limited to candidates that actually exist."
                ),
                "post_render_frame_limits": (
                    "Sampled final-output frames are silent still images: they contain no audible audio. "
                    "Do not infer which audio track is playing from visible source captions. Narration/TTS "
                    "timings are segment-level, not word-level, so an internal phrase cannot be assigned to "
                    "an exact sampled-frame timestamp."
                ),
                "narration_and_subtitle_gaps": (
                    "Generated narration/subtitle gaps may be deliberate original-audio blocks. In those "
                    "gaps the source audio and its existing captions can intentionally return; do not call "
                    "the absence of generated subtitles a defect without evidence of accidental silence."
                ),
                "actual_audio_timing": (
                    "narration.json start/end values are authoring slots, not proof that speech fills the "
                    "whole slot. For post-render timing, assembly_manifest audio_segments "
                    "actual_place_start/actual_place_end are authoritative. A sampled frame after "
                    "actual_place_end is in an original-audio gap."
                ),
            }
            if stage != "post_render":
                semantic.pop("final_output", None)
                # These files can be leftovers from an earlier assembly when narration is being revised.
                # Pre-assemble QC must evaluate current narration/TTS, not stale rendered captions.
                semantic.pop("generated_subtitles", None)
                semantic["evidence_roles"].pop("generated_subtitles", None)
                semantic["evidence_roles"].pop("post_render_frame_limits", None)
                semantic["evidence_roles"].pop("actual_audio_timing", None)
            else:
                semantic["final_output"] = {
                    "candidates": [
                        dict(candidate)
                        for candidate in semantic["final_output"]["candidates"]
                        if candidate["exists"]
                    ]
                }
            # Artifact summaries are already bounded at collection time; summarizing them again
            # would wrap their ``items`` arrays in another summary layer and obscure the values.
            return semantic
        
        
        def build_payload(
            evidence: Mapping[str, Any],
            *,
            stage: str = DEFAULT_STAGE,
            config: Mapping[str, Any] | None = None,
        ) -> dict[str, Any]:
            """Build the report-safe semantic payload (never contains image base64)."""
            cfg = safe_mimo_config(config)
            payload = {
                "stage": stage,
                "artifact": ARTIFACT_NAME,
                "model": cfg["model"],
                "config": cfg,
                "instructions": (
                    "你是 video-recap 的建议性质量审阅器。结合解说、剪辑计划、字幕/ASR、TTS、"
                    "组装元数据和抽样画面,指出语义或审美问题。只返回主观观察,不做自动修复,"
                    "只根据可见的实际字段判断,不得从文件字节数、截断或省略标记推断内容缺失;"
                    "source_asr 是源素材原声证据,不是生成后 TTS 的逐字转录;generated_subtitles 是本轮生成字幕,"
                    "二者角色不可混淆,解说改写与源台词不一致本身不是问题;"
                    "空 failures 表示没有失败,storyboard 的 labels_burned 仅表示诊断图上的时间标签,"
                    "不代表标签烧进最终视频;pre_assemble 阶段尚无 final_output 属于正常状态;"
                    "成片抽样画面是带时间戳的稀疏点样本,只能判断该具体时刻,不得用单帧否定整段内其他时刻的画面。"
                    "这些抽样帧是无声静帧,不能从画面中的源字幕推断当时实际播放哪条音轨;"
                    "源字幕遮罩策略为 off 时,保留的源字幕无需与解说逐字一致;解说/TTS 只有段级时间,"
                    "没有词级对齐,不得把句中某个短语强行对应到抽样帧的精确秒数。"
                    "生成解说和字幕之间的空档可能是刻意保留的原声块,此时原声和源字幕回归是正常设计,"
                    "没有意外静音证据时不得把生成字幕空档当成缺失。只输出可采取行动的疑似问题,"
                    "不要把正常、一致、相符或通过项作为 observation。不得提出阻断决定。"
                    "narration.json 的 start/end 是创作槽位,不表示旁白铺满整段;成片阶段必须以 "
                    "assembly_manifest.audio_segments 的 actual_place_start/actual_place_end 判断实际旁白。"
                    "抽样点超过 actual_place_end 时属于原声块;一个字幕时间段含多个分句时,静帧匹配其中任一分句都不算错位。"
                    "只有同一编号静帧中同时清晰可读两层不同字幕,才能报告字幕重叠;不同静帧分别出现解说字幕和源字幕不算重叠。"
                    "遮罩 opacity=1.0 表示源字幕像素已被不透明覆盖,不得臆测遮罩下仍有可见文字。"
                    "黑色遮罩带上与 generated_subtitles cue 一致的白字是本轮预期生成字幕,不是残留源字幕;"
                    "必须先逐字对照 generated_subtitles,再判断是否另有第二层源字幕。"
                    "每张抽样图都由同编号 BEGIN/END 文本包围,必须按编号和时间戳独立判断,"
                    "不得把一张图里的文字或人物归到另一张图。"
                    '返回 JSON:{"observations":[{"code":...,'
                    '"message":...,"category":"semantic|aesthetic",'
                    '"confidence":"low|medium|high","sample_policy":'
                    '"semantic|aesthetic|sampled","evidence":{...}}]}。最多 12 条。'
                ),
                "evidence": _semantic_evidence(evidence, stage=stage),
            }
            return payload
        
        
        def _request_payload(
            payload: Mapping[str, Any], frame_samples: Sequence[Mapping[str, Any]]
        ) -> dict[str, Any]:
            request_evidence = {
                "stage": payload["stage"],
                "instructions": payload["instructions"],
                "evidence": payload["evidence"],
            }
            content: list[dict[str, Any]] = [
                {
                    "type": "text",
                    "text": json.dumps(
                        request_evidence, ensure_ascii=False, separators=(",", ":")
                    ),
                }
            ]
            for index, sample in enumerate(frame_samples, start=1):
                timestamp_label = f"{sample['timestamp']:.3f}s"
                content.append(
                    {
                        "type": "text",
                        "text": (
                            f"BEGIN QC_FRAME_{index}: final-output timestamp {timestamp_label}. "
                            f"Judge only this image and do not transfer its content to another frame."
                        ),
                    }
                )
                content.append({"type": "image_url", "image_url": {"url": sample["data_url"]}})
                content.append(
                    {
                        "type": "text",
                        "text": f"END QC_FRAME_{index}: timestamp {timestamp_label}.",
                    }
                )
            return {
                "model": payload["model"],
                "messages": [{"role": "user", "content": content}],
                "max_completion_tokens": 1600,
                "thinking": {"type": "disabled"},
            }
        
        
        def _strip_json_fence(text: str) -> str:
            value = text.strip()
            if value.startswith("```") and value.endswith("```"):
                lines = value.splitlines()
                if len(lines) >= 3:
                    value = "\n".join(lines[1:-1]).strip()
            return value
        
        
        def _validated_live_output(response: Any) -> Any:
            """The one shape check on the third-party model response: a list of observations, or an
            object carrying `observations`/`findings`; anything else is a failed (fail-open) request."""
            try:
                content = response["choices"][0]["message"]["content"]
            except (KeyError, IndexError, TypeError):
                raise ValueError("malformed_response") from None
            if isinstance(content, str):
                try:
                    content = json.loads(_strip_json_fence(content))
                except (TypeError, ValueError):
                    raise ValueError("malformed_json_content") from None
            if isinstance(content, list):
                return content
            if isinstance(content, Mapping) and isinstance(content.get("observations"), list):
                return content
            if isinstance(content, Mapping) and isinstance(content.get("findings"), list):
                return content
            raise ValueError("malformed_observations")
        
      • mimo_report.py 10.3 KB
        """Sample frames and build aggregate MiMo QC reports."""
        
        from __future__ import annotations
        
        import base64
        import json
        import math
        import os
        import subprocess
        import tempfile
        from pathlib import Path
        from typing import Any
        from collections.abc import Callable, Mapping, Sequence
        
        import qc_contract
        from qc.mimo_client import mimo_qc_api_call
        from qc.mimo_evidence import (
            _cache_evidence,
            _effective_config,
            _existing_final_output,
            collect_evidence,
            safe_mimo_config,
        )
        from qc.mimo_observations import normalize_observations
        from qc.mimo_payload import _request_payload, _validated_live_output, build_payload
        from qc.mimo_contract import (
            ARTIFACT_NAME,
            DEFAULT_STAGE,
            FRAME_SAMPLER_VERSION,
            MAX_FRAME_DIMENSION,
            MAX_FRAMES,
        )
        
        JudgeCallable = Callable[
            [Mapping[str, Any]], Mapping[str, Any] | Sequence[Mapping[str, Any]]
        ]
        
        FrameSampler = Callable[..., Sequence[Mapping[str, Any]]]
        
        
        def _probe_duration(path: Path) -> float:
            """Media duration via ffprobe; 0.0 when ffprobe cannot read it (frame sampling is advisory)."""
            result = subprocess.run(
                [
                    "ffprobe",
                    "-v",
                    "error",
                    "-show_entries",
                    "format=duration",
                    "-of",
                    "json",
                    str(path),
                ],
                capture_output=True,
                text=True,
                timeout=20,
                check=False,
            )
            if result.returncode != 0:
                return 0.0
            try:
                return max(0.0, float(json.loads(result.stdout)["format"]["duration"]))
            except (KeyError, ValueError):
                return 0.0
        
        
        def sample_video_frames(
            video_path: str | Path,
            *,
            max_frames: int = MAX_FRAMES,
            max_dimension: int = MAX_FRAME_DIMENSION,
        ) -> list[dict[str, Any]]:
            """Extract at most six bounded JPEGs; returned base64 is request-only."""
            path = Path(video_path)
            if not path.is_file() or max_frames <= 0:
                return []
            try:
                duration = _probe_duration(path)
            except (OSError, subprocess.SubprocessError):
                return []
            if duration <= 0:
                return []
            count = min(max_frames, max(1, math.ceil(duration / 15.0)))
            samples = []
            with tempfile.TemporaryDirectory(prefix="video-recap-mimo-qc-") as tmp:
                for index in range(count):
                    timestamp = duration * (index + 0.5) / count
                    destination = Path(tmp) / f"frame-{index:02d}.jpg"
                    try:
                        result = subprocess.run(
                            [
                                "ffmpeg",
                                "-hide_banner",
                                "-loglevel",
                                "error",
                                "-y",
                                "-ss",
                                f"{timestamp:.3f}",
                                "-i",
                                str(path),
                                "-frames:v",
                                "1",
                                "-vf",
                                f"scale={max_dimension}:{max_dimension}:force_original_aspect_ratio=decrease",
                                "-q:v",
                                "4",
                                str(destination),
                            ],
                            capture_output=True,
                            timeout=30,
                            check=False,
                        )
                    except (OSError, subprocess.SubprocessError):
                        continue
                    if result.returncode != 0 or not destination.is_file():
                        continue
                    raw = destination.read_bytes()
                    samples.append(
                        {
                            "data_url": "data:image/jpeg;base64,"
                            + base64.b64encode(raw).decode("ascii"),
                            "bytes": len(raw),
                            "timestamp": round(timestamp, 3),
                        }
                    )
            return samples
        
        
        def _frame_metadata(samples: Sequence[Mapping[str, Any]]) -> dict[str, Any]:
            return {
                "count": len(samples),
                "max_frames": MAX_FRAMES,
                "max_dimension": MAX_FRAME_DIMENSION,
                "sampler_version": FRAME_SAMPLER_VERSION,
                "samples": [
                    {"bytes": sample["bytes"], "timestamp": sample["timestamp"]}
                    for sample in samples
                ],
            }
        
        
        def _cache_input(
            stage: str,
            payload: Mapping[str, Any],
            evidence: Mapping[str, Any],
            frames: Mapping[str, Any],
        ) -> dict[str, Any]:
            """Everything a stage report depends on; an existing report whose stored cache_input
            equals this dict is reused instead of a new request."""
            return {
                "stage": stage,
                "model": payload["model"],
                "config": payload["config"],
                "instructions": payload["instructions"],
                "evidence": _cache_evidence(evidence),
                "frames": frames,
                "contract": qc_contract.SCHEMA_VERSION,
            }
        
        
        def _stage_reports(path: Path) -> dict[str, dict[str, Any]]:
            """Per-stage reports from an existing aggregate mimo_qc.json, or {} before the first stage.
        
            The aggregate is this module's own atomically written, contract-validated report, so a
            corrupt one raises (recap_stage_qc keeps the pipeline fail-open around the whole stage)."""
            if not path.is_file():
                return {}
            return dict(json.loads(path.read_text(encoding="utf-8"))["metadata"]["stages"])
        
        
        def _error_name(exc: Exception) -> str:
            return (str(exc).strip() or type(exc).__name__)[:120]
        
        
        def build_report(
            work_dir: str | Path,
            *,
            stage: str = DEFAULT_STAGE,
            fixture: Any | None = None,
            dry_run: bool = False,
            judge: JudgeCallable | None = None,
            config: Mapping[str, Any] | None = None,
            final_output: str | Path | None = None,
            live: bool = False,
            refresh: bool = False,
            frame_sampler: FrameSampler | None = None,
            existing: Mapping[str, Any] | None = None,
            api_call: Callable[..., Any] | None = None,
        ) -> dict[str, Any]:
            """Build one validated stage report; all live failures remain successful QC."""
            root = Path(work_dir)
            evidence = collect_evidence(root, final_output=final_output)
            payload = build_payload(evidence, stage=stage, config=config)
            cfg = safe_mimo_config(config)
            samples: Sequence[Mapping[str, Any]] = []
            if stage == "post_render" and (fixture is not None or judge is not None or live):
                output = _existing_final_output(root, final_output)
                if output is not None:
                    sampler = frame_sampler or sample_video_frames
                    samples = list(
                        sampler(output, max_frames=MAX_FRAMES, max_dimension=MAX_FRAME_DIMENSION)
                    )[:MAX_FRAMES]
            frame_meta = _frame_metadata(samples)
            cache_input = _cache_input(stage, payload, evidence, frame_meta)
        
            if (
                not refresh
                and existing is not None
                and existing["metadata"]["cache_input"] == cache_input
            ):
                cached = json.loads(json.dumps(existing))
                cached["metadata"].update(status="cached", mode="live_cache", request_count=0)
                return cached
        
            status = "completed"
            error = None
            if fixture is not None:
                model_output = fixture
                mode = "fixture"
            elif judge is not None and not dry_run:
                mode = "injected_judge"
                try:
                    model_output = judge(payload)
                except Exception as exc:
                    model_output = {"observations": []}
                    status, error = "failed", _error_name(exc)
            elif live and not dry_run:
                mode = "live"
                if not cfg["key_present"]:
                    model_output = {"observations": []}
                    status, error = "unavailable", "missing_key"
                else:
                    # Fail-open boundary for the third-party request: transport errors arrive as
                    # MiMoQCRequestError (a RuntimeError), malformed responses as ValueError.
                    try:
                        response = (api_call or mimo_qc_api_call)(
                            _request_payload(payload, samples),
                            config=_effective_config(config),
                            timeout=60,
                        )
                        model_output = _validated_live_output(response)
                    except (RuntimeError, ValueError) as exc:
                        model_output = {"observations": []}
                        status, error = "failed", _error_name(exc)
            else:
                model_output = {"observations": []}
                mode = "dry_run"
                status = "dry_run"
        
            findings = normalize_observations(
                model_output, stage=stage, payload=payload, model_config=config
            )
            metadata = {
                "mode": mode,
                "status": status,
                "error": error,
                "report_only": True,
                "auto_repair": False,
                "pipeline_blocking": False,
                "deterministic": False,
                "request_count": 1
                if mode == "live" and status in {"completed", "failed"}
                else 0,
                "cache_input": cache_input,
                "frame_samples": frame_meta,
                "evidence": evidence,
                "payload": payload,
                "config": cfg,
            }
            return qc_contract.build_report(
                artifact=ARTIFACT_NAME,
                stage=stage,
                findings=findings,
                metadata=metadata,
            )
        
        
        def _aggregate_reports(
            stage_reports: Mapping[str, Mapping[str, Any]], current_stage: str
        ) -> dict[str, Any]:
            findings = [
                finding
                for stage_report in stage_reports.values()
                for finding in stage_report["findings"]
            ]
            metadata = dict(stage_reports[current_stage]["metadata"])
            metadata["stages"] = {
                stage: dict(report) for stage, report in stage_reports.items()
            }
            return qc_contract.build_report(
                artifact=ARTIFACT_NAME,
                stage=current_stage,
                findings=findings,
                metadata=metadata,
            )
        
        
        def write_report(
            work_dir: str | Path, report: Mapping[str, Any], *, output: str | Path | None = None
        ) -> tuple[Path, dict[str, Any]]:
            """Atomically merge one stage into mimo_qc.json and return the aggregate."""
            path = Path(output) if output else Path(work_dir) / ARTIFACT_NAME
            path.parent.mkdir(parents=True, exist_ok=True)
            stages = _stage_reports(path)
            stages[report["stage"]] = dict(report)
            aggregate = _aggregate_reports(stages, report["stage"])
            serialized = json.dumps(aggregate, ensure_ascii=False, indent=2) + "\n"
            descriptor, temp_name = tempfile.mkstemp(prefix=f".{path.name}.", dir=path.parent)
            try:
                with os.fdopen(descriptor, "w", encoding="utf-8") as handle:
                    handle.write(serialized)
                    handle.flush()
                    os.fsync(handle.fileno())
                os.replace(temp_name, path)
            except BaseException:
                Path(temp_name).unlink(missing_ok=True)
                raise
            return path, aggregate
        
      • mimo_runner.py 3.7 KB
        """Run and persist advisory MiMo QC stages."""
        
        from __future__ import annotations
        
        import argparse
        
        
        import json
        
        
        from pathlib import Path
        
        from typing import Any
        from collections.abc import Callable, Mapping, Sequence
        
        import qc_contract
        
        from qc.mimo_report import _stage_reports, build_report, write_report
        from qc.mimo_contract import ARTIFACT_NAME, DEFAULT_STAGE
        
        JudgeCallable = Callable[
            [Mapping[str, Any]], Mapping[str, Any] | Sequence[Mapping[str, Any]]
        ]
        
        FrameSampler = Callable[..., Sequence[Mapping[str, Any]]]
        
        
        def run(
            work_dir: str | Path,
            *,
            stage: str = DEFAULT_STAGE,
            fixture: Any | None = None,
            dry_run: bool = False,
            judge: JudgeCallable | None = None,
            config: Mapping[str, Any] | None = None,
            final_output: str | Path | None = None,
            output: str | Path | None = None,
            live: bool = False,
            refresh: bool = False,
            frame_sampler: FrameSampler | None = None,
            api_call: Callable[..., Any] | None = None,
        ) -> dict[str, Any]:
            path = Path(output) if output else Path(work_dir) / ARTIFACT_NAME
            existing = _stage_reports(path).get(stage)
            report = build_report(
                work_dir,
                stage=stage,
                fixture=fixture,
                dry_run=dry_run,
                judge=judge,
                config=config,
                final_output=final_output,
                live=live,
                refresh=refresh,
                frame_sampler=frame_sampler,
                existing=existing,
                api_call=api_call,
            )
            written_path, aggregate = write_report(work_dir, report, output=output)
            return {"path": str(written_path), "report": aggregate}
        
        
        def clear_report(work_dir: str | Path, *, output: str | Path | None = None) -> bool:
            path = Path(output) if output else Path(work_dir) / ARTIFACT_NAME
            try:
                path.unlink()
                return True
            except FileNotFoundError:
                return False
        
        
        def _load_fixture(path: str | Path | None) -> Any | None:
            if not path:
                return None
            with Path(path).open("r", encoding="utf-8") as handle:
                return json.load(handle)
        
        
        def main(
            argv: Sequence[str] | None = None,
            *,
            run_callable: Callable[..., dict[str, Any]] | None = None,
        ) -> int:
            parser = argparse.ArgumentParser(
                description="Write advisory MiMo QC from local recap evidence."
            )
            parser.add_argument("--work-dir", required=True)
            parser.add_argument(
                "--stage", default=DEFAULT_STAGE, choices=sorted(qc_contract.STAGES)
            )
            parser.add_argument("--fixture", help="offline model-response fixture JSON")
            parser.add_argument(
                "--live", action="store_true", help="make one live MiMo request for this stage"
            )
            parser.add_argument(
                "--refresh", action="store_true", help="ignore a matching stage cache"
            )
            parser.add_argument(
                "--dry-run",
                action="store_true",
                help="write evidence only; never access the network",
            )
            parser.add_argument("--model", help="override MIMO_QC_MODEL for this call")
            parser.add_argument("--final-output")
            parser.add_argument("--output")
            args = parser.parse_args(argv)
            fixture = _load_fixture(args.fixture)
            config = {"mimo_qc_model": args.model} if args.model else None
            result = (run_callable or run)(
                args.work_dir,
                stage=args.stage,
                fixture=fixture,
                live=args.live and fixture is None,
                refresh=args.refresh,
                dry_run=args.dry_run or (not args.live and fixture is None),
                config=config,
                final_output=args.final_output,
                output=args.output,
            )
            print(
                json.dumps(
                    {
                        "ok": result["report"]["ok"],
                        "status": result["report"]["metadata"]["status"],
                        "path": result["path"],
                    },
                    ensure_ascii=False,
                )
            )
            return 0
        
      • __init__.py 26 B
        """MiMo QC subpackage."""
        
    • doctor.py 20.4 KB
      #!/usr/bin/env python3
      """Environment doctor for the video-recap skill bundle.
      
      The pipeline runs on ffmpeg + MiMo for understanding; voiceover may explicitly use
      MiMo, Fish Audio, or a privately configured Index TTS endpoint.
      """
      from __future__ import annotations
      
      import argparse
      import json
      import os
      import shutil
      import subprocess
      import sys
      import urllib.parse
      from pathlib import Path
      
      from lib import CONFIG
      
      
      SCRIPT_DIR = Path(__file__).resolve().parent
      DEGRADED_GROUP = "warnings/degraded"
      TTS_PROVIDERS = ("auto", "mimo-tts", "fish-audio", "index-tts")
      
      
      def _command_path(name: str) -> str | None:
          return shutil.which(name)
      
      
      def _ffmpeg_filters() -> set[str]:
          """Filters the installed ffmpeg lists; empty when ffmpeg is absent.
      
          A present ffmpeg whose `-filters` fails or hangs is an environment fault and raises,
          so it is never misreported downstream as "filter absent"."""
          ffmpeg = _command_path("ffmpeg")
          if not ffmpeg:
              return set()
          try:
              result = subprocess.run(
                  [ffmpeg, "-hide_banner", "-filters"], text=True, capture_output=True, timeout=20
              )
          except (OSError, subprocess.SubprocessError) as exc:
              raise RuntimeError(f"`ffmpeg -filters` failed or hung: {exc}") from exc
          if result.returncode != 0:
              detail = (result.stderr or result.stdout or "").strip()[:300]
              raise RuntimeError(f"`ffmpeg -filters` failed (exit {result.returncode}): {detail}")
          filters = set()
          for line in result.stdout.splitlines():
              parts = line.split()
              if len(parts) >= 2 and parts[0] and parts[0][0] in ".TSCAPN|":
                  filters.add(parts[1])
          return filters
      
      
      def ffmpeg_has_subtitles_filter() -> bool:
          """True when this ffmpeg can burn subtitles — its filter list includes the libass
          `subtitles` filter. The render burns even the .ass file through `subtitles=` (see
          video-assemble assemble.py:_subtitle_burn_filter), so this — not the `ass` filter — is
          the exact capability `--burn-subtitles` needs. Reused by the orchestrator preflight
          (recap.py) to fail fast before any API spend."""
          return "subtitles" in _ffmpeg_filters()
      
      
      def _asr_status() -> dict[str, object]:
          configured = bool(CONFIG["mimo_asr_api_key"])
          return {
              "configured": configured,
              "available": configured,
              "mimo_asr_model": CONFIG["mimo_asr_model"],
              "mimo_asr_api_url": CONFIG["mimo_asr_api_url"],
              "mimo_asr_api_url_source": CONFIG["mimo_asr_api_url_source"],
              "mimo_asr_language": CONFIG["mimo_asr_language"],
              "mimo_asr_env_var": CONFIG["mimo_asr_env_var"],
              "note": "ASR uses MiMo (mimo-v2.5-asr); set MIMO_API_KEY, or run with --skip-asr.",
          }
      
      
      def _index_tts_status() -> dict[str, object]:
          """Validate private Index settings locally without exposing or contacting them."""
          endpoint = os.environ.get("INDEX_TTS_ENDPOINT", "")
          voice = os.environ.get("INDEX_TTS_VOICE", "").strip()
          has_forbidden_control = any(
              ord(char) < 32 or ord(char) == 127 for char in endpoint
          )
          try:
              parsed = urllib.parse.urlsplit(endpoint) if endpoint else None
              port = parsed.port if parsed else None
          except ValueError:
              parsed = None
              port = None
          endpoint_format_valid = bool(
              parsed
              and not has_forbidden_control
              and parsed.scheme in {"http", "https"}
              and parsed.hostname
              and "@" not in parsed.netloc
              and not parsed.query
              and not parsed.fragment
              and (port is None or 1 <= port <= 65535)
          )
          return {
              "index_tts_endpoint_set": bool(endpoint),
              "index_tts_endpoint_format_valid": endpoint_format_valid,
              "index_tts_voice_set": bool(voice),
              "index_tts_configured": endpoint_format_valid and bool(voice),
              "connectivity_checked": False,
              "validation_scope": "offline configuration only; connectivity and acoustic quality not checked",
          }
      
      
      def _capability(name: str, summary: str, *, detail: str = "", action: str = "") -> dict[str, str]:
          item = {"name": name, "summary": summary}
          if detail:
              item["detail"] = detail
          if action:
              item["action"] = action
          return item
      
      
      def _build_capability_menu(checks: dict) -> dict[str, list[dict[str, str]]]:
          """Human-ready preflight summary grouped by what can run, what blocks, and what degrades.
      
          This is intentionally a small rollup over the existing `checks` tree. It does not replace
          the raw machine checks, install anything, or introduce provider ranking.
          """
          system = checks["system_tools"]
          api = checks["api_config"]
          asr = checks["asr"]
          tts = checks["tts"]
      
          menu: dict[str, list[dict[str, str]]] = {
              "ready": [],
              "blocked": [],
              DEGRADED_GROUP: [],
              "optional_upgrades": [],
          }
      
          ffmpeg_ready = system["ffmpeg"]
          ffprobe_ready = system["ffprobe"]
          subtitles_ready = system["burn_subtitles_ready"]
          api_key_set = api["api_key_set"]
          asr_ready = asr["available"]
          tts_ready = tts["available"]
          vlm_ready = api["mimo_video_configured"]
          normal_core_ready = ffmpeg_ready and ffprobe_ready and api_key_set and vlm_ready and tts_ready
      
          if ffmpeg_ready and ffprobe_ready:
              menu["ready"].append(
                  _capability(
                      "core_media_tools",
                      "ffmpeg and ffprobe are available",
                      detail="Local probing, cutting, rendering, and duration checks can run.",
                  )
              )
          else:
              if not ffmpeg_ready:
                  menu["blocked"].append(
                      _capability("ffmpeg", "Missing ffmpeg", action="Install ffmpeg before running the recap pipeline.")
                  )
              if not ffprobe_ready:
                  menu["blocked"].append(
                      _capability("ffprobe", "Missing ffprobe", action="Install ffprobe before running media probing/export.")
                  )
      
          if api_key_set:
              menu["ready"].append(
                  _capability(
                      "mimo_credentials",
                      "MiMo API key is configured",
                      detail=f"Source: {api['api_env_var']}",
                  )
              )
          else:
              menu["blocked"].append(
                  _capability(
                      "mimo_credentials",
                      "Missing MIMO_API_KEY",
                      action="Set MIMO_API_KEY; the default ASR / VLM / TTS path depends on it.",
                  )
              )
      
          if vlm_ready:
              menu["ready"].append(
                  _capability(
                      "mimo_vlm",
                      "MiMo VLM/video understanding is configured",
                      detail=f"Model: {api['vlm_model']}",
                  )
              )
          elif api_key_set:
              menu["blocked"].append(
                  _capability(
                      "mimo_vlm",
                      "MiMo VLM/video understanding is not configured",
                      action="Set MIMO_VIDEO_API_KEY or the shared MIMO_API_KEY before video understanding.",
                  )
              )
      
          tts_provider = tts["provider"]
          if tts_provider == "index-tts":
              capability_name = "index_tts_configuration"
              action = "Set valid INDEX_TTS_ENDPOINT and INDEX_TTS_VOICE values before voiceover."
          elif tts_provider == "fish-audio":
              capability_name = "fish_audio_tts"
              action = "Set FISH_API_KEY before voiceover."
          else:
              capability_name = "mimo_tts"
              action = "Set MIMO_TTS_API_KEY or the shared MIMO_API_KEY before voiceover."
          if tts_ready:
              menu["ready"].append(
                  _capability(
                      capability_name,
                      (
                          "index-tts configuration is present"
                          if tts_provider == "index-tts"
                          else f"{tts_provider} is configured"
                      ),
                      detail=(
                          tts["validation_scope"]
                          if tts_provider == "index-tts"
                          else f"Model: {tts['model']}"
                      ),
                  )
              )
          elif api_key_set:
              menu["blocked"].append(
                  _capability(
                      capability_name,
                      f"{tts_provider} is not configured",
                      action=action,
                  )
              )
      
          if asr_ready:
              menu["ready"].append(
                  _capability(
                      "mimo_asr",
                      "MiMo ASR is configured",
                      detail=f"Language: {asr['mimo_asr_language']}; model: {asr['mimo_asr_model']}",
                  )
              )
          else:
              menu[DEGRADED_GROUP].append(
                  _capability(
                      "mimo_asr",
                      "ASR is unavailable; run only with --skip-asr",
                      action=asr["note"],
                  )
              )
      
          if subtitles_ready:
              menu["ready"].append(
                  _capability("subtitle_burn", "Subtitle burn-in is available", detail="ffmpeg has the subtitles/libass filter.")
              )
          elif ffmpeg_ready:
              menu[DEGRADED_GROUP].append(
                  _capability(
                      "subtitle_burn",
                      "Subtitle burn-in is unavailable",
                      action="Use --no-burn-subtitles or install an ffmpeg build with the subtitles/libass filter.",
                  )
              )
      
          if not normal_core_ready:
              menu["blocked"].append(
                  _capability(
                      "default_recap_pipeline",
                      "Default recap run is blocked",
                      detail="Resolve the blocking items above before a normal run.",
                  )
              )
          elif asr_ready and subtitles_ready:
              menu["ready"].append(
                  _capability(
                      "default_recap_pipeline",
                      (
                          "Recap prerequisites are configured"
                          if tts_provider == "index-tts"
                          else "Default recap run is ready"
                      ),
                      detail=(
                          "ASR, VLM, and media tools are configured; Index TTS passed offline configuration checks only."
                          if tts_provider == "index-tts"
                          else "ASR, VLM, TTS, and media tools are configured."
                      ),
                  )
              )
          else:
              actions = []
              if not asr_ready:
                  actions.append("run with --skip-asr")
              if not subtitles_ready:
                  actions.append("run with --no-burn-subtitles")
              menu[DEGRADED_GROUP].append(
                  _capability(
                      "recap_degraded_mode",
                      "Recap can run only in an explicit degraded mode",
                      detail="; ".join(actions),
                  )
              )
      
          menu["optional_upgrades"].append(
              _capability(
                  "jianying_export",
                  "Editable JianYing draft export can be requested with --export-jianying",
                  detail="No JianYing install is required to write the draft; ffprobe improves media metadata.",
              )
          )
          if subtitles_ready:
              menu["optional_upgrades"].append(
                  _capability(
                      "burned_subtitles",
                      "Burned subtitles are available and enabled by default",
                      action="Use --no-burn-subtitles if you prefer external subtitle files.",
                  )
              )
      
          return menu
      
      
      def build_report(*, tts_provider: str | None = None) -> dict[str, object]:
          filters = _ffmpeg_filters()
          ffmpeg_path = _command_path("ffmpeg") or ""
          ffprobe_path = _command_path("ffprobe") or ""
          mimo_video_configured = bool(CONFIG["mimo_video_api_key"])
          mimo_tts_configured = bool(CONFIG["mimo_tts_api_key"])
          fish_tts_configured = bool(CONFIG["fish_api_key"])
          index_tts = _index_tts_status()
          requested_tts_provider = tts_provider or CONFIG["tts_provider"]
          effective_tts_provider = requested_tts_provider
          if requested_tts_provider == "auto":
              effective_tts_provider = (
                  "mimo-tts" if mimo_tts_configured or not fish_tts_configured else "fish-audio"
              )
          if effective_tts_provider == "fish-audio":
              tts_configured = fish_tts_configured
          elif effective_tts_provider == "index-tts":
              tts_configured = index_tts["index_tts_configured"]
          else:
              tts_configured = mimo_tts_configured
          if effective_tts_provider == "fish-audio":
              tts_model = CONFIG["fish_tts_model"]
          elif effective_tts_provider == "index-tts":
              tts_model = "provider-managed"
          else:
              tts_model = CONFIG["mimo_tts_model"]
          subtitle_filter = "subtitles" in filters
          checks = {
              "system_tools": {
                  "ffmpeg": bool(ffmpeg_path),
                  "ffmpeg_path": ffmpeg_path,
                  "ffprobe": bool(ffprobe_path),
                  "ffprobe_path": ffprobe_path,
                  "ffmpeg_subtitles_filter": subtitle_filter,
                  "ffmpeg_ass_filter": "ass" in filters,
                  "burn_subtitles_ready": bool(ffmpeg_path and subtitle_filter),
              },
              "tts": {
                  "provider": effective_tts_provider,
                  "requested_provider": requested_tts_provider,
                  "mimo_tts_configured": mimo_tts_configured,
                  "mimo_tts_api_url": CONFIG["mimo_tts_api_url"],
                  "mimo_tts_api_url_source": CONFIG["mimo_tts_api_url_source"],
                  "mimo_tts_model": CONFIG["mimo_tts_model"],
                  "mimo_tts_model_source": CONFIG["mimo_tts_model_source"],
                  "mimo_tts_voice": CONFIG["mimo_tts_voice"],
                  "mimo_tts_voice_source": CONFIG["mimo_tts_voice_source"],
                  "fish_tts_configured": fish_tts_configured,
                  "fish_tts_api_url": CONFIG["fish_tts_api_url"],
                  "fish_tts_model": CONFIG["fish_tts_model"],
                  "fish_tts_reference_id_set": bool(CONFIG["fish_tts_reference_id"]),
                  "fish_tts_reference_id_source": CONFIG["fish_tts_reference_id_source"],
                  **index_tts,
                  "model": tts_model,
                  "available": tts_configured,
              },
              "asr": _asr_status(),
              "api_config": {
                  "api_provider": CONFIG["api_provider"],
                  "api_url": CONFIG["api_url"],
                  "api_url_source": CONFIG["api_url_source"],
                  "api_env_var": CONFIG["api_env_var"],
                  "api_key_set": bool(CONFIG["api_key"]),
                  "vlm_model": CONFIG["vlm_model"],
                  "vlm_model_source": CONFIG["vlm_model_source"],
                  "vlm_workers": CONFIG["vlm_workers"],
                  "mimo_video_configured": mimo_video_configured,
                  "mimo_video_api_url": CONFIG["mimo_video_api_url"],
                  "mimo_video_model": CONFIG["mimo_video_model"],
                  "mimo_video_model_source": CONFIG["mimo_video_model_source"],
              },
              "python": {
                  "executable": sys.executable,
                  "version": sys.version.split()[0],
              },
          }
      
          failures: list[str] = []
          warnings: list[str] = []
          tools = checks["system_tools"]
          for name in ("ffmpeg", "ffprobe"):
              if not tools[name]:
                  failures.append(f"Missing system tool: {name}")
          if requested_tts_provider not in TTS_PROVIDERS:
              failures.append(
                  "TTS_PROVIDER must be one of: " + ", ".join(TTS_PROVIDERS)
              )
          elif requested_tts_provider == "index-tts":
              if not index_tts["index_tts_endpoint_set"]:
                  failures.append("INDEX_TTS_ENDPOINT is not set")
              elif not index_tts["index_tts_endpoint_format_valid"]:
                  failures.append(
                      "INDEX_TTS_ENDPOINT must be an http/https URL with a hostname and no credentials, query, or fragment"
                  )
              if not index_tts["index_tts_voice_set"]:
                  failures.append("INDEX_TTS_VOICE is not set")
          if tools["ffmpeg"] and not tools["ffmpeg_subtitles_filter"]:
              warnings.append("ffmpeg lacks subtitles/libass filter; --burn-subtitles will fail")
          if not checks["api_config"]["api_key_set"]:
              failures.append("MIMO_API_KEY is not set; the default ASR / VLM path requires MiMo")
          if not checks["asr"]["available"]:
              warnings.append("ASR not configured (MIMO_API_KEY); pipeline can run with --skip-asr")
          return {
              "ok": not failures,
              "repo_root": str(SCRIPT_DIR.parents[2]),
              "checks": checks,
              "capability_menu": _build_capability_menu(checks),
              "failures": failures,
              "warnings": warnings,
          }
      
      
      def _status_icon(ok: bool, *, warning: bool = False) -> str:
          if ok:
              return "✓"
          return "!" if warning else "✗"
      
      
      def _print_human(report: dict) -> None:
          checks = report["checks"]
          print("video-recap doctor")
          print(f"Repo root: {report['repo_root']}")
      
          system = checks["system_tools"]
          print("\n[system]")
          print(f"{_status_icon(system['ffmpeg'])} ffmpeg: {system['ffmpeg_path'] or 'not found'}")
          print(f"{_status_icon(system['ffprobe'])} ffprobe: {system['ffprobe_path'] or 'not found'}")
          print(
              f"{_status_icon(system['ffmpeg_subtitles_filter'], warning=True)} "
              f"ffmpeg subtitles/libass filter: "
              f"{'available' if system['ffmpeg_subtitles_filter'] else 'missing'}"
          )
      
          api = checks["api_config"]
          print("\n[api]")
          print(f"✓ API provider: {api['api_provider']}")
          print(f"✓ API URL: {api['api_url']} (source: {api['api_url_source']})")
          print(
              f"{_status_icon(api['api_key_set'])} "
              f"{api['api_env_var']}: {'set' if api['api_key_set'] else 'not set'}"
          )
          print(f"✓ VLM model: {api['vlm_model']} (source: {api['vlm_model_source']})")
          print(f"✓ VLM_WORKERS: {api['vlm_workers']}")
      
          asr = checks["asr"]
          print("\n[asr]")
          print(
              f"{_status_icon(asr['available'], warning=True)} "
              f"MiMo ASR: {'configured' if asr['available'] else 'not configured'} "
              f"(key: {asr['mimo_asr_env_var']})"
          )
          print(f"✓ ASR model: {asr['mimo_asr_model']}")
          print(f"✓ ASR API URL: {asr['mimo_asr_api_url']} (source: {asr['mimo_asr_api_url_source']})")
          print(f"✓ ASR language: {asr['mimo_asr_language']}")
          if not asr["available"]:
              print(f"  note: {asr['note']}")
      
          tts = checks["tts"]
          print("\n[tts]")
          print(
              f"{_status_icon(tts['available'])} {tts['provider']}: "
              f"{'configured' if tts['available'] else 'not configured'}"
          )
          print(f"✓ TTS model: {tts['model']}")
          if tts["provider"] == "fish-audio":
              print(
                  "✓ TTS voice reference ID: "
                  f"{'set' if tts['fish_tts_reference_id_set'] else 'not set'} "
                  f"(source: {tts['fish_tts_reference_id_source']})"
              )
              print(f"✓ TTS API URL: {tts['fish_tts_api_url']}")
          elif tts["provider"] == "index-tts":
              print(
                  f"{_status_icon(tts['index_tts_endpoint_format_valid'])} "
                  "Index TTS endpoint: "
                  f"{'configured with valid format' if tts['index_tts_endpoint_format_valid'] else 'missing or invalid'}"
              )
              print(
                  f"{_status_icon(tts['index_tts_voice_set'])} Index TTS voice: "
                  f"{'configured' if tts['index_tts_voice_set'] else 'not set'}"
              )
              print(f"! Validation scope: {tts['validation_scope']}")
          else:
              print(f"✓ TTS voice: {tts['mimo_tts_voice']} (source: {tts['mimo_tts_voice_source']})")
              print(f"✓ TTS API URL: {tts['mimo_tts_api_url']} (source: {tts['mimo_tts_api_url_source']})")
      
          menu = report["capability_menu"]
          print("\n[capability menu]")
          for group in ("ready", "blocked", DEGRADED_GROUP, "optional_upgrades"):
              print(f"{group}:")
              items = menu[group]
              if not items:
                  print("  - none")
                  continue
              for item in items:
                  line = f"  - {item['name']}: {item['summary']}"
                  if "detail" in item:
                      line += f" ({item['detail']})"
                  print(line)
                  if "action" in item:
                      print(f"    action: {item['action']}")
      
          if report["warnings"]:
              print("\nWarnings:")
              for warning in report["warnings"]:
                  print(f"- {warning}")
          if report["failures"]:
              print("\nStatus: FAILED")
              for failure in report["failures"]:
                  print(f"- {failure}")
          else:
              print("\nStatus: OK")
      
      
      def main() -> int:
          parser = argparse.ArgumentParser(description="Check video-recap runtime prerequisites.")
          parser.add_argument("--json", action="store_true", help="Print machine-readable JSON")
          parser.add_argument(
              "--tts-provider",
              choices=TTS_PROVIDERS,
              default=None,
              help="override the TTS provider for this preflight report",
          )
          args = parser.parse_args()
      
          report = build_report(tts_provider=args.tts_provider)
          if args.json:
              print(json.dumps(report, ensure_ascii=False, indent=2))
              return 0 if report["ok"] else 1
      
          _print_human(report)
          return 0 if report["ok"] else 1
      
      
      if __name__ == "__main__":
          raise SystemExit(main())
      
    • final_qc.py 22.5 KB
      #!/usr/bin/env python3
      """Final post-render QC and golden-eval reports for video-recap.
      
      Local deterministic/report-only checks only: no network, no repair, no secrets.
      """
      from __future__ import annotations
      
      import argparse
      import json
      import math
      import shutil
      import subprocess
      from pathlib import Path
      from typing import Any
      from collections.abc import Callable, Mapping, Sequence
      
      import qc_contract
      from lib import load_json
      
      FINAL_QC_ARTIFACT = "final_qc.json"
      GOLDEN_EVAL_ARTIFACT = "golden_eval.json"
      POST_RENDER_STAGE = "post_render"
      GOLDEN_STAGE = "golden"
      _COLLECT_ARTIFACTS = (
          "assembly_manifest.json",
          "assembly_qc.json",
          "visual_qc.json",
          "preflight_qc.json",
          "mimo_qc.json",
      )
      # video-assemble QC artifacts: {"verdict", "blocking", "blocking_codes": [...]}.
      _UPSTREAM_QC_ARTIFACTS = ("assembly_qc.json", "visual_qc.json")
      ProbeRunner = Callable[[Path], Mapping[str, Any]]
      
      
      def _load_fixture(value: Any) -> Any:
          """A fixture is an in-memory JSON value or the path of a JSON file."""
          if isinstance(value, (Mapping, list)):
              return value
          return load_json(value)
      
      
      def _resolve_in_work_dir(work_dir: Path, path: str | Path) -> Path:
          p = Path(path)
          return p if p.is_absolute() else work_dir / p
      
      
      def _read_json_mapping(path: Path) -> Mapping[str, Any] | None:
          try:
              data = load_json(path)
          except (OSError, ValueError):
              return None
          return data if isinstance(data, Mapping) else None
      
      
      def _final_output_path(work_dir: Path, final_output: str | Path | None) -> Path | None:
          """The rendered final output: the caller's path, else assembly_manifest.final_output
          (video-assemble always writes it; output.mp4 is only the assembler's intermediate).
          None when neither exists yet, which the report surfaces as a missing final output."""
          if final_output is None:
              manifest = work_dir / "assembly_manifest.json"
              if not manifest.is_file():
                  return None
              final_output = load_json(manifest)["final_output"]
          return _resolve_in_work_dir(work_dir, final_output)
      
      
      def _file_metadata(path: Path | None, work_dir: Path) -> dict[str, Any]:
          if path is None:
              return {"path": "(assembly_manifest.json missing)", "exists": False, "bytes": 0}
          try:
              display = path.relative_to(work_dir).as_posix()
          except ValueError:  # final outputs normally live next to work_dir, not inside it
              display = str(path)
          exists = path.is_file()
          return {
              "path": display,
              "exists": exists,
              "bytes": path.stat().st_size if exists else 0,
          }
      
      
      def _artifact_summary(work_dir: Path, name: str) -> dict[str, Any]:
          path = work_dir / name
          meta = _file_metadata(path, work_dir)
          if meta["exists"]:
              data = _read_json_mapping(path)
              if data is None:
                  meta["summary"] = {"invalid": True}
              else:
                  meta["summary"] = {
                      key: data.get(key)
                      for key in ("schema_version", "artifact", "stage", "ok", "blocker_count", "finding_count")
                  }
          return meta
      
      
      def _finding(*, finding_id: str, code: str, message: str, category: str = "schema_invalid",
                   stage: str = POST_RENDER_STAGE, source: Mapping[str, Any] | None = None,
                   evidence: Mapping[str, Any] | None = None,
                   next_action: str = "manual_review") -> dict[str, Any]:
          return qc_contract.build_finding(
              finding_id=finding_id,
              stage=stage,
              severity="blocker",
              confidence="objective",
              sample_policy={"type": "deterministic"},
              category=category,
              code=code,
              message=message,
              deterministic=True,
              blocking=True,
              source=source,
              evidence=evidence,
              next_action=next_action,
              model_used="local_deterministic_final_qc_v1",
          )
      
      
      def _run_ffprobe(path: Path) -> Mapping[str, Any]:
          if shutil.which("ffprobe") is None:
              raise RuntimeError("ffprobe unavailable")
          cmd = [
              "ffprobe", "-v", "error", "-print_format", "json",
              "-show_format", "-show_streams", str(path),
          ]
          res = subprocess.run(cmd, capture_output=True, text=True)
          if res.returncode != 0:
              raise RuntimeError((res.stderr or res.stdout or "ffprobe failed").strip())
          try:
              return json.loads(res.stdout)
          except ValueError as exc:
              raise RuntimeError(f"ffprobe returned invalid JSON: {exc}") from exc
      
      
      def _tail_decode_check(path: Path) -> tuple[bool | None, str | None]:
          """Cheaply verify the final ~2s actually decodes. Returns (ok, detail).
      
          Header probing (ffprobe -show_format/-show_streams) passes a container-valid but
          media-truncated/corrupt payload (moov intact + mdat cut — realistic with +faststart on
          disk-full or partial upload). Decoding the tail catches it. Returns (None, ...) when ffmpeg
          is unavailable or the probe cannot run, so the caller skips rather than false-blocks.
          """
          if shutil.which("ffmpeg") is None:
              return None, "ffmpeg unavailable"
          try:
              res = subprocess.run(
                  ["ffmpeg", "-v", "error", "-xerror", "-sseof", "-2", "-i", str(path), "-f", "null", "-"],
                  capture_output=True, text=True, timeout=60,
              )
          except (OSError, subprocess.SubprocessError) as exc:
              return None, f"decode probe could not run: {exc}"
          if res.returncode != 0:
              return False, (res.stderr or "tail decode failed").strip()[:500]
          return True, None
      
      
      def _probe_metadata(path: Path, *, probe_fixture: Any = None, probe_runner: ProbeRunner | None = None) -> tuple[dict[str, Any] | None, dict[str, Any] | None]:
          """Return (probe, error); an ffprobe failure on existing media becomes a deterministic blocker."""
          try:
              raw = _load_fixture(probe_fixture) if probe_fixture is not None else (probe_runner or _run_ffprobe)(path)
              if not isinstance(raw, Mapping):
                  raise TypeError("probe metadata must be a JSON object")
              streams = raw.get("streams", [])
              format_info = raw.get("format", {})
              if not isinstance(streams, list) or any(not isinstance(stream, Mapping) for stream in streams):
                  raise TypeError("probe metadata streams must be an array of objects")
              if not isinstance(format_info, Mapping):
                  raise TypeError("probe metadata format must be an object")
              normalized = dict(raw)
              normalized["streams"] = [dict(stream) for stream in streams]
              normalized["format"] = dict(format_info)
              return normalized, None
          except (OSError, ValueError, RuntimeError, subprocess.SubprocessError, TypeError) as exc:
              return None, {"code": "probe_failed", "message": str(exc) or "ffprobe failed"}
      
      
      def _first_video_stream(probe: Mapping[str, Any]) -> Mapping[str, Any] | None:
          return next((s for s in probe.get("streams", []) if s.get("codec_type") == "video"), None)
      
      
      def _positive_finite(value: Any) -> float | None:
          """Parse an ffprobe number or rational ("30000/1001"); None unless positive and finite."""
          try:
              if isinstance(value, str) and "/" in value:
                  numerator, denominator = value.split("/", 1)
                  number = float(numerator) / float(denominator)
              else:
                  number = float(value)
          except (TypeError, ValueError, ZeroDivisionError):
              return None
          return number if math.isfinite(number) and number > 0 else None
      
      
      def _first_positive(candidates: Sequence[Any]) -> tuple[float | None, str | None, Any]:
          """(parsed, problem, raw): the first positive finite candidate wins; otherwise problem is
          'missing' (no candidate at all) or 'invalid' (with the first raw candidate)."""
          seen = [value for value in candidates if value not in (None, "")]
          for raw in seen:
              parsed = _positive_finite(raw)
              if parsed is not None:
                  return parsed, None, raw
          if not seen:
              return None, "missing", None
          return None, "invalid", seen[0]
      
      
      def _probe_duration(probe: Mapping[str, Any]) -> tuple[float | None, str | None, Any]:
          return _first_positive([probe.get("format", {}).get("duration")])
      
      
      def _probe_fps(video_stream: Mapping[str, Any] | None) -> tuple[float | None, str | None, Any]:
          if video_stream is None:
              return None, "missing", None
          return _first_positive([
              video_stream.get(key)
              for key in ("avg_frame_rate", "r_frame_rate", "fps", "frame_rate")
          ])
      
      
      def _probe_contract_findings(probe: Mapping[str, Any], *, final_meta: Mapping[str, Any]) -> list[dict[str, Any]]:
          findings: list[dict[str, Any]] = []
          source = {"artifact": final_meta["path"]}
          video_stream = _first_video_stream(probe)
          if video_stream is None:
              findings.append(_finding(
                  finding_id="final-qc-missing-video-stream",
                  code="missing_video_stream",
                  message="final output probe metadata has no video stream",
                  category="stream",
                  source=source,
                  evidence={"streams": probe.get("streams")},
                  next_action="rerender_final_output_with_video_stream",
              ))
      
          _duration, problem, raw = _probe_duration(probe)
          if problem is not None:
              findings.append(_finding(
                  finding_id=f"final-qc-{problem}-duration",
                  code=f"{problem}_duration",
                  message="final output probe metadata is missing a positive finite duration" if problem == "missing" else "final output probe metadata duration is not positive and finite",
                  category="duration",
                  source=source,
                  evidence={"duration": raw},
                  next_action="rerender_final_output_with_valid_duration",
              ))
      
          if not (video_stream or {}).get("codec_name"):
              findings.append(_finding(
                  finding_id="final-qc-missing-codec",
                  code="missing_codec",
                  message="final output probe metadata is missing a video codec",
                  category="stream",
                  source=source,
                  evidence={"video_stream": video_stream},
                  next_action="rerender_final_output_with_video_codec",
              ))
      
          _fps, problem, raw = _probe_fps(video_stream)
          if problem is not None:
              findings.append(_finding(
                  finding_id=f"final-qc-{problem}-fps",
                  code=f"{problem}_fps",
                  message="final output probe metadata is missing a positive finite video fps" if problem == "missing" else "final output probe metadata video fps is not positive and finite",
                  category="stream",
                  source=source,
                  evidence={"fps": raw, "video_stream": video_stream},
                  next_action="rerender_final_output_with_valid_fps",
              ))
          return findings
      
      
      def _upstream_blockers(work_dir: Path, artifact_name: str) -> list[dict[str, Any]]:
          """Roll a blocking video-assemble QC artifact into one final blocker per blocking code."""
          path = work_dir / artifact_name
          if not path.exists():
              return []
          data = _read_json_mapping(path)
          if data is None:  # unreadable upstream QC is itself a deterministic blocker
              return [_finding(
                  finding_id=f"final-qc-invalid-upstream-{artifact_name}",
                  code=f"upstream_{artifact_name.replace('.', '_')}_schema_invalid",
                  message=f"{artifact_name} is not a valid deterministic QC report",
                  source={"artifact": artifact_name},
                  evidence={"schema_invalid": True},
                  next_action="regenerate_upstream_qc",
              )]
          return [
              _finding(
                  finding_id=f"final-qc-upstream-{artifact_name}-{idx}",
                  code=f"upstream_{artifact_name.replace('.', '_')}_{code}",
                  message=f"{artifact_name} reported {code}",
                  source={"artifact": artifact_name},
                  evidence={"upstream_code": code, "upstream_verdict": data["verdict"]},
                  next_action="fix_upstream_qc_blocker",
              )
              for idx, code in enumerate(data["blocking_codes"])
          ]
      
      
      def collect_metadata(work_dir: str | Path, *, final_output: str | Path | None = None,
                           probe_fixture: Any = None, probe_runner: ProbeRunner | None = None) -> dict[str, Any]:
          root = Path(work_dir)
          selected = _final_output_path(root, final_output)
          probe = probe_error = None
          final_meta = _file_metadata(selected, root)
          if final_meta["exists"] and final_meta["bytes"] > 0:
              probe, probe_error = _probe_metadata(selected, probe_fixture=probe_fixture, probe_runner=probe_runner)
          return {
              "work_dir": str(root),
              "final_output": final_meta,
              "artifacts": {name: _artifact_summary(root, name) for name in _COLLECT_ARTIFACTS},
              "probe": probe,
              "probe_error": probe_error,
              # mimo_qc.json is advisory metadata only and is not rolled into final blockers.
              "auto_repair": False,
          }
      
      
      def build_final_qc(work_dir: str | Path, final_output: str | Path | None = None,
                         probe_fixture: Any = None, probe_runner: ProbeRunner | None = None,
                         decode_runner: Callable[[Path], tuple[bool | None, str | None]] | None = None) -> dict[str, Any]:
          root = Path(work_dir)
          selected = _final_output_path(root, final_output)
          metadata = collect_metadata(root, final_output=final_output, probe_fixture=probe_fixture, probe_runner=probe_runner)
          final_meta = metadata["final_output"]
          findings: list[dict[str, Any]] = []
          if not final_meta["exists"]:
              findings.append(_finding(
                  finding_id="final-qc-missing-final-output",
                  code="missing_final_output",
                  message="final output mp4 is missing",
                  category="missing_artifact",
                  source={"artifact": str(final_output) if final_output else "final_output"},
                  evidence={"final_output": final_meta},
                  next_action="render_final_output",
              ))
          elif final_meta["bytes"] == 0:
              findings.append(_finding(
                  finding_id="final-qc-empty-final-output",
                  code="empty_final_output",
                  message="final output mp4 is empty",
                  category="missing_artifact",
                  source={"artifact": final_meta["path"]},
                  evidence={"final_output": final_meta},
                  next_action="rerender_final_output",
              ))
          elif metadata["probe_error"] is not None:
              findings.append(_finding(
                  finding_id="final-qc-probe-failed",
                  code="probe_failed",
                  message="ffprobe failed or was unavailable for existing non-empty final output",
                  category="stream",
                  source={"artifact": final_meta["path"]},
                  evidence=metadata["probe_error"],
                  next_action="inspect_or_rerender_final_output",
              ))
          else:
              probe_findings = _probe_contract_findings(metadata["probe"], final_meta=final_meta)
              findings.extend(probe_findings)
              # Header probing cannot see a container-valid but media-truncated/corrupt payload.
              # A cheap tail decode catches it; skip for offline fixtures and when ffmpeg is absent
              # (decode_ok is None). Only add on a definite decode failure to avoid false-blocking.
              if probe_fixture is None and not probe_findings:
                  decode_ok, decode_detail = (decode_runner or _tail_decode_check)(selected)
                  if decode_ok is False:
                      findings.append(_finding(
                          finding_id="final-qc-undecodable-stream",
                          code="undecodable_stream",
                          message="final output tail failed to decode (truncated or corrupt media payload)",
                          category="stream",
                          source={"artifact": final_meta["path"]},
                          evidence={"decode_error": decode_detail},
                                  next_action="rerender_final_output",
                      ))
          for name in _UPSTREAM_QC_ARTIFACTS:
              findings.extend(_upstream_blockers(root, name))
          return qc_contract.build_report(
              artifact=FINAL_QC_ARTIFACT,
              stage=POST_RENDER_STAGE,
              findings=findings,
              metadata=metadata,
          )
      
      
      def _load_or_build_final_qc(work_dir: Path, final_qc_report: Mapping[str, Any] | None) -> dict[str, Any]:
          if final_qc_report is not None:
              return dict(final_qc_report)
          path = work_dir / FINAL_QC_ARTIFACT
          return load_json(path) if path.exists() else build_final_qc(work_dir)
      
      
      def build_golden_eval(work_dir: str | Path, final_qc_report: Mapping[str, Any] | None = None,
                            golden_fixture: Any = None) -> dict[str, Any]:
          root = Path(work_dir)
          final_report = _load_or_build_final_qc(root, final_qc_report)
          fixture = _load_fixture(golden_fixture) if golden_fixture is not None else {}
          metadata = {
              "work_dir": str(root),
              "fixture": fixture,
              "final_qc": {key: final_report[key] for key in ("ok", "blocker_count", "artifact", "stage")},
              "auto_repair": False,
          }
          findings: list[dict[str, Any]] = []
          expected_ok = fixture.get("expected_final_qc_ok", True)
          if final_report["ok"] != expected_ok:
              findings.append(_finding(
                  finding_id="golden-final-qc-ok-mismatch",
                  stage=GOLDEN_STAGE,
                  code="expected_final_qc_ok_mismatch",
                  message="final_qc ok state does not match golden expectation",
                  category="schema_invalid",
                  source={"artifact": FINAL_QC_ARTIFACT},
                  evidence={"expected": expected_ok, "actual": final_report["ok"]},
                  next_action="fix_final_qc_blockers",
              ))
          final_meta = final_report["metadata"]["final_output"]
          probe = final_report["metadata"]["probe"]
          duration = _probe_duration(probe)[0] if probe else None
          video_stream = _first_video_stream(probe) if probe else None
          codec = video_stream.get("codec_name") if video_stream else None
          min_duration = fixture.get("min_duration")
          if min_duration is not None and (duration is None or duration < min_duration):
              findings.append(_finding(
                  finding_id="golden-min-duration-mismatch",
                  stage=GOLDEN_STAGE,
                  code="min_duration_mismatch",
                  message="final output duration is below golden minimum",
                  category="duration",
                  source={"artifact": final_meta["path"]},
                  evidence={"expected_min_duration": min_duration, "actual_duration": duration},
                  next_action="adjust_render_duration",
              ))
          max_duration = fixture.get("max_duration")
          if max_duration is not None and (duration is None or duration > max_duration):
              findings.append(_finding(
                  finding_id="golden-max-duration-mismatch",
                  stage=GOLDEN_STAGE,
                  code="max_duration_mismatch",
                  message="final output duration is above golden maximum",
                  category="duration",
                  source={"artifact": final_meta["path"]},
                  evidence={"expected_max_duration": max_duration, "actual_duration": duration},
                  next_action="adjust_render_duration",
              ))
          expected_codec = fixture.get("expected_codec")
          if expected_codec is not None and codec != expected_codec:
              findings.append(_finding(
                  finding_id="golden-codec-mismatch",
                  stage=GOLDEN_STAGE,
                  code="codec_mismatch",
                  message="final output video codec does not match golden expectation",
                  category="stream",
                  source={"artifact": final_meta["path"]},
                  evidence={"expected_codec": expected_codec, "actual_codec": codec},
                  next_action="adjust_render_codec",
              ))
          for idx, name in enumerate(fixture.get("required_artifacts", [])):
              artifact_path = root / name
              if not artifact_path.is_file() or artifact_path.stat().st_size == 0:
                  findings.append(_finding(
                      finding_id=f"golden-required-artifact-missing-{idx}",
                      stage=GOLDEN_STAGE,
                      code="required_artifact_missing",
                      message="golden fixture requires an artifact that is missing or empty",
                      category="missing_artifact",
                      source={"artifact": name},
                      evidence={"required_artifact": name},
                      next_action="produce_required_artifact",
                  ))
          metadata["observed"] = {"duration": duration, "codec": codec, "final_output": final_meta}
          return qc_contract.build_report(
              artifact=GOLDEN_EVAL_ARTIFACT,
              stage=GOLDEN_STAGE,
              findings=findings,
              metadata=metadata,
          )
      
      
      def _write_report(path: Path, report: Mapping[str, Any]) -> None:
          path.write_text(json.dumps(report, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
      
      
      def run(work_dir: str | Path, final_output: str | Path | None = None,
              probe_fixture: Any = None, golden_fixture: Any = None,
              probe_runner: ProbeRunner | None = None, only: str = "all") -> dict[str, Any]:
          root = Path(work_dir)
          root.mkdir(parents=True, exist_ok=True)
          mode = {"final": "final_qc", "golden": "golden_eval"}.get(only, only)
          if mode not in {"all", "final_qc", "golden_eval"}:
              raise ValueError("only must be one of all, final_qc, golden_eval, final, golden")
          result: dict[str, Any] = {"work_dir": str(root), "written": []}
          final_report: dict[str, Any] | None = None
          if mode in {"all", "final_qc"}:
              final_report = build_final_qc(root, final_output=final_output, probe_fixture=probe_fixture, probe_runner=probe_runner)
              _write_report(root / FINAL_QC_ARTIFACT, final_report)
              result["final_qc"] = {"ok": final_report["ok"], "blocker_count": final_report["blocker_count"]}
              result["written"].append(FINAL_QC_ARTIFACT)
          if mode in {"all", "golden_eval"}:
              golden_report = build_golden_eval(root, final_qc_report=final_report, golden_fixture=golden_fixture)
              _write_report(root / GOLDEN_EVAL_ARTIFACT, golden_report)
              result["golden_eval"] = {"ok": golden_report["ok"], "blocker_count": golden_report["blocker_count"]}
              result["written"].append(GOLDEN_EVAL_ARTIFACT)
          return result
      
      
      def main(argv: Sequence[str] | None = None) -> int:
          ap = argparse.ArgumentParser(description="Write final_qc.json and golden_eval.json for a video-recap work_dir.")
          ap.add_argument("--work-dir", required=True)
          ap.add_argument("--final-output", default=None)
          ap.add_argument("--probe-fixture", default=None)
          ap.add_argument("--golden-fixture", default=None)
          ap.add_argument("--only", choices=["all", "final_qc", "golden_eval", "final", "golden"], default="all")
          args = ap.parse_args(argv)
          summary = run(
              args.work_dir,
              final_output=args.final_output,
              probe_fixture=args.probe_fixture,
              golden_fixture=args.golden_fixture,
              only=args.only,
          )
          print(json.dumps(summary, ensure_ascii=False, sort_keys=True))
          return 0
      
      
      if __name__ == "__main__":
          raise SystemExit(main())
      
    • lib.py 9.8 KB
      """Self-contained config and MiMo client for the video-recap orchestrator."""
      import json
      import os
      import socket
      import urllib.error
      import urllib.request
      from pathlib import Path
      
      
      # ── 配置 ──────────────────────────────────────────────────────────────
      
      DEFAULT_MIMO_API_URL = "https://api.xiaomimimo.com/v1"
      DEFAULT_MIMO_TOKEN_PLAN_CLUSTER = "cn"
      MIMO_TOKEN_PLAN_API_URLS = {
          "cn": "https://token-plan-cn.xiaomimimo.com/v1",
          "sgp": "https://token-plan-sgp.xiaomimimo.com/v1",
          "ams": "https://token-plan-ams.xiaomimimo.com/v1",
      }
      DEFAULT_MIMO_MODEL = "mimo-v2.5"          # VLM / chat (vision understanding)
      DEFAULT_MIMO_ASR_MODEL = "mimo-v2.5-asr"  # speech-to-text
      DEFAULT_MIMO_TTS_MODEL = "mimo-v2.5-tts"  # text-to-speech
      DEFAULT_FISH_TTS_API_URL = "https://api.fish.audio/v1/tts"
      DEFAULT_FISH_TTS_MODEL = "s2.1-pro-free"
      DEFAULT_FISH_TTS_REFERENCE_ID = "5653cea4ac83480aaf2bf45406556185"
      
      
      def normalize_api_url(raw_url):
          """Normalize a MiMo (OpenAI-compatible) base URL or chat/completions endpoint."""
          url = raw_url.rstrip("/")
          if url.endswith("/chat/completions"):
              return url
          return f"{url}/chat/completions"
      
      
      def is_mimo_token_plan_key(api_key):
          """Return True for Xiaomi MiMo Token Plan keys, which use token-plan base URLs."""
          return api_key.startswith("tp-")
      
      
      def default_mimo_api_url(is_token_plan):
          """Pick the correct MiMo base URL for pay-as-you-go vs Token Plan keys.
      
          MiMo uses independent credentials for pay-as-you-go (`sk-*`) and Token Plan
          (`tp-*`). Token Plan keys must be sent to the Token Plan cluster base URL,
          not the pay-as-you-go `api.xiaomimimo.com` endpoint.
      
          The caller classifies its own key with `is_mimo_token_plan_key` and passes only
          that bit: a credential never reaches a function whose return value is logged.
          """
          if not is_token_plan:
              return DEFAULT_MIMO_API_URL
          cluster = (os.environ.get("MIMO_TOKEN_PLAN_CLUSTER") or DEFAULT_MIMO_TOKEN_PLAN_CLUSTER).strip().lower()
          if cluster not in MIMO_TOKEN_PLAN_API_URLS:
              raise ValueError(
                  f"MIMO_TOKEN_PLAN_CLUSTER must be one of {sorted(MIMO_TOKEN_PLAN_API_URLS)}; got {cluster!r}"
              )
          return MIMO_TOKEN_PLAN_API_URLS[cluster]
      
      
      def env_int(name, default, *, minimum=None):
          """Read an integer env var; a malformed or out-of-range value is a clear error."""
          raw = os.environ.get(name)
          if raw is None or raw == "":
              return default
          try:
              value = int(raw)
          except ValueError:
              raise ValueError(f"{name} must be an integer; got {raw!r}") from None
          if minimum is not None and value < minimum:
              raise ValueError(f"{name} must be >= {minimum}; got {value}")
          return value
      
      
      def env_bool(name, default=False):
          """Read common boolean env var forms."""
          raw = os.environ.get(name)
          if raw is None or raw == "":
              return default
          return raw.strip().lower() in {"1", "true", "yes", "y", "on"}
      
      
      def load_json(path):
          return json.loads(Path(path).read_text(encoding="utf-8"))
      
      
      # Single MiMo credential powers ASR + VLM + TTS. Per-capability overrides
      # (MIMO_VIDEO_API_KEY / MIMO_TTS_API_KEY / MIMO_ASR_API_KEY and their *_API_URL forms)
      # are optional and fall back to MIMO_API_KEY / MIMO_API_URL. Token-Plan keys (tp-*) auto-
      # route to the Token-Plan cluster base URL; pay-as-you-go keys use api.xiaomimimo.com.
      _mimo_api_key = os.environ.get("MIMO_API_KEY", "")
      _mimo_video_api_key = os.environ.get("MIMO_VIDEO_API_KEY", "") or _mimo_api_key
      _mimo_tts_api_key = os.environ.get("MIMO_TTS_API_KEY", "") or _mimo_api_key
      _mimo_asr_api_key = os.environ.get("MIMO_ASR_API_KEY", "") or _mimo_api_key
      _raw_api_url = os.environ.get("MIMO_API_URL") or default_mimo_api_url(is_mimo_token_plan_key(_mimo_api_key))
      _raw_mimo_video_api_url = (
          os.environ.get("MIMO_VIDEO_API_URL")
          or os.environ.get("MIMO_API_URL")
          or default_mimo_api_url(is_mimo_token_plan_key(_mimo_video_api_key))
      )
      _raw_mimo_tts_api_url = (
          os.environ.get("MIMO_TTS_API_URL")
          or os.environ.get("MIMO_API_URL")
          or default_mimo_api_url(is_mimo_token_plan_key(_mimo_tts_api_key))
      )
      _raw_mimo_asr_api_url = (
          os.environ.get("MIMO_ASR_API_URL")
          or os.environ.get("MIMO_API_URL")
          or default_mimo_api_url(is_mimo_token_plan_key(_mimo_asr_api_key))
      )
      
      CONFIG = {
          "api_provider": "mimo",
          "api_url": normalize_api_url(_raw_api_url),
          "api_url_source": "env" if os.environ.get("MIMO_API_URL") else "default",
          "api_key": _mimo_api_key,
          "api_env_var": "MIMO_API_KEY",
          # Read through a COPY of CONFIG by qc.mimo_evidence._effective_config /
          # safe_mimo_config, which is why neither a `CONFIG.get(...)` grep nor live-dict
          # instrumentation sees them. They drive the QC model fallback chain and the
          # provenance recorded in the QC report.
          "mimo_qc_model": os.environ.get("MIMO_QC_MODEL") or os.environ.get("MIMO_VIDEO_MODEL")
          or os.environ.get("MIMO_MODEL", DEFAULT_MIMO_MODEL),
          "mimo_qc_model_source": "env" if os.environ.get("MIMO_QC_MODEL") else "fallback",
          "mimo_model": os.environ.get("MIMO_MODEL", DEFAULT_MIMO_MODEL),
          "mimo_model_source": "env" if os.environ.get("MIMO_MODEL") else "default",
          "mimo_api_url": normalize_api_url(_raw_api_url),
          "mimo_api_url_source": "env" if os.environ.get("MIMO_API_URL") else "default",
          "mimo_video_api_url_source": "env" if (
              os.environ.get("MIMO_VIDEO_API_URL") or os.environ.get("MIMO_API_URL")
          ) else "default",
          "mimo_disable_thinking": env_bool("MIMO_DISABLE_THINKING", True),
          "mimo_disable_thinking_source": "env" if os.environ.get("MIMO_DISABLE_THINKING") else "default",
          "mimo_media_resolution": os.environ.get("MIMO_MEDIA_RESOLUTION", "default"),
          "mimo_media_resolution_source": "env" if os.environ.get("MIMO_MEDIA_RESOLUTION") else "default",
          "mimo_api_key": _mimo_api_key,
          "mimo_video_api_url": normalize_api_url(_raw_mimo_video_api_url),
          "mimo_video_api_key": _mimo_video_api_key,
          "mimo_tts_api_url": normalize_api_url(_raw_mimo_tts_api_url),
          "mimo_tts_api_url_source": "env" if (
              os.environ.get("MIMO_TTS_API_URL") or os.environ.get("MIMO_API_URL")
          ) else "default",
          "mimo_tts_api_key": _mimo_tts_api_key,
          "mimo_asr_api_url": normalize_api_url(_raw_mimo_asr_api_url),
          "mimo_asr_api_url_source": "env" if (
              os.environ.get("MIMO_ASR_API_URL") or os.environ.get("MIMO_API_URL")
          ) else "default",
          "mimo_asr_api_key": _mimo_asr_api_key,
          "mimo_asr_env_var": "MIMO_ASR_API_KEY" if os.environ.get("MIMO_ASR_API_KEY") else "MIMO_API_KEY",
          "mimo_video_model": os.environ.get("MIMO_VIDEO_MODEL") or os.environ.get("MIMO_MODEL", DEFAULT_MIMO_MODEL),
          "mimo_video_model_source": "env" if (
              os.environ.get("MIMO_VIDEO_MODEL") or os.environ.get("MIMO_MODEL")
          ) else "default",
          "vlm_model": os.environ.get("MIMO_MODEL", DEFAULT_MIMO_MODEL),
          "vlm_model_source": "env" if os.environ.get("MIMO_MODEL") else "default",
          "mimo_asr_model": os.environ.get("MIMO_ASR_MODEL", DEFAULT_MIMO_ASR_MODEL),
          "mimo_asr_language": os.environ.get("MIMO_ASR_LANGUAGE", "auto"),  # auto | zh | en
          "mimo_tts_model": os.environ.get("MIMO_TTS_MODEL", DEFAULT_MIMO_TTS_MODEL),
          "mimo_tts_model_source": "env" if os.environ.get("MIMO_TTS_MODEL") else "default",
          "mimo_tts_voice": os.environ.get("MIMO_TTS_VOICE", "冰糖"),
          "mimo_tts_voice_source": "env" if os.environ.get("MIMO_TTS_VOICE") else "default",
          "tts_provider": os.environ.get("TTS_PROVIDER", "auto").strip().lower(),
          "fish_api_key": os.environ.get("FISH_API_KEY", ""),
          "fish_tts_api_url": os.environ.get("FISH_TTS_API_URL", DEFAULT_FISH_TTS_API_URL),
          "fish_tts_model": os.environ.get("FISH_TTS_MODEL", DEFAULT_FISH_TTS_MODEL),
          "fish_tts_reference_id": os.environ.get(
              "FISH_TTS_REFERENCE_ID", DEFAULT_FISH_TTS_REFERENCE_ID
          ).strip(),
          "fish_tts_reference_id_source": (
              "env" if os.environ.get("FISH_TTS_REFERENCE_ID") else "default"
          ),
          "vlm_workers": env_int("VLM_WORKERS", 8, minimum=1),  # VLM 并行分析线程数
      }
      
      
      class MiMoQCRequestError(RuntimeError):
          """Sanitized, fail-open transport error for the advisory QC request."""
      
      
      def mimo_qc_api_call(payload, *, config=None, timeout=60):
          """Send exactly one OpenAI-compatible MiMo request for one QC stage.
      
          Deliberately no retries: the QC feature is advisory, and the orchestrator's
          one-request-per-stage contract is more important than hiding 429/timeout
          behavior. Callers turn every failure into a non-blocking status report.
          """
          cfg = dict(CONFIG)
          if config:
              cfg.update(config)
          api_key = cfg.get("mimo_video_api_key") or cfg.get("mimo_api_key") or cfg.get("api_key")
          if not api_key:
              raise MiMoQCRequestError("missing_key")
          endpoint = normalize_api_url(
              cfg.get("mimo_video_api_url") or cfg.get("mimo_api_url") or cfg.get("api_url")
          )
          request = urllib.request.Request(
              endpoint,
              data=json.dumps(payload, ensure_ascii=False).encode("utf-8"),
              headers={
                  "Content-Type": "application/json",
                  "User-Agent": "video-recap/mimo-qc",
                  "api-key": api_key,
              },
              method="POST",
          )
          try:
              with urllib.request.urlopen(request, timeout=timeout) as response:
                  raw = response.read().decode("utf-8")
          except urllib.error.HTTPError as exc:
              raise MiMoQCRequestError(f"http_{exc.code}") from None
          except (TimeoutError, socket.timeout):
              raise MiMoQCRequestError("timeout") from None
          except (urllib.error.URLError, OSError):
              raise MiMoQCRequestError("network_error") from None
          try:
              result = json.loads(raw)
          except ValueError:
              raise MiMoQCRequestError("invalid_json") from None
          if not isinstance(result, dict):
              raise MiMoQCRequestError("invalid_response")
          return result
      
    • library.py 24.7 KB
      #!/usr/bin/env python3
      """Read-only resource / template / sample library for video-recap.
      
      The library shares its root with the material library (``--material-library-dir``):
      
          <library>/library.json
          <library>/resources/<kind>/<id>/resource.json
          <library>/templates/<kind>/<id>/v<version>/template.json
          <library>/samples/<id>/sample.json
      
      Nothing here writes to the library. ``check`` reports errors (the entry cannot be used)
      and warnings (usable, but a person should look: unknown licence, missing consent, a
      resource that changed since a template was adopted). The record format lives in
      ``references/resource-library.md``.
      """
      from __future__ import annotations
      
      import argparse
      import json
      import os
      import re
      import stat
      import sys
      from pathlib import Path
      
      from materials import file_identity
      
      LIBRARY_ENV = "VIDEO_RECAP_MATERIAL_LIBRARY_DIR"
      LIBRARY_SCHEMA = "video-recap.library.v1"
      RESOURCE_SCHEMA = "video-recap.resource.v1"
      TEMPLATE_SCHEMA = "video-recap.template.v1"
      SAMPLE_SCHEMA = "video-recap.sample.v1"
      
      AUDIO_EXTS = {".wav", ".mp3", ".m4a", ".aac", ".flac", ".ogg"}
      RESOURCE_EXTS = {
          "bgm": AUDIO_EXTS,
          "sfx": AUDIO_EXTS,
          "voice": AUDIO_EXTS,
          "font": {".ttf", ".otf", ".ttc"},
          "image": {".png", ".jpg", ".jpeg", ".webp"},
      }
      TEMPLATE_KINDS = ("subtitle_style", "packaging")
      LICENSE_STATUSES = ("unknown", "owned", "licensed", "restricted")
      CONSENT_STATUSES = ("unknown", "granted", "denied")
      TEMPLATE_STATUSES = ("draft", "adopted", "retired")
      PROVENANCES = ("measured", "fitted", "specified", "unknown")
      SUBTITLE_SIDE_MARGIN = 40  # video-assemble's default left/right subtitle margin at PlayRes = canvas
      SUBTITLE_PARAMS = {
          "font", "size_px", "outline_px", "shadow_px", "primary_color", "outline_color",
          "max_chars", "max_lines", "band",
      }
      VOICE_PROVIDERS = ("mimo-tts", "fish-audio", "index-tts")
      SAMPLE_EXTS = {".mp4", ".mov", ".mkv", ".webm"}
      
      RESOURCE_KEYS = {
          "schema", "id", "kind", "title", "files", "origin", "license", "tags", "notes",
          "voice", "consent", "font",
      }
      TEMPLATE_KEYS = {
          "schema", "id", "version", "kind", "title", "canvas", "params", "samples", "status",
          "adoption", "notes",
      }
      SAMPLE_KEYS = {
          "schema", "id", "title", "file", "canvas", "demonstrates", "not_reusable", "templates",
          "notes",
      }
      ID_RE = re.compile(r"^[a-z0-9][a-z0-9._-]{0,63}$")
      DATE_RE = re.compile(r"^\d{4}-\d{2}-\d{2}$")
      
      
      class Report:
          def __init__(self, root: Path):
              self.root = root
              self.errors: list[dict] = []
              self.warnings: list[dict] = []
      
          def _rel(self, path) -> str:
              try:
                  return Path(path).relative_to(self.root).as_posix()
              except ValueError:
                  return str(path)
      
          def error(self, path, code, message):
              self.errors.append({"path": self._rel(path), "code": code, "message": message})
      
          def warn(self, path, code, message):
              self.warnings.append({"path": self._rel(path), "code": code, "message": message})
      
      
      MAX_RECORD_BYTES = 2 * 1024 * 1024
      
      
      def _load(path: Path, schema: str, keys: set, report: Report):
          try:
              st = os.stat(path)
              if not stat.S_ISREG(st.st_mode) or st.st_size > MAX_RECORD_BYTES:
                  report.error(path, "unreadable", "记录必须是不超过 2 MB 的普通文件")
                  return None
              data = json.loads(path.read_text(encoding="utf-8"))
          except (OSError, ValueError) as exc:
              report.error(path, "unreadable", f"无法解析: {exc}")
              return None
          if not isinstance(data, dict):
              report.error(path, "not_object", "顶层必须是 JSON object")
              return None
          if data.get("schema") != schema:
              report.error(path, "schema", f"schema 必须是 {schema}")
              return None
          unknown = sorted(set(data) - keys)
          if unknown:
              report.error(path, "unknown_keys", f"未知字段: {', '.join(unknown)}")
          return data
      
      
      def _nonempty_str(value) -> bool:
          return isinstance(value, str) and bool(value.strip())
      
      
      def resolve_inside(root: Path, base: Path, rel) -> Path | None:
          """Resolve a record-relative path; None when it is absolute or escapes the library."""
          if not _nonempty_str(rel) or Path(rel).is_absolute():
              return None
          resolved = (base / rel).resolve()
          try:
              resolved.relative_to(root.resolve())
          except ValueError:
              return None
          return resolved
      
      
      def _check_resource(path: Path, data: dict, report: Report) -> dict | None:
          rdir = path.parent
          kind = rdir.parent.name
          if data.get("id") != rdir.name or not ID_RE.match(str(data.get("id", ""))):
              report.error(path, "id", f"id 必须与目录名 {rdir.name} 一致,且只用小写字母、数字、. _ -")
              return None
          if data.get("kind") != kind or kind not in RESOURCE_EXTS:
              report.error(path, "kind", f"kind 必须是所在目录 {kind},且属于 {sorted(RESOURCE_EXTS)}")
              return None
          if not _nonempty_str(data.get("title")):
              report.error(path, "title", "缺少 title")
          files = data.get("files", [])
          if not isinstance(files, list):
              report.error(path, "files", "files 必须是数组")
              files = []
          resolved = []
          for index, item in enumerate(files):
              rel = item.get("path") if isinstance(item, dict) else None
              target = resolve_inside(report.root, rdir, rel)
              if target is None:
                  report.error(path, "file_path", f"files[{index}].path 必须是库内相对路径")
                  continue
              if not _nonempty_str(item.get("role")):
                  report.error(path, "file_role", f"files[{index}] 缺少 role")
              if target.suffix.lower() not in RESOURCE_EXTS[kind]:
                  report.error(path, "file_ext", f"files[{index}] 扩展名不适用于 {kind}: {target.name}")
              if not target.is_file():
                  report.error(path, "file_missing", f"文件不存在: {rel}")
                  continue
              resolved.append({"role": item.get("role"), "path": str(target), **file_identity(target)})
          if kind != "voice" and not files:
              report.error(path, "files", f"{kind} 资源至少需要一个文件")
          licence = data.get("license")
          status = licence.get("status") if isinstance(licence, dict) else None
          if status not in LICENSE_STATUSES:
              report.error(path, "license", f"license.status 必须是 {LICENSE_STATUSES} 之一")
          elif status in {"unknown", "restricted"}:
              report.warn(path, f"license_{status}", f"授权状态为 {status},交付前需要人工确认")
          if kind == "voice":
              voice = data.get("voice")
              if not isinstance(voice, dict) or voice.get("provider") not in VOICE_PROVIDERS:
                  report.error(path, "voice", f"voice.provider 必须是 {VOICE_PROVIDERS} 之一")
              elif voice["provider"] != "mimo-tts" and not _nonempty_str(voice.get("voice_id")):
                  report.error(path, "voice", f"{voice['provider']} 音色需要 voice.voice_id(只有 mimo-tts 支持参考音频克隆)")
              elif not (_nonempty_str(voice.get("voice_id")) or files):
                  report.error(path, "voice", "voice 资源需要 voice.voice_id 或一个参考音频文件")
              if files:
                  consent = data.get("consent")
                  consent_status = consent.get("status") if isinstance(consent, dict) else None
                  if consent_status not in CONSENT_STATUSES:
                      report.error(path, "consent", f"参考音频需要 consent.status({CONSENT_STATUSES})")
                  elif consent_status != "granted":
                      report.warn(path, f"consent_{consent_status}", "参考音频的声音授权未确认")
          return {"id": data["id"], "kind": kind, "title": data.get("title", ""),
                  "license": status, "files": resolved, "record": str(path)}
      
      
      def _param_values(node, where="params"):
          """Yield (where, dict) for every nested dict in a template's params."""
          if isinstance(node, dict):
              yield where, node
              for key, value in node.items():
                  yield from _param_values(value, f"{where}.{key}")
          elif isinstance(node, list):
              for index, value in enumerate(node):
                  yield from _param_values(value, f"{where}[{index}]")
      
      
      def _check_rect(path, where, rect, canvas, report):
          fields = ("x", "y", "width", "height")
          if not isinstance(rect, dict) or not all(_is_int(rect.get(f)) for f in fields):
              report.error(path, "rect", f"{where} 需要整数 x / y / width / height")
              return
          if (rect["x"] < 0 or rect["y"] < 0 or rect["width"] <= 0 or rect["height"] <= 0
                  or rect["x"] + rect["width"] > canvas["width"]
                  or rect["y"] + rect["height"] > canvas["height"]):
              report.error(path, "rect_outside_canvas", f"{where} 超出画布 {canvas['width']}x{canvas['height']}")
      
      
      def _is_int(value) -> bool:
          return isinstance(value, int) and not isinstance(value, bool)
      
      
      def _param_number(params, key):
          node = params.get(key)
          value = node.get("value") if isinstance(node, dict) else None
          return value if isinstance(value, (int, float)) and not isinstance(value, bool) and value > 0 else None
      
      
      def _check_template(path: Path, data: dict, report: Report) -> dict | None:
          vdir = path.parent
          tid, kind = vdir.parent.name, vdir.parent.parent.name
          version = data.get("version")
          if data.get("id") != tid or not ID_RE.match(str(tid)):
              report.error(path, "id", f"id 必须与目录名 {tid} 一致")
              return None
          if not isinstance(version, int) or isinstance(version, bool) or version < 1 or vdir.name != f"v{version}":
              report.error(path, "version", f"version 必须是正整数,且目录名为 v<version>(当前 {vdir.name})")
              return None
          if data.get("kind") != kind or kind not in TEMPLATE_KINDS:
              report.error(path, "kind", f"kind 必须是所在目录 {kind},且属于 {TEMPLATE_KINDS}")
              return None
          canvas = data.get("canvas")
          if not (isinstance(canvas, dict) and all(isinstance(canvas.get(k), int) and canvas[k] > 0
                                                   for k in ("width", "height"))):
              report.error(path, "canvas", "canvas 需要正整数 width / height")
              canvas = None
          params = data.get("params")
          if not isinstance(params, dict):
              report.error(path, "params", "params 必须是 object")
              params = {}
          refs = []
          for where, node in _param_values(params):
              if "provenance" in node:
                  if node["provenance"] not in PROVENANCES or "value" not in node:
                      report.error(path, "provenance", f"{where} 需要 value 与 provenance({PROVENANCES})")
              if "resource" in node:
                  if _nonempty_str(node["resource"]):
                      refs.append((where, node["resource"]))
                  else:
                      report.error(path, "resource_ref", f"{where}.resource 必须是资源 id")
          if kind == "subtitle_style":
              unknown = sorted(set(params) - SUBTITLE_PARAMS)
              if unknown:
                  report.error(path, "unknown_params", f"subtitle_style 不认识的参数: {', '.join(unknown)}")
              font = params.get("font")
              if not (isinstance(font, dict) and (_nonempty_str(font.get("resource")) or _nonempty_str(font.get("family")))):
                  report.error(path, "font", "subtitle_style 需要 params.font.resource 或 params.font.family")
              if not isinstance(params.get("size_px"), dict):
                  report.error(path, "size_px", "subtitle_style 需要 params.size_px")
              for key in ("size_px", "max_chars", "max_lines"):
                  node = params.get(key)
                  if node is not None and not (isinstance(node, dict) and _is_int(node.get("value")) and node["value"] > 0):
                      report.error(path, "integer", f"params.{key}.value 必须是正整数(渲染按整数像素与字数处理)")
              for key in ("outline_px", "shadow_px"):
                  node = params.get(key)
                  value = node.get("value") if isinstance(node, dict) else None
                  if node is not None and not (isinstance(value, (int, float)) and not isinstance(value, bool) and value >= 0):
                      report.error(path, "number", f"params.{key}.value 必须是非负数")
              size, per_line = (_param_number(params, "size_px"), _param_number(params, "max_chars"))
              if per_line is None:
                  report.error(path, "max_chars", "subtitle_style 需要 params.max_chars:按最长一行在这块画布上校准")
              elif canvas and size and size * per_line > canvas["width"] - 2 * SUBTITLE_SIDE_MARGIN:
                  report.error(path, "line_too_wide",
                               f"{per_line} 字 × {size}px 超过画布宽 {canvas['width']} 减去两侧各 "
                               f"{SUBTITLE_SIDE_MARGIN}px 边距;减小字号或每行字数")
              band = params.get("band", {}).get("value") if isinstance(params.get("band"), dict) else None
              if band is not None and canvas and not (
                  isinstance(band, dict) and _is_int(band.get("y_top")) and _is_int(band.get("y_bot"))
                  and 0 <= band["y_top"] < band["y_bot"] <= canvas["height"]
              ):
                  report.error(path, "band", "params.band.value 需要 0 <= y_top < y_bot <= 画布高度")
          if kind == "packaging":
              layers = params.get("layers")
              if not isinstance(layers, list) or not layers:
                  report.error(path, "layers", "packaging 需要非空 params.layers")
                  layers = []
              names = [layer.get("name") for layer in layers if isinstance(layer, dict)]
              if len(names) != len(set(names)) or not all(_nonempty_str(n) for n in names):
                  report.error(path, "layers", "每个图层需要唯一的 name")
              for index, layer in enumerate(layers):
                  if not (isinstance(layer, dict) and isinstance(layer.get("image"), dict)
                          and _nonempty_str(layer["image"].get("resource"))):
                      report.error(path, "layer", f"params.layers[{index}] 需要 image.resource")
                      continue
                  if canvas:
                      _check_rect(path, f"params.layers[{index}].rect", layer.get("rect"), canvas, report)
              if canvas and "safe_rect" in params:
                  _check_rect(path, "params.safe_rect", params["safe_rect"], canvas, report)
          status = data.get("status")
          if status not in TEMPLATE_STATUSES:
              report.error(path, "status", f"status 必须是 {TEMPLATE_STATUSES} 之一")
          adoption = data.get("adoption")
          if status == "adopted":
              if not (isinstance(adoption, dict) and DATE_RE.match(str(adoption.get("date", "")))
                      and all(_nonempty_str(adoption.get(k)) for k in ("by", "statement", "scope"))):
                  report.error(path, "adoption", "adopted 模板需要 adoption.date(YYYY-MM-DD) / by / statement / scope")
          samples = data.get("samples", [])
          if not (isinstance(samples, list) and all(_nonempty_str(x) for x in samples)):
              report.error(path, "samples", "samples 必须是样片 id 数组")
              samples = []
          snapshot = adoption.get("resources") if isinstance(adoption, dict) else None
          if snapshot is not None and not (
              isinstance(snapshot, dict)
              and all(isinstance(v, list) and all(isinstance(f, dict) for f in v) for v in snapshot.values())
          ):
              report.error(path, "adoption_resources", "adoption.resources 必须是 {资源 id: [{path, size, mtime_ns}]}")
              adoption = {k: v for k, v in adoption.items() if k != "resources"}
          return {"id": tid, "version": version, "kind": kind, "title": data.get("title", ""),
                  "status": status, "canvas": canvas, "refs": refs,
                  "samples": samples, "adoption": adoption, "record": str(path)}
      
      
      def _check_sample(path: Path, data: dict, report: Report) -> dict | None:
          sdir = path.parent
          if data.get("id") != sdir.name or not ID_RE.match(sdir.name):
              report.error(path, "id", f"id 必须与目录名 {sdir.name} 一致")
              return None
          for key in ("title", "not_reusable"):
              if not _nonempty_str(data.get(key)):
                  report.error(path, key, f"缺少 {key}")
          if not (isinstance(data.get("demonstrates"), list) and data["demonstrates"]
                  and all(_nonempty_str(x) for x in data["demonstrates"])):
              report.error(path, "demonstrates", "demonstrates 需要非空字符串数组")
          rel = data.get("file", {}).get("path") if isinstance(data.get("file"), dict) else None
          target = None
          if _nonempty_str(rel) and Path(rel).is_absolute():
              target = Path(rel)
              if not target.is_file():
                  report.warn(path, "sample_offline", f"样片不在本机: {rel}")
                  target = None
          else:
              target = resolve_inside(report.root, sdir, rel)
              if target is None:
                  report.error(path, "file_path", "file.path 必须是库内相对路径或绝对路径")
              elif not target.is_file():
                  report.error(path, "file_missing", f"样片不存在: {rel}")
                  target = None
          if target is not None and target.suffix.lower() not in SAMPLE_EXTS:
              report.error(path, "file_ext", f"样片扩展名不支持: {target.name}")
          templates = data.get("templates", [])
          if not (isinstance(templates, list) and all(_nonempty_str(x) for x in templates)):
              report.error(path, "templates", "templates 必须是 id@vN 字符串数组")
              templates = []
          return {"id": data["id"], "title": data.get("title", ""), "file": str(target) if target else None,
                  "templates": templates, "record": str(path)}
      
      
      def _check_links(index: dict, report: Report):
          resources = {r["id"]: r for r in index["resources"]}
          templates = {f"{t['id']}@v{t['version']}": t for t in index["templates"]}
          samples = {s["id"] for s in index["samples"]}
          expected_kind = {"params.font": "font"}
          for template in index["templates"]:
              path = Path(template["record"])
              for where, ref in template["refs"]:
                  resource = resources.get(ref)
                  want = expected_kind.get(where) or ("image" if where.endswith(".image") else None)
                  if resource is None:
                      report.error(path, "resource_missing", f"{where} 引用的资源不存在: {ref}")
                  elif want and resource["kind"] != want:
                      report.error(path, "resource_kind", f"{where} 需要 {want} 资源,{ref} 是 {resource['kind']}")
                  elif want == "font" and not _nonempty_str(
                      (json.loads(Path(resource["record"]).read_text(encoding="utf-8")).get("font") or {}).get("family")
                  ):
                      report.error(path, "font_family", f"字体资源 {ref} 需要 font.family,渲染才能按名字找到这份字体")
              for sample_id in template["samples"]:
                  if sample_id not in samples:
                      report.error(path, "sample_missing", f"引用的样片不存在: {sample_id}")
              snapshot = (template["adoption"] or {}).get("resources") if isinstance(template["adoption"], dict) else None
              for res_id, files in (snapshot or {}).items():
                  current = {Path(f["path"]).relative_to(report.root.resolve()).as_posix(): f
                             for f in resources.get(res_id, {}).get("files", [])}
                  for recorded in files if isinstance(files, list) else []:
                      now = current.get(recorded.get("path"))
                      if now is None or (now["size"], now["mtime_ns"]) != (recorded.get("size"), recorded.get("mtime_ns")):
                          report.warn(path, "changed_since_adoption",
                                      f"采用后资源已变化: {res_id} {recorded.get('path')},需重新采用或出新版本")
          for sample in index["samples"]:
              for ref in sample["templates"]:
                  if ref not in templates:
                      report.error(Path(sample["record"]), "template_missing", f"引用的模板不存在: {ref}")
      
      
      def _guarded(check, path, data, report):
          """Run one record check; a shape the checks did not anticipate becomes an error, never a crash."""
          if not data:
              return None
          try:
              return check(path, data, report)
          except (TypeError, KeyError, AttributeError, ValueError, IndexError) as exc:
              report.error(path, "malformed", f"记录结构无法识别: {type(exc).__name__}: {exc}")
              return None
      
      
      def scan_library(root) -> tuple[dict, Report]:
          """Load and validate every record under ``root``; never writes."""
          root = Path(root).resolve()
          report = Report(root)
          index = {"resources": [], "templates": [], "samples": []}
          meta = root / "library.json"
          if meta.is_file():
              _load(meta, LIBRARY_SCHEMA, {"schema", "name", "notes"}, report)
          else:
              report.warn(meta, "no_library_json", "缺少 library.json(仍会扫描资源、模板与样片)")
          for path in sorted(root.glob("resources/*/*/resource.json")):
              data = _load(path, RESOURCE_SCHEMA, RESOURCE_KEYS, report)
              entry = _guarded(_check_resource, path, data, report)
              if entry:
                  index["resources"].append(entry)
          for path in sorted(root.glob("templates/*/*/v*/template.json")):
              data = _load(path, TEMPLATE_SCHEMA, TEMPLATE_KEYS, report)
              entry = _guarded(_check_template, path, data, report)
              if entry:
                  index["templates"].append(entry)
          for path in sorted(root.glob("samples/*/sample.json")):
              data = _load(path, SAMPLE_SCHEMA, SAMPLE_KEYS, report)
              entry = _guarded(_check_sample, path, data, report)
              if entry:
                  index["samples"].append(entry)
          try:
              _check_links(index, report)
          except (TypeError, KeyError, AttributeError, ValueError, IndexError, OSError) as exc:
              report.error(root, "malformed_links", f"交叉引用无法核对: {type(exc).__name__}: {exc}")
          return index, report
      
      
      def _library_root(value):
          root = value or os.environ.get(LIBRARY_ENV)
          if not root:
              raise SystemExit(f"需要 --library-dir 或环境变量 {LIBRARY_ENV}")
          root = Path(root).expanduser().resolve()
          if not root.is_dir():
              raise SystemExit(f"库目录不存在: {root}")
          return root
      
      
      def _print_issues(report: Report):
          for level, items in (("ERROR", report.errors), ("WARN", report.warnings)):
              for item in items:
                  print(f"{level} {item['path']} [{item['code']}] {item['message']}")
      
      
      def main(argv=None):
          ap = argparse.ArgumentParser(description="video-recap 资源库:只读列出、查看与校验。")
          ap.add_argument("--library-dir", help=f"库根目录(默认读取 {LIBRARY_ENV})")
          sub = ap.add_subparsers(dest="command", required=True)
          list_cmd = sub.add_parser("list", help="列出资源、模板与样片")
          list_cmd.add_argument("--kind", help="只看某一类,如 bgm / font / subtitle_style")
          list_cmd.add_argument("--json", action="store_true")
          show_cmd = sub.add_parser("show", help="查看一个条目(模板可写 id 或 id@vN)")
          show_cmd.add_argument("id")
          check_cmd = sub.add_parser("check", help="校验整个库;有错误时退出码为 1")
          check_cmd.add_argument("--json", action="store_true")
          args = ap.parse_args(argv)
          index, report = scan_library(_library_root(args.library_dir))
      
          if args.command == "check":
              if args.json:
                  print(json.dumps({"ok": not report.errors, "errors": report.errors,
                                    "warnings": report.warnings,
                                    "counts": {k: len(v) for k, v in index.items()}},
                                   ensure_ascii=False, indent=2))
              else:
                  _print_issues(report)
                  print(f"{len(report.errors)} error(s), {len(report.warnings)} warning(s); "
                        + ", ".join(f"{k} {len(v)}" for k, v in index.items()))
              return 1 if report.errors else 0
      
          rows = [{"type": "resource", "kind": r["kind"], "id": r["id"], "title": r["title"],
                   "status": r["license"]} for r in index["resources"]]
          rows += [{"type": "template", "kind": t["kind"], "id": f"{t['id']}@v{t['version']}",
                    "title": t["title"], "status": t["status"]} for t in index["templates"]]
          rows += [{"type": "sample", "kind": "sample", "id": s["id"], "title": s["title"],
                    "status": "online" if s["file"] else "offline"} for s in index["samples"]]
          if args.command == "list":
              rows = [r for r in rows if not args.kind or r["kind"] == args.kind]
              if args.json:
                  print(json.dumps(rows, ensure_ascii=False, indent=2))
              else:
                  for r in rows:
                      print(f"{r['type']:<9} {r['kind']:<15} {r['id']:<36} {r['status']:<10} {r['title']}")
              return 0
      
          matches = [e for group in index.values() for e in group
                     if args.id in {e["id"], f"{e['id']}@v{e.get('version')}"}]
          if not matches:
              print(f"未找到: {args.id}", file=sys.stderr)
              return 1
          for entry in matches:
              record = json.loads(Path(entry["record"]).read_text(encoding="utf-8"))
              print(json.dumps({"record": entry["record"], **record,
                                **({"resolved_files": entry["files"]} if "files" in entry else {})},
                               ensure_ascii=False, indent=2))
          return 0
      
      
      if __name__ == "__main__":
          sys.exit(main())
      
    • materials.py 16 KB
      """Filesystem material library helpers for video-recap.
      
      The library is intentionally grep-friendly: JSON/MD/JSONL files on disk, no DB,
      no embeddings, no raw-media copies. Current metadata lives in each material
      folder; the root ``materials_index.jsonl`` is an append-only journal for grep and
      history.
      
      Secret handling is BEST-EFFORT defense-in-depth, NOT a guarantee. ``_redact_text``
      masks common credential value shapes (``tp-``/``sk-``/``gh*_``/``AKIA``/JWT and
      ``KEY=VALUE`` assignments) and ``_redact_json`` drops the value of exact
      credential-named keys, but a secret in an unrecognized format can still slip
      through. Keep secrets (API keys, tokens) out of the analysis artifacts in the
      first place — the key is read from the environment/``.env`` and never needs to be
      written into scenes/ASR/VLM/summary JSON.
      """
      from __future__ import annotations
      
      import json
      import os
      import re
      import tempfile
      from datetime import datetime, timezone
      from pathlib import Path
      
      from lib import load_json
      
      ALLOWED_ARTIFACTS = {
          "scenes.json",
          "asr_result.json",
          "asr_clean.json",
          "vlm_analysis.json",
          "silence_periods.json",
          "speech_boundary_anchors.json",
          "speech_boundary_anchors_output.json",
          "timeline_fusion.json",
          "understanding_index.json",
          "understanding_index.md",
          "agent_narration_brief.md",
          "background_research.json",
          "reference_profile.json",
          "reference_match_report.json",
          "recap_run_manifest.json",
      }
      # Redaction targets credential VALUE shapes, not English/Chinese dictionary words — the
      # library must stay a faithful copy of the analysis. Bare words like "secret"/"token"
      # legitimately appear in transcripts/summaries and must NOT be touched.
      SECRET_VALUE_RES = (
          re.compile(r"\btp-[A-Za-z0-9_-]{8,}\b"),     # MiMo Token Plan keys
          re.compile(r"\bsk-[A-Za-z0-9_-]{16,}\b"),    # OpenAI-style keys
          re.compile(r"\bgh[pousr]_[A-Za-z0-9]{20,}\b"),  # GitHub tokens
          re.compile(r"\bAKIA[0-9A-Z]{16}\b"),         # AWS access key id
          re.compile(r"\beyJ[A-Za-z0-9_-]{6,}\.[A-Za-z0-9_-]{6,}\.[A-Za-z0-9_-]{6,}\b"),  # JWT
      )
      # `KEY=VALUE` / `"key": "value"` assignments whose key denotes a credential -> mask the VALUE.
      SECRET_ASSIGN_RE = re.compile(
          r"(?i)(\b(?:mimo(?:_\w+)?_api_key|api_key|secret_key|access_token|refresh_token|"
          r"authorization|password|passwd|bearer)\b\s*[:=]\s*)(\"?)([^\s\"',;]+)(\"?)"
      )
      # Exact JSON/dict key names whose VALUE is a credential and must be dropped (key name kept).
      SECRET_KEY_NAMES = frozenset({
          "api_key", "mimo_api_key", "mimo_asr_api_key", "mimo_tts_api_key", "mimo_video_api_key",
          "secret_key", "access_token", "refresh_token", "authorization", "password", "passwd",
      })
      
      
      def utc_now_iso() -> str:
          return datetime.now(timezone.utc).replace(microsecond=0).isoformat().replace("+00:00", "Z")
      
      
      def file_identity(path: str | Path) -> dict:
          """``{size, mtime_ns}`` of a file: the identity recap records and compares for a source
          video or adopted artifact. A file rewritten in place gets a new mtime_ns."""
          st = os.stat(os.fspath(path))
          return {"size": st.st_size, "mtime_ns": st.st_mtime_ns}
      
      
      def _slug(text: str, max_len: int = 48) -> str:
          raw = Path(text).stem.lower()
          raw = re.sub(r"[^a-z0-9\u4e00-\u9fff._-]+", "-", raw).strip("-._")
          return (raw or "material")[:max_len].strip("-._") or "material"
      
      
      def _id_stem(source_path: str | Path, max_len: int = 32) -> str:
          raw = re.sub(r"[^a-z0-9]+", "-", Path(source_path).stem.lower()).strip("-")
          return (raw or "source")[:max_len].strip("-") or "source"
      
      
      def source_id_for(source_path: str | Path) -> str:
          """``src_<stem>_<size>``: readable, stable across runs, and distinct for a different cut
          of the same title (the size changes)."""
          return f"src_{_id_stem(source_path)}_{os.stat(os.fspath(source_path)).st_size}"
      
      
      def assign_source_ids(sources: list[dict]) -> list[dict]:
          """Assign deterministic source_id values to manifest source records.
      
          Base id is ``source_id_for(source_path)``. When two records in one project share a
          base id, the first keeps it and later ones get ``_2``, ``_3``… suffixes in input order.
          """
          used: set[str] = set()
          assigned = []
          for raw in sources:
              item = dict(raw)
              path = str(Path(item["source_path"]).resolve())
              base = source_id_for(path)
              sid = base
              suffix = 2
              while sid in used:
                  sid = f"{base}_{suffix}"
                  suffix += 1
              used.add(sid)
              item["source_id"] = sid
              item["source_path"] = path
              assigned.append(item)
          return assigned
      
      
      def material_id_for(source_path: str | Path, source_identity: dict) -> str:
          return f"{_slug(str(source_path))}-{source_identity['size']}"
      
      
      def material_dir(library_dir: str | Path, material_id: str) -> Path:
          return Path(library_dir) / "materials" / material_id
      
      
      def _redact_text(text: str) -> str:
          """Redact credential value shapes only; leave ordinary words (secret/token/…) intact."""
          for rx in SECRET_VALUE_RES:
              text = rx.sub("[redacted-token]", text)
          return SECRET_ASSIGN_RE.sub(lambda m: f"{m.group(1)}{m.group(2)}[redacted-key]{m.group(4)}", text)
      
      
      def _redact_json(value):
          if isinstance(value, dict):
              out = {}
              for key, item in value.items():
                  # Drop only the value of an exact credential-named key; keep the key name and
                  # never coalesce distinct keys (so benign fields like token_economy survive).
                  if key.strip().lower() in SECRET_KEY_NAMES:
                      out[key] = "[redacted]"
                  else:
                      out[key] = _redact_json(item)
              return out
          if isinstance(value, list):
              return [_redact_json(item) for item in value]
          if isinstance(value, str):
              return _redact_text(value)
          return value
      
      
      def copy_artifact_redacted(src: Path, dst: Path) -> None:
          """Copy an allowed JSON/MD artifact without persisting obvious secret markers."""
          text = src.read_text(encoding="utf-8")
          if src.suffix.lower() == ".json":
              dst.write_text(json.dumps(_redact_json(json.loads(text)), ensure_ascii=False, indent=2), encoding="utf-8")
          else:
              dst.write_text(_redact_text(text), encoding="utf-8")
      
      
      def summarize_work_dir(work_dir: str | Path, *, source_name: str = "") -> dict:
          """Grep-friendly tags for material.md/index: character and entity names plus artifact counts."""
          work = Path(work_dir)
          summary = f"Analyzed video material: {source_name or work.name}"
          tags = []
          index_path = work / "understanding_index.json"
          if index_path.exists():
              # consolidate.py writes {characters, relationships, plot_points, entities,
              # research_glossary} and no prose summary. The list ITEMS are MiMo output that
              # consolidate only list-checks, so their shape is still tolerated here.
              index = load_json(index_path)
              for key in ("characters", "entities"):
                  for item in index[key][:12]:
                      text = item.get("name", "") if isinstance(item, dict) else item
                      tags.append(_redact_text(str(text))[:80])
          for name, label in (("scenes.json", "scenes"), ("asr_result.json", "asr")):
              if (work / name).exists():
                  tags.append(f"{label}:{len(load_json(work / name))}")
          return {
              "summary": summary,
              "tags": list(dict.fromkeys(tag for tag in tags if tag))[:20],
          }
      
      
      def allowed_artifact_paths(work_dir: str | Path) -> list[Path]:
          work = Path(work_dir)
          return [work / name for name in sorted(ALLOWED_ARTIFACTS) if (work / name).is_file()]
      
      
      def write_material_md(path: Path, metadata: dict, summary: str, tags: list[str]) -> None:
          artifact_lines = "\n".join(f"- `{a['name']}` → `{a['path']}`" for a in metadata["artifacts"])
          tags_text = ", ".join(tags) if tags else "(none)"
          text = f"""# Material: {metadata['source_name']}
      
      - material_id: `{metadata['material_id']}`
      - source: `{metadata['source_path']}`
      - source_identity: `{json.dumps(metadata['source_video_identity'], sort_keys=True)}`
      - settings: `{json.dumps(metadata['settings'], ensure_ascii=False, sort_keys=True)}`
      - updated_at: `{metadata['updated_at']}`
      - tags: {tags_text}
      
      ## Summary
      {summary}
      
      ## Artifacts
      {artifact_lines or '- (none)'}
      """
          path.write_text(text, encoding="utf-8")
      
      
      def _read_material_metadata(path: Path) -> dict | None:
          try:
              data = load_json(path)
          except (OSError, ValueError):
              return None
          return data if isinstance(data, dict) else None
      
      
      def _material_cache_entry(path: Path) -> dict | None:
          data = _read_material_metadata(path)
          if not data or not all(data.get(key) for key in (
              "source_path", "updated_at", "material_id",
          )) or not isinstance(data.get("artifacts"), list):
              return None
          if not isinstance(data.get("source_video_identity"), dict) \
                  or not isinstance(data.get("settings"), dict):
              return None
          return data
      
      
      def save_material(
          library_dir: str | Path,
          work_dir: str | Path,
          source_path: str | Path,
          source_identity: dict,
          settings: dict,
          *,
          duration: float | None = None,
          source_id: str | None = None,
          material_id: str | None = None,
          now: str | None = None,
      ) -> dict:
          """Persist small reusable analysis artifacts into the filesystem library."""
          lib = Path(library_dir)
          mid = material_id or material_id_for(source_path, source_identity)
          dest = material_dir(lib, mid)
          artifacts_dir = dest / "artifacts"
          artifacts_dir.mkdir(parents=True, exist_ok=True)
          now = now or utc_now_iso()
      
          sources = allowed_artifact_paths(work_dir)
          fresh_names = {src.name for src in sources}
          # Reconcile: drop allowed artifacts left from a previous (larger) save so the on-disk
          # artifacts/ dir always matches material.json — a stale orphan must not surface in greps.
          for stale in artifacts_dir.iterdir():
              if stale.is_file() and stale.name in ALLOWED_ARTIFACTS and stale.name not in fresh_names:
                  stale.unlink()
          copied = []
          for src in sources:
              dst = artifacts_dir / src.name
              copy_artifact_redacted(src, dst)
              copied.append({"name": src.name, "path": f"artifacts/{src.name}", "bytes": dst.stat().st_size})
      
          meta_path = dest / "material.json"
          previous = _read_material_metadata(meta_path)
          created_at = previous.get("created_at") if previous else None
          if not isinstance(created_at, str) or not created_at:
              created_at = now
          source_path = Path(source_path).resolve()
          summary_info = summarize_work_dir(work_dir, source_name=source_path.name)
          metadata = {
              "schema_version": 1,
              "material_id": mid,
              "source_id": source_id,
              "source_name": source_path.name,
              "source_path": str(source_path),
              "source_video_identity": dict(source_identity),
              "duration": duration,
              "settings": dict(settings),
              "artifacts": copied,
              "created_at": created_at,
              "updated_at": now,
          }
          meta_path.write_text(json.dumps(metadata, ensure_ascii=False, indent=2), encoding="utf-8")
          write_material_md(dest / "material.md", metadata, summary_info["summary"], summary_info["tags"])
      
          index_record = {
              "schema_version": 1,
              "event": "saved",
              "material_id": mid,
              "source_name": metadata["source_name"],
              "source_path": metadata["source_path"],
              "source_video_identity": dict(source_identity),
              "summary": summary_info["summary"],
              "tags": summary_info["tags"],
              "material_dir": str(dest),
              "updated_at": now,
          }
          lib.mkdir(parents=True, exist_ok=True)
          with (lib / "materials_index.jsonl").open("a", encoding="utf-8") as f:
              f.write(json.dumps(index_record, ensure_ascii=False, separators=(",", ":")) + "\n")
          return metadata
      
      
      def find_material_by_source(
          library_dir: str | Path, source_path: str | Path, source_identity: dict
      ) -> dict | None:
          """The newest material saved for this exact source file (resolved path + identity)."""
          root = Path(library_dir) / "materials"
          if not root.exists():
              return None
          resolved = str(Path(source_path).resolve())
          candidates = []
          for meta_path in root.glob("*/material.json"):
              data = _material_cache_entry(meta_path)
              if (
                  data is not None
                  and data["source_path"] == resolved
                  and data["source_video_identity"] == source_identity
              ):
                  data["material_dir"] = str(meta_path.parent)
                  candidates.append(data)
          if not candidates:
              return None
          # Deterministic fallback policy for legacy/manual callers that do not know
          # the expected material_id: newest wins, then material_id for stable ties.
          candidates.sort(key=lambda d: (d["updated_at"], d["material_id"]), reverse=True)
          return candidates[0]
      
      
      def restore_material(
          library_dir: str | Path,
          work_dir: str | Path,
          *,
          source_path: str | Path,
          source_identity: dict,
          settings: dict,
          material_id: str | None = None,
          overwrite: bool = True,
          prune_stale_allowed: bool = True,
      ) -> dict:
          """Restore allowed artifacts when source path, identity and settings all match.
      
          Returns a status dict and never partially restores on mismatch.
      
          By default the restored material is treated as the authoritative analysis
          snapshot for ``work_dir``: allowed analysis artifacts are staged first, then
          stale allowed artifacts in the destination are pruned before replacement.
          This prevents reused work dirs from mixing old scenes/ASR/VLM files with a
          newly restored material. Non-allowed files (for example narration.json or
          clip_plan.json) are never removed here.
          """
          lib = Path(library_dir)
          if material_id:
              meta_path = material_dir(lib, material_id) / "material.json"
              meta = _material_cache_entry(meta_path)
              if meta is None:
                  return {"restored": False, "reason": "material missing or invalid"}
              meta["material_dir"] = str(meta_path.parent)
          else:
              meta = find_material_by_source(lib, source_path, source_identity)
              if meta is None:
                  return {"restored": False, "reason": "material not found"}
          if meta["source_path"] != str(Path(source_path).resolve()) \
                  or meta["source_video_identity"] != source_identity:
              return {"restored": False, "reason": "source identity mismatch", "material_id": meta["material_id"]}
          if meta["settings"] != settings:
              return {"restored": False, "reason": "settings mismatch", "material_id": meta["material_id"]}
      
          src_dir = Path(meta["material_dir"]) / "artifacts"
          if not src_dir.exists():
              return {"restored": False, "reason": "material artifacts missing", "material_id": meta["material_id"]}
          dest = Path(work_dir)
          dest.mkdir(parents=True, exist_ok=True)
          with tempfile.TemporaryDirectory(prefix=".material_restore_", dir=str(dest)) as tmp_name:
              tmp = Path(tmp_name)
              staged = [artifact["name"] for artifact in meta["artifacts"]]  # save_material's own list
              if not staged or any(not (src_dir / name).is_file() for name in staged):
                  return {"restored": False, "reason": "material artifacts missing", "material_id": meta["material_id"]}
              for name in staged:
                  copy_artifact_redacted(src_dir / name, tmp / name)
      
              pruned = []
              if prune_stale_allowed:
                  # Never prune an artifact we are about to restore: the restore loop below honors
                  # `overwrite`, so pruning a staged name would (with overwrite=False) delete it and
                  # then skip the copy, losing the file.
                  for name in sorted(ALLOWED_ARTIFACTS - set(staged)):
                      out = dest / name
                      if out.exists():
                          out.unlink()
                          pruned.append(name)
      
              restored = []
              for name in staged:
                  out = dest / name
                  if out.exists() and not overwrite:
                      continue
                  (tmp / name).replace(out)
                  restored.append(name)
          return {
              "restored": bool(restored),
              "material_id": meta["material_id"],
              "artifacts": restored,
              "pruned_artifacts": pruned,
              "material": meta,
          }
      
    • mimo_qc.py 885 B
      #!/usr/bin/env python3
      """Public API and CLI entrypoint for advisory MiMo multimodal QC."""
      
      import qc.mimo_runner as mimo_qc_runner
      from qc.mimo_evidence import collect_evidence, safe_mimo_config
      from qc.mimo_observations import normalize_observations
      from qc.mimo_payload import build_payload
      from qc.mimo_client import mimo_qc_api_call
      
      clear_report = mimo_qc_runner.clear_report
      
      
      def run(work_dir, **kwargs):
          """mimo_qc_runner.run with the real MiMo transport bound at call time."""
          return mimo_qc_runner.run(work_dir, api_call=mimo_qc_api_call, **kwargs)
      
      
      def main(argv=None):
          return mimo_qc_runner.main(argv, run_callable=run)
      
      __all__ = [
          "build_payload",
          "clear_report",
          "collect_evidence",
          "main",
          "mimo_qc_api_call",
          "normalize_observations",
          "run",
          "safe_mimo_config",
      ]
      
      if __name__ == "__main__":
          raise SystemExit(main())
      
    • project_binding.py 11.8 KB
      """Resolve a ``recap_project.json`` binding into the settings stage skills already accept.
      
      Only video-recap reads the library. A bound subtitle style becomes ``SUBTITLE_*`` env for
      video-assemble, a bound voice becomes the voiceover provider / voice arguments, a bound
      BGM becomes ``BGM_PATH``. Stage skills keep seeing concrete paths and values; they never
      learn about templates. Anything the caller already set explicitly must agree with the
      binding — a silent override would make the delivered file disagree with the project.
      """
      from __future__ import annotations
      
      import json
      import os
      from pathlib import Path
      
      import library as library_lib
      
      PROJECT_FILE = "recap_project.json"
      PROJECT_SCHEMA = "video-recap.project.v1"
      PROJECT_KEYS = {"schema", "name", "library", "bindings", "notes"}
      BINDING_KINDS = ("subtitle_style", "packaging", "voice", "bgm")
      
      
      class BindingError(SystemExit):
          """A project binding that cannot be applied; exits before any stage runs."""
      
          def __init__(self, message):
              super().__init__(f"项目绑定无法使用: {message}")
      
      
      def project_path(value) -> Path:
          path = Path(value).expanduser().resolve()
          if path.is_dir():
              path = path / PROJECT_FILE
          if not path.is_file():
              raise BindingError(f"找不到 {path}")
          return path
      
      
      def _value(params, key):
          node = params.get(key)
          return node.get("value") if isinstance(node, dict) else None
      
      
      def _record(entry):
          return json.loads(Path(entry["record"]).read_text(encoding="utf-8"))
      
      
      def _subtitle_env(template, record, resources):
          params = record["params"]
          canvas = record["canvas"]
          env = {
              "SUBTITLE_PLAY_RES_X": canvas["width"],
              "SUBTITLE_PLAY_RES_Y": canvas["height"],
              "SUBTITLE_FONT_SIZE": _value(params, "size_px"),
              "SUBTITLE_OUTLINE": _value(params, "outline_px"),
              "SUBTITLE_SHADOW": _value(params, "shadow_px"),
              "SUBTITLE_PRIMARY_COLOR": _value(params, "primary_color"),
              "SUBTITLE_OUTLINE_COLOR": _value(params, "outline_color"),
              "SUBTITLE_MAX_CHARS": _value(params, "max_chars"),
              "SUBTITLE_MAX_LINES": _value(params, "max_lines"),
          }
          font = params["font"]
          if font.get("resource"):
              resource = resources[font["resource"]]
              env["SUBTITLE_FONT_NAME"] = (_record(resource).get("font") or {}).get("family")
              env["SUBTITLE_FONT_FILE"] = resource["files"][0]["path"]
          else:
              env["SUBTITLE_FONT_NAME"] = font["family"]
          band = _value(params, "band")
          if band:
              env["SUBTITLE_ALIGNMENT"] = 2
              env["SUBTITLE_MARGIN_V"] = canvas["height"] - band["y_bot"]
          return {k: str(v) for k, v in env.items() if v is not None}
      
      
      def _voice_updates(resource, args):
          record = _record(resource)
          voice = record.get("voice") or {}
          provider = voice.get("provider")
          consent = (record.get("consent") or {}).get("status")
          if resource["files"] and consent == "denied":
              raise BindingError(f"音色 {resource['id']} 的参考音频声音授权为 denied")
          if args.tts_provider not in ("auto", provider):
              raise BindingError(f"音色 {resource['id']} 属于 {provider},与 --tts-provider {args.tts_provider} 冲突")
          updates, env = {"tts_provider": provider}, {}
          if provider == "mimo-tts":
              if resource["files"]:
                  updates["voice_ref"] = resource["files"][0]["path"]
              else:
                  updates["mimo_tts_voice"] = voice.get("voice_id")
          elif provider == "fish-audio":
              env["FISH_TTS_REFERENCE_ID"] = voice.get("voice_id")
          else:
              env["INDEX_TTS_VOICE"] = voice.get("voice_id")
          return updates, env
      
      
      def _same_setting(key, current, value):
          if key.endswith(("_PATH", "_FILE", "VOICE_REF")):
              return Path(current).expanduser().resolve() == Path(value).expanduser().resolve()
          try:
              return float(current) == float(value)
          except ValueError:
              return current == value
      
      
      def _check_env_conflicts(env, environ):
          for key, value in env.items():
              current = environ.get(key, "").strip()
              if current and not _same_setting(key, current, value):
                  raise BindingError(f"{key}={current} 与项目绑定的 {value} 冲突;删掉其中一处")
      
      
      def _check_arg_conflicts(updates, args):
          for key, value in updates.items():
              current = getattr(args, key, None)
              if key == "tts_provider" or current in (None, ""):
                  continue
              if str(current) != str(value):
                  raise BindingError(f"--{key.replace('_', '-')} {current} 与项目绑定的 {value} 冲突")
      
      
      def resolve_project(value, args, environ=None, *, include_voice=True) -> dict:
          """Load and validate a project; return what to apply plus the lock's project block."""
          environ = os.environ if environ is None else environ
          path = project_path(value)
          try:
              project = json.loads(path.read_text(encoding="utf-8"))
          except ValueError as exc:
              raise BindingError(f"{path} 无法解析: {exc}") from None
          if not isinstance(project, dict) or project.get("schema") != PROJECT_SCHEMA:
              raise BindingError(f"{path} 的 schema 必须是 {PROJECT_SCHEMA}")
          unknown = sorted(set(project) - PROJECT_KEYS)
          bindings = project.get("bindings") or {}
          unknown += sorted(f"bindings.{k}" for k in set(bindings) - set(BINDING_KINDS))
          if unknown:
              raise BindingError(f"{path} 有未知字段: {', '.join(unknown)}")
          library_root = (path.parent / str(project.get("library", ""))).resolve()
          if not (library_root / "library.json").is_file():
              raise BindingError(f"库目录无效(缺少 library.json): {library_root}")
          index, report = library_lib.scan_library(library_root)
          resources = {r["id"]: r for r in index["resources"]}
          templates = {f"{t['id']}@v{t['version']}": t for t in index["templates"]}
          errors = {}
          for issue in report.errors:
              errors.setdefault(issue["path"], issue["message"])
      
          def require_valid(entry, label):
              rel = Path(entry["record"]).relative_to(library_root).as_posix()
              if rel in errors:
                  raise BindingError(f"{label}: {errors[rel]}(先运行 library.py check)")
      
          def template(kind):
              ref = bindings.get(kind)
              if not ref:
                  return None
              entry = templates.get(ref)
              if entry is None or entry["kind"] != kind:
                  raise BindingError(f"{kind} 模板 {ref} 不存在或无效(先运行 library.py check)")
              if entry["status"] != "adopted":
                  raise BindingError(f"{kind} 模板 {ref} 的状态是 {entry['status']},只有 adopted 能绑定")
              require_valid(entry, ref)
              for _, rid in entry["refs"]:
                  if rid not in resources:
                      raise BindingError(f"{ref} 引用的资源 {rid} 不存在或无效(先运行 library.py check)")
                  require_valid(resources[rid], f"{ref} 引用的资源 {rid}")
              return entry
      
          def resource(kind):
              rid = bindings.get(kind)
              if not rid:
                  return None
              entry = resources.get(rid)
              if entry is None or entry["kind"] != kind:
                  raise BindingError(f"{kind} 资源 {rid} 不存在或无效(先运行 library.py check)")
              require_valid(entry, f"{kind} 资源 {rid}")
              return entry
      
          env, updates, used_templates = {}, {}, []
          subtitle = template("subtitle_style")
          if subtitle:
              env.update(_subtitle_env(subtitle, _record(subtitle), resources))
              used_templates.append({"role": "subtitle_style", "id": subtitle["id"],
                                     "version": subtitle["version"], "status": subtitle["status"],
                                     "canvas": subtitle["canvas"]})
          packaging = template("packaging")
          if packaging:
              used_templates.append({"role": "packaging", "id": packaging["id"],
                                     "version": packaging["version"], "status": packaging["status"],
                                     "canvas": packaging["canvas"]})
          voice = resource("voice") if include_voice else None
          if voice:
              voice_updates, voice_env = _voice_updates(voice, args)
              updates.update(voice_updates)
              env.update(voice_env)
          bgm = resource("bgm")
          if bgm and getattr(args, "audio_mode", "narration") == "adopted-packet-copy":
              raise BindingError("adopted-packet-copy 冻结原音轨,不能再绑定 BGM;去掉 bindings.bgm 或换声音模式")
          if bgm:
              env["BGM_PATH"] = bgm["files"][0]["path"]
          material_env = environ.get("VIDEO_RECAP_MATERIAL_LIBRARY_DIR", "").strip()
          if (material_env and not getattr(args, "material_library_dir", None)
                  and Path(material_env).expanduser().resolve() != library_root):
              raise BindingError(
                  f"VIDEO_RECAP_MATERIAL_LIBRARY_DIR={material_env} 与项目的库 {library_root} 不是同一个目录;"
                  "素材库与资源库共用根目录,删掉其中一处"
              )
          _check_env_conflicts(env, environ)
          _check_env_conflicts(
              {name: updates[key] for key, name in (("mimo_tts_voice", "MIMO_TTS_VOICE"), ("voice_ref", "VOICE_REF"))
               if key in updates},
              environ,
          )
          _check_arg_conflicts(updates, args)
          return {
              "path": str(path),
              "name": project.get("name", ""),
              "library": str(library_root),
              "env": env,
              "arg_updates": updates,
              "templates": used_templates,
              "resources": [],
              "packaging": packaging and {"record": _record(packaging), "resources": resources},
          }
      
      
      def apply_project(resolved, args, environ=None) -> None:
          """Write the resolved settings where the stage skills will read them."""
          environ = os.environ if environ is None else environ
          for key, value in resolved["env"].items():
              environ[key] = value
          bound = set()
          for key, value in resolved["arg_updates"].items():
              if getattr(args, key, None) != value:
                  bound.add(key)
              setattr(args, key, value)
          if not getattr(args, "material_library_dir", None):
              args.material_library_dir = resolved["library"]
              bound.add("material_library_dir")
          args._bound_from_project = frozenset(bound)
          args.resolved_project = resolved
      
      
      def check_canvas(resolved, width, height) -> None:
          """Templates only apply to the canvas they were calibrated on."""
          for template in resolved["templates"]:
              canvas = template["canvas"]
              if (canvas["width"], canvas["height"]) != (width, height):
                  raise BindingError(
                      f"{template['role']} 模板 {template['id']}@v{template['version']} 按 "
                      f"{canvas['width']}x{canvas['height']} 校准,成片画布是 {width}x{height};换一个同画幅的模板"
                  )
      
      
      PACKAGING_LAYERS = "packaging_layers.json"
      _WRITTEN_BY = "video-recap --project"
      
      
      def sync_packaging_layers(work_dir, resolved) -> None:
          """Write the bound packaging template's layers for video-assemble, or retire our old copy.
      
          A caller-authored ``packaging_layers.json`` (no ``written_by`` marker) is left alone.
          """
          path = Path(work_dir) / PACKAGING_LAYERS
          packaging = (resolved or {}).get("packaging")
          if not packaging:
              if path.exists():
                  try:
                      plan = json.loads(path.read_text(encoding="utf-8"))
                  except ValueError:
                      plan = None
                  ours = isinstance(plan, dict) and plan.get("written_by") == _WRITTEN_BY
                  if ours:
                      path.unlink()
              return
          record, resources = packaging["record"], packaging["resources"]
          layers = [{"name": layer["name"],
                     "path": resources[layer["image"]["resource"]]["files"][0]["path"],
                     "rect": layer["rect"]}
                    for layer in record["params"]["layers"]]
          path.write_text(json.dumps({
              "written_by": _WRITTEN_BY,
              "template": {"id": record["id"], "version": record["version"]},
              "canvas": record["canvas"],
              "layers": layers,
          }, ensure_ascii=False, indent=2), encoding="utf-8")
      
    • qc_contract.py 11.7 KB
      """Shared shift-left QC report contract for video-recap artifacts.
      
      This module is intentionally standalone: it defines a small schema and local
      validation helpers only. It does not call MiMo, read credentials, connect the
      pipeline, or attempt automatic fixes.
      """
      from __future__ import annotations
      
      import re
      from pathlib import Path
      from typing import Any
      from collections.abc import Mapping, Sequence
      from urllib.parse import urlsplit, urlunsplit
      
      
      SCHEMA_VERSION = 1
      
      STAGES = frozenset({
          "pre_cut",
          "post_cut",
          "pre_tts",
          "post_tts",
          "pre_assemble",
          "post_render",
          "golden",
      })
      SEVERITIES = frozenset({"info", "advisory", "warning", "blocker"})
      CONFIDENCES = frozenset({"low", "medium", "high", "objective"})
      SAMPLE_POLICIES = frozenset({"all", "deterministic", "sampled", "semantic", "aesthetic"})
      ARTIFACTS = frozenset({"final_qc.json", "golden_eval.json", "mimo_qc.json", "preflight_qc.json"})
      
      DETERMINISTIC_CATEGORIES = frozenset({
          "missing_artifact",
          "duration",
          "stream",
          "subtitle",
          "tts_placement",
          "schema_invalid",
          "placement",
      })
      NON_DETERMINISTIC_CATEGORIES = frozenset({"semantic", "aesthetic", "mimo_semantic", "mimo_aesthetic"})
      BLOCKING_SEVERITIES = frozenset({"blocker"})
      _SECRET_KEY_RE = re.compile(r"(?i)(api[_-]?key|token|secret|password|authorization|credential)")
      _SECRET_VALUE_RE = re.compile(r"\b(?:sk|tp)-[A-Za-z0-9_-]{6,}\b")
      _URL_SCHEMES_TO_SCRUB = frozenset({"http", "https", "ws", "wss"})
      REDACTED = "<redacted>"
      
      
      def _redact_url(value: str) -> str:
          """Strip URL credentials, query, and fragment while preserving useful host/path."""
          try:
              parts = urlsplit(value)
          except ValueError:
              return value
          if parts.scheme.lower() not in _URL_SCHEMES_TO_SCRUB or not parts.netloc:
              return value
          host = parts.hostname or ""
          if not host:
              return value
          if ":" in host and not host.startswith("["):
              host = f"[{host}]"
          netloc = host
          try:
              port = parts.port
          except ValueError:
              port = None
          if port is not None:
              netloc = f"{netloc}:{port}"
          return urlunsplit((parts.scheme, netloc, parts.path, "", ""))
      
      
      def redact_secrets(value: Any) -> Any:
          """Return a JSON-safe copy with likely credentials removed.
      
          Centralized for all video-recap QC reports: redacts secret-looking keys,
          common synthetic key/token values, and URL userinfo/query/fragment fields.
          """
          if isinstance(value, Mapping):
              redacted: dict[str, Any] = {}
              for key, item in value.items():
                  key_s = str(key)
                  redacted[key_s] = REDACTED if _SECRET_KEY_RE.search(key_s) else redact_secrets(item)
              return redacted
          if isinstance(value, list):
              return [redact_secrets(item) for item in value]
          if isinstance(value, tuple):
              return [redact_secrets(item) for item in value]
          if isinstance(value, Path):
              return str(value)
          if isinstance(value, str):
              return _SECRET_VALUE_RE.sub(REDACTED, _redact_url(value))
          return value
      
      
      _REQUIRED_FINDING_FIELDS = frozenset({
          "finding_id",
          "stage",
          "severity",
          "blocking",
          "deterministic",
          "confidence",
          "rule_id",
          "decision_reason",
          "location",
          "evidence",
          "sample_policy",
          "model_used",
          "next_action",
          # Helper fields kept for current rule semantics and compatibility.
          "category",
          "code",
          "message",
          "source",
          "objective_corroboration",
      })
      _REQUIRED_LOCATION_FIELDS = frozenset({"timecode", "source_span"})
      _REQUIRED_REPORT_FIELDS = frozenset({
          "schema_version",
          "artifact",
          "stage",
          "ok",
          "blocker_count",
          "finding_count",
          "findings",
      })
      
      
      class QCContractError(ValueError):
          """Raised when a QC report or finding violates the shared contract."""
      
      
      def build_finding(
          *,
          stage: str,
          severity: str,
          confidence: str,
          sample_policy: str | Mapping[str, Any],
          category: str,
          code: str,
          message: str,
          finding_id: str | None = None,
          id: str | None = None,
          rule_id: str | None = None,
          decision_reason: str | None = None,
          model_used: str | None = None,
          next_action: str | None = None,
          deterministic: bool | None = None,
          blocking: bool | None = None,
          source: Mapping[str, Any] | None = None,
          location: Mapping[str, Any] | None = None,
          evidence: Mapping[str, Any] | None = None,
          objective_corroboration: Mapping[str, Any] | None = None,
          **extra: Any,
      ) -> dict[str, Any]:
          """Build and validate one normalized QC finding.
      
          Deterministic objective findings may block. MiMo semantic/aesthetic findings
          default to advisory and may never block. Objective checks belong in the
          deterministic QC producers instead of upgrading subjective model findings.
          """
          if finding_id is None:
              finding_id = id
          if finding_id is None:
              raise QCContractError("finding_id is required")
          if rule_id is None:
              rule_id = code
          if decision_reason is None:
              decision_reason = message
          if model_used is None:
              model_used = "local_deterministic"
          if next_action is None:
              next_action = "manual_review"
          if isinstance(sample_policy, str):
              sample_policy = {"type": sample_policy}
      
          if deterministic is None:
              deterministic = category in DETERMINISTIC_CATEGORIES
          if source is None:
              source = {}
          source = redact_secrets(source)
          if location is None:
              location = {"timecode": None, "source_span": None}
          else:
              location = {"timecode": location.get("timecode"), "source_span": location.get("source_span")}
          if evidence is None:
              evidence = {}
          evidence = redact_secrets(evidence)
          objective_corroboration = redact_secrets(objective_corroboration or {})
      
          is_non_deterministic = (not deterministic) or category in NON_DETERMINISTIC_CATEGORIES
          if is_non_deterministic and blocking is None:
              blocking = False
          elif blocking is None:
              blocking = severity in BLOCKING_SEVERITIES
      
          if is_non_deterministic and blocking:
              raise QCContractError(f"non-deterministic blocking finding is not allowed: {category}/{code}")
      
          finding = {
              "finding_id": finding_id,
              "id": finding_id,
              "stage": stage,
              "severity": severity,
              "blocking": bool(blocking),
              "deterministic": bool(deterministic),
              "confidence": confidence,
              "rule_id": rule_id,
              "decision_reason": decision_reason,
              "location": dict(location),
              "evidence": dict(evidence),
              "sample_policy": dict(sample_policy),
              "model_used": model_used,
              "next_action": next_action,
              "category": category,
              "code": code,
              "message": message,
              "source": dict(source),
              "objective_corroboration": dict(objective_corroboration),
          }
          finding.update(redact_secrets(extra))
          _validate_finding(finding)
          return finding
      
      
      def build_report(
          *,
          artifact: str,
          stage: str,
          findings: Sequence[Mapping[str, Any]] | None = None,
          metadata: Mapping[str, Any] | None = None,
      ) -> dict[str, Any]:
          """Build and validate a minimal report for a supported QC artifact."""
          normalized_findings = [redact_secrets(dict(f)) for f in (findings or [])]
          blocker_count = sum(1 for f in normalized_findings if f.get("blocking") is True)
          report = {
              "schema_version": SCHEMA_VERSION,
              "artifact": artifact,
              "stage": stage,
              "ok": blocker_count == 0,
              "blocker_count": blocker_count,
              "finding_count": len(normalized_findings),
              "findings": normalized_findings,
              "metadata": redact_secrets(dict(metadata or {})),
          }
          validate_report(report)
          return report
      
      
      def validate_report(report: Mapping[str, Any]) -> bool:
          """Validate report shape, value domains, and blocking semantics."""
          if not isinstance(report, Mapping):
              raise QCContractError("report must be an object")
          missing = _REQUIRED_REPORT_FIELDS - set(report)
          if missing:
              raise QCContractError(f"report missing required fields: {sorted(missing)}")
          if report["schema_version"] != SCHEMA_VERSION:
              raise QCContractError(f"unsupported schema_version: {report['schema_version']!r}")
          if report["artifact"] not in ARTIFACTS:
              raise QCContractError(f"unsupported artifact: {report['artifact']!r}")
          if report["stage"] not in STAGES:
              raise QCContractError(f"unsupported stage: {report['stage']!r}")
          if not isinstance(report["findings"], list):
              raise QCContractError("findings must be a list")
          for finding in report["findings"]:
              _validate_finding(finding)
          blocker_count = sum(1 for f in report["findings"] if f.get("blocking") is True)
          if report["blocker_count"] != blocker_count:
              raise QCContractError("blocker_count does not match findings")
          if report["finding_count"] != len(report["findings"]):
              raise QCContractError("finding_count does not match findings")
          if report["ok"] != (blocker_count == 0):
              raise QCContractError("ok must be true exactly when blocker_count is zero")
          return True
      
      
      def _validate_finding(finding: Mapping[str, Any]) -> None:
          if not isinstance(finding, Mapping):
              raise QCContractError("finding must be an object")
          missing = _REQUIRED_FINDING_FIELDS - set(finding)
          if missing:
              raise QCContractError(f"finding missing required fields: {sorted(missing)}")
          if finding["stage"] not in STAGES:
              raise QCContractError(f"unsupported finding stage: {finding['stage']!r}")
          if finding["severity"] not in SEVERITIES:
              raise QCContractError(f"unsupported severity: {finding['severity']!r}")
          if finding["confidence"] not in CONFIDENCES:
              raise QCContractError(f"unsupported confidence: {finding['confidence']!r}")
          sample_policy = finding["sample_policy"]
          if not isinstance(sample_policy, Mapping):
              raise QCContractError("sample_policy must be an object")
          sample_policy_type = sample_policy.get("type")
          if sample_policy_type not in SAMPLE_POLICIES:
              raise QCContractError(f"unsupported sample_policy.type: {sample_policy_type!r}")
          if not isinstance(finding["deterministic"], bool):
              raise QCContractError("deterministic must be boolean")
          if not isinstance(finding["blocking"], bool):
              raise QCContractError("blocking must be boolean")
          if finding["blocking"] and finding["severity"] != "blocker":
              raise QCContractError("blocking findings must use severity='blocker'")
          location = finding["location"]
          if not isinstance(location, Mapping):
              raise QCContractError("location must be an object")
          missing_location = _REQUIRED_LOCATION_FIELDS - set(location)
          if missing_location:
              raise QCContractError(f"location missing required fields: {sorted(missing_location)}")
          if finding["source"] is None or not isinstance(finding["source"], Mapping):
              raise QCContractError("source must be an object")
          if finding["evidence"] is None or not isinstance(finding["evidence"], Mapping):
              raise QCContractError("evidence must be an object")
          if finding["objective_corroboration"] is None or not isinstance(finding["objective_corroboration"], Mapping):
              raise QCContractError("objective_corroboration must be an object")
          for field in ("finding_id", "rule_id", "decision_reason", "model_used", "next_action"):
              if not isinstance(finding[field], str) or not finding[field]:
                  raise QCContractError(f"{field} must be a non-empty string")
      
          is_non_deterministic = (not finding["deterministic"]) or finding["category"] in NON_DETERMINISTIC_CATEGORIES
          if is_non_deterministic and finding["blocking"]:
              raise QCContractError(
                  f"non-deterministic blocking finding is not allowed: {finding['category']}/{finding['code']}"
              )
      
    • recap.py 183 B
      #!/usr/bin/env python3
      """CLI entrypoint for the self-contained video-recap orchestrator."""
      
      from recap_runner import main
      
      __all__ = ["main"]
      
      if __name__ == "__main__":
          main()
      
    • recap_cli.py 7.4 KB
      """Define the video-recap command-line contract."""
      
      import argparse
      import os
      
      from lib import env_bool
      from recap_source import AUDIO_MODES
      
      TTS_PROVIDERS = ("auto", "mimo-tts", "fish-audio", "index-tts")
      
      
      class _RecordExplicit:
          """Record the option argparse actually consumed, not the raw argv spelling."""
      
          def __call__(self, parser, namespace, values, option_string=None):
              if option_string is not None:
                  recorded = getattr(namespace, "_explicit_options", None)
                  if recorded is None:
                      recorded = set()
                      setattr(namespace, "_explicit_options", recorded)
                  recorded.add(option_string)
              super().__call__(parser, namespace, values, option_string)
      
      
      def _record_explicit_options(parser):
          """Make every optional action report itself, so guards never re-parse sys.argv."""
          tracked = {}
          for action in parser._actions:
              if not action.option_strings:
                  continue
              base = type(action)
              if not issubclass(base, _RecordExplicit):
                  action.__class__ = tracked.setdefault(
                      base, type(base.__name__, (_RecordExplicit, base), {})
                  )
      
      
      def parse_args(argv=None):
          parser = argparse.ArgumentParser(
              description="Full video recap orchestrator (video-* skill bundle).",
              allow_abbrev=False,
          )
          parser.add_argument("video", nargs="*")
      
          core = parser.add_argument_group("核心流程")
          core.add_argument("--work-dir", default=None)
          core.add_argument("--context", default="")
          core.add_argument("--scene-threshold", type=float, default=None)
          core.add_argument("--style", default="纪录片")
          core.add_argument(
              "--edit-mode",
              default=os.environ.get("EDIT_MODE", "full"),
              choices=["full", "cut", "dub"],
          )
          core.add_argument(
              "--target-duration", default=os.environ.get("TARGET_DURATION") or None
          )
          core.add_argument(
              "--allow-duration-drift",
              action="store_true",
              help="cut mode: accept clip duration drift from --target-duration (primary override)",
          )
          core.add_argument("--output-dir", default=None)
          core.add_argument("--skip-asr", action="store_true")
          core.add_argument("--mimo-video-overview", action="store_true")
          core.add_argument(
              "--consolidate",
              action=argparse.BooleanOptionalAction,
              default=True,
              help="build the understanding story index (Pass B); default ON, --no-consolidate to skip",
          )
          core.add_argument(
              "--consolidate-asr", action="store_true", help="also clean ASR (Pass A)"
          )
      
          voice = parser.add_argument_group("声音策略与配音")
          voice.add_argument("--audio-mode", choices=AUDIO_MODES, default="narration")
          voice.add_argument("--audio-stream-index", type=int, default=0)
          voice.add_argument(
              "--tts-provider",
              default=os.environ.get("TTS_PROVIDER", "auto"),
              choices=TTS_PROVIDERS,
              help="voiceover provider; auto prefers configured MiMo, then Fish Audio; Index is explicit",
          )
          voice.add_argument("--mimo-tts-voice", default=None, help="MiMo TTS voice")
          voice.add_argument(
              "--voice-ref",
              default=None,
              help="reference audio for cloned narration voice (mimo-v2.5-tts-voiceclone)",
          )
          voice.add_argument(
              "--preserve-approved-text",
              action="store_true",
              help="forward strict approved-text preservation to narration voiceover",
          )
          voice.add_argument(
              "--allow-partial-tts",
              action="store_true",
              help="allow video-voiceover to continue when some narration segments fail TTS",
          )
          voice.add_argument(
              "--burn-subtitles",
              action=argparse.BooleanOptionalAction,
              default=None,
              help="burn narration subtitles into the video (default on; --no-burn-subtitles to disable)",
          )
          voice.add_argument(
              "--subtitle-y-top",
              type=int,
              default=None,
              help="inclusive auto-rotated display-frame Y at the top of the measured subtitle band",
          )
          voice.add_argument(
              "--subtitle-y-bot",
              type=int,
              default=None,
              help="exclusive auto-rotated display-frame Y at the bottom of the measured subtitle band",
          )
      
          adoption = parser.add_argument_group("本地采用三件套(assembly-only)")
          adoption.add_argument(
              "--tts-meta", default=None,
              help="local adopted tts_meta.json; requires both adoption flags",
          )
          adoption.add_argument(
              "--narration-adoption", default=None,
              help="local narration_adoption v1; requires --tts-meta and --audio-mix-adoption",
          )
          adoption.add_argument(
              "--audio-mix-adoption", default=None,
              help="local audio_mix_adoption v1 for assembly-only full-sound rendering",
          )
      
          review = parser.add_argument_group("评审、QC 与导出")
          review.add_argument(
              "--review-narration",
              action=argparse.BooleanOptionalAction,
              default=None,
              help="run advisory narration quality review before TTS (default on; fail-open)",
          )
          review.add_argument(
              "--require-narration-review",
              action="store_true",
              help="make narration review a strict pre-TTS gate (also REQUIRE_NARRATION_REVIEW=1)",
          )
          review.add_argument(
              "--mimo-qc",
              default=os.environ.get("MIMO_QC", "off"),
              choices=["off", "pre-assemble", "post-render", "both"],
              help="optional advisory MiMo QC stage(s); never blocks the pipeline",
          )
          review.add_argument(
              "--mimo-qc-refresh",
              action="store_true",
              default=env_bool("MIMO_QC_REFRESH", False),
              help="ignore a matching MiMo QC stage cache",
          )
          review.add_argument(
              "--require-final-qc",
              action="store_true",
              help="full/cut: require literal passing final_qc and golden_eval summaries",
          )
          review.add_argument(
              "--export-jianying",
              action="store_true",
              help="also export an OPTIONAL 剪映/JianYing draft (decoupled; never required)",
          )
          review.add_argument(
              "--jianying-bundle-media",
              action="store_true",
              help="copy media into the 剪映 draft (default on; portable to another machine)",
          )
          review.add_argument(
              "--jianying-no-bundle-media",
              action="store_true",
              help="reference media in place instead of copying it into the draft",
          )
      
          materials = parser.add_argument_group("素材库")
          materials.add_argument(
              "--material-library-dir",
              default=None,
              help="filesystem material library dir (or VIDEO_RECAP_MATERIAL_LIBRARY_DIR)",
          )
          materials.add_argument(
              "--use-materials",
              action=argparse.BooleanOptionalAction,
              default=False,
              help="restore compatible analyzed artifacts from the material library",
          )
          materials.add_argument(
              "--save-materials",
              action="store_true",
              help="save analyzed JSON/MD artifacts into the material library",
          )
      
          materials.add_argument(
              "--project",
              default=None,
              help="recap_project.json(或其所在目录):绑定资源库里已采用的字幕样式、包装模板、音色与 BGM",
          )
      
          selfcheck = parser.add_argument_group("自检")
          selfcheck.add_argument("--doctor", action="store_true")
      
          _record_explicit_options(parser)
          args = parser.parse_args(argv)
          args._explicit_options = frozenset(getattr(args, "_explicit_options", ()))
          return parser, args
      
    • recap_inspect.py 18.3 KB
      #!/usr/bin/env python3
      """video-recap inspect — advisory, read-only orientation over a recap work_dir.
      
      Reads ONLY the JSON artifacts already in work_dir (no ffmpeg, no frame reads, no video
      probing, no new deps, no cross-skill import). A missing artifact is the legitimate
      "that stage has not run yet" state and is reported as such.
      
      Two subcommands:
        state     summarize the work_dir: source video, full|cut mode, which stage
                  artifacts are present vs missing, which file is the NEXT pause the pipeline waits
                  on, stale-manifest risk, and storyboard path(s) if present.
        clip-map  read clip_plan_validated.json and map a queried window between the OUTPUT and
                  SOURCE timelines using the SAME forward affine map cut.py uses
                  (output = clip.output_start + (src - clip.source_start), clamped to the clip),
                  reimplemented locally. Flags cross-clip boundaries and cut-out gaps.
      
      Output: markdown by default; --json for machine-readable; long free text is truncated
      unless --full is passed.
      """
      import argparse
      import json
      from pathlib import Path
      
      
      def load_json(path):
          return json.loads(Path(path).read_text(encoding="utf-8"))
      
      
      # --- artifact catalog --------------------------------------------------------
      # Stage artifacts probed by `state`. Order = rough pipeline order.
      _UNDERSTANDING_ARTIFACTS = [
          "scenes.json",
          "asr_result.json",
          "silence_periods.json",
          "speech_boundary_anchors.json",
          "vlm_analysis.json",
          "understanding_index.json",
          "agent_narration_brief.md",
      ]
      _CUT_ARTIFACTS = [
          "multi_source_manifest.json",
          "clip_plan.json",
          "clip_plan_validated.json",
          "edited_source.mp4",
      ]
      _SCRIPT_ARTIFACTS = [
          "narration.json",
          "narration_lint.json",
          "narration_review.json",
          "original_subtitles.json",
      ]
      _RENDER_ARTIFACTS = [
          "tts_meta.json",
          "timeline.json",
          "subtitles.srt",
          "subtitles.ass",
          "assembly_manifest.json",
      ]
      _STORYBOARD_ARTIFACTS = [
          "storyboard/source_storyboard.json",
          "storyboard/edited_storyboard.json",
      ]
      
      _COMPACT_TEXT_LIMIT = 80
      
      
      def _truncate(text, compact):
          text = str(text or "").strip()
          if not compact or len(text) <= _COMPACT_TEXT_LIMIT:
              return text
          return text[: _COMPACT_TEXT_LIMIT - 1] + "…"
      
      
      def _load_optional(path):
          """The artifact's JSON, or None when the stage that writes it has not run yet.
      
          A present-but-corrupt artifact raises: every file probed here is written by this
          pipeline or a sibling skill under its documented contract."""
          path = Path(path)
          return load_json(path) if path.exists() else None
      
      
      def _fmt_seconds(value):
          """mm:ss.ss for a seconds value."""
          sign = "-" if value < 0 else ""
          total = abs(value)
          return f"{sign}{int(total // 60):02d}:{total % 60:05.2f}"
      
      
      # --- source video discovery --------------------------------------------------
      def _discover_source(work_dir):
          """Find the source video path by reading recap artifacts, most authoritative first.
          Returns {path, origin} with path=None + origin="unknown" when nothing records it."""
          work_dir = Path(work_dir)
          # 1. recap_run_manifest.json — canonical: written before the first pause with both fields.
          #    A multi-source manifest carries `sources` instead of `source_video`.
          data = _load_optional(work_dir / "recap_run_manifest.json")
          if data is not None and data.get("source_video"):
              return {"path": data["source_video"], "origin": "recap_run_manifest.json"}
          # 2. assembly_manifest.json — late stage; carries input_video (+ source_video in cut mode).
          data = _load_optional(work_dir / "assembly_manifest.json")
          if data is not None:
              return {
                  "path": data.get("source_video") or data["input_video"],
                  "origin": "assembly_manifest.json",
              }
          # 3. edited_source.mp4.meta.json — video-cut records sources: {path: {size, mtime_ns}};
          #    a single-source cut has exactly one entry.
          data = _load_optional(work_dir / "edited_source.mp4.meta.json")
          sources = data.get("sources") if isinstance(data, dict) else None
          if isinstance(sources, dict) and len(sources) == 1:
              (path,) = sources
              return {"path": path, "origin": "edited_source.mp4.meta.json"}
          return {"path": None, "origin": "unknown"}
      
      
      def _discover_multi_source(work_dir):
          data = _load_optional(Path(work_dir) / "multi_source_manifest.json")
          if data is None:
              return None
          return {"schema_version": data["schema_version"], "sources": data["sources"]}
      
      
      # --- forward affine map --------------------------------------------------------
      def _source_to_output(src, clip):
          """The forward affine map cut.py uses, clamped to the clip's output span."""
          mapped = clip["output_start"] + (src - clip["source_start"])
          return round(max(clip["output_start"], min(mapped, clip["output_end"])), 3)
      
      
      def _output_to_source(out, clip):
          """Inverse of the forward affine map (same slope, clamped to the clip's source span)."""
          mapped = clip["source_start"] + (out - clip["output_start"])
          return round(max(clip["source_start"], min(mapped, clip["source_end"])), 3)
      
      
      def _overlap(a0, a1, b0, b1):
          """Intersection [lo, hi] of two closed ranges, or None when they do not overlap.
      
          A zero-width query (a0 == a1, e.g. --output-start 10 --output-end 10) is treated as a
          POINT lookup: it returns (a0, a0) when the point lies within [b0, b1], so a point that sits
          inside a clip is reported instead of silently falling through to "not in any clip".
          """
          if a0 == a1:
              return (a0, a0) if b0 <= a0 <= b1 else None
          lo, hi = max(a0, b0), min(a1, b1)
          return (lo, hi) if hi > lo else None
      
      
      # --- state subcommand --------------------------------------------------------
      def _present(work_dir, name):
          return (Path(work_dir) / name).exists()
      
      
      def _detect_mode(work_dir):
          """cut when any cut artifact is present, else full."""
          if any(_present(work_dir, n) for n in _CUT_ARTIFACTS):
              return "cut"
          return "full"
      
      
      def _next_pause(work_dir, mode):
          """The artifact the pipeline is currently waiting on the agent to write, or None when the
          next-needed input is already present. Mirrors recap.py's two-pause cut flow / single-pause full
          flow purely from file presence (advisory; the orchestrator owns the real decision)."""
          if mode == "cut":
              if not _present(work_dir, "clip_plan.json"):
                  return ("clip_plan.json", "pass1: 写剪辑计划(只写 clip_plan.json)")
              if not _present(work_dir, "narration.json"):
                  return ("narration.json", "pass2: 对着剪好的成片用 OUTPUT 时间轴写解说")
              return None
          if not _present(work_dir, "narration.json"):
              return ("narration.json", "写解说 narration.json")
          return None
      
      
      def _stale_manifest_note(work_dir, mode):
          """Advisory stale-manifest risks read purely from file presence. Surfaces the two desync traps recap.py guards: a cut
          narration without the phase ledger that ties it to a clip_plan, and a missing run manifest."""
          notes = []
          if not _present(work_dir, "recap_run_manifest.json"):
              notes.append("缺少 recap_run_manifest.json:无法证明 work_dir 属于当前视频/参数(resume 会被拒绝)。")
          if mode == "cut" and _present(work_dir, "narration.json") and not _present(work_dir, "recap_phase.json"):
              notes.append("有 narration.json 但缺少 recap_phase.json:无法判断解说是否对当前剪辑写的(可能 stale)。")
          return notes
      
      
      def _present_storyboards(work_dir):
          return [n for n in _STORYBOARD_ARTIFACTS if _present(work_dir, n)]
      
      
      def cmd_state(work_dir, compact=None):
          del compact
          work_dir = Path(work_dir)
          if not work_dir.exists():
              return {"error": f"work_dir 不存在: {work_dir}"}
      
          mode = _detect_mode(work_dir)
          groups = {
              "understanding": _UNDERSTANDING_ARTIFACTS,
              "cut": _CUT_ARTIFACTS,
              "script": _SCRIPT_ARTIFACTS,
              "render": _RENDER_ARTIFACTS,
          }
          artifacts = {
              group: {
                  "present": [n for n in names if _present(work_dir, n)],
                  "missing": [n for n in names if not _present(work_dir, n)],
              }
              for group, names in groups.items()
          }
          pause = _next_pause(work_dir, mode)
          return {
              "work_dir": str(work_dir),
              "mode": mode,
              "source_video": _discover_source(work_dir),
              "multi_source": _discover_multi_source(work_dir),
              "next_pause": {"artifact": pause[0], "hint": pause[1]} if pause else None,
              "artifacts": artifacts,
              "storyboards": _present_storyboards(work_dir),
              "stale_manifest_notes": _stale_manifest_note(work_dir, mode),
          }
      
      
      def _render_state_md(state, compact):
          if "error" in state:
              return state["error"]
          lines = [f"# recap work_dir 状态: {state['work_dir']}", ""]
          lines.append(f"模式: **{state['mode']}**")
          src = state["source_video"]
          lines.append(f"源视频: {_truncate(src['path'] or 'unknown', compact)}  (来源 {src['origin']})")
          if state["multi_source"]:
              lines.append("多源素材:")
              for s in state["multi_source"]["sources"]:
                  lines.append(
                      f"  - {s['source_id']}: {_truncate(s['source_path'], compact)} "
                      f"(work_dir {s['source_work_dir']}, material {s['material_id']})"
                  )
          lines.append("")
      
          if state["next_pause"]:
              np = state["next_pause"]
              lines.append(f"下一步暂停 → 等待写入 **{np['artifact']}**  ({np['hint']})")
          else:
              lines.append("下一步暂停 → 无(所需输入已就绪,可继续 voiceover/assemble)")
          lines.append("")
      
          lines.append("## 各阶段产物")
          for group, label in (("understanding", "理解"), ("cut", "剪辑"),
                               ("script", "解说"), ("render", "渲染")):
              info = state["artifacts"][group]
              lines.append(f"### {label}")
              lines.append(f"  有: {', '.join(info['present']) or '(无)'}")
              lines.append(f"  缺: {', '.join(info['missing']) or '(无)'}")
          lines.append("")
      
          if state["storyboards"]:
              lines.append("## storyboard")
              for s in state["storyboards"]:
                  lines.append(f"  - {s}")
              lines.append("")
      
          lines.append("## stale-manifest 风险")
          if state["stale_manifest_notes"]:
              for n in state["stale_manifest_notes"]:
                  lines.append(f"  ⚠️ {n}")
          else:
              lines.append("  无明显 stale 风险。")
          return "\n".join(lines)
      
      
      # --- clip-map subcommand -----------------------------------------------------
      def cmd_clip_map(work_dir, output_start, output_end, source_start, source_end, compact):
          work_dir = Path(work_dir)
          validated = work_dir / "clip_plan_validated.json"
          if not validated.exists():
              return {"error": "clip_plan_validated.json 不存在:这不是 cut 运行,或剪辑计划尚未被 cut.py 校验。"
                               "(full 模式没有源↔输出映射;cut 模式请先跑 video-cut。)"}
          # video-cut writes {"clips": [{clip_id, source_start, source_end, output_start,
          # output_end, duration, reason}]} (cut_contract); source_id/source_path only for
          # multi-source cuts.
          clips = load_json(validated)["clips"]
          if not clips:
              return {"error": "clip_plan_validated.json 没有有效的 clip。"}
      
          have_output = output_start is not None or output_end is not None
          have_source = source_start is not None or source_end is not None
          if not have_output and not have_source:
              return {"error": "请用 --output-start/--output-end 或 --source-start/--source-end 指定要查询的窗口。"}
      
          result = {
              "work_dir": str(work_dir),
              "clip_count": len(clips),
              "source_span": [clips[0]["source_start"], clips[-1]["source_end"]],
              "output_span": [clips[0]["output_start"], clips[-1]["output_end"]],
              "queries": [],
          }
          if have_output:
              result["queries"].append(
                  _map_output_window(output_start, output_end, clips, compact))
          if have_source:
              result["queries"].append(
                  _map_source_window(source_start, source_end, clips, compact))
          return result
      
      
      def _segment(clip, *, source, output, compact):
          # source_id/source_path are only written for multi-source cuts.
          return {
              "clip_id": clip["clip_id"],
              "source_id": clip.get("source_id"),
              "source_path": clip.get("source_path"),
              "output": output,
              "source": source,
              "reason": _truncate(clip["reason"], compact),
          }
      
      
      def _map_output_window(out_start, out_end, clips, compact):
          """Map a queried OUTPUT window to its SOURCE window(s), one per touched clip."""
          out_lo = clips[0]["output_start"] if out_start is None else out_start
          out_hi = clips[-1]["output_end"] if out_end is None else out_end
          if out_hi < out_lo:
              out_lo, out_hi = out_hi, out_lo
          segments = []
          for clip in clips:
              ov = _overlap(out_lo, out_hi, clip["output_start"], clip["output_end"])
              if not ov:
                  continue
              segments.append(_segment(
                  clip,
                  output=[round(ov[0], 3), round(ov[1], 3)],
                  source=[_output_to_source(ov[0], clip), _output_to_source(ov[1], clip)],
                  compact=compact,
              ))
          return {
              "direction": "output→source",
              "query": [round(out_lo, 3), round(out_hi, 3)],
              "clips_touched": [s["clip_id"] for s in segments],
              "cross_clip_boundary": len(segments) > 1,
              "segments": segments,
              # An OUTPUT window can't fall outside any clip (output is contiguous), so no cut-out gaps.
              "cut_out_source_gaps": [],
          }
      
      
      def _map_source_window(src_start, src_end, clips, compact):
          """Map a queried SOURCE window to its OUTPUT window(s), flagging source ranges in no clip."""
          src_lo = clips[0]["source_start"] if src_start is None else src_start
          src_hi = clips[-1]["source_end"] if src_end is None else src_end
          if src_hi < src_lo:
              src_lo, src_hi = src_hi, src_lo
          segments = []
          covered = []
          for clip in clips:
              ov = _overlap(src_lo, src_hi, clip["source_start"], clip["source_end"])
              if not ov:
                  continue
              covered.append(ov)
              segments.append(_segment(
                  clip,
                  source=[round(ov[0], 3), round(ov[1], 3)],
                  output=[_source_to_output(ov[0], clip), _source_to_output(ov[1], clip)],
                  compact=compact,
              ))
          # Cut-out gaps: parts of [src_lo, src_hi] covered by NO kept clip (footage that was cut away).
          gaps = []
          cursor = src_lo
          for lo, hi in sorted(covered):
              if lo > cursor:
                  gaps.append([round(cursor, 3), round(lo, 3)])
              cursor = max(cursor, hi)
          if cursor < src_hi:
              gaps.append([round(cursor, 3), round(src_hi, 3)])
          return {
              "direction": "source→output",
              "query": [round(src_lo, 3), round(src_hi, 3)],
              "clips_touched": [s["clip_id"] for s in segments],
              "cross_clip_boundary": len(segments) > 1,
              "segments": segments,
              "cut_out_source_gaps": gaps,
          }
      
      
      def _render_clip_map_md(result, compact):
          if "error" in result:
              return result["error"]
          lines = [f"# clip-map: {result['work_dir']}", ""]
          lines.append(
              f"{result['clip_count']} clips · "
              f"源 {_fmt_seconds(result['source_span'][0])}–{_fmt_seconds(result['source_span'][1])} · "
              f"输出 {_fmt_seconds(result['output_span'][0])}–{_fmt_seconds(result['output_span'][1])}")
          for q in result["queries"]:
              lines.append("")
              lines.append(f"## {q['direction']}  查询 "
                           f"{_fmt_seconds(q['query'][0])}–{_fmt_seconds(q['query'][1])}")
              if q["cross_clip_boundary"]:
                  lines.append(f"  ⚠️ 跨剪辑边界(涉及 clip {q['clips_touched']})")
              if not q["segments"]:
                  lines.append("  (该窗口不落在任何保留片段内)")
              for s in q["segments"]:
                  src = f"{_fmt_seconds(s['source'][0])}–{_fmt_seconds(s['source'][1])}"
                  out = f"{_fmt_seconds(s['output'][0])}–{_fmt_seconds(s['output'][1])}"
                  reason = f"  〔{s['reason']}〕" if s["reason"] else ""
                  sid = f" {s['source_id']}" if s["source_id"] else ""
                  spath = f" `{_truncate(s['source_path'], compact)}`" if s["source_path"] else ""
                  lines.append(f"  clip {s['clip_id']}{sid}{spath}: 源 {src}  ↔  输出 {out}{reason}")
              if q["cut_out_source_gaps"]:
                  lines.append("  ✂️ 被剪掉的源区间(不在成片里):")
                  for g in q["cut_out_source_gaps"]:
                      lines.append(f"     {_fmt_seconds(g[0])}–{_fmt_seconds(g[1])}")
          return "\n".join(lines)
      
      
      # --- CLI ---------------------------------------------------------------------
      def main(argv=None):
          parser = argparse.ArgumentParser(
              prog="recap_inspect.py",
              description="Advisory read-only inspection of a video-recap work_dir (pure JSON).")
          parser.add_argument("--work-dir", required=True, help="recap work_dir to inspect")
          parser.add_argument("--json", action="store_true", help="machine-readable JSON output")
          parser.add_argument("--full", dest="compact", action="store_false", default=True,
                              help="do not truncate long free text (truncated by default)")
          sub = parser.add_subparsers(dest="command", required=True)
      
          sub.add_parser("state", help="summarize work_dir + the next pause the pipeline is waiting on")
      
          cm = sub.add_parser("clip-map", help="map a window between OUTPUT and SOURCE timelines")
          cm.add_argument("--output-start", type=float, default=None)
          cm.add_argument("--output-end", type=float, default=None)
          cm.add_argument("--source-start", type=float, default=None)
          cm.add_argument("--source-end", type=float, default=None)
      
          args = parser.parse_args(argv)
      
          if args.command == "state":
              result = cmd_state(args.work_dir)
              rendered = _render_state_md(result, args.compact)
          else:
              result = cmd_clip_map(args.work_dir, args.output_start, args.output_end,
                                    args.source_start, args.source_end, args.compact)
              rendered = _render_clip_map_md(result, args.compact)
      
          if args.json:
              print(json.dumps(result, ensure_ascii=False, indent=2))
          else:
              print(rendered)
          return 0
      
      
      if __name__ == "__main__":
          raise SystemExit(main())
      
    • recap_review.py 2.8 KB
      """Apply the optional or strict narration-review gate before TTS."""
      
      import json
      from pathlib import Path
      
      from lib import env_bool
      
      
      def review_narration_enabled(args):
          if args.review_narration is not None:
              return args.review_narration
          return env_bool("REVIEW_NARRATION", True)
      
      
      def require_narration_review(args):
          return args.require_narration_review or env_bool("REQUIRE_NARRATION_REVIEW", False)
      
      
      def review_result_status(work_dir):
          """Gate on the review video-script wrote (findings: list of {severity in
          error|warning|suggestion}, already normalised by review_response)."""
          path = Path(work_dir) / "narration_review.json"
          if not path.exists():
              return {"ok": False, "reason": "missing narration_review.json"}
          data = json.loads(path.read_text(encoding="utf-8"))
          error_count = sum(1 for finding in data["findings"] if finding["severity"] == "error")
          if data.get("parse_error"):
              return {"ok": False, "reason": "parse_error", "review": data, "errors": error_count}
          if error_count:
              return {
                  "ok": False,
                  "reason": f"error {error_count}",
                  "review": data,
                  "errors": error_count,
              }
          # Strict mode gates only on parse errors or factual error findings. The model's
          # holistic verdict stays advisory because craft-class findings are warnings.
          return {"ok": True, "reason": "ok", "review": data, "errors": error_count}
      
      
      def clear_narration_review_artifacts(work_dir):
          """Remove prior review artifacts before a fresh pre-TTS review run."""
          for name in ("narration_review.json", "narration_review.md"):
              (Path(work_dir) / name).unlink(missing_ok=True)
      
      
      def run_narration_review(work_dir, args, *, run, timeline="source"):
          """Run the review gate using the caller's skill-command runner."""
          strict = require_narration_review(args)
          if not review_narration_enabled(args) and not strict:
              return False
          try:
              clear_narration_review_artifacts(work_dir)
              # Pin the grounding timeline so stale cut artifacts cannot alter auto-detection.
              review_args = ["--work-dir", work_dir, "--timeline", timeline]
              if strict and timeline == "cut_output":
                  review_args.append("--strict-evidence")
              run("video-script", "review.py", *review_args)
          except SystemExit as exc:
              if strict:
                  raise SystemExit(f"严格解说评审失败,已阻止 TTS: {exc}") from exc
              print(f"[video-recap] ⚠️ 建议性评审失败,继续执行 TTS: {exc}", flush=True)
              return False
      
          status = review_result_status(work_dir)
          if strict and not status["ok"]:
              raise SystemExit(f"严格解说评审未通过,已阻止 TTS: {status['reason']}")
          if strict:
              print("[video-recap] ✅ 严格解说评审通过,继续 TTS", flush=True)
          return True
      
    • recap_runner.py 28 KB
      """Orchestrate single- and multi-source recap pipelines."""
      
      import json
      import os
      from pathlib import Path
      
      import materials as material_lib
      import project_binding
      import resource_lock
      
      from recap_cli import TTS_PROVIDERS, parse_args
      from recap_review import run_narration_review
      from recap_runtime import (
          _analysis_settings,
          _build_multi_source_records,
          _coerce_videos,
          _entry,
          _optional_env_int,
          _preflight_burn_subtitles,
          _probe_display_height_or_raise,
          _probe_display_size_or_raise,
          _read_video_duration_or_raise,
          _run,
          _write_multi_source_manifest,
          _write_project_run_manifest,
          _write_run_manifest,
      )
      from recap_stage_qc import (
          _post_render_qc_metadata,
          _prepare_mimo_qc,
          _print_final_qc_pointer,
          _require_final_qc,
          _run_mimo_qc_stage,
          _tts_qc_metadata,
          _write_final_qc_reports,
          _write_shift_left_stage_qc,
      )
      from recap_source import (
          begin_local_adoption_qc,
          begin_non_narration_qc,
          extend_assemble_args,
          load_local_assembly_evidence,
          needs_voiceover,
          owned_local_delivery,
          reject_unbound_narration_workdir,
          reject_unsupported_subtitle_track,
          uses_local_adoption,
          uses_narration,
          validate_audio_routing,
          verify_local_assembly_evidence,
          remove_owned_local_delivery,
      )
      from recap_timeline import (
          _continuation_command,
          _cut_narration_is_stale,
          _manifest_mismatches,
          _material_library_dir,
          _materials_enabled,
          _multi_manifest_mismatches,
          _pause_for_agent,
          _print_narration_review_pointer,
          _read_assembly_output,
          _read_phase_ledger,
          _save_materials_enabled,
          _source_work_dir,
          _surface_cut_qc,
          _understand_args_for_source,
          _write_canonical_visual_overlays,
          _write_multi_source_clip_brief,
          _write_multi_source_output_brief,
          _write_phase_ledger,
      )
      
      RUN_MANIFEST = "recap_run_manifest.json"
      
      
      def _voiceover_args(work_dir, narration_path, args):
          """Build the shared full/cut invocation for video-voiceover."""
          result = ["--work-dir", str(work_dir), "--narration", str(narration_path)]
          if args.tts_provider != "auto":
              result += ["--tts-provider", args.tts_provider]
          if args.mimo_tts_voice:
              result += ["--mimo-voice", args.mimo_tts_voice]
          if args.voice_ref:
              result += ["--voice-ref", args.voice_ref]
          if args.allow_partial_tts:
              result.append("--allow-partial-tts")
          if args.preserve_approved_text:
              result.append("--preserve-approved-text")
          return result
      
      
      def _approved_validation_args(args):
          return ["--preserve-approved-text"] if args.preserve_approved_text else []
      
      
      def _record_resources(work_dir, args):
          """Write resource_lock.json for the finished render and surface anything needing a person."""
          library_dir = getattr(args, "material_library_dir", None) or os.environ.get(
              "VIDEO_RECAP_MATERIAL_LIBRARY_DIR"
          )
          try:
              lock = resource_lock.write_resource_lock(
                  work_dir, library_dir=library_dir, project=getattr(args, "resolved_project", None)
              )
          except (OSError, ValueError, TypeError, KeyError) as exc:  # a record must never fail a finished render
              print(f"[video-recap] ⚠ 未能写出 resource_lock.json: {type(exc).__name__}: {exc}", flush=True)
              return
          print(resource_lock.summary_line(lock), flush=True)
      
      
      def _finish_recap(work_dir, final_output, args):
          final_qc_result = _write_final_qc_reports(work_dir, final_output)
          if args.require_final_qc:
              _require_final_qc(final_qc_result, work_dir)
          print(f"[video-recap] ✅ 完成: {final_output}")
          _print_final_qc_pointer(final_qc_result)
      
      
      def _run_local_adoption(video, work_dir, args):
          """Run the strict local bundle directly through the authoritative assembler."""
          work_dir.mkdir(parents=True, exist_ok=False)
          manifest = _write_run_manifest(work_dir, video, args)
          def record_failure(exc):
              failed = {**manifest, "local_adoption_run": {
                  "status": "FAILED", "error_type": type(exc).__name__,
              }}
              (work_dir / RUN_MANIFEST).write_text(
                  json.dumps(failed, ensure_ascii=False, indent=2), encoding="utf-8"
              )
          begin_local_adoption_qc(work_dir, _write_shift_left_stage_qc)
          assemble_args = [
              str(video),
              "--work-dir", str(work_dir),
              "--recap-stem", video.stem,
              "--output-dir", args.output_dir,
              "--tts-meta", args.tts_meta,
              "--narration-adoption", args.narration_adoption,
              "--audio-mix-adoption", args.audio_mix_adoption,
          ]
          if args.burn_subtitles is not None:
              assemble_args.append(
                  "--burn-subtitles" if args.burn_subtitles else "--no-burn-subtitles"
              )
          if args.subtitle_y_top is not None:
              assemble_args += [
                  "--subtitle-y-top", str(args.subtitle_y_top),
                  "--subtitle-y-bot", str(args.subtitle_y_bot),
              ]
          owned_delivery = None
          try:
              _run("video-assemble", "assemble.py", *assemble_args)
              evidence = load_local_assembly_evidence(work_dir)
              owned_delivery = owned_local_delivery(evidence)
              verify_local_assembly_evidence(evidence, manifest)
          except BaseException as exc:
              remove_owned_local_delivery(owned_delivery)
              record_failure(exc)
              raise
          final_output = _read_assembly_output(work_dir)
          _record_resources(work_dir, args)
          _write_shift_left_stage_qc(
              work_dir,
              "post_render",
              metadata=_post_render_qc_metadata(work_dir, final_output),
          )
          _finish_recap(work_dir, final_output, args)
      
      
      def _run_or_restore_understanding(source_record, source_work_dir, args):
          """Run video-understanding for one source, or restore it from the material library."""
          source_work_dir = Path(source_work_dir)
          source_work_dir.mkdir(parents=True, exist_ok=True)
          source_path = source_record["source_path"]
          source_identity = source_record["source_video_identity"]
          settings = source_record["settings"]
          lib_dir = _material_library_dir(args)
          restored = False
          if _materials_enabled(args):
              result = material_lib.restore_material(
                  lib_dir,
                  source_work_dir,
                  source_path=source_path,
                  source_identity=source_identity,
                  settings=settings,
              )
              restored = result["restored"]
              if restored:
                  print(
                      f"[video-recap] ♻️  复用素材库: {result['material_id']} → {source_work_dir}",
                      flush=True,
                  )
              else:
                  print(
                      f"[video-recap] 素材库未复用 {source_record['source_name']}: {result['reason']}",
                      flush=True,
                  )
          if not restored:
              _run(
                  "video-understanding",
                  "understand.py",
                  *_understand_args_for_source(source_record, source_work_dir, args),
              )
          _write_run_manifest(source_work_dir, source_record["source_path"], args)
          if _save_materials_enabled(args):
              meta = material_lib.save_material(
                  lib_dir,
                  source_work_dir,
                  source_path,
                  source_identity,
                  settings,
                  source_id=source_record["source_id"],
                  material_id=source_record["material_id"],
              )
              source_record["material_id"] = meta["material_id"]
              print(
                  f"[video-recap] 💾 已沉淀素材: {meta['material_id']} → {lib_dir}",
                  flush=True,
              )
          return restored
      
      
      def _single_source_record(video, args):
          """Identity + settings for the one source of a single-video run."""
          identity = material_lib.file_identity(video)
          return {
              "source_id": material_lib.source_id_for(video),
              "source_path": str(video),
              "source_name": video.name,
              "source_video_identity": identity,
              "settings": _analysis_settings(args),
              "material_id": material_lib.material_id_for(video, identity),
          }
      
      
      def _rebuild_understanding_brief(source_record, source_work_dir, args):
          """Rebuild agent_narration_brief.md from cached/restored analysis only.
          Cut pass 2 needs an OUTPUT-time brief after edited_source.mp4 exists. A
          material restore may have supplied pass-1 analysis artifacts (and even a
          source-time brief), but it must not skip this phase-specific brief rebuild.
          """
          _run(
              "video-understanding",
              "understand.py",
              *_understand_args_for_source(source_record, source_work_dir, args),
              "--brief-only",
          )
      
      
      def _reject_stale(mismatches, label):
          if mismatches:
              details = "\n  - ".join(mismatches)
              raise SystemExit(
                  f"work_dir 与当前{label}recap 输入不匹配,拒绝复用既有 narration/clip_plan;"
                  "请使用新的 --work-dir,或删除旧产物后重新运行 Phase A。\n"
                  f"  - {details}"
              )
      
      
      def _reject_stale_multi_manifest(work_dir, videos, args, source_records):
          _reject_stale(_multi_manifest_mismatches(work_dir, videos, args, source_records), "多视频 ")
      
      
      def _render_cut(video_arg, work_dir, args, *extra):
          """Render edited_source.mp4 from clip_plan.json and record the post-cut QC stage."""
          crender = [str(video_arg), "--work-dir", str(work_dir), *extra]
          if args.target_duration:
              crender += ["--target-duration", args.target_duration]
          if args.allow_duration_drift:
              crender.append("--allow-duration-drift")
          _run("video-cut", "cut.py", *crender)
          cut_qc = _surface_cut_qc(work_dir)
          _write_shift_left_stage_qc(work_dir, "post_cut", metadata={"cut_qc": cut_qc})
      
      
      def _reject_stale_cut_narration(work_dir, clip_plan_identity):
          if _cut_narration_is_stale(_read_phase_ledger(work_dir), clip_plan_identity):
              raise SystemExit(
                  "clip_plan.json 已改变,但 narration.json 仍是对旧剪辑写的,会与剪后画面对不上。"
                  "请删除 narration.json,重跑后按新成片重新写解说。"
              )
      
      
      def _validate_cut_output_narration(work_dir, args):
          output_duration = _read_video_duration_or_raise(work_dir / "edited_source.mp4")
          _run(
              "video-script", "validate.py", "--work-dir", work_dir,
              "--mode", "cut_output", "--output-duration", f"{output_duration:.3f}",
              *_approved_validation_args(args),
          )
      
      
      def _narrate(work_dir, args, timeline):
          """Review -> TTS -> visual overlays for a validated narration.json."""
          narration_json = work_dir / "narration.json"
          review_ran = run_narration_review(work_dir, args, run=_run, timeline=timeline)
          _write_shift_left_stage_qc(
              work_dir, "pre_tts", metadata={"review_ran": review_ran, "timeline": timeline}
          )
          _run(
              "video-voiceover", "voiceover.py",
              *_voiceover_args(work_dir, narration_json, args),
          )
          _write_shift_left_stage_qc(work_dir, "post_tts", metadata=_tts_qc_metadata(work_dir))
          overlays_path = _write_canonical_visual_overlays(work_dir, narration_json)
          _write_shift_left_stage_qc(
              work_dir, "pre_assemble", metadata={"visual_overlays": str(overlays_path)}
          )
          _run_mimo_qc_stage(work_dir, args, "pre_assemble")
          return review_ran
      
      
      def _deliver(work_dir, args, assemble_video, recap_stem, timeline, extra_assemble_args=()):
          """Shared tail of every full/cut run: narration (if owned) -> assemble -> final QC."""
          project = getattr(args, "resolved_project", None)
          if project and project["templates"]:
              project_binding.check_canvas(project, *_probe_display_size_or_raise(assemble_video))
          project_binding.sync_packaging_layers(work_dir, project)
          review_ran = _narrate(work_dir, args, timeline) if uses_narration(args) else None
          aargs = [str(assemble_video), "--work-dir", str(work_dir), "--recap-stem", recap_stem]
          extend_assemble_args(aargs, args)
          if args.output_dir:
              aargs += ["--output-dir", args.output_dir]
          if args.burn_subtitles is not None:
              aargs.append(
                  "--burn-subtitles" if args.burn_subtitles else "--no-burn-subtitles"
              )
          if args.subtitle_y_top is not None:
              aargs += [
                  "--subtitle-y-top", str(args.subtitle_y_top),
                  "--subtitle-y-bot", str(args.subtitle_y_bot),
              ]
          # env-only burn intent (BURN_SUBTITLES) is propagated implicitly: assemble re-derives it
          # via the same env_bool default the preflight used, so the two agree by shared env.
          aargs += extra_assemble_args
          if args.export_jianying:  # env EXPORT_JIANYING is honored by assemble.py itself
              aargs.append("--export-jianying")
          if args.jianying_bundle_media:
              aargs.append("--jianying-bundle-media")
          if args.jianying_no_bundle_media:
              aargs.append("--jianying-no-bundle-media")
          _run("video-assemble", "assemble.py", *aargs)
      
          final_output = _read_assembly_output(work_dir)
          _record_resources(work_dir, args)
          _write_shift_left_stage_qc(
              work_dir,
              "post_render",
              metadata=_post_render_qc_metadata(work_dir, final_output),
          )
          if uses_narration(args):
              _run_mimo_qc_stage(work_dir, args, "post_render", final_output=final_output)
          _finish_recap(work_dir, final_output, args)
          if uses_narration(args):
              _print_narration_review_pointer(work_dir, review_ran=review_ran)
      
      
      def _run_multi_cut(videos, work_dir, args):
          """Multi-video MVP: cut-first/narrate-second over a project work_dir."""
          videos = _coerce_videos(videos)
          work_dir = Path(work_dir)
          work_dir.mkdir(parents=True, exist_ok=True)
          reject_unbound_narration_workdir(work_dir, args)
          reject_unsupported_subtitle_track(work_dir, args)
          source_records = _build_multi_source_records(videos, args)
          narration_json = work_dir / "narration.json"
          clip_plan_json = work_dir / "clip_plan.json"
          edited_source = work_dir / "edited_source.mp4"
          inspect_py = _entry("video-recap", "recap_inspect.py")
          # If a project manifest already exists, it must match before any Phase-B reuse.
          if (work_dir / RUN_MANIFEST).exists():
              _reject_stale_multi_manifest(work_dir, videos, args, source_records)
          manifest_path = _write_multi_source_manifest(work_dir, source_records)
          if not uses_narration(args):
              begin_non_narration_qc(work_dir, args, _write_shift_left_stage_qc)
          if not clip_plan_json.exists():
              for record in source_records:
                  _run_or_restore_understanding(
                      record, _source_work_dir(work_dir, record), args
                  )
              # Rewrite: _run_or_restore_understanding fills in material_id per record.
              manifest_path = _write_multi_source_manifest(work_dir, source_records)
              _write_project_run_manifest(work_dir, videos, args, source_records)
              _write_multi_source_clip_brief(work_dir, source_records, args)
              _pause_for_agent(
                  work_dir,
                  f"{clip_plan_json}(多视频剪辑计划;每个 clip 必须带 source_id)",
                  _continuation_command(videos, work_dir, args),
                  inspect_hint=f"python3 {inspect_py} --work-dir {work_dir} state",
              )
              return
      
          _reject_stale_multi_manifest(work_dir, videos, args, source_records)
          cp_identity = material_lib.file_identity(clip_plan_json)
          _render_cut(videos[0], work_dir, args, "--sources-manifest", str(manifest_path))
          if uses_narration(args):
              if not narration_json.exists():
                  _write_multi_source_output_brief(
                      work_dir, source_records, work_dir / "clip_plan_validated.json"
                  )
                  _write_phase_ledger(
                      work_dir,
                      clip_plan_identity=cp_identity,
                      edited_source_rendered=True,
                      multi_source=True,
                  )
                  _pause_for_agent(
                      work_dir,
                      f"{narration_json}(用成片 OUTPUT 时间轴写解说,对着 {edited_source})",
                      _continuation_command(videos, work_dir, args),
                      inspect_hint=(
                          f"python3 {inspect_py} --work-dir {work_dir} "
                          "clip-map --output-start <s> --output-end <e>"
                      ),
                  )
                  return
              _reject_stale_cut_narration(work_dir, cp_identity)
              _write_phase_ledger(
                  work_dir,
                  clip_plan_identity=cp_identity,
                  narration_written=True,
                  multi_source=True,
              )
              _validate_cut_output_narration(work_dir, args)
      
          _deliver(work_dir, args, edited_source, f"multi_{videos[0].stem}", "cut_output")
      
      
      def main():
          ap, args = parse_args()
      
          if args.require_final_qc and args.edit_mode == "dub":
              ap.error("--require-final-qc is only supported in full/cut modes, not dub")
          # argparse `choices` does not cover the TTS_PROVIDER env default; source-owned audio and
          # local adoption never synthesise, so an ambient provider is irrelevant there.
          if (args.doctor or needs_voiceover(args)) and args.tts_provider not in TTS_PROVIDERS:
              ap.error("TTS_PROVIDER/--tts-provider must be one of: " + ", ".join(TTS_PROVIDERS))
      
          if args.doctor:
              if any(
                  getattr(args, field) is not None
                  for field in ("tts_meta", "narration_adoption", "audio_mix_adoption")
              ):
                  ap.error("--doctor cannot be combined with local adoption inputs")
              doctor_args = []
              if args.tts_provider != "auto":
                  doctor_args += ["--tts-provider", args.tts_provider]
              _run("video-recap", "doctor.py", *doctor_args)
              return
          if not uses_local_adoption(args) and args.mimo_qc not in {
              "off", "pre-assemble", "post-render", "both"
          }:
              ap.error(
                  "MIMO_QC/--mimo-qc must be one of: off, pre-assemble, post-render, both"
              )
          if not args.video:
              ap.error("video is required (unless --doctor)")
          validate_audio_routing(ap, args)
          args.resolved_project = None
          if args.project:
              args.project = str(project_binding.project_path(args.project))
              if uses_local_adoption(args) or args.edit_mode == "dub":
                  ap.error("--project applies to full/cut runs, not dub or local adoption bundles")
              project_binding.apply_project(
                  project_binding.resolve_project(
                      args.project, args, include_voice=needs_voiceover(args)
                  ),
                  args,
              )
      
          if needs_voiceover(args) and args.voice_ref is None:
              args.voice_ref = os.environ.get("VOICE_REF", "").strip() or None
          elif not needs_voiceover(args):
              # Source-owned audio must not inherit ambient narration configuration. Explicit
              # narration flags were already rejected by validate_audio_routing().
              args.voice_ref = None
          try:
              if args.subtitle_y_top is None:
                  args.subtitle_y_top = _optional_env_int("SUBTITLE_Y_TOP")
              if args.subtitle_y_bot is None:
                  args.subtitle_y_bot = _optional_env_int("SUBTITLE_Y_BOT")
          except ValueError as exc:
              ap.error(str(exc))
      
          if needs_voiceover(args):
              explicit_mimo_voice = (
                  args.mimo_tts_voice or os.environ.get("MIMO_TTS_VOICE", "").strip()
              )
              if explicit_mimo_voice and args.voice_ref:
                  ap.error("--mimo-tts-voice and --voice-ref are mutually exclusive")
              if args.tts_provider in {"fish-audio", "index-tts"} and (
                  explicit_mimo_voice or args.voice_ref
              ):
                  ap.error(
                      "--mimo-tts-voice/--voice-ref are only supported by --tts-provider mimo-tts"
                  )
          if args.edit_mode == "dub" and args.voice_ref:
              ap.error(
                  "--voice-ref is only supported in full/cut modes; dub clones the source voice automatically"
              )
          if args.edit_mode == "dub" and args.tts_provider == "fish-audio":
              ap.error(
                  "--tts-provider fish-audio is only supported in full/cut modes; "
                  "dub uses MiMo voice cloning"
              )
          if args.edit_mode == "dub" and args.subtitle_y_top is not None:
              ap.error(
                  "--subtitle-y-top/--subtitle-y-bot are only supported in full/cut modes"
              )
          if args.edit_mode == "dub" and args.burn_subtitles:
              ap.error("--burn-subtitles is only supported in full/cut modes; dub never burns subtitles")
          if (args.subtitle_y_top is None) != (args.subtitle_y_bot is None):
              ap.error("--subtitle-y-top and --subtitle-y-bot must be provided together")
          if args.subtitle_y_top is not None:
              if args.subtitle_y_top < 0 or args.subtitle_y_bot <= args.subtitle_y_top:
                  ap.error("subtitle Y coordinates must satisfy 0 <= top < bot")
      
          videos = _coerce_videos(args.video)
          if len(videos) > 1 and args.edit_mode != "cut":
              raise SystemExit(
                  "多视频输入当前 MVP 只支持 --edit-mode cut;full/dub 请一次输入一个视频。"
              )
          if len(videos) > 1 and args.subtitle_y_top is not None:
              ap.error("多视频 cut 暂不支持全局 subtitle Y 坐标;各源字幕带可能不同")
          if args.subtitle_y_top is not None:
              canvas_height = _probe_display_height_or_raise(
                  videos[0], require_square_pixels=True
              )
              if args.subtitle_y_bot > canvas_height:
                  ap.error(
                      f"subtitle Y coordinates exceed display canvas height {canvas_height}: "
                      f"bot={args.subtitle_y_bot}"
                  )
          if args.voice_ref:
              voice_ref = Path(args.voice_ref).expanduser().resolve()
              if not voice_ref.is_file():
                  ap.error(f"reference audio does not exist or is not a file: {voice_ref}")
              args.voice_ref = str(voice_ref)
      
          _execute_pipeline(args, videos)
      
      
      def _execute_pipeline(args, videos):
          # Fail fast before any expensive understanding/VLM/ASR/TTS work if the run will burn
          # subtitles but this ffmpeg can't (otherwise it only blows up at the final render).
          _preflight_burn_subtitles(args)
      
          if uses_local_adoption(args):
              _run_local_adoption(videos[0], Path(args.work_dir), args)
              return
      
          if len(videos) > 1:
              work_dir = (
                  Path(args.work_dir).resolve()
                  if args.work_dir
                  else videos[0].parent / f"work_dir_multi_{videos[0].stem}"
              )
              # Validate the work-directory audio policy before any run-local QC state is reset.
              reject_unbound_narration_workdir(work_dir, args)
              _prepare_mimo_qc(work_dir, args)
              _run_multi_cut(videos, work_dir, args)
              return
      
          video = videos[0]
          work_dir = (
              Path(args.work_dir).resolve()
              if args.work_dir
              else video.parent / f"work_dir_{video.stem}"
          )
          work_dir.mkdir(parents=True, exist_ok=True)
          reject_unbound_narration_workdir(work_dir, args)
          reject_unsupported_subtitle_track(work_dir, args)
          _prepare_mimo_qc(work_dir, args)
          if args.edit_mode == "dub":
              _run_dub(video, work_dir, args)
          else:
              _run_single(video, work_dir, args)
      
      
      def _run_dub(video, work_dir, args):
          """EN->ZH translation-dub in the original cloned voice (replaces speech, not overlay).
          One pause: prepare (ASR + sentence-seg + reference) -> agent writes the Chinese
          translation (dub_script.json) -> render (clone TTS + full-replace mux)."""
          dub_script = work_dir / "dub_script.json"
          dub_args = ["--video", str(video), "--work-dir", str(work_dir)]
          if not dub_script.exists():
              _run("video-voiceover", "dub.py", "--stage", "prepare", *dub_args)
              _write_run_manifest(work_dir, video, args)
              cont = _continuation_command(video, work_dir, args)
              print("=" * 50)
              print(
                  f"[video-recap] ⏸  阅读 {work_dir / 'dub_brief.md'},把英文原声转写切分并翻译成中文,写入 {dub_script}"
              )
              print(
                  '[video-recap]    格式 [{"start": 起秒, "end": 止秒, "zh": "译文"}](按 start 升序);逐句忠实、跟原声节奏一致、保留原音色'
              )
              print(f"[video-recap]    写完后重跑继续: {cont}")
              print("=" * 50)
              return
          _reject_stale(_manifest_mismatches(work_dir, video, args), " ")
          _run("video-voiceover", "dub.py", "--stage", "render", *dub_args)
          print(f"[video-recap] ✅ 配音完成: {work_dir / ('dub_' + video.stem + '.mp4')}")
      
      
      def _run_single(video, work_dir, args):
          """Single-video full or cut run; returns early at each agent pause."""
          cut = args.edit_mode == "cut"
          narration_json = work_dir / "narration.json"
          clip_plan_json = work_dir / "clip_plan.json"
          edited_source = work_dir / "edited_source.mp4"
          inspect_py = _entry("video-recap", "recap_inspect.py")
          state_hint = f"python3 {inspect_py} --work-dir {work_dir} state"
      
          def _understand():
              _run_or_restore_understanding(_single_source_record(video, args), work_dir, args)
      
          def _pause(need_text, inspect_hint):
              # Same banner as the multi-source flow; only the continuation command differs.
              _pause_for_agent(
                  work_dir,
                  need_text,
                  _continuation_command(video, work_dir, args),
                  inspect_hint=inspect_hint,
              )
      
          def _reject_stale_manifest():
              _reject_stale(_manifest_mismatches(work_dir, video, args), " ")
      
          if not cut:
              # Full mode: a single pause (understand -> agent writes narration.json -> produce).
              if uses_narration(args):
                  if not narration_json.exists():
                      _understand()
                      _write_run_manifest(work_dir, video, args)
                      _pause(f"{narration_json}", state_hint)
                      return
                  _reject_stale_manifest()
                  _run(
                      "video-script", "validate.py", "--work-dir", work_dir,
                      "--mode", "full", *_approved_validation_args(args),
                  )
              elif (work_dir / RUN_MANIFEST).exists():
                  _reject_stale_manifest()
              else:
                  _write_run_manifest(work_dir, video, args)
              if not uses_narration(args):
                  begin_non_narration_qc(work_dir, args, _write_shift_left_stage_qc)
              _deliver(work_dir, args, video, video.stem, "source")
              return
      
          # Cut mode: cut-first / narrate-second (two pauses), so narration is authored against the
          # REAL output timeline — no source-time mapping exists that could drop/clamp/desync.
          if not clip_plan_json.exists():
              # PASS 1: understand -> agent writes clip_plan.json ONLY.
              if not uses_narration(args) and (work_dir / RUN_MANIFEST).exists():
                  _reject_stale_manifest()
              if not uses_narration(args):
                  begin_non_narration_qc(work_dir, args, _write_shift_left_stage_qc)
              _understand()
              _write_run_manifest(work_dir, video, args)
              _pause(f"{clip_plan_json}(只写剪辑计划;解说下一步对着剪好的成片写)", state_hint)
              return
          _reject_stale_manifest()
          if not uses_narration(args):
              begin_non_narration_qc(work_dir, args, _write_shift_left_stage_qc)
          cp_identity = material_lib.file_identity(clip_plan_json)
          _render_cut(video, work_dir, args)
          if uses_narration(args):
              if not narration_json.exists():
                  # PASS 2: rebuild the brief (now an OUTPUT-timeline variant) and pause for narration.
                  _rebuild_understanding_brief(_single_source_record(video, args), work_dir, args)
                  _write_phase_ledger(
                      work_dir, clip_plan_identity=cp_identity, edited_source_rendered=True
                  )
                  _pause(
                      f"{narration_json}(用成片 OUTPUT 时间轴写解说,对着 {edited_source})",
                      f"python3 {inspect_py} --work-dir {work_dir} "
                      "clip-map --output-start <s> --output-end <e>  # 核对输出↔原片时间轴",
                  )
                  return
              _reject_stale_cut_narration(work_dir, cp_identity)
              _write_phase_ledger(work_dir, clip_plan_identity=cp_identity, narration_written=True)
              _validate_cut_output_narration(work_dir, args)
          else:
              _write_phase_ledger(
                  work_dir,
                  clip_plan_identity=cp_identity,
                  edited_source_rendered=True,
                  audio_mode=args.audio_mode,
                  audio_stream_index=args.audio_stream_index,
              )
          # let the timeline / 剪映 export reference the original clips, not edited_source.mp4
          _deliver(
              work_dir, args, edited_source, video.stem, "cut_output",
              ["--source-video", str(video)],
          )
      
    • recap_runtime.py 9.3 KB
      """Provide subprocess, manifest, and preflight runtime helpers."""
      
      import json
      import math
      import shlex
      import shutil
      import subprocess
      import sys
      from pathlib import Path
      
      import materials as material_lib
      from doctor import ffmpeg_has_subtitles_filter
      from lib import env_bool, env_int, load_json
      from recap_source import audio_binding
      
      BUNDLE = Path(__file__).resolve().parents[2]  # the skills/ directory
      
      RUN_MANIFEST = "recap_run_manifest.json"
      
      MULTI_SOURCE_MANIFEST = "multi_source_manifest.json"
      
      
      def _run(skill, script, *cli_args):
          cmd = [sys.executable, str(_entry(skill, script)), *map(str, cli_args)]
          print(f"[video-recap] ▶ {skill}/{script}", flush=True)
          res = subprocess.run(cmd)
          if res.returncode != 0:
              raise SystemExit(f"{skill}/{script} 失败 (exit {res.returncode})")
      
      
      def _optional_env_int(name):
          """lib.env_int with -1 as the documented "unset" sentinel."""
          value = env_int(name, None)
          return None if value == -1 else value
      
      
      def _read_video_duration_or_raise(path):
          """Return media duration via ffprobe, or hard-fail before downstream TTS/render."""
          path = Path(path)
          cmd = [
              "ffprobe",
              "-v",
              "quiet",
              "-show_entries",
              "format=duration",
              "-of",
              "csv=p=0",
              str(path),
          ]
          res = subprocess.run(cmd, capture_output=True, text=True)
          if res.returncode != 0:
              detail = (res.stderr or res.stdout or "ffprobe failed").strip()
              raise SystemExit(f"无法读取成片时长: {path} ({detail})")
          try:
              duration = float(res.stdout.strip())
          except ValueError:
              raise SystemExit(f"无法读取成片时长: {path} (ffprobe 输出无效: {res.stdout!r})") from None
          if not math.isfinite(duration) or duration <= 0:
              raise SystemExit(f"无法读取成片时长: {path} (duration={duration:.3f})")
          return duration
      
      
      def _probe_display_height_or_raise(path, *, require_square_pixels=False):
          """Return ffmpeg's display-coordinate height, accounting for rotation and SAR."""
          return _probe_display_size_or_raise(path, require_square_pixels=require_square_pixels)[1]
      
      
      def _probe_display_size_or_raise(path, *, require_square_pixels=False):
          """Return ffmpeg's display-coordinate (width, height), accounting for rotation and SAR."""
          cmd = [
              "ffprobe",
              "-v",
              "error",
              "-select_streams",
              "v:0",
              "-show_entries",
              "stream=width,height,sample_aspect_ratio:stream_tags=rotate:stream_side_data=rotation",
              "-of",
              "json",
              str(path),
          ]
          res = subprocess.run(cmd, capture_output=True, text=True)
          try:
              stream = json.loads(res.stdout)["streams"][0]
              width, height = int(stream["width"]), int(stream["height"])
          except (IndexError, KeyError, ValueError):
              detail = (res.stderr or res.stdout or "ffprobe failed").strip()
              raise SystemExit(f"无法读取视频画布: {path} ({detail})") from None
          sar = stream.get("sample_aspect_ratio", "1:1")
          try:
              num, den = sar.split(":", 1)
              sar_ratio = float(num) / float(den)
          except (ValueError, ZeroDivisionError):
              sar_ratio = math.nan
          if require_square_pixels and (
              not math.isfinite(sar_ratio) or abs(sar_ratio - 1.0) >= 1e-9
          ):
              raise SystemExit(
                  f"subtitle Y coordinates currently require square-pixel video (SAR 1:1); got {sar}"
              )
          display_width = max(1, round(width * sar_ratio)) if math.isfinite(sar_ratio) else width
          rotation_values = [
              stream.get("tags", {}).get("rotate"),
              *(item.get("rotation") for item in stream.get("side_data_list", [])),
          ]
          rotation = next(
              (int(round(float(value))) % 360 for value in rotation_values if value not in (None, "")),
              0,
          )
          return (height, display_width) if rotation in {90, 270} else (display_width, height)
      
      
      def _analysis_settings(args):
          return {
              "context": args.context,
              "scene_threshold": args.scene_threshold,
              "style": args.style,
              "edit_mode": args.edit_mode,
              "target_duration": args.target_duration,
              "skip_asr": bool(args.skip_asr),
              "mimo_video_overview": bool(args.mimo_video_overview),
              "consolidate": bool(args.consolidate),
              "consolidate_asr": bool(args.consolidate_asr),
          }
      
      
      def _coerce_videos(video_or_videos):
          if isinstance(video_or_videos, (list, tuple)):
              return [Path(v).resolve() for v in video_or_videos]
          return [Path(video_or_videos).resolve()]
      
      
      def _run_manifest_payload(video, args):
          return {
              "schema_version": 1,
              "source_video": str(Path(video).resolve()),
              "source_video_identity": material_lib.file_identity(video),
              "settings": _analysis_settings(args),
              "audio": audio_binding(args),
          }
      
      
      def _write_run_manifest(work_dir, video, args):
          payload = _run_manifest_payload(video, args)
          (work_dir / RUN_MANIFEST).write_text(
              json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8"
          )
          return payload
      
      
      def _build_multi_source_records(videos, args):
          records = []
          for video in _coerce_videos(videos):
              identity = material_lib.file_identity(video)
              records.append(
                  {
                      "source_path": str(video),
                      "source_name": video.name,
                      "source_video_identity": identity,
                      "settings": _analysis_settings(args),
                      "material_id": material_lib.material_id_for(video, identity),
                  }
              )
          records = material_lib.assign_source_ids(records)
          for record in records:
              record["source_work_dir"] = f"sources/{record['source_id']}"
          return records
      
      
      def _multi_run_manifest_payload(videos, args, source_records):
          return {
              "schema_version": 2,
              "mode": "multi_source",
              "sources": [
                  {
                      "source_id": s["source_id"],
                      "source_path": s["source_path"],
                      "source_video_identity": s["source_video_identity"],
                      "source_work_dir": s["source_work_dir"],
                      "material_id": s["material_id"],
                  }
                  for s in source_records
              ],
              "source_videos": [str(v) for v in _coerce_videos(videos)],
              "settings": _analysis_settings(args),
              "audio": audio_binding(args),
          }
      
      
      def _write_project_run_manifest(work_dir, videos, args, source_records):
          payload = _multi_run_manifest_payload(videos, args, source_records)
          (work_dir / RUN_MANIFEST).write_text(
              json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8"
          )
      
      
      def _write_multi_source_manifest(work_dir, source_records):
          path = Path(work_dir) / MULTI_SOURCE_MANIFEST
          payload = {
              "schema_version": 1,
              "sources": [
                  {
                      "source_id": s["source_id"],
                      "source_path": s["source_path"],
                      "source_name": s["source_name"],
                      "source_video_identity": s["source_video_identity"],
                      "source_work_dir": s["source_work_dir"],
                      "material_id": s["material_id"],
                  }
                  for s in source_records
              ],
          }
          path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
          return path
      
      
      def _load_run_manifest(work_dir):
          """The run manifest, or None before Phase A has written one."""
          path = Path(work_dir) / RUN_MANIFEST
          return load_json(path) if path.exists() else None
      
      
      def _burn_subtitles_intended(args):
          """Effective burn-subtitles state at orchestrator level. Mirrors video-assemble's
          CONFIG default `env_bool("BURN_SUBTITLES", True)` (burn is ON by default); an explicit
          CLI flag (--burn-subtitles / --no-burn-subtitles) overrides the env."""
          if args.burn_subtitles is not None:
              return args.burn_subtitles
          return env_bool("BURN_SUBTITLES", True)
      
      
      def _ffmpeg_present_but_cannot_burn():
          """True only when ffmpeg EXISTS but lacks the libass `subtitles` filter — the specific
          "subtitle-burn environment unsupported" case. Returns False when ffmpeg is absent
          entirely: that is a more fundamental problem that surfaces at the first stage (understand
          calls ffprobe/ffmpeg) and is reported by doctor, so this guard stays narrow — and does
          not fire in mocked, ffmpeg-less test environments."""
          if shutil.which("ffmpeg") is None:
              return False
          return not ffmpeg_has_subtitles_filter()
      
      
      def _preflight_burn_subtitles(args):
          """Fail fast BEFORE any understanding/VLM/ASR/TTS spend when subtitle burn-in is on but
          this ffmpeg can't burn it. Without it the run only dies at the final assemble
          `-vf subtitles=` step — after the whole expensive pipeline has run. Dub renders through
          dub.py, which never burns subtitles, so it is exempt."""
          if args.edit_mode == "dub" or not _burn_subtitles_intended(args):
              return
          if _ffmpeg_present_but_cannot_burn():
              raise SystemExit(
                  "字幕烧录已开启,但当前 ffmpeg 不支持 subtitles/libass 滤镜,整条流程会跑到最后渲染才失败。\n"
                  "  解决其一:(1) 安装带 libass 的 ffmpeg;(2) 加 --no-burn-subtitles 关闭烧录"
                  "(仍输出 .srt 外挂字幕)。\n"
                  f"  自检:python3 {shlex.quote(str(_entry('video-recap', 'doctor.py')))}"
              )
      
      
      def _entry(skill, script):
          return BUNDLE / skill / "scripts" / script
      
    • recap_source.py 10.4 KB
      """Validate recap audio ownership without importing another skill."""
      
      import json
      from pathlib import Path
      
      
      AUDIO_MODES = ("narration", "source-mix", "adopted-packet-copy")
      
      _TTS_OPTIONS = frozenset({
          "--tts-provider",
          "--mimo-tts-voice",
          "--voice-ref",
          "--allow-partial-tts",
          "--preserve-approved-text",
          "--review-narration",
          "--no-review-narration",
          "--require-narration-review",
      })
      
      _LOCAL_ADOPTION_OPTIONS = (
          ("tts_meta", "--tts-meta"),
          ("narration_adoption", "--narration-adoption"),
          ("audio_mix_adoption", "--audio-mix-adoption"),
      )
      
      _LOCAL_ADOPTION_CONFLICTS = _TTS_OPTIONS | frozenset({
          "--mimo-qc",
          "--mimo-qc-refresh",
          "--export-jianying",
          "--jianying-bundle-media",
          "--jianying-no-bundle-media",
      })
      
      
      def uses_narration(args):
          return args.audio_mode == "narration"
      
      
      def uses_local_adoption(args):
          return all(getattr(args, field) is not None for field, _ in _LOCAL_ADOPTION_OPTIONS)
      
      
      def needs_voiceover(args):
          return uses_narration(args) and not uses_local_adoption(args)
      
      
      def audio_binding(args):
          binding = {
              "mode": args.audio_mode,
              "selected_stream_index": args.audio_stream_index,
          }
          if uses_local_adoption(args):
              binding["local_adoption"] = {
                  field: {"path": str(Path(getattr(args, field)).resolve())}
                  for field, _ in _LOCAL_ADOPTION_OPTIONS
              }
          return binding
      
      
      def _resolve_local_file(parser, value, option):
          path = Path(value).expanduser().resolve()
          if not path.is_file():
              parser.error(f"{option} does not exist or is not a file: {path}")
          return str(path)
      
      
      def validate_local_adoption(parser, args):
          """Validate and resolve the assembly-only local adoption boundary."""
          supplied = [getattr(args, field) is not None for field, _ in _LOCAL_ADOPTION_OPTIONS]
          if any(supplied) and not all(supplied):
              parser.error("--tts-meta, --narration-adoption and --audio-mix-adoption are all-or-none")
          if not all(supplied):
              return
          if args.edit_mode != "full" or args.audio_mode != "narration":
              parser.error("local audio adoption requires --edit-mode full and --audio-mode narration")
          if args.audio_stream_index != 0:
              parser.error("local audio adoption requires --audio-stream-index 0")
          if len(args.video) != 1:
              parser.error("local audio adoption requires exactly one input video")
          if args.work_dir is None:
              parser.error("local audio adoption requires explicit --work-dir")
          conflicts = sorted(set(args._explicit_options).intersection(_LOCAL_ADOPTION_CONFLICTS))
          if conflicts:
              parser.error("local audio adoption cannot use: " + ", ".join(conflicts))
          for field, option in _LOCAL_ADOPTION_OPTIONS:
              setattr(args, field, _resolve_local_file(parser, getattr(args, field), option))
      
          work_dir = Path(args.work_dir).expanduser().resolve()
          if work_dir.exists():
              parser.error("local audio adoption requires a new --work-dir")
          output_dir = (
              Path(args.output_dir).expanduser().resolve()
              if args.output_dir is not None else work_dir.parent
          )
          # Pre-run guard only: mirrors the assembler's `recap_<stem>.mp4` naming so an existing
          # delivery is refused before work starts. The authoritative published path is read back
          # from the child's assembly_manifest.json after the run.
          delivery = output_dir / f"recap_{Path(args.video[0]).stem}.mp4"
          if delivery.exists():
              parser.error(f"local audio adoption will not overwrite delivery: {delivery}")
          args.work_dir = str(work_dir)
          args.output_dir = str(output_dir)
      
          # Ambient authoring configuration is irrelevant because this route never synthesizes
          # or reviews narration. Explicit forms were rejected above.
          args.tts_provider = "auto"
          args.mimo_tts_voice = None
          args.voice_ref = None
          args.mimo_qc = "off"
      
      
      def load_local_assembly_evidence(work_dir):
          """Load the assembler's binding records for the adopted local bundle."""
          work_dir = Path(work_dir)
          try:
              return {
                  name: json.loads((work_dir / filename).read_text(encoding="utf-8"))
                  for name, filename in (
                      ("assembly", "assembly_manifest.json"),
                      ("narration", "narration_input_binding.json"),
                      ("mix", "audio_mix_binding.json"),
                  )
              }
          except (OSError, UnicodeDecodeError, json.JSONDecodeError) as exc:
              raise SystemExit("assembler did not publish readable local adoption evidence") from exc
      
      
      def _same_file(declared, expected):
          return Path(declared).resolve() == Path(expected).resolve()
      
      
      def verify_local_assembly_evidence(evidence, manifest):
          """The child bindings must reference the same adoption files and picture recap ran with."""
          try:
              local = manifest["audio"]["local_adoption"]
              narration = evidence["narration"]
              mix = evidence["mix"]
              matches = (
                  _same_file(narration["adoption"]["path"], local["narration_adoption"]["path"])
                  and _same_file(narration["adoption"]["tts_meta"]["path"], local["tts_meta"]["path"])
                  and _same_file(mix["adoption"]["path"], local["audio_mix_adoption"]["path"])
                  and _same_file(mix["picture"]["path"], manifest["source_video"])
              )
          except (KeyError, TypeError):
              matches = False
          if not matches:
              raise SystemExit("assembler bindings do not match the local adoption manifest")
      
      
      def owned_local_delivery(evidence):
          """Return a stable ownership token only when all child reports bind the published file."""
          try:
              expected = Path(evidence["assembly"]["final_output"]).resolve()
          except (KeyError, TypeError) as exc:
              raise SystemExit(
                  "assembly manifest does not declare final_output; cannot own the delivery"
              ) from exc
          try:
              declared = [
                  Path(evidence["narration"]["final_output"]["path"]).resolve(),
                  Path(evidence["mix"]["final_output"]["path"]).resolve(),
              ]
          except (KeyError, TypeError):
              return None
          if declared != [expected, expected] or not expected.is_file():
              return None
          stat = expected.stat()
          return {"path": expected, "device": stat.st_dev, "inode": stat.st_ino}
      
      
      def remove_owned_local_delivery(token):
          if token is None:
              return
          path = token["path"]
          if not path.is_file():
              return
          stat = path.stat()
          if (stat.st_dev, stat.st_ino) == (token["device"], token["inode"]):
              path.unlink()
      
      
      def reject_unbound_narration_workdir(work_dir, args):
          """Do not reinterpret an unbound narrated work directory as source-owned audio."""
          work_dir = Path(work_dir)
          if (
              not uses_narration(args)
              and not (work_dir / "recap_run_manifest.json").exists()
              and (work_dir / "narration.json").exists()
          ):
              raise SystemExit(
                  "当前 work_dir 含 narration.json 但缺少可验证的音频运行策略;"
                  "非解说模式请使用新的 --work-dir"
              )
      
      
      def validate_audio_routing(parser, args):
          """Reject combinations the current pipeline cannot execute truthfully."""
          validate_local_adoption(parser, args)
          mode = args.audio_mode
          stream = args.audio_stream_index
          if stream < 0:
              parser.error("--audio-stream-index must be a non-negative integer")
          if args.edit_mode == "dub" and mode != "narration":
              parser.error("--edit-mode dub cannot be combined with a non-narration --audio-mode")
          # index-tts + MiMo voice/--voice-ref is rejected once, in recap_runner.main, after the
          # ambient VOICE_REF/MIMO_TTS_VOICE have been folded in.
          if args.tts_provider == "index-tts" and args.edit_mode == "dub":
              parser.error("--edit-mode dub does not support --tts-provider index-tts")
          if mode == "narration":
              if stream != 0:
                  parser.error("narration currently requires --audio-stream-index 0")
              return
      
          conflicts = sorted(set(args._explicit_options).intersection(_TTS_OPTIONS))
          if conflicts:
              detail = ", ".join(conflicts)
              parser.error(f"--audio-mode {mode} cannot use TTS/strict narration options: {detail}")
          if args.mimo_qc != "off":
              parser.error(f"--audio-mode {mode} currently requires --mimo-qc off")
          if args.edit_mode == "cut" and stream != 0:
              parser.error("cut audio modes currently require --audio-stream-index 0")
          if args.export_jianying and stream != 0:
              parser.error("JianYing export currently requires --audio-stream-index 0")
          # Environment defaults are irrelevant once source audio owns the run. Normalizing
          # prevents an invalid ambient provider from leaking into later state or continuation.
          args.tts_provider = "auto"
      
      
      def extend_assemble_args(cli_args, args):
          mode = args.audio_mode
          stream = args.audio_stream_index
          if mode != "narration":
              cli_args += ["--audio-mode", mode]
          if stream != 0:
              cli_args += ["--audio-stream-index", str(stream)]
          return cli_args
      
      
      def reject_unsupported_subtitle_track(work_dir, args):
          if args.audio_mode == "source-mix" and (Path(work_dir) / "subtitle_track.json").exists():
              raise SystemExit(
                  "source-mix 当前不能绑定显式 subtitle_track.json;请使用新的 work_dir,"
                  "或选择 adopted-packet-copy"
              )
      
      
      def begin_non_narration_qc(work_dir, args, write_stage):
          """Start a clean run-local QC ledger without reading or deleting old TTS."""
          (Path(work_dir) / "preflight_qc.json").unlink(missing_ok=True)
          return write_stage(
              work_dir,
              "pre_assemble",
              metadata={
                  "audio_mode": args.audio_mode,
                  "selected_audio_stream_index": args.audio_stream_index,
                  "tts": "not_applicable",
                  "narration_validation": "not_applicable",
                  "narration_review": "not_applicable",
                  "visual_overlays": "preserved_not_authored_by_this_run",
              },
          )
      
      
      def begin_local_adoption_qc(work_dir, write_stage):
          """Record that authoring was intentionally skipped for adopted local assets."""
          return write_stage(
              work_dir,
              "pre_assemble",
              metadata={
                  "audio_mode": "narration",
                  "selected_audio_stream_index": 0,
                  "tts": "adopted_local_not_generated",
                  "narration_validation": "not_run",
                  "narration_review": "not_run",
                  "semantic_validation": "video-assemble",
                  "visual_overlays": "not_authored_not_present",
              },
          )
      
    • recap_stage_qc.py 5.7 KB
      """Write shift-left, MiMo, and final QC stage reports."""
      
      import json
      from pathlib import Path
      
      import final_qc
      import mimo_qc
      import qc_contract
      
      ASSEMBLY_MANIFEST = "assembly_manifest.json"
      PREFLIGHT_QC = "preflight_qc.json"
      
      
      def _load_preflight_stage_reports(work_dir):
          path = Path(work_dir) / PREFLIGHT_QC
          if not path.exists():
              return {}
          return json.loads(path.read_text(encoding="utf-8"))["metadata"]["stages"]
      
      
      def _write_shift_left_stage_qc(work_dir, stage, metadata, findings=None):
          """Write/roll up local shift-left QC for one pipeline stage.
      
          This is a local contract artifact only: no MiMo/deep eval calls, no repair, and no
          credential persistence (qc_contract redacts every report it builds).
          """
          stage_report = qc_contract.build_report(
              artifact=PREFLIGHT_QC, stage=stage, findings=findings, metadata=metadata
          )
          stage_reports = _load_preflight_stage_reports(work_dir)
          stage_reports[stage] = stage_report
          top_metadata = dict(stage_report["metadata"])
          top_metadata["latest_stage"] = stage
          top_metadata["stages"] = stage_reports
          top_report = qc_contract.build_report(
              artifact=PREFLIGHT_QC,
              stage=stage,
              findings=[f for report in stage_reports.values() for f in report["findings"]],
              metadata=top_metadata,
          )
          (Path(work_dir) / PREFLIGHT_QC).write_text(
              json.dumps(top_report, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
          )
          return top_report
      
      
      def _tts_qc_metadata(work_dir):
          """Called after video-voiceover exited 0, which always writes tts_meta.json."""
          work_dir = Path(work_dir)
          metadata = {"tts_meta": json.loads((work_dir / "tts_meta.json").read_text(encoding="utf-8"))}
          tts_dir = work_dir / "tts_segments"
          if tts_dir.is_dir():
              metadata["tts_segments"] = [
                  p.relative_to(work_dir).as_posix() for p in sorted(tts_dir.iterdir()) if p.is_file()
              ]
          return metadata
      
      
      def _post_render_qc_metadata(work_dir, final_output):
          """Called after video-assemble exited 0, which always writes assembly_manifest.json."""
          manifest = Path(work_dir) / ASSEMBLY_MANIFEST
          return {
              "final_output": str(final_output),
              "assembly_manifest": json.loads(manifest.read_text(encoding="utf-8")),
          }
      
      
      def _write_final_qc_reports(work_dir, final_output):
          """Write report-only final QC artifacts after render.
      
          final_qc.run converts ffprobe unavailability/failure into deterministic
          blockers; only unexpected schema/write errors propagate.
          """
          return final_qc.run(work_dir, final_output=final_output)
      
      
      def _print_final_qc_pointer(result):
          """Surface a report-only final_qc/golden_eval FAIL so the shift-left QC is
          not a silent no-op. Advisory only: it never changes the exit status."""
          problems = [
              f"{key} blocker_count={result[key].get('blocker_count', '?')}"
              for key in ("final_qc", "golden_eval")
              if result[key].get("ok") is False
          ]
          if problems:
              print(
                  "[video-recap] ⚠️  最终 QC 未通过(仅报告,不阻断): "
                  + "; ".join(problems)
                  + ";详见 final_qc.json / golden_eval.json"
              )
      
      
      def _require_final_qc(result, work_dir):
          """Fail closed unless both final summaries are literal blocker-free passes."""
          paths = [Path(work_dir) / name for name in ("final_qc.json", "golden_eval.json")]
          print("[video-recap] 最终 QC 报告: " + "; ".join(map(str, paths)))
          invalid = []
          for name in ("final_qc", "golden_eval"):
              summary = result.get(name) if isinstance(result, dict) else None
              blockers = summary.get("blocker_count") if isinstance(summary, dict) else None
              if not isinstance(summary, dict) or summary.get("ok") is not True or \
                      type(blockers) is not int or blockers != 0:
                  invalid.append(name)
          if invalid:
              raise SystemExit(
                  "严格最终 QC 未通过或摘要格式无效: " + ", ".join(invalid)
              )
      
      
      def _mimo_qc_stage_enabled(args, stage):
          return (
              args.mimo_qc == "both"
              or (args.mimo_qc == "pre-assemble" and stage == "pre_assemble")
              or (args.mimo_qc == "post-render" and stage == "post_render")
          )
      
      
      def _prepare_mimo_qc(work_dir, args):
          """Remove an old advisory artifact when this run has MiMo QC disabled."""
          if args.mimo_qc == "off":
              mimo_qc.clear_report(work_dir)
      
      
      def _print_mimo_qc_pointer(result, stage):
          report = result["report"]
          metadata = report["metadata"]
          status = metadata["status"]
          stage_findings = [f for f in report["findings"] if f["stage"] == stage]
          if status in {"failed", "unavailable"}:
              print(
                  f"[video-recap] ⚠ MiMo QC {stage}: {status} ({metadata['error']});建议性检查不可用,继续流水线"
              )
              return
          print(
              f"[video-recap] ℹ MiMo QC {stage}: {status}, {len(stage_findings)} 条建议;详见 {result['path']}"
          )
          for finding in stage_findings[:5]:
              print(f"[video-recap]   - {finding['message']}")
      
      
      def _run_mimo_qc_stage(work_dir, args, stage, *, final_output=None):
          """Run one selected advisory stage and never propagate a failure."""
          if not _mimo_qc_stage_enabled(args, stage):
              return None
          try:
              result = mimo_qc.run(
                  work_dir,
                  stage=stage,
                  live=True,
                  refresh=args.mimo_qc_refresh,
                  final_output=final_output,
              )
          except Exception as exc:
              print(
                  f"[video-recap] ⚠ MiMo QC {stage}: {type(exc).__name__};"
                  "建议性检查失败,继续流水线"
              )
              return None
          _print_mimo_qc_pointer(result, stage)
          return result
      
    • recap_timeline.py 28 KB
      """Own recap timeline artifacts, continuation state, and cut QC surfaces."""
      
      import json
      import os
      import shlex
      import sys
      from pathlib import Path
      
      from lib import load_json
      from recap_runtime import (
          _coerce_videos,
          _entry,
          _load_run_manifest,
          _multi_run_manifest_payload,
          _run_manifest_payload,
      )
      from recap_source import audio_binding, uses_narration
      
      ASSEMBLY_MANIFEST = "assembly_manifest.json"
      
      PHASE_LEDGER = "recap_phase.json"
      
      CUT_TIMELINE_CRAFT_BULLETS = [
          "- 片段顺序必须服务同一条故事主线,而不是无序高光;可使用 0–1 个 cold open,随后回到 setup → turn → escalation → payoff。",
          "- 每个片段必须对应 `recap_story_plan.json` 中的一个 change-based beat;删除后不损失因果、人物或情绪的片段通常不保留。",
          "- 已从源证据确认的必保问答、反应或兑现段,写入同一 `clip_plan.json.required_evidence`(nodes登记id/source绝对路径/start/end原片秒/track/content,before登记必要先后);只锁具体区间,不锁整个beat。剪点吸附后工具会核验,reason或花字不能代替缺失的源段。",
          "- `reason` 统一写成 `beat_id | function | change | POV | preferred moment | 入点 | 出点`,不能只写 hook、重要剧情或事件摘要。",
          "- 优先保留因果、揭示、决定、关系移动、情绪转向与不可替代的表演/反应;跳过片尾、广告、重复静态画面和水印废片段。",
          "- 片段追求最短但完整:建立镜头可以短,关键表演/反应允许多停一点;在完整台词、完整动作或自然声音边界结束,避免原声从半句中切入或切出。",
          "- 新人物/地点/情绪场景先给画面建立空间,再进旁白;不要在原片叠化或闪白中间再切一次。scene score 只用于定位候选接点,最终必须播放接点前后判断。",
          "- 对短时间内密集的 scene 候选先区分来源:原片的无关短镜头整段删除,相关短镜头扩展到完整动作/反应;本次拼接制造的切点优先移动边界、恢复同源连续运动或合并片段,能消除就不保留。",
      ]
      
      MULTI_SOURCE_NARRATION_CRAFT_BULLETS = [
          "- `narration.json` 只使用 edited_source.mp4 的 OUTPUT 时间线,不使用任一原片时间。",
          "- 先为每个 beat 指定 `audio_owner`,再决定是否需要旁白;允许 original_dialogue、action_sound、ambience、music、silence 或 narration 主导。",
          "- 旁白只承担 context、causal_link、foreshadow、interpretation 或 transition;`narration_job=none` 的 beat 不写旁白。",
          "- 7:3 不是配额,只是素材无法给出更好判断时的粗略回退;强对白、动作声、环境或沉默可以完整拥有一个 beat。",
          "- 旁白拥有 beat 时才写成一个连续思路的 BLOCK;句数不是目标,字幕拆 cue 也不能切碎 TTS。不要把每个 beat 都机械变成旁白块或固定原声留白。",
          "- 使用 kept-clip map 与每个 source work_dir 核对人物、事实、ASR 和上下文;跨源转场必须有明确叙事任务。",
      ]
      
      _ALLOWED_VISUAL_OVERLAY_TYPES = {"top_title", "inline_label_or_callout"}
      
      _VISUAL_OVERLAYS = "visual_overlays.json"
      
      
      def _canonical_visual_overlay(overlay, segment):
          """Return the first-release assemble overlay contract, or None for unsupported types.
      
          Recap owns the handoff artifact but does not invent richer overlay semantics: only the
          two overlay types implemented by assemble pass through, an overlay without its own
          start/end inherits the narration segment's, and renderer placement hints are preserved.
          """
          if overlay["type"] not in _ALLOWED_VISUAL_OVERLAY_TYPES:
              return None
          item = {
              "type": overlay["type"],
              "text": overlay["text"],
              "start": overlay.get("start", segment["start"]),
              "end": overlay.get("end", segment["end"]),
          }
          for key in ("anchor", "x", "y", "max_width", "style"):
              if key in overlay:
                  item[key] = overlay[key]
          return item
      
      
      def _write_canonical_visual_overlays(work_dir, narration_path):
          """Write assemble's canonical work_dir/visual_overlays.json recap handoff.
      
          Direct assemble still supports user-authored/manual visual_overlays.json files. Once
          recap owns the handoff for a narration, the canonical artifact is a deterministic
          reflection of the current narration so reused work_dirs cannot render stale overlays
          from a previous run; unsupported types are filtered out and represented as an explicit
          empty overlay list.
          """
          path = Path(work_dir) / _VISUAL_OVERLAYS
          overlays = [
              item
              for segment in load_json(narration_path)
              for overlay in segment.get("visual_overlays", [])
              if (item := _canonical_visual_overlay(overlay, segment)) is not None
          ]
          payload = {"schema_version": 1, "overlays": overlays}
          path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
          print(f"[video-recap] 🧩 visual overlays: {len(overlays)} → {path}", flush=True)
          return path
      
      
      def _print_grounding_qc_pointer(work_dir):
          qc_path = Path(work_dir) / "grounding_qc.json"
          if not qc_path.exists():
              return
          data = load_json(qc_path)
          ranges = data["review_coverage"]["time_ranges"]
          warnings = data["warnings"]
          suffix = f" · warnings {len(warnings)}" if warnings else ""
          print(
              f"[video-recap] 🧭 Grounding QC: {data['verdict']} · ranges {len(ranges)}{suffix} → {qc_path}",
              flush=True,
          )
      
      
      def _print_narration_review_pointer(work_dir, *, review_ran=True):
          """Surface the advisory narration review produced by this run, if any.
      
          Review is optional/fail-open: review_ran=False (disabled or failed) prints nothing so a
          stale narration_review.md from an older run is never surfaced. When it ran, video-script's
          review_runner has written both narration_review.json and .md.
          """
          _print_grounding_qc_pointer(work_dir)
          if not review_ran:
              return
          review_md = Path(work_dir) / "narration_review.md"
          data = load_json(Path(work_dir) / "narration_review.json")
          n_err = sum(1 for f in data["findings"] if f["severity"] == "error")
          print(
              f"[video-recap] 📋 解说评审(建议性,不拦截): {data['verdict']} · "
              f"{len(data['findings'])} 条意见(error {n_err})→ {review_md}"
          )
      
      
      def _settings_for_compare(settings):
          """Settings that, if changed, invalidate reusing an existing work_dir on resume.
      
          `consolidate`/`consolidate_asr` are EXCLUDED: they only ADD an optional understanding
          artifact and never re-run Phase A on a Phase-B resume, so a stored manifest carrying the
          old default (or missing the key entirely, pre-dating it) must still resume — otherwise
          flipping `--consolidate`'s default ON would hard-fail every in-flight work_dir.
          """
          s = dict(settings)
          s.pop("consolidate", None)
          s.pop("consolidate_asr", None)
          return s
      
      
      def _manifest_mismatches(work_dir, video, args):
          expected = _run_manifest_payload(video, args)
          actual = _load_run_manifest(work_dir)
          if actual is None:
              return ["缺少 recap_run_manifest.json;不能证明 work_dir 属于当前视频/参数"]
          # A multi-source manifest carries `sources` instead of these keys, so .get() reports it
          # as a mismatch rather than a crash.
          mismatches = [
              f"{key}: expected {expected[key]!r}, got {actual.get(key)!r}"
              for key in ("source_video", "source_video_identity")
              if actual.get(key) != expected[key]
          ]
          if _settings_for_compare(actual["settings"]) != _settings_for_compare(expected["settings"]):
              mismatches.append("settings: 当前 CLI/env 参数与 Phase A manifest 不匹配")
          actual_audio = actual.get("audio")
          expected_audio = audio_binding(args)
          if actual_audio != expected_audio:
              mismatches.append(
                  f"audio: expected {expected_audio!r}, got {actual_audio!r}"
              )
          return mismatches
      
      
      def _multi_manifest_mismatches(work_dir, videos, args, source_records):
          expected = _multi_run_manifest_payload(videos, args, source_records)
          actual = _load_run_manifest(work_dir)
          if actual is None:
              return ["缺少 recap_run_manifest.json;不能证明 work_dir 属于当前多视频/参数"]
          if actual.get("mode") != "multi_source":
              return [f"mode: expected 'multi_source', got {actual.get('mode')!r}"]
          mismatches = []
          identity = ("source_id", "source_path", "source_video_identity")
          # A 0.5.0 manifest carries `source_video_fingerprint` instead of `source_video_identity`;
          # .get() reports it as a mismatch rather than a crash.
          if [{k: s.get(k) for k in identity} for s in actual["sources"]] != [
              {k: s[k] for k in identity} for s in expected["sources"]
          ]:
              mismatches.append(
                  "sources: 当前输入视频列表/顺序/source_id/identity 与 Phase A manifest 不匹配"
              )
          if _settings_for_compare(actual["settings"]) != _settings_for_compare(expected["settings"]):
              mismatches.append("settings: 当前 CLI/env 参数与 Phase A manifest 不匹配")
          actual_audio = actual.get("audio")
          expected_audio = audio_binding(args)
          if actual_audio != expected_audio:
              mismatches.append(
                  f"audio: expected {expected_audio!r}, got {actual_audio!r}"
              )
          return mismatches
      
      
      def _read_assembly_output(work_dir):
          return Path(load_json(Path(work_dir) / ASSEMBLY_MANIFEST)["final_output"])
      
      
      def _read_phase_ledger(work_dir):
          """Phase ledger (cut mode): which artifacts exist and the clip_plan/narration they match.
      
          Lets resume be driven by recorded phase state rather than bare file existence — the
          prerequisite for the cut-first/narrate-second two-pause flow, and the guard that keeps a
          narration written for one clip_plan from silently driving a different cut into TTS.
          None before the first cut pass has recorded anything.
          """
          path = Path(work_dir) / PHASE_LEDGER
          return load_json(path) if path.exists() else None
      
      
      def _write_phase_ledger(work_dir, **fields):
          ledger = _read_phase_ledger(work_dir) or {}
          ledger.update(fields)
          (Path(work_dir) / PHASE_LEDGER).write_text(
              json.dumps(ledger, ensure_ascii=False, indent=2), encoding="utf-8"
          )
          return ledger
      
      
      def _cut_narration_is_stale(ledger, current_clip_plan_identity):
          """Two-pass cut: the narration is authored against the rendered cut shown at the A2 pause,
          i.e. against the clip_plan recorded in the ledger (its ``{size, mtime_ns}`` identity). If
          clip_plan.json was rewritten since (a re-cut) while that narration is still present, it
          describes the OLD cut — stale."""
          return ledger is not None and ledger["clip_plan_identity"] != current_clip_plan_identity
      
      
      def _continuation_command(video, work_dir, args):
          # Values a project binding filled in are re-derived from --project on resume.
          bound = getattr(args, "_bound_from_project", frozenset())
          parts = [
              sys.executable,
              str(_entry("video-recap", "recap.py")),
              *[str(v) for v in _coerce_videos(video)],
              "--work-dir",
              str(work_dir),
          ]
          if args.context:
              parts += ["--context", args.context]
          if args.scene_threshold is not None:
              parts += ["--scene-threshold", str(args.scene_threshold)]
          if args.style != "纪录片":
              parts += ["--style", args.style]
          if args.edit_mode != "full":
              parts += ["--edit-mode", args.edit_mode]
          if args.audio_mode != "narration":
              parts += ["--audio-mode", args.audio_mode]
          if args.audio_stream_index != 0:
              parts += ["--audio-stream-index", str(args.audio_stream_index)]
          if args.target_duration:
              parts += ["--target-duration", args.target_duration]
          if args.allow_duration_drift:
              parts.append("--allow-duration-drift")
          if args.skip_asr:
              parts.append("--skip-asr")
          if args.mimo_video_overview:
              parts.append("--mimo-video-overview")
          if args.mimo_qc != "off":
              parts += ["--mimo-qc", args.mimo_qc]
          if args.mimo_qc_refresh:
              parts.append("--mimo-qc-refresh")
          if not args.consolidate:  # default is ON; only the opt-out needs to round-trip
              parts.append("--no-consolidate")
          if args.consolidate_asr:
              parts.append("--consolidate-asr")
          if uses_narration(args):
              if args.mimo_tts_voice and "mimo_tts_voice" not in bound:
                  parts += ["--mimo-tts-voice", args.mimo_tts_voice]
              if args.tts_provider != "auto" and "tts_provider" not in bound:
                  parts += ["--tts-provider", args.tts_provider]
              if args.voice_ref and "voice_ref" not in bound:
                  parts += ["--voice-ref", args.voice_ref]
              if args.allow_partial_tts:
                  parts.append("--allow-partial-tts")
              if args.preserve_approved_text:
                  parts.append("--preserve-approved-text")
          if args.burn_subtitles is not None:
              parts.append("--burn-subtitles" if args.burn_subtitles else "--no-burn-subtitles")
          if args.subtitle_y_top is not None:
              parts += ["--subtitle-y-top", str(args.subtitle_y_top)]
          if args.subtitle_y_bot is not None:
              parts += ["--subtitle-y-bot", str(args.subtitle_y_bot)]
          if args.output_dir:
              parts += ["--output-dir", args.output_dir]
          if args.export_jianying:
              parts.append("--export-jianying")
          if args.jianying_bundle_media:
              parts.append("--jianying-bundle-media")
          if args.jianying_no_bundle_media:
              parts.append("--jianying-no-bundle-media")
          if uses_narration(args):
              if args.review_narration is not None:
                  parts.append(
                      "--review-narration" if args.review_narration else "--no-review-narration"
                  )
              if args.require_narration_review:
                  parts.append("--require-narration-review")
          if args.material_library_dir and "material_library_dir" not in bound:
              parts += ["--material-library-dir", args.material_library_dir]
          if getattr(args, "project", None):
              parts += ["--project", args.project]
          if args.use_materials:
              parts.append("--use-materials")
          if args.save_materials:
              parts.append("--save-materials")
          if args.require_final_qc:
              parts.append("--require-final-qc")
          return " ".join(shlex.quote(part) for part in parts)
      
      
      def _source_work_dir(project_work_dir, source_record):
          return Path(project_work_dir) / source_record["source_work_dir"]
      
      
      def _understand_args_for_source(source_record, source_work_dir, args):
          uargs = [
              source_record["source_path"],
              "--work-dir",
              str(source_work_dir),
              "--style",
              args.style,
              "--edit-mode",
              args.edit_mode,
          ]
          if args.context:
              uargs += ["--context", args.context]
          if args.scene_threshold is not None:
              uargs += ["--scene-threshold", str(args.scene_threshold)]
          if args.target_duration:
              uargs += ["--target-duration", args.target_duration]
          if args.skip_asr:
              uargs.append("--skip-asr")
          if args.mimo_video_overview:
              uargs.append("--mimo-video-overview")
          uargs.append("--consolidate" if args.consolidate else "--no-consolidate")
          if args.consolidate_asr:
              uargs.append("--consolidate-asr")
          return uargs
      
      
      def _brief_excerpt(path, limit=1200):
          text = Path(path).read_text(encoding="utf-8")
          if len(text) <= limit:
              return text
          marker = "\n…\n"
      
          def section(heading):
              lines = text.splitlines()
              try:
                  start = lines.index(heading)
              except ValueError:
                  return ""
              end = next(
                  (i for i in range(start + 1, len(lines)) if lines[i].startswith("## ")),
                  len(lines),
              )
              return "\n".join(lines[start:end]).strip()
      
          def clipped(value, budget):
              if len(value) <= budget:
                  return value
              usable = budget - len(marker)
              head = round(usable * 0.7)
              return value[:head].rstrip() + marker + value[len(value) - (usable - head):].lstrip()
      
          # The generic writing contract occupies the front of every generated brief. A raw
          # prefix therefore drops the source facts that multi-source planning actually needs.
          # Give evidence-bearing sections equal space and retain both their opening and tail.
          preferred = [
              value
              for value in (
                  section("## Story context (from background_research.json)"),
                  section("## Understanding index (from consolidate.py)"),
                  section("## ASR writing chunks (semantic windows)"),
                  section("## Scene timing guide"),
              )
              if value
          ]
          if preferred:
              per_section = (limit - 2 * (len(preferred) - 1)) // len(preferred)
              return "\n\n".join(clipped(value, per_section) for value in preferred)[:limit]
          usable = limit - len(marker)
          head = usable // 4
          return text[:head].rstrip() + marker + text[len(text) - (usable - head):].lstrip()
      
      
      def _write_multi_source_clip_brief(work_dir, source_records, args):
          lines = [
              "# Multi-source Clip Plan Brief",
              "",
              "你正在做多视频剪辑复盘。当前 MVP 只支持 `--edit-mode cut`:先做跨素材的故事与视听决定,再写 `clip_plan.json`;下一步会剪出 `edited_source.mp4`,最后按 OUTPUT 时间轴写 `narration.json`。",
              "",
              "## 创作决定",
              "",
              "先判断创作控制模式:CREATE 比较至少两个可行剪辑假设;DIRECTED 落实用户指定结构;REVISION 只改最新反馈点并冻结未点名内容。然后写或更新 `recap_story_plan.json` 与 `visual_audio_board.json`。前者记录观众承诺、POV、戏剧问题、选定主线及 change-based beats;后者记录每拍的具体画面/反应、入点/出点、原声锚点、`audio_owner` 与 `narration_job`。",
              "",
              "多视频不是把每个来源各做一段小总结。每个来源片段都必须服务同一条主线,并用 `source_id` 保留证据归属。这两份计划是 Agent 与建议型评审使用的工作记录,不是 CLI 渲染门禁。",
              "",
              "## 必须写入的格式",
              "",
              "```json",
              '{"target_duration":"10m","clips":[{"source_id":"src_xxx","start":12.0,"end":38.0,"reason":"b01 | hook | knowledge: unknown→threat | POV=主角 | 保留倾听反应 | 入点=问题已问出 | 出点=沉默落地"}]}',
              "```",
              "",
              "- 每个 clip 必须带 `source_id`。",
              "- `start`/`end` 是对应 source 原视频时间(秒)。",
              "- 不同 `source_id` 的相同时间段不算重叠;同一 `source_id` 内不要重复/重叠,除非你明确接受稀疏/重复剪辑风险。",
              '- 素材库是文件系统 JSON/MD/JSONL;需要找历史素材时直接 `grep -R "关键词" <material-library-dir>`。',
              "",
              "## 剪辑规则",
              *CUT_TIMELINE_CRAFT_BULLETS,
              "- 跨来源选择必须服务共享主线;除非本来就是 setup / turn / payoff 的设计,不要让某个来源变成脱节的小复盘。",
              "- 选片前使用下方 source work_dir(`sources/<source_id>`)核对 scenes.json、ASR、索引和逐来源 brief。",
          ]
          if args.target_duration:
              lines.append(f"- 目标时长:`{args.target_duration}`。")
          lines += ["", "## Sources", ""]
          for s in source_records:
              swd = _source_work_dir(work_dir, s)
              lines += [
                  f"### {s['source_id']} — {s['source_name']}",
                  f"- path: `{s['source_path']}`",
                  f"- work_dir: `{swd}`",
                  f"- material_id: `{s['material_id']}`",
              ]
              excerpt = _brief_excerpt(swd / "agent_narration_brief.md")
              if excerpt:
                  lines += ["", "#### per-source brief excerpt", "", excerpt, ""]
          (Path(work_dir) / "agent_narration_brief.md").write_text(
              "\n".join(lines).rstrip() + "\n", encoding="utf-8"
          )
      
      
      def _source_speech_rows(source_dir):
          """The cleaned transcript wins when it carries text; same precedence as
          audio_mix._handoff_speech_evidence and sentence_boundaries._load_source_speech_spans."""
          clean = source_dir / "asr_clean.json"
          if clean.exists():
              rows = load_json(clean)["segments"]
              if any(row["text"].strip() for row in rows):
                  return rows
          raw = source_dir / "asr_result.json"
          return load_json(raw) if raw.exists() else []
      
      
      def _write_multi_source_output_speech_evidence(work_dir, source_records, plan):
          """Map each source's speech and quiet evidence onto the combined output clock."""
          source_by_id = {row["source_id"]: row for row in source_records}
          cache = {}
      
          def load_source(source_id):
              if source_id not in cache:
                  source_dir = _source_work_dir(work_dir, source_by_id[source_id])
                  anchors_path = source_dir / "speech_boundary_anchors.json"
                  quiet_path = source_dir / "silence_periods.json"
                  cache[source_id] = (
                      load_json(anchors_path)["sentence_anchors"] if anchors_path.exists() else [],
                      _source_speech_rows(source_dir),
                      load_json(quiet_path) if quiet_path.exists() else [],
                  )
              return cache[source_id]
      
          mapped_anchors, mapped_speech, mapped_quiet = [], [], []
          for clip in plan["clips"]:
              source_id = clip["source_id"]
              source_start = float(clip["source_start"])
              source_end = float(clip["source_end"])
              output_start = float(clip["output_start"])
              anchors, speech_rows, quiet_rows = load_source(source_id)
              for anchor in anchors:
                  when = float(anchor["time"])
                  if not (source_start - 0.05 <= when <= source_end + 0.05):
                      continue
                  pause = max(source_start, min(float(anchor["pause_start"]), when))
                  item = dict(anchor)
                  item.update(
                      source_id=source_id,
                      source_time=round(when, 3),
                      time=round(output_start + when - source_start, 3),
                      source_pause_start=round(pause, 3),
                      pause_start=round(output_start + pause - source_start, 3),
                  )
                  mapped_anchors.append(item)
              for rows, destination, require_text in (
                  (speech_rows, mapped_speech, True),
                  (quiet_rows, mapped_quiet, False),
              ):
                  for row in rows:
                      if require_text and not row["text"].strip():
                          continue
                      if not require_text and row["has_speech"]:
                          continue
                      start = max(source_start, float(row["start"]))
                      end = min(source_end, float(row["end"]))
                      if end <= start:
                          continue
                      item = dict(row)
                      item.update(
                          source_id=source_id,
                          source_start=round(start, 3),
                          source_end=round(end, 3),
                          start=round(output_start + start - source_start, 3),
                          end=round(output_start + end - source_start, 3),
                      )
                      destination.append(item)
      
          payload = {
              "schema_version": 2,
              "artifact": "speech_boundary_anchors_output.json",
              "timeline": "cut_output",
              "source_artifact": "multi_source_manifest.json",
              "sentence_anchors": sorted(mapped_anchors, key=lambda row: row["time"]),
              "speech_spans": sorted(mapped_speech, key=lambda row: (row["start"], row["end"])),
              "quiet_windows": sorted(mapped_quiet, key=lambda row: (row["start"], row["end"])),
          }
          Path(work_dir, "speech_boundary_anchors_output.json").write_text(
              json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8"
          )
          return payload
      
      
      def _write_multi_source_output_brief(work_dir, source_records, validated_plan_path):
          plan = load_json(validated_plan_path)
          source_by_id = {s["source_id"]: s for s in source_records}
          speech_evidence = _write_multi_source_output_speech_evidence(
              work_dir, source_records, plan
          )
          lines = [
              "# Multi-source Output Narration Brief",
              "",
              "现在 `edited_source.mp4` 已经由多个源视频剪好。请对剪后成片的 OUTPUT 时间轴写 `narration.json`。",
              "",
              "## 更新创作决定",
              "",
              "先查看 `edited_source.mp4` 与下方 kept-clip map,在 `visual_audio_board.json` 中补齐 OUTPUT 起止时间,并根据实际成片重新确认每拍的 `audio_owner`、原声锚点与 `narration_job`。如剪后顺序改变了主线或情绪路径,同时更新 `recap_story_plan.json`。",
              "",
              "beat 对应关系保留在视听板中;`narration.json` 仍只承载时间、文本与朗读参数,CLI 不声称校验计划映射。",
              "",
              "## narration.json 格式",
              "",
              "```json",
              '[{"start":0.0,"end":4.0,"narration":"解说文本。","pause_after_ms":250,"overlaps_speech":true,"emotion":"平静"}]',
              "```",
              "",
              "注意:`start`/`end` 是剪后成片时间,不是原视频时间。",
              "",
              "## 输出时间线写作规则",
              *MULTI_SOURCE_NARRATION_CRAFT_BULLETS,
              "- 旁白增加上下文、因果、预期、证据支持的解释或跨源过渡,不复述画面像素;先保人物与原声,再润色句子。",
              "",
              "## Kept clips (output → source)",
          ]
          for c in plan["clips"]:
              src = source_by_id[c["source_id"]]
              reason = f" — {c['reason']}" if c["reason"] else ""
              lines.append(
                  f"- output {_fmt_range(c['output_start'], c['output_end'])} → "
                  f"{c['source_id']} `{src['source_path']}` "
                  f"source {_fmt_range(c['source_start'], c['source_end'])}{reason}"
              )
          anchors = speech_evidence["sentence_anchors"]
          if anchors:
              lines += ["", "## 原声句末安全切入点"]
              lines.extend(
                  f"- {row['time']:.3f}s ({row['source_id']})"
                  for row in anchors
                  if row["confidence"] in {"high", "medium"}
              )
          lines += ["", "## Source work dirs"]
          for s in source_records:
              lines.append(f"- {s['source_id']}: `{_source_work_dir(work_dir, s)}`")
          (Path(work_dir) / "agent_narration_brief.md").write_text(
              "\n".join(lines).rstrip() + "\n", encoding="utf-8"
          )
      
      
      def _cut_qc_summary_line(qc):
          geometry = qc["output_geometry"]
          parts = [
              f"target_duration_status={qc['target_duration_status']}",
              f"total_duration={qc['total_duration']}",
              f"clip_count={qc['clip_count']}",
              f"join_fade_ms={qc['join_fade_ms']}",
              f"output_geometry={geometry['width']}x{geometry['height']}@{geometry['fps']}fps"
              f" reason={qc['output_geometry_reason']}",
          ]
          if qc.get("warnings"):
              parts.append(f"warnings={len(qc['warnings'])}")
          return "[video-recap] cut QC: " + "; ".join(parts)
      
      
      def _surface_cut_qc(work_dir):
          """Print the cut QC video-cut recorded; cut.py itself already fails on blocking QC."""
          qc = load_json(Path(work_dir) / "clip_plan_validated.json")["qc"]
          print(_cut_qc_summary_line(qc), flush=True)
          return qc
      
      
      def _fmt_range(start, end):
          return f"{start:.3f}-{end:.3f}s"
      
      
      def _material_library_dir(args):
          return args.material_library_dir or os.environ.get("VIDEO_RECAP_MATERIAL_LIBRARY_DIR") or None
      
      
      def _materials_enabled(args):
          return bool(_material_library_dir(args) and args.use_materials)
      
      
      def _save_materials_enabled(args):
          return bool(_material_library_dir(args) and args.save_materials)
      
      
      def _pause_for_agent(work_dir, need_text, cont, inspect_hint=None):
          brief = Path(work_dir) / "agent_narration_brief.md"
          print("=" * 50)
          if "Research the story FIRST" in brief.read_text(encoding="utf-8"):
              print(
                  "[video-recap] ⚑ 理解素材偏薄:先按 brief 顶部「Research the story FIRST」调研并写 "
                  "background_research.json,再写解说,避免看图说话。"
              )
          print(f"[video-recap] ⏸  阅读 {brief}(按 video-script 规则)后写入 {need_text}")
          if inspect_hint:
              print(f"[video-recap]    先核对状态/时间轴(建议性): {inspect_hint}")
          print(f"[video-recap]    写完后重跑继续: {cont}")
          print("=" * 50)
      
    • resource_lock.py 7 KB
      """Record every resource and template a finished run used: ``work_dir/resource_lock.json``.
      
      Facts come from artifacts the stages already wrote (run manifest, ``tts_meta.json``,
      ``assembly_manifest.json``) plus the optional project binding. When a library is
      configured, each file-backed entry is matched to a registered resource by resolved path
      so its licence and consent status travel with the run. Nothing here changes a render.
      """
      from __future__ import annotations
      
      import json
      from datetime import datetime, timezone
      from pathlib import Path
      
      import library as library_lib
      from materials import file_identity
      
      LOCK_NAME = "resource_lock.json"
      LOCK_SCHEMA = "video-recap.resource-lock.v1"
      
      
      def _read(path: Path):
          try:
              return json.loads(path.read_text(encoding="utf-8"))
          except (OSError, ValueError):
              return None
      
      
      def _file_entry(role: str, path, detail=None) -> dict:
          entry = {"role": role, "path": None, "size": None, "mtime_ns": None,
                   "detail": detail or {}, "library": None}
          if path:
              resolved = Path(path).expanduser().resolve()
              entry["path"] = str(resolved)
              if resolved.is_file():
                  entry.update(file_identity(resolved))
          return entry
      
      
      def _source_entries(work_dir: Path) -> list[dict]:
          manifest = _read(work_dir / "recap_run_manifest.json") or {}
          if manifest.get("mode") == "multi_source":
              return [_file_entry("source_video", s.get("source_path"), {"source_id": s.get("source_id")})
                      for s in manifest.get("sources", [])]
          if manifest.get("source_video"):
              return [_file_entry("source_video", manifest["source_video"])]
          return []
      
      
      def _voice_entry(work_dir: Path) -> dict | None:
          meta = _read(work_dir / "tts_meta.json")
          if not isinstance(meta, dict):
              return None
          voice = meta.get("voice") or {"provider": meta.get("engine")}
          reference = voice.get("reference") or {}
          entry = _file_entry("voice", reference.get("path"), {
              "provider": voice.get("provider"), "model": voice.get("model"),
              "voice_id": voice.get("voice_id"),
          })
          return entry
      
      
      def _assembly_entries(work_dir: Path) -> list[dict]:
          manifest = _read(work_dir / "assembly_manifest.json") or {}
          settings = manifest.get("assembly_settings") or {}
          entries = []
          bgm = (settings.get("audio_mix") or {}).get("bgm_path")
          if bgm:
              entries.append(_file_entry("bgm", bgm))
          packaging = (settings.get("video_filters") or {}).get("packaging_layers") or {}
          for layer in packaging.get("layers", []):
              entries.append(_file_entry("packaging_layer", layer.get("path"), {"name": layer.get("name")}))
          style = settings.get("subtitle_style")
          if isinstance(style, dict):
              entries.append(_file_entry("subtitle_font", style.get("font_file"),
                                         {"family": style.get("font_name")}))
          return entries
      
      
      def _match_library(entries: list[dict], index: dict) -> None:
          by_path = {}
          for resource in index["resources"]:
              for f in resource["files"]:
                  by_path[str(Path(f["path"]).resolve())] = resource
          voices = [r for r in index["resources"] if r["kind"] == "voice"]
          records = {r["id"]: json.loads(Path(r["record"]).read_text(encoding="utf-8"))
                     for r in index["resources"]}
          for entry in entries:
              resource = by_path.get(entry["path"]) if entry["path"] else None
              if (resource is None and entry["role"] == "voice" and not entry["path"]
                      and isinstance(entry["detail"].get("voice_id"), str) and entry["detail"]["voice_id"]):
                  detail = entry["detail"]
                  resource = next((v for v in voices
                                   if (records[v["id"]].get("voice") or {}).get("provider") == detail.get("provider")
                                   and (records[v["id"]].get("voice") or {}).get("voice_id") == detail.get("voice_id")), None)
              if resource is None:
                  continue
              record = records[resource["id"]]
              entry["library"] = {"id": resource["id"], "kind": resource["kind"],
                                  "license": resource["license"],
                                  "consent": (record.get("consent") or {}).get("status")}
      
      
      def _attention(entries: list[dict], library_configured: bool) -> list[dict]:
          items = []
          for entry in entries:
              role, lib = entry["role"], entry["library"]
              if role == "source_video":
                  continue
              if lib is None:
                  if library_configured and (entry["path"] or role == "voice"):
                      items.append({"code": "unregistered", "role": role,
                                    "message": "资源库中没有登记这项资源,授权状态未知"})
                  continue
              if lib["license"] in {"unknown", "restricted"}:
                  items.append({"code": f"license_{lib['license']}", "role": role,
                                "message": f"{lib['id']} 的授权状态为 {lib['license']},交付前需要人工确认"})
              if role == "voice" and entry["path"] and lib.get("consent") != "granted":
                  items.append({"code": f"consent_{lib.get('consent') or 'unknown'}", "role": role,
                                "message": f"{lib['id']} 的参考音频声音授权未确认"})
          return items
      
      
      def build_resource_lock(work_dir, *, library_dir=None, project=None) -> dict:
          """Assemble the lock payload from a finished work_dir; ``project`` is a resolved binding."""
          work_dir = Path(work_dir)
          entries = _source_entries(work_dir)
          voice = _voice_entry(work_dir)
          if voice:
              entries.append(voice)
          entries += _assembly_entries(work_dir)
          entries += list((project or {}).get("resources", []))
          library_root = None
          if library_dir and Path(library_dir).is_dir():
              library_root = str(Path(library_dir).resolve())
              index, _ = library_lib.scan_library(library_root)
              _match_library(entries, index)
          return {
              "schema": LOCK_SCHEMA,
              "generated_at": datetime.now(timezone.utc).replace(microsecond=0).isoformat().replace("+00:00", "Z"),
              "work_dir": str(work_dir.resolve()),
              "library": library_root,
              "project": ({"path": project["path"], "name": project.get("name", "")} if project else None),
              "templates": list((project or {}).get("templates", [])),
              "resources": entries,
              "attention": _attention(entries, library_root is not None),
          }
      
      
      def write_resource_lock(work_dir, *, library_dir=None, project=None) -> dict:
          lock = build_resource_lock(work_dir, library_dir=library_dir, project=project)
          path = Path(work_dir) / LOCK_NAME
          path.write_text(json.dumps(lock, ensure_ascii=False, indent=2), encoding="utf-8")
          return lock
      
      
      def summary_line(lock: dict) -> str:
          registered = sum(1 for e in lock["resources"] if e["library"])
          text = (f"[video-recap] 资源记录: {LOCK_NAME}({len(lock['resources'])} 项,"
                  f"资源库已登记 {registered} 项)")
          if lock["attention"]:
              text += "\n" + "\n".join(f"[video-recap]    ⚠ {a['role']}: {a['message']}" for a in lock["attention"])
          return text
      
  • SKILL.md 15.1 KB
    ---
    name: video-recap
    description: >
     从输入视频生成中文解说成片或原声剧情短片。用户提供 .mp4 / .mov / .mkv / .webm,并要求剪辑、添加旁白、
     配音、总结、短剧/电视剧/电影/纪录片/科普解说时使用。负责编排 video-* 技能链:视频理解 →
     Agent 制定故事与视听方案 → 剪辑 → 配音 → 合成。触发词:视频解说、视频旁白、生成解说、
     视频 recap、video recap、voiceover、narration、auto-dub、recap。
    ---
    
    ## 1. 定位与流程
    
    本技能是五个独立技能的轻量编排器。各技能只通过 `work_dir` 中的 JSON / MP4 产物通信,不共享代码:
    
    ```text
    video-understanding ─▶ Agent 按 video-script 制定方案并写稿 ─▶ [video-cut] ─▶ video-voiceover ─▶ video-assemble
    ```
    
    流程支持断点续跑:写好 `narration.json` 后重复同一条命令即可继续。第二阶段会比对
    `recap_run_manifest.json` 记录的源视频路径、文件大小/修改时间与运行参数,拒绝复用来自其他源视频或其他参数的旧工作目录;视频理解产物也只在来源一致时复用。
    
    画面流程 `--edit-mode full|cut|dub` 与声音策略 `--audio-mode` 是两个独立开关;
    `narration` 保留上述解说流程,`source-mix` 不做配音,`adopted-packet-copy` 冻结当前输入的已采用 AAC 音轨。
    组合只有下面几条路径,其余组合在启动时直接报错:
    
    | 输入 | `--edit-mode` | `--audio-mode` | Agent 暂停点 | 流程 | 详见 |
    |---|---|---|---|---|---|
    | 单视频 | full | narration | 1:`narration.json` | 理解 → 写稿 → 校验 → 配音 → 合成 | §4 |
    | 单视频 | cut | narration | 2:`clip_plan.json`,再对着成片写 `narration.json` | 理解 → 剪辑 → 重建输出时间 brief → 写稿 → 配音 → 合成 | §4 |
    | 多视频 | cut | narration | 2:同上,clip 必须带 `source_id` | 逐源理解 → 剪辑 → 写稿 → 配音 → 合成 | §4.3 |
    | 单视频 | full | source-mix / adopted-packet-copy | 无 | 直接合成当前整段 | §4.7 |
    | 单 / 多视频 | cut | source-mix / adopted-packet-copy | 1:`clip_plan.json` | 理解 → 剪辑 → 合成,不写稿 | §4.7 |
    | 单视频 | dub | narration | 1:`dub_script.json` | 英文转写 → 译稿 → 克隆音色整轨替换 | §5 |
    | 已剪好的母版 | full | narration + 三个采用 JSON | 无 | 只做严格合成 | 下文 |
    
    所有 full/cut 路径共用同一段收尾:(有旁白时)评审 → TTS → 合成 → 成片 QC。使用原声模式时读
    `references/audio-routing.md`。
    
    已有预制画面和本地采用的完整声音三件套时,可走严格 assembly-only 路径:
    
    ```bash
    python3 scripts/recap.py picture.mp4 --edit-mode full --work-dir NEW_WORK \
      --output-dir DELIVERY \
      --tts-meta tts_meta.json \
      --narration-adoption narration_adoption.json \
      --audio-mix-adoption audio_mix_adoption.json
    ```
    
    三个 JSON 参数必须同时出现。该入口只接受单视频、full、narration、音轨 0、新工作目录和未存在的
    交付文件;不运行理解、写稿、解说评审、TTS、cut、MiMo QC 或剪映导出。语义与媒体形状仍由
    video-assemble 严格验证,recap 只核对子技能绑定记录引用的是同一批采用文件与母版路径,不把调用方
    采用的声音或混音声明成自动创作或发布批准。详见 `references/audio-routing.md`。
    
    这里的单视频是**已经剪好的母版**。重剪后可以复用未改动的 WAV 与 `tts_meta.json`,但必须按新母版
    重新写混音采用文件里的落点与准备好的音床;衔接步骤见 `references/audio-routing.md` 的 “Keep adopted voice after a cut”。
    
    ## 2. 创作职责
    
    这不是单纯的 JSON / 渲染流水线。Agent 是本次内容的创作负责人。先判断本轮的**创作控制模式**(CREATE / DIRECTED / REVISION,与 `--edit-mode` 无关),再在进入昂贵的下游处理前完成五次判断:
    
    1. **导演判断**:观众承诺、POV、戏剧问题、情绪终点与揭示节奏。
    2. **故事编辑**:beat 定义为“发生了什么变化”,不是场景摘要。
    3. **画面剪辑**:选择真正值得保留的具体时刻、反应、入点与出点。
    4. **声音/旁白**:先分配画面、原声、沉默和旁白的任务,再写解说词。
    5. **观众复核**:分别检查无旁白、只听声音和第一次观看的体验。
    
    三种控制模式的定义、REVISION 的修改/冻结规则、创作方法以及 `recap_story_plan.json` / `visual_audio_board.json` / `style_card.json` 的写法,全部按 `video-script` 执行;它会要求先读创作手册。这些文件只记录可审计的当前决定,不增加服务或渲染依赖。
    
    ## 3. 环境与脚本路径
    
    ```bash
    # ffmpeg: brew install ffmpeg | apt install ffmpeg | choco install ffmpeg
    export MIMO_API_KEY=***
    ```
    
    同一个 MiMo key 驱动:
    
    - ASR:`mimo-v2.5-asr`
    - VLM:`mimo-v2.5`
    - TTS:`mimo-v2.5-tts`
    
    TTS 供应商由 `--tts-provider mimo-tts|fish-audio|index-tts`(或 `TTS_PROVIDER`)透传给配音技能;Fish Audio 与自托管 index-tts 各自的环境变量、默认音色和能力限制见该技能。ASR/VLM 始终使用 MiMo。`--doctor` 只做离线配置检查。
    
    `tp-*` Token Plan 密钥默认使用中国区集群,可用 `MIMO_TOKEN_PLAN_CLUSTER` 覆盖。
    
    可选能力:
    
    - `--mimo-video-overview`:按场景块补充 MiMo 视频理解。
    - `--mimo-qc pre-assemble|post-render|both`:在合成前、成片后或两个阶段给出建议型复核。
    
    MiMo QC 默认关闭;每个选定阶段最多请求一次,写入 `mimo_qc.json`。任何凭证缺失、限流、超时、格式错误或采样失败都只记录状态,不阻断流程。可覆盖配置见 `references/config-playbook.md`,QC 报告的最小契约见 `references/shift-left-qc-schema.md`。
    
    下面的 `scripts/...` 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。脚本启动后会自行定位兄弟技能和资源。
    
    ## 4. 标准解说流程(audio-mode narration)
    
    ### 4.1 背景调研
    
    若能识别影片、剧集或主题,先按 video-understanding 技能的调研指南 `research-guide.md` 调研并写入
    `work_dir/background_research.json`。视频理解会把人物名和剧情背景折入 VLM 上下文,避免只得到“黑衣男子”一类模糊描述。无法识别来源时可跳过。
    
    ### 4.2 分析并暂停创作
    
    ```bash
    python3 scripts/recap.py <video> --work-dir <work_dir> --context "背景"
    ```
    
    命令完成视频理解、写出 `agent_narration_brief.md`,然后暂停。此时按以下顺序执行 `video-script`:
    
    1. 查看创作 brief 与原片故事板。
    2. 写 `recap_story_plan.json` 和 `visual_audio_board.json`。
    3. full 模式写 `narration.json`;cut 模式第一阶段只写 `clip_plan.json`。
    4. cut 模式第二阶段查看剪后故事板,补充输出时间与声音分工,再写 `narration.json`。
    
    不要从标题或旁白句子开始;先锁定故事体验和素材选择。
    
    时间线有两条不可降级的硬约束:原声只能在可靠句末/静音边界被切入、切出或恢复;旁白必须使用
    完整逐段音频,任何 clip 映射裁段、TTS 裁尾或剪映引用更长的加速前素材都阻断。Agent 收到
    `interrupts_source_sentence` / `unsafe_clip_sentence_boundary` / `no_safe_fit` /
    `timeline_audio_mismatch` 时,应移动边界、缩短整句或删除该块,而不是增加抢断 override。
    
    ### 4.3 多视频与素材库
    
    多视频只支持 cut 模式。项目 brief 会列出稳定的 `source_id`,`clip_plan.json` 中每个片段都必须填写来源:
    
    ```bash
    python3 scripts/recap.py ep1.mp4 ep2.mp4 --edit-mode cut --target-duration 10m --work-dir work_dir_multi_ep
    ```
    
    可选文件系统素材库:
    
    ```bash
    python3 scripts/recap.py ep1.mp4 --material-library-dir .video-materials --save-materials
    python3 scripts/recap.py ep1.mp4 ep2.mp4 --edit-mode cut --material-library-dir .video-materials --use-materials
    ```
    
    素材检索只是对 JSON / MD / JSONL 做 grep,例如 `grep -R "keyword" .video-materials`。当前版本不复制原始媒体,也不提供数据库、向量或语义搜索。
    
    同一根目录还可以登记可复用的资源(BGM、音效、音色、字体、图片)、带版本与采用记录的模板(字幕样式、包装图层)和样片。
    格式与 `scripts/library.py check|list|show` 只读工具见 `references/resource-library.md`。用 `--project recap_project.json` 把已采用的字幕样式、音色与 BGM 绑定到这次运行;
    每次 full / cut 合成后 `work_dir/resource_lock.json` 记下实际用到的资源与授权状态。
    
    ### 4.4 继续生成成片
    
    写好所需产物后,重复同一条命令:
    
    ```bash
    python3 scripts/recap.py <video> --work-dir <work_dir>  # 可追加 --edit-mode cut / --no-burn-subtitles
    ```
    
    流程会校验当前阶段的硬输入(`clip_plan.json` / `narration.json`);两份创作计划仍是 Agent 与建议型评审使用的工作记录,不是渲染门禁。cut 模式随后生成 `edited_source.mp4`,再合成旁白并输出 `recap_<name>.mp4`。
    
    若需要建议型 MiMo 复核:
    
    ```bash
    python3 scripts/recap.py <video> --work-dir <work_dir> --mimo-qc both
    ```
    
    合成前复核会读取脚本、计划和 TTS 元数据;成片后还会读取最多六张临时 JPEG。输入文件的大小/修改时间、模型与提示都未变时直接复用上次报告,`--mimo-qc-refresh` 可强制刷新。帧的 base64 与凭证不会写入磁盘。
    
    已有批准解说稿时加 `--preserve-approved-text`:校验与 TTS 原样保留批准稿(只更新 `overlaps_speech`),装不下时间窗即失败,不缩稿、不降级为部分成功。
    
    ### 4.5 字幕与克隆旁白
    
    若要把旁白字幕固定在原片字幕区域,先在仓库根目录运行:
    
    ```bash
    python3 tools/measure_subtitle.py <video>
    ```
    
    再传入测得的 `--subtitle-y-top/--subtitle-y-bot`。坐标基于 ffmpeg 自动旋转后的显示画布,区间为半开 `[top, bot)`,并要求底对齐 ASS 样式;显式设置后,该区域默认使用 60% 透明度的旁白窗口遮罩。
    
    解说模式如需克隆参考声音,使用 `--voice-ref <audio>`;它与 dub 模式不同。
    
    ### 4.6 最终观看与交付复核
    
    脚本、接点检测、样帧和 QC 报告都不能替代观看。每轮准备交付前,必须检查**本轮实际要交付的最终文件**,而不是旧别名、无字幕母版或中间代理:
    
    1. 正常速度完整播放一次短片,不边看边改;先记录真实观看问题。
    2. 播放每个拼接点前后约 0.5–1 秒,检查闪帧、原片叠化被截断、动作跳变和半句原声。
    3. 完整只听声音一次,检查旁白是否碎成一句一停、场景间声音是否接得上、关键原声是否完整。
    4. 单独复看开头、核心情绪/表演点和结尾,确认进入时机、回报停留和收束都成立。
    5. REVISION 分别验证本轮修改项已经改变、冻结项没有意外变化;然后再做解码、时长、音画规格等机械检查。
    
    scene score、亮度统计、contact sheet 与自动 QC 只负责定位候选问题;最终判断以真实播放为准。密集切点的来源判断与处理规则按剪辑技能执行。修复失败时回到剪点、声音或文案层,不用更多包装掩盖。
    
    full/cut 交付如需让确定性的最终检查影响命令退出状态,显式传
    `--require-final-qc`。只有 `final_qc.json` 与 `golden_eval.json` 的摘要均为
    `ok: true` 且整数 `blocker_count: 0` 才打印完成并返回成功;缺失、畸形或 blocker
    会保留报告和已渲染诊断媒体,但命令非零退出且不打印完成。默认仍是仅报告、不阻断。
    该参数不支持 `--edit-mode dub`;dub 未传该参数时的准备和渲染行为不变。
    
    ### 4.7 不需要解说的片子
    
    ```bash
    # 对当前整段输入直接合成;不隐式跑理解/ASR/TTS
    python3 scripts/recap.py locked_picture.mp4 --work-dir source_work --audio-mode source-mix
    # 剪辑计划仍按 cut 流程产生,剪完不再暂停等待 narration.json
    python3 scripts/recap.py ep1.mp4 ep2.mp4 --edit-mode cut --work-dir cut_work --audio-mode source-mix
    # 只换包装时冻结当前整片 AAC;不允许同时加 BGM/TTS
    python3 scripts/recap.py adopted.mp4 --work-dir packaging_work --audio-mode adopted-packet-copy
    ```
    
    `source-mix` 仍会混音和重编码;`cut + adopted-packet-copy` 冻结的是剪后中间片的声音,不是原片的 AAC 包。
    当前严格字幕轨只支持 adopted 模式;其他字幕来源没有因此变成精确对齐。切换声音模式须新工作目录,不得把旧 TTS、QC 或自动生成的解说花字混入本轮原声生产。细节见 `references/audio-routing.md`。
    
    ## 5. 英译中原声复刻模式
    
    `--edit-mode dub` 把英文视频翻译为中文,并用原说话者的克隆音色替换人声;它不是在压低原声上叠加解说。
    
    ```bash
    python3 scripts/recap.py <video> --edit-mode dub --work-dir <work_dir>
    ```
    
    准备阶段会转写英文、提取一段参考音频,并写出 `dub_brief.md` 与 `dub_transcript.json`。Agent 随后写:
    
    ```json
    [{"start": 0.0, "end": 2.0, "zh": "中文译文"}]
    ```
    
    要求:
    
    - 逐句忠实翻译,不删钩子、不合并、不擅自压缩;原文重复,译文也按时间重复。
    - 每句沿用原声 `[start, end]`,相邻句不重叠。
    - 译文尽量控制在约 5 字/秒,使其能在原时间窗内说完。
    
    重复同一命令后输出 `dub_<name>.mp4`。每句单独克隆并贴回原时间线;只有即将覆盖下一句时才局部加速。当前版本只支持单说话者、整轨替换,不分离背景音乐。
    
    ## 6. 自检与只读 dashboard
    
    ```bash
    python3 scripts/recap.py --doctor
    ```
    
    ### 6.1 只读 dashboard
    
    ```bash
    python3 scripts/dashboard_server.py --root <目录> [--port 0] [--open]
    ```
    
    前台运行并打印本机地址(只绑定 127.0.0.1,Agent 启动时放到后台)。它在 `--root` 下按 `library.json`、`recap_project.json`、
    `recap_run_manifest.json` 发现资源库、项目与运行,按阶段显示剪辑节奏、旁白、成片与时间线、QC 和 `resource_lock.json`。
    严格只读:只接受 GET / HEAD,不写任何文件;页面上的「复制给助手」只复制一句请求,改动回到对话里做。
    
    ## 7. 输出与参数
    
    主要输出:
    
    - `recap_<video>.mp4`:最终成片。
    - `subtitles.srt` / `subtitles.ass`:字幕。
    - `work_dir/`:全部中间产物,契约见 `references/data-schema.md`。
    - `work_dir/recap_story_plan.json` / `visual_audio_board.json`:Agent 创作意图与剪辑决定。
    - `work_dir/mimo_qc.json`:可选的建议型复核,不作为发布门禁。
    
    完整参数列表以 `python3 scripts/recap.py --help` 为准。`--style` 是原样传给 Agent 的自由文本指导,不是 preset、枚举、开关或有限风格分类。
    
    ## 8. 能力边界
    
    - 语义评审默认建议型、失败开放;只有调用方显式启用严格解说评审时,事实矛盾、残句或评审不可用才会在 TTS 前阻断。确定性校验阶段始终负责硬校验。
    - MiMo QC 不能阻断、自动修复或改变退出状态,只提供定位建议。
    - 宣发标题、花字或外部文案回填见 `video-script` 的 references/promotional-copy.md。
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related