Claude Skill

video-understanding

把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief。 用于理解、索引或总结视频,也作为后续创作前的分析阶段。输入视频文件;输出 scenes.json、 asr_result.json、vlm_analysis.json、silence_periods.json、timeline_fusion.json、agent_narration_brief.md。 触发词:视频理解、视频分析、视频索引、video understanding、analyze video、看懂视频。

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download zenstory-ai-oh-story-dsh-packages_knowledge_video-recap_skills_video-understanding-d734089.zip · 110 KB
Part of zenstory-ai/oh-story-dsh — 31 skills

Install

skills CLI npx skills add https://github.com/zenstory-ai/oh-story-dsh/tree/main/packages/knowledge/video-recap/skills/video-understanding
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install zenstory-ai-oh-story-dsh@llmmart
Git git clone https://github.com/zenstory-ai/oh-story-dsh.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole zenstory-ai/oh-story-dsh collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

1. 定位

本技能把源视频转成 Agent 与下游阶段可读取的理解索引。它的创作角色是素材观察员 / 场记,不是导演:

  • 先观察,再解释;事实与推断分开。
  • 除了“发生了什么”,还要让下游看见知识、权力、目标、关系或情绪在哪一刻变化。
  • 标出由谁的 POV 承载变化、哪个反应或表演不可替代,以及哪里存在完整台词/动作的自然剪辑边界。
  • 证据不足时保留不确定性,不制造戏剧结论。

2. 处理阶段

  1. 场景检测:写 scenes.json,包含切点、时长和废片段过滤结果。
  2. 抽帧:为视觉分析提取代表帧。
  3. ASR:通过 mimo-v2.5-asr 写粗分段对白 asr_result.json,并写 asr_timing_evidence.json 说明可用性、有限时间精度与文本修正来源。
  4. 静音检测:写 silence_periods.json,标注安静窗口与 has_speech。
  5. VLM 观察:写 vlm_analysis.json,包含场景描述、深层分析和 frame_facts。
  6. 时间线融合与创作 brief:写 timeline_fusion.json、asr_writing_chunks.json 和 agent_narration_brief.md。

各阶段只有在输出产物与 provenance sidecar 同时匹配当前视频及影响结果的设置时才会复用;--force 强制重算。

3. 环境要求

# ffmpeg: brew install ffmpeg | apt install ffmpeg | choco install ffmpeg
export MIMO_API_KEY=***

ASR 使用 mimo-v2.5-asr;VLM 使用 mimo-v2.5。--skip-asr 可跳过对白转写,但完整理解仍需要 MIMO_API_KEY 运行 VLM。--mimo-video-overview 可开启按场景块的视频概览。

若 work_dir/background_research.json 存在,本技能会把剧情梗概和角色名折入 VLM 上下文;--context 可补充一条简短提示。

下面的 scripts/... 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。

4. 运行命令

python3 scripts/understand.py <video> --work-dir <work_dir> \
  [--context "节目名/角色名"] [--scene-threshold 0.1] [--skip-asr] [--mimo-video-overview] [--force]

5. 输出契约

文件 内容
scenes.json 场景切点、起止时间与时长
asr_result.json [{start, end, text}] 时间戳对白
asr_timing_evidence.json ASR 可用性状态、粗窗口精度、glossary 前后文本,以及它所描述的源视频/音频/结果文件(路径存在性 + size/mtime)
vlm_analysis.json 逐场景描述、深层分析与 frame_facts
silence_periods.json [{start, end, duration, has_speech}] 安静窗口
timeline_fusion.json VLM、ASR 与静音信息的统一时间线
asr_writing_chunks.json 按句界和场景切分的 ASR 写作块
agent_narration_brief.md Agent 首先阅读的创作简报

后续写作阶段根据创作简报与索引制定方案并写 narration.json。

6. 参考资料

  • 背景调研:references/research-guide.md,产出 background_research.json。
  • JSON 结构:references/data-schema.md。

7. 能力边界

  • 不写解说词,也不做解说评分;只负责生成理解索引与创作简报。
  • 不编造信号无法支持的剧情;当 ASR / VLM 过薄时输出素材警告。
  • MiMo ASR 的 start/end 是固定分片形成的粗窗口,不是词级对齐;空文本只表示原因未知, 不能当作已证实静音。asr_timing_evidence.json 的状态字段见 references/data-schema.md。
Files (oh-story-dsh)
  • references
    • data-schema.md 6.8 KB
      # 数据格式(中间 JSON)
      
      所有中间文件均在 pipeline 工作目录(work_dir/)下。
      
      ## vlm_analysis.json
      
      每场景的 VLM 分析结果,数组格式:
      
      ```json
      [
        {
          "scene_id": 1,
          "start": 5.0,
          "end": 15.0,
          "description": "男子闯入房间",
          "depth_analysis": "角色情绪分析...",
          "frame_facts": {
            "5.0": ["男子闯入房间, 头发蓬乱表情紧张"],
            "10.0": ["男子俯身盯着床上男孩, 男孩睁眼惊醒"]
          }
        }
      ]
      ```
      
      | 字段 | 类型 | 说明 |
      |------|------|------|
      | `scene_id` | int | 场景编号 |
      | `start` | float | 开始时间(秒) |
      | `end` | float | 结束时间(秒) |
      | `description` | string | 画面简述(≤80字) |
      | `depth_analysis` | string | 深层分析(情绪/关系/潜台词) |
      | `frame_facts` | object | 帧级事实,key 为时间戳字符串 |
      
      ## asr_result.json
      
      语音转文字结果:
      
      ```json
      [
        {"start": 0.0, "end": 3.5, "text": "What are you doing here?"}
      ]
      ```
      
      `start/end` 仅表示送入 MiMo ASR 的固定粗分片窗口。该列表为兼容产物,不能据此声称词级
      时间、准确对白边界或静音证明;空 `text` 的原因未知。
      
      ## asr_timing_evidence.json
      
      ASR 的独立证据 sidecar,不改变 `asr_result.json` 的既有数组结构。它记录自己描述的是哪一份源视频、
      `audio.wav`(存在时)和 `asr_result.json`(各自的 `size` + `mtime_ns`),并明确当前精度边界:
      
      ```json
      {
        "schema_version": 2,
        "status": "AVAILABLE_COARSE",
        "source_video": {"size": 123456, "mtime_ns": 1700000000000000000},
        "audio": {"size": 2048, "mtime_ns": 1700000001000000000},
        "asr_result": {"size": 512, "mtime_ns": 1700000002000000000},
        "glossary": {
          "names": ["叶轻眉"],
          "name_count": 1
        },
        "precision": {
          "window_timing": "COARSE_SEGMENT_WINDOWS",
          "dialogue_boundaries": "NOT_VERIFIED",
          "word_alignment": "NOT_PERFORMED",
          "empty_text_meaning": "UNKNOWN_NOT_PROVEN_SILENCE"
        },
        "windows": [{
          "index": 0,
          "start": 0.0,
          "end": 3.5,
          "text_availability": "AVAILABLE",
          "observed_text": "她叫叶青眉",
          "post_glossary_text": "她叫叶轻眉",
          "glossary_modified": true
        }]
      }
      ```
      
      `status` 可为 `AVAILABLE_COARSE`、`EXPLICITLY_SKIPPED`、`UNAVAILABLE_NO_KEY`、
      `UNAVAILABLE_NO_DURATION`、`FAILED_AUDIO_EXTRACTION`、`FAILED_PROVIDER`、`EMPTY_UNKNOWN`
      或 `LEGACY_UNVERIFIED`。`LEGACY_UNVERIFIED` 标记没有旧 sidecar 的兼容缓存,可离线复用但
      `observed_text`/`glossary_modified` 为 `null`,且始终保持 legacy 身份;非 legacy sidecar 记录
      当时参与修正的人名/别名列表(`glossary.names`),人名表变化或所描述的文件被重写(size/mtime 不再
      一致)都会使 ASR 缓存失效。
      `UNAVAILABLE_NO_DURATION` 与 `EMPTY_UNKNOWN` 是可重试的不可用结果,不作为缓存命中;写作
      brief 会校验 sidecar 并打印当前状态,缺失或与当前文件不一致时显示 `MISSING_OR_STALE`。
      
      ## asr_writing_chunks.json
      
      由 CLI 在生成 `agent_narration_brief.md` 时自动写出。它把长 ASR 按句子边界拆成适合 Agent 消化的语义块;中文按字符计数,非 CJK 文本按词数计数,并尽量保留 scene 对齐。
      
      ```json
      [
        {
          "chunk_id": 0,
          "start": 0.0,
          "end": 28.5,
          "scene_ids": [0, 1],
          "char_count": 642,
          "text": "第一段对白……",
          "segments": [
            {"start": 0.0, "end": 3.5, "text": "第一句。", "char_count": 4}
          ]
        }
      ]
      ```
      
      ## silence_periods.json
      
      静音窗口列表(适合放解说):
      
      ```json
      [
        {"start": 2.0, "end": 8.5, "duration": 6.5, "has_speech": false}
      ]
      ```
      
      `has_speech` 标记该窗口是否与检测到的 ASR 语音重叠;下游(pipeline / narration)只把 `has_speech=false` 的窗口当作可放解说的安静窗口。
      
      ## timeline_fusion.json
      
      由 CLI 在生成 brief 时自动写出。它把 VLM 场景、ASR 对白和静音窗口按时间轴 overlap 合并,减少写稿时手工推断“这一幕有没有对白/能不能插解说”的成本。
      
      ```json
      [
        {
          "scene_id": 0,
          "time_range": [0.0, 10.0],
          "visual_description": "两人在门口对峙",
          "depth_analysis": "关系紧张",
          "frame_facts": {"1.0": ["女子回头"]},
          "dialogue_segments": [
            {"start": 2.0, "end": 4.0, "overlap_seconds": 2.0, "text": "你到底是谁"}
          ],
          "dialogue_overlap_seconds": 2.0,
          "narration_slots": [
            {"start": 5.0, "end": 7.0, "duration": 2.0, "char_budget": 5}
          ],
          "recommended_mode": "ducked-bed"
        }
      ]
      ```
      
      ## deslop_qc_requirements.json(工具/brief 生成的运行契约)
      
      `deslop_qc_requirements.json` 是 tool/brief generated run contract:工具或 brief 生成本次运行的 QC 要求,供 `deslop_qc` 读取,不由 Agent 手写。字段为 `schema_version` 与 `style_card_required`。
      
      `style_card_required` 默认 `false`(advisory):缺少 `style_card.json` 只是 warning,不阻断出片。将来的 opt-in 运行可把它设为 `true`,让 `style_card.json` 成为硬性要求——`deslop_qc` 只读这个字段判断缺少 `style_card.json` 是否是 blocker,不扫描 `agent_narration_brief.md` 的 prompt wording 来推断。如果 requirements 文件缺失或损坏,按 legacy/migration advisory 处理,不作为 hard failure。
      
      该契约不改变 `--style`:`--style` 仍是 freeform verbatim guidance,不增加固定风格档位。它也不改变 `deslop_qc` 边界:仍然是 report-only,不是 AIGC detector,不自动改写。
      
      最小示例:
      
      ```json
      {
        "schema_version": 1,
        "style_card_required": false
      }
      ```
      
      ## background_research.json
      
      可选的背景调研结果(由 Agent 使用任意可用搜索/浏览方式整理):
      
      ```json
      {
        "synopsis": "剧情概要",
        "characters": {"角色名": "角色简介"},
        "worldbuilding": "世界观设定",
        "episode_context": "集数上下文",
        "character_details": {
          "角色名": {
            "aliases": ["别名/昵称"],
            "role": "主角|配角|反派|次要角色",
            "relationships": ["与XX是夫妻", "与YY是师徒"]
          }
        },
        "plot_arcs": [
          {"name": "线索名称", "description": "简要描述", "status": "进行中|已解决|伏笔"}
        ],
        "cultural_notes": [
          {"item": "文化梗/典故/时代背景", "explanation": "解释"}
        ]
      }
      ```
      
      > `character_details`、`plot_arcs`、`cultural_notes` 为(可选,新增)字段。仅含 `synopsis`、`characters`、`worldbuilding`、`episode_context` 四个原始字段的旧 JSON 仍然有效。
      
      ## 其他产物
      
      `narration.json`、`narration_lint.json`、`style_card.json`、`packaging_plan.json`、`deslop_qc.json`、`clip_plan.json`、`clip_plan_validated.json` 由后续的写稿与剪辑阶段读写,本技能既不生成也不校验它们;其格式以本技能生成的 `agent_narration_brief.md` 和编排器的中间产物契约为准。
      
    • prompt-templates.md 1.7 KB
      # Prompt 模板
      
      > 本文件被 `load_prompt()` 按 `### NAME` 切片读取。当前 CLI 只读取 VLM 模板;解说词由 Agent 根据 `agent_narration_brief.md` 撰写。
      
      ## 目录
      
      | 模板名 | 用途 | 使用阶段 |
      |--------|------|----------|
      | VLM_DEPTH_PROMPT | 场景画面深度分析 | VLM 分析 |
      
      ### VLM_DEPTH_PROMPT
      仔细观察这些视频帧,分析这个场景。分三部分输出:
      
      【描述】
      不超过80字,描述画面中正在发生什么(场景、人物、动作、表情、文字、事件)。
      
      【帧标签】
      逐帧列出每帧可见的人物动作和物品关系。每帧一行,格式:
      时间点 | 该帧可见的关键动作(1-2个完整短句,包含谁在做什么、物品如何被使用,如有字幕/文字也一并列出)
      示例:
      5.0s | 男子闯入房间, 头发蓬乱表情紧张
      10.0s | 男子俯身盯着床上男孩, 男孩睁眼惊醒
      15.0s | 男孩从床上坐起双手抱住头上的瓷枕准备防御
      
      【深层分析】
      不超过180字,分析:
      1. 角色的真实情绪或心理状态(通过表情、肢体语言、眼神判断)
      2. 人物之间的关系动态(亲密、紧张、疏离、试探等)
      3. 从场景开头到结尾发生了什么变化(知识、权力、目标、关系或情绪)
      4. 谁的视角最能承载这次变化;说话者、倾听者或旁观者中,谁的反应最值得保留
      5. 对话背后的潜台词(他们没说出口但由可见证据支持的东西)
      
      要求:
      - 必须基于画面可见的证据推断,不做无根据的猜测
      - 不要用"象征""隐喻"等文学词汇
      - 用口语化表达:"他明显是在试探她的反应"
      - 如果画面信息不足,直接说"信息不足"
      
    • research-guide.md 2.8 KB
      # 背景调研指南(可选)
      
      Agent 在开始视频理解**之前**先做背景调研,确认作品名、角色关系、剧情背景或科普概念,写入 `work_dir/background_research.json`。理解阶段会把它并入 VLM 的画面分析上下文,让场景描述能直接叫出人物名字、带着剧情读画面,而不是把人都标成「黑衣男子」。**调研不是必需步骤**;没有浏览器工具、网络不可用或用户未提供明确片名时,直接跳过,不阻塞视频生成。
      
      ## 工具选择
      
      使用当前环境里可用的任意检索方式即可,例如:
      
      - Codex / Claude 自带的联网搜索或网页浏览工具
      - 用户已经配置好的浏览器自动化或页面读取工具
      - 普通浏览器手动查找后把结果整理进 JSON
      - 已知资料、用户给的剧情背景或项目内已有文档
      
      ## 通用流程
      
      1. 从 `--context`、文件名或用户描述里提取可搜索关键词。
      2. 搜索 1-3 个高价值问题,不要过度调研。
      3. 只记录对解说有帮助的事实:人物关系、剧情前提、世界观、当前集上下文、术语解释。
      4. 写入 `work_dir/background_research.json`。
      5. 如果查不到可靠信息,跳过调研并继续视频理解;不要在理解索引生成前提前写 `narration.json`。
      
      ## 按视频类型搜索策略
      
      ### 短剧 / 电视剧
      
      | 搜索关键词 | 目标字段 |
      |-----------|---------|
      | `{作品名} 剧情 介绍 人物` | synopsis, characters |
      | `{作品名} 人物 关系` | characters |
      | `{作品名} 第{N}集 剧情` | episode_context |
      
      ### 电影
      
      | 搜索关键词 | 目标字段 |
      |-----------|---------|
      | `{电影名} 剧情 简介` | synopsis |
      | `{电影名} 影评 解读` | cultural_notes |
      
      ### 纪录片 / 科普视频
      
      | 搜索关键词 | 目标字段 |
      |-----------|---------|
      | `{主题} 背景 知识` | worldbuilding |
      | `{核心概念} 解释` | synopsis |
      | `{主题} 最新 进展` | cultural_notes |
      
      ## 写入格式
      
      写入 `work_dir/background_research.json`,只填搜到的字段,未搜到的字段省略,不要写 `null` 或空字符串。
      
      ```json
      {
        "synopsis": "...",
        "characters": {
          "角色名": "简介或关系"
        },
        "worldbuilding": "...",
        "episode_context": "...",
        "cultural_notes": [
          {"item": "...", "explanation": "..."}
        ]
      }
      ```
      
      ## 错误处理
      
      | 场景 | 处理 |
      |------|------|
      | 没有浏览器 / 搜索工具 | 跳过调研,继续视频理解 |
      | 搜索无结果 | 换一组关键词重试一次,仍无结果则继续视频理解 |
      | 结果互相矛盾 | 只使用用户提供的上下文或画面/ASR 可验证的信息 |
      | 信息可能剧透后续 | 只用于理解人物关系,不在解说里剧透用户没要求的后续剧情 |
      
      **原则**:背景资料只能补上下文,不能替代 `vlm_analysis.json` / `asr_result.json` 里的画面和对白证据。
      
  • scripts
    • briefing
      • builder.py 26.5 KB
        """Build the agent narration brief from validated local evidence."""
        
        from pathlib import Path
        
        from lib import CONFIG, log
        
        from agent_text import _chunk_asr_for_writing, _format_frame_facts
        from briefing.context import (
            _format_asr_chunks_for_brief,
            _format_background_research,
            _format_consolidation,
            _format_substrate_warning,
            _format_timeline_fusion_for_brief,
            _load_background_research,
            _load_consolidation,
            _write_json_artifact,
            assess_understanding_substrate,
        )
        from briefing.inputs import (
            _format_optional_stage_warnings,
            _load_clean_asr,
            _load_mimo_overview_for_brief,
        )
        from briefing.timeline import (
            _format_output_clip_list,
            _format_research_directive,
            _format_sentence_entry_anchors_for_brief,
            _load_cut_output_spans_for_brief,
            _parse_target_seconds,
            _remap_brief_evidence_to_output_timeline,
            _write_deslop_qc_requirements,
        )
        from timeline_fusion import (
            _build_timeline_fusion,
            _quiet_windows_for_scene,
            _scene_asr_lines,
        )
        
        def build_agent_brief(
            scenes_analysis,
            asr_result,
            silence_periods,
            video_duration,
            work_dir,
            style="纪录片",
            *,
            mimo_overview_enabled=None,
            mimo_overview_video_path=None,
            asr_evidence=None,
        ):
            """Write a compact brief that tells the agent exactly how to author recap artifacts."""
            # account for the global narration atempo (CONFIG['narration_speed']) so a beat's text
            # is budgeted against the FINAL sped-up audio, not the raw TTS rate — otherwise windows
            # are over-sized and the bed shows long silent gaps between sentences.
            effective_rate = (
                CONFIG["speech_rate"]
                * CONFIG["speech_safety_margin"]
                * CONFIG["narration_speed"]
            )
            breath_sec = CONFIG["breath_ms"] / 1000
            target_pause_ms = CONFIG["breath_ms"]
            edit_mode = CONFIG["edit_mode"]
            target_duration = CONFIG["target_duration"] or "(not set)"
            # Cut mode sizes narration to the OUTPUT (the kept clips), not the full source.
            # In pass 2, prefer the actual validated edited_source duration; target_duration is
            # only a planning goal and can differ after clip snapping or under/over-selection.
            output_seconds = video_duration
            if edit_mode == "cut":
                spans = _load_cut_output_spans_for_brief(work_dir, required=False)
                if spans:
                    output_seconds = max(span["output_end"] for span in spans)
                else:
                    target_seconds = _parse_target_seconds(CONFIG["target_duration"])
                    if target_seconds:
                        output_seconds = min(video_duration, target_seconds)
            # A loose first-draft timing fallback, not a creative quota. The Agent's beat map and audio-owner
            # decisions determine the real count; this only prevents accidental per-sentence fragmentation.
            cov_target = CONFIG["narration_coverage_target"]
            block_seconds = CONFIG["narration_block_seconds"]
            target_count = max(1, round(output_seconds * cov_target / block_seconds))
        
            has_story_context = (Path(work_dir) / "background_research.json").exists()
            substrate = assess_understanding_substrate(
                scenes_analysis, asr_result, has_story_context=has_story_context
            )
            thin_substrate = substrate["level"] in ("thin", "empty")
            beat_count_phrase = (
                f"at most ~{target_count}" if thin_substrate else f"roughly {target_count}"
            )
            output_label = (
                f"{output_seconds / 60:.0f}min"
                if output_seconds >= 60
                else f"{output_seconds:.0f}s"
            )
            source_label = (
                f"{video_duration / 60:.0f}min"
                if video_duration >= 60
                else f"{video_duration:.0f}s"
            )
        
            lines = [
                "# Agent Narration Brief",
                "",
                "Write the required JSON artifact(s) manually from the analysis files in this work directory.",
                "The CLI will not generate final narration text; it will only validate timing, run TTS, and assemble the video.",
                "",
                f"- Style (--style, freeform verbatim guidance): {style}",
                "- Do not translate `--style` into a preset, enum, switch, or fallback ladder; synthesize the voice from this freeform text plus evidence.",
                f"- Edit mode: {edit_mode}",
                f"- Source video duration: {video_duration:.1f}s",
                f"- Target duration (cut mode): {target_duration}",
                f"- Effective speech budget: {effective_rate:.2f} Chinese chars/sec after {breath_sec:.2f}s pause allowance",
            ]
            # Understanding validates evidence before calling; the standalone script skill
            # has no producer dependency and must not invent an available ASR status.
            if asr_evidence is None:
                asr_evidence = {"status": "MISSING_OR_STALE", "glossary_modifications": None}
        
            lines.extend(
                [
                    "",
                    "## ASR timing evidence",
                    "",
                    f"- Status: {asr_evidence['status']}",
                    f"- Glossary modifications: {asr_evidence.get('glossary_modifications') or '(unverified)'}; details: asr_timing_evidence.json",
                    "- Empty text: UNKNOWN_NOT_PROVEN_SILENCE; these windows locate a search region, not subtitle onsets.",
                    "- Timing: coarse provider windows; word alignment: NOT_PERFORMED; dialogue boundaries: NOT_VERIFIED.",
                    "- ASR window ends and timeline-fusion/quiet-window recommendations are advisory only. "
                    "They must not be treated as verified dialogue boundaries or safe cut points; use direct listening "
                    "and picture review before preserving or cutting original dialogue.",
                    "",
                ]
            )
            if thin_substrate:
                lines.append(
                    f"- Narration density: substrate is {substrate['level']} — do NOT chase a beat count. "
                    f"Write fewer, grounded blocks; skipping a stretch beats narrating pixels."
                )
            else:
                lines.append(
                    "- Content-led audio allocation. Assign every beat to picture/original dialogue/action sound/"
                    "ambience/music/silence/narration BEFORE writing prose. When narration owns a beat, write it as "
                    "one fluent BLOCK. A 7:3 narration/original split is only a rough fallback when the material gives "
                    "no clearer answer; it is not a quota or quality target."
                )
                lines.append(
                    "- Narration must have a job: context, causal_link, foreshadow, interpretation, or transition. "
                    "Use none when picture, original audio, or silence already carries the beat. Avoid per-sentence "
                    "stutter and wall-to-wall talk; size each authored block to its text and dramatic task."
                )
            if edit_mode == "cut":
                lines.append(
                    f"- Timing fallback only: {beat_count_phrase} narration BLOCKS across the ~{output_label} CUT OUTPUT "
                    f"(sized to the kept clips, NOT the {source_label} source). The beat map may justify fewer or more; never pad to hit this number."
                )
            else:
                lines.append(
                    f"- Timing fallback only: {beat_count_phrase} narration BLOCKS across the timeline. "
                    "The beat map and audio owners decide the real count; never pad to hit this number."
                )
            lines.extend(
                [
                    f"- Default pause between beats: {target_pause_ms}ms",
                    f"- Context: {CONFIG['context_info'] or '(none)'}",
                    "",
                    "## Creative decisions before expression / packaging",
                    "",
                    "First determine the creative-control mode from the current user request and existing artifacts: CREATE explores a new cut; DIRECTED implements user-specified structure/shots/lines; REVISION changes an existing viewed version while freezing everything not named in the latest feedback. This is separate from full/cut/dub.",
                    "- CREATE compares at least two real editorial hypotheses. DIRECTED/REVISION inherit the user-specified or accepted spine and must not invent alternatives to satisfy a template.",
                    "- In REVISION, state the current change list and frozen list first. Update `style_card.json` for expression/subtitle feedback and `visual_audio_board.json` for picture/timing/audio feedback; update `recap_story_plan.json` only when the promise, POV, spine, or story beats change. Remove stale descriptions when content is deleted.",
                    "Before `clip_plan.json` or final narration, author or update `recap_story_plan.json` and `visual_audio_board.json` from the applicable director intent, beat changes, and picture/audio decisions described below.",
                    "- `recap_story_plan.json` owns director intent, the chosen POV/spine, beats defined by changes in knowledge/power/goal/relationship/emotion/risk, and competing hypotheses only when exploration is appropriate.",
                    "- `visual_audio_board.json` owns the exact picture/performance/reaction, entry/exit reason, original-audio anchor, `audio_owner`, and `narration_job` for each beat.",
                    "- Only after the story/edit decisions are coherent, author `style_card.json` from `--style`, evidence, and user preference. It owns voice and pacing, not story structure, and is not a preset enum, fixed taxonomy, title plan, or packaging promise.",
                    "- `packaging_plan.json` is optional and deferred until content lock unless the user explicitly asks for packaging. It must express the story's truthful promise, never drive or distort it.",
                    "- `deslop_qc.json` is deterministic report-only tool QC: do not hand-author it, do not treat it as an AIGC detector, and do not auto-rewrite from it. Corrections remain human/agent rewrite work guided by objective blockers and advisory readability signals.",
                    "",
                ]
            )
        
            consolidation_index = _load_consolidation(work_dir, scenes_analysis)
            mimo_overview = _load_mimo_overview_for_brief(
                work_dir,
                scenes_analysis,
                enabled=mimo_overview_enabled,
                video_path=mimo_overview_video_path,
            )
            lines.extend(
                _format_optional_stage_warnings(
                    work_dir,
                    mimo_overview_enabled=mimo_overview_enabled,
                    mimo_overview=mimo_overview,
                    consolidation_index=consolidation_index,
                )
            )
            lines.extend(_format_substrate_warning(substrate))
            lines.extend(_format_research_directive(work_dir, substrate))
            lines.extend(_format_background_research(_load_background_research(work_dir)))
            lines.extend(_format_consolidation(consolidation_index))
        
            asr_for_chunks = _load_clean_asr(work_dir, asr_result) or asr_result
            chunk_scenes, chunk_asr = scenes_analysis, asr_for_chunks
            fusion_scenes, fusion_asr, fusion_silence = (
                scenes_analysis,
                asr_result,
                silence_periods,
            )
            if edit_mode == "cut" and (Path(work_dir) / "edited_source.mp4").exists():
                chunk_scenes, chunk_asr, _ = _remap_brief_evidence_to_output_timeline(
                    work_dir, scenes_analysis, asr_for_chunks, [], required=True
                )
                fusion_scenes, fusion_asr, fusion_silence = (
                    _remap_brief_evidence_to_output_timeline(
                        work_dir, scenes_analysis, asr_result, silence_periods, required=True
                    )
                )
            asr_chunks = _chunk_asr_for_writing(chunk_asr, chunk_scenes)
            timeline_fusion = _build_timeline_fusion(fusion_scenes, fusion_asr, fusion_silence)
            _write_json_artifact(work_dir, "asr_writing_chunks.json", asr_chunks)
            _write_json_artifact(work_dir, "timeline_fusion.json", timeline_fusion)
            lines.extend(_format_asr_chunks_for_brief(asr_chunks))
            lines.extend(_format_timeline_fusion_for_brief(timeline_fusion))
            lines.extend(_format_sentence_entry_anchors_for_brief(work_dir, edit_mode))
        
            overview_text = (mimo_overview or {}).get("content", "").strip()
            if overview_text:
                lines.extend(
                    [
                        "## MiMo scene-chunk video overview",
                        "",
                        overview_text[:2000],
                        "",
                    ]
                )
        
            if edit_mode == "cut":
                cut_target_example = (
                    target_duration if target_duration != "(not set)" else "30m"
                )
                if not (Path(work_dir) / "edited_source.mp4").exists():
                    # PASS 1 of 2 (cut-first): pick the footage. Narration comes AFTER the cut is
                    # rendered, so it can be written against the real OUTPUT timeline — no source->output
                    # mapping, no silent drop/clamp, no desync.
                    lines.extend(
                        [
                            "## Cut mode — step 1 of 2: direct the story, then write `clip_plan.json` ONLY",
                            "",
                            f"Goal: a ~{output_label} recap cut from a {source_label} source. In CREATE compare two editorial hypotheses;",
                            "in DIRECTED/REVISION inherit the requested or accepted spine. Then confirm the viewer promise/POV/dramatic question and write or update `recap_story_plan.json` + `visual_audio_board.json`.",
                            "Then choose the footage; the CLI",
                            "then renders the cut and asks you to narrate against that real output. Do NOT write narration.json yet.",
                            "How to choose clips (use the Scene timing guide + ASR + index below):",
                            "- Build ONE complete arc: a hook, the key turns of the plot, and a cliffhanger/payoff at the end — not a flat highlights reel.",
                            "- Keep clips that carry causality, a reveal, a decision, or a strong emotional beat; cut establishing/transition/repeated/static shots.",
                            "- SKIP non-story footage: 片头/片尾 credits, 演职员表, 广告/赞助, 台标/水印 stretches, and any scene the analysis marks rejected/无法描述 (often a watermark) — they look bad on screen and add nothing.",
                            "- Prefer scenes that have a real visual description in the guide; favor faces, action and dialogue over scenery.",
                            "- Clip order is the story spine, not unordered highlights: you may use 0–1 optional cold-open/high-impact clip first, then return to the coherent main arc (setup → turn → payoff) and escalate to the ending.",
                            "- Use `reason` to preserve the actual edit decision: `beat_id | function | change | POV | preferred moment | 入点 | 出点`, not merely 'important plot'. Function is `cold_open`, `setup`, `turn`, `escalation`, or `payoff`.",
                            "- Clip length follows the moment. Vary pace; after any cold-open, order clips by causality so the cut reads as one coherent story, not a flat highlights reel.",
                            "- Inspect dense scene-change candidates before locking boundaries. For source-authored cuts, delete irrelevant short shots and extend relevant shots to a complete action/reaction; for edit-created joins, move boundaries, restore same-source motion, or merge clips so the artificial cut disappears where possible. Do not hide a bad join with a transition.",
                            "- Preserve complete spoken lines, but do not infer completeness from ASR window ends or quiet-window suggestions. Directly listen and inspect picture around each proposed boundary, then place the cut after the verified utterance/reaction; automatic snapping is only an advisory candidate.",
                            "- Advisory does not mean absent: the cut CLI still snaps clip ends to sentence boundaries by default and blocks a cut that lands mid-sentence, using `silence_periods.json` and `speech_boundary_anchors.json` as the executable safety net under your listening.",
                            "",
                            "### clip_plan.json shape (original source timestamps)",
                            "",
                            "```json",
                            "{",
                            f'  "target_duration": "{cut_target_example}",',
                            '  "clips": [',
                            '    {"start": 12.0, "end": 38.0, "reason": "b01 | hook | knowledge: unknown→threat | POV=主角 | 保留倾听反应 | 入点=问题已问出 | 出点=沉默落地"}',
                            "  ]",
                            "}",
                            "```",
                        ]
                    )
                else:
                    # PASS 2 of 2: the cut is rendered (edited_source.mp4); narrate in OUTPUT time.
                    lines.extend(_format_output_clip_list(work_dir))
                    lines.extend(
                        [
                            "## Cut mode — step 2 of 2: write `narration.json` in OUTPUT time",
                            "",
                            "Inspect the edited storyboard first. Update `visual_audio_board.json` with OUTPUT ranges and re-assign `audio_owner` / `narration_job` based on the cut that actually exists.",
                            f"The cut is rendered as `edited_source.mp4` (~{output_label}). Write narration timed to THAT output",
                            "timeline (0 .. total), NOT the original source — your timestamps play exactly where you put them, with",
                            "no mapping and no dropping. Use the kept-clip OUTPUT ranges above to know what is on screen when, tell",
                            "one continuous arc across the cut, following the planned beat and audio-owner decisions.",
                            "",
                            "### narration.json shape (OUTPUT timestamps, 0..total)",
                            "",
                            "Each authored narration item is one fluent BLOCK; beats owned by picture, original audio, or silence",
                            "have no narration item. Size end-start to the text and preserve every planned original-audio anchor.",
                            "```json",
                            "[",
                            '  {"start": 2.0, "end": 13.0, "narration": "【主角】表面只是旁观者,暗中却握着关键线索。这一次,TA要赌上全部去查清旧案真相。", "pause_after_ms": 250, "overlaps_speech": true, "emotion": "紧张", "source_entry_policy": "sentence_boundary"}',
                            "]",
                            "```",
                        ]
                    )
            else:
                lines.extend(
                    [
                        "## Required JSON shape",
                        "",
                        "Before narration, write `recap_story_plan.json` and `visual_audio_board.json`, then use the board to decide which planned beats need narration.",
                        "beat 对应关系记录在 `visual_audio_board.json`;`narration.json` 仍只承载时间、文本与朗读参数,不声明 CLI 会校验计划映射。",
                        "Each authored narration item is one fluent BLOCK; beats owned by picture, original audio, or silence have no narration item.",
                        "```json",
                        "[",
                        '  {"start": 5.0, "end": 16.0, "narration": "【主角】表面只是旁观者,暗中却握着关键线索。这一次,TA要赌上全部去查清旧案真相。", "pause_after_ms": 250, "overlaps_speech": true, "emotion": "平静", "source_entry_policy": "sentence_boundary"}',
                        "]",
                        "```",
                    ]
                )
        
            lines.extend(
                [
                    "",
                    "## Writing rules (creative decisions first; blocks are delivery form)",
                    "",
                    "1. Assign `audio_owner` and `narration_job` before prose. No clear narration job means no narration for that beat.",
                    "2. When narration owns a beat, write one fluent BLOCK that completes a continuous thought (often premise -> trigger/action -> change/meaning) for one TTS call. Sentence count is not a target: use one or a few complete connected sentences, and never split TTS merely because captions need shorter cues.",
                    "3. 7:3 is a rough fallback, never a coverage quota. A strong dialogue/performance/action/silence beat may contain no narration; an exposition bridge may be narration-led.",
                    "4. Default `overlaps_speech` to true only for authored narration windows. If source speech has already started, use a sentence end confirmed by direct listening and picture review; ASR coarse ends are not proof. Do not cover a must-hear original-audio anchor; leave that beat un-narrated or place narration around it.",
                    "5. Do not describe what the viewer can already see; narration may add context, causal links, foreshadowing, evidence-grounded interpretation, or transitions.",
                    "6. Keep timing visually local: anchor each block to the planned beat and exact footage it covers; don't let prose run past the change it explains.",
                    "7. Preserve performance: consider the listener/reaction instead of the speaker/action, and leave enough time for an irreplaceable look, pause, mistake, or action sound to land.",
                    "8. Give every block an `emotion` that fits its whole arc; keep it STEADY across the block and shift only at a real emotional turn between blocks.",
                    "9. Bridge narration and original audio as one dramatic beat: tee up a must-hear moment before it plays, then let the next block react to the change it caused.",
                    "10. 不要在解说文本里使用破折号(——、—):破折号烧进字幕里很突兀,该停顿就用逗号,该断句就用句号;同理 `original_subtitles.json` 里也不要用破折号。",
                    "11. `不是 A,而是 B` is useful only when the material genuinely corrects a plausible prior belief. Do not invent a false contrast for insight; prefer direct action, causality, and consequence.",
                    "12. 写完后停止并把控制权交还调用方;本技能只写工作产物,不调用其他技能脚本。",
                    "",
                    "## 原声留白字幕 `original_subtitles.json`(校对原声台词)",
                    "",
                    '你在解说块之间留出的原声留白,会把【原声台词】烧成字幕(和解说字幕用 「」 区分开)。请额外写一个 `original_subtitles.json`,把每段留白里真正听得到的原声台词,按 OUTPUT 时间轴写成 `[{"start": 秒, "end": 秒, "text": "台词"}]`:',
                    "- 只写留白里【实际出声】的台词;被解说盖过、或已经被剪掉的句子,不要写进来(这正是自动 ASR 兜底会出错的地方)。",
                    "- 订正 ASR 的错字和人名(例:叶青眉 → 叶轻眉),删掉口胡和语气词。",
                    "- 每条短到一行(≤ 约 20 字),`start`/`end` 对齐它在留白里出声的时间;拿不准就贴着所在留白的区间写。",
                    "- 没有清晰原声的留白可以不写;整个文件也可省略——省略时系统会用 ASR 粗略兜底(可能偏多偏乱)。",
                    "",
                    "## Per-block emotion (`emotion` field → MiMo TTS instruct)",
                    "",
                    "Each block's `emotion` is a short Chinese tone tag MiMo-v2.5-tts follows for the whole utterance. Pick 1-2 that fit the block:",
                    "- 基础情绪: 开心 悲伤 愤怒 恐惧 惊讶 兴奋 委屈 平静 冷漠",
                    "- 复合情绪: 怅然 欣慰 无奈 愧疚 释然 嫉妒 厌倦 忐忑 动情",
                    "- 整体语调: 温柔 高冷 活泼 严肃 慵懒 俏皮 深沉 干练 凌厉",
                    'You may combine, e.g. "紧张 深沉" or "无奈". Default to 平静 only for neutral setup; a recap mostly lives in 紧张/深沉/惊讶/悲伤/动情.',
                    "",
                    "## Recap craft (what separates a real recap from captions)",
                    "",
                    "- Hook: the opening must create a truthful dramatic question or stakes that this edit actually pays off; do not manufacture an unrelated retention line.",
                    "- Through-line: follow the chosen POV/spine from `recap_story_plan.json`; every beat must change knowledge, power, goal, relationship, emotion, or risk.",
                    "- Escalation: raise the stakes or reveal new information as you go; later beats should land harder than earlier ones.",
                    '- Curiosity gaps: tease consequences before they happen ("他还不知道,这一步会要命") and pay them off later.',
                    "- Payoff: the final 1-2 beats must resolve or twist the spine, leaving an aftertaste — never trail off on a generic line.",
                    "- Information, not narration of pixels: every beat should add something the picture alone can't tell (who, why, what's at stake).",
                    '- Voice: concrete nouns and verbs, specific names; cut adjectives and vague grandeur ("危机四伏"/"震撼人心" are filler).',
                    "- Use the real names, relationships and stakes from the Story context / index above — never generic labels like 男子/白衣女子.",
                    "- Show motive and consequence, not actions: say WHY a character does it and what it costs, not what they are doing on screen.",
                    "- Performance: when the reaction carries more emotion than the line, let the reaction own the picture and avoid explaining it away.",
                    "- 衔接 hand-off: a narration block and the original-audio gap beside it are ONE beat — tee up the original before it plays, then have the next block answer what it showed; never let a block end self-contained and the original come in cold.",
                    "- Counterfactual review before TTS: remove each beat, compare speaker vs listener, mute narration, listen audio-only, and replace prose with original audio/silence where that is stronger. Apply the 1-3 highest-return changes first.",
                    "",
                    "看图说话 (bad) vs recap (good) — same shot:",
                    '- ✗ "一个蒙眼的男人抱着一个篮子走在雨里。"  (just describes the frame)',
                    '- ✓ "护送者本可以独自离开,却为了保护那个孩子,主动把追兵引向自己。"  (who, why, stakes)',
                    "",
                    "## Scene timing guide",
                    "",
                ]
            )
        
            for scene in scenes_analysis:
                duration = scene["end"] - scene["start"]
                max_chars = max(5, int(max(1.0, duration - breath_sec) * effective_rate))
                quiets = _quiet_windows_for_scene(silence_periods, scene)
                quiet_text = ", ".join(f"{s:.1f}-{e:.1f}s" for s, e in quiets) or "none"
                lines.extend(
                    [
                        f"### Scene {scene['scene_id'] + 1}: {scene['start']:.1f}-{scene['end']:.1f}s",
                        f"- Duration: {duration:.1f}s; max budget if fully narrated: {max_chars} chars",
                        f"- Quiet windows: {quiet_text}",
                        f"- Description: {scene.get('description', '')}",
                    ]
                )
                if scene.get("depth_analysis"):
                    lines.append(f"- Deeper analysis: {scene['depth_analysis']}")
                facts = _format_frame_facts(scene)
                if facts:
                    lines.append(facts.rstrip())
                asr_lines = _scene_asr_lines(asr_result, scene)
                if asr_lines:
                    lines.append("- ASR overlap:")
                    lines.extend(asr_lines[:8])
                lines.append("")
        
            brief_path = Path(work_dir) / "agent_narration_brief.md"
            brief_path.write_text("\n".join(lines), encoding="utf-8")
            _write_deslop_qc_requirements(work_dir)
            log(f"已写入 Agent 解说写作 brief: {brief_path}")
            return brief_path
        
      • context.py 11.6 KB
        """Load and format research, consolidation, and substrate context."""
        
        import json
        import re
        from pathlib import Path
        
        from lib import CONFIG, file_identity
        
        
        def _load_background_research(work_dir):
            """Load the agent-authored background_research.json ({} when absent)."""
            path = Path(work_dir) / "background_research.json"
            if not path.exists():
                return {}
            return json.loads(path.read_text(encoding="utf-8"))
        
        
        def _clip_text(text, limit):
            """Collapse whitespace and clip. Research fields are optional, so None clips to ""."""
            return re.sub(r"\s+", " ", str(text or "")).strip()[:limit]
        
        
        def _format_background_research(research, limit=1800):
            """Render background_research.json into a bounded Story-context brief section."""
            if not research:
                return []
            lines = [
                "## Story context (from background_research.json)",
                "",
                "Research is context_only: use it for aliases/names/world terms and weak background only; do NOT reveal future plot or upgrade research-only relationships/causes into current on-screen facts without visual/ASR support.",
                "",
            ]
            for key, label in (
                ("synopsis", "Synopsis"),
                ("worldbuilding", "Worldbuilding"),
                ("episode_context", "Episode context"),
            ):
                value = _clip_text(research.get(key), 500)
                if value:
                    lines.append(f"- {label}: {value}")
        
            characters = research.get("characters", {})
            if characters:
                lines.append("- Characters:")
                for name, desc in list(characters.items())[:12]:
                    lines.append(f"    - {_clip_text(name, 60)}: {_clip_text(desc, 160)}")
        
            details = research.get("character_details", {})
            if details:
                lines.append("- Character details:")
                for name, info in list(details.items())[:8]:
                    bits = []
                    aliases = [_clip_text(alias, 40) for alias in info.get("aliases", [])[:4]]
                    if aliases:
                        bits.append("别名 " + "/".join(aliases))
                    role = _clip_text(info.get("role"), 80)
                    if role:
                        bits.append(role)
                    rels = [_clip_text(rel, 80) for rel in info.get("relationships", [])[:4]]
                    if rels:
                        bits.append(";".join(rels))
                    if bits:
                        lines.append(f"    - {_clip_text(name, 60)}: {'; '.join(bits)}")
        
            arcs = research.get("plot_arcs", [])
            if arcs:
                lines.append("- Plot arcs:")
                for arc in arcs[:8]:
                    name = _clip_text(arc.get("name"), 80)
                    desc = _clip_text(arc.get("description"), 180)
                    status = _clip_text(arc.get("status"), 40)
                    if name or desc:
                        tail = f" [{status}]" if status else ""
                        lines.append(f"    - {name}: {desc}{tail}".rstrip())
        
            notes = research.get("cultural_notes", [])
            if notes:
                lines.append("- Cultural notes:")
                for note in notes[:6]:
                    item = _clip_text(note.get("item"), 80)
                    expl = _clip_text(note.get("explanation"), 160)
                    if item and expl:
                        lines.append(f"    - {item}: {expl}")
                    elif item or expl:
                        lines.append(f"    - {item or expl}")
        
            lines.extend(
                [
                    "",
                    'Use these names, relationships, and stakes in the narration instead of generic labels like "男子"/"白发女子".',
                    "",
                ]
            )
            text = "\n".join(lines)
            if len(text) <= limit:
                return lines
            clipped = text[:limit].rsplit("\n", 1)[0].rstrip()
            return [
                *clipped.splitlines(),
                "",
                "[Story context clipped to keep ASR/visual evidence in context]",
                "",
            ]
        
        
        def assess_understanding_substrate(
            scenes_analysis, asr_result, *, has_story_context=False
        ):
            """Measure how much real signal the writing agent has to work with.
        
            Recap quality collapses to generic 看图说话 when the agent has no story spine and
            only literal frame descriptions to paraphrase. A spine is substantial dialogue
            (ASR) OR researched/given story context — frame-fact VOLUME alone (a visually busy
            but storyless clip, e.g. an anime whose dialogue ASR could not read) is NOT a spine,
            so it must not grade as "rich"; otherwise the sparse-substrate warning, the research
            directive, and the density relief never fire for exactly that cold-narration case.
            """
            asr_chars = sum(len(seg["text"]) for seg in asr_result)
            scenes_with_facts = sum(1 for s in scenes_analysis if s.get("frame_facts"))
            desc_lens = [len(s["description"]) for s in scenes_analysis]
            avg_desc = sum(desc_lens) // len(desc_lens) if desc_lens else 0
        
            has_asr = asr_chars >= 20
            has_facts = scenes_with_facts > 0
            has_story_spine = asr_chars >= 200 or bool(has_story_context)
            if not has_asr and not has_facts and avg_desc < 25:
                level = "empty"
            elif (has_asr or has_facts) and has_story_spine:
                level = "rich"
            else:
                level = "thin"
            return {
                "level": level,
                "asr_chars": asr_chars,
                "scene_count": len(scenes_analysis),
                "scenes_with_frame_facts": scenes_with_facts,
                "avg_description_len": avg_desc,
                "has_story_context": bool(has_story_context),
            }
        
        
        def _format_substrate_warning(assessment):
            """Render a loud brief banner when the understanding substrate is weak."""
            if assessment["level"] == "rich":
                return []
            if assessment["level"] == "empty":
                head = "⚠️ UNDERSTANDING SUBSTRATE IS EMPTY — narration will be generic guesswork unless you fix this first."
            else:
                head = '⚠️ Understanding substrate is THIN — narration risks generic "看图说话" without more grounding.'
            return [
                head,
                f"  ASR chars: {assessment['asr_chars']} | scenes: {assessment['scene_count']} | "
                f"scenes with frame_facts: {assessment['scenes_with_frame_facts']} | avg description: {assessment['avg_description_len']} chars",
                "  Before writing: do background research (write background_research.json with names/relationships/plot),",
                "  lean on any --context provided, and use ASR dialogue + frame_facts as the factual spine.",
                "  Do NOT invent plot. If you truly have nothing, keep beats sparse and factual rather than fabricating drama.",
                "",
            ]
        
        
        def _write_json_artifact(work_dir, name, payload):
            path = Path(work_dir) / name
            path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
            return path
        
        
        def _format_asr_chunks_for_brief(chunks, max_chunks=24):
            if not chunks:
                return []
            lines = [
                "## ASR writing chunks (semantic windows)",
                "",
                "Use these chunks as the dialogue spine instead of swallowing the full transcript at once.",
                "Full data: `asr_writing_chunks.json`.",
                "",
            ]
            for chunk in chunks[:max_chunks]:
                scene_ids = ",".join(str(sid) for sid in chunk["scene_ids"]) or "n/a"
                text = chunk["text"]
                if len(text) > 900:
                    text = text[:897] + "..."
                lines.extend(
                    [
                        f"### ASR chunk {chunk['chunk_id'] + 1}: {chunk['start']:.1f}-{chunk['end']:.1f}s | scenes {scene_ids} | {chunk['char_count']} units",
                        text or "(empty transcript chunk)",
                        "",
                    ]
                )
            if len(chunks) > max_chunks:
                lines.append(
                    f"... {len(chunks) - max_chunks} more chunks in `asr_writing_chunks.json`."
                )
                lines.append("")
            return lines
        
        
        def _format_timeline_fusion_for_brief(fusion, max_items=40):
            if not fusion:
                return []
            lines = [
                "## Timeline fusion (VLM + ASR + quiet slots)",
                "",
                "This is the pre-aligned multimodal view. Use `narration_slots` when available; otherwise write ducked-bed beats around dialogue.",
                "Full data: `timeline_fusion.json`.",
                "",
            ]
            for item in fusion[:max_items]:
                start, end = item["time_range"]
                slot_text = (
                    ", ".join(
                        f"{slot['start']:.1f}-{slot['end']:.1f}s/{slot['char_budget']}字"
                        for slot in item["narration_slots"][:4]
                    )
                    or "none"
                )
                dialogue_text = (
                    "; ".join(
                        f"{seg['start']:.1f}-{seg['end']:.1f}s {seg['text'][:80]}"
                        for seg in item["dialogue_segments"][:3]
                        if seg["text"]
                    )
                    or "none"
                )
                lines.extend(
                    [
                        f"### Fusion scene {item['scene_id']}: {start:.1f}-{end:.1f}s ({item['recommended_mode']})",
                        f"- Visual: {item['visual_description']}",
                        f"- Dialogue overlap: {item['dialogue_overlap_seconds']:.1f}s | {dialogue_text}",
                        f"- Narration slots: {slot_text}",
                        "",
                    ]
                )
            if len(fusion) > max_items:
                lines.append(
                    f"... {len(fusion) - max_items} more fused scenes in `timeline_fusion.json`."
                )
                lines.append("")
            return lines
        
        
        def _consolidation_model():
            return CONFIG["vlm_model"]
        
        
        def _load_consolidation(work_dir, scenes_analysis):
            """Load consolidate.py's understanding_index.json only when its provenance matches
            the current vlm_analysis.json and model ({} otherwise)."""
            work_dir = Path(work_dir)
            path = work_dir / "understanding_index.json"
            meta_path = work_dir / "understanding_index.json.meta.json"
            if not path.exists() or not meta_path.exists():
                return {}
            try:
                meta = json.loads(meta_path.read_text(encoding="utf-8"))
                source = file_identity(work_dir / "vlm_analysis.json")
            except (OSError, json.JSONDecodeError):
                return {}
            if not isinstance(meta, dict):
                return {}
            expected = {
                "source": source,
                "model": _consolidation_model(),
            }
            if scenes_analysis:
                expected["scene_count"] = len(scenes_analysis)
            if not meta.items() >= expected.items():
                return {}
            try:
                index = json.loads(path.read_text(encoding="utf-8"))
            except (OSError, json.JSONDecodeError):
                return {}
            if not isinstance(index, dict) or not all(
                isinstance(index.get(key), list)
                for key in ("characters", "relationships", "plot_points", "entities")
            ):
                return {}
            if not all(isinstance(item, dict) for item in index["characters"]):
                return {}
            if not all(isinstance(item, dict) for item in index["relationships"]):
                return {}
            return index
        
        
        def _format_consolidation(index):
            """Render consolidate.py's understanding_index.json into a compact brief section.
        
            plot_points/entities items are model output and may be bare strings or objects."""
            if not index:
                return []
            lines = ["## Understanding index (from consolidate.py)", ""]
            chars = index["characters"]
            if chars:
                lines.append("- Characters:")
                for c in chars:
                    lines.append(
                        f"    - {c.get('name', '?')}: {str(c.get('description', '')).strip()}"
                    )
            rels = index["relationships"]
            if rels:
                lines.append("- Relationships:")
                for r in rels:
                    lines.append(
                        f"    - {r.get('a', '?')} — {r.get('relation', '?')} — {r.get('b', '?')}"
                    )
            plot = index["plot_points"]
            if plot:
                lines.append("- Plot spine:")
                lines.extend(
                    f"    {i + 1}. {p.get('text', p) if isinstance(p, dict) else p}"
                    for i, p in enumerate(plot)
                )
            ents = index["entities"]
            if ents:
                ent_names = [str(e.get("name", e) if isinstance(e, dict) else e) for e in ents]
                lines.append(f"- Entities: {', '.join(ent_names)}")
            lines.append("")
            return lines
        
      • inputs.py 9.2 KB
        """Validate optional ASR, MiMo, and stage-status inputs for the brief."""
        
        import json
        from pathlib import Path
        
        from lib import CONFIG, file_identity
        from briefing.context import _consolidation_model
        
        # Shared with consolidate.py, which this byte-identical copy cannot import (the sibling
        # skill ships no consolidate.py). Keep both literals in sync.
        _ASR_SPAN_TOL = 0.05
        
        _MIMO_REJECTION_MARKERS = (
            "request was rejected",
            "considered high risk",
            "high risk",
            "content policy",
            "cannot process",
            "无法处理",
            "内容审核",
            "违规",
        )
        
        
        def _load_clean_asr(work_dir, asr_result):
            """Return consolidate.py's cleaned ASR segments, or None to fall back to raw asr_result.
        
            Accepted only when the file is at least as fresh as asr_result.json, its provenance
            (source identity / model) matches, and every segment keeps its original span
            within _ASR_SPAN_TOL."""
            if not asr_result:
                return None
            work_dir = Path(work_dir)
            clean_path = work_dir / "asr_clean.json"
            src_path = work_dir / "asr_result.json"
            try:
                fresh = (
                    clean_path.exists()
                    and clean_path.stat().st_mtime >= src_path.stat().st_mtime
                )
            except OSError:
                return None
            if not fresh:
                return None
            try:
                payload = json.loads(clean_path.read_text(encoding="utf-8"))
                source = file_identity(src_path)
            except (OSError, json.JSONDecodeError):
                return None
            if not isinstance(payload, dict):
                return None
            provenance = {"source": source, "model": _consolidation_model()}
            if not payload.items() >= provenance.items():
                return None
            segments = payload.get("segments")
            if not isinstance(segments, list) or len(segments) != len(asr_result):
                return None
            for orig, clean in zip(asr_result, segments):
                if (
                    not isinstance(clean, dict)
                    or not isinstance(clean.get("start"), (int, float))
                    or not isinstance(clean.get("end"), (int, float))
                ):
                    return None
                if (
                    abs(clean["start"] - orig["start"]) > _ASR_SPAN_TOL
                    or abs(clean["end"] - orig["end"]) > _ASR_SPAN_TOL
                ):
                    return None
            return segments
        
        
        def _is_mimo_chunk_usable(content):
            """A chunk is usable only if MiMo returned real analysis (not empty / a moderation refusal)."""
            text = (content or "").strip()
            if not text:
                return False
            low = text.lower()
            return not any(marker in low for marker in _MIMO_REJECTION_MARKERS)
        
        
        def _mimo_video_settings():
            """Non-secret MiMo video-overview settings that affect generated content."""
            return {
                "model": CONFIG["mimo_video_model"],
                "mimo_video_api_url": CONFIG["mimo_video_api_url"],
                "mimo_video_fps": CONFIG["mimo_video_fps"],
                "mimo_media_resolution": CONFIG["mimo_media_resolution"],
                "mimo_video_chunk_max_seconds": CONFIG["mimo_video_chunk_max_seconds"],
                "mimo_video_chunk_min_seconds": CONFIG["mimo_video_chunk_min_seconds"],
                "mimo_video_base64_max_mb": CONFIG["mimo_video_base64_max_mb"],
                "mimo_video_prompt": CONFIG["mimo_video_prompt"],
                "mimo_disable_thinking": CONFIG["mimo_disable_thinking"],
            }
        
        
        def _mimo_chunk_cache_key(chunk):
            """Stable identifier for a MiMo chunk (index + scene span) for partial-cache reuse."""
            return f"{chunk['chunk_id']}|{chunk['scene_id']}|{chunk['start']:.3f}-{chunk['end']:.3f}"
        
        
        def _mimo_video_chunks(scenes):
            """Split scene spans into MiMo video chunks. scenes.json rows carry no scene_id, so the
            row index stands in for it there."""
            max_seconds = CONFIG["mimo_video_chunk_max_seconds"]
            min_seconds = CONFIG["mimo_video_chunk_min_seconds"]
            chunks = []
            for scene_index, scene in enumerate(scenes):
                start, end = scene["start"], scene["end"]
                scene_id = scene.get("scene_id", scene_index)
                cursor = start
                while cursor < end:
                    chunk_end = min(end, cursor + max_seconds)
                    if end - chunk_end < min_seconds and chunk_end < end:
                        chunk_end = end
                    chunks.append(
                        {
                            "chunk_id": len(chunks),
                            "scene_id": scene_id,
                            "start": round(cursor, 3),
                            "end": round(chunk_end, 3),
                        }
                    )
                    cursor = chunk_end
            return chunks
        
        
        def _mimo_overview_matches_current_inputs(overview, scenes, video_path=None):
            """True only when mimo_video_overview.json was produced from the current settings,
            source video and scene plan, and none of its chunks was moderation-rejected.
            Provenance keys are read with .get: a stale or partial file simply does not match."""
            if not isinstance(overview, dict) or overview.get("input") != "scene_chunks":
                return False
            settings = _mimo_video_settings()
            expected_keys = [
                _mimo_chunk_cache_key(chunk) for chunk in _mimo_video_chunks(scenes)
            ]
            chunks = overview.get("chunks")
            if not isinstance(chunks, list) or not all(
                isinstance(chunk, dict)
                and isinstance(chunk.get("content"), str)
                and _is_mimo_chunk_usable(chunk["content"])
                for chunk in chunks
            ):
                return False
            if video_path is not None:
                try:
                    if overview.get("source_video_identity") != file_identity(video_path):
                        return False
                except OSError:
                    return False
            try:
                cached_keys = [_mimo_chunk_cache_key(chunk) for chunk in chunks]
            except (KeyError, TypeError, ValueError):
                return False
            return overview.get("settings") == settings and cached_keys == expected_keys
        
        
        def _load_mimo_overview_for_brief(work_dir, scenes, enabled=None, video_path=None):
            """The current MiMo overview for the brief, or None when disabled, absent or stale."""
            if not (CONFIG["mimo_video_overview"] if enabled is None else enabled):
                return None
            path = Path(work_dir) / "mimo_video_overview.json"
            if not path.exists():
                return None
            try:
                overview = json.loads(path.read_text(encoding="utf-8"))
            except (OSError, json.JSONDecodeError):
                return None
            if not _mimo_overview_matches_current_inputs(
                overview, scenes, video_path=video_path
            ):
                return None
            return overview
        
        
        def _load_optional_stage_status(work_dir, filename):
            """Optional-stage status sidecar, or None when the stage never ran."""
            path = Path(work_dir) / filename
            if not path.exists():
                return None
            try:
                status = json.loads(path.read_text(encoding="utf-8"))
            except (OSError, json.JSONDecodeError):
                return None
            if not (
                isinstance(status, dict)
                and isinstance(status.get("enabled"), bool)
                and isinstance(status.get("status"), str)
                and isinstance(status.get("message"), str)
            ):
                return None
            return status
        
        
        def _optional_stage_warning(stage, status, message):
            line = f"- {stage}: {status}"
            return f"{line} — {message[:180]}" if message else line
        
        
        def _format_optional_stage_warnings(
            work_dir,
            *,
            mimo_overview_enabled=None,
            mimo_overview=None,
            consolidation_index=None,
        ):
            """Surface fail-open optional-stage loss near the top of the brief."""
            work_dir = Path(work_dir)
            warnings = []
        
            overview_enabled = (
                CONFIG["mimo_video_overview"]
                if mimo_overview_enabled is None
                else mimo_overview_enabled
            )
            overview_status = _load_optional_stage_status(
                work_dir, "mimo_video_overview.status.json"
            )
            if (
                overview_status is not None
                and overview_status["enabled"]
                and overview_status["status"] in {"failed", "skipped_no_key"}
            ):
                warnings.append(
                    _optional_stage_warning(
                        "mimo_video_overview",
                        overview_status["status"],
                        overview_status["message"],
                    )
                )
            elif overview_enabled and not mimo_overview:
                warnings.append(
                    _optional_stage_warning(
                        "mimo_video_overview",
                        "missing_artifact",
                        "enabled but no valid mimo_video_overview.json is available to this brief",
                    )
                )
        
            consolidation_status = _load_optional_stage_status(
                work_dir, "consolidation.status.json"
            )
            if consolidation_status is not None and consolidation_status["enabled"]:
                if consolidation_status["status"] == "failed":
                    warnings.append(
                        _optional_stage_warning(
                            "consolidation", "failed", consolidation_status["message"]
                        )
                    )
                elif consolidation_status.get("do_index") is True and not consolidation_index:
                    warnings.append(
                        _optional_stage_warning(
                            "consolidation",
                            "missing_index",
                            "enabled but no valid understanding_index.json is available to this brief",
                        )
                    )
        
            if not warnings:
                return []
            return [
                "## Optional stage warnings",
                "",
                "These stages are fail-open; continue, but do not assume their missing context exists.",
                *warnings,
                "",
            ]
        
      • timeline.py 13.4 KB
        """Remap cut evidence and format output-timeline brief directives."""
        
        import json
        import math
        import re
        from pathlib import Path
        
        from lib import CONFIG, file_identity
        
        
        def _parse_target_seconds(value):
            """Parse a cut-mode target duration ("30m" / "600" / "1h5m" / "00:30:00") to seconds.
        
            None or blank means the knob is unset (None). Any other value that does not parse to
            a positive duration is a typo in user input and raises ValueError.
            """
            if value is None:
                return None
            if isinstance(value, (int, float)):
                seconds = float(value)
            else:
                text = str(value).strip().lower()
                if not text:
                    return None
                if ":" in text:
                    parts = [float(p) for p in text.split(":")]
                    if len(parts) == 2:
                        seconds = parts[0] * 60 + parts[1]
                    elif len(parts) == 3:
                        seconds = parts[0] * 3600 + parts[1] * 60 + parts[2]
                    else:
                        raise ValueError(f"unparseable target duration: {value!r}")
                    if any(p < 0 for p in parts):
                        raise ValueError(f"unparseable target duration: {value!r}")
                else:
                    factors = {"ms": 0.001, "s": 1, "m": 60, "h": 3600}
                    seconds = 0.0
                    pos = 0
                    for m in re.finditer(r"([0-9]+(?:\.[0-9]+)?)(ms|s|m|h)?", text):
                        if m.start() != pos:
                            break
                        pos = m.end()
                        seconds += float(m.group(1)) * factors[m.group(2) or "s"]
                    if pos == 0 or pos != len(text):
                        raise ValueError(f"unparseable target duration: {value!r}")
            if not seconds > 0:
                raise ValueError(f"target duration must be positive: {value!r}")
            return seconds
        
        
        def _format_research_directive(work_dir, substrate):
            """A loud, actionable research-first directive — but ONLY when the substrate is too thin
            for real commentary (no dialogue/story spine) and no background_research.json exists yet.
        
            The narration's quality ceiling is how much story context the agent has: with only frame
            descriptions it can only narrate pixels, so research the title FIRST (see
            references/research-guide.md). Fires only for thin/empty substrate — NOT merely because a
            title was given — so a dialogue-rich titled run is never nagged.
            """
            if (Path(work_dir) / "background_research.json").exists():
                return []  # already researched; _format_background_research surfaces it
            if substrate["level"] not in ("thin", "empty"):
                return []  # rich enough to write from dialogue/spine; do not nag
            context = CONFIG["context_info"].strip()
            return [
                "## ⚑ Research the story FIRST (do this before writing narration)",
                "",
                "Reason: the understanding substrate is thin — no dialogue/story spine, only frame",
                "descriptions, so without research the narration can only describe pixels.",
                "1. Pull the title/keywords from `--context`, the filename, or the user's description"
                + (f" (context: {context})." if context else "."),
                "2. Use any available web-search/browser tool to look up synopsis, characters, and",
                "   relationships (see `references/research-guide.md`).",
                "3. Write `work_dir/background_research.json`, then re-read this brief and write narration",
                "   that names people and reads the picture through the plot.",
                "4. If no tool/network or nothing found: skip — keep beats sparse and strictly grounded in",
                "   the visible ASR/frame evidence rather than inventing drama.",
                "",
            ]
        
        
        def _load_cut_output_spans_for_brief(work_dir, *, required=False):
            """Load fresh source→output spans for cut pass 2 brief evidence.
        
            Pass 2 narration is authored against edited_source.mp4's OUTPUT timeline, so
            ASR chunks and timeline_fusion must use the same OUTPUT clock. Before pass 2
            (no edited_source.mp4 yet), callers may fall back to source-time evidence; once
            pass 2 exists, missing/stale validated spans are a hard contract failure.
            """
        
            def fail(reason):
                if required:
                    raise SystemExit(
                        "cut pass2 brief requires fresh clip_plan_validated.json with explicit "
                        f"source/output spans ({reason})"
                    )
                return None
        
            work_dir = Path(work_dir)
            if not (work_dir / "edited_source.mp4").exists():
                return None
            validated_path = work_dir / "clip_plan_validated.json"
            if not validated_path.exists():
                return fail("missing clip_plan_validated.json")
            plan = json.loads(validated_path.read_text(encoding="utf-8"))
            raw_path = work_dir / "clip_plan.json"
            if (
                raw_path.exists()
                and validated_path.stat().st_mtime_ns < raw_path.stat().st_mtime_ns
            ):
                return fail("stale clip_plan_validated.json")
            spans = []
            for clip in plan["clips"]:
                span = {
                    key: clip[key]
                    for key in ("source_start", "source_end", "output_start", "output_end")
                }
                if not all(math.isfinite(value) for value in span.values()):
                    return fail("non-finite clip span")
                if span["source_end"] <= span["source_start"] or span["output_end"] <= span["output_start"]:
                    return fail("non-positive clip span")
                spans.append(span)
            return spans or fail("no clips")
        
        
        def _source_output_overlaps_for_brief(start, end, spans):
            overlaps = []
            for span in spans:
                source_start = max(start, span["source_start"])
                source_end = min(end, span["source_end"])
                if source_end <= source_start:
                    continue
                output_start = span["output_start"] + (source_start - span["source_start"])
                output_end = span["output_start"] + (source_end - span["source_start"])
                overlaps.append(
                    {
                        "source_start": source_start,
                        "source_end": source_end,
                        "output_start": output_start,
                        "output_end": output_end,
                    }
                )
            return overlaps
        
        
        def _remap_frame_facts_for_brief(frame_facts, overlap):
            out = {}
            for raw_ts, vals in frame_facts.items():
                ts = float(raw_ts)
                if not (overlap["source_start"] <= ts <= overlap["source_end"]):
                    continue
                out_ts = overlap["output_start"] + (ts - overlap["source_start"])
                out[f"{out_ts:.3f}"] = vals
            return out
        
        
        def _remap_scenes_to_output_for_brief(scenes, spans):
            out = []
            for scene in scenes:
                overlaps = _source_output_overlaps_for_brief(scene["start"], scene["end"], spans)
                for part_idx, overlap in enumerate(overlaps):
                    item = dict(scene)
                    item["start"] = round(overlap["output_start"], 3)
                    item["end"] = round(overlap["output_end"], 3)
                    item["frame_facts"] = _remap_frame_facts_for_brief(
                        scene.get("frame_facts", {}), overlap
                    )
                    if len(overlaps) > 1:
                        item["scene_id"] = f"{scene['scene_id']}.{part_idx}"
                    out.append(item)
            out.sort(key=lambda x: (x["start"], x["end"]))
            return out
        
        
        def _remap_segments_to_output_for_brief(segments, spans):
            out = []
            for seg in segments:
                for overlap in _source_output_overlaps_for_brief(seg["start"], seg["end"], spans):
                    item = dict(seg)
                    item["start"] = round(overlap["output_start"], 3)
                    item["end"] = round(overlap["output_end"], 3)
                    if "duration" in item:
                        item["duration"] = round(item["end"] - item["start"], 3)
                    out.append(item)
            out.sort(key=lambda x: (x["start"], x["end"]))
            return out
        
        
        def _remap_brief_evidence_to_output_timeline(
            work_dir, scenes_analysis, asr_result, silence_periods, *, required=False
        ):
            spans = _load_cut_output_spans_for_brief(work_dir, required=required)
            if not spans:
                return scenes_analysis, asr_result, silence_periods
            return (
                _remap_scenes_to_output_for_brief(scenes_analysis, spans),
                _remap_segments_to_output_for_brief(asr_result, spans),
                _remap_segments_to_output_for_brief(silence_periods, spans),
            )
        
        
        def _read_json(path, default):
            """JSON artifact, or `default` when the optional file was never written."""
            return json.loads(path.read_text(encoding="utf-8")) if path.exists() else default
        
        
        def _sentence_entry_anchors_for_brief(work_dir, edit_mode):
            """Load sentence anchors and remap them to cut OUTPUT time when needed."""
            work_dir = Path(work_dir)
            anchors = [
                dict(item)
                for item in _read_json(
                    work_dir / "speech_boundary_anchors.json", {"sentence_anchors": []}
                )["sentence_anchors"]
            ]
            if edit_mode != "cut" or not (work_dir / "edited_source.mp4").exists():
                return anchors
        
            spans = _load_cut_output_spans_for_brief(work_dir, required=True)
            remapped = []
            for anchor in anchors:
                source_time = anchor["time"]
                for span in spans:
                    if span["source_start"] - 0.05 <= source_time <= span["source_end"] + 0.05:
                        item = dict(anchor)
                        item["source_time"] = round(source_time, 3)
                        item["time"] = round(
                            span["output_start"] + source_time - span["source_start"], 3
                        )
                        # Preserve the measured safe pause in OUTPUT time too, so cut-mode lint
                        # compares one clock.
                        source_pause_start = max(
                            span["source_start"], min(anchor.get("pause_start", source_time - 0.12), source_time)
                        )
                        item["source_pause_start"] = round(source_pause_start, 3)
                        item["pause_start"] = round(
                            span["output_start"] + source_pause_start - span["source_start"], 3
                        )
                        remapped.append(item)
        
            speech_rows = [
                row for row in _read_json(work_dir / "asr_result.json", []) if row["text"]
            ]
            quiet_rows = [
                row
                for row in _read_json(work_dir / "silence_periods.json", [])
                if not row["has_speech"]
            ]
            out_payload = {
                "schema_version": 2,
                "artifact": "speech_boundary_anchors_output.json",
                "timeline": "cut_output",
                "source_artifact": "speech_boundary_anchors.json",
                "clip_plan_identity": file_identity(work_dir / "clip_plan_validated.json"),
                "sentence_anchors": sorted(remapped, key=lambda item: item["time"]),
                "speech_spans": _remap_segments_to_output_for_brief(speech_rows, spans),
                "quiet_windows": _remap_segments_to_output_for_brief(quiet_rows, spans),
            }
            (work_dir / "speech_boundary_anchors_output.json").write_text(
                json.dumps(out_payload, ensure_ascii=False, indent=2), encoding="utf-8"
            )
            return out_payload["sentence_anchors"]
        
        
        def _format_sentence_entry_anchors_for_brief(work_dir, edit_mode):
            anchors = [
                anchor
                for anchor in _sentence_entry_anchors_for_brief(work_dir, edit_mode)
                if anchor["confidence"] in {"high", "medium"}
            ]
            if not anchors:
                return []
            lines = [
                "## 原声句末安全切入点",
                "",
                "这些时间是 ASR 句末标点与短声学停顿对齐后的旁白安全入口。旁白在原声已开始后切入时,"
                "必须从其中一个点开始;否则会在 TTS 前被 `interrupts_source_sentence` 硬阻断。",
                "- 调整方式:优先把 `start` 移到建议锚点;放不下时缩短文本、移动整块或删除该旁白,不能让脚本静默挪音频。",
                '- 原声句子完整性是硬约束:切入块写 `"source_entry_policy": "sentence_boundary"`;'
                "不存在 `intentional_interrupt` 绕过方式。没有后续可靠句末锚点时,移动、缩短或删除旁白块。",
            ]
            for anchor in anchors:
                source_suffix = (
                    f" (SOURCE {anchor['source_time']:.2f}s)" if "source_time" in anchor else ""
                )
                text_tail = str(anchor.get("text_tail", "")).strip()
                lines.append(
                    f"- {anchor['time']:.2f}s [{anchor['confidence']}]{source_suffix} {text_tail}".rstrip()
                )
            lines.append("")
            return lines
        
        
        def _format_output_clip_list(work_dir):
            """List the kept clips on the OUTPUT timeline (cut-first/narrate-second pass 2), so the
            agent narrates against the real rendered cut instead of the source timeline."""
            path = Path(work_dir) / "clip_plan_validated.json"
            if not path.exists():
                return []
            plan = json.loads(path.read_text(encoding="utf-8"))
            out = ["## Kept clips on the OUTPUT timeline", ""]
            for c in plan["clips"]:
                reason = c.get("reason", "")
                out.append(
                    f"- OUTPUT {c['output_start']:.1f}–{c['output_end']:.1f}s ← SOURCE[{c.get('source_id', '0')}] "
                    f"{c['source_start']:.1f}–{c['source_end']:.1f}s (clip_id={c.get('clip_id', '?')})"
                    + (f" — {reason}" if reason else "")
                )
            out.append("")
            return out if len(out) > 2 else []
        
        
        def _write_deslop_qc_requirements(work_dir):
            """Write the stable deslop QC contract consumed by deslop_qc.py.
        
            style_card_required defaults to False (advisory): a missing style_card.json is a
            warning, not a render-blocking error. A future opt-in run can set it True to make
            style_card.json a hard requirement — deslop_qc.py reads that field.
            """
            payload = {
                "schema_version": 1,
                "style_card_required": False,
            }
            path = Path(work_dir) / "deslop_qc_requirements.json"
            path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
            return path
        
      • __init__.py 101 B
        """Brief-generation subpackage: builds the creative agent brief from ASR/scene/timeline evidence."""
        
    • agent_text.py 10.8 KB
      """Normalize, split, budget, and de-duplicate narration text."""
      
      import copy
      import re
      
      from lib import CONFIG, log
      from deslop_qc import _sentence_pieces, _text_units
      
      
      def _format_frame_facts(scene):
          """将帧动作描述格式化为可注入 agent brief 的文本。"""
          facts = scene.get("frame_facts", {})
          if not facts:
              return ""
          lines = [f"    {ts}s: {'; '.join(facts[ts])}" for ts in sorted(facts, key=float)]
          return "\n  帧动作:\n" + "\n".join(lines)
      
      
      def _text_char_count(text):
          """计算文本的有效字数(去除标点和空白,这些不占 TTS 朗读时间)。"""
          return len(
              re.sub(
                  r'[,。!?、;:…“”‘’《》〈〉\s"\'「」『』()()【】\[\]—~·,.!?;:\\-]',
                  "",
                  text,
              )
          )
      
      
      def _split_text_by_sentence_windows(text, min_chars=500, max_chars=800):
          """Clipto-style three-tier sentence boundary splitting for long ASR text."""
          text = re.sub(r"\s+", " ", text).strip()
          if not text:
              return []
          if _text_units(text) <= max_chars:
              return _sentence_pieces(text)
      
          result = []
          rest = text
          sentence_marks = "。!?!?;;."
          while _text_units(rest) > max_chars:
              # The min/max thresholds are unit-based but punctuation is char-indexed.
              # For Chinese (the dominant recap target) these are nearly identical; for
              # non-CJK this remains a safe sentence-boundary heuristic around words.
              char_min = min(len(rest), max(1, min_chars))
              char_max = min(len(rest), max(1, max_chars))
              window = rest[:char_max]
              cut = max(window.rfind(mark) for mark in sentence_marks)
              if cut + 1 < char_min:
                  outside = -1
                  for i, ch in enumerate(rest[char_max:], start=char_max):
                      if ch in sentence_marks:
                          outside = i
                          break
                  cut = outside if outside >= 0 and outside + 1 <= len(rest) else char_max - 1
              piece = rest[: cut + 1].strip()
              if not piece:
                  piece = rest[:char_max].strip()
                  cut = char_max - 1
              result.append(piece)
              rest = rest[cut + 1 :].strip()
          if rest:
              result.append(rest)
          return result
      
      
      def _timed_sentence_pieces(seg, min_chars, max_chars):
          """Split one ASR segment into timed sentence pieces with approximate spans."""
          text = seg["text"].strip()
          if not text:
              return []
          start, end = seg["start"], seg["end"]
          pieces = []
          for sentence in _sentence_pieces(text):
              if _text_units(sentence) > max_chars:
                  pieces.extend(
                      _split_text_by_sentence_windows(
                          sentence, min_chars=min_chars, max_chars=max_chars
                      )
                  )
              else:
                  pieces.append(sentence)
          total_units = sum(max(1, _text_units(piece)) for piece in pieces)
          duration = end - start
          cursor = start
          timed = []
          for idx, piece in enumerate(pieces):
              units = max(1, _text_units(piece))
              piece_end = end if idx == len(pieces) - 1 else cursor + duration * units / total_units
              timed.append(
                  {
                      "start": round(cursor, 2),
                      "end": round(piece_end, 2),
                      "text": piece,
                      "char_count": _text_units(piece),
                  }
              )
              cursor = piece_end
          return timed
      
      
      def _scene_ids_for_range(scenes, start, end):
          duration = max(0.001, end - start)
          scene_ids = []
          for scene in scenes:
              overlap = _overlap_seconds(start, end, scene["start"], scene["end"])
              # Ignore tiny boundary tails from approximate ASR sentence timing. A
              # scene id should mean the chunk materially belongs to that scene.
              if overlap and (overlap >= duration * 0.2 or overlap >= 3.0):
                  scene_ids.append(scene["scene_id"])
          return scene_ids
      
      
      def _chunk_asr_for_writing(segments, scenes_analysis, min_chars=None, max_chars=None):
          """Chunk ASR into semantic windows before an agent writes long-dialogue recaps.
      
          The strategy mirrors Clipto's segment splitter: accumulate a window, prefer
          the last sentence boundary inside max length, allow a slightly longer first
          boundary outside the window, and fall back to the remaining text. CJK text is
          measured by characters; non-CJK text is measured by words.
          """
          min_chars = CONFIG["asr_chunk_min_chars"] if min_chars is None else min_chars
          max_chars = CONFIG["asr_chunk_max_chars"] if max_chars is None else max_chars
          pieces = [
              piece
              for seg in segments
              for piece in _timed_sentence_pieces(seg, min_chars, max_chars)
          ]
      
          chunks = []
          current = []
          current_units = 0
          current_scene_ids = set()
      
          def flush():
              nonlocal current, current_units, current_scene_ids
              if not current:
                  return
              chunks.append(
                  {
                      "chunk_id": len(chunks),
                      "start": current[0]["start"],
                      "end": current[-1]["end"],
                      "scene_ids": sorted(
                          current_scene_ids, key=lambda sid: (isinstance(sid, str), sid)
                      ),
                      "char_count": current_units,
                      "text": " ".join(piece["text"] for piece in current).strip(),
                      "segments": current,
                  }
              )
              current = []
              current_units = 0
              current_scene_ids = set()
      
          for piece in pieces:
              units = max(1, piece["char_count"])
              piece_scene_ids = set(
                  _scene_ids_for_range(scenes_analysis, piece["start"], piece["end"])
              )
              crosses_scene = (
                  current
                  and current_scene_ids
                  and piece_scene_ids
                  and not (current_scene_ids & piece_scene_ids)
              )
              if crosses_scene and current_units >= min_chars:
                  flush()
              if current and current_units >= min_chars and current_units + units > max_chars:
                  flush()
              current.append(piece)
              current_scene_ids.update(piece_scene_ids)
              current_units += units
              if current_units >= max_chars:
                  flush()
          flush()
          return chunks
      
      
      def _truncate_at_sentence(text, max_chars):
          """在句子边界截断,不产生残句。max_chars 按有效字符计(不含标点空白)。"""
          if _text_char_count(text) <= max_chars:
              return text
          eff = 0
          cutoff = len(text)
          for i, ch in enumerate(text):
              eff += 1 if _text_char_count(ch) else 0
              if eff > max_chars:
                  cutoff = i + 1
                  break
          idx = max(text[:cutoff].rfind(sep) for sep in ["。", "!", "?", "!", "?"])
          if idx > 0:
              return text[: idx + 1]
          idx = max(text[:cutoff].rfind(sep) for sep in [",", "、", ";", ","])
          if idx > 3:
              return text[:idx] + "。"
          return ""
      
      
      def _char_bigrams(text):
          return {text[i : i + 2] for i in range(len(text) - 1) if text[i : i + 2].strip()}
      
      
      def _post_dedup_narration(narration):
          """去除相邻相似解说段(bigram 重叠 >60% 则合并)。"""
          if len(narration) < 2:
              return narration
          result = [narration[0]]
          for seg in narration[1:]:
              prev = result[-1]
              set_a, set_b = _char_bigrams(prev["narration"]), _char_bigrams(seg["narration"])
              if not set_a or not set_b:
                  result.append(seg)
                  continue
              overlap = len(set_a & set_b) / min(len(set_a), len(set_b))
              # Only merge near-identical adjacent beats. Short Chinese beats share many
              # bigrams by chance, so a low threshold collapses intentional parallel beats
              # ("他不再试探" / "他直接赌上全力") and silently drops density below target.
              if overlap > 0.6:
                  # Validation is also the handoff boundary for renderer metadata: keep the
                  # overlays authored on either merged beat.
                  merged_overlays = copy.deepcopy(
                      prev.get("visual_overlays", []) + seg.get("visual_overlays", [])
                  )
                  if len(seg["narration"]) > len(prev["narration"]):
                      prev["narration"] = seg["narration"]
                  prev["end"] = seg["end"]
                  prev["pause_after_ms"] = seg["pause_after_ms"]
                  if merged_overlays:
                      prev["visual_overlays"] = merged_overlays
                  log(f"  去重合并: {prev['start']:.0f}-{prev['end']:.0f}s")
              else:
                  result.append(seg)
          removed = len(narration) - len(result)
          if removed:
              log(f"  去重: {len(narration)} → {len(result)} 段 (合并 {removed} 段)")
          return result
      
      
      def _scene_available_seconds(start, end):
          return max(0.0, end - start - CONFIG["narration_tail_pad_seconds"])
      
      
      def _recommended_char_budget(start, end):
          # account for the global narration atempo (CONFIG['narration_speed']) so a beat's text
          # is budgeted against the FINAL sped-up audio, not the raw TTS rate — otherwise windows
          # are over-sized and the bed shows long silent gaps between sentences.
          effective_rate = (
              CONFIG["speech_rate"] * CONFIG["speech_safety_margin"] * CONFIG["narration_speed"]
          )
          return int(_scene_available_seconds(start, end) * effective_rate)
      
      
      def _find_scene_for_midpoint(scenes_analysis, start, end):
          mid = (start + end) / 2
          for scene in scenes_analysis:
              if scene["start"] <= mid <= scene["end"]:
                  return scene
          return None
      
      
      def _normalise_narration_segment(seg):
          """Normalize one lint-validated narration segment for the full-mode rewrite.
      
          Authored timing is a delivery contract: scene boundaries are approximate
          visual-analysis buckets, so start/end are never clamped here.
          """
          item = {
              "start": round(seg["start"], 2),
              "end": round(seg["end"], 2),
              "narration": str(seg["narration"]).strip(),
              "pause_after_ms": int(seg.get("pause_after_ms", CONFIG["breath_ms"])),
              "overlaps_speech": bool(seg.get("overlaps_speech", True)),
          }
          for optional_key in (
              "source_start",
              "source_end",
              "source_clip_id",
              "source_entry_policy",
              "source_entry_reason",
          ):
              if optional_key in seg:
                  item[optional_key] = seg[optional_key]
          # carry the per-beat emotion/tone tag (MiMo TTS instruct) through lint untouched
          if seg.get("emotion"):
              item["emotion"] = seg["emotion"].strip()
          # Renderer-owned metadata survives the validation rewrite; the recap orchestrator
          # filters supported overlay kinds later.
          if "visual_overlays" in seg:
              item["visual_overlays"] = copy.deepcopy(seg["visual_overlays"])
          return item
      
      
      def _clean_narration_punctuation(text):
          text = re.sub(r"\s+", " ", text).strip()
          text = re.sub(r'[,:、;,]["\']?[。!?]', "。", text)
          return re.sub(r'["\']。$', "。", text)
      
      
      def _overlap_seconds(start, end, other_start, other_end):
          return max(0.0, min(end, other_end) - max(start, other_start))
      
    • asr.py 10.9 KB
      import base64
      import json
      import os
      import re
      import time
      from pathlib import Path
      
      from lib import CONFIG
      from lib import log, run_cmd, get_video_duration, mimo_asr_api_call
      from detect import _audio_meta_path, _write_audio_meta
      from asr_timing_evidence import (
          EVIDENCE_FILENAME,
          load_glossary_names,
          write_asr_timing_evidence,
      )
      
      # ── Step 3: ASR 转录(MiMo mimo-v2.5-asr,云端 API)────────────────────────
      
      _ASR_AUDIO_MIME = "audio/wav"
      
      
      class ASRProviderError(RuntimeError):
          """Provider/API response failed, so an empty transcript is not cacheable success."""
      
      
      def _load_name_glossary(work_dir):
          """从 background_research.json 收集已知人名(characters 键 + character_details 键及别名)。
      
          返回去重后、长度 >=2 的人名列表(按长度降序,长名优先匹配)。文件缺失或无名字时返回 []。
          """
          return load_glossary_names(work_dir)
      
      
      def _correct_text_with_glossary(text, names):
          """用人名表修正 ASR 同音字错误(如 叶青眉 → 叶轻眉)。
      
          对每个已知人名,扫描文本中所有等长窗口;若某窗口与人名恰好相差一个字符(仅一处不同),
          则替换为该人名。严格约束为「恰好一字之差」,避免过度纠正(叶轻风 与 叶轻眉 也是一字之差,
          但这正是限制为单字替换的边界——只有当窗口本身不是任何已知人名时才会被改写)。
          """
          if not text or not names:
              return text
          name_set = set(names)
          for name in names:
              n = len(name)
              if n < 2 or len(text) < n:
                  continue
              i = 0
              while i <= len(text) - n:
                  window = text[i:i + n]
                  # Only rewrite a one-char-off window when it is NOT itself a known name —
                  # otherwise a distinct real name one char away (叶轻风 vs 叶轻眉) would be corrupted.
                  if window != name and window not in name_set and _one_char_diff(window, name):
                      text = text[:i] + name + text[i + n:]
                      i += n
                  else:
                      i += 1
          return text
      
      
      def _one_char_diff(a, b):
          """两个等长字符串是否恰好相差一个字符(仅一处不同)。"""
          if len(a) != len(b):
              return False
          diff = 0
          for ca, cb in zip(a, b):
              if ca != cb:
                  diff += 1
                  if diff > 1:
                      return False
          return diff == 1
      
      
      def _apply_glossary_corrections(segments, work_dir):
          """对已转录的 segments 就地应用人名表修正;无人名表时为 no-op。"""
          names = _load_name_glossary(work_dir)
          if not names:
              return segments
          for seg in segments:
              original = seg["text"]
              corrected = _correct_text_with_glossary(original, names)
              if corrected != original:
                  seg["text"] = corrected
          return segments
      
      
      
      def transcribe_audio(video_path, work_dir):
          """提取音频并用 MiMo ASR 分段转录,通过分段合成时间戳。"""
          work_dir = Path(work_dir)
          asr_file = work_dir / "asr_result.json"
          (work_dir / EVIDENCE_FILENAME).unlink(missing_ok=True)
      
          if not CONFIG["mimo_asr_api_key"]:
              key_name = CONFIG["mimo_asr_env_var"]
              log(f"ASR 跳过:未设置 {key_name}(MiMo ASR 需要;VLM/TTS 也需要同一个 key)。"
                  f"如不需要对白可加 --skip-asr")
              asr_file.write_text(json.dumps([], ensure_ascii=False, indent=2), encoding="utf-8")
              write_asr_timing_evidence(
                  work_dir, video_path, "UNAVAILABLE_NO_KEY", final_segments=[]
              )
              return []
      
          # 提取音频
          audio_wav = work_dir / "audio.wav"
          audio_meta = _audio_meta_path(work_dir)
          audio_wav.unlink(missing_ok=True)
          audio_meta.unlink(missing_ok=True)
          cmd = ["ffmpeg", "-y", "-i", str(video_path), "-vn",
                 "-ar", "16000", "-ac", "1", str(audio_wav)]
          try:
              result = run_cmd(cmd)
          except Exception as exc:
              audio_wav.unlink(missing_ok=True)
              audio_meta.unlink(missing_ok=True)
              asr_file.unlink(missing_ok=True)
              write_asr_timing_evidence(work_dir, video_path, "FAILED_AUDIO_EXTRACTION")
              raise RuntimeError("音频提取失败: ffmpeg 无法完成") from exc
          if result.returncode != 0:
              audio_wav.unlink(missing_ok=True)
              audio_meta.unlink(missing_ok=True)
              asr_file.unlink(missing_ok=True)
              write_asr_timing_evidence(work_dir, video_path, "FAILED_AUDIO_EXTRACTION")
              raise RuntimeError(f"音频提取失败: {result.stderr}")
          _write_audio_meta(work_dir, video_path)
      
          # 获取音频时长;ffprobe 失败时不伪造时长(否则会向 asr_result.json 写入虚构时间戳),
          # 而是记录 UNAVAILABLE_NO_DURATION 证据并跳过转录
          try:
              duration = get_video_duration(audio_wav)
          except RuntimeError as exc:
              log(f"ASR 警告: 无法获取音频时长,跳过 ASR 转录: {exc}")
              asr_file.write_text(json.dumps([], ensure_ascii=False, indent=2), encoding="utf-8")
              write_asr_timing_evidence(
                  work_dir,
                  video_path,
                  "UNAVAILABLE_NO_DURATION",
                  final_segments=[],
                  audio_path=audio_wav,
              )
              return []
      
          segments_dir = work_dir / "audio_segments"
          segments_dir.mkdir(exist_ok=True)
      
          segment_length = int(CONFIG["asr_segment_seconds"])
          try:
              if duration <= segment_length:
                  # 短音频,整段转录
                  text = _run_asr(audio_wav)
                  asr_result = [{"start": 0.0, "end": round(duration, 2), "text": text}]
              else:
                  # 长音频,分段转录(固定粗窗口,不是词级或对白边界对齐)
                  asr_result = _segment_and_transcribe(
                      audio_wav, segments_dir, duration, segment_length
                  )
          except ASRProviderError:
              asr_file.unlink(missing_ok=True)
              write_asr_timing_evidence(
                  work_dir, video_path, "FAILED_PROVIDER", audio_path=audio_wav
              )
              raise
      
          # 用 background_research.json 的人名表修正 ASR 同音字错误(如 叶青眉 → 叶轻眉);无人名表时为 no-op
          observed_result = [dict(segment) for segment in asr_result]
          _apply_glossary_corrections(asr_result, work_dir)
      
          # 保存
          asr_file.write_text(json.dumps(asr_result, ensure_ascii=False, indent=2), encoding="utf-8")
          status = "AVAILABLE_COARSE" if any(s["text"] for s in asr_result) else "EMPTY_UNKNOWN"
          write_asr_timing_evidence(
              work_dir,
              video_path,
              status,
              observed_segments=observed_result,
              final_segments=asr_result,
              audio_path=audio_wav,
          )
      
          total_text = " ".join(s["text"] for s in asr_result if s["text"])
          empty = sum(1 for s in asr_result if not s["text"])
          suffix = f"({empty} 段无文本:原因未知,不代表已证实静音)" if empty else ""
          log(f"ASR 转录完成: {len(asr_result)} 段, 共 {len(total_text)} 字{suffix}")
          return asr_result
      
      
      def _strip_reasoning_residue(text):
          """Remove MiMo reasoning-model <think>…</think> leakage from ASR content.
      
          Thinking-disable is not applied to -asr models (lib._prepare_api_payload), so the reasoning
          model can leak a <think> block — full, truncated/unclosed, or a leading orphan/residual tag
          (a bare "think>" prefix) — into the transcript. The dubbing implementation carries its own
          copy of this policy; keep behavior aligned through tests.
          """
          text = re.sub(r"(?is)<think\b.*?</think\s*>", "", text)  # full <think>…</think> block
          text = re.sub(r"(?is)<think\b.*\Z", "", text)            # unclosed/truncated <think tail
          return re.sub(r"(?i)^\s*<?/?think\s*>\s*", "", text)     # leading orphan/residual think tag
      
      
      def _run_asr(wav_path):
          """用 MiMo ASR (mimo-v2.5-asr) 转录单个 wav 文件,返回纯文本。
      
          音频以 base64 data-URI 放进 OpenAI 风格的 chat/completions 消息里,转写文本回到
          choices[0].message.content。API/响应结构失败会抛错,避免把瞬时失败缓存成空转写;
          只有无音频、超体积等确定不可发送的片段返回空串。
          """
          try:
              raw = Path(wav_path).read_bytes()
          except OSError as e:
              log(f"ASR 警告: 无法读取音频 {wav_path}: {e}")
              return ""
          if not raw:
              return ""
      
          b64 = base64.b64encode(raw).decode("ascii")
          max_b64_bytes = int(float(CONFIG["mimo_asr_base64_max_mb"]) * 1024 * 1024)
          if len(b64) > max_b64_bytes:
              log(f"ASR 警告: 分片 base64 体积 {len(b64) / 1024 / 1024:.1f}MB 超过 MiMo 上限 "
                  f"{CONFIG['mimo_asr_base64_max_mb']}MB,跳过该段;可调小 ASR_SEGMENT_SECONDS")
              return ""
      
          payload = {
              "model": CONFIG["mimo_asr_model"],
              "messages": [{
                  "role": "user",
                  "content": [{
                      "type": "input_audio",
                      "input_audio": {"data": f"data:{_ASR_AUDIO_MIME};base64,{b64}"},
                  }],
              }],
              "asr_options": {"language": CONFIG["mimo_asr_language"]},
          }
          try:
              resp = mimo_asr_api_call(payload)
          except Exception as e:
              raise ASRProviderError(f"MiMo ASR 调用失败: {e}") from e
          try:
              return _strip_reasoning_residue(str(resp["choices"][0]["message"]["content"] or "")).strip()
          except (KeyError, IndexError, TypeError):
              raise ASRProviderError(
                  f"MiMo ASR 返回结构异常: {json.dumps(resp, ensure_ascii=False)[:200]}"
              )
      
      
      def _segment_and_transcribe(audio_wav, segments_dir, total_duration, segment_length=None):
          """分段转录长音频"""
          if segment_length is None:
              segment_length = int(CONFIG["asr_segment_seconds"])
          # 长视频 ASR 是顺序调用;可选节流让调用间隔开,降低踩到集群限流的频率(默认 0=不节流)
          try:
              throttle = max(0.0, float(os.environ.get("ASR_THROTTLE_SECONDS", "0") or 0))
          except ValueError:
              throttle = 0.0
          results = []
      
          for i, start in enumerate(range(0, int(total_duration), segment_length)):
              if throttle and i:
                  time.sleep(throttle)
              end = min(start + segment_length, total_duration)
              seg_wav = segments_dir / f"seg_{i:03d}.wav"
      
              cmd = ["ffmpeg", "-y", "-i", str(audio_wav),
                     "-ss", str(start), "-to", str(end),
                     "-ar", "16000", "-ac", "1", str(seg_wav)]
              cut = run_cmd(cmd)
              if cut.returncode != 0:
                  # 切分失败时不要对磁盘上的陈旧/残缺音频转录,否则会得到错位文本
                  log(f"  段 {i+1}: 切分失败,跳过转录 ({cut.stderr.strip()[:200]})")
                  text = ""
              else:
                  text = _run_asr(seg_wav)
              results.append({
                  "start": round(start, 2),
                  "end": round(end, 2),
                  "text": text,
              })
              log(f"  段 {i+1}: {start:.0f}s-{end:.0f}s => {len(text)} 字")
      
          return results
      
    • asr_timing_evidence.py 10.3 KB
      """Write and validate honest timing/provenance evidence for coarse MiMo ASR."""
      
      import json
      import math
      import os
      import tempfile
      from pathlib import Path
      
      from lib import file_identity, load_background_research
      
      
      EVIDENCE_FILENAME = "asr_timing_evidence.json"
      SCHEMA_VERSION = 2
      VALID_STATUSES = {
          "AVAILABLE_COARSE",
          "EXPLICITLY_SKIPPED",
          "UNAVAILABLE_NO_KEY",
          "UNAVAILABLE_NO_DURATION",
          "FAILED_AUDIO_EXTRACTION",
          "FAILED_PROVIDER",
          "EMPTY_UNKNOWN",
          "LEGACY_UNVERIFIED",
      }
      _TOP_KEYS = {
          "schema_version", "status", "source_video", "audio", "asr_result",
          "glossary", "precision", "windows",
      }
      _WINDOW_KEYS = {
          "index", "start", "end", "text_availability", "observed_text",
          "post_glossary_text", "glossary_modified",
      }
      _GLOSSARY_KEYS = {"names", "name_count"}
      _PRECISION = {
          "window_timing": "COARSE_SEGMENT_WINDOWS",
          "dialogue_boundaries": "NOT_VERIFIED",
          "word_alignment": "NOT_PERFORMED",
          "empty_text_meaning": "UNKNOWN_NOT_PROVEN_SILENCE",
      }
      
      
      def _identity(path):
          """{size, mtime_ns} of a bound file, or None when it does not exist."""
          path = Path(path)
          return file_identity(path) if path.exists() else None
      
      
      def load_glossary_names(work_dir):
          """Return normalized names that can actually affect ASR correction."""
          data = load_background_research(work_dir)
          names = set()
          characters = data.get("characters")
          if isinstance(characters, dict):
              names.update(characters.keys())
          details = data.get("character_details")
          if isinstance(details, dict):
              for name, info in details.items():
                  names.add(name)
                  if isinstance(info, dict) and isinstance(info.get("aliases"), list):
                      names.update(alias for alias in info["aliases"] if isinstance(alias, str))
          return sorted(
              {name for name in names if isinstance(name, str) and len(name) >= 2},
              key=lambda value: (-len(value), value),
          )
      
      
      def _glossary_binding(work_dir, *, legacy=False):
          """The glossary names that could have corrected this transcription (legacy: unknown)."""
          if legacy:
              return {"names": None, "name_count": None}
          names = load_glossary_names(work_dir)
          return {"names": names, "name_count": len(names)}
      
      
      def _atomic_json_write(path, payload):
          path = Path(path)
          path.parent.mkdir(parents=True, exist_ok=True)
          fd, temporary = tempfile.mkstemp(prefix=f".{path.name}.", dir=path.parent)
          try:
              with os.fdopen(fd, "w", encoding="utf-8") as handle:
                  json.dump(payload, handle, ensure_ascii=False, indent=2)
                  handle.write("\n")
              os.replace(temporary, path)
          finally:
              try:
                  os.unlink(temporary)
              except FileNotFoundError:
                  pass
      
      
      def _window_evidence(observed_segments, final_segments, legacy):
          observed_segments = observed_segments or []
          windows = []
          for index, final in enumerate(final_segments or []):
              observed = observed_segments[index] if index < len(observed_segments) else None
              observed_text = None if legacy or observed is None else str(observed.get("text") or "")
              final_text = str(final.get("text") or "")
              windows.append({
                  "index": index,
                  "start": final.get("start"),
                  "end": final.get("end"),
                  "text_availability": "AVAILABLE" if final_text else "UNAVAILABLE_UNKNOWN",
                  "observed_text": observed_text,
                  "post_glossary_text": final_text,
                  "glossary_modified": None if observed_text is None else observed_text != final_text,
              })
          return windows
      
      
      def write_asr_timing_evidence(
          work_dir, video_path, status, *, observed_segments=None, final_segments=None,
          audio_path=None,
      ):
          """Write a sidecar recording which source/audio/result files it describes, without
          changing the legacy result."""
          if status not in VALID_STATUSES:
              raise ValueError(f"unknown ASR timing evidence status: {status}")
          work_dir = Path(work_dir)
          legacy = status == "LEGACY_UNVERIFIED"
          payload = {
              "schema_version": SCHEMA_VERSION,
              "status": status,
              "source_video": _identity(video_path),
              "audio": _identity(audio_path) if audio_path else None,
              "asr_result": _identity(work_dir / "asr_result.json"),
              "glossary": _glossary_binding(work_dir, legacy=legacy),
              "precision": dict(_PRECISION),
              "windows": _window_evidence(observed_segments, final_segments, legacy),
          }
          path = work_dir / EVIDENCE_FILENAME
          _atomic_json_write(path, payload)
          return path
      
      
      def _is_int(value):
          return isinstance(value, int) and not isinstance(value, bool)
      
      
      def _is_number(value):
          return isinstance(value, (int, float)) and not isinstance(value, bool) and math.isfinite(value)
      
      
      def _is_identity(value):
          return (
              isinstance(value, dict) and set(value) == {"size", "mtime_ns"}
              and _is_int(value["size"]) and _is_int(value["mtime_ns"])
          )
      
      
      def _valid_glossary(payload, work_dir, legacy):
          if not isinstance(payload, dict) or set(payload) != _GLOSSARY_KEYS:
              return False
          if legacy:
              return payload == _glossary_binding(work_dir, legacy=True)
          return payload == _glossary_binding(work_dir) and _is_int(payload.get("name_count"))
      
      
      def _valid_status_relationships(status, result, audio):
          texts = [str(segment.get("text") or "") for segment in result]
          if status == "AVAILABLE_COARSE":
              return bool(result) and any(texts) and _is_identity(audio)
          if status == "EMPTY_UNKNOWN":
              return bool(result) and not any(texts) and _is_identity(audio)
          if status == "UNAVAILABLE_NO_DURATION":
              return result == [] and _is_identity(audio)
          if status in {"EXPLICITLY_SKIPPED", "UNAVAILABLE_NO_KEY", "LEGACY_UNVERIFIED"}:
              return audio is None and (status == "LEGACY_UNVERIFIED" or result == [])
          return False
      
      
      def validate_asr_timing_evidence(evidence_path, video_path, asr_result_path):
          """Return whether a sidecar still describes the current source/audio/result files
          (by size + mtime_ns) and the limited timing it claims."""
          evidence_path = Path(evidence_path)
          result_path = Path(asr_result_path)
          try:
              evidence = json.loads(evidence_path.read_text(encoding="utf-8"))
              result = json.loads(result_path.read_text(encoding="utf-8"))
          except (OSError, ValueError, TypeError):
              return False
          if not isinstance(evidence, dict) or set(evidence) != _TOP_KEYS:
              return False
          if not isinstance(result, list) or not all(isinstance(item, dict) for item in result):
              return False
          if not _is_int(evidence.get("schema_version")) or evidence["schema_version"] != SCHEMA_VERSION:
              return False
          status = evidence.get("status")
          if status not in VALID_STATUSES or status.startswith("FAILED_"):
              return False
          if evidence.get("precision") != _PRECISION:
              return False
          source_identity = _identity(video_path)
          result_identity = _identity(result_path)
          if not _is_identity(source_identity) or evidence.get("source_video") != source_identity:
              return False
          if not _is_identity(result_identity) or evidence.get("asr_result") != result_identity:
              return False
          audio = evidence.get("audio")
          if audio is not None and not _is_identity(audio):
              return False
          if not _valid_status_relationships(status, result, audio):
              return False
          legacy = status == "LEGACY_UNVERIFIED"
          if not _valid_glossary(evidence.get("glossary"), evidence_path.parent, legacy):
              return False
      
          if audio is not None:
              audio_path = evidence_path.parent / "audio.wav"
              if audio != _identity(audio_path):
                  return False
              try:
                  audio_meta = json.loads(
                      (evidence_path.parent / "audio.wav.meta.json").read_text(encoding="utf-8")
                  )
              except (OSError, ValueError, TypeError):
                  return False
              if not isinstance(audio_meta, dict):
                  return False
              if audio_meta.get("source_video_identity") != source_identity:
                  return False
      
          windows = evidence.get("windows")
          if not isinstance(windows, list) or len(windows) != len(result):
              return False
          any_modified = False
          for index, (window, segment) in enumerate(zip(windows, result)):
              if not isinstance(window, dict) or set(window) != _WINDOW_KEYS:
                  return False
              start, end = segment.get("start"), segment.get("end")
              text = str(segment.get("text") or "")
              observed = window.get("observed_text")
              modified = window.get("glossary_modified")
              if not _is_number(start) or not _is_number(end) or end <= start:
                  return False
              if (
                  not _is_int(window.get("index")) or window["index"] != index
                  or not _is_number(window.get("start")) or not _is_number(window.get("end"))
                  or window["start"] != start or window["end"] != end
                  or not isinstance(window.get("post_glossary_text"), str)
                  or window["post_glossary_text"] != text
                  or window.get("text_availability") != ("AVAILABLE" if text else "UNAVAILABLE_UNKNOWN")
              ):
                  return False
              if legacy:
                  if observed is not None or modified is not None:
                      return False
              elif (
                  not isinstance(observed, str) or not isinstance(modified, bool)
                  or modified != (observed != text)
              ):
                  return False
              any_modified = any_modified or modified is True
          if any_modified and evidence["glossary"]["name_count"] == 0:
              return False
          return True
      
      
      def asr_evidence_summary_for_brief(work_dir, video_path):
          """Return a compact validated banner; never upgrade missing/stale evidence."""
          work_dir = Path(work_dir)
          evidence_path = work_dir / EVIDENCE_FILENAME
          result_path = work_dir / "asr_result.json"
          valid = bool(video_path) and validate_asr_timing_evidence(
              evidence_path, video_path, result_path
          )
          if valid:
              payload = json.loads(evidence_path.read_text(encoding="utf-8"))
              return {"status": payload["status"],
                      "glossary_modifications": None if payload["status"] == "LEGACY_UNVERIFIED"
                      else sum(w["glossary_modified"] is True for w in payload["windows"])}
          return {"status": "MISSING_OR_STALE", "glossary_modifications": None}
      
    • brief.py 250 B
      """Public narration-brief API for this self-contained skill."""
      
      from briefing.builder import build_agent_brief
      from briefing.context import assess_understanding_substrate
      
      __all__ = [
          "assess_understanding_substrate",
          "build_agent_brief",
      ]
      
    • consolidate.py 22.7 KB
      #!/usr/bin/env python3
      """video-understanding consolidation / 整理 (index build-up).
      
      Optional, synchronous, in-pipeline. Two independent LLM passes over the video's own signal:
        - Pass B (index, default): roll up per-scene vlm_analysis into a global
          character / relationship / plot index -> understanding_index.json (+ .md).
        - Pass A (asr cleanup, opt-in): clean/punctuate/lightly speaker-attribute the raw run-on
          ASR text -> asr_clean.json. Timing is preserved BY CONSTRUCTION: the model returns cleaned
          TEXT per segment index only; each segment's original start/end is re-attached here, so a
          cleanup pass can never shift the timing-bearing spans downstream chunking depends on.
      
      Both passes are idempotent (a fresh artifact is reused); a failure raises and the runner
      records it in consolidation.status.json. Mirrors review.py's pure-seam + thin-driver shape so it is
      unit-testable with a mocked api_call. NON-required: the pipeline runs unchanged without it.
      """
      
      import argparse
      import json
      import re
      from pathlib import Path
      
      from lib import CONFIG, log, api_call, file_identity, load_background_research
      from understanding_cache import _fresh
      
      # Shared tolerance for the per-segment span check. The brief-side gate inlines the SAME
      # literal (it cannot import this module without breaking the brief/narration byte-parity).
      # Keep both in sync. Consolidate preserves spans exactly, so this only guards hand-edited files.
      _ASR_SPAN_TOL = 0.05
      ASR_CLEAN_POSTPROCESS_VERSION = 2
      
      CLEAN_PROMPT = """你在清洗中文视频的 ASR 逐段转写。对【每一段】做:补标点、修明显同音/错别字、(能判断时)在句首轻标说话人,让长段连读文本变成清晰可读的句子。
      铁律:
      - 不要合并或拆分段落,输出段数必须与输入完全一致,顺序一致。
      - 不要改时间,不要输出 start/end(时间由程序保留)。
      - 只清洗 text,不要增删事实、不要脑补画面。
      - 必须把相邻两段连起来检查标点是否连续;切窗处不是自然句尾时,不要凭单段补句号。
      - 分段音频窗口互不重叠。相邻段首尾字面相同可能是合法复沓(例如“真的很难。难得让人想放弃”),不能仅凭文字相同删除任何一边的词;没有音频重叠证据时保留各段词汇内容。
      只返回 JSON:{"segments":[{"i":0,"text":"清洗后的文本","speaker":"可选说话人"}, ...]},i 为输入段的下标。"""
      
      INDEX_SCHEMA_VERSION = 2
      
      INDEX_PROMPT = """你在根据逐场景画面分析、ASR对白、background_research术语表,为一个视频建立【全局理解索引】,供后续写解说词时保持人物/关系/主线一致。
      规则:
      - visual/asr 是当前视频事实证据;每个人物、关系、剧情节点、物件尽量给 evidence_ids。
      - background_research 只能用于人名/别名/术语/身份消歧,默认 support=context_only;不能把后续剧情或未出现关系升级为当前画面事实。
      - ASR 中出现但画面描述未命名的人名,应进入 characters[*].asr_mentions。
      只返回 JSON:
      {"characters":[{"name":"角色名或外观指代","description":"身份/特征","aliases":[],"visual_descriptions":[],"asr_mentions":[],"research_role":"","evidence_ids":[],"confidence":"high|medium|low"}],
       "relationships":[{"a":"角色","b":"角色","relation":"关系","evidence_ids":[],"support":"direct|indirect|context_only"}],
       "plot_points":[{"time":"00:00","text":"按时间顺序的关键剧情节点","evidence_ids":[]}],
       "entities":[{"name":"重要物件/地点/线索","evidence_ids":[]}],
       "research_glossary":[{"name":"名字/术语","aliases":[],"role":"说明","support":"context_only"}]}"""
      
      
      # ── pure seams (no I/O; unit-testable) ────────────────────────────────────────
      
      
      def build_clean_messages(asr_result):
          lines = []
          for i, seg in enumerate(asr_result or []):
              lines.append(f"{i}. {seg['text'].strip()}")
          user = f"{CLEAN_PROMPT}\n\n## 逐段转写(共 {len(lines)} 段)\n" + "\n".join(lines)
          return [{"role": "user", "content": user}]
      
      
      def parse_clean_response(text, asr_result):
          """Zip cleaned TEXT onto the original segments (start/end preserved by construction).
          Returns the original asr_result UNCHANGED on any parse / shape / count problem."""
          base = list(asr_result or [])
          if not base:
              return asr_result
          data = _extract_json(text)
          if not isinstance(data, dict):
              return asr_result
          segs = data.get("segments")
          if not isinstance(segs, list) or len(segs) != len(base):
              return asr_result  # count mismatch -> reject wholesale (idempotent no-op)
          by_index = {}
          for item in segs:
              if isinstance(item, dict) and isinstance(item.get("i"), int):
                  by_index[item["i"]] = item
          if len(by_index) != len(base):
              return asr_result
          out = []
          for i, seg in enumerate(base):
              cleaned = by_index.get(i, {})
              new_text = str(cleaned.get("text", "")).strip() or seg["text"]
              merged = {"start": seg["start"], "end": seg["end"], "text": new_text}
              speaker = str(cleaned.get("speaker", "")).strip()
              if speaker:
                  merged["speaker"] = speaker
              out.append(merged)
          return out
      
      
      def _input_identity(work_dir, name):
          """{size, mtime_ns} of one input artifact, or None when it is absent."""
          path = Path(work_dir) / name
          return file_identity(path) if path.exists() else None
      
      
      def _research_glossary_from_context(background_research):
          glossary = []
          if not isinstance(background_research, dict):
              return glossary
          chars = background_research.get("characters")
          if isinstance(chars, dict):
              for name, role in chars.items():
                  if str(name).strip():
                      glossary.append(
                          {
                              "name": str(name).strip(),
                              "aliases": [],
                              "role": str(role).strip(),
                              "support": "context_only",
                          }
                      )
          details = background_research.get("character_details")
          if isinstance(details, dict):
              for name, info in details.items():
                  if not isinstance(info, dict):
                      continue
                  aliases = (
                      [str(a).strip() for a in (info.get("aliases") or []) if str(a).strip()]
                      if isinstance(info.get("aliases"), list)
                      else []
                  )
                  role = str(info.get("role", "")).strip()
                  if str(name).strip() or aliases or role:
                      glossary.append(
                          {
                              "name": str(name).strip(),
                              "aliases": aliases,
                              "role": role,
                              "support": "context_only",
                          }
                      )
          return glossary
      
      
      def _dialogue_source(asr_result, asr_clean):
          """asr_clean.json segments when present (consolidate's own output), else raw asr_result rows."""
          return asr_clean["segments"] if asr_clean else (asr_result or [])
      
      
      def _dialogue_segments_for_index(asr_result=None, asr_clean=None):
          out = []
          for i, seg in enumerate(_dialogue_source(asr_result, asr_clean)):
              text = seg["text"].strip()
              if not text:
                  continue
              out.append(
                  {
                      "id": f"asr:{i}",
                      "start": seg["start"],
                      "end": seg["end"],
                      "text": text,
                  }
              )
          return out
      
      
      def _stable_unique(values):
          seen = set()
          out = []
          for value in values or []:
              key = (
                  json.dumps(value, ensure_ascii=False, sort_keys=True)
                  if isinstance(value, dict)
                  else str(value)
              )
              if key in seen:
                  continue
              seen.add(key)
              out.append(value)
          return out
      
      
      def _apply_deterministic_asr_research_fallback(
          index, asr_result=None, asr_clean=None, background_research=None
      ):
          """Guarantee ASR/research mentions survive even when the LLM omits them.
      
          The LLM still owns synthesis, but cacheable deterministic fields make ASR contribution
          observable: aliases/names from background_research are matched against ASR/asr_clean text
          and written to characters[*].asr_mentions/evidence_ids without inventing relationships or
          plot facts. Research remains context_only via research_glossary.
          """
          index = dict(index or {})
          for key in (
              "characters",
              "relationships",
              "plot_points",
              "entities",
              "research_glossary",
          ):
              index.setdefault(key, [])
          glossary = index.get("research_glossary") or _research_glossary_from_context(
              background_research
          )
          if not glossary:
              glossary = _research_glossary_from_context(background_research)
          for item in glossary:
              if isinstance(item, dict):
                  item["support"] = "context_only"
          index["research_glossary"] = glossary
      
          segments = _dialogue_segments_for_index(asr_result=asr_result, asr_clean=asr_clean)
          char_by_name = {}
          for c in index.get("characters") or []:
              if isinstance(c, dict) and str(c.get("name", "")).strip():
                  char_by_name[str(c.get("name")).strip()] = c
          for g in glossary:
              if not isinstance(g, dict):
                  continue
              name = str(g.get("name", "")).strip()
              aliases = [str(a).strip() for a in (g.get("aliases") or []) if str(a).strip()]
              terms = [t for t in [name, *aliases] if t]
              if not terms:
                  continue
              matches = []
              evidence_ids = []
              for seg in segments:
                  text = seg["text"]
                  hit_terms = [term for term in terms if term and term in text]
                  if hit_terms:
                      matches.append(
                          {
                              "text": text,
                              "evidence_id": seg["id"],
                              "matched_aliases": hit_terms,
                          }
                      )
                      evidence_ids.append(seg["id"])
              if not matches:
                  continue
              char = char_by_name.get(name)
              if char is None:
                  char = {
                      "name": name or matches[0]["matched_aliases"][0],
                      "description": "",
                      "aliases": aliases,
                      "visual_descriptions": [],
                      "asr_mentions": [],
                      "research_role": str(g.get("role", "")).strip(),
                      "evidence_ids": [],
                      "confidence": "medium",
                  }
                  index["characters"].append(char)
                  char_by_name[char["name"]] = char
              char.setdefault("aliases", [])
              char["aliases"] = _stable_unique(list(char.get("aliases") or []) + aliases)
              char.setdefault("asr_mentions", [])
              char["asr_mentions"] = _stable_unique(
                  list(char.get("asr_mentions") or []) + matches
              )
              char.setdefault("evidence_ids", [])
              char["evidence_ids"] = _stable_unique(
                  list(char.get("evidence_ids") or []) + evidence_ids
              )
              if not char.get("research_role"):
                  char["research_role"] = str(g.get("role", "")).strip()
              char.setdefault("confidence", "medium")
          return index
      
      
      def build_index_messages(
          vlm_analysis, asr_result=None, asr_clean=None, background_research=None
      ):
          lines = []
          for i, scene in enumerate(vlm_analysis or []):
              sid = scene["scene_id"]
              start = float(scene["start"])
              end = float(scene["end"])
              desc = scene["description"].strip().replace("\n", " ")
              facts = scene.get("frame_facts")  # vlm.py only writes the key when non-empty
              fact_txt = ""
              if facts:
                  actions = []
      
                  def fact_key(x):
                      # timestamps are regex-captured from the VLM reply ("[\d.]+"), so "1.2.3" is possible
                      try:
                          return (0, float(x))
                      except ValueError:
                          return (1, str(x))
      
                  for ts in sorted(facts.keys(), key=fact_key):
                      actions.extend(facts[ts])
                  if actions:
                      fact_txt = " | 帧实: " + ";".join(a for a in actions[:6] if a)
              lines.append(f"[visual:{i} 场景{sid} {start:.0f}-{end:.0f}s] {desc}{fact_txt}")
          asr_lines = []
          for i, seg in enumerate(_dialogue_source(asr_result, asr_clean)):
              text = seg["text"].strip()
              if not text:
                  continue
              asr_lines.append(
                  f"[asr:{i} {float(seg['start']):.0f}-{float(seg['end']):.0f}s] {text}"
              )
          glossary = _research_glossary_from_context(background_research)
          glossary_lines = []
          for i, item in enumerate(glossary[:80]):
              aliases = "/".join(item.get("aliases") or [])
              alias_txt = f" aliases={aliases}" if aliases else ""
              glossary_lines.append(
                  f"[research:{i} support=context_only] {item.get('name', '')}{alias_txt}: {item.get('role', '')}"
              )
          user = (
              f"{INDEX_PROMPT}\n\n"
              f"## 逐场景画面分析(共 {len(lines)} 段)\n" + "\n".join(lines) + "\n\n"
              f"## ASR / cleaned dialogue(共 {len(asr_lines)} 段)\n"
              + ("\n".join(asr_lines) or "(无)")
              + "\n\n"
              "## Research glossary(clock=null/context_only,不得升级为当前剧情事实)\n"
              + ("\n".join(glossary_lines) or "(无)")
          )
          return [{"role": "user", "content": user}]
      
      
      def parse_index_response(text):
          data = _extract_json(text)
          if not isinstance(data, dict):
              return {
                  "schema_version": INDEX_SCHEMA_VERSION,
                  "characters": [],
                  "relationships": [],
                  "plot_points": [],
                  "entities": [],
                  "research_glossary": [],
              }
          out = {"schema_version": INDEX_SCHEMA_VERSION}
          for key in (
              "characters",
              "relationships",
              "plot_points",
              "entities",
              "research_glossary",
          ):
              val = data.get(key)
              out[key] = val if isinstance(val, list) else []
          for item in out["research_glossary"]:
              if isinstance(item, dict):
                  item["support"] = "context_only"
          return out
      
      
      def format_index_md(index):
          out = ["# Understanding index (from consolidate.py)", ""]
          chars = index.get("characters") or []
          if chars:
              out.append("## Characters")
              for c in chars:
                  if isinstance(c, dict):
                      out.append(
                          f"- **{c.get('name', '?')}** — {c.get('description', '')}".rstrip()
                      )
                  else:
                      out.append(f"- {c}")
              out.append("")
          rels = index.get("relationships") or []
          if rels:
              out.append("## Relationships")
              for r in rels:
                  if isinstance(r, dict):
                      out.append(
                          f"- {r.get('a', '?')} — {r.get('relation', '?')} — {r.get('b', '?')}"
                      )
                  else:
                      out.append(f"- {r}")
              out.append("")
          plot = index.get("plot_points") or []
          if plot:
              out.append("## Plot spine")
              for i, p in enumerate(plot):
                  out.append(f"{i + 1}. {p.get('text', p) if isinstance(p, dict) else p}")
              out.append("")
          ents = index.get("entities") or []
          if ents:
              out.append("## Entities")
              for e in ents:
                  out.append(f"- {e.get('name', e) if isinstance(e, dict) else e}")
              out.append("")
          glossary = index.get("research_glossary") or []
          if glossary:
              out.append("## Research glossary (context_only)")
              for g in glossary:
                  if isinstance(g, dict):
                      aliases = "/".join(g.get("aliases") or [])
                      alias_txt = f" ({aliases})" if aliases else ""
                      out.append(
                          f"- {g.get('name', '?')}{alias_txt}: {g.get('role', '')} [{g.get('support', 'context_only')}]"
                      )
                  else:
                      out.append(f"- {g}")
              out.append("")
          return "\n".join(out).rstrip() + "\n"
      
      
      def _extract_json(text):
          raw = str(text or "")
          fence = re.search(r"```(?:json)?\s*(\{.*?\})\s*```", raw, re.DOTALL)
          candidate = fence.group(1) if fence else raw
          if not fence:
              first, last = candidate.find("{"), candidate.rfind("}")
              if first != -1 and last > first:
                  candidate = candidate[first : last + 1]
          try:
              return json.loads(candidate)
          except ValueError:
              return None
      
      
      def _index_meta_path(work_dir):
          return Path(work_dir) / "understanding_index.json.meta.json"
      
      
      def _index_meta(work_dir, vlm_analysis):
          """Provenance of understanding_index.json: `source` (vlm_analysis.json identity) + `model`
          are what brief consumers check; the other inputs and the prompt text are the producer's
          own rebuild gate."""
          return {
              "schema_version": INDEX_SCHEMA_VERSION,
              "source": _input_identity(work_dir, "vlm_analysis.json"),
              "inputs": {
                  "asr_result": _input_identity(work_dir, "asr_result.json"),
                  "asr_clean": _input_identity(work_dir, "asr_clean.json"),
                  "background_research": _input_identity(work_dir, "background_research.json"),
              },
              "scene_count": len([s for s in (vlm_analysis or []) if isinstance(s, dict)]),
              "model": CONFIG["vlm_model"],
              "prompt": INDEX_PROMPT,
          }
      
      
      def _write_index_meta(work_dir, vlm_analysis):
          _index_meta_path(work_dir).write_text(
              json.dumps(_index_meta(work_dir, vlm_analysis), ensure_ascii=False, indent=2),
              encoding="utf-8",
          )
      
      
      def _index_cache_matches(work_dir, vlm_analysis):
          meta_path = _index_meta_path(work_dir)
          if not meta_path.exists():
              return False
          try:
              meta = json.loads(meta_path.read_text(encoding="utf-8"))
          except (json.JSONDecodeError, OSError):
              return False
          return meta == _index_meta(work_dir, vlm_analysis)
      
      
      # ── thin drivers (I/O + api_call) ─────────────────────────────────────────────
      
      
      def _load(work_dir, name):
          """Read one of this skill's own JSON artifacts: absent → None, corrupt → raise."""
          path = Path(work_dir) / name
          if not path.exists():
              return None
          return json.loads(path.read_text(encoding="utf-8"))
      
      
      def consolidate_transcript(work_dir):
          work_dir = Path(work_dir)
          asr_result = _load(work_dir, "asr_result.json")
          if not asr_result:
              log("consolidate(asr): 无 asr_result.json,跳过")
              return None
          out_path = work_dir / "asr_clean.json"
          if _fresh(out_path, work_dir / "asr_result.json"):
              existing = _load(work_dir, "asr_clean.json") or {}
              if (
                  existing.get("source") == _input_identity(work_dir, "asr_result.json")
                  and existing.get("model") == CONFIG["vlm_model"]
                  and existing.get("prompt") == CLEAN_PROMPT
                  and existing.get("postprocess_version") == ASR_CLEAN_POSTPROCESS_VERSION
              ):
                  log("consolidate(asr): asr_clean.json 已最新,跳过")
                  return existing
          resp = api_call(
              {
                  "model": CONFIG["vlm_model"],
                  "messages": build_clean_messages(asr_result),
                  "max_tokens": 4000,
                  "temperature": 0.2,
              }
          )
          content = _response_text(resp)
          segments = parse_clean_response(content, asr_result)
          payload = {
              "source": _input_identity(work_dir, "asr_result.json"),
              "model": CONFIG["vlm_model"],
              "prompt": CLEAN_PROMPT,
              "postprocess_version": ASR_CLEAN_POSTPROCESS_VERSION,
              "segments": segments,
          }
          out_path.write_text(
              json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8"
          )
          log(f"consolidate(asr): 写出 asr_clean.json({len(segments)} 段)")
          return payload
      
      
      def consolidate_index(work_dir):
          work_dir = Path(work_dir)
          vlm_analysis = _load(work_dir, "vlm_analysis.json")
          if not vlm_analysis:
              log("consolidate(index): 无 vlm_analysis.json,跳过")
              return None
          out_path = work_dir / "understanding_index.json"
          if _fresh(out_path, work_dir / "vlm_analysis.json") and _index_cache_matches(
              work_dir, vlm_analysis
          ):
              log("consolidate(index): understanding_index.json 已最新,跳过")
              return _load(work_dir, "understanding_index.json")
          asr_result = _load(work_dir, "asr_result.json") or []
          asr_clean = _load(work_dir, "asr_clean.json") or {}
          background_research = load_background_research(work_dir)
          resp = api_call(
              {
                  "model": CONFIG["vlm_model"],
                  "messages": build_index_messages(
                      vlm_analysis,
                      asr_result=asr_result,
                      asr_clean=asr_clean,
                      background_research=background_research,
                  ),
                  "max_tokens": 3000,
                  "temperature": 0.2,
              }
          )
          index = parse_index_response(_response_text(resp))
          index = _apply_deterministic_asr_research_fallback(
              index,
              asr_result=asr_result,
              asr_clean=asr_clean,
              background_research=background_research,
          )
          out_path.write_text(
              json.dumps(index, ensure_ascii=False, indent=2), encoding="utf-8"
          )
          _write_index_meta(work_dir, vlm_analysis)
          (work_dir / "understanding_index.md").write_text(
              format_index_md(index), encoding="utf-8"
          )
          log(
              f"consolidate(index): 写出 understanding_index.json(角色 {len(index['characters'])})"
          )
          return index
      
      
      def consolidate(work_dir, do_asr=False, do_index=True):
          """Default = index-only (Pass B, zero timing risk). Pass A (asr) is opt-in.
      
          Failures propagate; understanding_runner catches them and writes the "failed" status."""
          result = {}
          if do_asr:
              result["asr_clean"] = consolidate_transcript(work_dir)
          if do_index:
              result["index"] = consolidate_index(work_dir)
          return result
      
      
      def _response_text(resp):
          try:
              return resp["choices"][0]["message"]["content"]
          except (KeyError, IndexError, TypeError):
              log("consolidate: API 返回结构异常")
              return ""
      
      
      def main():
          ap = argparse.ArgumentParser(
              description="Consolidate the understanding index (and optionally clean ASR)."
          )
          ap.add_argument("--work-dir", required=True)
          ap.add_argument("--asr", action="store_true", help="also run Pass A (ASR cleanup)")
          ap.add_argument("--no-index", action="store_true", help="skip Pass B (index)")
          args = ap.parse_args()
          res = consolidate(args.work_dir, do_asr=args.asr, do_index=not args.no_index)
          print(
              json.dumps(
                  {
                      "status": "consolidated",
                      "index": bool(res.get("index")),
                      "asr_clean": bool(res.get("asr_clean")),
                  },
                  ensure_ascii=False,
              )
          )
      
      
      if __name__ == "__main__":
          main()
      
    • deslop_qc.py 10.1 KB
      """Deterministic narration readability/QC scanner.
      
      This module is intentionally report-only. It flags local readability and
      packaging hygiene risks for video narration; it is not an AIGC detector and it
      never rewrites text.
      """
      
      from __future__ import annotations
      
      import json
      import re
      from pathlib import Path
      from typing import Any
      
      CONTRACT = (
          "Local readability/QC report only: this is not an AIGC detector, does not "
          "claim AI-generation accuracy, and never rewrites text. Corrections remain "
          "human/agent rewrite work."
      )
      
      EM_DASH_RE = re.compile(r"——|—")
      NEGATIVE_FLIP_RE = re.compile(r"并?不是[^。!?!?;;\.]{0,40}而是")
      # Catch verbatim copy of the brief's example scaffolding without over-matching real text.
      # The bracketed role token 【主角】 is the placeholder; anchor the pronoun signal to the
      # example's exact phrase "TA要赌上" so the legitimate gender-neutral pronoun "TA" (他/她) +
      # 要 (idiomatic 解说 suspense device "凶手就在其中,TA要做的下一件事…") is NOT hard-blocked.
      PLACEHOLDER_RE = re.compile(r"【主角】|TA要赌上")
      
      CLICHE_TERMS = [
          "命运的齿轮", "殊不知", "与此同时", "一场阴谋", "背后真相", "不为人知",
          "故事就此展开", "命运", "真相", "秘密", "危机", "反转", "救赎",
      ]
      ABSTRACT_TERMS = ["人性", "成长", "救赎", "命运", "时代", "宿命", "选择", "意义", "价值", "情感"]
      REASONING_MARKERS = ["因为", "所以", "因此", "也就是说", "换句话说", "这意味着", "原因是", "可见"]
      METAPHOR_MARKERS = ["像", "仿佛", "好似", "犹如", "如同", "一把", "一场", "一张", "棋局", "风暴"]
      CONNECTIVE_MARKERS = ["但", "却", "而", "于是", "随后", "直到", "结果", "因为", "所以", "这时", "最后"]
      
      
      def _sentence_pieces(text: str) -> list[str]:
          """Split text into sentence-like pieces while keeping terminal punctuation."""
          text = re.sub(r"\s+", " ", text).strip()
          if not text:
              return []
          parts = re.split(r"([。!?!?;;.])", text)
          out: list[str] = []
          for idx in range(0, len(parts), 2):
              body = parts[idx].strip()
              punct = parts[idx + 1] if idx + 1 < len(parts) else ""
              if body or punct:
                  out.append((body + punct).strip())
          return out or [text]
      
      
      def _text_units(text: str) -> int:
          """Length unit for prose: CJK characters, otherwise words."""
          if re.search(r"[\u3400-\u4dbf\u4e00-\u9fff\uf900-\ufaff]", text):
              return len(re.sub(r"\s+", "", text))
          return len(re.findall(r"\b\w+\b", text))
      
      
      def _normalise_segments(payload: list[dict[str, Any]], *, source: str) -> list[dict[str, Any]]:
          key = "narration" if source == "narration" else "text"
          segments: list[dict[str, Any]] = []
          for idx, item in enumerate(payload):
              text = str(item.get(key, "")).strip()
              if text:
                  segments.append({"source": source, "index": idx, "text": text})
          return segments
      
      
      def _load_original_subtitles(work_dir: Path | None) -> list[dict[str, Any]]:
          if work_dir is None:
              return []
          path = work_dir / "original_subtitles.json"
          if not path.exists():
              return []
          return _normalise_segments(json.loads(path.read_text(encoding="utf-8")), source="original_subtitles")
      
      
      def _style_card_requirement(work_dir: Path | None) -> tuple[bool, str]:
          """Read the explicit requirements contract; workspaces without one are advisory-only."""
          if work_dir is None:
              return False, "legacy_default"
          path = work_dir / "deslop_qc_requirements.json"
          if not path.exists():
              return False, "legacy_default"
          data = json.loads(path.read_text(encoding="utf-8"))
          return data["style_card_required"], "deslop_qc_requirements.json"
      
      
      def _style_card_issue(work_dir: Path | None, required: bool) -> dict[str, Any] | None:
          if work_dir is None:
              return None
          severity = "blocker" if required else "advisory"
          path = work_dir / "style_card.json"
          if not path.exists():
              return {
                  "severity": severity,
                  "code": "missing_style_card",
                  "source": "style_card",
                  "index": None,
                  "message": (
                      "style_card.json is required by this expression/packaging run but is missing"
                      if required else
                      "style_card.json is absent; legacy/migration workspaces may continue, but new expression-special runs should author it"
                  ),
              }
          data = json.loads(path.read_text(encoding="utf-8"))
          if not isinstance(data, dict) or not data:
              return {
                  "severity": severity,
                  "code": "malformed_style_card",
                  "source": "style_card",
                  "index": None,
                  "message": (
                      "style_card.json is required but empty or not a JSON object"
                      if required else
                      "style_card.json is present but empty or not a JSON object; treat as migration warning unless the run requires it"
                  ),
              }
          return None
      
      
      def _add_issue(bucket: list[dict[str, Any]], severity: str, code: str, source: str, index: int | None, message: str, **extra: Any) -> None:
          issue = {"severity": severity, "code": code, "source": source, "index": index, "message": message}
          issue.update({k: v for k, v in extra.items() if v is not None})
          bucket.append(issue)
      
      
      def analyze_deslop_qc(narration: list[dict[str, Any]], *, work_dir: str | Path | None = None, original_subtitles: list[dict[str, Any]] | None = None) -> dict[str, Any]:
          """Return a structured report; never mutates or rewrites input text."""
          work_path = Path(work_dir) if work_dir is not None else None
          segments = _normalise_segments(narration, source="narration")
          if original_subtitles is None:
              segments.extend(_load_original_subtitles(work_path))
          else:
              segments.extend(_normalise_segments(original_subtitles, source="original_subtitles"))
      
          blockers: list[dict[str, Any]] = []
          advisories: list[dict[str, Any]] = []
      
          required, requirement_source = _style_card_requirement(work_path)
          style_issue = _style_card_issue(work_path, required)
          if style_issue:
              (blockers if style_issue["severity"] == "blocker" else advisories).append(style_issue)
      
          all_text = "\n".join(seg["text"] for seg in segments)
          total_units = max(1, _text_units(all_text))
          sentences: list[str] = []
          for seg in segments:
              text = seg["text"]
              if EM_DASH_RE.search(text):
                  _add_issue(blockers, "blocker", "em_dash", seg["source"], seg["index"], "破折号(—/——)不得出现在 narration/original_subtitles 中", matches=EM_DASH_RE.findall(text))
              if PLACEHOLDER_RE.search(text):
                  _add_issue(blockers, "blocker", "placeholder_leakage", seg["source"], seg["index"], "示例占位内容泄漏到成稿中(如【主角】/TA要赌上)", matches=PLACEHOLDER_RE.findall(text))
              for sentence in _sentence_pieces(text):
                  sentences.append(sentence)
                  if NEGATIVE_FLIP_RE.search(sentence):
                      _add_issue(advisories, "advisory", "negative_positive_flip", seg["source"], seg["index"], "“不是…而是…”模板化转折偏多,建议改成更具体的因果/行动表达", sentence=sentence[:120])
      
          def term_count(terms: list[str]) -> int:
              return sum(all_text.count(term) for term in terms)
      
          cliche_count = term_count(CLICHE_TERMS)
          abstract_count = term_count(ABSTRACT_TERMS)
          reasoning_count = term_count(REASONING_MARKERS)
          metaphor_count = term_count(METAPHOR_MARKERS)
          connective_count = term_count(CONNECTIVE_MARKERS)
      
          if cliche_count >= 4 or (cliche_count / total_units) > 0.025:
              _add_issue(advisories, "advisory", "cliche_density", "narration", None, "套话/高频抽象词偏密,建议换成具体行动、选择和后果", count=cliche_count)
          if abstract_count >= 5 or (abstract_count / total_units) > 0.03:
              _add_issue(advisories, "advisory", "abstract_summary", "narration", None, "抽象总结词偏多,容易像概括论文;补足画面证据和具体抉择", count=abstract_count)
          if reasoning_count >= 4:
              _add_issue(advisories, "advisory", "reasoning_chain", "narration", None, "解释链标记偏多,建议减少“因为/所以/这意味着”等讲理口吻", count=reasoning_count)
          if metaphor_count >= 4:
              _add_issue(advisories, "advisory", "metaphor_markers", "narration", None, "比喻/包装化标记偏多,确认没有遮住剧情事实", count=metaphor_count)
          if total_units >= 180 and connective_count <= 1:
              _add_issue(advisories, "advisory", "low_connective_density", "narration", None, "长文本缺少转折/因果连接,可能像平铺摘要", connective_count=connective_count)
      
          long_segments = [seg for seg in segments if seg["source"] == "narration" and _text_units(seg["text"]) > 180]
          if long_segments:
              _add_issue(advisories, "advisory", "overlong_narration_block", "narration", long_segments[0]["index"], "单个解说块过长,建议拆成更可听的段落", max_units=max(_text_units(seg["text"]) for seg in long_segments))
          long_sentences = [s for s in sentences if _text_units(s) > 90]
          if long_sentences:
              _add_issue(advisories, "advisory", "long_paragraph", "narration", None, "长句/长段落偏多,字幕阅读压力较大", max_units=max(_text_units(s) for s in long_sentences))
      
          return {
              "ok": not blockers,
              "contract": CONTRACT,
              "scanner": "deslop_qc.py",
              "style_card_required": required,
              "style_card_requirement_source": requirement_source,
              "blocker_count": len(blockers),
              "advisory_count": len(advisories),
              "blockers": blockers,
              "advisories": advisories,
              "metrics": {
                  "segments_scanned": len(segments),
                  "text_units": total_units if all_text else 0,
                  "sentence_count": len(sentences),
                  "cliche_count": cliche_count,
                  "abstract_count": abstract_count,
                  "reasoning_marker_count": reasoning_count,
                  "metaphor_marker_count": metaphor_count,
                  "connective_count": connective_count,
              },
          }
      
      
      __all__ = ["CONTRACT", "analyze_deslop_qc"]
      
    • detect.py 21 KB
      import json
      import os
      import re
      import subprocess
      from concurrent.futures import ThreadPoolExecutor
      from pathlib import Path
      
      from lib import CONFIG
      from lib import log, run_cmd, get_video_duration, file_identity
      
      # ── Step 2: 场景检测 ──────────────────────────────────────────────────
      
      def detect_scenes(video_path, work_dir, threshold=None):
          """使用 ffmpeg scdet 滤镜检测场景切换"""
          threshold = CONFIG["scene_threshold"] if threshold is None else threshold
          scdet_threshold = int(threshold * 100)
      
          cmd = ["ffmpeg", "-i", str(video_path),
                 "-vf", f"scdet=threshold={scdet_threshold}",
                 "-f", "null", "-"]
          result = run_cmd(cmd)
          if result.returncode != 0:
              raise RuntimeError(f"场景检测失败: {result.stderr}")
      
          # 解析 lavfi.scd.time 和 lavfi.scd.score
          times = []
          for line in result.stderr.split("\n"):
              match = re.search(r"lavfi\.scd\.time[:=]\s*(\S+)", line)
              if match:
                  times.append(float(match.group(1)))
      
          if not times:
              # 没有检测到场景切换,整个视频作为一个场景
              duration = get_video_duration(video_path)
              scenes = [{"start": 0.0, "end": duration}]
              log(f"未检测到场景切换,整个视频作为一个场景 ({duration:.1f}s)")
          else:
              scenes = []
              prev = 0.0
              for t in times:
                  scenes.append({"start": round(prev, 2), "end": round(t, 2)})
                  prev = t
              # 最后一个场景到视频结束
              duration = get_video_duration(video_path)
              scenes.append({"start": round(prev, 2), "end": round(duration, 2)})
      
          log(f"检测到 {len(scenes)} 个场景")
      
          # 先过滤黑/白帧过渡场景,再合并短场景
          # (合并会保留短场景的 start,若黑场被并入长场景,单点采样会误删整段,故顺序在前)
          if CONFIG["scene_junk_filter"]:
              scenes = _filter_junk_scenes(scenes, video_path)
          # 合并短场景(< 3s 合并到相邻场景)
          scenes = _merge_short_scenes(scenes, min_duration=CONFIG["scene_merge_min"])
      
          # 保存
          scenes_file = work_dir / "scenes.json"
          scenes_file.write_text(json.dumps(scenes, ensure_ascii=False, indent=2), encoding="utf-8")
          for i, s in enumerate(scenes):
              log(f"  场景 {i+1}: {s['start']:.1f}s - {s['end']:.1f}s ({s['end']-s['start']:.1f}s)")
      
          return scenes
      
      
      def _merge_short_scenes(scenes, min_duration=4.0):
          """合并过短的场景到相邻场景"""
          if len(scenes) <= 1:
              return scenes
          merged = [scenes[0]]
          for s in scenes[1:]:
              prev = merged[-1]
              # 任一侧过短就把两段并成一段 [prev.start, s.end]。(原来写成 if/elif 两个分支,
              # 但两个分支体完全一样,读起来像是有两种不同处理,实际没有区别。)
              too_short = (
                  prev["end"] - prev["start"] < min_duration
                  or s["end"] - s["start"] < min_duration
              )
              if too_short:
                  merged[-1] = {"start": prev["start"], "end": s["end"]}
              else:
                  merged.append(s)
          # 如果最后一个太短,已在前一步合并
          log(f"合并短场景后: {len(scenes)} → {len(merged)} 个场景")
          return merged
      
      
      def _sample_frame_luma(video_path, timestamp, sample_size=64):
          """Extract one frame via ffmpeg and return luma values without extra deps."""
          cmd = [
              "ffmpeg",
              "-hide_banner",
              "-loglevel",
              "error",
              "-ss",
              f"{max(0.0, float(timestamp)):.3f}",
              "-i",
              str(video_path),
              "-frames:v",
              "1",
              "-vf",
              f"scale={sample_size}:{sample_size},format=rgb24",
              "-f",
              "rawvideo",
              "-",
          ]
          result = subprocess.run(cmd, capture_output=True)
          if result.returncode != 0 or not result.stdout:
              raise RuntimeError(result.stderr.decode("utf-8", errors="replace")[:300])
          data = result.stdout
          return [
              0.2126 * data[i] + 0.7152 * data[i + 1] + 0.0722 * data[i + 2]
              for i in range(0, len(data) - 2, 3)
          ]
      
      
      def _is_junk_scene(video_path, timestamp, threshold_dark=None, threshold_bright=None):
          """Return True for near-black or near-white scene-start frames."""
          threshold_dark = CONFIG["scene_junk_dark_luma"] if threshold_dark is None else threshold_dark
          threshold_bright = CONFIG["scene_junk_bright_luma"] if threshold_bright is None else threshold_bright
          pixel_ratio = min(1.0, float(CONFIG["scene_junk_pixel_ratio"]))
          try:
              lumas = _sample_frame_luma(video_path, timestamp)
          except Exception as exc:
              log(f"场景亮度采样失败,保留场景 {timestamp:.1f}s: {exc}")
              return False
          if not lumas:
              return False
          avg_luma = sum(lumas) / len(lumas)
          dark_ratio = sum(1 for value in lumas if value <= threshold_dark) / len(lumas)
          bright_ratio = sum(1 for value in lumas if value >= threshold_bright) / len(lumas)
          return (
              avg_luma <= threshold_dark and dark_ratio >= pixel_ratio
          ) or (
              avg_luma >= threshold_bright and bright_ratio >= pixel_ratio
          )
      
      
      def _scene_is_junk(scene, video_path):
          """Whether every probe point of one scene is a black/white transition frame."""
          start = float(scene["start"])
          end = float(scene["end"])
          # 多点采样:起始、中点、结尾各探一帧;只有全部为垃圾帧才删除,
          # 含任意非垃圾帧的场景必须保留(避免短黑场并入长真实场景后被整段误删)
          probe_times = [
              min(end, start + 0.1),
              (start + end) / 2.0,
              max(start, end - 0.1),
          ]
          # `all` 的短路很重要:绝大多数场景第一探就否掉,只花一次 ffmpeg。
          return all(_is_junk_scene(video_path, t) for t in probe_times)
      
      
      def _filter_junk_scenes(scenes, video_path):
          """Filter black/white transition scenes while never deleting the whole video."""
          if len(scenes) <= 1:
              return scenes
          # 每个探点都是一次 ffmpeg 进程(seek + 解一帧),一部 30 分钟片有几百个场景,串行下来
          # 光进程启动就要几十秒。按场景并行:每个 worker 内部仍然短路,所以既保留了「多数场景只
          # 探一次」,又把进程启动开销摊平。判定逻辑与探点完全不变。
          workers = int(CONFIG["scene_junk_workers"])
          with ThreadPoolExecutor(max_workers=min(workers, len(scenes))) as executor:
              verdicts = list(executor.map(lambda s: _scene_is_junk(s, video_path), scenes))
          filtered = []
          removed = []
          for scene, is_junk in zip(scenes, verdicts):
              if is_junk:
                  removed.append(scene)
              else:
                  filtered.append(scene)
          if not filtered:
              log("场景黑/白帧过滤会删除全部场景,已放弃过滤")
              return scenes
          if removed:
              log(f"过滤黑/白帧过渡场景: {len(scenes)} → {len(filtered)}")
          return filtered
      
      
      
      # ── Step 3.5: 静音检测 ─────────────────────────────────────────────
      
      # The silence-detection audio is always extracted below as 16 kHz mono signed-16 PCM.
      _SILENCE_AUDIO_BYTES_PER_SECOND = 16000 * 1 * 2
      
      
      def _wav_seconds_estimate(audio_path):
          """Approximate duration from file size — no ffprobe subprocess on this path."""
          try:
              return max(0.0, Path(audio_path).stat().st_size / _SILENCE_AUDIO_BYTES_PER_SECOND)
          except OSError:
              return 0.0
      
      
      def _compact_ffmpeg_error(stderr, limit=400):
          """Keep the actionable tail without dumping ffmpeg build/configuration banners."""
          text = " ".join(str(stderr or "").split())
          if len(text) <= limit:
              return text
          return "…" + text[-limit:]
      
      def annotate_quiet_windows_with_asr(periods, asr_result=None, *, video_duration=None, configured_segment_seconds=None):
          """Pure helper: annotate quiet windows with ASR-overlap confidence and QC.
      
          Coarse grid ASR (large synthetic windows from chunking) must not turn every quiet
          window into speech. The return value is (annotated_periods, qc). Inputs are copied.
          """
          out = [dict(p) for p in (periods or [])]
          qc = {"coarse_asr_windows": 0, "low_confidence_speech_flags": 0, "asr_granularity": "none"}
          for qp in out:
              qp.setdefault("speech_overlap_ratio", 0.0)
              qp.setdefault("asr_overlap_seconds", 0.0)
              qp.setdefault("asr_granularity", "none")
              qp.setdefault("has_speech", False)
              qp.setdefault("has_speech_reason", "no_asr_overlap")
          if not asr_result:
              return out, qc
      
          valid_segments = []
          for seg in asr_result:
              ss, se = float(seg["start"]), float(seg["end"])
              if se > ss:
                  valid_segments.append((ss, se))
          asr_coverage = sum(se - ss for ss, se in valid_segments)
          avg_seg_dur = asr_coverage / len(valid_segments) if valid_segments else 0
          if video_duration is None:
              ends = [float(p["end"]) for p in out] + [se for _, se in valid_segments]
              video_duration = max(ends or [0.0])
          configured = float(configured_segment_seconds if configured_segment_seconds is not None else CONFIG["asr_segment_seconds"])
          coarse_asr = (
              (len(valid_segments) <= 5 and asr_coverage > float(video_duration) * 0.8) or
              asr_coverage > float(video_duration) * 1.5 or
              avg_seg_dur > max(45.0, configured * 1.5) or
              (avg_seg_dur >= configured * 0.9 and asr_coverage > float(video_duration) * 0.7)
          )
          granularity = "coarse_grid" if coarse_asr else "segment"
          qc["asr_granularity"] = granularity
          qc["avg_asr_segment_seconds"] = round(avg_seg_dur, 3)
          for qp in out:
              overlap_seconds = 0.0
              for ss, se in valid_segments:
                  overlap_seconds += max(0.0, min(float(qp["end"]), se) - max(float(qp["start"]), ss))
              ratio = overlap_seconds / max(0.001, float(qp["duration"]))
              qp["asr_overlap_seconds"] = round(overlap_seconds, 3)
              qp["speech_overlap_ratio"] = round(ratio, 4)
              qp["asr_granularity"] = granularity
              if coarse_asr:
                  qp["has_speech"] = False
                  qp["has_speech_reason"] = "coarse_asr_overlap_ignored" if overlap_seconds > 0 else "coarse_asr_no_overlap"
                  if overlap_seconds > 0:
                      qc["coarse_asr_windows"] += 1
                  continue
              if ratio >= 0.3:
                  qp["has_speech"] = True
                  qp["has_speech_reason"] = "asr_overlap_high_confidence"
              elif overlap_seconds > 0:
                  qp["has_speech"] = False
                  qp["has_speech_reason"] = "asr_overlap_low_confidence_quiet"
                  qc["low_confidence_speech_flags"] += 1
              else:
                  qp["has_speech"] = False
                  qp["has_speech_reason"] = "no_asr_overlap"
          return out, qc
      
      
      def detect_silence_periods(video_path, work_dir, asr_result=None):
          """用 ffmpeg silencedetect 检测安静时段,作为解说插入的候选窗口"""
          audio_path = work_dir / "audio.wav"
          if not _audio_cache_matches(audio_path, video_path):
              if audio_path.exists():
                  audio_path.unlink()
              # 提取到临时文件,成功后原子移动到位,避免被中断的 -y 运行留下半截 audio.wav
              tmp_path = work_dir / "audio.wav.tmp"
              extract = run_cmd([
                  "ffmpeg", "-y", "-i", str(video_path),
                  "-vn", "-ar", "16000", "-ac", "1",
                  "-f", "wav", str(tmp_path)  # .tmp extension hides the format from ffmpeg; state it
              ])
              if extract.returncode != 0 or not tmp_path.exists():
                  log(
                      "音频提取失败,无法检测静音窗口(视频可能无音轨): "
                      f"{_compact_ffmpeg_error(extract.stderr)}"
                  )
                  if tmp_path.exists():
                      tmp_path.unlink()
                  return []
              os.replace(str(tmp_path), str(audio_path))
              _write_audio_meta(work_dir, video_path)
      
          # Sentence-entry anchors use short acoustic pauses aligned to terminal ASR punctuation.
          # They are deliberately separate from silence_periods.json: a 200ms sentence pause is a
          # safe place to ENTER narration, but not a multi-second quiet window that can own a block.
          detect_speech_boundary_anchors(work_dir, asr_result or [])
      
          noise = CONFIG["silence_noise_threshold"]
          min_dur = CONFIG["silence_min_duration"]
          cmd = ["ffmpeg", "-i", str(audio_path),
                 "-af", f"silencedetect=noise={noise}:d={min_dur}",
                 "-f", "null", "-"]
          # Everything else in this function degrades to `return []` (narration then falls back to
          # non-quiet placement). A timeout must degrade the same way instead of raising
          # TimeoutExpired through the whole understanding stage. Scale the budget with the audio
          # so a long feature is not cut off by a constant tuned for short clips — estimated from
          # the file size (this is the 16 kHz mono s16 WAV we just wrote), not from another probe.
          timeout = max(120.0, _wav_seconds_estimate(audio_path) * 0.5)
          try:
              result = run_cmd(cmd, timeout=timeout)
          except subprocess.TimeoutExpired:
              log(f"静音检测超时(>{timeout:.0f}s),跳过安静窗口检测")
              return []
          if result.returncode != 0:
              log(f"静音检测失败: {_compact_ffmpeg_error(result.stderr)}")
              return []
          output = result.stderr
      
          # 解析 silence_start / silence_end
          starts = [float(m) for m in re.findall(r'silence_start:\s*([\d.]+)', output)]
          ends = [float(m) for m in re.findall(r'silence_end:\s*([\d.]+)', output)]
      
          # 配对所有静音段(不过滤时长,后面合并后再过滤)
          raw_periods = []
          for s, e in zip(starts, ends):
              raw_periods.append({"start": round(s, 2), "end": round(e, 2),
                                  "duration": round(e - s, 2)})
          # 末尾静音
          if len(starts) > len(ends):
              dur_start = starts[len(ends)]
              total_dur = get_video_duration(str(audio_path))
              raw_periods.append({"start": round(dur_start, 2), "end": round(total_dur, 2),
                                  "duration": round(total_dur - dur_start, 2)})
      
          # 合并相邻静音段(间隔 < merge_gap 的合并为一个大窗口)
          merge_gap = CONFIG["silence_merge_gap"]
          merged = []
          for rp in sorted(raw_periods, key=lambda x: x["start"]):
              if merged and rp["start"] - merged[-1]["end"] < merge_gap:
                  merged[-1]["end"] = rp["end"]
                  merged[-1]["duration"] = round(merged[-1]["end"] - merged[-1]["start"], 2)
              else:
                  merged.append({"start": rp["start"], "end": rp["end"],
                                 "duration": rp["duration"], "has_speech": False})
      
          # 过滤最短时长
          quiet_min = CONFIG["quiet_window_min"]
          periods = [p for p in merged if p["duration"] >= quiet_min]
      
          # 与 ASR 交叉验证:标记有语音的窗口,并记录可观测 confidence/reason。
          periods, qc = annotate_quiet_windows_with_asr(
              periods,
              asr_result,
              video_duration=get_video_duration(str(audio_path)) if asr_result else None,
              configured_segment_seconds=CONFIG["asr_segment_seconds"],
          )
      
          (work_dir / "silence_periods.qc.json").write_text(
              json.dumps(qc, ensure_ascii=False, indent=2), encoding="utf-8")
      
          # 保存
          (work_dir / "silence_periods.json").write_text(
              json.dumps(periods, ensure_ascii=False, indent=2), encoding="utf-8")
          log(f"检测到 {len(periods)} 个安静窗口 (≥{quiet_min}s)")
          for qp in periods:
              flag = " [有语音]" if qp["has_speech"] else ""
              log(f"  {qp['start']:.1f}s-{qp['end']:.1f}s ({qp['duration']:.1f}s){flag}")
          return periods
      
      
      def detect_speech_boundary_anchors(work_dir, asr_result):
          """Write sentence-end entry anchors by aligning ASR punctuation to short pauses.
      
          MiMo ASR timestamps are window-level, not word-level. We therefore estimate each terminal
          punctuation time from its character position inside the ASR window, then snap it to the
          closest short acoustic pause. The output is guidance + a deterministic pre-TTS gate; it
          never rewrites narration timing on its own.
          """
          work_dir = Path(work_dir)
          audio_path = work_dir / "audio.wav"
          out_path = work_dir / "speech_boundary_anchors.json"
          segments = list(asr_result or [])
          report = {
              "schema_version": 1,
              "artifact": "speech_boundary_anchors.json",
              "status": "completed",
              "detector": {
                  "noise_threshold": CONFIG["source_boundary_noise_threshold"],
                  "min_pause_seconds": float(CONFIG["source_boundary_min_pause"]),
                  "alignment": "terminal_punctuation_to_nearest_acoustic_pause",
              },
              "sentence_anchors": [],
              "acoustic_pauses": [],
          }
          if not audio_path.exists() or not segments:
              report["status"] = "unavailable"
              report["reason"] = "missing_audio_or_asr"
              out_path.write_text(json.dumps(report, ensure_ascii=False, indent=2), encoding="utf-8")
              return report
      
          result = run_cmd([
              "ffmpeg", "-hide_banner", "-nostats", "-i", str(audio_path),
              "-af", (
                  "silencedetect="
                  f"noise={CONFIG['source_boundary_noise_threshold']}:"
                  f"d={float(CONFIG['source_boundary_min_pause'])}"
              ),
              "-f", "null", "-",
          ], timeout=120)
          if result.returncode != 0:
              report["status"] = "failed"
              report["reason"] = "silencedetect_failed"
              out_path.write_text(json.dumps(report, ensure_ascii=False, indent=2), encoding="utf-8")
              return report
      
          starts = [float(value) for value in re.findall(r"silence_start:\s*([\d.]+)", result.stderr or "")]
          ends = [float(value) for value in re.findall(r"silence_end:\s*([\d.]+)", result.stderr or "")]
          pauses = []
          for index, (start, end) in enumerate(zip(starts, ends)):
              if end <= start:
                  continue
              pauses.append({
                  "index": index,
                  "start": round(start, 3),
                  "end": round(end, 3),
                  "midpoint": round((start + end) / 2.0, 3),
                  "duration": round(end - start, 3),
              })
          report["acoustic_pauses"] = pauses
      
          max_error = float(CONFIG["source_boundary_max_alignment_error"])
          used = set()
          anchors = []
          for asr_index, segment in enumerate(segments):
              seg_start = float(segment["start"])
              seg_end = float(segment["end"])
              text = segment["text"].strip()
              if seg_end <= seg_start or not text:
                  continue
              last_midpoint = seg_start - 1e-6
              for match in re.finditer(r"[。!?!?;;]", text):
                  expected = seg_start + (seg_end - seg_start) * (match.end() / max(1, len(text)))
                  candidates = [
                      pause for pause in pauses
                      if pause["index"] not in used
                      and seg_start - 0.3 <= pause["midpoint"] <= seg_end + 0.3
                      and pause["midpoint"] > last_midpoint
                      and abs(pause["midpoint"] - expected) <= max_error
                  ]
                  if not candidates:
                      continue
                  pause = min(candidates, key=lambda item: abs(item["midpoint"] - expected))
                  used.add(pause["index"])
                  last_midpoint = pause["midpoint"]
                  error = abs(pause["midpoint"] - expected)
                  confidence = "high" if error <= 0.6 else ("medium" if error <= 1.2 else "low")
                  anchors.append({
                      "time": pause["end"],
                      "pause_start": pause["start"],
                      "pause_end": pause["end"],
                      "expected_time": round(expected, 3),
                      "alignment_error": round(error, 3),
                      "confidence": confidence,
                      "punctuation": match.group(0),
                      "text_tail": text[max(0, match.end() - 32):match.end()],
                      "asr_segment_index": asr_index,
                  })
          report["sentence_anchors"] = sorted(anchors, key=lambda item: item["time"])
          out_path.write_text(json.dumps(report, ensure_ascii=False, indent=2), encoding="utf-8")
          log(f"检测到 {len(anchors)} 个原声句末安全切入点")
          return report
      
      
      def _audio_meta_path(work_dir):
          return Path(work_dir) / "audio.wav.meta.json"
      
      
      def _audio_cache_matches(audio_path, video_path):
          audio_path = Path(audio_path)
          if not audio_path.exists():
              return False
          meta_path = _audio_meta_path(audio_path.parent)
          if not meta_path.exists():
              return False
          try:
              meta = json.loads(meta_path.read_text(encoding="utf-8"))
          except (OSError, ValueError, TypeError):
              return False
          try:
              expected = file_identity(video_path)
          except OSError:
              return False
          return (
              isinstance(meta, dict)
              and meta.get("source_video") == str(Path(video_path).resolve())
              and meta.get("source_video_identity") == expected
          )
      
      
      def _write_audio_meta(work_dir, video_path):
          _audio_meta_path(work_dir).write_text(
              json.dumps({
                  "schema_version": 1,
                  "source_video": str(Path(video_path).resolve()),
                  "source_video_identity": file_identity(video_path),
                  "audio": "audio.wav",
              }, ensure_ascii=False, indent=2),
              encoding="utf-8",
          )
      
    • extract.py 2.5 KB
      from lib import CONFIG
      from lib import log, run_cmd
      
      # ── Step 1: 帧提取 ───────────────────────────────────────────────────
      
      # ffmpeg 的 `fps` 滤镜从源时间 0 开始输出第一帧,而 image2 muxer 的 `-start_number`
      # 默认是 1。所以 frame_00001.jpg 对应 t=0,第 n 号帧对应 (n-1)/fps —— 不是 n/fps。
      # 这个 off-by-one 会把所有画面锚点(frame_facts、场景归属、storyboard 标签)整体推后
      # 1/fps 秒(长视频默认 fps=1,即整整 1 秒)。这里是该约定的唯一定义处。
      FRAME_START_NUMBER = 1
      
      # 帧编号↔时间的换算规则变更时递增;进入各阶段缓存 key,强制重算旧产物。
      FRAME_TIME_CONVENTION_VERSION = 2
      
      
      def frame_time(frame_number, fps):
          """帧文件序号 → 源视频时间(秒)。frame_00001 → 0.0。"""
          fps = float(fps)
          if fps <= 0:
              raise ValueError("fps 必须大于 0")
          return (int(frame_number) - FRAME_START_NUMBER) / fps
      
      
      def frame_number_for_time(timestamp, fps):
          """源视频时间(秒)→ 对应的帧文件序号(可能是小数,供最近邻查找用)。"""
          fps = float(fps)
          if fps <= 0:
              raise ValueError("fps 必须大于 0")
          return float(timestamp) * fps + FRAME_START_NUMBER
      
      
      def parse_frame_number(frame_path):
          """从 frame_NNNNN.jpg 解析序号;命名不符合约定时返回 None。"""
          parts = frame_path.stem.split("_")
          if len(parts) != 2 or not parts[1].isdigit():
              return None
          return int(parts[1])
      
      
      def extract_frames(video_path, work_dir, fps=None):
          """提取视频帧"""
          fps = CONFIG["fps"] if fps is None else fps
          if fps <= 0:
              raise ValueError("fps 必须大于 0;完整 pipeline 会自动计算 fps")
          frames_dir = work_dir / "frames"
          frames_dir.mkdir(exist_ok=True)
      
          # 清理上一次(可能更高 fps)残留的帧,避免陈旧帧泄漏进本次结果
          for stale in frames_dir.glob("frame_*.jpg"):
              stale.unlink()
      
          output_pattern = str(frames_dir / "frame_%05d.jpg")
          cmd = ["ffmpeg", "-y", "-i", str(video_path),
                 "-vf", f"fps={fps}", "-q:v", "2",
                 "-start_number", str(FRAME_START_NUMBER), output_pattern]
          result = run_cmd(cmd)
          if result.returncode != 0:
              raise RuntimeError(f"帧提取失败: {result.stderr}")
      
          frames = sorted(frames_dir.glob("frame_*.jpg"))
          log(f"提取了 {len(frames)} 帧 ({fps}fps)")
          return frames
      
    • lib.py 21.6 KB
      """Self-contained config + utilities for this skill (no cross-skill imports).
      Merged from the shared core; reads the same env vars as the rest of the bundle."""
      import json
      import math
      import os
      import re
      import subprocess
      import time
      import urllib.request
      import urllib.error
      from pathlib import Path
      from datetime import datetime, timezone
      from email.utils import parsedate_to_datetime
      
      
      # ── 配置 ──────────────────────────────────────────────────────────────
      
      DEFAULT_MIMO_API_URL = "https://api.xiaomimimo.com/v1"
      DEFAULT_MIMO_TOKEN_PLAN_CLUSTER = "cn"
      MIMO_TOKEN_PLAN_API_URLS = {
          "cn": "https://token-plan-cn.xiaomimimo.com/v1",
          "sgp": "https://token-plan-sgp.xiaomimimo.com/v1",
          "ams": "https://token-plan-ams.xiaomimimo.com/v1",
      }
      DEFAULT_MIMO_MODEL = "mimo-v2.5"          # VLM / chat (vision understanding)
      DEFAULT_MIMO_ASR_MODEL = "mimo-v2.5-asr"  # speech-to-text
      
      
      def normalize_api_url(raw_url):
          """Normalize a MiMo (OpenAI-compatible) base URL or chat/completions endpoint."""
          url = raw_url.rstrip("/")
          if url.endswith("/chat/completions"):
              return url
          return f"{url}/chat/completions"
      
      
      def is_mimo_token_plan_key(api_key):
          """Return True for Xiaomi MiMo Token Plan keys, which use token-plan base URLs."""
          return str(api_key or "").strip().startswith("tp-")
      
      
      def default_mimo_api_url(is_token_plan, cluster=None):
          """Pick the correct MiMo base URL for pay-as-you-go vs Token Plan keys.
      
          MiMo uses independent credentials for pay-as-you-go (`sk-*`) and Token Plan
          (`tp-*`). Token Plan keys must be sent to the Token Plan cluster base URL,
          not the pay-as-you-go `api.xiaomimimo.com` endpoint.
      
          The caller classifies its own key with `is_mimo_token_plan_key` and passes only
          that bit: a credential never reaches a function whose return value is logged.
          """
          if is_token_plan:
              cluster_name = (cluster or os.environ.get("MIMO_TOKEN_PLAN_CLUSTER") or DEFAULT_MIMO_TOKEN_PLAN_CLUSTER)
              cluster_name = str(cluster_name).strip().lower()
              if cluster_name not in MIMO_TOKEN_PLAN_API_URLS:
                  raise ValueError(
                      f"MiMo token-plan cluster must be one of {sorted(MIMO_TOKEN_PLAN_API_URLS)}; "
                      f"got {cluster_name!r}"
                  )
              return MIMO_TOKEN_PLAN_API_URLS[cluster_name]
          return DEFAULT_MIMO_API_URL
      
      
      def _env_number(name, default, cast, minimum):
          raw = os.environ.get(name)
          if raw is None or raw == "":
              return default
          try:
              value = cast(raw)
          except ValueError as exc:
              raise ValueError(f"{name} must be a {cast.__name__}; got {raw!r}") from exc
          if isinstance(value, float) and not math.isfinite(value):
              raise ValueError(f"{name} must be finite; got {raw!r}")
          if minimum is not None and value < minimum:
              raise ValueError(f"{name} must be >= {minimum}; got {value}")
          return value
      
      
      def env_int(name, default, *, minimum=None):
          """Read an integer env var, rejecting malformed or below-minimum values."""
          return _env_number(name, default, int, minimum)
      
      
      def env_bool(name, default=False):
          """Read common boolean env var forms."""
          raw = os.environ.get(name)
          if raw is None or raw == "":
              return default
          return raw.strip().lower() in {"1", "true", "yes", "y", "on"}
      
      
      def env_float(name, default, *, minimum=None):
          """Read a float env var, rejecting malformed or below-minimum values."""
          return _env_number(name, default, float, minimum)
      
      
      # Single MiMo credential powers ASR + VLM + TTS. Per-capability overrides
      # (MIMO_VIDEO_API_KEY / MIMO_TTS_API_KEY / MIMO_ASR_API_KEY and their *_API_URL forms)
      # are optional and fall back to MIMO_API_KEY / MIMO_API_URL. Token-Plan keys (tp-*) auto-
      # route to the Token-Plan cluster base URL; pay-as-you-go keys use api.xiaomimimo.com.
      _mimo_api_key = os.environ.get("MIMO_API_KEY", "")
      _mimo_video_api_key = os.environ.get("MIMO_VIDEO_API_KEY", "") or _mimo_api_key
      _mimo_asr_api_key = os.environ.get("MIMO_ASR_API_KEY", "") or _mimo_api_key
      _raw_api_url = os.environ.get("MIMO_API_URL") or default_mimo_api_url(is_mimo_token_plan_key(_mimo_api_key))
      _raw_mimo_video_api_url = (
          os.environ.get("MIMO_VIDEO_API_URL")
          or os.environ.get("MIMO_API_URL")
          or default_mimo_api_url(is_mimo_token_plan_key(_mimo_video_api_key))
      )
      _raw_mimo_asr_api_url = (
          os.environ.get("MIMO_ASR_API_URL")
          or os.environ.get("MIMO_API_URL")
          or default_mimo_api_url(is_mimo_token_plan_key(_mimo_asr_api_key))
      )
      
      CONFIG = {
          "api_url": normalize_api_url(_raw_api_url),
          "api_key": _mimo_api_key,
          "api_env_var": "MIMO_API_KEY",
          "mimo_api_url": normalize_api_url(_raw_api_url),
          "mimo_api_key": _mimo_api_key,
          "mimo_video_api_url": normalize_api_url(_raw_mimo_video_api_url),
          "mimo_video_api_key": _mimo_video_api_key,
          "mimo_asr_api_url": normalize_api_url(_raw_mimo_asr_api_url),
          "mimo_asr_api_key": _mimo_asr_api_key,
          "mimo_asr_env_var": "MIMO_ASR_API_KEY" if os.environ.get("MIMO_ASR_API_KEY") else "MIMO_API_KEY",
          "mimo_model": os.environ.get("MIMO_MODEL", DEFAULT_MIMO_MODEL),
          "mimo_video_model": os.environ.get("MIMO_VIDEO_MODEL") or os.environ.get("MIMO_MODEL", DEFAULT_MIMO_MODEL),
          "vlm_model": os.environ.get("MIMO_MODEL", DEFAULT_MIMO_MODEL),
          "mimo_asr_model": os.environ.get("MIMO_ASR_MODEL", DEFAULT_MIMO_ASR_MODEL),
          "mimo_asr_language": os.environ.get("MIMO_ASR_LANGUAGE", "auto"),  # auto | zh | en
          "mimo_asr_base64_max_mb": env_float("MIMO_ASR_BASE64_MAX_MB", 10.0, minimum=1.0),
          # ASR 分段窗口秒数。越小 → 长视频的对白时间戳越精细(默认 15s)。旧值 180s 会把 >3min
          # 视频的对白塌缩成一个时间戳,既让 brief 无法定位对白,又触发 detect.py 的粗粒度跳过,
          # 使 overlaps_speech/安静窗口判断失真。代价是更多 ASR 调用;ASR 慢时可调大。
          "asr_segment_seconds": env_float("ASR_SEGMENT_SECONDS", 15.0, minimum=5.0),
          "scene_threshold": 0.1,
          "mimo_media_resolution": os.environ.get("MIMO_MEDIA_RESOLUTION", "default"),
          "mimo_video_overview": env_bool("MIMO_VIDEO_OVERVIEW", False),  # opt-in (--mimo-video-overview / =1); when on it becomes the PRIMARY per-scene description, frames stay the anchor/fallback
          "mimo_video_fps": env_float("MIMO_VIDEO_FPS", 3.0, minimum=0.1),
          "mimo_video_chunk_max_seconds": env_float("MIMO_VIDEO_CHUNK_MAX_SECONDS", 20.0, minimum=1.0),
          "mimo_video_chunk_min_seconds": env_float("MIMO_VIDEO_CHUNK_MIN_SECONDS", 1.0, minimum=0.2),
          "mimo_video_chunk_timeout": env_int("MIMO_VIDEO_CHUNK_TIMEOUT", 180, minimum=1),
          "mimo_video_base64_max_mb": env_float("MIMO_VIDEO_BASE64_MAX_MB", 45.0, minimum=1.0),
          # Per-scene frame VLM sampling — scale frames with scene length instead of a hard cap of 6
          "vlm_seconds_per_frame": env_float("VLM_SECONDS_PER_FRAME", 4.0, minimum=0.5),
          "vlm_max_frames": env_int("VLM_MAX_FRAMES", 16, minimum=3),
          "vlm_max_tokens": env_int("VLM_MAX_TOKENS", 1500, minimum=200),
          "mimo_video_prompt": os.environ.get(
              "MIMO_VIDEO_PROMPT",
              "请用中文分析这个视频分片的主要人物、场景变化、关键动作、情绪走向和剧情冲突,"
              "重点提取适合写短视频解说的故事线索。不要泛泛复述画面,要标出对后续写稿有用的信息。",
          ),
          "mimo_disable_thinking": env_bool("MIMO_DISABLE_THINKING", True),
          "fps": 0,  # 0 = 自动(≤60s→2fps, ≤5min→1.5fps, >5min→1fps)
          # Storyboard contact sheets are advisory and generated locally where frames/fps are owned.
          # Keep these keys only in the stage that consumes them; other stages do not read storyboard_*.
          "storyboard": env_bool("STORYBOARD", True),  # generate source/edited storyboard contact sheets
          "storyboard_max_tiles": env_int("STORYBOARD_MAX_TILES", 30, minimum=1),  # cap tiles per sheet for legibility
          "storyboard_columns": env_int("STORYBOARD_COLUMNS", 6, minimum=1),  # tile grid columns
          "storyboard_rows_per_page": env_int("STORYBOARD_ROWS_PER_PAGE", 5, minimum=1),  # tile grid rows before spilling to the next sheet
          "storyboard_long_scene_seconds": env_float("STORYBOARD_LONG_SCENE_SECONDS", 6.0, minimum=0.1),  # scenes at/above this also sample +1/3 & +2/3
          # TTS 语速(字符/秒)。实测 mimo-tts 冰糖音色中位 ~3.9 字/秒,可用 SPEECH_RATE 覆盖
          # 生成解说时使用 speech_rate * safety_margin 作为约束
          "speech_rate": env_float("SPEECH_RATE", 3.9, minimum=0.5),  # 旧值 3.5 系统性偏低 ~10-17%
          "speech_safety_margin": env_float("SPEECH_SAFETY_MARGIN", 0.85, minimum=0.1),  # 保守系数:TTS 实际语速有 ±20% 波动
          # Block-coverage lint thresholds — promoted from inline .get() literals to real CONFIG keys (tunable; defaults unchanged)
          "narration_coverage_target": 0.7,   # rough first-draft/diagnostic fallback; content-led audio decisions may differ (not a quota)
          "narration_coverage_max": 0.85,     # above this coverage → no_original_blocks (narration is wall-to-wall)
          "narration_coverage_min": 0.5,      # below this coverage → under_narrated
          "narration_block_seconds": 9.0,     # block cadence used to derive target block count
          "original_block_min_seconds": 2.5,  # a deliberate original-audio gap must be at least this long
          "narration_block_min_chars": 16,    # below this avg block size → fragmented_beats
          "breath_ms": 250,  # 段间呼吸空间(ms);block recap 块内连贯、块间留原声呼吸
          "narration_speed": env_float("NARRATION_SPEED", 1.15, minimum=0.5),  # 解说整体提速(atempo),默认回到可懂区间;长片可设 1.0
          "narration_tail_pad_seconds": 0.1,  # 解说尾部最少留白;短 slot 会自动压低 delay 避免截断
          "quiet_overlap_min_ratio": 0.8,  # 解说段至少多少比例落在安静窗口内才标记为非对白重叠
          "visual_beat_max_seconds": 18.0,  # 单段解说超过该时长且跨多个帧锚点时给 lint 提醒
          "visual_beat_max_facts": 3,  # 单段解说最多建议覆盖的 frame_facts 锚点数量
          "asr_chunk_min_chars": env_int("ASR_CHUNK_MIN_CHARS", 500, minimum=1),  # brief 中 ASR 写作分块最小字数/词数
          "asr_chunk_max_chars": env_int("ASR_CHUNK_MAX_CHARS", 800, minimum=1),  # brief 中 ASR 写作分块最大字数/词数
          "silence_noise_threshold": "-25dB",  # ffmpeg silencedetect 噪声阈值
          # Short acoustic pauses are not quiet windows; they are sentence-boundary candidates.
          # A looser threshold plus ASR-punctuation alignment lets the writer enter after a full
          # source sentence instead of ducking it halfway through.
          "source_boundary_noise_threshold": os.environ.get("SOURCE_BOUNDARY_NOISE_THRESHOLD", "-18dB"),
          "source_boundary_min_pause": env_float("SOURCE_BOUNDARY_MIN_PAUSE", 0.12, minimum=0.05),
          "source_boundary_max_alignment_error": env_float(
              "SOURCE_BOUNDARY_MAX_ALIGNMENT_ERROR", 2.1, minimum=0.2
          ),
          "silence_min_duration": 0.3,     # 静音最短持续秒数
          "quiet_window_min": 1.0,         # 可放解说的安静窗口最短秒数
          "silence_merge_gap": 0.5,        # 相邻静音段间隔<此值时合并
          "scene_merge_min": 4.0,         # 场景合并最短时长,<此值的场景合并到相邻场景
          "scene_junk_filter": env_bool("SCENE_JUNK_FILTER", True),  # 过滤连续黑/白帧无效过渡场景
          "scene_junk_workers": env_int("SCENE_JUNK_WORKERS", 8, minimum=1),  # 黑/白帧探测并行度(每个探点一次 ffmpeg)
          "scene_junk_dark_luma": env_float("SCENE_JUNK_DARK_LUMA", 8.0, minimum=0.0),
          "scene_junk_bright_luma": env_float("SCENE_JUNK_BRIGHT_LUMA", 245.0, minimum=0.0),
          "scene_junk_pixel_ratio": env_float("SCENE_JUNK_PIXEL_RATIO", 0.995, minimum=0.0),
          "context_info": "",              # 额外上下文(节目名、角色名等)
          "vlm_workers": env_int("VLM_WORKERS", 8, minimum=1),  # VLM 并行分析线程数
          "edit_mode": os.environ.get("EDIT_MODE", "full"),  # full | cut
          "target_duration": os.environ.get("TARGET_DURATION", ""),  # cut 模式目标成片时长,如 10m
      }
      
      SCRIPT_DIR = Path(__file__).parent
      PROMPTS_DIR = SCRIPT_DIR.parent / "references"
      
      def log(msg):
          print(f"[video-recap] {msg}", flush=True)
      
      def run_cmd(cmd, **kwargs):
          """运行命令,返回 CompletedProcess"""
          if isinstance(cmd, list):
              display_parts = []
              for part in cmd:
                  text = str(part)
                  display_parts.append(text if len(text) <= 240 else text[:237] + "...")
              display = " ".join(display_parts)
          else:
              display = str(cmd)
              if len(display) > 2000:
                  display = display[:1997] + "..."
          log(f"运行: {display}")
          return subprocess.run(cmd, capture_output=True, text=True, **kwargs)
      
      def get_video_duration(video_path):
          """获取视频时长(秒)"""
          cmd = ["ffprobe", "-v", "quiet", "-show_entries", "format=duration",
                 "-of", "csv=p=0", str(video_path)]
          result = run_cmd(cmd)
          if result.returncode != 0:
              raise RuntimeError(f"ffprobe 无法读取时长 {video_path}: {result.stderr.strip()[-500:]}")
          try:
              return float(result.stdout.strip())
          except ValueError as exc:
              raise RuntimeError(f"ffprobe 时长输出无法解析 {video_path}: {result.stdout.strip()[:100]!r}") from exc
      
      
      def load_background_research(work_dir):
          """Load the agent-authored background_research.json: missing → {}, malformed → raise."""
          path = Path(work_dir) / "background_research.json"
          if not path.exists():
              return {}
          try:
              data = json.loads(path.read_text(encoding="utf-8"))
          except ValueError as exc:
              raise ValueError(f"background_research.json 不是合法 JSON,请修复后重试: {exc}") from exc
          if not isinstance(data, dict):
              raise ValueError("background_research.json 顶层必须是 JSON 对象({...})")
          return data
      
      def file_identity(path):
          """{"size", "mtime_ns"} of a file: the cache identity recorded instead of content hashes."""
          st = os.stat(os.fspath(path))
          return {"size": st.st_size, "mtime_ns": st.st_mtime_ns}
      
      
      def _retry_after_seconds(value, fallback):
          """Parse Retry-After seconds or HTTP-date; return fallback on malformed input."""
          if not value:
              return fallback
          try:
              return max(fallback, max(0, int(value)))
          except (TypeError, ValueError):
              pass
          try:
              retry_at = parsedate_to_datetime(value)
              if retry_at.tzinfo is None:
                  retry_at = retry_at.replace(tzinfo=timezone.utc)
              return max(fallback, max(0, int((retry_at - datetime.now(timezone.utc)).total_seconds())))
          except (TypeError, ValueError, IndexError, OverflowError):
              return fallback
      
      
      _ERROR_DATA_URL_RE = re.compile(
          r"data:(?:audio|video|image)/[^;,\s\"'<>]+;base64,[A-Za-z0-9+/=]+",
          re.IGNORECASE,
      )
      _ERROR_KEY_RE = re.compile(r"\b(?:tp|sk)-[A-Za-z0-9_-]{8,}\b")
      
      
      def _sanitize_api_error(value, limit=500):
          """Bound transport diagnostics without echoing request media or credentials."""
          text = _ERROR_DATA_URL_RE.sub("<redacted-data-url>", str(value or ""))
          text = _ERROR_KEY_RE.sub("<redacted-key>", text)
          return text[:limit]
      
      def _api_headers(api_provider=None, api_url=None, api_key=None):
          """Build MiMo auth headers (OpenAI-compatible chat/completions with an api-key header)."""
          del api_provider, api_url  # MiMo is the only provider; signature kept for call sites
          key = CONFIG["api_key"] if api_key is None else api_key
          return {
              "Content-Type": "application/json",
              "User-Agent": "video-recap/1.0",
              "api-key": key,
          }
      
      def _prepare_api_payload(payload, api_provider=None, api_url=None):
          """Normalize payload fields for MiMo's OpenAI-compatible chat/completions API."""
          del api_provider, api_url
          normalized = dict(payload)
          if "max_tokens" in normalized and "max_completion_tokens" not in normalized:
              normalized["max_completion_tokens"] = normalized.pop("max_tokens")
          model = str(normalized.get("model") or "")
          if (
              CONFIG["mimo_disable_thinking"]
              and not model.endswith(("-tts", "-asr"))
              and "thinking" not in normalized
          ):
              # MiMo V2.5 may spend small max_completion_tokens budgets on reasoning_content.
              # The recap pipeline needs visible text, so disable thinking unless set explicitly.
              normalized["thinking"] = {"type": "disabled"}
          return normalized
      
      def _mimo_endpoint(kind):
          """Return per-capability MiMo endpoint settings (video understanding / TTS / ASR)."""
          by_kind = {
              "video": ("mimo_video_api_url", "mimo_video_api_key", "mimo_video_env_var"),
              "tts": ("mimo_tts_api_url", "mimo_tts_api_key", "mimo_tts_env_var"),
              "asr": ("mimo_asr_api_url", "mimo_asr_api_key", "mimo_asr_env_var"),
          }
          if kind not in by_kind:
              raise ValueError(f"Unsupported MiMo endpoint kind: {kind}")
          url_key, key_key, src_key = by_kind[kind]
          return {
              "api_url": CONFIG.get(url_key) or CONFIG.get("mimo_api_url"),
              "api_key": CONFIG.get(key_key) or CONFIG.get("mimo_api_key"),
              "api_env_var": CONFIG.get(src_key, "MIMO_API_KEY"),
          }
      
      def _call_mimo_endpoint(kind, payload, max_retries=10):
          settings = _mimo_endpoint(kind)
          return api_call(
              payload,
              max_retries=max_retries,
              api_provider="mimo",
              api_url=settings["api_url"],
              api_key=settings["api_key"],
              api_env_var=settings["api_env_var"],
          )
      
      def mimo_video_api_call(payload, max_retries=10):
          """Call the MiMo video-understanding endpoint."""
          return _call_mimo_endpoint("video", payload, max_retries=max_retries)
      
      def mimo_asr_api_call(payload, max_retries=10):
          """Call the MiMo speech-recognition (ASR) endpoint."""
          return _call_mimo_endpoint("asr", payload, max_retries=max_retries)
      
      def api_call(payload, max_retries=8, *, api_provider=None, api_url=None, api_key=None, api_env_var=None):
          """调用 OpenAI-compatible API,带重试。
      
          长视频理解会发出数百次 VLM/ASR 调用,集群的 429 限流是常态而非错误,所以重试更耐心
          (更多次数 + 退避封顶 60s + 遵从 Retry-After),避免一次瞬时限流就中止整段理解。
          集群的配额窗口常以分钟计,所以 429 在没有 Retry-After 时也至少等 10s,给窗口时间复位。
          """
          endpoint = normalize_api_url(api_url if api_url is not None else CONFIG["api_url"])
          headers = _api_headers(api_provider=api_provider, api_url=endpoint, api_key=api_key)
          data = json.dumps(_prepare_api_payload(payload, api_provider=api_provider, api_url=endpoint)).encode("utf-8")
      
          for attempt in range(max_retries):
              try:
                  req = urllib.request.Request(endpoint, data=data, headers=headers)
                  with urllib.request.urlopen(req, timeout=300) as resp:
                      return json.loads(resp.read().decode("utf-8"))
              except urllib.error.HTTPError as e:
                  body = _sanitize_api_error(e.read().decode("utf-8", errors="replace"))
                  wait = min(2 ** attempt, 60)
                  if e.code == 429:
                      retry_after = e.headers.get("Retry-After")
                      wait = _retry_after_seconds(retry_after, max(wait, 10))
                      log(f"API 速率限制 (尝试 {attempt+1}/{max_retries}), 等待 {wait}s")
                  elif e.code == 401:
                      key_name = api_env_var or CONFIG["api_env_var"]
                      raise RuntimeError(f"API 认证失败 (401)。请检查 {key_name} 和 API URL 是否匹配。")
                  elif e.code == 403:
                      hint = "API 访问被拒绝 (403)。"
                      if "1010" in body or "cloudflare" in body.lower():
                          hint += "IP 被 Cloudflare 限流,请等待几分钟后重试。"
                          raise RuntimeError(hint)
                      hint += "请检查 API key 权限和 API URL 设置。"
                      raise RuntimeError(hint)
                  elif e.code == 405:
                      raise RuntimeError("API 端点不可用 (405),可能被 WAF 拦截。请检查 MIMO_API_URL 或稍后重试。")
                  elif e.code == 503:
                      log(f"API 服务暂不可用 (503),等待 {wait}s (尝试 {attempt+1}/{max_retries})")
                  elif e.code == 524:
                      # Cloudflare 超时:服务端处理超时,需要更长退避
                      wait = max(wait, 4 * (attempt + 1))
                      log(f"API 超时 (524),等待 {wait}s (尝试 {attempt+1}/{max_retries})")
                  else:
                      log(f"API 调用失败 (尝试 {attempt+1}/{max_retries}): HTTP {e.code} — {body}")
                  if attempt < max_retries - 1:
                      time.sleep(wait)
                  else:
                      raise RuntimeError(f"API 调用失败 {max_retries} 次: HTTP {e.code} — {body}")
              except Exception as e:  # noqa: BLE001 - transport/decode faults are all retryable here
                  # (Deliberately broad, but no longer written as `(URLError, Exception)`, which
                  # read as a tuple while `Exception` already subsumed the first member.)
                  wait = min(2 ** attempt, 60)
                  safe_error = _sanitize_api_error(e)
                  log(f"API 调用失败 (尝试 {attempt+1}/{max_retries}): {safe_error}")
                  if attempt < max_retries - 1:
                      log(f"等待 {wait}s 后重试...")
                      time.sleep(wait)
                  else:
                      raise RuntimeError(f"API 调用失败 {max_retries} 次: {safe_error}")
      
      def load_prompt(name):
          """加载 prompt 模板"""
          path = PROMPTS_DIR / "prompt-templates.md"
          if not path.exists():
              return None
          content = path.read_text(encoding="utf-8")
          # 用 ### NAME 和 ### 分隔提取对应 prompt
          pattern = rf"### {name}\s*\n(.*?)(?=\n### |\Z)"
          m = re.search(pattern, content, re.DOTALL)
          return m.group(1).strip() if m else None
      
    • narration_lint.py 28.2 KB
      """Enforce narration timing, evidence, budget, and sentence-integrity rules.
      
      This is the single validator for the agent-authored narration.json: shape, numeric
      timing and text checks live here, and every later consumer trusts its result.
      """
      
      import json
      import math
      from pathlib import Path
      
      from lib import CONFIG, log
      from agent_text import (
          _clean_narration_punctuation,
          _find_scene_for_midpoint,
          _normalise_narration_segment,
          _post_dedup_narration,
          _recommended_char_budget,
          _scene_available_seconds,
          _text_char_count,
          _truncate_at_sentence,
      )
      from deslop_qc import analyze_deslop_qc
      from speech_ownership import (
          entry_overlaps_source_speech,
          load_source_sentence_evidence,
      )
      
      
      def _lint_issue(level, index, code, message, **extra):
          issue = {"level": level, "index": index, "code": code, "message": message}
          issue.update(extra)
          return issue
      
      
      def _is_number(value):
          return (
              isinstance(value, (int, float))
              and not isinstance(value, bool)
              and math.isfinite(value)
          )
      
      
      def _visual_overlay_issues(index, seg):
          """Shape-check the renderer handoff recap reads verbatim from a validated segment.
      
          recap owns which overlay `type` values it actually renders and drops the rest; lint owns
          the shape, because nothing between here and that filter re-checks it — and the filter runs
          after TTS, so a malformed overlay caught there would already have cost a synthesis pass.
          """
          overlays = seg.get("visual_overlays")
          if overlays is None:
              return []
          if not isinstance(overlays, list):
              return [
                  _lint_issue(
                      "error", index, "invalid_visual_overlays", "visual_overlays must be an array"
                  )
              ]
          issues = []
          for position, overlay in enumerate(overlays):
              if not isinstance(overlay, dict):
                  issues.append(
                      _lint_issue(
                          "error", index, "invalid_visual_overlay",
                          "visual_overlays entries must be objects",
                          overlay_index=position,
                      )
                  )
                  continue
              for key in ("type", "text"):
                  value = overlay.get(key)
                  if not isinstance(value, str) or not value.strip():
                      issues.append(
                          _lint_issue(
                              "error", index, "invalid_visual_overlay",
                              f"visual_overlays[].{key} must be a non-empty string",
                              overlay_index=position,
                          )
                      )
              for key in ("start", "end"):
                  if key in overlay and not _is_number(overlay[key]):
                      issues.append(
                          _lint_issue(
                              "error", index, "invalid_visual_overlay",
                              f"visual_overlays[].{key} must be numeric",
                              overlay_index=position,
                          )
                      )
          return issues
      
      
      def _scene_bounds_for_midpoint(scenes_analysis, start, end):
          scene = _find_scene_for_midpoint(scenes_analysis, start, end)
          if not scene:
              return None
          return scene["start"], scene["end"]
      
      
      def _frame_fact_times_for_segment(scenes_analysis, start, end):
          """Return frame-fact timestamps covered by a narration segment.
      
          These timestamps are cheap visual anchors. A narration slot that spans too
          many anchors is more likely to drift into "general story summary" instead
          of staying attached to what is on screen.
          """
          scene = _find_scene_for_midpoint(scenes_analysis, start, end)
          if not scene:
              return []
          return sorted(
              ts
              for ts in map(float, scene.get("frame_facts", {}))
              if start <= ts <= end
          )
      
      
      def _clip_span(clip):
          """A validated plan clip carries source_start/source_end; a raw agent plan start/end."""
          if "source_start" in clip:
              return clip["source_start"], clip["source_end"]
          return clip["start"], clip["end"]
      
      
      def _clip_matches_for_segment(seg, clip_plan, requested_clip_id):
          if not clip_plan:
              return []
          clips = clip_plan["clips"]
          midpoint = (seg["start"] + seg["end"]) / 2
          if requested_clip_id is not None:
              clips = [clip for clip in clips if clip.get("clip_id") == requested_clip_id]
          return [
              clip
              for clip in clips
              if _clip_span(clip)[0] <= midpoint <= _clip_span(clip)[1]
          ]
      
      
      def _source_sentence_entry_issue(index, start, anchors, speech_owned):
          """Return a blocking issue when narration enters midway through source speech.
      
          The validator reports a correction instead of silently moving audio. Source
          sentence integrity is invariant; an editorial policy string cannot bypass it.
          """
          if not speech_owned:
              return None
          # A cold-open at the first frame is not an interruption: the source sentence
          # has not been allowed to start. All later entries must use a measured anchor.
          if start <= 0.25:
              return None
          if not anchors:
              return _lint_issue(
                  "error",
                  index,
                  "source_sentence_anchors_unavailable",
                  "Source speech owns this entry but no verified sentence-end anchor is available. "
                  "Move/remove the narration or regenerate output speech evidence before TTS.",
                  entry_time=round(start, 3),
                  suggested_start=None,
              )
          # Only the measured acoustic pause owns a safe entry. A tiny 80ms post-anchor
          # tolerance covers timestamp/sample rounding; the old +450ms allowance could
          # already be several Chinese syllables into the next sentence.
          for anchor in anchors:
              when = anchor["time"]
              pause_start = anchor.get("pause_start", when - 0.12)
              if pause_start - 0.05 <= start <= when + 0.08:
                  return None
      
          suggested = next(
              (anchor for anchor in anchors if anchor["time"] > start + 0.08), None
          )
          return _lint_issue(
              "error",
              index,
              "interrupts_source_sentence",
              "Narration enters before the source sentence finishes. Move this block to the suggested "
              "sentence-end anchor, or shorten/move/remove it when there is no later verified anchor, "
              "then rerun lint before TTS. Source sentence interruption has no override.",
              entry_time=round(start, 3),
              suggested_start=round(suggested["time"], 3) if suggested else None,
              source_text_tail=str(suggested.get("text_tail", "")).strip() if suggested else "",
              anchor_confidence=suggested["confidence"] if suggested else None,
          )
      
      
      def _has_connected_predecessor(narration, idx, start):
          """A back-to-back narration handoff (<=150ms gap) does not expose the source track,
          so the next TTS block is not a new source-speech entry."""
          for other_idx, other in enumerate(narration):
              if other_idx == idx or not isinstance(other, dict):
                  continue
              other_start, other_end = other.get("start"), other.get("end")
              if not _is_number(other_start) or not _is_number(other_end):
                  continue
              if other_start < start and -0.001 <= start - other_end <= 0.15:
                  return True
          return False
      
      
      def lint_narration(
          narration, scenes_analysis=None, *, clip_plan=None, mode="full", work_dir=None,
          require_chronological=False,
      ):
          """Preflight-check agent narration before TTS; write narration_lint.json when work_dir is set.
      
          Segments are sorted by start for the timing checks; ``require_chronological`` (the
          --preserve-approved-text contract) additionally rejects input that is not already in order."""
          scenes_analysis = scenes_analysis or []
          errors = []
          warnings = []
          normalized = []
          source_evidence = load_source_sentence_evidence(work_dir, mode=mode)
          source_sentence_anchors = source_evidence["anchors"]
          if not isinstance(narration, list):
              errors.append(
                  _lint_issue(
                      "error",
                      None,
                      "invalid_json_shape",
                      "narration.json must be a JSON array",
                  )
              )
          else:
              previous_start = None
              for idx, seg in enumerate(narration):
                  if not isinstance(seg, dict):
                      errors.append(
                          _lint_issue(
                              "error",
                              idx,
                              "invalid_segment",
                              "Narration segment must be an object",
                          )
                      )
                      continue
                  start, end = seg.get("start"), seg.get("end")
                  if not _is_number(start) or not _is_number(end):
                      errors.append(
                          _lint_issue(
                              "error", idx, "invalid_time", "start/end must be finite numbers"
                          )
                      )
                      continue
                  if require_chronological and previous_start is not None and start < previous_start:
                      errors.append(
                          _lint_issue(
                              "error",
                              idx,
                              "out_of_order",
                              "Segments must be listed in chronological order",
                              start=start,
                              previous_start=previous_start,
                          )
                      )
                  previous_start = start
                  text = seg.get("narration")
                  if not isinstance(text, str):
                      errors.append(
                          _lint_issue(
                              "error",
                              idx,
                              "invalid_narration",
                              "narration must be a string",
                              start=start,
                              end=end,
                          )
                      )
                      continue
                  text = text.strip()
                  pause = seg.get("pause_after_ms", CONFIG["breath_ms"])
                  if not isinstance(pause, int) or isinstance(pause, bool) or pause < 0:
                      errors.append(
                          _lint_issue(
                              "error",
                              idx,
                              "invalid_pause",
                              "pause_after_ms must be a non-negative integer number of milliseconds",
                          )
                      )
                      continue
                  errors.extend(_visual_overlay_issues(idx, seg))
                  if end <= start:
                      errors.append(
                          _lint_issue(
                              "error",
                              idx,
                              "invalid_time_range",
                              "end must be greater than start",
                              start=start,
                              end=end,
                          )
                      )
                      continue
                  if not text:
                      errors.append(
                          _lint_issue(
                              "error",
                              idx,
                              "empty_narration",
                              "narration text must not be empty",
                              start=start,
                              end=end,
                          )
                      )
                      continue
      
                  char_count = _text_char_count(text)
                  budget = _recommended_char_budget(start, end)
                  # estimate at the REAL playback rate (after the narration_speed atempo); otherwise a
                  # beat sized to its 1.3x-sped slot looks "over budget" when it actually fits.
                  play_rate = CONFIG["speech_rate"] * CONFIG["narration_speed"]
                  estimated_tts_seconds = char_count / play_rate
                  slot_seconds = _scene_available_seconds(start, end)
                  if budget < 5:
                      warnings.append(
                          _lint_issue(
                              "warning",
                              idx,
                              "slot_too_short",
                              "Narration slot is very short; TTS may be clipped",
                              start=start,
                              end=end,
                              budget_chars=budget,
                          )
                      )
                  elif estimated_tts_seconds > slot_seconds:
                      warnings.append(
                          _lint_issue(
                              "warning",
                              idx,
                              "over_budget",
                              "Text may exceed the available TTS slot",
                              start=start,
                              end=end,
                              budget_chars=budget,
                              actual_chars=char_count,
                              estimated_tts_seconds=round(estimated_tts_seconds, 2),
                              slot_seconds=round(slot_seconds, 2),
                          )
                      )
                  if text[-1] not in "。!?!?….":
                      warnings.append(
                          _lint_issue(
                              "warning",
                              idx,
                              "incomplete_sentence",
                              "Narration should end with a complete sentence punctuation",
                              text_tail=text[-8:],
                          )
                      )
      
                  if not _has_connected_predecessor(narration, idx, start):
                      entry_issue = _source_sentence_entry_issue(
                          idx,
                          start,
                          source_sentence_anchors,
                          entry_overlaps_source_speech(
                              start,
                              source_evidence,
                              authored_overlap=seg.get("overlaps_speech", True),
                          ),
                      )
                      if entry_issue:
                          errors.append(entry_issue)
      
                  scene_bounds = _scene_bounds_for_midpoint(scenes_analysis, start, end)
                  if scenes_analysis and not scene_bounds:
                      warnings.append(
                          _lint_issue(
                              "warning",
                              idx,
                              "outside_scene",
                              "Narration midpoint does not match any detected scene",
                              start=start,
                              end=end,
                          )
                      )
                  elif scene_bounds and (start < scene_bounds[0] or end > scene_bounds[1]):
                      warnings.append(
                          _lint_issue(
                              "warning",
                              idx,
                              "crosses_scene_boundary",
                              "Narration extends outside its midpoint scene boundary",
                              start=start,
                              end=end,
                              scene_start=scene_bounds[0],
                              scene_end=scene_bounds[1],
                          )
                      )
      
                  frame_fact_times = _frame_fact_times_for_segment(
                      scenes_analysis, start, end
                  )
                  if (end - start) > CONFIG["visual_beat_max_seconds"] and len(
                      frame_fact_times
                  ) > CONFIG["visual_beat_max_facts"]:
                      warnings.append(
                          _lint_issue(
                              "warning",
                              idx,
                              "visual_beat_too_broad",
                              "Narration spans many visual anchors; split or tighten timing so the voiceover stays tied to current pictures",
                              start=start,
                              end=end,
                              duration=round(end - start, 2),
                              frame_fact_times=[round(ts, 2) for ts in frame_fact_times[:8]],
                          )
                      )
      
                  if mode == "cut":
                      requested_clip_id = seg.get("source_clip_id")
                      if requested_clip_id is not None:
                          try:
                              requested_clip_id = int(requested_clip_id)
                          except (TypeError, ValueError):
                              errors.append(
                                  _lint_issue(
                                      "error",
                                      idx,
                                      "invalid_source_clip_id",
                                      "source_clip_id must be an integer",
                                  )
                              )
                              continue
                      matches = _clip_matches_for_segment(seg, clip_plan, requested_clip_id)
                      if not matches:
                          errors.append(
                              _lint_issue(
                                  "error",
                                  idx,
                                  "outside_clip_plan",
                                  "Cut-mode narration must fall inside a selected clip",
                                  start=start,
                                  end=end,
                              )
                          )
                      elif len(matches) > 1 and requested_clip_id is None:
                          errors.append(
                              _lint_issue(
                                  "error",
                                  idx,
                                  "ambiguous_source_clip",
                                  "Repeated/overlapping clips require source_clip_id",
                                  start=start,
                                  end=end,
                              )
                          )
                      if len(matches) == 1:
                          clip_start, clip_end = _clip_span(matches[0])
                          if start < clip_start or end > clip_end:
                              warnings.append(
                                  _lint_issue(
                                      "warning",
                                      idx,
                                      "crosses_clip_boundary",
                                      "Narration extends beyond its clip; it will be trimmed to the clip and may describe footage that was cut",
                                      start=start,
                                      end=end,
                                      clip_start=round(clip_start, 3),
                                      clip_end=round(clip_end, 3),
                                  )
                              )
      
                  normalized.append(
                      {"index": idx, "start": start, "end": end, "char_count": char_count}
                  )
      
          sorted_segments = sorted(normalized, key=lambda item: item["start"])
          if isinstance(narration, list) and not sorted_segments:
              errors.append(
                  _lint_issue(
                      "error",
                      None,
                      "empty_narration_file",
                      "narration.json must contain at least one valid narration segment",
                  )
              )
          for prev, curr in zip(sorted_segments, sorted_segments[1:]):
              if curr["start"] < prev["end"]:
                  errors.append(
                      _lint_issue(
                          "error",
                          curr["index"],
                          "time_overlap",
                          "Segment overlaps the previous narration segment",
                          previous_index=prev["index"],
                          previous_end=prev["end"],
                          start=curr["start"],
                          end=curr["end"],
                      )
                  )
      
          # Block-coverage check (full mode only; cut-mode density is measured on the mapped output
          # timeline, not the source timestamps used here). Coverage is diagnostic, not a creative quota:
          # flag wall-to-wall or sparse drafts so the Agent consciously checks the visual/audio board,
          # plus missing original-audio gaps and fragmented one-sentence TTS blocks.
          metrics = {}
          if mode == "full" and len(sorted_segments) >= 2:
              timeline_start = 0.0
              timeline_end = sorted_segments[-1]["end"]
              scene_ends = [scene["end"] for scene in scenes_analysis]
              if scene_ends:
                  timeline_end = max(timeline_end, max(scene_ends))
              span = max(0.0, timeline_end - timeline_start)
              # Score coverage at the SAME conservative rate the writer is budgeted at
              # (_recommended_char_budget uses speech_rate * speech_safety_margin * narration_speed),
              # else the metric scores ~18% stricter than its own budget and false-flags under_narrated.
              play_rate = (
                  CONFIG["speech_rate"] * CONFIG["speech_safety_margin"] * CONFIG["narration_speed"]
              )
              spoken = [s["char_count"] / play_rate for s in sorted_segments]
              spoken_ends = [
                  min(sorted_segments[i]["end"], sorted_segments[i]["start"] + spoken[i])
                  for i in range(len(sorted_segments))
              ]
              # Measure the UNION of estimated playback intervals. Adjacent authored
              # blocks can overlap the conservative text-duration estimate; summing them
              # double-counted time and falsely called a recap wall-to-wall.
              estimated_intervals = sorted(
                  (sorted_segments[i]["start"], spoken_ends[i])
                  for i in range(len(sorted_segments))
                  if spoken_ends[i] > sorted_segments[i]["start"]
              )
              merged_intervals = []
              for start, end in estimated_intervals:
                  if merged_intervals and start <= merged_intervals[-1][1]:
                      merged_intervals[-1][1] = max(merged_intervals[-1][1], end)
                  else:
                      merged_intervals.append([start, end])
              narrated_seconds = sum(end - start for start, end in merged_intervals)
              coverage = narrated_seconds / span if span > 0 else 0.0
              orig_min = CONFIG["original_block_min_seconds"]
              orig_gaps = [sorted_segments[0]["start"] - timeline_start]
              orig_gaps.extend(
                  sorted_segments[i + 1]["start"] - spoken_ends[i]
                  for i in range(len(sorted_segments) - 1)
              )
              orig_gaps.append(timeline_end - spoken_ends[-1])
              original_blocks = sum(1 for g in orig_gaps if g >= orig_min)
              avg_chars = sum(s["char_count"] for s in sorted_segments) / len(sorted_segments)
              cov_target = CONFIG["narration_coverage_target"]
              cov_max = CONFIG["narration_coverage_max"]
              cov_min = CONFIG["narration_coverage_min"]
              block_min_chars = CONFIG["narration_block_min_chars"]
              metrics = {
                  "segment_count": len(sorted_segments),
                  "timeline_span_seconds": round(span, 2),
                  "narrated_seconds": round(narrated_seconds, 2),
                  "narration_coverage": round(coverage, 2),
                  "coverage_target": cov_target,
                  "original_block_count": original_blocks,
                  "avg_block_chars": round(avg_chars, 1),
              }
              if coverage > cov_max:
                  warnings.append(
                      _lint_issue(
                          "warning",
                          None,
                          "no_original_blocks",
                          "Narration is nearly wall-to-wall — the original audio never gets to breathe. Pull back at a "
                          "few strong moments and write NO narration there so the original plays at full volume. "
                          "Choose those moments by audio_owner, not by a fixed ratio.",
                          narration_coverage=round(coverage, 2),
                          coverage_max=cov_max,
                      )
                  )
              elif coverage < cov_min:
                  warnings.append(
                      _lint_issue(
                          "warning",
                          None,
                          "under_narrated",
                          "Narration coverage is sparse. This is not automatically wrong: verify that every long gap is "
                          "intentionally owned by original dialogue/action/ambience/music/silence in visual_audio_board.json. "
                          "Only add a block when it has a specific narration job.",
                          narration_coverage=round(coverage, 2),
                          coverage_min=cov_min,
                      )
                  )
              if original_blocks == 0 and span >= 3 * orig_min:
                  warnings.append(
                      _lint_issue(
                          "warning",
                          None,
                          "no_original_breaks",
                          "No deliberate original-audio blocks — narration runs end-to-end with no gap for a strong "
                          "original moment (a key line, an action beat, the music). Leave a few multi-second gaps "
                          "between blocks where the original plays alone.",
                          original_block_min_seconds=orig_min,
                      )
                  )
              if len(sorted_segments) >= 8 and avg_chars < block_min_chars:
                  warnings.append(
                      _lint_issue(
                          "warning",
                          None,
                          "fragmented_beats",
                          "Beats are fragmented into single short sentences; each is synthesized as a separate TTS "
                          "utterance, which sounds choppy. Merge adjacent related sentences into one fluent BLOCK "
                          "that completes a continuous thought; sentence count is not the target.",
                          avg_block_chars=round(avg_chars, 1),
                          block_min_chars=block_min_chars,
                      )
                  )
      
          deslop_qc = analyze_deslop_qc(
              [seg for seg in narration if isinstance(seg, dict)]
              if isinstance(narration, list)
              else [],
              work_dir=work_dir,
          )
          for blocker in deslop_qc["blockers"]:
              errors.append(
                  _lint_issue(
                      "error",
                      blocker["index"],
                      blocker["code"],
                      blocker["message"],
                      source=blocker["source"],
                      matches=blocker.get("matches"),
                      sentence=blocker.get("sentence"),
                  )
              )
      
          report = {
              "ok": not errors,
              "error_count": len(errors),
              "warning_count": len(warnings),
              "metrics": metrics,
              "deslop_qc": deslop_qc,
              "errors": errors,
              "warnings": warnings,
          }
          if work_dir is not None:
              Path(work_dir, "deslop_qc.json").write_text(
                  json.dumps(deslop_qc, ensure_ascii=False, indent=2), encoding="utf-8"
              )
              Path(work_dir, "narration_lint.json").write_text(
                  json.dumps(report, ensure_ascii=False, indent=2), encoding="utf-8"
              )
          return report
      
      
      def validate_narration_or_raise(
          narration, scenes_analysis=None, *, clip_plan=None, mode="full", work_dir=None,
          require_chronological=False,
      ):
          report = lint_narration(
              narration, scenes_analysis, clip_plan=clip_plan, mode=mode, work_dir=work_dir,
              require_chronological=require_chronological,
          )
          if report["errors"]:
              sample = "; ".join(
                  f"#{e['index']}: {e['code']}" for e in report["errors"][:3]
              )
              raise ValueError(f"narration.json 预检失败: {sample}; 详见 narration_lint.json")
          if report["warnings"]:
              log(
                  f"narration lint: {len(report['warnings'])} warnings (see narration_lint.json)"
              )
          else:
              log("narration lint: ok")
          return report
      
      
      def _validate_narration_budget(narration, scenes_analysis):
          """Trim lint-validated narration to its timing budgets; drop what cannot be spoken."""
          del scenes_analysis  # scene boundaries are advisory (lint warns); authored timing is kept
          cleaned = []
          for raw in narration:
              item = _normalise_narration_segment(raw)
              max_chars = _recommended_char_budget(item["start"], item["end"])
              if max_chars < 5:
                  log(f"  丢弃过短解说段 {item['start']:.1f}-{item['end']:.1f}s")
                  continue
              if _text_char_count(item["narration"]) > max_chars * 1.25:
                  truncated = _truncate_at_sentence(item["narration"], max_chars)
                  if truncated and _text_char_count(truncated) >= 5:
                      log(f"  解说超预算,已截短: {item['start']:.1f}-{item['end']:.1f}s")
                      item["narration"] = truncated
                  else:
                      log(
                          f"  解说超预算且无法安全截断,已丢弃: {item['start']:.1f}-{item['end']:.1f}s"
                      )
                      continue
              item["narration"] = _clean_narration_punctuation(item["narration"])
              stripped = item["narration"].strip()
              if stripped and stripped[-1] in ",:、;,—":
                  item["narration"] = stripped.rstrip(",:、;,—") + "。"
              cleaned.append(item)
      
          cleaned.sort(key=lambda n: n["start"])
          deduped = []
          for item in cleaned:
              if deduped and item["start"] < deduped[-1]["end"]:
                  prev = deduped[-1]
                  log(
                      f"  解说时间重叠: {item['start']:.1f}-{item['end']:.1f}s vs "
                      f"{prev['start']:.1f}-{prev['end']:.1f}s"
                  )
                  if _text_char_count(item["narration"]) > _text_char_count(
                      prev["narration"]
                  ):
                      deduped[-1] = item
              else:
                  deduped.append(item)
          return _post_dedup_narration(deduped)
      
    • speech_ownership.py 4.9 KB
      """Load measured source-speech evidence and classify narration ownership."""
      
      import json
      from pathlib import Path
      
      from lib import CONFIG, file_identity
      
      
      def _empty_evidence(mode):
          return {
              "anchors": [],
              "speech_spans": [],
              "quiet_windows": [],
              "require_measured": mode == "cut_output",
          }
      
      
      def _read_json(path):
          """JSON artifact, or None when the optional file was never written."""
          return json.loads(path.read_text(encoding="utf-8")) if path.exists() else None
      
      
      def _output_payload_is_current(payload, work_dir):
          plan_path = Path(work_dir) / "clip_plan_validated.json"
          return (
              payload is not None
              and plan_path.exists()
              and payload["clip_plan_identity"] == file_identity(plan_path)
          )
      
      
      def load_source_sentence_evidence(work_dir, mode="full"):
          """Load sentence boundaries plus speech/quiet spans on the narration clock."""
          if work_dir is None:
              return _empty_evidence(mode)
          work_dir = Path(work_dir)
          if mode == "full":
              payload = _read_json(work_dir / "speech_boundary_anchors.json") or {
                  "sentence_anchors": []
              }
              speech_spans = [
                  row for row in _read_json(work_dir / "asr_result.json") or [] if row["text"]
              ]
              quiet_windows = [
                  row
                  for row in _read_json(work_dir / "silence_periods.json") or []
                  if not row["has_speech"]
              ]
          else:
              payload = _read_json(work_dir / "speech_boundary_anchors_output.json")
              if mode == "cut_output" and not _output_payload_is_current(payload, work_dir):
                  return _empty_evidence(mode)
              if payload is None:
                  payload = {"sentence_anchors": [], "speech_spans": [], "quiet_windows": []}
              speech_spans = payload["speech_spans"]
              quiet_windows = payload["quiet_windows"]
          anchors = [
              anchor
              for anchor in payload["sentence_anchors"]
              if anchor["confidence"] in {"high", "medium"}
          ]
          return {
              "anchors": sorted(anchors, key=lambda item: item["time"]),
              "speech_spans": speech_spans,
              "quiet_windows": quiet_windows,
              "require_measured": mode == "cut_output",
          }
      
      
      def _merged_intervals(start, end, rows):
          intervals = sorted(
              (max(start, row["start"]), min(end, row["end"]))
              for row in rows
              if row["end"] > start and row["start"] < end
          )
          merged = []
          for left, right in intervals:
              if right <= left:
                  continue
              if merged and left <= merged[-1][1]:
                  merged[-1] = (merged[-1][0], max(merged[-1][1], right))
              else:
                  merged.append((left, right))
          return merged
      
      
      def _interval_overlap(start, end, rows):
          return sum(right - left for left, right in _merged_intervals(start, end, rows))
      
      
      def _speech_overlap_excluding_quiet(start, end, speech, quiet):
          speech_intervals = _merged_intervals(start, end, speech)
          quiet_intervals = _merged_intervals(start, end, quiet)
          overlap = sum(right - left for left, right in speech_intervals)
          for speech_left, speech_right in speech_intervals:
              overlap -= sum(
                  max(0.0, min(speech_right, quiet_right) - max(speech_left, quiet_left))
                  for quiet_left, quiet_right in quiet_intervals
              )
          return max(0.0, overlap)
      
      
      def segment_overlaps_source_speech(seg, evidence):
          """Classify aggregate mix ownership across the complete narration interval."""
          start, end = seg["start"], seg["end"]
          quiet = evidence["quiet_windows"]
          quiet_min = max(0.3, (end - start) * CONFIG["quiet_overlap_min_ratio"])
          speech = evidence["speech_spans"]
          if speech:
              return _speech_overlap_excluding_quiet(start, end, speech, quiet) > 0.05
          if quiet and _interval_overlap(start, end, quiet) >= quiet_min:
              return False
          if evidence["anchors"] or evidence["require_measured"]:
              return True
          return bool(seg.get("overlaps_speech", True))
      
      
      def entry_overlaps_source_speech(start, evidence, *, authored_overlap=True, tolerance=0.05):
          """Classify the entry instant; later quiet time cannot erase an unsafe start."""
          if any(
              row["start"] - tolerance <= start <= row["end"] + tolerance
              for row in evidence["quiet_windows"]
          ):
              return False
          if any(
              row["start"] - tolerance <= start < row["end"] - tolerance
              for row in evidence["speech_spans"]
          ):
              return True
          if evidence["speech_spans"]:
              return False
          if evidence["anchors"] or evidence["require_measured"]:
              return True
          return bool(authored_overlap)
      
      
      def measure_narration_speech_ownership(narration, work_dir, mode="full"):
          """Return narration copies with aggregate ownership derived from evidence."""
          evidence = load_source_sentence_evidence(work_dir, mode=mode)
          return [
              {**seg, "overlaps_speech": segment_overlaps_source_speech(seg, evidence)}
              for seg in narration
          ]
      
    • storyboard.py 18.5 KB
      #!/usr/bin/env python3
      """Storyboard (contact-sheet) generation for video-understanding.
      
      Two ADVISORY artifacts that help the writing agent orient on the timeline by SCANNING
      one image instead of opening dozens of frames:
      
        - source_storyboard.{jpg,json}  — scene-anchored tiles over the SOURCE timeline.
        - edited_storyboard.{jpg,json}  — one row per kept clip over the cut OUTPUT timeline,
                                          each tile dual-labelled (output time / source time).
      
      Both reuse the frames already extracted by understand.py (frames/frame_*.jpg at CONFIG["fps"]).
      Nothing here re-extracts video. Every function returns dict|None and degrades to
      None + log(...) on ANY failure (no frames, ffmpeg missing/non-zero, font probe raises),
      so a storyboard quirk can NEVER block the pipeline (Principle 1: advisory, never blocking).
      
      drawtext "mm:ss" labels are attempted when a usable font is found (labels_burned:true);
      when no font is available (or drawtext errors) the sheet is still produced UNLABELLED
      (labels_burned:false) and the JSON sidecar stays authoritative for all timestamps.
      """
      import json
      import math
      import shutil
      import subprocess
      from pathlib import Path
      
      from extract import frame_number_for_time, parse_frame_number
      from lib import CONFIG, run_cmd, log, get_video_duration
      
      
      # Candidate font files probed (in order) for burning mm:ss labels. The first that exists
      # AND that drawtext can actually load wins. A probe that RAISES must never abort the sheet.
      _FONT_CANDIDATES = (
          "/System/Library/Fonts/Supplemental/Arial.ttf",
          "/System/Library/Fonts/Helvetica.ttc",
          "/Library/Fonts/Arial.ttf",
          "/usr/share/fonts/truetype/dejavu/DejaVuSans.ttf",
          "/usr/share/fonts/dejavu/DejaVuSans.ttf",
          "/usr/share/fonts/truetype/liberation/LiberationSans-Regular.ttf",
          "/usr/share/fonts/TTF/DejaVuSans.ttf",
      )
      
      
      def _fmt_mmss(seconds):
          """Format a timestamp as mm:ss (clamped to >= 0)."""
          total = int(round(max(0.0, float(seconds))))
          return f"{total // 60:02d}:{total % 60:02d}"
      
      
      def _ffmpeg_available():
          return shutil.which("ffmpeg") is not None
      
      
      def _probe_font():
          """Return a usable font file path, or None. Any exception → None (never raises out).
      
          Probe order: explicit candidate files first, then `fc-match` if available. The path
          is only returned when the file exists on disk; we do NOT shell out to drawtext here —
          a render-time drawtext error is caught separately and also degrades to unlabelled.
          """
          try:
              for candidate in _FONT_CANDIDATES:
                  if Path(candidate).is_file():
                      return candidate
              fc_match = shutil.which("fc-match")
              if fc_match:
                  result = subprocess.run(
                      [fc_match, "-f", "%{file}", "sans"],
                      capture_output=True, text=True, timeout=10,
                  )
                  path = (result.stdout or "").strip()
                  if result.returncode == 0 and path and Path(path).is_file():
                      return path
          except Exception as exc:  # noqa: BLE001 - a font probe must NEVER abort the sheet
              log(f"storyboard 字体探测异常(降级为不烧时间戳): {exc}")
              return None
          return None
      
      
      def _frame_index(work_dir):
          """Return (sorted_frame_paths, sorted_numbers) for frames/frame_*.jpg, or ([], [])."""
          frames_dir = Path(work_dir) / "frames"
          if not frames_dir.is_dir():
              return [], []
          pairs = []
          for path in frames_dir.glob("frame_*.jpg"):
              number = parse_frame_number(path)
              if number is None:
                  continue
              pairs.append((number, path))
          pairs.sort(key=lambda item: item[0])
          numbers = [num for num, _ in pairs]
          paths = [path for _, path in pairs]
          return paths, numbers
      
      
      def _nearest_existing_frame(timestamp, fps, paths, numbers):
          """Map a SOURCE timestamp → the nearest EXISTING frame file, clamped to [first,last].
      
          Frames are named frame_{n:05d}.jpg; the number↔time mapping is owned by extract.py
          (frame_00001 is t=0, so n = t*fps + 1). Rounding a timestamp blindly can yield a frame
          number that was never written (fps boundary / last frame gap), so we resolve to the
          closest number that actually exists on disk.
          """
          if not numbers:
              return None
          fps = float(fps)
          if fps <= 0:
              return None
          target = frame_number_for_time(timestamp, fps)
          # Clamp into the real extracted range so out-of-range timestamps pin to first/last frame.
          if target <= numbers[0]:
              return paths[0]
          if target >= numbers[-1]:
              return paths[-1]
          # numbers is sorted; find the closest by absolute distance (ties → earlier frame).
          best_idx = 0
          best_dist = None
          for idx, num in enumerate(numbers):
              dist = abs(num - target)
              if best_dist is None or dist < best_dist:
                  best_dist = dist
                  best_idx = idx
              elif num - target > best_dist:
                  break  # sorted: distance only grows from here
          return paths[best_idx]
      
      
      def _scene_anchor_timestamps(scenes, max_tiles):
          """Scene-anchored sample timestamps: each scene midpoint; long scenes also +1/3 & +2/3.
      
          Returns a list of (scene_id, timestamp). A scene is "long" when adding the thirds gives
          materially distinct sample points; we treat scenes longer than `long_scene_seconds` as
          long. Total is capped at max_tiles by evenly subsampling the ordered anchor list, so the
          sheet stays legible (D2: one bounded contact sheet, not dozens of frame reads).
          """
          long_scene_seconds = float(CONFIG["storyboard_long_scene_seconds"])
          anchors = []
          for scene_id, scene in enumerate(scenes or []):
              start = float(scene["start"])
              end = float(scene["end"])
              if end <= start:
                  continue
              mid = (start + end) / 2.0
              if (end - start) >= long_scene_seconds:
                  third = start + (end - start) / 3.0
                  two_third = start + 2.0 * (end - start) / 3.0
                  points = [third, mid, two_third]
              else:
                  points = [mid]
              for ts in points:
                  anchors.append((scene_id, round(ts, 3)))
          if max_tiles and len(anchors) > max_tiles:
              # Evenly subsample to the cap, preserving timeline order and scene spread.
              step = len(anchors) / float(max_tiles)
              anchors = [anchors[int(i * step)] for i in range(max_tiles)]
          return anchors
      
      
      def _labelled_frame(frame_path, label, font_path, scratch_dir, out_name):
          """Burn `label` onto a copy of frame_path via drawtext; return the labelled path or None.
      
          None signals the caller to fall back to the original frame (and flip labels_burned off).
          A drawtext failure here is non-fatal: the unlabelled frame still tiles fine.
          """
          out_path = scratch_dir / out_name
          safe_label = label.replace("\\", "\\\\").replace(":", "\\:").replace("'", "’")
          drawtext = (
              f"drawtext=fontfile='{font_path}':text='{safe_label}':"
              "x=8:y=8:fontsize=28:fontcolor=white:"
              "box=1:boxcolor=black@0.55:boxborderw=6"
          )
          cmd = ["ffmpeg", "-y", "-i", str(frame_path), "-vf", drawtext, "-frames:v", "1", str(out_path)]
          try:
              result = run_cmd(cmd)
          except Exception as exc:  # noqa: BLE001 - a label render must never abort the sheet
              log(f"storyboard drawtext 异常(降级为不烧时间戳): {exc}")
              return None
          if result.returncode != 0 or not out_path.exists():
              return None
          return out_path
      
      
      def _tile_pages(frame_paths, columns, out_dir, out_stem, scratch_dir):
          """Tile frame_paths into one or more contact-sheet pages; return the list of page paths.
      
          ffmpeg's tile filter lays one grid per page. We page so each sheet holds at most
          columns*rows tiles where rows is chosen to keep the grid roughly square but capped, then
          spill into _001.jpg, _002.jpg… Returns [] on any ffmpeg failure (caller degrades to None).
          """
          columns = max(1, int(columns))
          rows_per_page = int(CONFIG["storyboard_rows_per_page"])
          per_page = columns * rows_per_page
          pages = []
          total = len(frame_paths)
          page_count = max(1, math.ceil(total / per_page))
          for page_idx in range(page_count):
              chunk = frame_paths[page_idx * per_page:(page_idx + 1) * per_page]
              if not chunk:
                  continue
              cols = min(columns, len(chunk))
              rows = max(1, math.ceil(len(chunk) / cols))
              if page_count == 1:
                  page_path = out_dir / f"{out_stem}.jpg"
              else:
                  page_path = out_dir / f"{out_stem}_{page_idx + 1:03d}.jpg"
              # Stage the chunk as a contiguous numbered sequence so ffmpeg's image2 demuxer +
              # tile filter consume EXACTLY these frames (the source frame numbers are sparse).
              seq_dir = scratch_dir / f"seq_{out_stem}_{page_idx:03d}"
              seq_dir.mkdir(parents=True, exist_ok=True)
              for seq_idx, src in enumerate(chunk):
                  shutil.copyfile(src, seq_dir / f"f_{seq_idx:05d}.jpg")
              cmd = [
                  "ffmpeg", "-y",
                  "-framerate", "1",
                  "-i", str(seq_dir / "f_%05d.jpg"),
                  "-frames:v", "1",
                  "-vf", f"tile={cols}x{rows}",
                  str(page_path),
              ]
              result = run_cmd(cmd)
              if result.returncode != 0 or not page_path.exists():
                  log(f"storyboard tile 失败: {result.stderr[-300:]}")
                  return []
              pages.append(page_path)
          return pages
      
      
      def _render_storyboard(work_dir, tiles, out_stem):
          """Shared render path: optionally burn labels, tile to pages, return (page_paths, labels_burned).
      
          `tiles` is a list of dicts that ALREADY carry a resolved `frame_file` (absolute path) and a
          `label` string. Returns (None, _) on hard failure so callers degrade to None.
          """
          if not _ffmpeg_available():
              log("storyboard 跳过:未找到 ffmpeg")
              return None, False
          storyboard_dir = Path(work_dir) / "storyboard"
          storyboard_dir.mkdir(parents=True, exist_ok=True)
          scratch_dir = storyboard_dir / f".scratch_{out_stem}"
          if scratch_dir.exists():
              shutil.rmtree(scratch_dir, ignore_errors=True)
          scratch_dir.mkdir(parents=True, exist_ok=True)
      
          font_path = _probe_font()
          labels_burned = bool(font_path)
          render_frames = []
          if labels_burned:
              for idx, tile in enumerate(tiles):
                  labelled = _labelled_frame(
                      Path(tile["frame_file"]), tile["label"], font_path, scratch_dir,
                      f"lbl_{out_stem}_{idx:05d}.jpg",
                  )
                  if labelled is None:
                      # First failure → abandon labelling entirely so the WHOLE sheet is consistent
                      # (no half-labelled pages). The JSON sidecar still carries every timestamp.
                      labels_burned = False
                      break
                  render_frames.append(labelled)
          if not labels_burned:
              render_frames = [Path(tile["frame_file"]) for tile in tiles]
      
          try:
              pages = _tile_pages(
                  render_frames, CONFIG["storyboard_columns"],
                  storyboard_dir, out_stem, scratch_dir,
              )
          finally:
              shutil.rmtree(scratch_dir, ignore_errors=True)
          if not pages:
              return None, labels_burned
          return pages, labels_burned
      
      
      def build_source_storyboard(work_dir, video_path, scenes, fps):
          """Scene-anchored SOURCE-timeline contact sheet. Returns dict|None.
      
          Samples each scene midpoint (long scenes also +1/3,+2/3), maps every sample to the
          nearest EXISTING extracted frame (clamped), tiles them, and writes
          storyboard/source_storyboard.json. Returns None + log on any failure.
          """
          try:
              work_dir = Path(work_dir)
              paths, numbers = _frame_index(work_dir)
              if not paths:
                  log("storyboard 跳过 source:frames/ 为空或缺失")
                  return None
              max_tiles = int(CONFIG["storyboard_max_tiles"])
              columns = int(CONFIG["storyboard_columns"])
              anchors = _scene_anchor_timestamps(scenes, max_tiles)
              if not anchors:
                  log("storyboard 跳过 source:无可用场景锚点")
                  return None
      
              tiles = []
              for tile_id, (scene_id, ts) in enumerate(anchors):
                  frame = _nearest_existing_frame(ts, fps, paths, numbers)
                  if frame is None:
                      continue
                  tiles.append({
                      "tile_id": tile_id,
                      "timestamp": round(float(ts), 3),
                      "label": _fmt_mmss(ts),
                      "scene_id": scene_id,
                      "frame_file": str(frame),
                  })
              if not tiles:
                  log("storyboard 跳过 source:未解析到任何帧")
                  return None
      
              pages, labels_burned = _render_storyboard(work_dir, tiles, "source_storyboard")
              if not pages:
                  log("storyboard 跳过 source:拼贴失败")
                  return None
      
              for tile in tiles:
                  tile["frame_file"] = Path(tile["frame_file"]).name
              payload = {
                  "schema_version": 1,
                  "timeline": "source",
                  "video_path": str(video_path),
                  "fps": float(fps) if fps else None,
                  "labels_burned": labels_burned,
                  "page_images": [str(p) for p in pages],
                  "sample_policy": {
                      "max_tiles": max_tiles,
                      "columns": columns,
                      "anchors": "scene_midpoint+long_scene_thirds",
                  },
                  "tiles": tiles,
              }
              try:
                  payload["duration"] = round(get_video_duration(video_path), 3)
              except RuntimeError as exc:
                  log(f"storyboard source:无法读取视频时长,sidecar 省略 duration(忽略): {exc}")
              json_path = work_dir / "storyboard" / "source_storyboard.json"
              json_path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
              log(f"storyboard source: {len(tiles)} tiles → {len(pages)} page(s), labels_burned={labels_burned}")
              return payload
          except Exception as exc:  # noqa: BLE001 - advisory: never propagate
              log(f"storyboard source 失败(忽略): {exc}")
              return None
      
      
      def _source_to_output(source_time, clip):
          """Forward affine source→output map for ONE clip (reimplements cut.py:319 locally).
      
          output = clip.output_start + (src − clip.source_start), clamped to [output_start, output_end].
          Read the authoritative numbers from clip_plan_validated.json; do NOT import cut.py.
          """
          src = float(source_time)
          out_start = float(clip["output_start"])
          out_end = float(clip["output_end"])
          mapped = out_start + (src - float(clip["source_start"]))
          return round(max(out_start, min(mapped, out_end)), 3)
      
      
      def build_edited_storyboard(work_dir, source_video_path, clip_plan_validated, fps):
          """OUTPUT-timeline contact sheet: one row per kept clip over the cut. Returns dict|None.
      
          For each kept clip, samples source start / mid / (end − 0.5s), maps each to the nearest
          EXISTING SOURCE frame (SOURCE fps — frames are reused, NO re-extraction), and FRAME-IDENTITY
          dedupes so a ≤1s clip yields 1-2 tiles (not 3 identical). Each tile is dual-labelled with
          both `output_timestamp` and `source_timestamp` (+ `source_clip_id`). Writes
          storyboard/edited_storyboard.json. Returns None + log on any failure.
          """
          try:
              work_dir = Path(work_dir)
              paths, numbers = _frame_index(work_dir)
              if not paths:
                  log("storyboard 跳过 edited:frames/ 为空或缺失")
                  return None
              clips = clip_plan_validated["clips"]
              if not clips:
                  log("storyboard 跳过 edited:clip_plan_validated 无 clips")
                  return None
              columns = int(CONFIG["storyboard_columns"])
              max_tiles = int(CONFIG["storyboard_max_tiles"])
      
              tiles = []
              seen_frames = set()  # frame-identity dedupe (NOT luma de-dupe; that stays deferred)
              tile_id = 0
              for clip in clips:
                  try:
                      source_start = float(clip["source_start"])
                      source_end = float(clip["source_end"])
                      clip_id = clip.get("clip_id")
                  except (KeyError, TypeError, ValueError):
                      continue
                  if source_end <= source_start:
                      continue
                  mid = (source_start + source_end) / 2.0
                  end_sample = max(source_start, source_end - 0.5)
                  for src_ts in (source_start, mid, end_sample):
                      frame = _nearest_existing_frame(src_ts, fps, paths, numbers)
                      if frame is None:
                          continue
                      key = (clip_id, str(frame))
                      if key in seen_frames:
                          continue  # same clip resolving to the same frame → drop the duplicate tile
                      seen_frames.add(key)
                      out_ts = _source_to_output(src_ts, clip)
                      tiles.append({
                          "tile_id": tile_id,
                          "output_timestamp": out_ts,
                          "source_timestamp": round(float(src_ts), 3),
                          "source_clip_id": clip_id,
                          "label": f"out {_fmt_mmss(out_ts)} / src {_fmt_mmss(src_ts)}",
                          "frame_file": str(frame),
                      })
                      tile_id += 1
              if not tiles:
                  log("storyboard 跳过 edited:未解析到任何帧")
                  return None
              if len(tiles) > max_tiles:
                  tiles = tiles[:max_tiles]
      
              pages, labels_burned = _render_storyboard(work_dir, tiles, "edited_storyboard")
              if not pages:
                  log("storyboard 跳过 edited:拼贴失败")
                  return None
      
              for tile in tiles:
                  tile["frame_file"] = Path(tile["frame_file"]).name
              edited_source = Path(work_dir) / "edited_source.mp4"
              payload = {
                  "schema_version": 1,
                  "timeline": "output",
                  "source_video_path": str(source_video_path),
                  "edited_video_path": str(edited_source) if edited_source.exists() else None,
                  "labels_burned": labels_burned,
                  "page_images": [str(p) for p in pages],
                  "sample_policy": {
                      "max_tiles": max_tiles,
                      "columns": columns,
                      "per_clip": "source_start+mid+(end-0.5s), frame-identity deduped",
                  },
                  "tiles": tiles,
              }
              json_path = work_dir / "storyboard" / "edited_storyboard.json"
              json_path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
              log(f"storyboard edited: {len(tiles)} tiles → {len(pages)} page(s), labels_burned={labels_burned}")
              return payload
          except Exception as exc:  # noqa: BLE001 - advisory: never propagate
              log(f"storyboard edited 失败(忽略): {exc}")
              return None
      
    • timeline_fusion.py 4.2 KB
      """Fuse scenes, ASR, and quiet windows for narration planning."""
      
      from lib import CONFIG
      from agent_text import _overlap_seconds, _recommended_char_budget
      from narration_lint import _validate_narration_budget
      
      
      def _quiet_windows(silence_periods):
          return [qp for qp in silence_periods if not qp["has_speech"]]
      
      
      def _align_narration_to_quiet(narration, scenes_analysis, silence_periods):
          """Recompute overlaps_speech from real quiet windows; keep the agent's timing.
      
          The dense continuous-bed design places narration ON the pictured beat over a
          ducked original bed, so segments are never relocated into silence gaps. Only
          the overlaps_speech flag that the ducking stage consumes is corrected, leaving
          the agent's start/end (and text) intact.
      
          Budget/dedup runs FIRST so a dedup-merged beat's overlaps_speech reflects its
          extended timing, not its original shorter span.
          """
          aligned = _validate_narration_budget(narration, scenes_analysis)
          quiet_windows = _quiet_windows(silence_periods)
          quiet_ratio_min = CONFIG["quiet_overlap_min_ratio"]
          for n in aligned:
              seg_dur = n["end"] - n["start"]
              quiet_overlap = sum(
                  _overlap_seconds(n["start"], n["end"], qw["start"], qw["end"])
                  for qw in quiet_windows
              )
              n["overlaps_speech"] = quiet_overlap < max(0.3, seg_dur * quiet_ratio_min)
          return aligned
      
      
      def _scene_asr_lines(asr_result, scene):
          lines = []
          for seg in asr_result:
              if scene["start"] < seg["end"] and scene["end"] > seg["start"] and seg["text"]:
                  lines.append(f"    [{seg['start']:.1f}-{seg['end']:.1f}] {seg['text']}")
          return lines
      
      
      def _quiet_windows_for_scene(silence_periods, scene):
          windows = []
          for qp in _quiet_windows(silence_periods):
              if qp["start"] < scene["end"] and qp["end"] > scene["start"]:
                  start = max(qp["start"], scene["start"])
                  end = min(qp["end"], scene["end"])
                  if end > start:
                      windows.append((start, end))
          return windows
      
      
      def _build_timeline_fusion(scenes, asr_segments, silence_periods):
          """Fuse VLM scenes, ASR dialogue and quiet narration slots on one timeline."""
          fusion = []
          quiet_windows = _quiet_windows(silence_periods)
          for scene in scenes:
              start, end = scene["start"], scene["end"]
              dialogue_segments = []
              dialogue_overlap = 0.0
              for seg in asr_segments:
                  overlap = _overlap_seconds(start, end, seg["start"], seg["end"])
                  if overlap <= 0:
                      continue
                  dialogue_overlap += overlap
                  dialogue_segments.append(
                      {
                          "start": round(seg["start"], 2),
                          "end": round(seg["end"], 2),
                          "overlap_seconds": round(overlap, 2),
                          "text": seg["text"],
                      }
                  )
      
              narration_slots = []
              for window in quiet_windows:
                  overlap = _overlap_seconds(start, end, window["start"], window["end"])
                  if overlap <= 0:
                      continue
                  slot_start = max(start, window["start"])
                  slot_end = min(end, window["end"])
                  narration_slots.append(
                      {
                          "start": round(slot_start, 2),
                          "end": round(slot_end, 2),
                          "duration": round(slot_end - slot_start, 2),
                          "char_budget": _recommended_char_budget(slot_start, slot_end),
                      }
                  )
      
              fusion.append(
                  {
                      "scene_id": scene["scene_id"],
                      "time_range": [round(start, 2), round(end, 2)],
                      "visual_description": scene["description"],
                      "depth_analysis": scene.get("depth_analysis", ""),
                      "frame_facts": scene.get("frame_facts", {}),
                      "dialogue_segments": dialogue_segments,
                      "dialogue_overlap_seconds": round(dialogue_overlap, 2),
                      "narration_slots": narration_slots,
                      "recommended_mode": "quiet-slot"
                      if narration_slots and dialogue_overlap < (end - start) * 0.4
                      else "ducked-bed",
                  }
              )
          return fusion
      
    • understand.py 192 B
      #!/usr/bin/env python3
      """CLI entrypoint for the self-contained video-understanding skill."""
      
      from understanding_runner import main
      
      __all__ = ["main"]
      
      if __name__ == "__main__":
          main()
      
    • understanding_brief.py 7.4 KB
      """Format research context and rebuild briefs from cached analysis."""
      
      import json
      
      from pathlib import Path
      
      from lib import CONFIG, log, load_background_research
      from asr_timing_evidence import asr_evidence_summary_for_brief
      
      
      from detect import detect_speech_boundary_anchors
      
      
      from brief import build_agent_brief, assess_understanding_substrate
      
      
      from understanding_cache import _load_json, _merge_overview_into_scenes
      from understanding_storyboard import (
          _generate_edited_storyboard,
          _generate_source_storyboard,
          _prepend_storyboard_brief_header,
      )
      
      
      def _clip_text(text, limit):
          value = " ".join(str(text or "").split()).strip()
          return value[:limit]
      
      
      def _research_context(work_dir):
          """Fold background_research.json into a compact context string for the VLM prompt.
      
          The agent does story research first (per references/research-guide.md) and writes
          work_dir/background_research.json; this surfaces synopsis, named characters,
          relationships, plot arcs, and cultural notes so scene VLM analysis can name people
          and read scenes with plot knowledge instead of labelling everyone "黑衣男子".
          Returns "" when no usable research file is present, so behaviour is unchanged.
          """
          data = load_background_research(work_dir)
          if not data:
              return ""
          parts = []
          for key in ("synopsis", "episode_context", "worldbuilding"):
              val = _clip_text(data.get(key), 320)
              if val:
                  parts.append(val)
          chars = data.get("characters")
          if isinstance(chars, dict) and chars:
              named = ";".join(
                  f"{_clip_text(name, 40)}({_clip_text(desc, 120)})"
                  for name, desc in list(chars.items())[:12]
                  if _clip_text(name, 40)
              )
              if named:
                  parts.append("主要人物:" + named)
          details = data.get("character_details")
          if isinstance(details, dict) and details:
              detail_lines = []
              for name, info in list(details.items())[:8]:
                  if not isinstance(info, dict):
                      continue
                  bits = []
                  aliases = info.get("aliases")
                  if isinstance(aliases, list) and aliases:
                      clean_aliases = [_clip_text(alias, 30) for alias in aliases[:3]]
                      clean_aliases = [alias for alias in clean_aliases if alias]
                      if clean_aliases:
                          bits.append("别名" + "/".join(clean_aliases))
                  role = _clip_text(info.get("role"), 60)
                  if role:
                      bits.append(role)
                  rels = info.get("relationships")
                  if isinstance(rels, list) and rels:
                      clean_rels = [_clip_text(rel, 60) for rel in rels[:3]]
                      bits.extend(rel for rel in clean_rels if rel)
                  clean_name = _clip_text(name, 40)
                  if clean_name and bits:
                      detail_lines.append(f"{clean_name}({';'.join(bits)})")
              if detail_lines:
                  parts.append("人物关系:" + ";".join(detail_lines))
          arcs = data.get("plot_arcs")
          if isinstance(arcs, list) and arcs:
              arc_lines = []
              for arc in arcs[:6]:
                  if not isinstance(arc, dict):
                      val = _clip_text(arc, 120)
                      if val:
                          arc_lines.append(val)
                      continue
                  name = _clip_text(arc.get("name"), 50)
                  desc = _clip_text(arc.get("description"), 120)
                  status = _clip_text(arc.get("status"), 30)
                  if name or desc:
                      tail = f"[{status}]" if status else ""
                      arc_lines.append(f"{name}:{desc}{tail}".strip(":"))
              if arc_lines:
                  parts.append("剧情线:" + ";".join(arc_lines))
          notes = data.get("cultural_notes")
          if isinstance(notes, list) and notes:
              note_lines = []
              for note in notes[:4]:
                  if not isinstance(note, dict):
                      val = _clip_text(note, 100)
                      if val:
                          note_lines.append(val)
                      continue
                  item = _clip_text(note.get("item"), 50)
                  expl = _clip_text(note.get("explanation"), 100)
                  if item or expl:
                      note_lines.append(f"{item}:{expl}".strip(":"))
              if note_lines:
                  parts.append("背景注释:" + ";".join(note_lines))
          return " ".join(parts).strip()[:1200]
      
      
      def _load_understanding_artifacts_for_brief(work_dir):
          """Load existing analysis artifacts for a brief-only regeneration pass.
      
          Material-library restores are allowed to reuse expensive analysis artifacts,
          but cut pass 2 must still rebuild `agent_narration_brief.md` against the
          rendered OUTPUT timeline. This helper deliberately performs no extraction,
          ASR, VLM, or external API calls; it only reads already-present JSON. An absent
          artifact contributes [] (the brief documents the thin substrate); a corrupt one raises.
          """
          work_dir = Path(work_dir)
      
          def _own(name):
              path = work_dir / name
              return _load_json(path) if path.exists() else []
      
          return _own("vlm_analysis.json"), _own("asr_result.json"), _own("silence_periods.json")
      
      
      def _write_brief_from_existing_artifacts(video, work_dir, args, video_duration):
          """Regenerate only the agent brief/timeline-fusion from existing artifacts."""
          scenes, asr_result, silence_periods = _load_understanding_artifacts_for_brief(
              work_dir
          )
          if not (Path(work_dir) / "speech_boundary_anchors.json").exists():
              detect_speech_boundary_anchors(work_dir, asr_result)
          overview_path = Path(work_dir) / "mimo_video_overview.json"
          if CONFIG["mimo_video_overview"]:
              scenes = _merge_overview_into_scenes(scenes, overview_path)
      
          source_storyboard = None
          edited_storyboard = None
          scenes_json = Path(work_dir) / "scenes.json"
          if scenes_json.exists():
              source_storyboard = _generate_source_storyboard(
                  work_dir, Path(video), scenes, scenes_json, force=False
              )
          edited_storyboard = _generate_edited_storyboard(work_dir, video, force=False)
          cut_mode = (Path(work_dir) / "clip_plan_validated.json").exists()
      
          substrate = assess_understanding_substrate(scenes, asr_result)
          if substrate["level"] != "rich":
              banner = "理解素材为空" if substrate["level"] == "empty" else "理解素材偏薄"
              log(
                  f"⚠️  {banner}:ASR {substrate['asr_chars']} 字 | 场景 {substrate['scene_count']} | "
                  f"带 frame_facts 的场景 {substrate['scenes_with_frame_facts']} | 平均画面描述 {substrate['avg_description_len']} 字"
              )
          brief_path = build_agent_brief(
              scenes,
              asr_result,
              silence_periods,
              video_duration,
              work_dir,
              args.style,
              mimo_overview_enabled=CONFIG["mimo_video_overview"],
              mimo_overview_video_path=video,
              asr_evidence=asr_evidence_summary_for_brief(work_dir, video),
          )
          _prepend_storyboard_brief_header(
              brief_path, source_storyboard, edited_storyboard, cut_mode=cut_mode
          )
          log("=" * 50)
          log(f"brief-only 完成。写作 brief: {brief_path}")
          print(
              json.dumps(
                  {
                      "status": "brief_only",
                      "work_dir": str(work_dir),
                      "brief": str(brief_path),
                      "substrate": substrate["level"],
                      "scenes": len(scenes),
                      "asr_segments": len(asr_result),
                  },
                  ensure_ascii=False,
              )
          )
      
    • understanding_cache.py 12 KB
      """Own understanding-stage cache keys, provenance, and statuses."""
      
      import json
      
      from pathlib import Path
      
      from asr_timing_evidence import EVIDENCE_FILENAME, validate_asr_timing_evidence
      from extract import FRAME_TIME_CONVENTION_VERSION
      from lib import CONFIG, log, file_identity, load_prompt
      
      
      from vlm import (
          _is_mimo_chunk_usable,
      )
      
      
      def _fresh(out, *inputs):
          out = Path(out)
          if not out.exists():
              return False
          ins = [Path(p) for p in inputs if p and Path(p).exists()]
          if not ins:
              return out.stat().st_size > 0
          return out.stat().st_mtime >= max(p.stat().st_mtime for p in ins)
      
      
      def _artifact_meta_path(artifact_path):
          artifact_path = Path(artifact_path)
          return artifact_path.with_name(f"{artifact_path.name}.meta.json")
      
      
      def _load_json(path):
          return json.loads(Path(path).read_text(encoding="utf-8"))
      
      
      def _video_input(video_path):
          """The source video as a cache input: resolved path plus {size, mtime_ns}, so two
          different files written in the same timestamp tick never share a stage output."""
          return {"path": str(Path(video_path).resolve()), **file_identity(video_path)}
      
      
      def _artifact_identity(path):
          """{size, mtime_ns} of an input artifact, or None when it does not exist."""
          path = Path(path)
          return file_identity(path) if path.exists() else None
      
      
      def _stage_cache_valid(artifact_path, expected_meta):
          """A stage output is reusable when it exists, is non-empty, its sidecar equals the
          expected inputs/settings dict, and the output itself is the one the sidecar recorded."""
          artifact_path = Path(artifact_path)
          if not artifact_path.exists() or artifact_path.stat().st_size == 0:
              return False
          meta_path = _artifact_meta_path(artifact_path)
          if not meta_path.exists():
              return False
          try:
              meta = _load_json(meta_path)
          except (OSError, ValueError, TypeError):
              return False
          if not isinstance(meta, dict) or meta.get("artifact") != file_identity(artifact_path):
              return False
          expected = dict(expected_meta)
          expected["artifact"] = meta["artifact"]
          return meta == expected
      
      
      def _write_stage_meta(artifact_path, meta):
          meta_path = _artifact_meta_path(artifact_path)
          payload = dict(meta)
          payload["artifact"] = _artifact_identity(artifact_path)
          meta_path.write_text(
              json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8"
          )
      
      
      def _remove_stage_meta(artifact_path):
          _artifact_meta_path(artifact_path).unlink(missing_ok=True)
      
      
      def _short_status_message(message, limit=240):
          """Compact optional-stage messages for sidecars; no tracebacks or bulky payloads."""
          text = " ".join(str(message or "").split())
          return text[:limit]
      
      
      def _write_optional_stage_status(work_dir, filename, payload):
          path = Path(work_dir) / filename
          safe = dict(payload)
          safe["message"] = _short_status_message(safe.get("message", ""))
          path.write_text(json.dumps(safe, ensure_ascii=False, indent=2), encoding="utf-8")
          return path
      
      
      def _write_mimo_overview_status(
          work_dir, status, message="", artifact=None, *, enabled=None
      ):
          return _write_optional_stage_status(
              work_dir,
              "mimo_video_overview.status.json",
              {
                  "stage": "mimo_video_overview",
                  "enabled": bool(CONFIG["mimo_video_overview"])
                  if enabled is None
                  else bool(enabled),
                  "status": status,
                  "message": message,
                  "artifact": artifact,
              },
          )
      
      
      def _merge_overview_into_scenes(scenes, overview_path):
          """Make the MiMo video-overview the PRIMARY per-scene description when present.
      
          The frame VLM still provides `frame_facts` (timestamped grounding) and `depth_analysis`;
          this only replaces the per-scene `description` with the motion-aware video-overview analysis,
          keeping the original frame description under `frame_description` for provenance/fallback.
          Scenes whose overview chunk was missing or moderation-rejected keep the frame description.
      
          In-memory only — `vlm_analysis.json` on disk stays the pure frame-VLM product, so the VLM
          cache stays coherent and the merge is re-derived (frames + overview) every run. Because
          `frame_facts` is untouched, `assess_understanding_substrate` (which grades on frame_facts +
          ASR) cannot regress; richer descriptions can only help.
          """
          if not Path(overview_path).exists():
              return scenes
          overview = _load_json(overview_path)
          by_scene = {}
          for chunk in overview["chunks"]:
              content = chunk["content"].strip()
              if _is_mimo_chunk_usable(content):
                  by_scene.setdefault(chunk["scene_id"], []).append(content)
          if not by_scene:
              return scenes
          enriched = 0
          for scene in scenes:
              contents = by_scene.get(scene["scene_id"])
              if not contents:
                  continue
              scene.setdefault("frame_description", scene.get("description", ""))
              scene["description"] = "\n".join(contents)
              scene["description_source"] = "mimo_video_overview"
              enriched += 1
          if enriched:
              log(
                  f"已用 MiMo 视频概览增强 {enriched} 个场景的描述(逐帧 frame_facts 保留作锚点)"
              )
          return scenes
      
      
      def _write_consolidation_status(
          work_dir,
          status,
          message="",
          artifacts=None,
          *,
          enabled=True,
          do_asr=False,
          do_index=True,
      ):
          return _write_optional_stage_status(
              work_dir,
              "consolidation.status.json",
              {
                  "stage": "consolidation",
                  "enabled": bool(enabled),
                  "do_asr": bool(do_asr),
                  "do_index": bool(do_index),
                  "status": status,
                  "message": message,
                  "artifacts": list(artifacts or []),
              },
          )
      
      
      def _present_consolidation_artifacts(work_dir):
          work_dir = Path(work_dir)
          return [
              name
              for name in ("understanding_index.json", "asr_clean.json")
              if (work_dir / name).exists()
          ]
      
      
      def _scene_cache_payload(video_path):
          return {
              "schema_version": 1,
              "stage": "scenes",
              "inputs": {"video": _video_input(video_path)},
              "settings": {
                  "scene_threshold": CONFIG.get("scene_threshold"),
                  "scene_junk_filter": CONFIG.get("scene_junk_filter"),
                  "scene_merge_min": CONFIG.get("scene_merge_min"),
                  "scene_junk_dark_luma": CONFIG.get("scene_junk_dark_luma"),
                  "scene_junk_bright_luma": CONFIG.get("scene_junk_bright_luma"),
                  "scene_junk_pixel_ratio": CONFIG.get("scene_junk_pixel_ratio"),
              },
          }
      
      
      def _asr_cache_payload(video_path, *, skip_asr=False):
          return {
              "schema_version": 1,
              "stage": "asr",
              "inputs": {"video": _video_input(video_path)},
              "settings": {
                  "skip_asr": bool(skip_asr),
                  "mimo_asr_api_key_present": bool(CONFIG.get("mimo_asr_api_key")),
                  "mimo_asr_api_url": CONFIG.get("mimo_asr_api_url"),
                  "mimo_asr_model": CONFIG.get("mimo_asr_model"),
                  "mimo_asr_language": CONFIG.get("mimo_asr_language"),
                  "mimo_asr_base64_max_mb": CONFIG.get("mimo_asr_base64_max_mb"),
                  "asr_segment_seconds": CONFIG.get("asr_segment_seconds"),
              },
          }
      
      
      def _asr_cache_state(artifact_path, expected_meta, video_path):
          """Classify an ASR cache without upgrading stale evidence into new authority."""
          artifact_path = Path(artifact_path)
          if not _stage_cache_valid(artifact_path, expected_meta):
              return "MISS"
          evidence_path = artifact_path.parent / EVIDENCE_FILENAME
          if not evidence_path.exists():
              return "LEGACY_UNVERIFIED"
          if validate_asr_timing_evidence(evidence_path, video_path, artifact_path):
              evidence = _load_json(evidence_path)
              if evidence.get("status") == "LEGACY_UNVERIFIED":
                  return "LEGACY_UNVERIFIED"
              # An all-empty transcription is an unexplained outcome, not proven silence; treating it
              # as fresh would make one bad run a permanent cache hit.
              if evidence.get("status") in {"UNAVAILABLE_NO_DURATION", "EMPTY_UNKNOWN"}:
                  return "MISS"
              return "FRESH"
          return "MISS"
      
      
      def _silence_cache_payload(video_path, asr_json):
          return {
              "schema_version": 1,
              "stage": "silence",
              "inputs": {
                  "video": _video_input(video_path),
                  "asr_result": _artifact_identity(asr_json),
              },
              "asr_meta": _load_json(_artifact_meta_path(asr_json))
              if _artifact_meta_path(asr_json).exists()
              else None,
              "settings": {
                  "silence_noise_threshold": CONFIG.get("silence_noise_threshold"),
                  "silence_min_duration": CONFIG.get("silence_min_duration"),
                  "quiet_window_min": CONFIG.get("quiet_window_min"),
                  "silence_merge_gap": CONFIG.get("silence_merge_gap"),
                  "source_boundary_noise_threshold": CONFIG.get(
                      "source_boundary_noise_threshold"
                  ),
                  "source_boundary_min_pause": CONFIG.get("source_boundary_min_pause"),
                  "source_boundary_max_alignment_error": CONFIG.get(
                      "source_boundary_max_alignment_error"
                  ),
              },
          }
      
      
      def _vlm_prompt_payload():
          """The exact prompt/context text the VLM stage sends; compared by equality."""
          prompt = load_prompt("VLM_DEPTH_PROMPT")
          if not prompt:
              prompt = (
                  "仔细观察这些视频帧。分两部分输出:\n"
                  "【描述】不超过80字,描述画面中正在发生什么。\n"
                  "【深层分析】不超过120字,分析角色情绪、关系动态、潜台词。"
              )
          context = CONFIG.get("context_info", "")
          if context:
              prompt = f"已知信息:{context}\n\n{prompt}"
          return {"prompt_text": prompt, "context_info": context}
      
      
      def _vlm_cache_payload(video_path, work_dir, scenes_json, frames):
          return {
              "schema_version": 1,
              "stage": "vlm",
              "inputs": {
                  "video": _video_input(video_path),
                  "scenes": _artifact_identity(scenes_json),
                  "background_research": _artifact_identity(
                      Path(work_dir) / "background_research.json"
                  ),
              },
              "frames": _frame_cache_payload(video_path, CONFIG.get("fps"), frames),
              "prompt": _vlm_prompt_payload(),
              "settings": {
                  "fps": CONFIG.get("fps"),
                  "vlm_model": CONFIG.get("vlm_model"),
                  "api_url": CONFIG.get("api_url"),
                  "mimo_disable_thinking": CONFIG.get("mimo_disable_thinking", True),
                  "mimo_media_resolution": CONFIG.get("mimo_media_resolution"),
              },
          }
      
      
      def _frames_manifest_path(work_dir):
          return Path(work_dir) / "frames" / "frames_manifest.json"
      
      
      def _frame_cache_payload(video_path, fps, frames):
          frame_names = [Path(frame).name for frame in frames]
          return {
              # v2: frame number→time is (n-1)/fps, not n/fps. Bumping invalidates frames/VLM
              # artifacts produced under the old off-by-one so they are recomputed, not reused.
              "schema_version": FRAME_TIME_CONVENTION_VERSION,
              "inputs": {"video": _video_input(video_path)},
              "fps": float(fps),
              "frame_count": len(frames),
              "frames": frame_names,
          }
      
      
      def _write_frames_manifest(work_dir, video_path, fps, frames):
          manifest_path = _frames_manifest_path(work_dir)
          manifest_path.parent.mkdir(parents=True, exist_ok=True)
          manifest_path.write_text(
              json.dumps(
                  _frame_cache_payload(video_path, fps, frames), ensure_ascii=False, indent=2
              ),
              encoding="utf-8",
          )
      
      
      def _frames_cache_valid(video_path, work_dir, fps):
          frames_dir = Path(work_dir) / "frames"
          frames = sorted(frames_dir.glob("frame_*.jpg"))
          if not frames:
              return False
          manifest_path = _frames_manifest_path(work_dir)
          if not manifest_path.exists():
              return False
          try:
              manifest = json.loads(manifest_path.read_text(encoding="utf-8"))
              expected = _frame_cache_payload(video_path, fps, frames)
          except (OSError, ValueError, TypeError):
              return False
          return manifest == expected
      
    • understanding_runner.py 14.7 KB
      """Orchestrate the video-understanding stages."""
      
      import argparse
      
      
      import json
      
      from pathlib import Path
      
      from lib import CONFIG, log, get_video_duration, api_call
      
      from extract import extract_frames
      
      from detect import detect_scenes, detect_silence_periods, detect_speech_boundary_anchors
      
      from asr import transcribe_audio
      from asr_timing_evidence import write_asr_timing_evidence, asr_evidence_summary_for_brief
      
      from vlm import (
          analyze_scenes,
          analyze_video_overview,
          mimo_video_overview_cache_fresh,
      )
      
      from brief import build_agent_brief, assess_understanding_substrate
      
      
      from understanding_brief import _research_context, _write_brief_from_existing_artifacts
      from understanding_cache import (
          _asr_cache_payload,
          _asr_cache_state,
          _frames_cache_valid,
          _load_json,
          _merge_overview_into_scenes,
          _present_consolidation_artifacts,
          _remove_stage_meta,
          _scene_cache_payload,
          _silence_cache_payload,
          _stage_cache_valid,
          _vlm_cache_payload,
          _write_consolidation_status,
          _write_frames_manifest,
          _write_mimo_overview_status,
          _write_stage_meta,
      )
      from understanding_storyboard import (
          _generate_edited_storyboard,
          _generate_source_storyboard,
          _prepend_storyboard_brief_header,
      )
      
      
      def main():
          ap = argparse.ArgumentParser(
              description="Analyze a video into an understanding index + narration brief."
          )
          ap.add_argument("video")
          ap.add_argument("--work-dir", required=True)
          ap.add_argument(
              "--context", default="", help="extra context (show name, character names, ...)"
          )
          ap.add_argument("--scene-threshold", type=float, default=None)
          ap.add_argument("--style", default="纪录片")
          ap.add_argument(
              "--edit-mode",
              default=None,
              choices=["full", "cut"],
              help="recap mode to document in the writing brief",
          )
          ap.add_argument(
              "--target-duration",
              default=None,
              help="cut-mode target duration to document in the writing brief",
          )
          ap.add_argument("--skip-asr", action="store_true")
          ap.add_argument("--mimo-video-overview", action="store_true")
          ap.add_argument(
              "--force", action="store_true", help="ignore cached artifacts and recompute"
          )
          ap.add_argument(
              "--brief-only",
              action="store_true",
              help="rebuild agent_narration_brief.md from existing artifacts only; no extraction/API",
          )
          ap.add_argument(
              "--consolidate",
              action=argparse.BooleanOptionalAction,
              default=True,
              help="build the global understanding story index (Pass B); default ON, --no-consolidate to skip",
          )
          ap.add_argument(
              "--consolidate-asr",
              action="store_true",
              help="also clean the ASR transcript (Pass A)",
          )
          args = ap.parse_args()
      
          video = args.video
          work_dir = Path(args.work_dir)
          work_dir.mkdir(parents=True, exist_ok=True)
          # Story research (if the agent wrote background_research.json first) feeds the VLM
          # context, so scene analysis can name characters and read scenes with plot knowledge.
          research_ctx = _research_context(work_dir)
          context_parts = [p for p in (research_ctx, args.context) if p and p.strip()]
          if context_parts:
              CONFIG["context_info"] = " ".join(context_parts)
          if research_ctx:
              log(f"已并入 background_research.json 到理解上下文({len(research_ctx)} 字)")
          if args.scene_threshold is not None:
              CONFIG["scene_threshold"] = args.scene_threshold
          if args.edit_mode is not None:
              CONFIG["edit_mode"] = args.edit_mode
          if args.target_duration is not None:
              CONFIG["target_duration"] = args.target_duration
          if args.mimo_video_overview:
              CONFIG["mimo_video_overview"] = True
      
          video_duration = get_video_duration(video)
          if CONFIG["fps"] <= 0:
              CONFIG["fps"] = (
                  2 if video_duration <= 60 else (1.5 if video_duration <= 300 else 1)
              )
          log(f"FPS: {CONFIG['fps']} (视频时长: {video_duration:.1f}s)")
      
          if args.brief_only:
              _write_brief_from_existing_artifacts(video, work_dir, args, video_duration)
              return
      
          scenes_json = work_dir / "scenes.json"
          asr_json = work_dir / "asr_result.json"
          silence_json = work_dir / "silence_periods.json"
          vlm_json = work_dir / "vlm_analysis.json"
          frames_dir = work_dir / "frames"
      
          # Step 1: frame extraction
          if not args.force and _frames_cache_valid(video, work_dir, CONFIG["fps"]):
              frames = sorted(frames_dir.glob("frame_*.jpg"))
              log(f"跳过帧提取(缓存匹配 {len(frames)} 帧)")
          else:
              frames = extract_frames(video, work_dir)
              _write_frames_manifest(work_dir, video, CONFIG["fps"], frames)
      
          # Step 2: scene detection
          scenes_meta = _scene_cache_payload(video)
          if not args.force and _stage_cache_valid(scenes_json, scenes_meta):
              scenes = _load_json(scenes_json)
              log(f"跳过场景检测(已存在 {len(scenes)} 个场景)")
          else:
              scenes = detect_scenes(video, work_dir, CONFIG["scene_threshold"])
              _write_stage_meta(scenes_json, scenes_meta)
      
          # Step 3: ASR
          asr_meta = _asr_cache_payload(video, skip_asr=args.skip_asr)
          cache_state = None
          if not args.skip_asr and not args.force:
              cache_state = _asr_cache_state(asr_json, asr_meta, video)
          if args.skip_asr:
              asr_result = []
              asr_json.write_text(
                  json.dumps(asr_result, ensure_ascii=False, indent=2), encoding="utf-8"
              )
              _write_stage_meta(asr_json, asr_meta)
              write_asr_timing_evidence(
                  work_dir, video, "EXPLICITLY_SKIPPED", final_segments=asr_result
              )
              log("跳过 ASR(--skip-asr)")
          elif cache_state in {
              "FRESH",
              "LEGACY_UNVERIFIED",
          }:
              asr_result = _load_json(asr_json)
              if cache_state == "LEGACY_UNVERIFIED":
                  write_asr_timing_evidence(
                      work_dir,
                      video,
                      "LEGACY_UNVERIFIED",
                      final_segments=asr_result,
                  )
                  log(f"复用旧 ASR({len(asr_result)} 段;时间/声学证据未经验证)")
              else:
                  log(f"跳过 ASR(证据匹配,已存在 {len(asr_result)} 段)")
          else:
              try:
                  asr_result = transcribe_audio(video, work_dir)
              except Exception as e:
                  _remove_stage_meta(asr_json)
                  asr_json.unlink(missing_ok=True)
                  raise RuntimeError(
                      f"ASR 失败;未写入可复用缓存,请修复后重试或显式使用 --skip-asr: {e}"
                  ) from e
              _write_stage_meta(asr_json, asr_meta)
      
          # Step 3.5: silence detection
          silence_meta = _silence_cache_payload(video, asr_json)
          if not args.force and _stage_cache_valid(silence_json, silence_meta):
              silence_periods = _load_json(silence_json)
              log(f"跳过静音检测(已存在 {len(silence_periods)} 个窗口)")
              if not (work_dir / "speech_boundary_anchors.json").exists():
                  detect_speech_boundary_anchors(work_dir, asr_result)
          else:
              silence_periods = detect_silence_periods(video, work_dir, asr_result)
              _write_stage_meta(silence_json, silence_meta)
      
          # Step 4: VLM analysis (the only stage that requires the chat API key)
          vlm_meta = _vlm_cache_payload(video, work_dir, scenes_json, frames)
          if not args.force and _stage_cache_valid(vlm_json, vlm_meta):
              vlm_analysis = _load_json(vlm_json)
              log(f"跳过 VLM 分析(已存在 {len(vlm_analysis)} 个场景)")
          else:
              if not CONFIG["api_key"]:
                  key_name = CONFIG["api_env_var"]
                  raise SystemExit(f"请设置 {key_name} 环境变量(VLM 画面分析需要)")
              log("VLM API 连通性预检...")
              api_call(
                  {
                      "model": CONFIG["vlm_model"],
                      "messages": [{"role": "user", "content": "hi"}],
                      "max_tokens": 5,
                  }
              )
              vlm_analysis = analyze_scenes(scenes, frames, work_dir, resume=not args.force)
              _write_stage_meta(vlm_json, vlm_meta)
      
          # Step 4.1: optional MiMo scene-chunk video understanding
          overview_path = work_dir / "mimo_video_overview.json"
          if CONFIG["mimo_video_overview"]:
              if not CONFIG["mimo_video_api_key"]:
                  log("跳过 MiMo 分片视频概览:未设置 MIMO_API_KEY")
                  overview_path.unlink(missing_ok=True)
                  _write_mimo_overview_status(
                      work_dir,
                      "skipped_no_key",
                      "未设置 MIMO_API_KEY,MiMo 分片视频概览未运行",
                      None,
                  )
              elif mimo_video_overview_cache_fresh(overview_path, video, scenes):
                  log("跳过 MiMo 分片视频概览(缓存匹配)")
                  _write_mimo_overview_status(
                      work_dir, "cached", "缓存匹配", overview_path.name
                  )
              else:
                  overview_path.unlink(missing_ok=True)
                  try:
                      overview = analyze_video_overview(video, work_dir, scenes)
                  except Exception as e:
                      log(f"MiMo 分片视频概览失败(忽略): {e}")
                      _write_mimo_overview_status(work_dir, "failed", e, None)
                  else:
                      if overview:
                          _write_mimo_overview_status(
                              work_dir, "ok", "MiMo 分片视频概览完成", overview_path.name
                          )
                      else:
                          _write_mimo_overview_status(
                              work_dir, "failed", "MiMo 分片视频概览未产出有效 artifact", None
                          )
          else:
              overview_path.unlink(missing_ok=True)
              _write_mimo_overview_status(
                  work_dir, "disabled", "MiMo 分片视频概览未启用", None, enabled=False
              )
      
          # Make the video-overview the primary per-scene description (frame_facts stay the anchor).
          # No-op/revert when overview is absent, so disabling it cleanly returns to frame descriptions.
          vlm_analysis = _merge_overview_into_scenes(vlm_analysis, overview_path)
      
          # optional consolidation (整理): build the understanding index before the brief folds it in
          if args.consolidate or args.consolidate_asr:
              from consolidate import consolidate
      
              try:
                  consolidate(
                      work_dir, do_asr=args.consolidate_asr, do_index=args.consolidate
                  )
              except Exception as e:
                  log(f"consolidate 跳过(忽略): {e}")
                  _write_consolidation_status(
                      work_dir,
                      "failed",
                      e,
                      _present_consolidation_artifacts(work_dir),
                      do_asr=args.consolidate_asr,
                      do_index=args.consolidate,
                  )
              else:
                  expected = []
                  skipped = []
                  if args.consolidate:
                      if vlm_analysis:
                          expected.append("understanding_index.json")
                      else:
                          skipped.append("无 vlm_analysis,跳过 index")
                  if args.consolidate_asr:
                      if asr_result:
                          expected.append("asr_clean.json")
                      else:
                          skipped.append("无 ASR 文本,跳过 ASR 清洗")
                  artifacts = _present_consolidation_artifacts(work_dir)
                  missing = [name for name in expected if name not in artifacts]
                  if missing:
                      _write_consolidation_status(
                          work_dir,
                          "failed",
                          f"未产出预期 artifact: {', '.join(missing)}",
                          artifacts,
                          do_asr=args.consolidate_asr,
                          do_index=args.consolidate,
                      )
                  elif expected:
                      _write_consolidation_status(
                          work_dir,
                          "ok",
                          "consolidation 完成",
                          artifacts,
                          do_asr=args.consolidate_asr,
                          do_index=args.consolidate,
                      )
                  else:
                      _write_consolidation_status(
                          work_dir,
                          "skipped",
                          ";".join(skipped) or "无可整理输入",
                          artifacts,
                          do_asr=args.consolidate_asr,
                          do_index=args.consolidate,
                      )
          else:
              _write_consolidation_status(
                  work_dir,
                  "disabled",
                  "consolidation 未启用",
                  [],
                  enabled=False,
                  do_asr=False,
                  do_index=False,
              )
      
          # Storyboard contact sheets (advisory, never blocking). Source uses scene anchors over the
          # source timeline; edited is gated on clip_plan_validated.json file-presence (NOT edit_mode —
          # recap.py forwards --edit-mode cut in BOTH passes, so the validated plan is the only reliable
          # pass2 signal). Both cache via _write_stage_meta/_stage_cache_valid with fps + frame-set in the key.
          source_storyboard = _generate_source_storyboard(
              work_dir, Path(video), scenes, scenes_json, force=args.force
          )
          edited_storyboard = _generate_edited_storyboard(work_dir, video, force=args.force)
          cut_mode = (work_dir / "clip_plan_validated.json").exists()
      
          # understanding substrate warning + writing brief
          substrate = assess_understanding_substrate(vlm_analysis, asr_result)
          if substrate["level"] != "rich":
              banner = "理解素材为空" if substrate["level"] == "empty" else "理解素材偏薄"
              log(
                  f"⚠️  {banner}:ASR {substrate['asr_chars']} 字 | 场景 {substrate['scene_count']} | "
                  f"带 frame_facts 的场景 {substrate['scenes_with_frame_facts']} | 平均画面描述 {substrate['avg_description_len']} 字"
              )
          brief_path = build_agent_brief(
              vlm_analysis,
              asr_result,
              silence_periods,
              video_duration,
              work_dir,
              args.style,
              mimo_overview_enabled=CONFIG["mimo_video_overview"],
              mimo_overview_video_path=video,
              asr_evidence=asr_evidence_summary_for_brief(work_dir, video),
          )
          # C1: post-process the RETURNED brief FILE (not brief.py) so the brief⇄narration twin stays
          # byte-identical. Prepends a storyboard header pointing the agent at the sheet(s).
          _prepend_storyboard_brief_header(
              brief_path, source_storyboard, edited_storyboard, cut_mode=cut_mode
          )
      
          log("=" * 50)
          log(f"理解完成。写作 brief: {brief_path}")
          print(
              json.dumps(
                  {
                      "status": "analyzed",
                      "work_dir": str(work_dir),
                      "brief": str(brief_path),
                      "substrate": substrate["level"],
                      "scenes": len(scenes),
                      "asr_segments": len(asr_result),
                  },
                  ensure_ascii=False,
              )
          )
      
    • understanding_storyboard.py 6.8 KB
      """Generate and cache source and edited storyboard sheets."""
      
      from pathlib import Path
      
      from lib import CONFIG, log
      
      
      from storyboard import build_source_storyboard, build_edited_storyboard
      
      
      from understanding_cache import (
          _artifact_identity,
          _video_input,
          _frames_manifest_path,
          _load_json,
          _stage_cache_valid,
          _write_stage_meta,
      )
      
      
      def _storyboard_sample_policy():
          return {
              "max_tiles": CONFIG["storyboard_max_tiles"],
              "columns": CONFIG["storyboard_columns"],
          }
      
      
      def _edited_storyboard_meta(clip_plan_validated_json, frames_manifest_path):
          """Cache key for the edited storyboard: clip plan + fps + frame-set all in the key, so an
          fps change OR a re-validated plan invalidates it. font availability is deliberately NOT in
          the key (the JSON `labels_burned` flag surfaces it instead)."""
          return {
              "schema_version": 1,
              "stage": "edited_storyboard",
              "inputs": {
                  "clip_plan_validated": _artifact_identity(clip_plan_validated_json),
                  "frames_manifest": _artifact_identity(frames_manifest_path),
              },
              "fps": float(CONFIG.get("fps") or 0),
              "sample_policy": _storyboard_sample_policy(),
          }
      
      
      def _generate_source_storyboard(
          work_dir, video_path, scenes, scenes_json, *, force=False
      ):
          """Generate (or reuse cached) the source storyboard. Advisory: returns dict|None, never raises.
      
          Cached via _write_stage_meta/_stage_cache_valid on storyboard/source_storyboard.json; the
          meta includes fps + the frames-manifest identity so an fps-change resume rebuilds (Principle 5).
          If frames/ is absent (cache hit skipped extraction / cleaned) → skip + log; pipeline continues.
          """
          if not CONFIG["storyboard"]:
              return None
          frames_dir = Path(work_dir) / "frames"
          if not frames_dir.is_dir() or not any(frames_dir.glob("frame_*.jpg")):
              log("storyboard 跳过 source:frames/ 缺失(缓存命中跳过了帧提取?)")
              return None
          json_path = Path(work_dir) / "storyboard" / "source_storyboard.json"
          meta = {
              "schema_version": 1,
              "stage": "source_storyboard",
              "inputs": {
                  "video": _video_input(video_path),
                  "scenes": _artifact_identity(scenes_json),
                  "frames_manifest": _artifact_identity(_frames_manifest_path(work_dir)),
              },
              "fps": float(CONFIG.get("fps") or 0),
              "sample_policy": _storyboard_sample_policy(),
          }
          if not force and _stage_cache_valid(json_path, meta):
              try:
                  cached = _load_json(json_path)
              except (OSError, ValueError):
                  log(
                      "storyboard source 缓存命中但文件损坏,重建"
                  )  # advisory: never raise out
              else:
                  log("storyboard 跳过 source(缓存匹配)")
                  return cached
          result = build_source_storyboard(work_dir, video_path, scenes, CONFIG["fps"])
          if result is not None and json_path.exists():
              _write_stage_meta(json_path, meta)
          return result
      
      
      def _generate_edited_storyboard(work_dir, source_video_path, *, force=False):
          """Generate (or reuse cached) the edited storyboard, GATED on clip_plan_validated.json
          file-presence (NOT on edit_mode — recap.py forwards --edit-mode cut in BOTH passes, so the
          validated plan presence is the only reliable pass2 signal). Advisory: returns dict|None.
          """
          if not CONFIG["storyboard"]:
              return None
          clip_plan_validated_json = Path(work_dir) / "clip_plan_validated.json"
          if not clip_plan_validated_json.exists():
              return None  # pass1 (no validated plan yet) → no edited storyboard
          frames_dir = Path(work_dir) / "frames"
          if not frames_dir.is_dir() or not any(frames_dir.glob("frame_*.jpg")):
              log("storyboard 跳过 edited:frames/ 缺失(缓存命中跳过了帧提取?)")
              return None
          json_path = Path(work_dir) / "storyboard" / "edited_storyboard.json"
          meta = _edited_storyboard_meta(
              clip_plan_validated_json, _frames_manifest_path(work_dir)
          )
          if not force and _stage_cache_valid(json_path, meta):
              try:
                  cached = _load_json(json_path)
              except (OSError, ValueError):
                  log(
                      "storyboard edited 缓存命中但文件损坏,重建"
                  )  # advisory: never raise out
              else:
                  log("storyboard 跳过 edited(缓存匹配)")
                  return cached
          try:
              clip_plan_validated = _load_json(clip_plan_validated_json)
          except (OSError, ValueError):
              log("storyboard 跳过 edited:clip_plan_validated.json 无法解析")
              return None
          result = build_edited_storyboard(
              work_dir, source_video_path, clip_plan_validated, CONFIG["fps"]
          )
          if result is not None and json_path.exists():
              _write_stage_meta(json_path, meta)
          return result
      
      
      def _prepend_storyboard_brief_header(
          brief_path, source_storyboard, edited_storyboard, *, cut_mode
      ):
          """Post-process the RETURNED brief markdown FILE (C1): prepend a short storyboard header.
      
          Editing the brief FILE on disk (not brief.py) keeps the brief⇄narration twin byte-identical.
          Branches the edited-storyboard line on clip_plan_validated presence (edited_storyboard truthy)
          so pass1 never prints a not-yet-existing path. If labels_burned:false, point to inspect clip-map.
          """
          if not source_storyboard and not edited_storyboard:
              return
          try:
              brief_path = Path(brief_path)
              lines = ["## Storyboard(先看 storyboard 再写)", ""]
              any_labels_missing = False
              if source_storyboard:
                  pages = source_storyboard.get("page_images") or []
                  lines.append(
                      f"- 源时间线 storyboard: {', '.join(pages)}(tiles 时间戳=原片时间)"
                  )
                  if not source_storyboard.get("labels_burned", False):
                      any_labels_missing = True
              if cut_mode and edited_storyboard:
                  pages = edited_storyboard.get("page_images") or []
                  lines.append(
                      f"- 成片(output)时间线 storyboard: {', '.join(pages)}"
                      "(每块双标 out 时间 / src 原片时间;注意区分两条时间线)"
                  )
                  if not edited_storyboard.get("labels_burned", False):
                      any_labels_missing = True
              if any_labels_missing:
                  lines.append(
                      "- 时间戳未烧入 → 用 `inspect clip-map` 查时间(JSON sidecar 仍为权威时间源)"
                  )
              lines.append("")
              header = "\n".join(lines) + "\n"
              existing = brief_path.read_text(encoding="utf-8") if brief_path.exists() else ""
              brief_path.write_text(header + existing, encoding="utf-8")
          except OSError as exc:
              log(f"storyboard brief 头部写入失败(忽略): {exc}")
      
    • vlm.py 28.1 KB
      import base64
      import json
      import mimetypes
      import re
      from collections import OrderedDict
      from pathlib import Path
      from concurrent.futures import ThreadPoolExecutor, as_completed
      from threading import Lock
      
      from extract import (
          FRAME_TIME_CONVENTION_VERSION,
          frame_time,
          parse_frame_number,
      )
      from lib import CONFIG
      from lib import log, api_call, load_prompt, mimo_video_api_call, run_cmd, file_identity
      
      # ── Step 4: VLM 视觉分析 ─────────────────────────────────────────────
      
      def _parse_vlm_depth_response(raw_text):
          """解析 VLM 深度分析响应,提取【描述】、【帧标签】和【深层分析】"""
          if not raw_text or not raw_text.strip():
              return "(VLM 无法识别此场景画面)", "", {}
      
          # 提取【描述】
          desc_match = re.search(r'【描述】\s*\n?(.*?)(?=【帧标签】|【深层分析】|$)', raw_text, re.DOTALL)
          description = desc_match.group(1).strip() if desc_match else raw_text.strip()
      
          # 提取【帧标签】
          frame_facts = {}
          facts_match = re.search(r'【帧标签】\s*\n?(.*?)(?=【深层分析】|$)', raw_text, re.DOTALL)
          if facts_match:
              for line in facts_match.group(1).strip().split("\n"):
                  line = line.strip()
                  if not line:
                      continue
                  # 格式: "12.0s | 男子拿起茶壶对嘴喝, 满脸疲惫"
                  m = re.match(r'([\d.]+)\s*s?\s*\|\s*(.+)', line)
                  if m:
                      ts = m.group(1)
                      actions = [a.strip() for a in re.split(r"[,,;;、]+", m.group(2)) if a.strip()]
                      if actions:
                          frame_facts[ts] = actions
      
          # 提取【深层分析】
          depth_match = re.search(r'【深层分析】\s*\n?(.*?)$', raw_text, re.DOTALL)
          depth_analysis = depth_match.group(1).strip() if depth_match else ""
      
          if not description:
              description = "(VLM 无法识别此场景画面)"
      
          return description, depth_analysis, frame_facts
      
      
      def _max_frames_for_duration(duration):
          """Frames the VLM sees for one scene: ~1 per `vlm_seconds_per_frame`, floor 3, capped by
          `vlm_max_frames`. Replaces the old hard cap of 6 that starved long/merged scenes."""
          spf = float(CONFIG["vlm_seconds_per_frame"])
          ceiling = int(CONFIG["vlm_max_frames"])
          return max(3, min(ceiling, round(max(0.0, float(duration)) / spf)))
      
      
      def _vlm_scene_cache_path(work_dir):
          return Path(work_dir) / "vlm_scene_cache.json"
      
      
      def _load_vlm_scene_cache(work_dir, settings):
          """Per-scene VLM resume cache (scene_key -> analysis) written under `settings`.
      
          Tolerant: {} if absent/corrupt, and {} when the stored request settings differ, so a
          partial-failure cache is never resumed after an output-affecting setting changed."""
          path = _vlm_scene_cache_path(work_dir)
          if not path.exists():
              return {}
          try:
              data = json.loads(path.read_text(encoding="utf-8"))
          except (OSError, ValueError):
              return {}
          if not isinstance(data, dict) or data.get("settings") != settings:
              return {}
          scenes = data.get("scenes")
          return scenes if isinstance(scenes, dict) else {}
      
      
      def _flush_vlm_scene_cache(work_dir, settings, cache):
          """Persist the resume cache atomically (temp + rename) so an abort/crash never corrupts it."""
          path = _vlm_scene_cache_path(work_dir)
          try:
              tmp = path.with_name(path.name + ".tmp")
              tmp.write_text(
                  json.dumps({"settings": settings, "scenes": cache}, ensure_ascii=False),
                  encoding="utf-8",
              )
              tmp.replace(path)
          except OSError as exc:
              log(f"VLM 场景缓存写入失败(忽略): {exc}")
      
      
      def _looks_rate_limited(message):
          """Heuristic: was a scene failure a transient rate-limit (worth a low-concurrency retry) rather
          than a persistent error (empty response / parse failure) that re-running would not fix?"""
          m = str(message).lower()
          return "429" in m or "too many requests" in m or "rate limit" in m or "限流" in m
      
      
      def analyze_scenes(scenes, frames, work_dir, *, resume=True):
          """对每个场景的关键帧调用 VLM 进行视觉分析(并行)"""
          if not scenes:
              analyses = []
              vlm_file = work_dir / "vlm_analysis.json"
              vlm_file.write_text(json.dumps(analyses, ensure_ascii=False, indent=2), encoding="utf-8")
              log("VLM 分析完成: 0 个场景")
              return analyses
      
          if not frames:
              raise RuntimeError("VLM 分析需要先提取至少一帧;frames 为空")
      
          fps = CONFIG["fps"]
          if fps <= 0:
              raise ValueError("CONFIG['fps'] 必须大于 0;请先运行完整 pipeline 或指定 --fps")
      
          vlm_prompt = load_prompt("VLM_DEPTH_PROMPT")
          if not vlm_prompt:
              vlm_prompt = "仔细观察这些视频帧。分两部分输出:\n【描述】不超过80字,描述画面中正在发生什么。\n【深层分析】不超过120字,分析角色情绪、关系动态、潜台词。"
      
          ctx = CONFIG["context_info"]
          if ctx:
              vlm_prompt = f"已知信息:{ctx}\n\n{vlm_prompt}"
      
          # 构建帧时间映射 (frame_NNNNN.jpg -> time in seconds)。换算规则由 extract.py 独家定义:
          # ffmpeg 首帧在 t=0 而文件编号从 1 起,所以是 (n-1)/fps,不是 n/fps。
          frame_times = {}
          for f in frames:
              num = parse_frame_number(f)
              if num is None:
                  continue
              frame_times[f] = frame_time(num, fps)
      
          # base64 编码缓存(LRU)。原来是无上界的 dict:整个 VLM 阶段会把用到的每一帧的 base64
          # 都常驻内存,而 base64 比原图还大 1/3。40 分钟视频 fps=1 约 2400 帧 → 数百 MB 常驻。
          # 相邻场景才会复用同一帧,容量取「并发数 × 单场景最大帧数」就够,再多也命中不了。
          b64_capacity = max(
              8,
              int(CONFIG["vlm_max_frames"]) * int(CONFIG["vlm_workers"]),
          )
          b64_cache = OrderedDict()
          b64_lock = Lock()
      
          def _get_b64(frame_path):
              with b64_lock:
                  cached = b64_cache.get(frame_path)
                  if cached is not None:
                      b64_cache.move_to_end(frame_path)
                      return cached
              encoded = base64.b64encode(frame_path.read_bytes()).decode()
              with b64_lock:
                  b64_cache[frame_path] = encoded
                  b64_cache.move_to_end(frame_path)
                  while len(b64_cache) > b64_capacity:
                      b64_cache.popitem(last=False)
              return encoded
      
          def _analyze_single_scene(i, scene):
              """分析单个场景,返回 (scene_id, result_dict)"""
              scene_frames = [f for f, t in frame_times.items()
                              if scene["start"] <= t <= scene["end"]]
              if not scene_frames:
                  mid = (scene["start"] + scene["end"]) / 2
                  scene_frames = [min(frames, key=lambda f: abs(frame_times.get(f, 999) - mid))]
      
              duration = scene["end"] - scene["start"]
              max_frames = _max_frames_for_duration(duration)
      
              if len(scene_frames) > max_frames:
                  step = len(scene_frames) / max_frames
                  scene_frames = [scene_frames[int(j * step)] for j in range(max_frames)]
              else:
                  scene_frames = scene_frames[:max_frames]
      
              content_parts = []
              for f in scene_frames:
                  b64 = _get_b64(f)
                  content_parts.append({
                      "type": "image_url",
                      "image_url": {"url": f"data:image/jpeg;base64,{b64}"}
                  })
      
              # 帧级事实标签:将帧时间注入prompt
              frame_ts_list = [f"{frame_times[f]:.1f}s" for f in scene_frames]
              frame_ts_text = "帧时间点: " + ", ".join(frame_ts_list)
              content_parts.append({"type": "text", "text": frame_ts_text + "\n\n" + vlm_prompt})
      
              payload = {
                  "model": CONFIG["vlm_model"],
                  "messages": [{"role": "user", "content": content_parts}],
                  "max_tokens": int(CONFIG["vlm_max_tokens"]),
              }
      
              log(f"VLM 分析场景 {i+1}/{len(scenes)} ({len(scene_frames)} 帧)...")
      
              raw_response = ""
              for attempt in range(3):
                  resp = api_call(payload)
                  try:
                      msg = resp["choices"][0]["message"]
                      raw_response = (msg.get("content") or msg.get("reasoning_content") or "")
                  except (KeyError, IndexError):
                      log(f"VLM 返回异常: {json.dumps(resp, ensure_ascii=False)[:200]}")
      
                  if raw_response.strip():
                      break
      
                  if attempt < 2:
                      log(f"  场景 {i+1} VLM 返回空,重试 ({attempt+2}/3)...")
                      retry_parts = content_parts[:-1]
                      retry_parts.append({"type": "text", "text": frame_ts_text + "\n\n" + vlm_prompt + "\n请务必按格式输出,不要留空。"})
                      payload = {
                          "model": CONFIG["vlm_model"],
                          "messages": [{"role": "user", "content": retry_parts}],
                          "max_tokens": int(CONFIG["vlm_max_tokens"]),
                      }
      
              if not raw_response.strip():
                  raise RuntimeError("VLM 连续 3 次返回空内容")
      
              # 解析 【描述】、【帧标签】和【深层分析】
              description, depth_analysis, frame_facts = _parse_vlm_depth_response(raw_response)
      
              result = {
                  "scene_id": i,
                  "start": scene["start"],
                  "end": scene["end"],
                  "description": description,
                  "depth_analysis": depth_analysis,
              }
              if frame_facts:
                  result["frame_facts"] = frame_facts
      
              return i, result
      
          # 断点续传:逐场景持久化分析结果,避免少数场景失败(多为 429 限流)时整批画面理解全部作废。
          # The resume cache is valid only under the SAME output-affecting settings the outer stage
          # gate tracks (understanding_cache._vlm_cache_payload). vlm_prompt already folds in
          # context_info (and background_research). The request/endpoint settings below are NOT in
          # vlm_prompt, so they are stored explicitly — otherwise a partial-failure cache reused after
          # flipping e.g. mimo_disable_thinking would yield a stale/mixed analysis.
          cache_settings = {
              "vlm_prompt": vlm_prompt,
              "vlm_model": CONFIG.get("vlm_model"),
              "vlm_max_tokens": CONFIG.get("vlm_max_tokens"),
              "vlm_seconds_per_frame": CONFIG.get("vlm_seconds_per_frame"),
              "vlm_max_frames": CONFIG.get("vlm_max_frames"),
              "fps": round(float(fps), 3),
              "frame_time_convention": FRAME_TIME_CONVENTION_VERSION,
              "api_url": CONFIG.get("api_url"),
              "mimo_disable_thinking": CONFIG.get("mimo_disable_thinking", True),
              "mimo_media_resolution": CONFIG.get("mimo_media_resolution"),
          }
          cache = _load_vlm_scene_cache(work_dir, cache_settings) if resume else {}
      
          def _scene_cache_key(i, scene):
              return "|".join(str(x) for x in (
                  i, round(float(scene["start"]), 3), round(float(scene["end"]), 3),
              ))
      
          analyses = [None] * len(scenes)
          todo = []
          for i, s in enumerate(scenes):
              analyses[i] = cache.get(_scene_cache_key(i, s))
              if analyses[i] is None:
                  todo.append(i)
          if todo and len(todo) < len(scenes):
              log(f"VLM 复用 {len(scenes) - len(todo)} 个已缓存场景,待分析 {len(todo)} 个")
      
          def _run_pass(indices, workers):
              """分析给定场景索引;结果按批持久化到续传缓存(在主线程,无需加锁)。返回 (i, err) 失败列表。
      
              续传缓存每次都是整份重写,原来每完成一个场景就刷一次 → 写入量随场景数平方增长
              (300 个场景约 75MB 冗余写),而且序列化发生在收结果的主循环上。改成按批刷:
              崩溃最多丢失不到一批的进度,而这些场景本来就会在重跑时被重新分析。
              """
              if not indices:
                  return []
              failures = []
              pending = 0
              # 与并发度对齐:一批大致就是「同时在飞的那几个场景」。
              flush_every = max(1, min(16, workers))
              with ThreadPoolExecutor(max_workers=max(1, workers)) as executor:
                  futures = {executor.submit(_analyze_single_scene, i, scenes[i]): i for i in indices}
                  try:
                      for future in as_completed(futures):
                          i = futures[future]
                          try:
                              idx, result = future.result()
                              analyses[idx] = result
                              cache[_scene_cache_key(idx, scenes[idx])] = result
                              pending += 1
                              if pending >= flush_every:
                                  _flush_vlm_scene_cache(work_dir, cache_settings, cache)
                                  pending = 0
                          except Exception as e:  # noqa: BLE001 - 单个场景失败不能拖垮其余已完成的
                              log(f"VLM 场景 {i+1} 分析失败: {e}")
                              failures.append((i, str(e)))
                  finally:
                      # 无论正常结束、失败还是中断,都把这一轮已完成的场景落盘。
                      if pending:
                          _flush_vlm_scene_cache(work_dir, cache_settings, cache)
              return failures
      
          base_workers = min(len(todo) or 1, int(CONFIG["vlm_workers"]))
          log(f"VLM 并行分析 {len(todo)} 个场景 (workers={base_workers})...")
          failures = _run_pass(todo, base_workers)
          if failures:
              # 限流(429)多是并发瞬时拥塞,降并发重试一轮(已成功的从缓存跳过);其它错误(空响应/解析失败)
              # 重试无益,直接保留为失败,不浪费一轮调用。
              rate_limited = [i for i, msg in failures if _looks_rate_limited(msg)]
              persistent = [(i, msg) for i, msg in failures if not _looks_rate_limited(msg)]
              if rate_limited:
                  retry_workers = max(1, base_workers // 4)
                  log(f"VLM {len(rate_limited)} 个场景疑似限流(429),降并发到 {retry_workers} 重试...")
                  persistent += _run_pass(rate_limited, retry_workers)
              failures = persistent
      
          if failures:
              sample = "; ".join(f"场景 {i+1}: {msg}" for i, msg in failures[:3])
              raise RuntimeError(
                  f"VLM 分析失败 {len(failures)}/{len(scenes)} 个场景。其余已缓存到 vlm_scene_cache.json,"
                  f"重跑可断点续传(只重试失败场景)。示例: {sample}"
              )
      
          vlm_file = work_dir / "vlm_analysis.json"
          vlm_file.write_text(json.dumps(analyses, ensure_ascii=False, indent=2), encoding="utf-8")
          _vlm_scene_cache_path(work_dir).unlink(missing_ok=True)  # 完整理解已生成,续传缓存可清理
          log(f"VLM 分析完成: {len(analyses)} 个场景")
          return analyses
      
      
      def _video_data_url(video_path):
          """Return a MiMo-compatible data URL for a local video chunk, or None when too large."""
          max_bytes = int(float(CONFIG["mimo_video_base64_max_mb"]) * 1024 * 1024)
          encoded_size = int(video_path.stat().st_size * 4 / 3) + 128
          if encoded_size > max_bytes:
              log(
                  "MiMo 视频分片超过 base64 上限: "
                  f"编码后约 {encoded_size / 1024 / 1024:.1f}MB,超过限制 {max_bytes / 1024 / 1024:.1f}MB;"
                  "请降低 MIMO_VIDEO_CHUNK_MAX_SECONDS 或 MIMO_VIDEO_FPS"
              )
              return None
          mime_type = mimetypes.guess_type(str(video_path))[0] or "video/mp4"
          encoded = base64.b64encode(video_path.read_bytes()).decode("ascii")
          return f"data:{mime_type};base64,{encoded}"
      
      
      def _mimo_video_chunks(scenes):
          """Build local MiMo video-understanding chunks from ffmpeg scene boundaries."""
          if not scenes:
              raise RuntimeError("MiMo 视频分片理解需要 scenes;请先运行 ffmpeg scene/scdet 场景检测")
      
          max_seconds = float(CONFIG["mimo_video_chunk_max_seconds"])
          min_seconds = float(CONFIG["mimo_video_chunk_min_seconds"])
          chunks = []
          for scene_index, scene in enumerate(scenes):
              start = float(scene["start"])
              end = float(scene["end"])
              if end <= start:
                  continue
              scene_id = scene.get("scene_id", scene_index)
              cursor = start
              while cursor < end:
                  chunk_end = min(end, cursor + max_seconds)
                  if end - chunk_end < min_seconds and chunk_end < end:
                      chunk_end = end
                  if chunk_end > cursor:
                      chunks.append({
                          "chunk_id": len(chunks),
                          "scene_id": scene_id,
                          "start": round(cursor, 3),
                          "end": round(chunk_end, 3),
                      })
                  cursor = chunk_end
          if not chunks:
              raise RuntimeError("MiMo 视频分片理解没有可用分片;请检查 scenes.json")
          return chunks
      
      
      def _extract_video_chunk(video_path, chunk, output_path):
          """Cut one scene-based chunk into a compact local MP4 for MiMo video_url data URL."""
          start = float(chunk["start"])
          duration = max(0.1, float(chunk["end"]) - start)
          fps = float(CONFIG["mimo_video_fps"])
          cmd = [
              "ffmpeg", "-y",
              "-ss", f"{start:.3f}",
              "-t", f"{duration:.3f}",
              "-i", str(video_path),
              "-map", "0:v:0",
              "-an",
              "-vf", f"fps={fps:g}",
              "-c:v", "libx264",
              "-preset", "veryfast",
              "-crf", "30",
              "-pix_fmt", "yuv420p",
              "-movflags", "+faststart",
              str(output_path),
          ]
          result = run_cmd(cmd, timeout=CONFIG["mimo_video_chunk_timeout"])
          if result.returncode != 0:
              raise RuntimeError(f"MiMo 视频分片裁剪失败: {result.stderr[-500:]}")
          return output_path
      
      
      def _mimo_chunk_prompt(chunk):
          return (
              f"这是原视频 {chunk['start']:.1f}s-{chunk['end']:.1f}s 的场景分片,"
              f"scene_id={chunk['scene_id']}。"
              f"{CONFIG['mimo_video_prompt']}"
          )
      
      
      def _mimo_video_model():
          """MiMo 视频理解使用的模型:优先 mimo_video_model,回退 mimo_model,再回退 vlm_model。"""
          return CONFIG.get("mimo_video_model") or CONFIG.get("mimo_model") or CONFIG["vlm_model"]
      
      
      def mimo_video_settings():
          """Return non-secret MiMo video-overview settings that affect generated content."""
          return {
              "model": _mimo_video_model(),
              "mimo_video_api_url": CONFIG.get("mimo_video_api_url"),
              "mimo_video_fps": CONFIG.get("mimo_video_fps", 2.0),
              "mimo_media_resolution": CONFIG.get("mimo_media_resolution", "default"),
              "mimo_video_chunk_max_seconds": CONFIG.get("mimo_video_chunk_max_seconds", 20.0),
              "mimo_video_chunk_min_seconds": CONFIG.get("mimo_video_chunk_min_seconds", 1.0),
              "mimo_video_base64_max_mb": CONFIG.get("mimo_video_base64_max_mb", 45.0),
              "mimo_video_prompt": CONFIG.get("mimo_video_prompt", ""),
              "mimo_disable_thinking": CONFIG.get("mimo_disable_thinking", True),
          }
      
      
      def _mimo_chunk_cache_key(chunk):
          """Stable identifier for a MiMo chunk (index + scene span) for partial-cache reuse."""
          return (
              f"{chunk['chunk_id']}|{chunk['scene_id']}|"
              f"{float(chunk['start']):.3f}-{float(chunk['end']):.3f}"
          )
      
      
      def _mimo_partial_provenance(video_path, scenes):
          return {
              "source_video_identity": file_identity(video_path),
              "chunks": [_mimo_chunk_cache_key(chunk) for chunk in _mimo_video_chunks(scenes)],
          }
      
      
      def _load_mimo_partial(partial_path, video_path=None, scenes=None):
          """Load the internal partial chunk cache, keyed by chunk identifier.
      
          Returns {} when missing/unreadable or when settings/source/chunk provenance differs,
          so a changed source video or scene plan cannot reuse paid chunks from another run.
          """
          if not partial_path.exists():
              return {}
          try:
              partial = json.loads(partial_path.read_text(encoding="utf-8"))
          except (json.JSONDecodeError, OSError):
              return {}
          if not isinstance(partial, dict):
              return {}
          if partial.get("settings") != mimo_video_settings():
              return {}
          if video_path is not None or scenes is not None:
              try:
                  if partial.get("provenance") != _mimo_partial_provenance(video_path, scenes):
                      return {}
              except (OSError, RuntimeError, TypeError, ValueError):
                  return {}
          done = partial.get("chunks")
          return done if isinstance(done, dict) else {}
      
      
      def _save_mimo_partial(partial_path, done, video_path=None, scenes=None):
          """Persist completed chunk results incrementally so paid chunks survive a mid-loop failure."""
          payload = {
              "settings": mimo_video_settings(),
              "chunks": done,
          }
          if video_path is not None or scenes is not None:
              payload["provenance"] = _mimo_partial_provenance(video_path, scenes)
          partial_path.write_text(json.dumps(payload, ensure_ascii=False, indent=2), encoding="utf-8")
      
      
      def _mimo_chunks_match(cached_chunks, expected_chunks):
          if not isinstance(cached_chunks, list) or len(cached_chunks) != len(expected_chunks):
              return False
          try:
              cached_keys = [_mimo_chunk_cache_key(chunk) for chunk in cached_chunks]
              expected_keys = [_mimo_chunk_cache_key(chunk) for chunk in expected_chunks]
          except (KeyError, TypeError, ValueError):
              return False
          return cached_keys == expected_keys
      
      
      def mimo_video_overview_cache_fresh(overview_path, video_path, scenes):
          """Return True only when the final MiMo overview matches current inputs/settings."""
          overview_path = Path(overview_path)
          if not overview_path.exists():
              return False
          try:
              overview = json.loads(overview_path.read_text(encoding="utf-8"))
          except (json.JSONDecodeError, OSError):
              return False
          if not isinstance(overview, dict) or overview.get("input") != "scene_chunks":
              return False
          if overview.get("settings") != mimo_video_settings():
              return False
          chunks = overview.get("chunks")
          if not isinstance(chunks, list) or not all(
              isinstance(chunk, dict) and _is_mimo_chunk_usable(chunk.get("content"))
              for chunk in chunks
          ):
              return False  # a moderation-rejected chunk is retried, never served from cache
          try:
              if overview.get("source_video_identity") != file_identity(video_path):
                  return False
              expected_chunks = _mimo_video_chunks(scenes)
          except (OSError, RuntimeError, TypeError, ValueError):
              return False
          return _mimo_chunks_match(chunks, expected_chunks)
      
      
      def _analyze_mimo_video_chunk(chunk_path, chunk):
          video_url = _video_data_url(chunk_path)
          if not video_url:
              raise RuntimeError(f"MiMo 视频分片 {chunk['chunk_id'] + 1} 超过 data URL 上限")
      
          content_parts = [
              {
                  "type": "video_url",
                  "video_url": {"url": video_url},
                  "fps": CONFIG["mimo_video_fps"],
                  "media_resolution": CONFIG["mimo_media_resolution"],
              },
              {"type": "text", "text": _mimo_chunk_prompt(chunk)},
          ]
          model = _mimo_video_model()
          payload = {
              "model": model,
              "messages": [{"role": "user", "content": content_parts}],
              "max_tokens": 1200,
          }
          resp = mimo_video_api_call(payload)
          try:
              msg = resp["choices"][0]["message"]
          except (KeyError, IndexError, TypeError) as exc:
              raise RuntimeError("MiMo 视频分片理解响应缺少 choices[0].message") from exc
          return {
              "chunk_id": chunk["chunk_id"],
              "scene_id": chunk["scene_id"],
              "start": chunk["start"],
              "end": chunk["end"],
              "model": resp.get("model", model),
              "content": msg.get("content", ""),
              "reasoning_content": msg.get("reasoning_content", ""),
              "usage": resp.get("usage", {}),
              "clip_path": f"mimo_video_chunks/{chunk_path.name}",
          }
      
      
      _MIMO_REJECTION_MARKERS = (
          "request was rejected", "considered high risk", "high risk",
          "content policy", "cannot process", "无法处理", "内容审核", "违规",
      )
      
      
      def _is_mimo_chunk_usable(content):
          """A chunk is usable only if MiMo returned real analysis (not empty / a moderation refusal)."""
          text = str(content or "").strip()
          if not text:
              return False
          low = text.lower()
          return not any(marker in low for marker in _MIMO_REJECTION_MARKERS)
      
      
      def analyze_video_overview(video_path, work_dir, scenes=None):
          """Use MiMo video understanding over local ffmpeg scene chunks."""
          if not CONFIG["mimo_video_overview"]:
              return None
          if not CONFIG["mimo_video_api_key"]:
              log("MiMo 视频概览已启用,但未设置 MIMO_VIDEO_API_KEY/MIMO_API_KEY,跳过")
              return None
      
          chunks = _mimo_video_chunks(scenes)
          chunks_dir = work_dir / "mimo_video_chunks"
          chunks_dir.mkdir(exist_ok=True)
      
          # 增量缓存:已完成的分片先落盘,避免一段失败时丢弃所有已付费分片
          partial_path = work_dir / "mimo_video_overview.partial.json"
          done = _load_mimo_partial(partial_path, video_path, scenes)
      
          log(f"MiMo 视频理解:按 ffmpeg scene 分片分析 {len(chunks)} 段...")
          chunk_results = []
          unusable_chunks = []
          for chunk in chunks:
              cache_key = _mimo_chunk_cache_key(chunk)
              cached = done.get(cache_key)
              if cached is not None:  # the partial only ever stores usable chunks
                  log(
                      f"  MiMo 分片 {chunk['chunk_id'] + 1}/{len(chunks)}: "
                      f"{chunk['start']:.1f}-{chunk['end']:.1f}s(命中增量缓存,跳过)"
                  )
                  chunk_results.append(cached)
                  continue
              chunk_path = chunks_dir / (
                  f"chunk_{chunk['chunk_id']:03d}_scene_{chunk['scene_id']}_"
                  f"{chunk['start']:.2f}-{chunk['end']:.2f}.mp4"
              )
              _extract_video_chunk(video_path, chunk, chunk_path)
              log(
                  f"  MiMo 分片 {chunk['chunk_id'] + 1}/{len(chunks)}: "
                  f"{chunk['start']:.1f}-{chunk['end']:.1f}s"
              )
              chunk_result = _analyze_mimo_video_chunk(chunk_path, chunk)
              if _is_mimo_chunk_usable(chunk_result["content"]):
                  chunk_results.append(chunk_result)
                  done[cache_key] = chunk_result
                  _save_mimo_partial(partial_path, done, video_path, scenes)
              else:
                  unusable_chunks.append(chunk)
                  log(f"  MiMo 分片 {chunk['chunk_id'] + 1}: 未返回有效内容,保留为待重试")
      
          if not chunk_results:
              log(f"MiMo 视频概览:{len(chunks)} 段均无有效内容(疑似被内容审核拦截),跳过概览")
              partial_path.unlink(missing_ok=True)
              return None
      
          if unusable_chunks:
              # Degrade gracefully instead of aborting the whole understanding: some chunks get
              # moderation-rejected (e.g. a burned-in watermark / violent frames). Write the overview
              # from the usable chunks; the scenes whose chunks were rejected simply fall back to the
              # frame-VLM description downstream. (Aborting here would make overview unsafe to enable
              # by default on any moderated source.)
              sample = ", ".join(str(chunk["chunk_id"] + 1) for chunk in unusable_chunks[:5])
              log(
                  f"MiMo 视频概览:{len(unusable_chunks)}/{len(chunks)} 段无有效内容(疑似内容审核),"
                  f"以可用分片降级产出,未覆盖场景回退到逐帧描述。样例分片: {sample}"
              )
      
          content = "\n\n".join(
              f"### 分片 {item['chunk_id'] + 1} "
              f"(scene {item['scene_id']}, {item['start']:.1f}-{item['end']:.1f}s)\n"
              f"{item['content'].strip()}"
              for item in chunk_results
          )
      
          overview = {
              "model": _mimo_video_model(),
              "content": content,
              "chunks": chunk_results,
              "chunk_count": len(chunk_results),
              "fps": CONFIG["mimo_video_fps"],
              "media_resolution": CONFIG["mimo_media_resolution"],
              "input": "scene_chunks",
              "partial": bool(unusable_chunks),
              "unusable_chunk_count": len(unusable_chunks),
              "source_video_identity": file_identity(video_path),
              "chunk_max_seconds": CONFIG["mimo_video_chunk_max_seconds"],
              "settings": mimo_video_settings(),
          }
          overview_path = work_dir / "mimo_video_overview.json"
          overview_path.write_text(json.dumps(overview, ensure_ascii=False, indent=2), encoding="utf-8")
          # 所有分片完成后清理增量缓存,保持 work_dir 仅有规范产物
          partial_path.unlink(missing_ok=True)
          log(f"MiMo 分片视频概览完成: {overview_path}")
          return overview
      
  • SKILL.md 4 KB
    ---
    name: video-understanding
    user-invocable: false
    description: >
     把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief。
     用于理解、索引或总结视频,也作为后续创作前的分析阶段。输入视频文件;输出 scenes.json、
     asr_result.json、vlm_analysis.json、silence_periods.json、timeline_fusion.json、agent_narration_brief.md。
     触发词:视频理解、视频分析、视频索引、video understanding、analyze video、看懂视频。
    ---
    
    ## 1. 定位
    
    本技能把源视频转成 Agent 与下游阶段可读取的理解索引。它的创作角色是**素材观察员 / 场记**,不是导演:
    
    - 先观察,再解释;事实与推断分开。
    - 除了“发生了什么”,还要让下游看见知识、权力、目标、关系或情绪在哪一刻变化。
    - 标出由谁的 POV 承载变化、哪个反应或表演不可替代,以及哪里存在完整台词/动作的自然剪辑边界。
    - 证据不足时保留不确定性,不制造戏剧结论。
    
    ## 2. 处理阶段
    
    1. **场景检测**:写 `scenes.json`,包含切点、时长和废片段过滤结果。
    2. **抽帧**:为视觉分析提取代表帧。
    3. **ASR**:通过 `mimo-v2.5-asr` 写粗分段对白 `asr_result.json`,并写
       `asr_timing_evidence.json` 说明可用性、有限时间精度与文本修正来源。
    4. **静音检测**:写 `silence_periods.json`,标注安静窗口与 `has_speech`。
    5. **VLM 观察**:写 `vlm_analysis.json`,包含场景描述、深层分析和 `frame_facts`。
    6. **时间线融合与创作 brief**:写 `timeline_fusion.json`、`asr_writing_chunks.json` 和 `agent_narration_brief.md`。
    
    各阶段只有在输出产物与 provenance sidecar 同时匹配当前视频及影响结果的设置时才会复用;`--force` 强制重算。
    
    ## 3. 环境要求
    
    ```bash
    # ffmpeg: brew install ffmpeg | apt install ffmpeg | choco install ffmpeg
    export MIMO_API_KEY=***
    ```
    
    ASR 使用 `mimo-v2.5-asr`;VLM 使用 `mimo-v2.5`。`--skip-asr` 可跳过对白转写,但完整理解仍需要 `MIMO_API_KEY` 运行 VLM。`--mimo-video-overview` 可开启按场景块的视频概览。
    
    若 `work_dir/background_research.json` 存在,本技能会把剧情梗概和角色名折入 VLM 上下文;`--context` 可补充一条简短提示。
    
    下面的 `scripts/...` 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。
    
    ## 4. 运行命令
    
    ```bash
    python3 scripts/understand.py <video> --work-dir <work_dir> \
      [--context "节目名/角色名"] [--scene-threshold 0.1] [--skip-asr] [--mimo-video-overview] [--force]
    ```
    
    ## 5. 输出契约
    
    | 文件 | 内容 |
    |------|------|
    | `scenes.json` | 场景切点、起止时间与时长 |
    | `asr_result.json` | `[{start, end, text}]` 时间戳对白 |
    | `asr_timing_evidence.json` | ASR 可用性状态、粗窗口精度、glossary 前后文本,以及它所描述的源视频/音频/结果文件(路径存在性 + size/mtime) |
    | `vlm_analysis.json` | 逐场景描述、深层分析与 `frame_facts` |
    | `silence_periods.json` | `[{start, end, duration, has_speech}]` 安静窗口 |
    | `timeline_fusion.json` | VLM、ASR 与静音信息的统一时间线 |
    | `asr_writing_chunks.json` | 按句界和场景切分的 ASR 写作块 |
    | `agent_narration_brief.md` | Agent 首先阅读的创作简报 |
    
    后续写作阶段根据创作简报与索引制定方案并写 `narration.json`。
    
    ## 6. 参考资料
    
    - 背景调研:`references/research-guide.md`,产出 `background_research.json`。
    - JSON 结构:`references/data-schema.md`。
    
    ## 7. 能力边界
    
    - 不写解说词,也不做解说评分;只负责生成理解索引与创作简报。
    - 不编造信号无法支持的剧情;当 ASR / VLM 过薄时输出素材警告。
    - MiMo ASR 的 `start/end` 是固定分片形成的**粗窗口**,不是词级对齐;空文本只表示原因未知,
      不能当作已证实静音。`asr_timing_evidence.json` 的状态字段见 `references/data-schema.md`。
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related