Claude Skill

video-cut

把长视频按 Agent 选择的原片区间剪成短片。作为两阶段创作流程中的剪辑环节,读取 clip_plan.json 与源视频, 输出 edited_source.mp4;随后 Agent 按输出时间线写 narration.json。支持单视频与多视频(sources manifest)拼剪, 本工具不读取、不映射旁白。 触发词:视频剪辑、剪辑式解说、video cut、clip plan、拼剪。

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download zenstory-ai-oh-story-dsh-packages_knowledge_video-recap_skills_video-cut-d734089.zip · 39 KB
Part of zenstory-ai/oh-story-dsh — 31 skills

Install

skills CLI npx skills add https://github.com/zenstory-ai/oh-story-dsh/tree/main/packages/knowledge/video-recap/skills/video-cut
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install zenstory-ai-oh-story-dsh@llmmart
Git git clone https://github.com/zenstory-ai/oh-story-dsh.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole zenstory-ai/oh-story-dsh collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

1. 定位

本技能只执行 Agent 已经做出的剪辑决定:

  1. 校验并补全 clip_plan.json,写出带 clip_id、原片/输出时间与时长的 clip_plan_validated.json。
  2. 先避开原片硬切附近的闪帧风险,最后把边界吸附到可靠句末/自然停顿;声音完整性拥有最终优先级。
  3. 拼接选定区间,输出 edited_source.mp4。
  4. 到此停止,由 Agent 按真实输出时间线写 narration.json;本工具不读取旁白,也不做原片→输出映射。

相同输入会得到相同输出。edited_source.mp4.meta.json 记录标准化 clips、渲染设置和每个源文件的 size/mtime_ns;三者与当前一致且 edited_source.mp4 存在非空才复用,任一不同即重渲染。只有 sidecar 而没有媒体文件不复用。

2. 输入契约

work_dir/clip_plan.json 可以是数组,也可以是 {"clips": [...]}:

{"start": 12.0, "end": 28.5, "reason": "b02 | turn | power: A→B | POV=女主 | 保留反应 | 入点=问题落下 | 出点=沉默结束"}
  • start / end 是原片秒数;也接受 source_start / source_end 或 in / out。
  • 顶层可选 target_duration,例如 "10m"。
  • 多视频项目的每个片段还必须填写 source_id。
  • speech_boundary_anchors.json 与 ASR 时间段由理解阶段提供;Agent 先写大致区间,工具会尝试吸附并把仍在讲话区间内的入/出点作为 blocker 返回。

3. 剪辑意图契约

工具不会替 Agent 做创作选择。写片段前先完成本节的剪辑意图检查,并让每个区间映射到 recap_story_plan.json 的一个 beat。

使用现有自由文本 reason 保存简洁决定:

beat_id | function | change | POV | preferred moment | 入点 reason | 出点 reason

不要因为“事件重要”就保留整段;要保留最能让 change 成立的具体表演、反应、动作或揭示。理解与情绪允许时晚进早出,同时保证台词、动作和技术边界完整。

对不能删去的问答、反应或动作兑现,先核源证据,再在同一 clip_plan.json 登记精确区间:

{
  "clips": [{"start": 12, "end": 18}],
  "required_evidence": {
    "nodes": [
      {"id": "refusal", "source": "/media/episode.mp4", "start": 12.25, "end": 14.5, "track": "audio", "content": "对方拒绝请求"},
      {"id": "response", "source": "/media/episode.mp4", "start": 15, "end": 17.5, "track": "video", "content": "听到拒绝后的反应与决定"}
    ],
    "before": [["refusal", "response"]]
  }
}

source 使用实际源文件绝对路径,start/end 是原片秒;多源可另填 source_id 消歧。只登记确实需要保留的具体时刻,不将整个 beat 默认锁死。before 只登记本片必需的先后关系;无需约束顺序时写 before: []。

工具在全部画面/句界吸附后检查每个必保时刻至少有一处完整连续保留、来源和先后;音频节点还检查源音轨是否存在。每次结果出现(包括局部片段)都需满足其声明的前提,不能用后面的完整段替开头缺前提的片段过关。结果写入 clip_plan_validated.json.qc.required_evidence;缺段、错序或无效声明会在预检、缓存复用和渲染前阻断,时长放宽选项不会跳过。该结果验证选段保留,实际语义与最终混音仍按审片步骤核对。

下面的 scripts/... 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。

4. 运行命令

python3 scripts/cut.py <video> --work-dir <work_dir> \
  [--target-duration 10m] [--clip-padding 0] [--allow-overlap]

5. 输出契约

  • clip_plan_validated.json:标准化片段,包含 clip_id、source_start/end、output_start/end 与 duration。
  • edited_source.mp4:按计划拼接后的短视频。
  • shot_review.json:仅 --review-shots 开启后生成的实际视频短镜/密集切镜候选;不会更改计划。

下游把 edited_source.mp4 当作视频,把 Agent 按输出时间写的 narration.json 当作旁白。

6. 边界与时间线规则

  • clip_plan.json 使用原片时间;narration.json 直接使用剪后输出时间,不存在原片 → 输出的旁白映射。
  • 默认禁止重叠或重复原片区间;--allow-overlap 开启后才允许。
  • 片段起点只能位于源头、可靠句末/静音窗,或与上一片段构成无损同源连续连接;片段终点同理。ASR 判定仍在讲话且无法吸附时写入 unsafe_clip_sentence_boundary 并阻断。
  • SCENE_CUT_SNAP 默认开启:先按画面把 source start 向后、source end 向前吸附到附近硬切,随后句末吸附再做最终修正,避免视觉修正重新制造半句原声。默认范围为 SCENE_CUT_SNAP_MARGIN=0.5 秒,检测阈值为 SCENE_CUT_DETECT_THRESHOLD=0.4。
  • scene-change score 只提供接点候选,不证明接点自然。先检查短时间窗内是否出现密集候选,再区分来源:原片自带的无关短镜头整段删除;相关但短到像闪帧的镜头通过扩展 IN/OUT 保留完整动作、反应或台词,不用定格/慢放伪造时长;由本次拼接制造的切点则优先移动边界、恢复同源连续运动、合并相邻片段或改用更自然的连接,尽量消除。成片后仍要逐个播放接点前后约 0.5–1 秒;白闪或曝光叠化再结合逐帧亮度定位,不能为了通过视觉检测切断完整台词,也不能用转场遮掩坏接点。
  • 修短残镜时不得仅为压低 scene 分数而对接点附近施加与所属镜头不连续的极端放大或位移;取景复核与修复验证流程见 references/shot-review.md。
  • 连续同源片段的无损连接不做句中双侧音频淡出;非连续片段仍在安全停顿内做防爆音淡入淡出。

需要检查短时间频繁切镜时,先用 ffmpeg scene filter 召回候选时间:

ffmpeg -i input.mp4 -vf "select='gt(scene,0.35)',showinfo" -an -f null -

0.35 是起始阈值,不是质量判据;大幅运动、闪白和叠化都可能误报。把候选映射回原片 shot 与本次拼接边界后,按上面的来源分类处理,并以正常速度播放决定是否保留。

需要精确到实际帧、检查长区间内部残镜并记录所用计划路径时,使用 scripts/shot_review.py 或 cut.py --review-shots(有黑边或包装时加 --roi / --shot-roi); 详见 references/shot-review.md。

7. 能力边界

  • 不做语义理解,不写旁白,不替 Agent 选择片段;scene filter 只承担技术边界候选检测。
  • 只做生成 edited_source.mp4 所需的剪切、拼接与一次中间编码,不承担字幕包装或最终交付压缩。
Files (oh-story-dsh)
  • references
    • shot-review.md 5.4 KB
      # 成片内部短镜头与密集切镜召回
      
      `shot_review.py` 是只读检查器,不改剪点、不触发吸附、不重编码视频,也不把短镜头自动删掉。
      EDL 的一个长区间内仍可能藏着 2、12、14 帧反打;因此检查对象是**本轮实际渲染的视频**,不是只看计划里每段有多长。
      
      ```bash
      # 默认 --threshold 0.35:
      python3 scripts/shot_review.py edited_source.mp4 --output review/shot_review.json
      # 只有当前 video-cut 的输出与 sidecar 齐全,才额外关联当前剪辑计划:
      python3 scripts/shot_review.py edited_source.mp4 --plan clip_plan_validated.json \
        --output review/shot_review_planned.json
      # 怀疑漏检时另跑一轮更低阈值做对照;0.08 只是一个更高召回的取值示例,不是规定的第二遍:
      python3 scripts/shot_review.py edited_source.mp4 \
        --threshold 0.08 --output review/shot_review_sensitive.json
      # 有黑边/包装时,按实测主画窗填入像素坐标,保持阈值不变做对照:
      python3 scripts/shot_review.py packaged.mp4 --roi "$X" "$Y" "$WIDTH" "$HEIGHT" \
        --output review/shot_review_picture.json
      # 正常 cut 后显式开启;缓存命中也查实际文件;normalize-only 不扫描:
      python3 scripts/cut.py source.mp4 --work-dir work --review-shots
      # 同一画窗/阈值参数也可传给 cut.py --review-shots --shot-roi X Y WIDTH HEIGHT --shot-scene-threshold T
      ```
      
      `--roi` 仅裁检测输入,不改视频文件或剪辑计划。坐标采用 FFmpeg 自动转正后的原生像素,
      不是播放器按 SAR 拉伸后的显示尺寸;矩形需完整位于画布内。报告的 `scene_roi` 保存实际
      检测范围,未指定时为 `null`(全画布)。先排除黑边及标题/花字区,再调召回阈值;
      同时保留全画布或更大范围的对照,不能通过缩小范围隐藏问题。画窗随镜头变化时,
      一个固定 ROI 只能覆盖该区域,不能据此宣称已检查所有画面。
      
      ## 精度和边界
      
      - `ffprobe` 完整解码收集真实帧 PTS;scene filter 使用整数 PTS 与过滤器 timebase,不用 seek 后的浮点 `pts_time` 反推帧号。
      - 帧号从 0 开始,区间半开 `[start_frame,end_frame)`;首镜、末镜也检查。VFR 同时记录真实 rational PTS 时长,不能用平均 fps 乘秒数。
      - 默认候选规则:镜头不超过 **1 秒**,或 **2 秒内至少 4 个切点**。帧数上限缺省由实测帧钟换算(`round(fps × max_short_seconds)`),不写死某个帧率;需要固定帧数时用 `--max-short-frames` 显式覆盖。参数只控制召回,不是统一剪辑标准。
      - `0.35` 是默认 scene 起始阈值。加黑边、包装占比大、低反差的画面会系统性少报。**某个阈值下零候选不是没有闪帧的证据。** 用一段已知有坏短镜的窗口校准阈值,保留每一轮的报告,不静默换阈值只报告“通过”。降低阈值也可能增加曝光、运动和动效误报。
      - 未提供计划时来源为 `UNKNOWN`。提供计划时先核对计划 clips、渲染设置与 `edited_source.mp4.meta.json` 一致、源文件 `size`/`mtime_ns` 未变、输出存在非空;失配直接失败,不猜“最新版本”。报告 `plan_binding.path` 只记录所用计划路径。
      - 靠近量化后 EDL 拼接点一帧以内仅标 `EDIT_JOIN_CANDIDATE`。内部候选可以标所在已绑定 clip、估计原片秒数,但仍为 `UNKNOWN`;没有独立源片证据,不能认定是原片自带切镜。
      - 未知时码、末帧时长无法覆盖到流末尾、解码失败均阻断扫描。报告先置 `SCANNING`,失败写 `SCAN_FAILED`,不留下旧成功报告冒充本轮结果。仅新路径或可识别为本工具 schema 的旧报告可写;报告路径不能覆盖计划与元数据双方声明的视频/源文件,即使二者已经过期失配。
      
      ## 如何处理候选
      
      对每个候选看前后 0.5–1 秒,核对角色表演、完整动作、对白与原片切镜;未实际播放就保留 `NEEDS_DYNAMIC_REVIEW`。
      先对照源画面和取景参数:同一原生镜头中途换 crop/画窗,或曝光闪光,都可能产生切点分数;两个候选间的帧数不一定是一段真实短镜头。
      修复后用同样参数重跑对照:候选消失才算真的修掉;仍被保留的短镜逐个判断,不因“短”就判错。
      判断无关残镜后,在作者计划里删除整个无关镜头或恢复同源连续画面;相关反应可回原片扩完整,不能用定格补长、转场掩盖或只修导出文件。
      
      ## 修短残镜时的取景复核
      
      修短残镜时,同时核对原片镜头变化与输出 `crop/window` 变化。不得仅为压低 scene 分数,对接点前后少量帧施加与所属镜头不连续的极端放大或位移,造成关键表情、动作被裁或清晰度明显下降。先在原片自然切点上分别确定相邻镜头各自连续、清晰、主体完整的取景;不同镜头不要求景别或裁幅一致,但尺度变化、主体位置及视线/运动方向须在正常速度下复核。保存修复前后相同输出帧号的邻帧对照,再核声音、字幕、总帧数与播放速度。候选数量或 scene 分数下降,只说明检测结果改变,不证明接点已经修复;没有正常速度观看能力时保留未检查状态。
      
      `NO_CANDIDATES` 仅表示该阈值下没有短镜/密集切镜候选;不代表故事、听感、动态观感或发布授权通过。报告始终保留 `normal_speed_review=NOT_CHECKED`,不会写入已有 `clip_plan_validated.json['qc']`。
      
  • scripts
    • cut.py 1.1 KB
      #!/usr/bin/env python3
      """Public API and CLI entrypoint for the self-contained video-cut skill."""
      
      from cut_cli import main
      from cut_contract import (
          edited_source_render_cache_payload,
          normalize_clip_plan,
          normalize_multi_source_clip_plan,
          parse_duration_seconds,
          should_reuse_edited_source,
      )
      from cut_render import build_edited_source_video
      from media_geometry import VideoGeometry
      from narration_mapping import update_cut_qc
      from sentence_boundaries import (
          enforce_clip_sentence_boundaries,
          snap_clip_ends_to_lines,
          snap_clip_starts_to_lines,
          snap_clips_off_shot_changes,
          snap_multi_source_clips,
      )
      
      __all__ = [
          "VideoGeometry",
          "build_edited_source_video",
          "edited_source_render_cache_payload",
          "enforce_clip_sentence_boundaries",
          "main",
          "normalize_clip_plan",
          "normalize_multi_source_clip_plan",
          "parse_duration_seconds",
          "should_reuse_edited_source",
          "snap_clip_ends_to_lines",
          "snap_clip_starts_to_lines",
          "snap_clips_off_shot_changes",
          "snap_multi_source_clips",
          "update_cut_qc",
      ]
      
      if __name__ == "__main__":
          main()
      
    • cut_cli.py 12.1 KB
      """Command-line orchestration for the video-cut skill."""
      
      import json
      import math
      
      
      from pathlib import Path
      
      from lib import CONFIG, get_video_duration, log
      import shot_review
      
      from cut_contract import (
          _write_edited_source_meta,
          load_clip_plan,
          normalize_clip_plan,
          normalize_multi_source_clip_plan,
          parse_duration_seconds,
          should_reuse_edited_source,
      )
      from cut_render import (
          build_edited_source_video,
          update_delivery_qc,
          write_cut_delivery_qc,
      )
      from media_geometry import _has_audio_stream, _select_output_geometry
      from narrative_selection import check_required_evidence
      from narration_mapping import update_cut_qc
      from sentence_boundaries import (
          _combine_boundary_windows,
          _load_sentence_boundary_windows,
          _load_silence_for_source,
          _load_source_speech_spans,
          enforce_clip_sentence_boundaries,
          snap_clip_ends_to_lines,
          snap_clip_starts_to_lines,
          snap_clips_off_shot_changes,
          snap_multi_source_clips,
      )
      
      
      def main():
          import argparse
      
          parser = argparse.ArgumentParser(
              description="video-cut: build edited_source.mp4 from an agent clip plan; narration is authored afterwards on the output timeline."
          )
          parser.add_argument("video", help="source video path")
          parser.add_argument(
              "--work-dir",
              required=True,
              help="dir holding clip_plan.json",
          )
          parser.add_argument(
              "--clip-plan",
              default=None,
              help="clip plan json (default: <work-dir>/clip_plan.json)",
          )
          parser.add_argument(
              "--sources-manifest",
              default=None,
              help="multi-source manifest json mapping source_id values to source media",
          )
          parser.add_argument(
              "--target-duration",
              default=None,
              help="target output duration, e.g. 10m / 600 / 00:10:00",
          )
          parser.add_argument(
              "--clip-padding",
              type=float,
              default=None,
              help="seconds to pad each clip on both ends (default: CLIP_PADDING env, else 0)",
          )
          parser.add_argument(
              "--allow-overlap",
              action="store_true",
              help="allow overlapping/duplicate source ranges",
          )
          parser.add_argument(
              "--normalize-only",
              action="store_true",
              help="only normalize the clip plan -> clip_plan_validated.json (no render); "
              "lets validate lint the SAME padded/pruned plan the render uses",
          )
          parser.add_argument(
              "--review-shots", action="store_true",
              help="scan actual rendered/reused video for internal short-shot and dense-cut candidates; never repair",
          )
          parser.add_argument(
              "--shot-scene-threshold", type=float, default=None,
              help="explicit scene recall threshold for --review-shots (default 0.35; not an acceptance criterion)",
          )
          parser.add_argument(
              "--shot-roi", nargs=4, type=int, metavar=("X", "Y", "WIDTH", "HEIGHT"),
              help="scan only this pixel rectangle with --review-shots; never crop the rendered video",
          )
          parser.add_argument(
              "--allow-duration-drift",
              action="store_true",
              help="do not block when validated clip duration is far from --target-duration",
          )
          args = parser.parse_args()
          if args.shot_scene_threshold is not None and (
              not args.review_shots or not math.isfinite(args.shot_scene_threshold)
              or not 0 <= args.shot_scene_threshold <= 1
          ):
              parser.error("--shot-scene-threshold requires --review-shots and a finite value in [0,1]")
          if args.shot_roi is not None and (
              not args.review_shots or min(args.shot_roi[:2]) < 0 or min(args.shot_roi[2:]) <= 0
          ):
              parser.error("--shot-roi requires --review-shots, nonnegative X/Y and positive WIDTH/HEIGHT")
      
          # CLIP_PADDING is declared in every skill's CONFIG, but video-cut is the only place that
          # implements padding — and it used to read the CLI flag alone, so setting the env var did
          # nothing at all while `clip_padding_source: "env"` reported otherwise. CLI still wins.
          clip_padding = (
              args.clip_padding if args.clip_padding is not None else CONFIG["clip_padding"]
          )
      
          work_dir = Path(args.work_dir)
          work_dir.mkdir(parents=True, exist_ok=True)
          clip_plan_path = (
              Path(args.clip_plan) if args.clip_plan else work_dir / "clip_plan.json"
          )
          raw_plan = load_clip_plan(clip_plan_path)
      
          target_seconds = (
              parse_duration_seconds(args.target_duration) if args.target_duration else None
          )
          sources_manifest = (
              json.loads(Path(args.sources_manifest).read_text(encoding="utf-8"))
              if args.sources_manifest
              else None
          )
          if sources_manifest is not None:
              validated_plan = normalize_multi_source_clip_plan(
                  raw_plan,
                  sources_manifest,
                  target_duration=target_seconds,
                  clip_padding=clip_padding,
                  allow_overlap=args.allow_overlap,
              )
              video_duration = None
          else:
              video_duration = get_video_duration(args.video)
              validated_plan = normalize_clip_plan(
                  raw_plan,
                  video_duration,
                  target_duration=target_seconds,
                  clip_padding=clip_padding,
                  allow_overlap=args.allow_overlap,
              )
      
          # Keep boundaries off the original footage's hard cuts (avoids 闪烁 at the edit point).
          # This visual-only pass runs FIRST. The sentence/quiet pass below is the final authority:
          # a prettier edit point must never move the final boundary back inside a spoken sentence.
          if sources_manifest is None and CONFIG["scene_cut_snap"]:
              validated_plan = snap_clips_off_shot_changes(
                  validated_plan,
                  args.video,
                  margin=CONFIG["scene_cut_snap_margin"],
                  threshold=CONFIG["scene_cut_detect_threshold"],
              )
      
          if sources_manifest is None:
              safe_boundaries = _combine_boundary_windows(
                  _load_silence_for_source(work_dir, None),
                  _load_sentence_boundary_windows(work_dir),
              )
              if CONFIG["snap_clip_line_end"]:
                  validated_plan = snap_clip_starts_to_lines(
                      validated_plan,
                      safe_boundaries,
                      video_duration,
                      CONFIG["clip_start_snap_max_prepend"],
                      max_trim=CONFIG["clip_start_snap_max_trim"],
                  )
                  validated_plan = snap_clip_ends_to_lines(
                      validated_plan,
                      safe_boundaries,
                      video_duration,
                      CONFIG["clip_snap_max_extend"],
                  )
              validated_plan = enforce_clip_sentence_boundaries(
                  validated_plan,
                  safe_boundaries,
                  _load_source_speech_spans(work_dir),
                  video_duration,
              )
      
          # Multi-source: snap each clip against ITS OWN source's pauses/shot-changes (single-source
          # snaps above can't, since silence_periods.json and args.video are per-project, not per-source).
          if sources_manifest is not None:
              validated_plan = snap_multi_source_clips(
                  validated_plan,
                  validated_plan["sources"],
                  work_dir,
                  line_max_extend=CONFIG["clip_snap_max_extend"],
                  scene_margin=CONFIG["scene_cut_snap_margin"],
                  scene_threshold=CONFIG["scene_cut_detect_threshold"],
                  do_line_snap=CONFIG["snap_clip_line_end"],
                  do_scene_snap=CONFIG["scene_cut_snap"],
                  start_max_prepend=CONFIG["clip_start_snap_max_prepend"],
                  start_max_trim=CONFIG["clip_start_snap_max_trim"],
              )
      
          validated_plan.setdefault("qc", {})["join_fade_ms"] = round(
              CONFIG["clip_join_audio_fade_ms"], 3
          )
          # Single-source clips carry no source_path; the CLI video is the only input.
          source_paths = list(
              dict.fromkeys(
                  clip["source_path"] for clip in validated_plan["clips"] if "source_path" in clip
              )
          ) or [str(args.video)]
          _, _, _, geometry_qc = _select_output_geometry(source_paths, validated_plan["clips"])
          validated_plan["qc"]["output_geometry"] = geometry_qc
          validated_plan["qc"]["output_geometry_reason"] = geometry_qc["reason"]
          update_cut_qc(
              validated_plan,
              allow_duration_drift=bool(args.allow_duration_drift),
              duration_drift_allowed_by="--allow-duration-drift" if args.allow_duration_drift else None,
          )
          if isinstance(raw_plan, dict) and 'required_evidence' in raw_plan:
              # Re-evaluate the final snapped ranges even when the media cache can be reused.
              # A prior rendered receipt must not survive a failed revision preflight.
              (work_dir / 'cut_delivery_qc.json').unlink(missing_ok=True)
              contract = raw_plan['required_evidence']
              plan_sources = {str(Path(path).resolve()): path for path in source_paths}
              source_audio = {}
      
              def has_source_audio(source):
                  # Probed lazily, once per source, only for validated audio nodes.
                  if source not in source_audio:
                      source_audio[source] = source in plan_sources and _has_audio_stream(plan_sources[source])
                  return source_audio[source]
      
              report = check_required_evidence(contract, validated_plan, input_video=args.video,
                                               source_audio=has_source_audio)
              validated_plan['qc']['required_evidence'] = {**report, 'contract': contract}
              if report['selection_status'] == 'BLOCK':
                  validated_plan['qc'].setdefault('blocking', []).extend(report['findings'])
          update_delivery_qc(
              validated_plan,
              source_paths=source_paths,
              output_path=work_dir / "edited_source.mp4",
          )
          (work_dir / "clip_plan_validated.json").write_text(
              json.dumps(validated_plan, ensure_ascii=False, indent=2), encoding="utf-8"
          )
          if validated_plan["qc"].get("blocking"):
              raise SystemExit(
                  "clip_plan QC blocking: fix required source evidence, unsafe sentence boundaries or target-duration drift. "
                  "Only duration drift can be explicitly accepted with --allow-duration-drift; "
                  "sentence truncation is never allowed. See clip_plan_validated.json['qc']."
              )
          if args.normalize_only:
              # normalize-only produces planned delivery facts in clip_plan_validated.json, but no
              # rendered/reused media exists in this run, so remove any stale final delivery artifact.
              (work_dir / "cut_delivery_qc.json").unlink(missing_ok=True)
              print(
                  json.dumps(
                      {
                          "status": "normalized",
                          "clips": len(validated_plan["clips"]),
                          "total_duration": validated_plan["total_duration"],
                      },
                      ensure_ascii=False,
                  )
              )
              return
      
          edited_source_path = work_dir / "edited_source.mp4"
          if should_reuse_edited_source(edited_source_path, validated_plan, args.video):
              log(f"复用剪辑源视频: {edited_source_path}")
              update_delivery_qc(
                  validated_plan,
                  source_paths=source_paths,
                  output_path=edited_source_path,
                  rendered=True,
              )
              write_cut_delivery_qc(work_dir, validated_plan)
              _write_edited_source_meta(edited_source_path, validated_plan, args.video)
              (work_dir / "clip_plan_validated.json").write_text(
                  json.dumps(validated_plan, ensure_ascii=False, indent=2), encoding="utf-8"
              )
          else:
              build_edited_source_video(
                  args.video, validated_plan, work_dir, edited_source_path
              )
              (work_dir / "clip_plan_validated.json").write_text(
                  json.dumps(validated_plan, ensure_ascii=False, indent=2), encoding="utf-8"
              )
      
          if args.review_shots:
              review_options = {"plan_path": work_dir / "clip_plan_validated.json"}
              if args.shot_scene_threshold is not None:
                  review_options["threshold"] = args.shot_scene_threshold
              if args.shot_roi is not None:
                  review_options["roi"] = args.shot_roi
              shot_review.write_scan(
                  edited_source_path, work_dir / "shot_review.json",
                  **review_options,
              )
      
          log(
              f"剪辑模式: {len(validated_plan['clips'])} 个片段 → {validated_plan['total_duration']:.1f}s"
          )
      
    • cut_contract.py 16 KB
      """Normalize cut plans and maintain the render-cache sidecar."""
      
      import json
      import re
      
      from pathlib import Path
      
      from lib import CONFIG, file_identity, get_video_duration, log
      
      
      def parse_duration_seconds(value):
          """Parse seconds, 10m/1h forms, or HH:MM:SS into seconds."""
          if value is None or value == "":
              return None
          if isinstance(value, (int, float)):
              seconds = float(value)
              if seconds <= 0:
                  raise ValueError("duration must be positive")
              return seconds
      
          text = str(value).strip().lower()
          if not text:
              return None
      
          if ":" in text:
              parts = text.split(":")
              if len(parts) not in (2, 3):
                  raise ValueError(f"invalid duration: {value}")
              try:
                  nums = [float(p) for p in parts]
              except ValueError as exc:
                  raise ValueError(f"invalid duration: {value}") from exc
              if any(n < 0 for n in nums):
                  raise ValueError("duration must be positive")
              if nums[-1] >= 60 or (len(nums) == 3 and nums[-2] >= 60):
                  raise ValueError(f"invalid duration: {value}")
              if len(nums) == 2:
                  seconds = nums[0] * 60 + nums[1]
              else:
                  seconds = nums[0] * 3600 + nums[1] * 60 + nums[2]
              if seconds <= 0:
                  raise ValueError("duration must be positive")
              return seconds
      
          # One or more <number><unit> tokens: "600", "10m", "500ms", "2m30s", "1h5m30s".
          # A bare number is read as seconds; units may be combined (compound durations).
          factors = {"ms": 0.001, "s": 1, "m": 60, "h": 3600}
          sign = 1.0
          body = text
          if body[:1] in "+-":
              sign = -1.0 if body[0] == "-" else 1.0
              body = body[1:]
          token_re = re.compile(r"([0-9]+(?:\.[0-9]+)?)(ms|s|m|h)?")
          pos = 0
          seconds = 0.0
          matched = False
          for m in token_re.finditer(body):
              if m.start() != pos:
                  break
              pos = m.end()
              matched = True
              seconds += float(m.group(1)) * factors[m.group(2) or "s"]
          if not matched or pos != len(body):
              raise ValueError(f"invalid duration: {value}")
          seconds *= sign
          if seconds <= 0:
              raise ValueError("duration must be positive")
          return seconds
      
      
      def _overlaps_authored_range(ranges, start, end):
          """Whether [start,end) collides with an already-accepted clip, as AUTHORED.
      
          Overlap is judged on the agent's own in/out points, never on the padded ones.
          `clip_padding` deliberately widens every clip by the same amount on both ends, so
          judging padded ranges makes any two back-to-back clips (…, 10) and (10, …) look like
          duplicate footage and hard-fails a perfectly ordinary plan. Padding is an output
          nicety; only what the agent actually asked for defines duplication.
          """
          return any(start < other_end and end > other_start for other_start, other_end in ranges)
      
      
      def _clip_value(raw, *names):
          for name in names:
              if name in raw:
                  return raw[name]
          return None
      
      
      def load_clip_plan(path):
          """Load `clip_plan.json`, accepting either a list or {"clips": [...]} object."""
          return json.loads(Path(path).read_text(encoding="utf-8"))
      
      
      def _edited_source_meta_path(output_path):
          return Path(str(output_path) + ".meta.json")
      
      
      def _load_edited_source_meta(output_path):
          """None when never rendered; a sidecar this skill wrote but cannot parse raises."""
          meta_path = _edited_source_meta_path(output_path)
          if not meta_path.exists():
              return None
          return json.loads(meta_path.read_text(encoding="utf-8"))
      
      
      def _source_identities_for_plan(validated_plan, input_video=None):
          """{path: {size, mtime_ns}} for every media file that can affect edited_source.mp4."""
          # Single-source plans carry no per-clip source_path; the CLI video is the only input.
          paths = {clip["source_path"] for clip in validated_plan["clips"] if "source_path" in clip}
          if not paths and input_video is not None:
              paths.add(str(input_video))
          return {str(Path(path)): file_identity(path) for path in sorted(paths)}
      
      
      def edited_source_render_cache_payload():
          """Render-affecting settings that invalidate edited_source.mp4 cache reuse.
      
          Keep this payload limited to inputs that can change rendered media bytes.
          Observational QC produced after validation/render is intentionally excluded.
          """
          return {"clip_join_audio_fade_ms": round(CONFIG["clip_join_audio_fade_ms"], 3)}
      
      
      def _write_edited_source_meta(output_path, validated_plan, input_video=None):
          meta = {
              "schema_version": 3,
              "plan": validated_plan["clips"],
              "render_cache": edited_source_render_cache_payload(),
              "sources": _source_identities_for_plan(validated_plan, input_video),
              "total_duration": validated_plan["total_duration"],
              "clip_count": len(validated_plan["clips"]),
          }
          _edited_source_meta_path(output_path).write_text(
              json.dumps(meta, ensure_ascii=False, indent=2), encoding="utf-8"
          )
      
      
      def should_reuse_edited_source(output_path, validated_plan, input_video=None):
          """True only when a non-empty edited_source.mp4 matches the plan, settings and sources."""
          output_path = Path(output_path)
          if not output_path.exists() or output_path.stat().st_size == 0:
              return False
          meta = _load_edited_source_meta(output_path)
          if not isinstance(meta, dict) or meta.get("schema_version") != 3:
              return False  # never rendered, or a sidecar from an older schema: re-render
          return (
              meta["plan"] == validated_plan["clips"]
              and meta["render_cache"] == edited_source_render_cache_payload()
              and meta["sources"] == _source_identities_for_plan(validated_plan, input_video)
          )
      
      
      def _manifest_source_entries(sources_manifest):
          """Return source rows from common multi-source manifest shapes."""
          if isinstance(sources_manifest, dict):
              if isinstance(sources_manifest.get("sources"), list):
                  return sources_manifest["sources"]
              rows = []
              for source_id, value in sources_manifest.items():
                  if source_id in {"schema_version", "version"}:
                      continue
                  if isinstance(value, dict):
                      row = dict(value)
                      row.setdefault("source_id", source_id)
                      rows.append(row)
              if rows:
                  return rows
          elif isinstance(sources_manifest, list):
              return sources_manifest
          raise ValueError(
              "sources manifest must be a list, a {sources:[...]} object, or a source_id map"
          )
      
      
      def normalize_sources_manifest(sources_manifest):
          """Normalize source manifest rows to {source_id: {source_path, duration}}."""
          sources = {}
          for idx, raw in enumerate(_manifest_source_entries(sources_manifest)):
              if not isinstance(raw, dict):
                  raise ValueError(f"source #{idx + 1} must be an object")
              source_id = raw.get("source_id", raw.get("id", raw.get("name")))
              if source_id in (None, ""):
                  raise ValueError(f"source #{idx + 1} is missing source_id")
              source_id = str(source_id)
              source_path = raw.get(
                  "source_path",
                  raw.get("path", raw.get("video_path", raw.get("video", raw.get("file")))),
              )
              if not source_path:
                  raise ValueError(f"source {source_id} is missing source_path/path")
              duration = raw.get(
                  "duration", raw.get("duration_seconds", raw.get("source_duration"))
              )
              if duration in (None, ""):
                  duration = get_video_duration(source_path)
              try:
                  duration = float(duration)
              except (TypeError, ValueError) as exc:
                  raise ValueError(f"source {source_id} has invalid duration") from exc
              if duration <= 0:
                  raise ValueError(f"source {source_id} has invalid duration")
              sources[source_id] = {
                  "source_id": source_id,
                  "source_path": str(source_path),
                  "duration": duration,
              }
              if raw.get("source_work_dir") not in (None, ""):
                  sources[source_id]["source_work_dir"] = str(raw["source_work_dir"])
          if not sources:
              raise ValueError("sources manifest has no sources")
          return sources
      
      
      def normalize_multi_source_clip_plan(
          raw_plan,
          sources_manifest,
          target_duration=None,
          clip_padding=0.0,
          min_clip_duration=0.3,
          allow_overlap=False,
      ):
          """Validate a multi-source clip plan and map source_id clips to source paths/durations.
      
          Clip order follows the raw plan; overlap validation is isolated per source_id.
          """
          sources = normalize_sources_manifest(sources_manifest)
          if isinstance(raw_plan, dict):
              raw_clips = raw_plan.get("clips", [])
              plan_target = raw_plan.get("target_duration") or raw_plan.get(
                  "target_duration_seconds"
              )
              if target_duration is None and plan_target not in (None, ""):
                  target_duration = parse_duration_seconds(plan_target)
          elif isinstance(raw_plan, list):
              raw_clips = raw_plan
          else:
              raise ValueError(
                  "clip_plan.json must be a JSON array or an object with a clips array"
              )
      
          if not isinstance(raw_clips, list):
              raise ValueError("clip_plan.json field `clips` must be an array")
      
          padding = max(0.0, clip_padding)
          min_duration = max(0.05, min_clip_duration)
          clips = []
          source_ranges = {}
          cursor = 0.0
      
          for idx, raw in enumerate(raw_clips):
              if not isinstance(raw, dict):
                  log(f"  跳过无效 clip #{idx + 1}: not an object")
                  continue
              source_id = raw.get("source_id", raw.get("id"))
              if source_id in (None, ""):
                  raise ValueError(f"clip #{idx + 1} is missing source_id")
              source_id = str(source_id)
              source = sources.get(source_id)
              if not source:
                  raise ValueError(
                      f"clip #{idx + 1} references unknown source_id: {source_id}"
                  )
              try:
                  raw_start = float(_clip_value(raw, "start", "source_start", "in"))
                  raw_end = float(_clip_value(raw, "end", "source_end", "out"))
              except (TypeError, ValueError):
                  log(f"  跳过无效 clip #{idx + 1}: missing numeric start/end")
                  continue
              if raw_end - raw_start < min_duration:
                  log(f"  跳过过短 clip #{idx + 1}: {raw_start:.1f}-{raw_end:.1f}s")
                  continue
              source_duration = source["duration"]
              start = round(max(0.0, min(raw_start - padding, source_duration)), 3)
              end = round(max(0.0, min(raw_end + padding, source_duration)), 3)
              if end - start < min_duration:
                  log(f"  跳过过短 clip #{idx + 1}: {start:.1f}-{end:.1f}s")
                  continue
              ranges = source_ranges.setdefault(source_id, [])
              if not allow_overlap and _overlaps_authored_range(ranges, raw_start, raw_end):
                  raise ValueError(
                      f"clip #{idx + 1} overlaps an earlier source range for source_id {source_id}; "
                      "split or remove duplicate source footage before mapping narration"
                  )
              ranges.append((raw_start, raw_end))
      
              duration = round(end - start, 3)
              clip = {
                  "clip_id": len(clips),
                  "source_id": source_id,
                  "source_path": source["source_path"],
                  "source_start": start,
                  "source_end": end,
                  "output_start": round(cursor, 3),
                  "output_end": round(cursor + duration, 3),
                  "duration": duration,
                  "reason": str(raw.get("reason", raw.get("note", ""))).strip(),
              }
              clips.append(clip)
              cursor += duration
      
          if not clips:
              raise ValueError("clip_plan.json has no valid clips")
      
          total_duration = round(sum(c["duration"] for c in clips), 3)
          plan = {
              "clips": clips,
              "total_duration": total_duration,
              "target_duration": round(target_duration, 3) if target_duration else None,
              "sources": {
                  sid: {
                      "source_path": s["source_path"],
                      "duration": round(s["duration"], 3),
                      **(
                          {"source_work_dir": s["source_work_dir"]}
                          if s.get("source_work_dir")
                          else {}
                      ),
                  }
                  for sid, s in sources.items()
              },
              "allow_overlap": bool(allow_overlap),
          }
          if target_duration and total_duration > target_duration * 1.15:
              plan["warning"] = (
                  f"validated clips total {total_duration:.1f}s exceeds target "
                  f"{target_duration:.1f}s by more than 15%"
              )
              log(f"警告: {plan['warning']}")
          return plan
      
      
      def normalize_clip_plan(
          raw_plan,
          video_duration,
          target_duration=None,
          clip_padding=0.0,
          min_clip_duration=0.3,
          allow_overlap=False,
      ):
          """Validate and enrich an agent-authored clip plan.
      
          Returns a dict with validated `clips`, `total_duration`, and target metadata.
          Clip order follows the agent-provided order, so montage ordering is possible.
          """
          if isinstance(raw_plan, dict):
              raw_clips = raw_plan.get("clips", [])
              plan_target = raw_plan.get("target_duration") or raw_plan.get(
                  "target_duration_seconds"
              )
              if target_duration is None and plan_target not in (None, ""):
                  target_duration = parse_duration_seconds(plan_target)
          elif isinstance(raw_plan, list):
              raw_clips = raw_plan
          else:
              raise ValueError(
                  "clip_plan.json must be a JSON array or an object with a clips array"
              )
      
          if not isinstance(raw_clips, list):
              raise ValueError("clip_plan.json field `clips` must be an array")
      
          padding = max(0.0, clip_padding)
          min_duration = max(0.05, min_clip_duration)
          clips = []
          source_ranges = []
          cursor = 0.0
      
          for idx, raw in enumerate(raw_clips):
              if not isinstance(raw, dict):
                  log(f"  跳过无效 clip #{idx + 1}: not an object")
                  continue
              try:
                  raw_start = float(_clip_value(raw, "start", "source_start", "in"))
                  raw_end = float(_clip_value(raw, "end", "source_end", "out"))
              except (TypeError, ValueError):
                  log(f"  跳过无效 clip #{idx + 1}: missing numeric start/end")
                  continue
              if raw_end - raw_start < min_duration:
                  log(f"  跳过过短 clip #{idx + 1}: {raw_start:.1f}-{raw_end:.1f}s")
                  continue
              start = round(max(0.0, min(raw_start - padding, video_duration)), 3)
              end = round(max(0.0, min(raw_end + padding, video_duration)), 3)
              if end - start < min_duration:
                  log(f"  跳过过短 clip #{idx + 1}: {start:.1f}-{end:.1f}s")
                  continue
              if not allow_overlap and _overlaps_authored_range(source_ranges, raw_start, raw_end):
                  raise ValueError(
                      f"clip #{idx + 1} overlaps an earlier source range; "
                      "split or remove duplicate source footage before mapping narration"
                  )
              source_ranges.append((raw_start, raw_end))
      
              duration = round(end - start, 3)
              clip = {
                  "clip_id": len(clips),
                  "source_start": start,
                  "source_end": end,
                  "output_start": round(cursor, 3),
                  "output_end": round(cursor + duration, 3),
                  "duration": duration,
                  "reason": str(raw.get("reason", raw.get("note", ""))).strip(),
              }
              clips.append(clip)
              cursor += duration
      
          if not clips:
              raise ValueError("clip_plan.json has no valid clips")
      
          total_duration = round(sum(c["duration"] for c in clips), 3)
          plan = {
              "clips": clips,
              "total_duration": total_duration,
              "target_duration": round(target_duration, 3) if target_duration else None,
              "source_duration": round(video_duration, 3),
              "allow_overlap": bool(allow_overlap),
          }
          if target_duration and total_duration > target_duration * 1.15:
              plan["warning"] = (
                  f"validated clips total {total_duration:.1f}s exceeds target "
                  f"{target_duration:.1f}s by more than 15%"
              )
              log(f"警告: {plan['warning']}")
          return plan
      
    • cut_render.py 11 KB
      """Render edited source media and write delivery QC."""
      
      import json
      
      
      from pathlib import Path
      
      from lib import CONFIG, filter_file_args, get_video_duration, log, run_cmd
      
      from cut_contract import _write_edited_source_meta
      from media_geometry import _has_audio_stream
      from sentence_boundaries import _continuous_source_join
      
      
      def _audio_segment_filter(
          label_in, label_out, start, end, duration, fade_in_ms, fade_out_ms, extra_filters=""
      ):
          max_fade = duration / 2
          fade_in = max(0.0, min(fade_in_ms / 1000.0, max_fade))
          fade_out = max(0.0, min(fade_out_ms / 1000.0, max_fade))
          base = f"{label_in}atrim=start={start:.3f}:end={end:.3f},asetpts=PTS-STARTPTS"
          if fade_in > 0:
              base += f",afade=t=in:st=0:d={fade_in:.3f}"
          if fade_out > 0:
              base += f",afade=t=out:st={max(0.0, duration - fade_out):.3f}:d={fade_out:.3f}"
          if extra_filters:
              base += f",{extra_filters}"
          return f"{base}{label_out}"
      
      
      def _clip_audio_edge_fades(clips, idx, fade_ms):
          """Do not create an audible dip where adjacent clips are a lossless source continuation."""
          fade_in = (
              0.0
              if idx > 0 and _continuous_source_join(clips[idx - 1], clips[idx])
              else fade_ms
          )
          fade_out = (
              0.0
              if idx + 1 < len(clips) and _continuous_source_join(clips[idx], clips[idx + 1])
              else fade_ms
          )
          return fade_in, fade_out
      
      
      def _probe_audio_sample_rate(video_path):
          """Sample rate of the first audio stream, or None when it cannot be observed.
      
          Delivery QC is observational and runs even on a plan that was never rendered, so an
          ffprobe that is absent (OSError) reads the same as one that reports no audio stream.
          """
          cmd = [
              "ffprobe",
              "-v",
              "error",
              "-select_streams",
              "a:0",
              "-show_entries",
              "stream=sample_rate",
              "-of",
              "csv=p=0",
              str(video_path),
          ]
          try:
              result = run_cmd(cmd)
          except OSError:
              return None
          text = result.stdout.strip()
          if result.returncode != 0 or not text:
              return None
          return int(float(text))
      
      
      def _delivery_reencode_reason(source_paths, clips):
          reasons = ["trim_concat_filter_requires_reencode"]
          if len(source_paths) > 1:
              reasons.append("multi_source_geometry_audio_normalization")
          if any("source_path" not in clip for clip in clips):
              reasons.append("single_source_filter_concat_no_stream_copy")
          return "+".join(reasons)
      
      
      def update_delivery_qc(validated_plan, *, source_paths, output_path=None, rendered=False):
          """Attach cut delivery facts to qc.delivery_qc without writing visual_qc."""
          qc = validated_plan["qc"]
          clips = validated_plan["clips"]
          target_sample_rate = 48000
          probed_sample_rate = (
              _probe_audio_sample_rate(output_path)
              if output_path and Path(output_path).exists()
              else None
          )
          delivery_qc = {
              "schema_version": 1,
              "video_encode_passes": 1,
              "reencode_reason": _delivery_reencode_reason(source_paths, clips),
              "stream_copy_risk": {
                  "status": "avoided",
                  "reason": "cut uses trim/concat/filtergraph with explicit libx264/aac encode; no risky stream-copy path",
              },
              "audio_sample_rate": {
                  "target": target_sample_rate,
                  "probed": probed_sample_rate,
              },
              "final_compat_notes": [
                  "video encoded with libx264/yuv420p-compatible filter path",
                  "audio encoded as AAC with 48000 Hz target for delivery compatibility",
                  "edited_source.mp4 is an intermediate; downstream assembly may perform another intentional encode",
              ],
              "output_geometry": qc["output_geometry"],
              "rendered": rendered,
              "planned": not rendered,
          }
          if probed_sample_rate and probed_sample_rate != target_sample_rate:
              delivery_qc["final_compat_notes"].append(
                  f"probed audio sample rate {probed_sample_rate} differs from target {target_sample_rate}"
              )
          qc["delivery_qc"] = delivery_qc
          return delivery_qc
      
      
      def write_cut_delivery_qc(work_dir, validated_plan):
          path = Path(work_dir) / "cut_delivery_qc.json"
          path.write_text(
              json.dumps(validated_plan["qc"]["delivery_qc"], ensure_ascii=False, indent=2),
              encoding="utf-8",
          )
          return path
      
      
      def build_edited_source_video(input_video, validated_plan, work_dir, output_path=None):
          """Build `edited_source.mp4` by concatenating validated source ranges.
      
          `validated_plan["qc"]["output_geometry"]` is required: the caller (cut_cli, or any
          public user of this API) selects the canvas with `_select_output_geometry` first, so the
          same geometry is recorded in clip_plan_validated.json and used for the render."""
          work_dir = Path(work_dir)
          output_path = Path(output_path or work_dir / "edited_source.mp4")
          clips = validated_plan["clips"]
          qc = validated_plan["qc"]
          if "output_geometry" not in qc:
              raise KeyError(
                  "validated_plan['qc']['output_geometry'] is required: select the canvas with "
                  "media_geometry._select_output_geometry before build_edited_source_video"
              )
      
          source_paths = []
          for clip in clips:
              source_path = clip.get("source_path", str(input_video))
              if source_path not in source_paths:
                  source_paths.append(source_path)
          source_index = {path: idx for idx, path in enumerate(source_paths)}
          audio_by_input = {path: _has_audio_stream(path) for path in source_paths}
          join_fade_ms = CONFIG["clip_join_audio_fade_ms"]
          qc["join_fade_ms"] = round(join_fade_ms, 3)
      
          parts = []
          concat_inputs = []
          extra_inputs = []
          if len(source_paths) > 1:
              # Distinct sources almost always differ in resolution/SAR/fps/pixel-format (and
              # some may lack audio), which the bare concat filter rejects. Normalize every video
              # segment to one canvas and give every clip an audio segment (real or synthesized
              # silence) so concat always succeeds with a continuous track and no source's audio
              # is dropped just because a sibling source is silent.
              geometry_qc = qc["output_geometry"]
              canvas_w, canvas_h, canvas_fps = (
                  geometry_qc["width"],
                  geometry_qc["height"],
                  geometry_qc["fps"],
              )
              vnorm = (
                  f"scale={canvas_w}:{canvas_h}:force_original_aspect_ratio=decrease,"
                  f"pad={canvas_w}:{canvas_h}:(ow-iw)/2:(oh-ih)/2,setsar=1,"
                  f"fps={canvas_fps},format=yuv420p"
              )
              for clip_pos, clip in enumerate(clips):
                  idx = clip["clip_id"]
                  clip_source = clip.get("source_path", str(input_video))
                  input_idx = source_index[clip_source]
                  start = clip["source_start"]
                  end = clip["source_end"]
                  dur = end - start
                  fade_in_ms, fade_out_ms = _clip_audio_edge_fades(clips, clip_pos, join_fade_ms)
                  parts.append(
                      f"[{input_idx}:v]trim=start={start:.3f}:end={end:.3f},setpts=PTS-STARTPTS,{vnorm}[v{idx}]"
                  )
                  if audio_by_input[clip_source]:
                      parts.append(
                          _audio_segment_filter(
                              f"[{input_idx}:a]",
                              f"[a{idx}]",
                              start,
                              end,
                              dur,
                              fade_in_ms,
                              fade_out_ms,
                              extra_filters="aresample=48000,aformat=sample_rates=48000:channel_layouts=stereo",
                          )
                      )
                  else:
                      parts.append(
                          f"anullsrc=r=48000:cl=stereo,atrim=duration={dur:.3f},asetpts=PTS-STARTPTS,"
                          f"aformat=sample_rates=48000:channel_layouts=stereo[a{idx}]"
                      )
                  concat_inputs.append(f"[v{idx}][a{idx}]")
              parts.append("".join(concat_inputs) + f"concat=n={len(clips)}:v=1:a=1[v][a]")
              maps = ["-map", "[v]", "-map", "[a]"]
          else:
              has_audio = all(audio_by_input.values())
              for clip_pos, clip in enumerate(clips):
                  idx = clip["clip_id"]
                  input_idx = source_index[clip.get("source_path", str(input_video))]
                  start = clip["source_start"]
                  end = clip["source_end"]
                  parts.append(
                      f"[{input_idx}:v]trim=start={start:.3f}:end={end:.3f},setpts=PTS-STARTPTS[v{idx}]"
                  )
                  concat_inputs.append(f"[v{idx}]")
                  if has_audio:
                      fade_in_ms, fade_out_ms = _clip_audio_edge_fades(
                          clips, clip_pos, join_fade_ms
                      )
                      parts.append(
                          _audio_segment_filter(
                              f"[{input_idx}:a]",
                              f"[a{idx}]",
                              start,
                              end,
                              end - start,
                              fade_in_ms,
                              fade_out_ms,
                          )
                      )
                      concat_inputs.append(f"[a{idx}]")
      
              if has_audio:
                  parts.append(
                      "".join(concat_inputs) + f"concat=n={len(clips)}:v=1:a=1[v][a]"
                  )
                  maps = ["-map", "[v]", "-map", "[a]"]
              else:
                  parts.append("".join(concat_inputs) + f"concat=n={len(clips)}:v=1:a=0[v]")
                  maps = ["-map", "[v]", "-map", f"{len(source_paths)}:a", "-shortest"]
                  extra_inputs = [
                      "-f",
                      "lavfi",
                      "-t",
                      f"{validated_plan['total_duration']:.3f}",
                      "-i",
                      "anullsrc=channel_layout=stereo:sample_rate=48000",
                  ]
      
          filter_complex = ";".join(parts)
          if len(filter_complex.encode("utf-8")) > 7000:
              filter_script = work_dir / "edit_filter_complex.txt"
              filter_script.write_text(filter_complex, encoding="utf-8")
              filter_args = filter_file_args("filter_complex", filter_script)
          else:
              filter_args = ["-filter_complex", filter_complex]
      
          input_args = []
          for source_path in source_paths:
              input_args.extend(["-i", str(source_path)])
          cmd = [
              "ffmpeg",
              "-y",
              *input_args,
              *extra_inputs,
              *filter_args,
              *maps,
              "-c:v",
              "libx264",
              "-preset",
              "veryfast",
              "-crf",
              "18",
              "-pix_fmt",
              "yuv420p",
              "-c:a",
              "aac",
              "-b:a",
              "192k",
              "-ar",
              "48000",
              "-movflags",
              "+faststart",
              str(output_path),
          ]
          result = run_cmd(cmd)
          if result.returncode != 0:
              raise RuntimeError(f"剪辑源视频失败: {result.stderr}")
      
          update_delivery_qc(
              validated_plan,
              source_paths=source_paths,
              output_path=output_path,
              rendered=True,
          )
          write_cut_delivery_qc(work_dir, validated_plan)
          _write_edited_source_meta(output_path, validated_plan, input_video)
          duration = get_video_duration(output_path)
          log(f"剪辑源视频: {output_path} ({duration:.1f}s, {len(clips)} clips)")
          return output_path
      
    • lib.py 4.7 KB
      """Self-contained utilities for the video-cut skill (no cross-skill imports)."""
      import functools
      import math
      import os
      import shutil
      import subprocess
      import tempfile
      
      
      def log(msg):
          print(f"[video-cut] {msg}", flush=True)
      
      
      def env_bool(name, default):
          """Read an env var as a boolean (1/true/yes → True; 0/false/no → False)."""
          val = os.environ.get(name)
          if val is None:
              return default
          text = val.strip().lower()
          if text in ("1", "true", "yes"):
              return True
          if text in ("0", "false", "no"):
              return False
          raise ValueError(f"{name} must be 1/true/yes or 0/false/no, got {val!r}")
      
      
      def env_float(name, default, min_val=None):
          """Read an env var as a float, rejecting malformed or below-minimum values."""
          val = os.environ.get(name)
          if val is None:
              return default
          try:
              result = float(val)
          except ValueError as exc:
              raise ValueError(f"{name} must be a number, got {val!r}") from exc
          if not math.isfinite(result):
              raise ValueError(f"{name} must be finite, got {val!r}")
          if min_val is not None and result < min_val:
              raise ValueError(f"{name} must be >= {min_val}, got {val!r}")
          return result
      
      
      CONFIG = {
          "snap_clip_line_end": env_bool("SNAP_CLIP_LINE_END", True),
          "clip_snap_max_extend": env_float("CLIP_SNAP_MAX_EXTEND", 2.0, min_val=0.0),
          "clip_start_snap_max_prepend": env_float("CLIP_START_SNAP_MAX_PREPEND", 1.8, min_val=0.0),
          "clip_start_snap_max_trim": env_float("CLIP_START_SNAP_MAX_TRIM", 0.35, min_val=0.0),
          "clip_join_audio_fade_ms": env_float("CLIP_JOIN_AUDIO_FADE_MS", 30.0, min_val=0.0),
          # video-cut is the only skill that implements clip padding, so it is the only skill that
          # may declare the knob. cut_cli reads this as the default for --clip-padding.
          "clip_padding": env_float("CLIP_PADDING", 0.0, min_val=0.0),
          # Keep clip boundaries off the ORIGINAL footage's hard cuts: a clip that opens/closes a few
          # tenths of a second from a source shot-change shows a brief sliver of the adjacent shot that
          # then hard-cuts again — a visible 闪烁/flicker at the edit point. Snap source_start forward
          # past (and source_end back before) any shot-change within the margin.
          "scene_cut_snap": env_bool("SCENE_CUT_SNAP", True),
          "scene_cut_snap_margin": env_float("SCENE_CUT_SNAP_MARGIN", 0.5, min_val=0.0),    # 边界±此秒内有切镜头才避让
          "scene_cut_detect_threshold": env_float("SCENE_CUT_DETECT_THRESHOLD", 0.4, min_val=0.0),  # ffmpeg scene 分数阈值(硬切)
      }
      
      
      def file_identity(path):
          """{size, mtime_ns} — the cache identity of an input file (no content read)."""
          st = os.stat(os.fspath(path))
          return {"size": st.st_size, "mtime_ns": st.st_mtime_ns}
      
      
      def run_cmd(cmd, **kwargs):
          """Run a command list and return the CompletedProcess (stdout/stderr captured)."""
          display = " ".join(
              str(part) if len(str(part)) <= 240 else str(part)[:237] + "..." for part in cmd
          )
          log(f"运行: {display}")
          return subprocess.run(cmd, capture_output=True, text=True, **kwargs)
      
      
      # ffmpeg 7 added `-/option path` to read any option's value from a file; ffmpeg 9 removed the
      # older `-filter_complex_script` / `-filter_script` spellings, which are all ffmpeg <= 6 knows.
      _LEGACY_FILTER_FILE_OPTIONS = {
          "filter_complex": "-filter_complex_script",
          "filter:v:0": "-filter_script:v:0",
      }
      
      
      @functools.lru_cache(maxsize=None)
      def _ffmpeg_reads_option_files():
          """Whether the ffmpeg on PATH accepts `-/option path` (asked once per process)."""
          if shutil.which("ffmpeg") is None:
              return False
          with tempfile.TemporaryDirectory() as tmp:
              graph = os.path.join(tmp, "probe_filter.txt")
              with open(graph, "w", encoding="utf-8") as fh:
                  fh.write("null")
              result = subprocess.run(["ffmpeg", "-hide_banner", "-/filter_complex", graph],
                                      stdin=subprocess.DEVNULL, capture_output=True, text=True,
                                      timeout=20)
          return "Unrecognized option" not in result.stderr
      
      
      def filter_file_args(option, path):
          """ffmpeg arguments that load `option`'s filtergraph from `path`, spelled for this ffmpeg."""
          if _ffmpeg_reads_option_files():
              return [f"-/{option}", str(path)]
          return [_LEGACY_FILTER_FILE_OPTIONS[option], str(path)]
      
      
      def get_video_duration(video_path):
          """Return media duration in seconds via ffprobe."""
          cmd = ["ffprobe", "-v", "error", "-show_entries", "format=duration",
                 "-of", "csv=p=0", str(video_path)]
          result = run_cmd(cmd)
          if result.returncode != 0:
              raise RuntimeError(f"ffprobe 无法读取时长: {video_path}: {result.stderr.strip()}")
          return float(result.stdout.strip())
      
    • media_geometry.py 8.4 KB
      """Probe display geometry and select a stable output canvas."""
      
      import json
      
      from lib import run_cmd
      
      
      def _has_audio_stream(video_path):
          cmd = [
              "ffprobe",
              "-v",
              "error",
              "-select_streams",
              "a:0",
              "-show_entries",
              "stream=index",
              "-of",
              "csv=p=0",
              str(video_path),
          ]
          result = run_cmd(cmd)
          return result.returncode == 0 and bool(result.stdout.strip())
      
      
      class VideoGeometry(tuple):
          """Tuple-compatible (width, height, fps) with probe facts attached for QC callers."""
      
          def __new__(cls, width, height, fps, facts):
              obj = super().__new__(cls, (width, height, fps))
              obj.facts = facts
              return obj
      
      
      def _parse_ratio(value):
          """ffprobe aspect ratios are 'N:D' (or 'N/D'); '0:1' and 'N/A' mean unknown."""
          if value in (None, "", "0:1", "0/1", "N/A"):
              return None
          left, _, right = str(value).replace("/", ":").partition(":")
          num, den = float(left), float(right)
          return num / den if den > 0 and num > 0 else None
      
      
      def _stream_rotation(stream):
          """Rotation from the legacy `rotate` tag or the display-matrix side data, else 0."""
          tags = stream.get("tags", {})
          if "rotate" in tags:
              return int(round(float(tags["rotate"]))) % 360
          for side_data in stream.get("side_data_list", []):
              if "rotation" in side_data:
                  return int(round(float(side_data["rotation"]))) % 360
          return 0
      
      
      def _fps_from_rate(rate):
          """ffprobe frame rates are 'N/D' fractions; '0/0' means unknown."""
          num, _, den = rate.partition("/")
          return float(num) / float(den) if float(den) > 0 else 0.0
      
      
      def _geometry_from_stream(stream):
          coded_width, coded_height = stream["width"], stream["height"]
          parsed_sar = _parse_ratio(stream.get("sample_aspect_ratio"))
          dar = _parse_ratio(stream.get("display_aspect_ratio"))
          rotation = _stream_rotation(stream)
          display_height = float(coded_height)
          if parsed_sar:
              sar = parsed_sar
              display_width = float(coded_width) * sar
              aspect_source = "sample_aspect_ratio"
          elif dar:
              sar = 1.0
              display_width = display_height * dar
              aspect_source = "display_aspect_ratio_fallback"
          else:
              sar = 1.0
              display_width = float(coded_width)
              aspect_source = "square_pixel_fallback"
          rotation_swaps_axes = rotation in {90, 270}
          if rotation_swaps_axes:
              display_width, display_height = display_height, display_width
      
          width, height = _clamp_even_geometry(round(display_width), round(display_height))
          fps = _fps_from_rate(stream["r_frame_rate"]) or _fps_from_rate(
              stream.get("avg_frame_rate", "0/0")
          )
          if not 0 < fps <= 120:
              fps = 30.0
          facts = {
              "coded_width": coded_width,
              "coded_height": coded_height,
              "width": width,
              "height": height,
              "fps": round(fps, 3),
              "sample_aspect_ratio": stream.get("sample_aspect_ratio", "1:1"),
              "sample_aspect_ratio_float": round(sar, 6),
              "display_aspect_ratio": stream.get("display_aspect_ratio"),
              "display_aspect_ratio_float": round(dar or 0.0, 6),
              "display_aspect_source": aspect_source,
              "display_width": width,
              "display_height": height,
              "rotation": rotation,
              "rotation_swaps_axes": rotation_swaps_axes,
          }
          return VideoGeometry(width, height, round(fps, 3), facts)
      
      
      def _probe_video_geometry(video_path):
          """Display geometry (width, height, fps) of the first video stream, rotation/SAR/DAR-aware.
      
          Unpacks like a 3-tuple while exposing `.facts` for QC. Used to normalize heterogeneous
          multi-source segments to one square-pixel geometry before concat (ffmpeg's concat filter
          rejects mismatched width/height/SAR/pixel-format/fps).
          """
          cmd = [
              "ffprobe",
              "-v",
              "error",
              "-select_streams",
              "v:0",
              "-show_entries",
              "stream=width,height,r_frame_rate,avg_frame_rate,sample_aspect_ratio,display_aspect_ratio:stream_tags=rotate:stream_side_data=rotation",
              "-of",
              "json",
              str(video_path),
          ]
          result = run_cmd(cmd)
          if result.returncode != 0:
              raise RuntimeError(f"ffprobe 无法读取视频几何信息: {video_path}: {result.stderr.strip()}")
          streams = json.loads(result.stdout).get("streams", [])
          if not streams:
              raise RuntimeError(f"没有视频流: {video_path}")
          return _geometry_from_stream(streams[0])
      
      
      def _orientation(width, height):
          if width > height:
              return "landscape"
          if height > width:
              return "portrait"
          return "square"
      
      
      def _fps_bucket(fps):
          common = [23.976, 24.0, 25.0, 29.97, 30.0, 50.0, 59.94, 60.0]
          nearest = min(common, key=lambda x: abs(fps - x))
          bucket = nearest if abs(fps - nearest) <= 0.15 else round(fps)
          return max(1.0, min(60.0, bucket))
      
      
      def _clamp_even_geometry(width, height):
          return max(2, width - width % 2), max(2, height - height % 2)
      
      
      def _select_output_geometry(source_paths, clips):
          """Deterministically select canvas/fps from all used sources, not just the first."""
          used = {}
          for clip in clips:
              if "source_path" in clip:
                  used[clip["source_path"]] = used.get(clip["source_path"], 0.0) + clip["duration"]
          if not used:  # single-source plans carry no per-clip source_path
              used = {str(path): 0.0 for path in source_paths}
          rows = []
          for path in sorted(used):
              probed = _probe_video_geometry(path)
              width, height, fps = probed
              facts = probed.facts
              rows.append(
                  {
                      "path": path,
                      "source_id": next(
                          (c["source_id"] for c in clips if c.get("source_path") == path), None
                      ),
                      "used_duration": round(used[path], 3),
                      "width": width,
                      "height": height,
                      "coded_width": facts["coded_width"],
                      "coded_height": facts["coded_height"],
                      "display_width": facts["display_width"],
                      "display_height": facts["display_height"],
                      "area": width * height,
                      "fps": fps,
                      "fps_bucket": _fps_bucket(fps),
                      "orientation": _orientation(width, height),
                      "rotation": facts["rotation"],
                      "sample_aspect_ratio": facts["sample_aspect_ratio"],
                      "sample_aspect_ratio_float": facts["sample_aspect_ratio_float"],
                      "display_aspect_ratio": facts["display_aspect_ratio"],
                      "rotation_swaps_axes": facts["rotation_swaps_axes"],
                  }
              )
      
          orientation_duration = {}
          for row in rows:
              orientation_duration[row["orientation"]] = (
                  orientation_duration.get(row["orientation"], 0.0) + row["used_duration"]
              )
          chosen_orientation = sorted(
              orientation_duration.items(),
              key=lambda kv: (
                  kv[1],
                  max(r["area"] for r in rows if r["orientation"] == kv[0]),
                  kv[0],
              ),
              reverse=True,
          )[0][0]
          eligible = [r for r in rows if r["orientation"] == chosen_orientation]
          selected = sorted(
              eligible, key=lambda r: (-r["area"], r["source_id"] or "", r["path"])
          )[0]
      
          fps_duration = {}
          for row in rows:
              fps_duration[row["fps_bucket"]] = (
                  fps_duration.get(row["fps_bucket"], 0.0) + row["used_duration"]
              )
          fps = sorted(fps_duration.items(), key=lambda kv: (kv[1], kv[0]), reverse=True)[0][0]
      
          width, height = selected["width"], selected["height"]
          reason = {
              "width": width,
              "height": height,
              "fps": round(fps, 3),
              "reason": "weighted_orientation_area_fps",
              "source_id": selected["source_id"],
              "source_path": selected["path"],
              "orientation": chosen_orientation,
              "orientation_used_duration": round(orientation_duration[chosen_orientation], 3),
              "fps_bucket_used_duration": round(fps_duration[fps], 3),
              "rotation": selected["rotation"],
              "sample_aspect_ratio": selected["sample_aspect_ratio"],
              "display_aspect_ratio": selected["display_aspect_ratio"],
              "coded_width": selected["coded_width"],
              "coded_height": selected["coded_height"],
              "display_width": selected["display_width"],
              "display_height": selected["display_height"],
              "sources": rows,
          }
          return width, height, round(fps, 3), reason
      
    • narration_mapping.py 2.6 KB
      """Cut QC bookkeeping for clip_plan_validated.json."""
      
      
      def update_cut_qc(plan, *, allow_duration_drift=False, duration_drift_allowed_by=None):
          """Populate clip_plan_validated.json['qc'] as the single cut QC source."""
          qc = dict(plan.get("qc", {}))
          warnings = list(qc.get("warnings", []))
          blocking = list(qc.get("blocking", []))
          total = plan["total_duration"]
          target = plan["target_duration"]
          if target is None:
              target_status = "missing"
              target_qc = {
                  "status": target_status,
                  "target_duration": None,
                  "total_duration": round(total, 3),
              }
          else:
              ratio = total / target  # parse_duration_seconds guarantees target > 0
              if ratio < 0.85:
                  target_status = "under"
              elif ratio > 1.15:
                  target_status = "over"
              else:
                  target_status = "ok"
              severity = None
              if ratio < 0.60 or ratio > 1.40:
                  severity = "blocking"
              elif target_status in {"under", "over"}:
                  severity = "warning"
              target_qc = {
                  "status": target_status,
                  "target_duration": round(target, 3),
                  "total_duration": round(total, 3),
                  "ratio": round(ratio, 3),
                  "warning_thresholds": {"under": 0.85, "over": 1.15},
                  "blocking_thresholds": {"under": 0.60, "over": 1.40},
              }
              if severity:
                  warning = {
                      "code": "target_duration_drift",
                      "status": target_status,
                      "severity": "warning" if allow_duration_drift else severity,
                      "target_duration": round(target, 3),
                      "total_duration": round(total, 3),
                      "ratio": round(ratio, 3),
                  }
                  if allow_duration_drift:
                      warning["allowed"] = True
                      warning["duration_drift_allowed_by"] = (
                          duration_drift_allowed_by or "--allow-duration-drift"
                      )
                      target_qc["duration_drift_allowed_by"] = warning[
                          "duration_drift_allowed_by"
                      ]
                  warnings.append(warning)
                  if severity == "blocking" and not allow_duration_drift:
                      blocking.append(warning)
          qc["target_duration_status"] = target_status
          qc["target_duration"] = target_qc
          qc.setdefault("boundary_status", {})
          qc["clip_count"] = len(plan["clips"])
          qc["total_duration"] = round(total, 3)
          if warnings:
              qc["warnings"] = warnings
          if blocking:
              qc["blocking"] = blocking
          else:
              qc.pop("blocking", None)
          plan["qc"] = qc
          return plan
      
    • narrative_selection.py 8.8 KB
      """Validate explicitly required source spans in a normalized constant-speed cut plan.
      
      This checks source-span selection and ordering only.  It does not claim that the
      rendered mix is audible or that the declared content is semantically correct.
      """
      
      import math
      import os
      from typing import TypeGuard
      
      
      _EPSILON = 1e-9
      
      
      def _invalid(message):
          return {
              "selection_status": "BLOCK",
              "nodes": [],
              "findings": [
                  {
                      "code": "REQUIRED_EVIDENCE_INVALID",
                      "message": message,
                  }
              ],
              "semantic_status": "NOT_CHECKED",
          }
      
      
      def _realpath(value, *, require_absolute=True):
          if not isinstance(value, str) or (require_absolute and not os.path.isabs(value)):
              raise ValueError("required evidence source must be an absolute path")
          return os.path.realpath(value)
      
      
      def _finite_nonnegative(value: object) -> TypeGuard[int | float]:
          return (
              isinstance(value, (int, float))
              and not isinstance(value, bool)
              and math.isfinite(value)
              and value >= 0
          )
      
      
      def _validate_contract(contract):
          if not isinstance(contract, dict):
              raise ValueError("required_evidence must be an object")
          nodes = contract.get("nodes")
          if not isinstance(nodes, list) or not nodes:
              raise ValueError("required_evidence.nodes must be a non-empty array")
          if not isinstance(contract.get("before"), list):
              raise ValueError("required_evidence.before must be an array")
      
          normalized = []
          ids = set()
          for index, node in enumerate(nodes, 1):
              if not isinstance(node, dict):
                  raise ValueError(f"required evidence node #{index} must be an object")
              node_id = node.get("id")
              if not isinstance(node_id, str) or not node_id.strip() or node_id in ids:
                  raise ValueError("required evidence node ids must be unique non-empty strings")
              ids.add(node_id)
              track = node.get("track")
              if track not in {"audio", "video"}:
                  raise ValueError(f"required evidence node {node_id} has invalid track")
              content = node.get("content")
              if not isinstance(content, str) or not content.strip():
                  raise ValueError(f"required evidence node {node_id} has empty content")
              start, end = node.get("start"), node.get("end")
              if not (_finite_nonnegative(start) and _finite_nonnegative(end) and end > start):
                  raise ValueError(f"required evidence node {node_id} has invalid start/end")
              source_id = node.get("source_id")
              if "source_id" in node and (
                  not isinstance(source_id, str) or not source_id.strip()
              ):
                  raise ValueError(f"required evidence node {node_id} has invalid source_id")
      
              item = {
                  "id": node_id,
                  "track": track,
                  "content": content,
                  "source": _realpath(node.get("source")),
                  "start": start,
                  "end": end,
              }
              if "source_id" in node:
                  item["source_id"] = source_id
              normalized.append(item)
      
          edges = []
          for edge in contract["before"]:
              if (
                  not isinstance(edge, list)
                  or len(edge) != 2
                  or edge[0] not in ids
                  or edge[1] not in ids
                  or edge[0] == edge[1]
              ):
                  raise ValueError("required_evidence.before contains an invalid node pair")
              edges.append((edge[0], edge[1]))
          return normalized, edges
      
      
      def _same(value, expected):
          return abs(value - expected) <= _EPSILON
      
      
      def _matching_segments(node, validated_plan, input_video):
          segments = []
          for clip in validated_plan["clips"]:
              source = _realpath(
                  os.fspath(clip.get("source_path", input_video)), require_absolute=False
              )
              if source != node["source"]:
                  continue
              if "source_id" in node and clip.get("source_id") != node["source_id"]:
                  continue
      
              clip_start = clip["source_start"]
              clip_end = clip["source_end"]
              start = max(node["start"], clip_start)
              end = min(node["end"], clip_end)
              if end <= start:
                  continue
              output_start = clip["output_start"] + start - clip_start
              segments.append(
                  {
                      "source_id": clip.get("source_id"),
                      "source_start": start,
                      "source_end": end,
                      "output_start": output_start,
                      "output_end": output_start + end - start,
                  }
              )
          return sorted(segments, key=lambda item: item["output_start"])
      
      
      def _occurrences(node, validated_plan, input_video):
          segments = _matching_segments(node, validated_plan, input_video)
          occurrences = []
          for index, segment in enumerate(segments):
              if not _same(segment["source_start"], node["start"]):
                  continue
              source_cursor = segment["source_end"]
              output_cursor = segment["output_end"]
              for following in segments[index + 1 :]:
                  if source_cursor >= node["end"] - _EPSILON:
                      break
                  if (
                      following["source_id"] == segment["source_id"]
                      and _same(following["source_start"], source_cursor)
                      and _same(following["output_start"], output_cursor)
                  ):
                      source_cursor = following["source_end"]
                      output_cursor = following["output_end"]
              if _same(source_cursor, node["end"]):
                  occurrences.append(
                      {"start": segment["output_start"], "end": output_cursor}
                  )
          return occurrences
      
      
      def check_required_evidence(contract, validated_plan, *, input_video, source_audio) -> dict:
          """Return bounded source-selection QC for a declared required-evidence contract.
      
          `source_audio(realpath) -> bool` is consulted only for audio nodes, after the contract
          has been validated, so callers never pre-walk the raw contract to decide whether to probe."""
          try:
              nodes, edges = _validate_contract(contract)
          except (AttributeError, KeyError, TypeError, ValueError) as exc:
              return _invalid(str(exc))
      
          findings = []
          occurrences_by_id = {}
          fragment_starts_by_id = {}
          report_nodes = []
          for node in nodes:
              if node["track"] == "audio" and source_audio(node["source"]) is not True:
                  occurrences = []
                  fragment_starts = []
                  findings.append(
                      {
                          "code": "REQUIRED_EVIDENCE_AUDIO_UNAVAILABLE",
                          "message": (
                              f"audio node {node['id']} requires a source with an audio stream"
                          ),
                          "node_id": node["id"],
                      }
                  )
              else:
                  fragment_starts = [
                      segment["output_start"]
                      for segment in _matching_segments(node, validated_plan, input_video)
                  ]
                  occurrences = _occurrences(node, validated_plan, input_video)
                  if not occurrences:
                      findings.append(
                          {
                              "code": "REQUIRED_EVIDENCE_MISSING",
                              "message": (
                                  f"node {node['id']} is not retained as one continuous source "
                                  "and output span"
                              ),
                              "node_id": node["id"],
                          }
                      )
              occurrences_by_id[node["id"]] = occurrences
              fragment_starts_by_id[node["id"]] = fragment_starts
              report_nodes.append({**node, "occurrences": occurrences})
      
          for premise_id, result_id in edges:
              premise_occurrences = occurrences_by_id[premise_id]
              result_fragment_starts = fragment_starts_by_id[result_id]
              unpreceded_starts = [
                  result_start
                  for result_start in result_fragment_starts
                  if not any(
                      premise["end"] <= result_start + _EPSILON
                      for premise in premise_occurrences
                  )
              ]
              ordered = bool(result_fragment_starts) and not unpreceded_starts
              if not ordered:
                  detail = (
                      f"; earliest unpreceded output start={min(unpreceded_starts):.9g}s"
                      if unpreceded_starts
                      else "; no matching result fragment was retained"
                  )
                  findings.append(
                      {
                          "code": "REQUIRED_EVIDENCE_ORDER",
                          "message": (
                              f"every retained fragment of {result_id} must follow a complete "
                              f"occurrence of {premise_id}{detail}"
                          ),
                          "node_id": result_id,
                      }
                  )
      
          return {
              "selection_status": "BLOCK" if findings else "PASS",
              "nodes": report_nodes,
              "findings": findings,
              "semantic_status": "NOT_CHECKED",
          }
      
    • sentence_boundaries.py 22.3 KB
      """Snap clip boundaries to complete speech and clean shot transitions."""
      
      import json
      import re
      import subprocess
      from pathlib import Path
      
      from lib import log
      
      
      def _find_source_artifact(work_dir, filename, source_id=None, source_work_dir=None):
          """First existing copy of a per-source understanding artifact, by layout precedence."""
          candidates = []
          if source_work_dir:
              candidates.append(Path(work_dir) / source_work_dir / filename)
          if source_id is not None:
              candidates.append(Path(work_dir) / "sources" / source_id / filename)
          candidates.append(Path(work_dir) / filename)
          return next((path for path in candidates if path.exists()), None)
      
      
      def _load_sentence_boundary_windows(work_dir, source_id=None, source_work_dir=None):
          """Load reliable sentence-end pause windows produced by video-understanding.
      
          A sentence anchor's `time` is the acoustic pause end, while `pause_start` is already
          after the final spoken sample. Any cut within that closed interval preserves the sentence.
          Low-confidence anchors are deliberately excluded from the hard-safety path.
          """
          path = _find_source_artifact(
              work_dir, "speech_boundary_anchors.json", source_id, source_work_dir
          )
          if path is None:
              return []
          payload = json.loads(path.read_text(encoding="utf-8"))
          windows = [
              {
                  "start": round(anchor["pause_start"], 3),
                  "end": round(anchor["time"], 3),
                  "kind": "sentence_anchor",
                  "confidence": anchor["confidence"],
              }
              for anchor in payload["sentence_anchors"]
              if anchor["confidence"] in {"high", "medium"}
          ]
          return sorted(windows, key=lambda row: (row["start"], row["end"]))
      
      
      def _load_source_speech_spans(work_dir, source_id=None, source_work_dir=None):
          """Merged ASR speech spans (asr_clean.json wins over asr_result.json).
      
          Only used to decide whether an unsafe edge blocks; a missing transcript means unchecked.
          """
          rows = []
          for filename in ("asr_clean.json", "asr_result.json"):
              path = _find_source_artifact(work_dir, filename, source_id, source_work_dir)
              if path is not None:
                  payload = json.loads(path.read_text(encoding="utf-8"))
                  rows = payload["segments"] if filename == "asr_clean.json" else payload
                  break
          spans = sorted(
              ({"start": row["start"], "end": row["end"]} for row in rows if row["text"].strip()),
              key=lambda row: (row["start"], row["end"]),
          )
          merged = []
          for span in spans:
              if merged and span["start"] <= merged[-1]["end"] + 0.05:
                  merged[-1]["end"] = max(merged[-1]["end"], span["end"])
              else:
                  merged.append(span)
          return merged
      
      
      def _combine_boundary_windows(*groups):
          unique = {
              (round(row["start"], 3), round(row["end"], 3)) for group in groups for row in group
          }
          return [{"start": start, "end": end} for start, end in sorted(unique)]
      
      
      def _same_source(left, right):
          # Single-source normalized plans omit source identity; both sides then read as None.
          return left.get("source_id") == right.get("source_id") and left.get(
              "source_path"
          ) == right.get("source_path")
      
      
      def _continuous_source_join(left, right, tolerance=0.05):
          return (
              _same_source(left, right)
              and abs(left["source_end"] - right["source_start"]) <= tolerance
              and abs(left["output_end"] - right["output_start"]) <= tolerance
          )
      
      
      def enforce_clip_sentence_boundaries(
          plan, boundary_windows, speech_spans, video_duration, tolerance=0.05
      ):
          """Block any audible clip edge that falls inside detected source speech.
      
          Safe edges are: source start/end, a reliable sentence/quiet pause, or a truly contiguous
          same-source join (no media is removed). Missing ASR timing degrades to `unchecked` rather
          than inventing speech. Once ASR says an edge is speech-owned, failure to snap is blocking.
          """
          clips = plan["clips"]
          checks, new_blockers = [], []
      
          def inside(rows, ts):
              return any(row["start"] - tolerance <= ts <= row["end"] + tolerance for row in rows)
      
          for idx, clip in enumerate(clips):
              for edge, ts in (("start", clip["source_start"]), ("end", clip["source_end"])):
                  contiguous = (
                      edge == "start"
                      and idx > 0
                      and _continuous_source_join(clips[idx - 1], clip, tolerance)
                  ) or (
                      edge == "end"
                      and idx + 1 < len(clips)
                      and _continuous_source_join(clip, clips[idx + 1], tolerance)
                  )
                  if edge == "start" and ts <= tolerance:
                      status, reason = "safe", "source_start"
                  elif edge == "end" and ts >= video_duration - tolerance:
                      status, reason = "safe", "source_end"
                  elif contiguous:
                      status, reason = "safe", "continuous_source_join"
                  elif inside(boundary_windows, ts):
                      status, reason = "safe", "sentence_or_quiet_boundary"
                  elif not speech_spans:
                      status, reason = "unchecked", "speech_timing_unavailable"
                  elif not inside(speech_spans, ts):
                      status, reason = "safe", "outside_detected_speech"
                  else:
                      status, reason = "blocking", "inside_detected_speech"
                  check = {
                      "clip_id": clip["clip_id"],
                      "source_id": clip.get("source_id"),
                      "edge": edge,
                      "time": round(ts, 3),
                      "status": status,
                      "reason": reason,
                  }
                  checks.append(check)
                  if status == "blocking":
                      new_blockers.append(
                          {
                              "code": "unsafe_clip_sentence_boundary",
                              **check,
                              "message": "剪辑边界仍落在原声讲话区间内,必须移动到句末锚点,不能截断原声句子。",
                          }
                      )
      
          qc = plan.setdefault("qc", {})
          qc.setdefault("boundary_status", {})["sentence_checks"] = checks
          existing = [
              row
              for row in qc.get("blocking", [])
              if row["code"] != "unsafe_clip_sentence_boundary"
          ]
          if existing or new_blockers:
              qc["blocking"] = existing + new_blockers
          else:
              qc.pop("blocking", None)
          return plan
      
      
      def _recompute_clip_timeline(clips):
          cursor = 0.0
          for clip in clips:
              duration = round(clip["source_end"] - clip["source_start"], 3)
              clip["duration"] = duration
              clip["output_start"] = round(cursor, 3)
              clip["output_end"] = round(cursor + duration, 3)
              cursor += duration
          return round(cursor, 3)
      
      
      def _plan_with_snapped_clips(plan, clips, boundary_key, events):
          """Copy of `plan` carrying the snapped clips, a cursor-based output timeline, and the
          per-pass boundary events under qc.boundary_status[boundary_key]."""
          result = dict(plan)
          result["clips"] = clips
          result["total_duration"] = _recompute_clip_timeline(clips)
          qc = dict(plan.get("qc", {}))
          boundary = dict(qc.get("boundary_status", {}))
          boundary[boundary_key] = events
          qc["boundary_status"] = boundary
          result["qc"] = qc
          return result
      
      
      def _candidate_overlaps(clips, idx, new_start, new_end):
          return any(
              new_start < other["source_end"] and new_end > other["source_start"]
              for j, other in enumerate(clips)
              if j != idx
          )
      
      
      def snap_clip_starts_to_lines(
          plan,
          silence_periods,
          video_duration,
          max_prepend,
          max_trim=0.35,
          min_clip_duration=0.3,
      ):
          """Snap clip starts to natural quiet boundaries, preferring safe prepend over trim.
      
          Policy:
          - start already inside a quiet window: keep.
          - speech start: prepend to nearest prior quiet window end if within max_prepend.
          - no usable prior quiet: keep and warn, except an extremely near next quiet start
            (<= max_trim) may trim forward if duration/overlap safety holds.
          - never overlap/collapse when allow_overlap is false; unsafe attempts keep-and-warn.
          """
          if not silence_periods:
              return plan
      
          clips = [dict(c) for c in plan["clips"]]
          allow_overlap = plan["allow_overlap"]
          events = []
      
          for i, clip in enumerate(clips):
              original_start = clip["source_start"]
              source_end = clip["source_end"]
              event = {
                  "clip_id": clip["clip_id"],
                  "source_id": clip.get("source_id"),
                  "original_start": round(original_start, 3),
                  "action": "kept",
              }
      
              if any(w["start"] <= original_start <= w["end"] for w in silence_periods):
                  event["reason"] = "already_quiet"
                  events.append(event)
                  continue
      
              prior_ends = [w["end"] for w in silence_periods if w["end"] <= original_start]
              if prior_ends:
                  candidate_start = max(prior_ends)
                  if original_start - candidate_start <= max_prepend:
                      candidate_start = round(candidate_start, 3)
                      safe = candidate_start < source_end - min_clip_duration + 1e-9
                      if safe and not allow_overlap:
                          safe = not _candidate_overlaps(clips, i, candidate_start, source_end)
                      if safe:
                          clip["source_start"] = candidate_start
                          event.update(
                              {
                                  "action": "prepended",
                                  "new_start": candidate_start,
                                  "delta": round(original_start - candidate_start, 3),
                              }
                          )
                          events.append(event)
                          continue
                      reason = "overlap_or_collapse"
                  else:
                      reason = "prior_quiet_too_far"
              else:
                  reason = "no_prior_quiet"
      
              next_starts = [w["start"] for w in silence_periods if w["start"] >= original_start]
              if next_starts:
                  candidate_start = min(next_starts)
                  trim_delta = candidate_start - original_start
                  if 0 < trim_delta <= max_trim:
                      candidate_start = round(min(video_duration, candidate_start), 3)
                      safe = source_end - candidate_start >= min_clip_duration
                      if safe and not allow_overlap:
                          safe = not _candidate_overlaps(clips, i, candidate_start, source_end)
                      if safe:
                          clip["source_start"] = candidate_start
                          event.update(
                              {
                                  "action": "trimmed",
                                  "new_start": candidate_start,
                                  "delta": round(trim_delta, 3),
                                  "fallback_from": reason,
                              }
                          )
                          events.append(event)
                          continue
                      reason = "unsafe_forward_trim"
      
              event["start_unsnapped_reason"] = reason
              event["warning_code"] = "clip_start_unsnapped"
              events.append(event)
      
          result = _plan_with_snapped_clips(plan, clips, "start_snaps", events)
          warnings = list(result["qc"].get("warnings", []))
          for event in events:
              if "warning_code" in event:
                  warnings.append(
                      {
                          "code": event["warning_code"],
                          "clip_id": event["clip_id"],
                          "source_id": event["source_id"],
                          "start_unsnapped_reason": event["start_unsnapped_reason"],
                      }
                  )
          if warnings:
              result["qc"]["warnings"] = warnings
          return result
      
      
      def snap_clip_ends_to_lines(plan, silence_periods, video_duration, max_extend):
          """Extend each clip's source_end forward to the next natural pause, preventing mid-sentence cuts.
      
          - No quiet windows: the plan is returned unchanged.
          - A source_end already inside a quiet window is left alone.
          - Otherwise extend to the next quiet window start, capped by max_extend and video_duration.
          - When plan["allow_overlap"] is False, never extend into another clip's source range.
          - Recomputes output_start/output_end/duration for all clips cursor-based.
          """
          if not silence_periods:
              return plan
      
          clips = [dict(c) for c in plan["clips"]]
          allow_overlap = plan["allow_overlap"]
          events = []
      
          for i, clip in enumerate(clips):
              source_end = clip["source_end"]
              event = {
                  "clip_id": clip["clip_id"],
                  "source_id": clip.get("source_id"),
                  "original_end": round(source_end, 3),
                  "action": "kept",
              }
      
              if any(w["start"] <= source_end <= w["end"] for w in silence_periods):
                  event["reason"] = "already_quiet"
                  events.append(event)
                  continue
      
              candidates = [w["start"] for w in silence_periods if w["start"] >= source_end]
              if not candidates:
                  event["end_unsnapped_reason"] = "no_next_quiet"
                  events.append(event)
                  continue
              next_quiet_start = min(candidates)
      
              if next_quiet_start > source_end + max_extend:
                  event["end_unsnapped_reason"] = "next_quiet_too_far"
                  events.append(event)
                  continue
              candidate_end = min(next_quiet_start, video_duration)
              if candidate_end <= source_end:
                  event["end_unsnapped_reason"] = "non_forward_candidate"
                  events.append(event)
                  continue
      
              # When overlaps are forbidden, cap against every other clip's source range.
              if not allow_overlap:
                  other_starts = [
                      c["source_start"]
                      for j, c in enumerate(clips)
                      if j != i and c["source_start"] > source_end
                  ]
                  if other_starts:
                      candidate_end = min(candidate_end, min(other_starts))
                  if candidate_end <= source_end:
                      event["end_unsnapped_reason"] = "overlap_or_collapse"
                      events.append(event)
                      continue
      
              clip["source_end"] = round(candidate_end, 3)
              event.update(
                  {
                      "action": "extended",
                      "new_end": clip["source_end"],
                      "delta": round(clip["source_end"] - source_end, 3),
                  }
              )
              events.append(event)
      
          return _plan_with_snapped_clips(plan, clips, "end_snaps", events)
      
      
      def _detect_shot_changes(video, win_start, win_end, threshold, lead=0.25):
          """Absolute source-time hard cuts inside [win_start, win_end] via ffmpeg's scene metric.
      
          Input-seek to a little before the window: the rebased output PTS restarts at ~0 at the seek
          target, so `seek + pts_time` recovers absolute source time. The `lead` keeps the seek/keyframe
          settling artifact frame outside [win_start, win_end] so it is filtered out, not mistaken for a
          cut.
          """
          if win_end - win_start < 1e-3:
              return []
          seek = max(0.0, win_start - lead)
          dur = (win_end - seek) + 0.1
          cmd = [
              "ffmpeg",
              "-hide_banner",
              "-nostats",
              "-ss",
              f"{seek:.3f}",
              "-i",
              str(video),
              "-t",
              f"{dur:.3f}",
              "-an",
              "-sn",
              "-filter:v",
              f"select='gt(scene,{threshold})',showinfo",
              "-f",
              "null",
              "-",
          ]
          proc = subprocess.run(cmd, capture_output=True, text=True)
          if proc.returncode != 0:
              raise RuntimeError(f"ffmpeg 切镜头检测失败: {video}: {proc.stderr.strip()[-500:]}")
          changes = set()
          for m in re.finditer(r"pts_time:([0-9.]+)", proc.stderr):
              t = seek + float(m.group(1))
              if win_start <= t <= win_end:
                  changes.add(round(t, 3))
          return sorted(changes)
      
      
      def snap_clips_off_shot_changes(plan, video, margin, threshold, min_keep=0.5):
          """Nudge each clip's boundaries clear of the ORIGINAL footage's hard cuts to avoid 闪烁.
      
          A clip whose source_start sits just before a shot-change opens on a brief sliver of the old
          shot that then hard-cuts again; one whose source_end sits just after a shot-change closes on a
          sliver of the next shot. Both flash at the edit point. So:
            - move source_start FORWARD onto a shot-change in (start, start+margin]  → clean open
            - move source_end   BACK   onto a shot-change in [end-margin, end)       → clean close
          Boundaries already on a cut, or with no nearby cut, are left untouched. Snaps that would shrink
          a clip below `min_keep` are skipped. Recomputes the output timeline cursor-based.
          """
          if margin <= 0:
              return plan
          clips = [dict(c) for c in plan["clips"]]
          n_start = n_end = 0
          events = []
          for clip in clips:
              s = clip["source_start"]
              e = clip["source_end"]
              new_s, new_e = s, e
              event = {
                  "clip_id": clip["clip_id"],
                  "source_id": clip.get("source_id"),
                  "original_start": round(s, 3),
                  "original_end": round(e, 3),
                  "start_action": "kept",
                  "end_action": "kept",
              }
              # Opening: a shot-change just AFTER source_start leaves an old-shot sliver before it.
              start_changes = [
                  c
                  for c in _detect_shot_changes(video, s, min(e, s + margin), threshold)
                  if c > s + 1e-3
              ]
              if start_changes:
                  cand = max(start_changes)  # open after the last rapid cut in the window
                  if cand < e - min_keep:
                      new_s = round(cand, 3)
                      event["start_action"] = "moved_forward"
                      event["new_start"] = new_s
                  else:
                      event["start_unsnapped_reason"] = "collapse"
              # Closing: a shot-change just BEFORE source_end leaves a next-shot sliver after it.
              end_changes = [
                  c
                  for c in _detect_shot_changes(video, max(new_s, e - margin), e, threshold)
                  if c < e - 1e-3
              ]
              if end_changes:
                  cand = min(end_changes)  # close before the first rapid cut in the window
                  if cand > new_s + min_keep:
                      new_e = round(cand, 3)
                      event["end_action"] = "moved_back"
                      event["new_end"] = new_e
                  else:
                      event["end_unsnapped_reason"] = "collapse"
              n_start += new_s != s
              n_end += new_e != e
              clip["source_start"] = new_s
              clip["source_end"] = new_e
              events.append(event)
      
          if n_start or n_end:
              log(
                  f"避让原片切镜头: {n_start} 个起点前移、{n_end} 个终点回收 (margin={margin}s, 阈值={threshold})"
              )
          return _plan_with_snapped_clips(plan, clips, "shot_snaps", events)
      
      
      def _load_silence_for_source(work_dir, source_id, source_work_dir=None):
          """Read a source's silence_periods.json (project layout when source_id is None); [] when absent."""
          path = _find_source_artifact(work_dir, "silence_periods.json", source_id, source_work_dir)
          return json.loads(path.read_text(encoding="utf-8")) if path else []
      
      
      def snap_multi_source_clips(
          plan,
          sources,
          work_dir,
          *,
          line_max_extend,
          scene_margin,
          scene_threshold,
          start_max_prepend,
          start_max_trim,
          do_line_snap=True,
          do_scene_snap=True,
      ):
          """Per-source line/shot snapping for a multi-source validated plan.
      
          Each clip is snapped using ITS OWN source's silence windows / shot changes and duration
          (a clip in source B never constrains a clip in source A), then the global OUTPUT timeline
          is recomputed once in plan order. Missing silence data leaves a boundary unchanged.
          Mirrors the single-source snap_clip_ends_to_lines + snap_clips_off_shot_changes.
          """
          clips = plan["clips"]
          allow_overlap = plan["allow_overlap"]
          groups = {}
          for clip in clips:
              groups.setdefault(clip["source_id"], []).append(clip)
          boundary_accum = {
              "start_snaps": [],
              "end_snaps": [],
              "shot_snaps": [],
              "sentence_checks": [],
          }
          blocking_accum = []
          for sid, group in groups.items():
              source = sources[sid]
              duration = source["duration"]
              mini = {"clips": [dict(c) for c in group], "allow_overlap": allow_overlap}
              # Visual cleanup goes first. Sentence/quiet snapping is the final authority because
              # a clean picture is never allowed to reintroduce a mid-sentence audio cut.
              if do_scene_snap:
                  mini = snap_clips_off_shot_changes(
                      mini, source["source_path"], margin=scene_margin, threshold=scene_threshold
                  )
              source_work_dir = source.get("source_work_dir")
              boundaries = _combine_boundary_windows(
                  _load_silence_for_source(work_dir, sid, source_work_dir),
                  _load_sentence_boundary_windows(work_dir, sid, source_work_dir),
              )
              if do_line_snap:
                  mini = snap_clip_starts_to_lines(
                      mini, boundaries, duration, start_max_prepend, max_trim=start_max_trim
                  )
                  mini = snap_clip_ends_to_lines(mini, boundaries, duration, line_max_extend)
              mini = enforce_clip_sentence_boundaries(
                  mini,
                  boundaries,
                  _load_source_speech_spans(work_dir, sid, source_work_dir),
                  duration,
              )
              mini_boundary = mini["qc"]["boundary_status"]
              for key, events in boundary_accum.items():
                  events.extend(mini_boundary.get(key, []))
              blocking_accum.extend(mini["qc"].get("blocking", []))
              for original, snapped in zip(group, mini["clips"]):
                  original["source_start"] = snapped["source_start"]
                  original["source_end"] = snapped["source_end"]
          # Recompute the global output timeline cursor-based, in plan order (not group order).
          plan["total_duration"] = _recompute_clip_timeline(clips)
      
          qc = plan.setdefault("qc", {})
          boundary = qc.setdefault("boundary_status", {})
          for key, events in boundary_accum.items():
              if events:
                  boundary.setdefault(key, []).extend(events)
          warnings = qc.setdefault("warnings", [])
          for event in boundary_accum["start_snaps"]:
              if "warning_code" in event:
                  warnings.append(
                      {
                          "code": event["warning_code"],
                          "clip_id": event["clip_id"],
                          "source_id": event["source_id"],
                          "start_unsnapped_reason": event["start_unsnapped_reason"],
                      }
                  )
          # A mid-sentence cut must never be dropped because some sibling QC list came back empty.
          if blocking_accum:
              qc.setdefault("blocking", []).extend(blocking_accum)
          return plan
      
    • shot_review.py 17 KB
      #!/usr/bin/env python3
      """Read-only internal-shot recall on the actual rendered video; never repair an EDL."""
      import argparse
      from bisect import bisect_left
      from fractions import Fraction
      import json
      import math
      import os
      from pathlib import Path
      import re
      import subprocess
      import tempfile
      
      from cut_contract import edited_source_render_cache_payload
      from lib import file_identity
      
      
      def _positive(value, name, *, integer=False):
          if isinstance(value, bool) or not isinstance(value, (int, float)):
              raise ValueError(f"{name} must be a positive number")
          if not math.isfinite(value) or value <= 0 or (integer and not isinstance(value, int)):
              raise ValueError(f"invalid {name}")
      
      
      def _validate_clock(pts, end):
          if not pts or any(not isinstance(p, Fraction) for p in [*pts, end]):
              raise ValueError("frame clock requires exact rational PTS")
          if pts[0] != 0 or any(a >= b for a, b in zip(pts, pts[1:])) or end <= pts[-1]:
              raise ValueError("invalid/nonmonotonic frame clock")
      
      
      def summarize_candidates(pts, end, cut_frames, *, max_short_frames=None,
                               max_short_seconds=1.0, dense_window_seconds=2.0,
                               min_dense_cuts=4):
          """Half-open frame spans including the head and tail; thresholds are recall policy."""
          _validate_clock(pts, end)
          _positive(max_short_seconds, "max_short_seconds")
          if max_short_frames is None:
              # A fixed frame count only coincides with the seconds limit at one frame rate, so
              # derive it from the measured clock; --max-short-frames stays an explicit override.
              measured_fps = Fraction(len(pts), 1) / end
              max_short_frames = max(1, round(measured_fps * Fraction(str(max_short_seconds))))
          _positive(max_short_frames, "max_short_frames", integer=True)
          _positive(dense_window_seconds, "dense_window_seconds")
          _positive(min_dense_cuts, "min_dense_cuts", integer=True)
          if any(type(f) is not int or not 0 < f < len(pts) for f in cut_frames):
              raise ValueError("invalid candidate frame index")
          cuts = sorted(set(cut_frames))
          boundaries = [0, *cuts, len(pts)]
          clock = [*pts, end]
          short = []
          for start, stop in zip(boundaries, boundaries[1:]):
              duration = clock[stop] - clock[start]
              if stop - start <= max_short_frames and duration <= Fraction(str(max_short_seconds)):
                  short.append({
                      "start_frame": start, "end_frame": stop, "frame_count": stop - start,
                      "start_exact": str(clock[start]), "end_exact": str(clock[stop]),
                      "duration_exact": str(duration), "duration_seconds": float(duration),
                      "review_status": "NEEDS_DYNAMIC_REVIEW",
                  })
          windows = []
          right = 0
          window_duration = Fraction(str(dense_window_seconds))
          for left in range(len(cuts)):
              right = max(right, left)
              while right < len(cuts) and pts[cuts[right]] - pts[cuts[left]] <= window_duration:
                  right += 1
              if right - left < min_dense_cuts:
                  continue
              members = cuts[left:right]
              if windows and members[0] <= windows[-1]["cut_frames"][-1]:
                  windows[-1]["cut_frames"] = sorted(set(windows[-1]["cut_frames"]) | set(members))
              else:
                  windows.append({"cut_frames": members, "review_status": "NEEDS_DYNAMIC_REVIEW"})
          for window in windows:
              window.update({"start_exact": str(pts[window["cut_frames"][0]]),
                             "end_exact": str(pts[window["cut_frames"][-1]])})
          return {
              "status": "NEEDS_REVIEW" if short or windows else "NO_CANDIDATES",
              "normal_speed_review": "NOT_CHECKED", "automatic_repairs": [],
              "candidates": [{"frame": f, "pts_exact": str(pts[f]), "origin": "UNKNOWN"} for f in cuts],
              "short_spans": short, "dense_windows": windows,
              "policy": {"max_short_frames": max_short_frames, "max_short_seconds": max_short_seconds,
                         "dense_window_seconds": dense_window_seconds, "min_dense_cuts": min_dense_cuts},
          }
      
      
      def probe_frame_clock(video):
          result = subprocess.run([
              "ffprobe", "-v", "error", "-select_streams", "v:0", "-show_frames",
              "-show_entries", "stream=time_base,start_pts,duration_ts:frame=pts,duration,pkt_duration",
              "-of", "json", str(video),
          ], capture_output=True, text=True, timeout=600)
          if result.returncode or result.stderr.strip():
              raise RuntimeError(f"frame decode probe failed: {result.stderr[-2000:]}")
          try:
              data = json.loads(result.stdout)
              stream = data["streams"][0]
              tb = Fraction(stream["time_base"])
              if tb <= 0:
                  raise ValueError("nonpositive time base")
              frames = data["frames"]
              absolute = [int(f["pts"]) * tb for f in frames]
              origin = absolute[0]
              pts = [p - origin for p in absolute]
              last_duration = frames[-1].get("duration", frames[-1].get("pkt_duration"))
              if last_duration is None or int(last_duration) <= 0:
                  raise ValueError("missing/invalid terminal frame duration")
              decoded_end = pts[-1] + int(last_duration) * tb
              if "duration_ts" in stream:
                  end = (int(stream.get("start_pts", frames[0]["pts"])) + int(stream["duration_ts"])) * tb - origin
                  if abs(end - decoded_end) > tb:
                      raise ValueError("decoded frame coverage does not reach declared stream end")
              else:
                  end = decoded_end
              _validate_clock(pts, end)
          except (KeyError, IndexError, ValueError, TypeError, ZeroDivisionError) as exc:
              raise ValueError(f"cannot establish exact frame clock: {exc}") from exc
          return pts, end, origin
      
      
      def _validate_scene_roi(roi, video):
          """Validate an ROI (four ints from the CLI) against the video's auto-oriented native frame."""
          if len(roi) != 4:
              raise ValueError("scene ROI must contain exactly four integers")
          x, y, width, height = roi
          if x < 0 or y < 0 or width <= 0 or height <= 0:
              raise ValueError("scene ROI requires x/y >= 0 and width/height > 0")
      
          from media_geometry import _probe_video_geometry
          facts = _probe_video_geometry(video).facts
          rotation = facts["rotation"]
          if rotation not in {0, 90, 180, 270}:
              raise ValueError("scene ROI requires a right-angle video rotation")
          canvas_width, canvas_height = (
              (facts["coded_height"], facts["coded_width"]) if rotation in {90, 270}
              else (facts["coded_width"], facts["coded_height"])
          )
          if x + width > canvas_width or y + height > canvas_height:
              raise ValueError(
                  f"scene ROI exceeds auto-oriented native frame {canvas_width}x{canvas_height}"
              )
      
      
      def detect_scene_pts(video, threshold, roi=None):
          """No seek or float pts_time: showinfo integer pts + its actual filter timebase.
      
          The single library-side check of threshold/ROI; the CLIs pre-check with parser.error."""
          if isinstance(threshold, bool) or not isinstance(threshold, (int, float)) or not math.isfinite(threshold) or not 0 <= threshold <= 1:
              raise ValueError("scene threshold must be finite and in [0,1]")
          filters = []
          if roi is not None:
              _validate_scene_roi(roi, video)
              x, y, width, height = roi
              filters.append(f"crop={width}:{height}:{x}:{y}:exact=1")
          filters.extend([f"select='gt(scene,{threshold})'", "showinfo"])
          result = subprocess.run([
              "ffmpeg", "-hide_banner", "-nostdin", "-v", "info", "-xerror", "-copyts",
              "-threads", "2", "-i", str(video), "-map", "0:v:0",
              "-vf", ",".join(filters),
              "-an", "-fps_mode", "passthrough", "-f", "null", "-",
          ], capture_output=True, text=True, timeout=600)
          if result.returncode:
              raise RuntimeError(f"scene decode failed: {result.stderr[-2000:]}")
          bases = set(re.findall(r"config in time_base:\s*([0-9]+/[0-9]+)", result.stderr))
          if len(bases) != 1:
              raise ValueError("ambiguous or missing scene filter timebase")
          tb = Fraction(bases.pop())
          if tb <= 0:
              raise ValueError("invalid scene filter timebase")
          values = re.findall(r"\bn:\s*\d+\s+pts:\s*(-?\d+)\s+pts_time:", result.stderr)
          return [int(p) * tb for p in values]
      
      
      def load_bound_plan(video, plan_path):
          """Only associate with the current cut-render cache: plan, render settings, sources."""
          from cut_contract import _edited_source_meta_path
          try:
              plan = json.loads(Path(plan_path).read_text(encoding="utf-8"))
              meta = json.loads(_edited_source_meta_path(video).read_text(encoding="utf-8"))
              sources = meta["sources"]
              if not isinstance(sources, dict) or not sources:
                  raise ValueError("source identities missing")
              if not Path(video).is_file() or Path(video).stat().st_size == 0:
                  raise ValueError("rendered media missing")
              if (meta["plan"] != plan["clips"]
                      or meta["render_cache"] != edited_source_render_cache_payload()
                      or any(file_identity(p) != identity for p, identity in sources.items())):
                  raise ValueError("stale media, plan or render settings")
              declared = {c["source_path"] for c in plan["clips"] if "source_path" in c}
              if declared and declared != set(sources):
                  raise ValueError("source set mismatch")
              if not declared and len(sources) != 1:
                  raise ValueError("ambiguous single source")
          except (OSError, KeyError, ValueError, TypeError) as exc:
              raise ValueError(f"cut plan binding failed: {exc}") from exc
          return plan
      
      
      def _associate_plan(report, plan, pts, end):
          """Legacy seconds are only a coarse locator, not exact source-frame proof."""
          spans = []
          previous = Fraction(0)
          for clip in plan["clips"]:
              start, stop = Fraction(str(clip["output_start"])), Fraction(str(clip["output_end"]))
              if start != previous or stop <= start:
                  raise ValueError("bound cut plan must be contiguous and ordered")
              spans.append((start, stop, clip))
              previous = stop
          tolerance = max(b - a for a, b in zip(pts, [*pts[1:], end]))
          if not spans or abs(previous - end) > tolerance:
              raise ValueError("bound plan duration disagrees with actual frame clock")
          joins = [bisect_left(pts, s[0]) for s in spans[1:]]
          for candidate in report["candidates"]:
              f, t = candidate["frame"], pts[candidate["frame"]]
              if any(abs(f - join) <= 1 for join in joins):
                  candidate["origin"] = "EDIT_JOIN_CANDIDATE"
              else:
                  for start, stop, clip in spans:
                      if start < t < stop:
                          candidate["inside_bound_clip"] = clip["clip_id"]
                          candidate["source_time_estimate"] = str(Fraction(str(clip["source_start"])) + t - start)
                          candidate["mapping_precision"] = "legacy_seconds_estimate"
                          # No source scan: motion, exposure, overlay animation and native cuts
                          # remain indistinguishable from output scene score alone.
                          break
          report["plan_joins"] = joins
      
      
      def scan_video(video, *, threshold=0.35, plan_path=None, roi=None, **policy):
          video = Path(video).resolve()
          plan = load_bound_plan(video, plan_path) if plan_path is not None else None
          pts, end, origin = probe_frame_clock(video)
          scenes = detect_scene_pts(video, threshold, roi)
          frames_by_pts = {p + origin: i for i, p in enumerate(pts)}
          if any(p not in frames_by_pts for p in scenes):
              raise ValueError("scene candidate does not match a unique decoded frame PTS")
          report = summarize_candidates(pts, end, [frames_by_pts[p] for p in scenes if frames_by_pts[p] > 0], **policy)
          if plan is not None:
              _associate_plan(report, plan, pts, end)
          report.update({
              "schema_version": 1, "artifact": "shot_review", "algorithm": "scene-frame-recall-v1", "scan_complete": True,
              "media": {"path": str(video), "frame_count": len(pts),
                        "origin_pts_exact": str(origin), "duration_exact": str(end)},
              "plan_binding": {"path": str(Path(plan_path).resolve())} if plan is not None else None,
              "scene_threshold": threshold,
              "scene_roi": None if roi is None else list(roi),
              "limits": ["scene score is not a confirmed shot or flash-frame defect",
                         "source-origin confirmation requires independent source footage review",
                         "NO_CANDIDATES is not perceptual approval or listening evidence"],
          })
          return report
      
      
      def _atomic_json(path, data):
          path = Path(path)
          path.parent.mkdir(parents=True, exist_ok=True)
          temp = None
          try:
              with tempfile.NamedTemporaryFile("w", encoding="utf-8", dir=path.parent, delete=False) as handle:
                  temp = Path(handle.name)
                  json.dump(data, handle, ensure_ascii=False, indent=2)
                  handle.flush()
                  os.fsync(handle.fileno())
              os.replace(temp, path)
          finally:
              if temp is not None:
                  temp.unlink(missing_ok=True)
      
      
      def _report_target(video, output, plan_path):
          """Unknown existing files are never report targets, even with corrupt source metadata."""
          from cut_contract import _edited_source_meta_path
          protected = {Path(video).resolve(), _edited_source_meta_path(video).resolve()}
          if plan_path is not None:
              protected.add(Path(plan_path).resolve())
          target = Path(output).resolve()
          if target in protected:
              raise ValueError("report must not overwrite media, source, plan or render metadata")
          if Path(output).exists():
              try:
                  old = json.loads(Path(output).read_text(encoding="utf-8"))
                  if (not isinstance(old, dict) or type(old.get("schema_version")) is not int
                          or old["schema_version"] != 1 or old.get("artifact") != "shot_review"):
                      raise ValueError("not a shot-review artifact")
              except (OSError, UnicodeError, ValueError) as exc:
                  raise ValueError("report must not overwrite an unknown existing file") from exc
          return target
      
      
      def _protect_declared_sources(video, plan_path, target):
          """Use both plan and metadata declarations, even when they disagree."""
          from cut_contract import _edited_source_meta_path
          if plan_path is None:
              return
          paths, errors = set(), []
          for path, kind in [(Path(plan_path), "plan"), (_edited_source_meta_path(video), "meta")]:
              try:
                  data = json.loads(path.read_text(encoding="utf-8"))
                  if kind == "plan":
                      paths.update(c["source_path"] for c in data["clips"] if "source_path" in c)
                  else:
                      paths.update(data["sources"])
              except (OSError, ValueError, KeyError, TypeError) as exc:
                  errors.append(exc)
          if target in {Path(p).resolve() for p in paths}:
              raise UnsafeReportTarget("report must not overwrite a declared source")
          if errors:
              raise ValueError(f"cannot verify declared source protection: {errors[0]}")
      
      
      class UnsafeReportTarget(ValueError):
          """Do not write even failure evidence to an input file."""
      
      
      def write_scan(video, output, **options):
          target = _report_target(video, output, options.get("plan_path"))
          roi = options.get("roi")
          base = {"schema_version": 1, "artifact": "shot_review", "scan_complete": False,
                  "normal_speed_review": "NOT_CHECKED",
                  "scene_threshold": options.get("threshold", 0.35),
                  "scene_roi": None if roi is None else list(roi)}
          try:
              _protect_declared_sources(video, options.get("plan_path"), target)
              _atomic_json(output, {**base, "status": "SCANNING"})
              report = scan_video(video, **options)
          except UnsafeReportTarget:
              raise
          except (OSError, ValueError, RuntimeError, KeyError, TypeError, subprocess.SubprocessError) as exc:
              _atomic_json(output, {**base, "status": "SCAN_FAILED", "error": str(exc)})
              raise
          _atomic_json(output, report)
          return report
      
      
      def main():
          parser = argparse.ArgumentParser(description=__doc__)
          parser.add_argument("video")
          parser.add_argument("--output", required=True)
          parser.add_argument("--plan", default=None, help="optional current clip_plan_validated.json; stale bindings fail")
          parser.add_argument("--threshold", type=float, default=0.35)
          parser.add_argument("--roi", nargs=4, type=int, metavar=("X", "Y", "WIDTH", "HEIGHT"))
          parser.add_argument("--max-short-frames", type=int, default=None,
                              help="explicit frame cap; default derives round(fps * max_short_seconds)")
          parser.add_argument("--max-short-seconds", type=float, default=1.0)
          parser.add_argument("--dense-window-seconds", type=float, default=2.0)
          parser.add_argument("--min-dense-cuts", type=int, default=4)
          args = parser.parse_args()
          report = write_scan(args.video, args.output, threshold=args.threshold, plan_path=args.plan, roi=args.roi,
                              max_short_frames=args.max_short_frames, max_short_seconds=args.max_short_seconds,
                              dense_window_seconds=args.dense_window_seconds, min_dense_cuts=args.min_dense_cuts)
          print(json.dumps({"status": report["status"], "short_spans": len(report["short_spans"]),
                            "dense_windows": len(report["dense_windows"]), "report": args.output}))
      
      
      if __name__ == "__main__":
          main()
      
  • SKILL.md 7.1 KB
    ---
    name: video-cut
    user-invocable: false
    description: >
     把长视频按 Agent 选择的原片区间剪成短片。作为两阶段创作流程中的剪辑环节,读取 clip_plan.json 与源视频,
     输出 edited_source.mp4;随后 Agent 按输出时间线写 narration.json。支持单视频与多视频(sources manifest)拼剪,
     本工具不读取、不映射旁白。
     触发词:视频剪辑、剪辑式解说、video cut、clip plan、拼剪。
    ---
    
    ## 1. 定位
    
    本技能只执行 Agent 已经做出的剪辑决定:
    
    1. 校验并补全 `clip_plan.json`,写出带 `clip_id`、原片/输出时间与时长的 `clip_plan_validated.json`。
    2. 先避开原片硬切附近的闪帧风险,最后把边界吸附到可靠句末/自然停顿;声音完整性拥有最终优先级。
    3. 拼接选定区间,输出 `edited_source.mp4`。
    4. 到此停止,由 Agent 按真实输出时间线写 `narration.json`;本工具不读取旁白,也不做原片→输出映射。
    
    相同输入会得到相同输出。`edited_source.mp4.meta.json` 记录标准化 clips、渲染设置和每个源文件的 `size`/`mtime_ns`;三者与当前一致且 `edited_source.mp4` 存在非空才复用,任一不同即重渲染。只有 sidecar 而没有媒体文件不复用。
    
    ## 2. 输入契约
    
    `work_dir/clip_plan.json` 可以是数组,也可以是 `{"clips": [...]}`:
    
    ```json
    {"start": 12.0, "end": 28.5, "reason": "b02 | turn | power: A→B | POV=女主 | 保留反应 | 入点=问题落下 | 出点=沉默结束"}
    ```
    
    - `start` / `end` 是原片秒数;也接受 `source_start` / `source_end` 或 `in` / `out`。
    - 顶层可选 `target_duration`,例如 `"10m"`。
    - 多视频项目的每个片段还必须填写 `source_id`。
    - `speech_boundary_anchors.json` 与 ASR 时间段由理解阶段提供;Agent 先写大致区间,工具会尝试吸附并把仍在讲话区间内的入/出点作为 blocker 返回。
    
    ## 3. 剪辑意图契约
    
    工具不会替 Agent 做创作选择。写片段前先完成本节的剪辑意图检查,并让每个区间映射到 `recap_story_plan.json` 的一个 beat。
    
    使用现有自由文本 `reason` 保存简洁决定:
    
    ```text
    beat_id | function | change | POV | preferred moment | 入点 reason | 出点 reason
    ```
    
    不要因为“事件重要”就保留整段;要保留最能让 change 成立的具体表演、反应、动作或揭示。理解与情绪允许时晚进早出,同时保证台词、动作和技术边界完整。
    
    对不能删去的问答、反应或动作兑现,先核源证据,再在同一 `clip_plan.json` 登记精确区间:
    
    ```json
    {
      "clips": [{"start": 12, "end": 18}],
      "required_evidence": {
        "nodes": [
          {"id": "refusal", "source": "/media/episode.mp4", "start": 12.25, "end": 14.5, "track": "audio", "content": "对方拒绝请求"},
          {"id": "response", "source": "/media/episode.mp4", "start": 15, "end": 17.5, "track": "video", "content": "听到拒绝后的反应与决定"}
        ],
        "before": [["refusal", "response"]]
      }
    }
    ```
    
    `source` 使用实际源文件绝对路径,`start/end` 是原片秒;多源可另填 `source_id` 消歧。只登记确实需要保留的具体时刻,不将整个 beat 默认锁死。`before` 只登记本片必需的先后关系;无需约束顺序时写 `before: []`。
    
    工具在全部画面/句界吸附后检查每个必保时刻至少有一处完整连续保留、来源和先后;音频节点还检查源音轨是否存在。每次结果出现(包括局部片段)都需满足其声明的前提,不能用后面的完整段替开头缺前提的片段过关。结果写入 `clip_plan_validated.json.qc.required_evidence`;缺段、错序或无效声明会在预检、缓存复用和渲染前阻断,时长放宽选项不会跳过。该结果验证选段保留,实际语义与最终混音仍按审片步骤核对。
    
    下面的 `scripts/...` 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。
    
    ## 4. 运行命令
    
    ```bash
    python3 scripts/cut.py <video> --work-dir <work_dir> \
      [--target-duration 10m] [--clip-padding 0] [--allow-overlap]
    ```
    
    ## 5. 输出契约
    
    - `clip_plan_validated.json`:标准化片段,包含 `clip_id`、`source_start/end`、`output_start/end` 与 `duration`。
    - `edited_source.mp4`:按计划拼接后的短视频。
    - `shot_review.json`:仅 `--review-shots` 开启后生成的实际视频短镜/密集切镜候选;不会更改计划。
    
    下游把 `edited_source.mp4` 当作视频,把 Agent 按输出时间写的 `narration.json` 当作旁白。
    
    ## 6. 边界与时间线规则
    
    - `clip_plan.json` 使用原片时间;`narration.json` 直接使用剪后输出时间,不存在原片 → 输出的旁白映射。
    - 默认禁止重叠或重复原片区间;`--allow-overlap` 开启后才允许。
    - 片段起点只能位于源头、可靠句末/静音窗,或与上一片段构成无损同源连续连接;片段终点同理。ASR 判定仍在讲话且无法吸附时写入 `unsafe_clip_sentence_boundary` 并阻断。
    - `SCENE_CUT_SNAP` 默认开启:先按画面把 source start 向后、source end 向前吸附到附近硬切,随后句末吸附再做最终修正,避免视觉修正重新制造半句原声。默认范围为 `SCENE_CUT_SNAP_MARGIN=0.5` 秒,检测阈值为 `SCENE_CUT_DETECT_THRESHOLD=0.4`。
    - scene-change score 只提供接点候选,不证明接点自然。先检查短时间窗内是否出现密集候选,再区分来源:原片自带的无关短镜头整段删除;相关但短到像闪帧的镜头通过扩展 IN/OUT 保留完整动作、反应或台词,不用定格/慢放伪造时长;由本次拼接制造的切点则优先移动边界、恢复同源连续运动、合并相邻片段或改用更自然的连接,尽量消除。成片后仍要逐个播放接点前后约 0.5–1 秒;白闪或曝光叠化再结合逐帧亮度定位,不能为了通过视觉检测切断完整台词,也不能用转场遮掩坏接点。
    - 修短残镜时不得仅为压低 scene 分数而对接点附近施加与所属镜头不连续的极端放大或位移;取景复核与修复验证流程见 `references/shot-review.md`。
    - 连续同源片段的无损连接不做句中双侧音频淡出;非连续片段仍在安全停顿内做防爆音淡入淡出。
    
    需要检查短时间频繁切镜时,先用 ffmpeg scene filter 召回候选时间:
    
    ```bash
    ffmpeg -i input.mp4 -vf "select='gt(scene,0.35)',showinfo" -an -f null -
    ```
    
    `0.35` 是起始阈值,不是质量判据;大幅运动、闪白和叠化都可能误报。把候选映射回原片 shot 与本次拼接边界后,按上面的来源分类处理,并以正常速度播放决定是否保留。
    
    需要精确到实际帧、检查长区间内部残镜并记录所用计划路径时,使用
    `scripts/shot_review.py` 或 `cut.py --review-shots`(有黑边或包装时加 `--roi` / `--shot-roi`);
    详见 `references/shot-review.md`。
    
    ## 7. 能力边界
    
    - 不做语义理解,不写旁白,不替 Agent 选择片段;scene filter 只承担技术边界候选检测。
    - 只做生成 `edited_source.mp4` 所需的剪切、拼接与一次中间编码,不承担字幕包装或最终交付压缩。
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related