Claude Skill

video-assemble

合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录, 最后做响度标准化。作为最终合成阶段使用。输入源视频、tts_meta.json 与旁白位置; 输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download zenstory-ai-oh-story-dsh-packages_knowledge_video-recap_skills_video-assemble-d734089.zip · 158 KB
Part of zenstory-ai/oh-story-dsh — 31 skills

Install

skills CLI npx skills add https://github.com/zenstory-ai/oh-story-dsh/tree/main/packages/knowledge/video-recap/skills/video-assemble
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install zenstory-ai-oh-story-dsh@llmmart
Git git clone https://github.com/zenstory-ai/oh-story-dsh.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole zenstory-ai/oh-story-dsh collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

1. 定位

本技能负责最终合成:

  1. 把各段旁白音频放到视频时间线上。
  2. 在旁白窗口内压低原声,支持 fixed / sidechain / zone 模式。
  3. 根据旁白位置生成 subtitles.srt;默认同时生成并烧录 subtitles.ass,--no-burn-subtitles 可关闭。
  4. 可选把最终响度标准化到目标 LUFS。

2. 声音收尾契约

合成阶段只实现创作决定,不凭空制造决定。Agent 在写旁白位置前,已在 visual_audio_board.json 为每个 beat 指定 audio_owner:

  • original_dialogue
  • action_sound
  • ambience / music
  • silence
  • narration

因此,旁白间隙是主动选择,不是必须填满的空白。不要为了“更满”而加入通用 BGM、压住必须听见的台词或消除有意义的沉默。

当前渲染器不解析 visual_audio_board.json;Agent 通过旁白时间、overlaps_speech、原声留白与现有混音参数落实这些决定。

3. 输入契约

  • <video>:源视频;cut 模式下为 edited_source.mp4。
  • work_dir/tts_meta.json:默认 narration 模式必需;配音阶段写出的 {segments: [...]}。每段包含 audio_path、时间、pause_after_ms、overlaps_speech 和用于混音/字幕的位置。显式 source-mix / adopted-packet-copy 模式不读取它。
  • 已采用的配音使用显式 --tts-meta 和 --narration-adoption:后者由调用方独立确认文字、请求的引擎/声线和速度策略,不能从待消费元数据自动“批准”出来。完整格式与记录边界见 references/narration-adoption.md。
  • 已采用的完整声音底轨与逐段配音可再传 --audio-mix-adoption;严格格式、48 kHz 声道矩阵和双 binding 事务见 references/explicit-audio-mix.md。

下面的 scripts/... 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。

4. 运行命令

python3 scripts/assemble.py <video> --work-dir <work_dir> \
  [--audio-mode narration|source-mix|adopted-packet-copy] [--audio-stream-index <N>] \
  [--tts-meta <tts_meta.json> --narration-adoption <narration_adoption.json>] \
  [--audio-mix-adoption <audio_mix_adoption.json>] \
  [--recap-stem <name>] [--output-dir <dir>] [--no-burn-subtitles] \
  [--subtitle-y-top <inclusive-y> --subtitle-y-bot <exclusive-y>] \
  [--source-video <orig.mp4>] [--export-jianying [--jianying-out <dir>]]

5. 输出契约

  • recap_<stem>.mp4:稳定的最终输出别名;每次运行覆盖更新。
  • work_dir/output.mp4:工作目录内成片。
  • subtitles.srt:旁白字幕;烧录时另有 subtitles.ass。
  • timeline.json:后端无关的多轨模型,包含视频、原声、旁白、BGM、字幕和 ducking 自动化。
  • _placed_*.wav:实际写入主混音的完整逐段旁白 PCM;时间线与剪映只引用这些文件。
  • narration_input_binding.json:旁白输入、转换、实际放置、旁白总轨和最终音轨的消费记录(路径、PCM 参数、packet 计数)。区分未采用与已绑定采用决定两种状态;不等于声线鉴定或听审。
  • audio_mix_binding.json:显式完整声音分支消费的画面时钟、底轨、48 kHz 配音放置、premaster、固定 master gain、最终 PCM/AAC 事实与 narration binding 路径的记录。
  • assembly_manifest.json:输入来源、cut 来源标识(路径、大小、mtime)、渲染设置与最终输出路径。
  • assembly_qc.json:旁白完整性、原声句末交接、时间线素材时长与交付质量的发布门禁。
  • 剪映草稿目录:仅 --export-jianying 时生成,包含 draft_content.json、draft_info.json 与 draft_meta_info.json。

6. 合成规则

  • 音频模式的处理与冻结语义见 references/audio-modes.md。默认仍为 narration;另外两种模式必须显式选择。
  • --audio-mix-adoption 只与显式 --tts-meta、--narration-adoption 同时使用;它保留 narration 模式名,但跳过旧速度/适配、原声 handoff、环境 BGM、duck、loudnorm 和 limiter。
  • 音频按轨道混合:原声、可选 BGM 与旁白各自独立。
  • 旁白不做任何容差裁尾;温和加速后仍放不下即 no_safe_fit。每段 _placed_*.wav 必须与序列化后的时间线区间等长或更短,否则 timeline_audio_mismatch 阻断。
  • 已采用配音的 v1 合同只支持原速、禁止段内适速;不能让环境默认 1.15 倍速或旧缓存覆盖它。放不下就修订安排,不裁尾。严格运行使用新工作目录与新输出路径;输入/实际混音来源变动或 QC 失败时,不发布候选成片。没有采用文件的旧入口仍是兼容模式,不自动获得同等证据。
  • 原声在旁白结束后保持压低到下一可靠句末的 pause_start,只在实测停顿内渐强, 于 source_restore_at 回满;无后续锚点时保持压低到时间线末端,而不是放出半句。
  • --export-jianying / EXPORT_JIANYING=1 可把 timeline.json 导出为可编辑草稿。cut 模式应传 --source-video <orig>,让草稿引用真实原片区间。
  • 剪映导出默认把视频、音频与图片复制到 Resources/local/{video,audio,image},保持草稿可搬迁;--jianying-no-bundle-media 只适合原路径始终可访问的情况。
  • 重叠覆盖物会拆到编号轨道;非空目标目录不会覆盖,而会创建编号兄弟目录。
  • 常速、倒放、变换、富文本、转场、蒙版、LUT、绿幕复合草稿及显式特效轨道通过 timeline v2 扩展表达。需要素材包的功能只接受调用方合法提供的离线资源。
  • 剪映草稿引用未烧录的源视频,因此原片硬字幕仍会保留,必要时在剪映内另行遮罩。
  • 字幕外观可用 SUBTITLE_FONT_SIZE、SUBTITLE_MARGIN_V、SUBTITLE_MAX_CHARS 等控制。
  • SUBTITLE_Y_TOP/BOT 把 ASS 基线放到测得的原片字幕区域,坐标为半开 [top, bot);显式遮罩策略下默认 SUBTITLE_MASK_OPACITY=0.6,SOURCE_SUBTITLE_MASK_TIMING=narration。
  • 原声在旁白间隙回到 IDLE_ORIG_VOLUME,旁白下压到 SPEECH_DUCKING_VOLUME;DUCK_FADE_SECONDS 控制过渡。还可配置 DUCKING_MODE、ZONE_DUCKING_VOLUME、FINAL_LOUDNORM 与 TARGET_LUFS。
  • 可通过 BGM_PATH 指定 BGM;它会循环到成片长度,并按 BGM_VOLUME / BGM_DUCKING_VOLUME 混音。不要在没有创作依据时设置通用 BGM。
  • 烧录字幕需要带 subtitles / libass 的 ffmpeg;合成阶段会预检并在缺失时明确失败。
  • 原声留白中的对白字幕优先读取 Agent 校对的 original_subtitles.json;否则保守映射 ASR。只有遮罩覆盖留白或用户字幕明确要求替换时才烧录原声对白,并用 「」 与旁白区分。

按原片区间准备声音,而不是整体压低旧成片

已有多段原声取舍和独立 BGM 决定时,先用 references/source-score.md 的独立 source_score.py 从原片声音流按精确帧区间重建原声轨、音乐轨及两者之和;它只输出 声音底轨和来源回执。要与逐段已采用配音合成,再由调用方提供 references/explicit-audio-mix.md 的严格 adoption;不要将底轨塞入旧入口自动 duck,也不要从含旧解说的成片取整条声音冒充干净原声。

7. 字幕与可选包装

先锁定画面、剪点、旁白和混音,再投入字幕动画或边框包装;字幕样式不能掩盖叙事、剪点或声音问题。 普通交付优先使用现有 ASS 路径;只有用户需要逐 cue 排版、动画或透明图层时,才用项目级代码渲染器, 并按 references/foreground-compose.md 把它生成的 RGBA 序列叠到锁定母版。包装顺序与样帧抽检清单见 references/packaging.md。

8. 能力边界

  • 不生成旁白文字,不合成 TTS,不重新转写视频。
  • 字幕烧录默认开启;关闭时不会重编码绘制字幕区域。

显式输出轴字幕轨的独立合同、完整替换语义和当前边界见 references/subtitle-track.md。

画面回原片重建后,若需保留另一文件中的已采用完整混音,先按 references/pair-media.md 显式配对独立画面与音轨。配对只复制流,不补字幕或片名卡; 后续字幕轨必须重新绑定配对后的容器与 a:0,不能继续沿用旧版本身份。

Files (oh-story-dsh)
  • references
    • jianying
      • empty_draft_meta_info.json 1.7 KB
        {
            "cloud_package_completed_time": "",
            "draft_cloud_capcut_purchase_info": "",
            "draft_cloud_last_action_download": false,
            "draft_cloud_materials": [],
            "draft_cloud_purchase_info": "",
            "draft_cloud_template_id": "",
            "draft_cloud_tutorial_info": "",
            "draft_cloud_videocut_purchase_info": "",
            "draft_cover": "",
            "draft_deeplink_url": "",
            "draft_enterprise_info": {
                "draft_enterprise_extra": "",
                "draft_enterprise_id": "",
                "draft_enterprise_name": "",
                "enterprise_material": []
            },
            "draft_fold_path": "",
            "draft_id": "",
            "draft_is_ai_packaging_used": false,
            "draft_is_ai_shorts": false,
            "draft_is_article_video_draft": false,
            "draft_is_from_deeplink": "false",
            "draft_is_invisible": false,
            "draft_materials": [
                {
                    "type": 0,
                    "value": []
                },
                {
                    "type": 1,
                    "value": []
                },
                {
                    "type": 2,
                    "value": []
                },
                {
                    "type": 3,
                    "value": []
                },
                {
                    "type": 6,
                    "value": []
                },
                {
                    "type": 7,
                    "value": []
                },
                {
                    "type": 8,
                    "value": []
                }
            ],
            "draft_materials_copied_info": [],
            "draft_name": "",
            "draft_new_version": "",
            "draft_removable_storage_device": "",
            "draft_root_path": "",
            "draft_segment_extra_info": [],
            "draft_timeline_materials_size_": 9734196,
            "draft_type": "",
            "tm_draft_cloud_completed": "",
            "tm_draft_cloud_modified": 0,
            "tm_draft_create": 1765203775977,
            "tm_draft_modified": 1765203775977,
            "tm_draft_removed": 0,
            "tm_duration": 249100000
        }
      • empty_jy_combination_segment.json 1.6 KB
        {
          "is_tone_modify": false,
          "enable_adjust_mask": false,
          "keyframe_refs": [],
          "enable_adjust": true,
          "enable_video_mask": true,
          "enable_hsl": false,
          "speed": 1,
          "clip": {
            "rotation": 0,
            "transform": {
              "y": 0,
              "x": 0
            },
            "flip": {
              "vertical": false,
              "horizontal": false
            },
            "scale": {
              "y": 1,
              "x": 1
            },
            "alpha": 1
          },
          "intensifies_audio": false,
          "uniform_scale": {
            "on": true,
            "value": 1
          },
          "responsive_layout": {
            "vertical_pos_layout": 0,
            "target_follow": "",
            "enable": false,
            "horizontal_pos_layout": 0,
            "size_layout": 0
          },
          "enable_hsl_curves": true,
          "desc": "",
          "state": 0,
          "enable_lut": true,
          "reverse": false,
          "target_timerange": {
            "start": 0,
            "duration": 0
          },
          "cartoon": false,
          "enable_color_correct_adjust": false,
          "last_nonzero_volume": 1,
          "enable_smart_color_adjust": false,
          "source": "segmentsourcenormal",
          "render_timerange": {
            "start": 0,
            "duration": 0
          },
          "render_index": 0,
          "group_id": "",
          "enable_color_match_adjust": false,
          "enable_color_curves": true,
          "raw_segment_id": "",
          "volume": 1,
          "visible": true,
          "track_render_index": 0,
          "track_attribute": 0,
          "is_loop": false,
          "id": "",
          "enable_color_wheels": true,
          "extra_material_refs": [],
          "template_id": "",
          "template_scene": "default",
          "material_id": "",
          "common_keyframes": [],
          "digital_human_template_group_id": "",
          "color_correct_alg_result": "",
          "hdr_settings": {
            "mode": 1,
            "intensity": 1,
            "nits": 1000
          },
          "is_placeholder": false,
          "source_timerange": {
            "start": 0,
            "duration": 0
          }
        }
      • empty_jy_combination_video_material.json 2.2 KB
        {
          "category_name": "",
          "aigc_type": "none",
          "duration": 0,
          "category_id": "",
          "beauty_face_auto_preset_infos": [],
          "beauty_body_preset_id": "",
          "picture_from": "none",
          "cartoon_path": "",
          "intensifies_audio_path": "",
          "crop_ratio": "free",
          "source_platform": 0,
          "local_material_from": "",
          "media_path": "",
          "is_unified_beauty_mode": false,
          "width": 0,
          "video_algorithm": {
            "algorithms": [],
            "path": "",
            "story_video_modify_video_config": {
              "is_overwrite_last_video": false,
              "tracker_task_id": "",
              "task_id": ""
            },
            "ai_background_configs": [],
            "ai_in_painting_config": [],
            "gameplay_configs": []
          },
          "type": "video",
          "height": 0,
          "material_id": "",
          "has_audio": true,
          "is_copyright": true,
          "source": 0,
          "intensifies_path": "",
          "crop_scale": 1,
          "live_photo_cover_path": "",
          "reverse_path": "",
          "reverse_intensifies_path": "",
          "picture_set_category_id": "",
          "crop": {
            "upper_right_x": 1,
            "lower_right_y": 1,
            "upper_right_y": 0,
            "upper_left_y": 0,
            "lower_left_y": 1,
            "lower_right_x": 1,
            "lower_left_x": 0,
            "upper_left_x": 0
          },
          "request_id": "",
          "local_material_id": "",
          "picture_set_category_name": "",
          "path": "",
          "is_text_edit_overdub": false,
          "check_flag": 62978047,
          "origin_material_id": "",
          "material_url": "",
          "beauty_face_auto_preset": {
            "preset_id": "",
            "scene": "",
            "name": "",
            "rate_map": ""
          },
          "is_ai_generate_content": false,
          "extra_type_option": 2,
          "id": "",
          "aigc_item_id": "",
          "material_name": "复合片段1",
          "beauty_face_preset_infos": [],
          "team_id": "",
          "local_id": "",
          "matting": {
            "expansion": 0,
            "feather": 0,
            "flag": 0,
            "reverse": false,
            "strokes": [],
            "custom_matting_id": "",
            "has_use_quick_eraser": false,
            "has_use_quick_brush": false,
            "path": "",
            "interactiveTime": [],
            "enable_matting_stroke": false
          },
          "formula_id": "",
          "has_sound_separated": false,
          "live_photo_timestamp": -1,
          "stable": {
            "stable_level": 0,
            "time_range": {
              "start": 0,
              "duration": 0
            },
            "matrix_path": ""
          },
          "aigc_history_id": ""
        }
      • empty_jy_draft.json 5.4 KB
        {
          "draft_cover_path": "",
          "type": "combination",
          "combination_id": "",
          "draft_config_path": "",
          "draft": {
            "version": 360000,
            "new_version": "111.0.0",
            "source": "default",
            "update_time": 0,
            "keyframe_graph_list": [],
            "duration": 86600000,
            "uneven_animation_template_info": {
              "order": "",
              "sub_template_info_list": [],
              "content": "",
              "composition": ""
            },
            "tracks": [],
            "canvas_config": {
              "height": 1080,
              "ratio": "original",
              "width": 1920
            },
            "relationships": [],
            "last_modified_platform": {
              "app_version": "",
              "os_version": "",
              "os": "",
              "hard_disk_id": "",
              "device_id": "",
              "mac_address": "",
              "app_id": 0,
              "app_source": ""
            },
            "color_space": -1,
            "static_cover_image_path": "",
            "render_index_track_mode_on": true,
            "keyframes": {
              "handwrites": [],
              "videos": [],
              "texts": [],
              "audios": [],
              "adjusts": [],
              "stickers": [],
              "filters": [],
              "effects": []
            },
            "platform": {
              "app_version": "9.6.0",
              "os_version": "15.1",
              "os": "mac",
              "hard_disk_id": "9d5cea4f22458a4e59d643d07162d324",
              "device_id": "0b02b9bd41947815545481af8a5bde46",
              "mac_address": "f8e784e2995a6aed158ef584d6245be8",
              "app_id": 3704,
              "app_source": "lv"
            },
            "smart_ads_info": {
              "page_from": "",
              "routine": "",
              "draft_url": ""
            },
            "path": "",
            "id": "AFCFCFAC-468E-46B6-92DC-8F2C66411CA3",
            "name": "",
            "lyrics_effects": [],
            "fps": 30,
            "config": {
              "use_float_render": false,
              "adjust_max_index": 1,
              "multi_language_list": [],
              "multi_language_main": "none",
              "lyrics_taskinfo": [],
              "lyrics_recognition_id": "",
              "system_font_list": [],
              "sticker_max_index": 1,
              "multi_language_mode": "none",
              "material_save_mode": 0,
              "attachment_info": [],
              "video_mute": false,
              "maintrack_adsorb": false,
              "multi_language_current": "none",
              "subtitle_sync": true,
              "subtitle_recognition_id": "",
              "combination_max_index": 1,
              "extract_audio_last_index": 1,
              "subtitle_taskinfo": [],
              "record_audio_last_index": 1,
              "lyrics_sync": true,
              "original_sound_last_index": 1
            },
            "create_time": 0,
            "draft_type": "video",
            "function_assistant_info": {
              "auto_adjust": false,
              "auto_caption": false,
              "enhande_voice": false,
              "auto_adjust_fixed": false,
              "fixed_rec_applied": false,
              "retouch_segid_list": [],
              "retouch": false,
              "enhande_voice_fixed": false,
              "enhance_quality_segid_list": [],
              "deflicker_segid_list": [],
              "auto_caption_template_id": "",
              "color_correction": false,
              "caption_opt": false,
              "audio_noise_segid_list": [],
              "auto_caption_segid_list": [],
              "normalize_loudness_fixed": false,
              "normalize_loudness_segid_list": [],
              "eye_correction": false,
              "enhance_quality_fixed": false,
              "color_correction_fixed_value": 50,
              "normalize_loudness_audio_denoise_segid_list": [],
              "video_noise_segid_list": [],
              "eye_correction_segid_list": [],
              "retouch_fixed": false,
              "auto_adjust_segid_list": [],
              "enhance_voice_segid_list": [],
              "smooth_slow_motion_fixed": false,
              "caption_opt_segid_list": [],
              "color_correction_segid_list": [],
              "smooth_slow_motion": false,
              "smart_segid_list": [],
              "smart_rec_applied": false,
              "color_correction_fixed": false,
              "fps": {
                "den": 1,
                "num": 0
              },
              "auto_adjust_fixed_value": 50,
              "normalize_loudness": false,
              "enhance_quality": false
            },
            "materials": {
              "audio_fades": [],
              "images": [],
              "flowers": [],
              "smart_crops": [],
              "common_mask": [],
              "time_marks": [],
              "shapes": [],
              "vocal_beautifys": [],
              "plugin_effects": [],
              "audios": [],
              "smart_relights": [],
              "video_radius": [],
              "vocal_separations": [],
              "video_effects": [],
              "video_shadows": [],
              "videos": [],
              "manual_deformations": [],
              "video_strokes": [],
              "audio_effects": [],
              "realtime_denoises": [],
              "handwrites": [],
              "texts": [],
              "text_templates": [],
              "hsl_curves": [],
              "tail_leaders": [],
              "manual_beautys": [],
              "material_colors": [],
              "multi_language_refs": [],
              "material_animations": [],
              "chromas": [],
              "speeds": [],
              "audio_pannings": [],
              "sound_channel_mappings": [],
              "ai_translates": [],
              "video_trackings": [],
              "primary_color_wheels": [],
              "placeholders": [],
              "stickers": [],
              "placeholder_infos": [],
              "beats": [],
              "audio_pitch_shifts": [],
              "color_curves": [],
              "green_screens": [],
              "drafts": [],
              "digital_humans": [],
              "effects": [],
              "digital_human_model_dressing": [],
              "audio_balances": [],
              "hsl": [],
              "log_color_wheels": [],
              "canvases": [],
              "loudnesses": [],
              "transitions": [],
              "audio_track_indexes": []
            },
            "is_drop_frame_timecode": false,
            "free_render_index_mode_on": false
          },
          "category_id": "",
          "precompile_combination": false,
          "combination_type": "none",
          "formula_id": "",
          "name": "",
          "id": "",
          "draft_file_path": "",
          "category_name": ""
        }
      • empty_jy_material_video.json 1.8 KB
        {
          "aigc_history_id": "",
          "aigc_item_id": "",
          "aigc_type": "none",
          "audio_fade": null,
          "cartoon_path": "",
          "category_id": "",
          "category_name": "",
          "check_flag": 63487,
          "crop": {
            "lower_left_x": 0.0,
            "lower_left_y": 1.0,
            "lower_right_x": 1.0,
            "lower_right_y": 1.0,
            "upper_left_x": 0.0,
            "upper_left_y": 0.0,
            "upper_right_x": 1.0,
            "upper_right_y": 0.0
          },
          "crop_ratio": "free",
          "crop_scale": 1.0,
          "duration": 0,
          "extra_type_option": 0,
          "formula_id": "",
          "freeze": null,
          "has_audio": true,
          "height": 1920,
          "id": "",
          "intensifies_audio_path": "",
          "intensifies_path": "",
          "is_ai_generate_content": false,
          "is_copyright": false,
          "is_text_edit_overdub": false,
          "is_unified_beauty_mode": false,
          "local_id": "",
          "local_material_id": "",
          "material_id": "",
          "material_name": "",
          "material_url": "",
          "matting": {
            "flag": 0,
            "has_use_quick_brush": false,
            "has_use_quick_eraser": false,
            "interactiveTime": [],
            "path": "",
            "strokes": []
          },
          "media_path": "",
          "object_locked": null,
          "origin_material_id": "",
          "path": "",
          "picture_from": "none",
          "picture_set_category_id": "",
          "picture_set_category_name": "",
          "request_id": "",
          "reverse_intensifies_path": "",
          "reverse_path": "",
          "smart_motion": null,
          "source": 0,
          "source_platform": 0,
          "stable": {
            "matrix_path": "",
            "stable_level": 0,
            "time_range": {
              "duration": 0,
              "start": 0
            }
          },
          "team_id": "",
          "type": "video",
          "video_algorithm": {
            "algorithms": [],
            "complement_frame_config": null,
            "deflicker": null,
            "gameplay_configs": [],
            "motion_blur_config": null,
            "noise_reduction": null,
            "path": "",
            "quality_enhance": null,
            "time_range": null
          },
          "width": 1080
        }
      • empty_jy_meta_material_value.json 358 B
        {
          "duration": 0,
          "height": 0,
          "id": "",
          "md5": "",
          "metetype": "",
          "type": 0,
          "width": 1080,
          "create_time": 0,
          "extra_info": "",
          "file_Path": "",
          "import_time": 0,
          "import_time_ms": 0,
          "item_source": 1,
          "roughcut_time_range": {
            "duration": 0,
            "start": 0
          },
          "sub_time_range": {
            "duration": -1,
            "start": -1
          }
        }
      • empty_jy_project_info.json 3.6 KB
        {
            "canvas_config": {
                "height": 1080,
                "ratio": "original",
                "width": 1920
            },
            "color_space": 0,
            "config": {
                "adjust_max_index": 1,
                "attachment_info": [],
                "combination_max_index": 1,
                "export_range": null,
                "extract_audio_last_index": 1,
                "lyrics_recognition_id": "",
                "lyrics_sync": true,
                "lyrics_taskinfo": [],
                "maintrack_adsorb": true,
                "material_save_mode": 0,
                "multi_language_current": "none",
                "multi_language_list": [],
                "multi_language_main": "none",
                "multi_language_mode": "none",
                "original_sound_last_index": 1,
                "record_audio_last_index": 1,
                "sticker_max_index": 1,
                "subtitle_keywords_config": null,
                "subtitle_recognition_id": "",
                "subtitle_sync": true,
                "subtitle_taskinfo": [],
                "system_font_list": [],
                "video_mute": false,
                "zoom_info_params": null
            },
            "cover": null,
            "create_time": 0,
            "duration": 0,
            "extra_info": null,
            "fps": 30.0,
            "free_render_index_mode_on": false,
            "group_container": null,
            "id": "",
            "keyframe_graph_list": [],
            "keyframes": {
                "adjusts": [],
                "audios": [],
                "effects": [],
                "filters": [],
                "handwrites": [],
                "stickers": [],
                "texts": [],
                "videos": []
            },
            "last_modified_platform": {
                "app_id": 3704,
                "app_source": "lv",
                "app_version": "5.9.5-beta1",
                "device_id": "0b02b9bd41947815545481af8a5bde46",
                "hard_disk_id": "9d5cea4f22458a4e59d643d07162d324",
                "mac_address": "046f91694f9a5ed2c112de702af7131b",
                "os": "mac",
                "os_version": "15.1"
            },
            "materials": {
                "ai_translates": [],
                "audio_balances": [],
                "audio_effects": [],
                "audio_fades": [],
                "audio_track_indexes": [],
                "audios": [],
                "beats": [],
                "canvases": [],
                "chromas": [],
                "color_curves": [],
                "digital_humans": [],
                "drafts": [],
                "effects": [],
                "flowers": [],
                "green_screens": [],
                "handwrites": [],
                "hsl": [],
                "images": [],
                "log_color_wheels": [],
                "loudnesses": [],
                "manual_deformations": [],
                "masks": [],
                "common_mask": [],
                "material_animations": [],
                "material_colors": [],
                "multi_language_refs": [],
                "placeholders": [],
                "plugin_effects": [],
                "primary_color_wheels": [],
                "realtime_denoises": [],
                "shapes": [],
                "smart_crops": [],
                "smart_relights": [],
                "sound_channel_mappings": [],
                "speeds": [],
                "stickers": [],
                "tail_leaders": [],
                "text_templates": [],
                "texts": [],
                "time_marks": [],
                "transitions": [],
                "video_effects": [],
                "video_trackings": [],
                "videos": [],
                "vocal_beautifys": [],
                "vocal_separations": []
            },
            "mutable_config": null,
            "name": "",
            "new_version": "111.0.0",
            "platform": {
                "app_id": 3704,
                "app_source": "lv",
                "app_version": "5.9.5-beta1",
                "device_id": "0b02b9bd41947815545481af8a5bde46",
                "hard_disk_id": "46f90855e144229994b84640e2c31384",
                "mac_address": "046f91694f9a5ed2c112de702af7131b",
                "os": "mac",
                "os_version": "15.1"
            },
            "relationships": [],
            "render_index_track_mode_on": false,
            "retouch_cover": null,
            "source": "default",
            "static_cover_image_path": "",
            "time_marks": null,
            "tracks": [],
            "update_time": 0,
            "version": 360000
        }
      • empty_jy_segment.json 1.3 KB
        {
          "caption_info": null,
          "cartoon": false,
          "clip": {
            "alpha": 1.0,
            "flip": {
              "horizontal": false,
              "vertical": false
            },
            "rotation": 0.0,
            "scale": {
              "x": 1.0,
              "y": 1.0
            },
            "transform": {
              "x": 0.0,
              "y": 0.0
            }
          },
          "common_keyframes": [],
          "enable_adjust": true,
          "enable_color_correct_adjust": false,
          "enable_color_curves": true,
          "enable_color_match_adjust": false,
          "enable_color_wheels": true,
          "enable_lut": true,
          "enable_smart_color_adjust": false,
          "extra_material_refs": [],
          "group_id": "",
          "hdr_settings": {
            "intensity": 1.0,
            "mode": 1,
            "nits": 1000
          },
          "id": "",
          "intensifies_audio": false,
          "is_placeholder": false,
          "is_tone_modify": false,
          "keyframe_refs": [],
          "last_nonzero_volume": 1.0,
          "material_id": "",
          "render_index": 0,
          "responsive_layout": {
            "enable": false,
            "horizontal_pos_layout": 0,
            "size_layout": 0,
            "target_follow": "",
            "vertical_pos_layout": 0
          },
          "reverse": false,
          "source_timerange": {
            "duration": 0,
            "start": 0
          },
          "speed": 1.0,
          "target_timerange": {
            "duration": 0,
            "start": 0
          },
          "template_id": "",
          "template_scene": "default",
          "track_attribute": 0,
          "track_render_index": 0,
          "uniform_scale": {
            "on": true,
            "value": 1.0
          },
          "visible": true,
          "volume": 1.0
        }
      • empty_jy_text_styles.json 440 B
        {
          "styles": [
            {
              "fill": {
                "alpha": 1.0,
                "content": {
                  "render_type": "solid",
                  "solid": {
                    "alpha": 1.0,
                    "color": [
                      1.0,
                      1.0,
                      1.0
                    ]
                  }
                }
              },
              "font": {
                "id": "",
                "path": ""
              },
              "range": [
                0,
                4
              ],
              "size": 15.0
            }
          ],
          "text": ""
        }
        
      • empty_yj_material_audio.json 865 B
        {
          "aigc_history_id": "",
          "aigc_item_id": "",
          "app_id": 0,
          "category_id": "",
          "category_name": "",
          "check_flag": 1,
          "copyright_limit_type": "none",
          "duration": 0,
          "effect_id": "",
          "formula_id": "",
          "id": "",
          "intensifies_path": "",
          "is_ai_clone_tone": false,
          "is_text_edit_overdub": false,
          "is_ugc": false,
          "local_material_id": "",
          "music_id": "",
          "name": "",
          "path": "",
          "query": "",
          "request_id": "",
          "resource_id": "",
          "search_id": "",
          "source_from": "",
          "source_platform": 0,
          "team_id": "",
          "text_id": "",
          "tone_category_id": "",
          "tone_category_name": "",
          "tone_effect_id": "",
          "tone_effect_name": "",
          "tone_platform": "",
          "tone_second_category_id": "",
          "tone_second_category_name": "",
          "tone_speaker": "",
          "tone_type": "",
          "type": "extract_music",
          "video_id": "",
          "wave_points": []
        }
      • empty_yj_material_text.json 2.4 KB
        {
          "add_type": 0,
          "alignment": 1,
          "background_alpha": 1.0,
          "background_color": "",
          "background_height": 0.0,
          "background_horizontal_offset": 0.0,
          "background_round_radius": 0.0,
          "background_style": 0,
          "background_vertical_offset": 0.0,
          "background_width": 0.0,
          "base_content": "",
          "bold_width": 0.0,
          "border_alpha": 1.0,
          "border_color": "",
          "border_width": 0.08,
          "caption_template_info": {
            "category_id": "",
            "category_name": "",
            "effect_id": "",
            "is_new": false,
            "path": "",
            "request_id": "",
            "resource_id": "",
            "resource_name": "",
            "source_platform": 0
          },
          "check_flag": 7,
          "combo_info": {
            "text_templates": []
          },
          "content": "",
          "fixed_height": -1.0,
          "fixed_width": -1.0,
          "font_category_id": "",
          "font_category_name": "",
          "font_id": "",
          "font_name": "",
          "font_path": "",
          "font_resource_id": "",
          "font_size": 15.0,
          "font_source_platform": 0,
          "font_team_id": "",
          "font_title": "none",
          "font_url": "",
          "fonts": [],
          "force_apply_line_max_width": false,
          "global_alpha": 1.0,
          "group_id": "",
          "has_shadow": false,
          "id": "",
          "initial_scale": 1.0,
          "inner_padding": -1.0,
          "is_rich_text": false,
          "italic_degree": 0,
          "ktv_color": "",
          "language": "",
          "layer_weight": 1,
          "letter_spacing": 0.0,
          "line_feed": 1,
          "line_max_width": 0.82,
          "line_spacing": 0.02,
          "multi_language_current": "none",
          "name": "",
          "original_size": [],
          "preset_category": "",
          "preset_category_id": "",
          "preset_has_set_alignment": false,
          "preset_id": "",
          "preset_index": 0,
          "preset_name": "",
          "recognize_task_id": "",
          "recognize_type": 0,
          "relevance_segment": [],
          "shadow_alpha": 0.9,
          "shadow_angle": -45.0,
          "shadow_color": "",
          "shadow_distance": 5.0,
          "shadow_point": {
            "x": 0.0,
            "y": 0.0
          },
          "shadow_smoothing": 0.45,
          "shape_clip_x": false,
          "shape_clip_y": false,
          "source_from": "",
          "style_name": "",
          "sub_type": 0,
          "subtitle_keywords": null,
          "subtitle_template_original_fontsize": 0.0,
          "text_alpha": 1.0,
          "text_color": "#FFFFFF",
          "text_curve": null,
          "text_preset_resource_id": "",
          "text_size": 30,
          "text_to_audio_ids": [],
          "tts_auto_update": false,
          "type": "text",
          "typesetting": 0,
          "underline": false,
          "underline_offset": 0.22,
          "underline_width": 0.05,
          "use_effect_default_color": true,
          "words": {
            "end_time": [],
            "start_time": [],
            "text": []
          }
        }
      • LICENSE.duo-video 1 KB · in bundle
      • SOURCE.md 728 B
        # JianYing protocol templates
        
        These JSON protocol templates are pinned to
        [`duoec/duo-video`](https://github.com/duoec/duo-video) commit
        `ef4eb46c823910553f901649f2f13fd7575e748f`, under its MIT license. They are
        data/schema baselines, not executable upstream code. Runtime builders deep-copy
        the templates and replace authored values such as IDs, paths, timings, canvas
        dimensions, and resource configuration.
        
        The required copyright and permission notice is preserved in
        `LICENSE.duo-video` in this directory.
        
        The exporter never uses duo-video's embedded example credentials. Resource-ID
        features accept an offline `material` or `resource_config` payload so official
        resource packages can be supplied legally by the caller.
        
    • audio-modes.md 3.5 KB
      # Assembly audio modes
      
      `assemble.py` keeps `narration` as its API and CLI default. Non-narration
      behavior is opt-in with `--audio-mode`; `--audio-stream-index N` is the
      zero-based audio ordinal used by FFmpeg's `0:a:N` selector.
      
      ## `narration`
      
      - Requires non-empty `tts_meta.json` segments, as before.
      - Builds/places narration WAV, performs source ducking and optional BGM mix,
        then applies the configured final loudness/limiter stage and AAC encoding.
      - Missing source audio may use the existing synthetic-silence fallback.
      - With the additional strict `--audio-mix-adoption`, narration instead consumes an
        adopted prepared bed plus complete direct-to-48-kHz narration placements. This is
        still narration mode, but it bypasses ambient BGM, ducking, speed/fit, loudnorm, and
        limiter operations. See `explicit-audio-mix.md`.
      
      ## `source-mix`
      
      - Does not read or require `tts_meta.json` and never creates narration audio.
      - Uses the selected input audio stream, applies `IDLE_ORIG_VOLUME`, mixes a
        declared `BGM_PATH` when present, and runs the configured final
        loudness/limiter stage before AAC encoding.
      - A declared but missing `BGM_PATH` is an error in this new mode (the legacy
        narration-mode warn-and-skip behavior remains unchanged).
      - This is a processed mix, not frozen audio. `assembly_qc.json` reports the
        operations that actually ran.
      
      ## `adopted-packet-copy`
      
      - First implementation supports a selected AAC stream from `<video>` only.
      - Rejects TTS segments, explicit `--tts-meta`, any configured `BGM_PATH`, a
        missing/wrong stream, unsupported codec, or audio/picture interval mismatch.
      - Maps that stream with `-c:a copy`. It does not build narration, mix, duck,
        normalize, limit, resample, change tempo, or add silence.
      - It deliberately omits `-t` and `-shortest`, preserving AAC priming and tail
        packets. After rendering, the output is probed and compared with the input:
        decoder parameters, packet count, total payload bytes, the packet-clock span,
        and every packet's size and PTS/DTS/duration converted to rational time must
        match. Packet side data also must match, including AAC skip-sample/
        discard-padding values and their reason fields.
      - QC/manifest evidence records selected stream, absolute stream index, codec,
        time base, sample rate, channels/layout, packet count, payload bytes, packet
        details, and actual operation flags.
      
      The adopted-copy settings payload excludes narration/mix/loudness defaults, so
      irrelevant ambient settings do not invalidate a frozen-audio render. Video
      filters and video re-encoding remain allowed; packet verification must still
      pass afterward.
      
      These modes do not create a commercial quality profile, automatic speech
      alignment, listening approval, or release approval. An explicit subtitle track
      is only a version/media-bound timing declaration; its evidence labels and
      `NOT_CHECKED` acoustic/listening status remain authoritative.
      Matching packets and decoder parameters show the encoded audio was copied; they
      do not prove perceptual quality or that a person listened to the result.
      
      `timeline.json` represents either non-narration mode as one complete clip from
      the current input video and records the selected stream. Adopted copy uses gain
      1 with no automation. Source-mix records processed/non-frozen delivery and is
      not claimed to be reconstructable by an editor. The optional JianYing exporter
      currently supports only selected stream 0; requesting export with another
      stream fails instead of silently substituting the default stream.
      
    • explicit-audio-mix.md 3.7 KB
      # Explicit adopted full-sound mix
      
      Use this path only when the picture, a completed `prepared_bed_receipt`, an exact
      narration adoption, narration placements/gains, and one fixed master gain have already
      been independently selected. It is still `audio_mode=narration`, but it bypasses the
      legacy narration timing, ducking, ambient BGM, loudness-normalization, and limiter path.
      
      ```bash
      python3 scripts/assemble.py picture.mp4 --work-dir NEW_WORK \
        --tts-meta /local/tts_meta.json \
        --narration-adoption /local/narration_adoption.json \
        --audio-mix-adoption /local/audio_mix_adoption.json
      ```
      
      All strict output paths, including the CLI delivery alias, must be new. The CLI copies
      to a hidden delivery stage and creates the alias with an exclusive atomic link, so a
      concurrent or existing file is never overwritten.
      
      ## Adoption schema v1
      
      ```json
      {
        "artifact": "audio_mix_adoption",
        "schema_version": 1,
        "prepared_receipt": {"path": "/local/prepared_bed_receipt.json"},
        "format": {"sample_rate": 48000, "channels": 2, "total_samples": 1856000},
        "segments": [
          {"index": 0, "output_start_sample": 366000, "gain": 0.4251421093940735}
        ],
        "master_gain_db": 0.75
      }
      ```
      
      Unknown or missing fields fail; legacy `*_sha256` keys are ignored. The format must
      equal the actual zero-origin picture frame clock at 48 kHz and the prepared receipt.
      The receipt, all three float PCM beds (format, sample count, finite PCM), the picture,
      the narration adoption file and the ordered segment indices are checked before
      narration snapshots are written.
      
      ## Exact audio operations
      
      Each narration snapshot is decoded directly and completely to `pcm_f32le`, 48 kHz
      stereo. Version 1 uses one fixed channel matrix, recorded per segment as
      `mono_equal_power` or `stereo_identity`:
      
      - mono is panned to left and right at `1/sqrt(2)` per channel;
      - stereo preserves independent left and right channels at unit gain;
      - inputs with more than two channels are rejected.
      
      There is no intermediate 44.1 kHz mono placement, speed change, fade, trim, tail pad,
      normalization, or automatic fit. Every complete converted WAV must fit its adopted
      integer `output_start_sample`; placements may not overlap. Adopted per-segment gains
      form a float `voice_bus.wav`. The producer's `prepared_bed.wav` and voice bus form a
      float premaster, then the sole adopted `master_gain_db` forms the float master consumed
      by the final AAC encode. The picture's old audio is never a mixer input.
      
      ## Records and publication
      
      `narration_input_binding.json` records the actual direct 48 kHz placed files and voice
      bus, marking the bus `CONSUMED_BY_EXPLICIT_MIX`. `audio_mix_binding.json` records the
      mix adoption path, picture path and clock, prepared receipt path and stems, converted
      placements (path, channel matrix, PCM facts), voice bus, premaster, master, the
      narration binding path, the final decoded PCM facts, the final AAC decoder/packet
      count/payload bytes, and the actual output picture decoder/frame clock. The output
      picture must preserve fps, frame count, zero start, and duration. A video-copy path
      reports `packet_identity: EXACT` when decoder parameters and packet sizes/timestamps
      are unchanged; an allowed visual re-encode reports `REENCODED_CLOCK_MATCH`. Timeline,
      settings, QC, and manifest identify the explicit path rather than reporting ambient
      ducking/BGM/loudnorm operations.
      
      The candidate render is written to a hidden file. Both bindings are written, QC must
      pass, the media is published, and a second QC runs against the published path. Any
      render, binding, or QC failure removes the candidate, final-named media and the
      bindings written by this run instead of overwriting an older success.
      Diagnostic/staging audio may remain in the unique work directory.
      
    • foreground-compose.md 5.1 KB
      # Caller-rendered foreground composition
      
      `compose_foreground.py` performs one narrow operation: it overlays a caller-rendered,
      full-canvas RGBA PNG sequence on an already locked H264/CFR/AAC base, optionally
      replaces a declared tail with an explicit endcard asset, and copies `a:0` without
      decoding or rewriting it. The PNG pixels are the source schema. This is not a CSS,
      font, text, subtitle, logo, animation, or plugin renderer.
      
      ## Command
      
      ```bash
      python3 scripts/compose_foreground.py foreground_plan.json \
        --output-dir a-new-directory [--plan-only]
      ```
      
      The output directory must not exist. A normal run publishes `foreground.mp4` only
      after the staged media's frame clock, canvas/color metadata, AAC decoder/packets/PTS/
      side data, and a full audio/video decode pass verification.
      A failure retains `foreground_run.json` and available logs but removes staged/final
      media. `--plan-only` validates the complete plan and inputs and writes only
      `foreground_run.json`; it does not render media or update any current pointer.
      
      ## Strict plan schema v1
      
      ```json
      {
        "artifact": "foreground_compose_plan",
        "schema_version": 1,
        "base": {"path": "/local/base.mp4"},
        "video": {"fps": "30/1", "width": 1280, "height": 720, "total_frames": 300},
        "foreground": {
          "directory": "/local/foreground_sequence",
          "pattern": "frame_%06d.png",
          "start_frame": 0,
          "end_frame": 270
        },
        "endcard": {
          "kind": "sequence",
          "directory": "/local/endcard_sequence",
          "pattern": "frame_%06d.png",
          "start_frame": 270,
          "end_frame": 300
        },
        "producer_receipt": {"path": "/local/producer_receipt.json"}
      }
      ```
      
      `endcard` is an exact tagged union. The alternative static form is:
      
      ```json
      {
        "kind": "still",
        "path": "/local/endcard.png",
        "start_frame": 270,
        "end_frame": 300
      }
      ```
      
      When the foreground itself covers the entire declared frame clock, the only accepted
      no-tail form is exactly:
      
      ```json
      {"kind": "none"}
      ```
      
      It accepts no additional fields. It is invalid when the foreground reserves any tail;
      there is no implicit filler, repeated last frame, or generated endcard.
      
      Use `still` only when a genuinely static full-canvas endcard is intended. A static
      PNG does not reconstruct a delivered fade or other animated endcard; supply the
      caller-rendered `sequence` form for that case.
      
      Both sequences use **local zero-based filenames** regardless of their output range.
      Thus the endcard example contains `frame_000000.png` through `frame_000029.png`,
      mapped to output frames `[270,300)`. The foreground must start at output frame zero.
      Its end must equal the endcard start, and the endcard end must equal the independently
      probed base frame count. With `kind: "none"`, the foreground end must instead equal
      that full base frame count. Bounds are half-open.
      
      The only accepted filename pattern is the literal `frame_%06d.png`. The directory
      must contain exactly the expected contiguous entries—no missing or extra files. Every
      actual image must probe as PNG, `rgba`, and the exact declared full canvas. Repeated
      files and symlinks are allowed so transparent holds need not duplicate storage. The
      run report records each sequence's directory and validated `frame_count`. Legacy
      `sha256` / `ordered_sha256` keys in a plan are ignored.
      
      ## Base and output invariants
      
      The first version accepts one selected `v:0` H264, zero-origin CFR, YUV420P,
      BT.709/TV-range base and adopted AAC `a:0`. Declared fps, canvas, and frame count are
      checked independently against actual decoded frame PTS. Picture and audio intervals
      must satisfy the narrow `pair_media` timing contract.
      
      Composition explicitly overlays in RGB, converts to BT.709 limited-range YUV420P,
      limits the final video filter to the declared half-open frame clock, and re-encodes
      only video with fixed `libx264 -preset fast -crf 18` settings. A no-tail run uses only
      the base and foreground inputs; it does not synthesize or pad a tail. The filter EOF
      prevents still-image input loops without imposing an output-level video frame cap that
      could stop accepted trailing AAC packets from being copied. The run
      report records that encoder, preset, CRF, pixel format, color contract, and audio-copy
      mode explicitly. H.264 CRF encoding is lossy: preserved transparent regions are
      visually checked against tolerances, not claimed pixel-identical to decoded base
      pixels. The compositor does not use `-r`, `-shortest`, or `-t`. Audio is mapped from
      base `a:0` with `-c:a copy`; output verification requires matching decoder
      parameters, packet count, payload bytes, PTS/DTS/duration/size, and side data.
      
      ## Provenance and review boundary
      
      The producer receipt is a caller artifact that must exist when the plan is read.
      Its declarations may
      describe roles or content decisions such as title/note/brand, `dialogues=[]`, or
      hidden markers. Core records those declarations as `DECLARED_NOT_CHECKED`; it does
      not infer semantic truth from pixels or claim that dialogue/subtitle duplication is
      absent. A separate producer/reviewer must establish that evidence.
      
      Every run reports `direct_listening` and `normal_speed_review` as `NOT_CHECKED` and
      `release_approved` as false. Pixel validation and deterministic rendering are not
      commercial-quality or release approval.
      
    • narration-adoption.md 4.7 KB
      # Narration adoption and the consumed-input record
      
      Narration assembly accepts legacy `tts_meta.json`, but a current file by itself is
      not an adoption decision. Use an explicit adoption document when the selected spoken
      text, requested provider/voice, and tempo policy must be bound to the actual final mix.
      
      ```bash
      python3 scripts/assemble.py input.mp4 --work-dir work \
        --tts-meta /local/tts_meta.json \
        --narration-adoption /local/narration_adoption.json
      ```
      
      `--narration-adoption` is narration-only and requires an explicitly supplied
      `--tts-meta`; it never adopts an ambient work-directory file automatically.
      
      ## Strict v1 schema
      
      ```json
      {
        "artifact": "narration_adoption",
        "schema_version": 1,
        "segments": [
          {
            "index": 0,
            "spoken_text": "exact selected words",
            "requested_provider": "caller-selected-provider",
            "requested_voice": "caller-selected-voice"
          }
        ],
        "tempo_policy": {
          "global_atempo": 1.0,
          "bounded_segment_fit": false,
          "segment_tempo_max": 1.0,
          "cumulative_tempo_max": 1.0,
          "cumulative_tempo_hard_max": 1.0
        }
      }
      ```
      
      That policy is the conservative default. An adoption may declare its own values and
      they are honoured: `global_atempo` must be positive and no greater than
      `cumulative_tempo_hard_max`, `segment_tempo_max` and `cumulative_tempo_max` must be at
      least `1.0`, `cumulative_tempo_hard_max` must not be below `cumulative_tempo_max`, and
      `bounded_segment_fit` must be a boolean. Malformed or out-of-bounds policies fail;
      ambient defaults such as a global 1.15 speed never override an adoption. With
      `bounded_segment_fit` false, an adopted segment that does not fit its authored window
      blocks without time trimming or bounded fit.
      
      The adoption's ordered segment indices and spoken text must match both the segment
      list in the supplied `tts_meta.json` and the in-memory list passed to
      `assemble_video`; those two lists must be equal to each other. Unknown fields fail;
      legacy `tts_meta_sha256` / `processed_wav_sha256` keys are ignored. The consumer never
      creates an adoption from metadata that it is about to consume.
      
      ## Validation and snapshots
      
      Every adopted `audio_path` must exist before any media is probed or rendered. The
      adopted inputs are then copied to per-run snapshots under
      `work/.narration_input_snapshots/`; the original files are never modified, and the
      render reads only the snapshots. Every conversion, placed WAV, and the completed
      narration bus is recorded with its path and probed PCM facts before the final FFmpeg
      command runs.
      
      On the legacy narration-mix path, Python's standard WAV reader cannot open every valid post-processed WAV encoding.
      Noncanonical input, including `pcm_f32le`, 48 kHz, or stereo WAV, is explicitly
      decoded by the existing FFmpeg executable to 44.1 kHz mono PCM16 before placement.
      The conversion path and actual PCM facts are recorded; the consumer does not claim
      that converted samples equal the original samples.
      
      The explicit full-sound path described in `explicit-audio-mix.md` deliberately does
      not use that 44.1 kHz mono conversion. It consumes the snapshots directly as complete
      48 kHz float placements, preserving native stereo channels.
      
      An adopted render writes to a hidden candidate path. The binding is written, QC runs
      against the candidate, the candidate is renamed to the final path, and QC runs once
      more against the published file. A blocking QC removes the candidate and the binding
      instead of publishing; the run records the final output as absent rather than
      claiming a missing file was published. Non-final work or diagnostic artifacts may
      remain for investigation.
      
      ## Binding report
      
      After successful assembly, `work/narration_input_binding.json` records:
      
      - original input path and snapshot path;
      - any explicit conversion path and PCM parameters;
      - each complete placed WAV path and PCM parameters;
      - `narration.wav` path and PCM parameters;
      - the final output path and its encoded audio-stream facts (decoder parameters,
        packet count, payload bytes, start time, duration);
      - the adoption path, its `tts_meta` path, and the exact tempo policy when supplied.
      
      The report uses one of two identity statuses:
      
      - `UNADOPTED`: no adoption was supplied; the legacy inputs were consumed as given;
      - `BOUND_TO_ADOPTION`: the adoption's selection and tempo policy were carried
        through the final mix.
      
      `BOUND_TO_ADOPTION` records what was consumed, not provider truth, acoustic speaker
      identity, direct listening, naturalness, or release approval; voice authentication
      and direct listening stay `NOT_CHECKED`. QC, manifest, and settings records reference
      a binding only while its recorded final output path still exists; source-mix and
      adopted-packet-copy modes do not reuse narration binding evidence.
      
    • packaging.md 2.8 KB
      # 字幕与包装的推荐顺序
      
      先锁定画面、剪点、旁白和混音,再投入字幕动画或边框包装。字幕样式不能掩盖叙事、剪点或声音问题。
      
      普通交付优先使用现有 ASS 路径。只有用户需要更精细的逐 cue 排版、动画或透明图层时,才使用 Remotion
      或其他代码渲染器作为**项目级可选实现**,不要把特定框架、字体、颜色或黄字写成核心依赖。
      
      ## 推荐顺序
      
      1. 输出无包装的锁定母版,确认音画内容不再变化。
      2. 从实际 TTS / 时间线生成 captions;TTS 块保持连续思路,字幕可以按阅读宽度拆 cue。
      3. 先抽检开头、亮背景、暗背景、人物近景和长字幕样帧,确定字号、安全区、描边、阴影及是否需要底板。
      4. 渲染完整透明字幕层,再 overlay 到锁定母版;合成后确认音轨未被意外改写。
      5. 完整播放实际最终文件,并复查字幕遮脸、跳字、断行、首尾帧和边界处残影。
      
      已有项目级渲染器能输出精确包装时,使用 `foreground-compose.md` 将它生成的 RGBA 序列叠到锁定母版,
      而不是用通用白字黑框近似品牌样式。该操作保留实际帧钟和 AAC 包,不生成字体或文案;只有通过验证的
      新文件才写入新目录。片名卡有渐显或动画时须提供完整序列,不能冻结最后一张图代替。
      
      ## 静态包装图层 `packaging_layers.json`
      
      包框、标题条、角标 logo 这类整片不动的图片,可以直接在合成时叠加,不必先渲染成逐帧序列。
      在 `work_dir` 写 `packaging_layers.json`:
      
      ```json
      {
        "canvas": {"width": 1080, "height": 1920},
        "layers": [
          {"name": "frame", "path": "/abs/frame.png", "rect": {"x": 0, "y": 0, "width": 1080, "height": 1920}}
        ],
        "template": {"id": "brand-frame", "version": 2}
      }
      ```
      
      - `canvas` 必须等于成片画布,否则合成前报错;`rect` 必须在画布内,图片缩放到 `rect` 大小(透明区域保持透明)。
      - 叠加顺序:遮原字幕 → 包装图层(按数组顺序)→ 画面文字 → 解说字幕 → 缩放。
      - `timeline.json` 同时得到对应的 image 轨,剪映草稿里的包框可单独编辑。
      - `assembly_manifest.json` 的 `video_filters.packaging_layers` 记录每张图片的路径与 `{size, mtime_ns}`。
      - 编排器按项目绑定的 `packaging` 模板写出这个文件时会带 `written_by` 标记;
        手写的文件不会被它删除。有渐显或动画的包装仍走 `foreground-compose.md`。
      
      ## 原则
      
      包装价值来自稳定、可读、与内容一致的排版,不来自效果数量。先建立统一字体、颜色、描边/阴影和轻量动效;
      底板、边框、花字与音效只有解决具体可读性或叙事任务时才加入。一个样帧好看不代表全片成立。
      
    • pair-media.md 4.1 KB
      # Pair rebuilt picture with a retained adopted soundtrack
      
      When the picture is rebuilt/reframed but the adopted mix must not change, use
      `pair_media.py` before `assemble.py --audio-mode adopted-packet-copy`. This is a
      real two-input stream copy, not re-encoding an old finished video's picture.
      It does not generate narration, align speech, design packaging, or create an
      end card. It does not decide which inputs the editor intended to adopt.
      
      ```json
      {
        "artifact": "media_pair",
        "schema_version": 1,
        "picture": {"path": "/project/picture.mp4"},
        "audio": {"path": "/project/adopted.m4a", "selected_stream": 0}
      }
      ```
      
      Both paths must be explicit local files. `selected_stream` is the audio ordinal
      (`a:N`), not the absolute stream index. The audio donor may be an old MP4 with
      unrelated video: only the selected audio is used. The picture's audio is ignored.
      Both files must exist when the plan is read. The plan is strict: unknown fields
      do not silently request unsupported retime, gain, trimming or offset operations
      (legacy `sha256` keys are ignored).
      
      ```bash
      python3 scripts/pair_media.py pair.json --output-dir new-pair-directory --plan-only
      # A PLANNED directory is not reusable as a rendered candidate: choose a NEW one.
      python3 scripts/pair_media.py pair.json --output-dir new-render-directory
      ```
      
      The output directory must not already exist. No existing input, final asset,
      current pointer, or user-approved version is overwritten. `--plan-only` checks
      actual inputs and timing but creates no video and invokes no mux. A rendered run
      provides:
      
      - `paired.mp4`: selected picture and audio, both compressed-stream copied.
      - `pair_run.json`: `PLANNED`, `PAIR_RENDERED`, or `FAILED`, explicit input and
        plan/output paths, frame count, timing tolerance and output stream numbering.
      - `picture_identity.json`: decoder, geometry/color, complete actual frame clock
        and compressed packet size/timing/side-data facts.
      - `adopted_audio_identity.json`: donor and output AAC packet/decoder facts.
      - mux command/log for reconstruction and diagnosis.
      
      The first implementation deliberately accepts H264 or HEVC in an MP4-family container,
      complete zero-origin CFR picture (1–120 fps), and contiguous AAC packets with
      known nominal sample duration. No VFR, retime, offset, format conversion, padding,
      trimming, `-shortest`, normalization or gain is inferred. Packet priming and skip
      metadata are retained. The packet clock must also agree with the declared audio
      stream interval (at most one nominal priming packet before the start; the packet
      end matches the stream end within one sample). This is checked again on the
      actual muxed output. Picture/audio starts and ends must differ by no more than
      the larger of one picture frame or one nominal AAC packet. The interval check is
      compatibility, **not perceptual synchronization or acoustic alignment**. A tiny
      container-tail tolerance does not authorize cutting a word.
      
      Muxing goes to a staging file. The output's video decoder/packet sizes and
      timestamps/full frame clock/geometry/color and adopted AAC packets must match
      the inputs, the output must contain only `v:0,a:0`, and full decode must pass
      before final publication. Failure leaves a FAILED record and logs, not a
      final `paired.mp4`. Rebuilding source geometry is certified by its own upstream
      render/map evidence, not by the existence of a successfully paired container.
      
      ## Subtitle integration: bind AFTER pairing
      
      The output audio ordinal is always **0**, even when the donor used `a:1`.
      Create the complete output-clock `subtitle_track.json` against
      `subtitles.track_binding.current_bindings(paired_video, 0)`, then run the existing
      adopted assembly path. Do not reuse a binding computed against the picture-only
      file or the donor's old stream ordinal. A present stale subtitle track fails rather than falling
      back to estimated timings. See [subtitle-track.md](subtitle-track.md).
      
      Pairing preserves the input picture, including any explicit black tail; it does
      not turn that black tail into a branded end card. Listening, normal-speed review,
      editorial intent and release approval remain separate from this mechanical proof.
      
    • source-score.md 6.7 KB
      # Source and score bed producer
      
      `source_score.py` is a standalone producer for explicitly authored source-audio
      intervals and one continuous score playhead. It does not mix narration, inspect old
      masters, infer dialogue, choose music, retime sound, normalize loudness, limit peaks,
      or publish a final video.
      
      ```bash
      python3 scripts/source_score.py source_score_plan.json --output-dir a-new-directory
      ```
      
      The output directory must not exist. Success creates:
      
      - `source_bed.wav` — reordered/pre-gained source intervals plus explicit silence;
      - `score_bed.wav` — one continuous raw-score window or an adopted frozen stem;
      - `prepared_bed.wav` — source plus score, without master gain/limiting;
      - `prepared_bed_receipt.json` — input paths, clocks, sample ranges, PCM facts and QC facts.
      
      All three beds are `pcm_f32le`, 48 kHz, stereo. Float output avoids repeated integer
      quantization and preserves values outside `[-1,1]`; the receipt records actual peak,
      finite-sample status, and `FLOAT_PRESERVED_NO_MASTER`. A later explicitly adopted
      consumer owns final master gain and delivery encoding.
      
      ## Strict plan schema v1
      
      ```json
      {
        "artifact": "source_score_plan",
        "schema_version": 1,
        "output": {"sample_rate": 48000, "channels": 2, "total_samples": 1920000},
        "source_segments": [
          {
            "id": "protected-dialogue-1",
            "path": "/local/source.mov",
            "audio_stream": 0,
            "source_fps": "24/1",
            "source_start_frame": 120,
            "source_end_frame": 180,
            "output_start_sample": 240000,
            "gain": 1.0,
            "fade_in_samples": 0,
            "fade_out_samples": 96,
            "fade_shape": "linear",
            "role": "protected_original"
          }
        ],
        "source_silence": [
          {"output_start_sample": 0, "output_end_sample": 240000, "role": "silence"},
          {"output_start_sample": 360000, "output_end_sample": 1920000, "role": "silence"}
        ],
        "score": {
          "kind": "raw",
          "path": "/local/score.wav",
          "audio_stream": 0,
          "source_offset_sample": 384000,
          "gain": 0.18,
          "fade_in_samples": 33600,
          "fade_out_samples": 120000,
          "fade_shape": "half_cosine"
        }
      }
      ```
      
      Every source range is half-open in CFR picture frames. Version 1 accepts H.264 and
      HEVC picture sources whose complete packet PTS/duration grid proves one zero-origin
      packet per same-speed frame. Frame boundaries are projected onto the 48 kHz clock by
      rounding the exact position half-up once, so `48000/fps` need not be an integer and a
      broadcast rate such as 30000/1001 is accepted; every derived bound comes from those
      rounded positions. Other codecs, incomplete packet timing, VFR, retime fields,
      nominal time seeks, and unknown fields are rejected rather than called verified
      (legacy `sha256` keys are ignored).
      Each unique selected source stream is decoded from its actual PTS clock to canonical
      48 kHz stereo exactly once. All authored intervals are then cut from that canonical
      sample array—never independently decoded or `-ss`-seeked per edit.
      
      The source segment duration is `(source_end_frame-source_start_frame)*48000/fps`.
      Its output end is derived, not declared. Source segments and explicit silence ranges
      must form an exact, nonoverlapping partition of `[0,total_samples)`. There are no
      implicit gaps. Supported roles are `protected_original` and
      `mixed_original_under_narration`; their gains and labels are retained unchanged.
      
      Source fades are explicitly `linear`. A fade of `N>=2` samples uses inclusive
      endpoints: fade-in gain is `i/(N-1)`, and fade-out gain is `(N-1-i)/(N-1)`. Thus the
      first/last samples are exactly zero and the opposite endpoints exactly one. Zero
      means no fade; a one-sample fade is rejected as ambiguous. Fade windows may not
      overlap within a segment.
      
      ## Score modes
      
      `raw` decodes the selected score stream once, advances one playhead from
      `source_offset_sample`, and takes exactly `total_samples`. The input must be long
      enough; there is no implicit loop. Gain and one pair of whole-score fades are applied
      once, so score playback never resets at picture/source cuts. `linear` and
      `half_cosine` fades use inclusive endpoints.
      
      When the adopted source bed is already the complete soundtrack and no additional
      music is selected, use the exact no-score form:
      
      ```json
      {"kind":"none"}
      ```
      
      It accepts no other fields. `score_bed.wav` is real all-zero float PCM at the exact
      output length, while `prepared_bed.wav` is an exact canonical PCM copy of
      `source_bed.wav`; the receipt records `kind:none` rather than inventing a music asset.
      
      An adopted score that already contains its chosen offset/gain/fades uses the exact
      alternative schema:
      
      ```json
      {"kind":"frozen","path":"/local/frozen.wav","audio_stream":0}
      ```
      
      Frozen input must be PCM16, PCM24, or PCM float, 48 kHz stereo, and exactly
      `total_samples`. PCM16/24 values convert exactly to float canonical representation.
      No offset, gain, fade, loop, normalization or other processing fields are accepted.
      The receipt records the frozen file path and its probed PCM facts.
      
      ## Receipt and failure boundary
      
      `prepared_bed_receipt` schema version 1 records the plan path, each input path,
      selected stream, actual CFR/audio clocks, canonical decoded PCM facts, resolved
      input/output sample ranges, gains, fades, roles, and each output's path, byte size,
      format, sample count, finite status and peak.
      
      Every declared input must exist and probe as declared before work starts. A selected
      source range must fit the actual canonical decoded samples; silence padding can never
      conceal a short source. WAVs remain hidden staging artifacts until every output
      validates. FFmpeg failure, invalid ranges, non-finite PCM, or an existing target
      directory cannot publish final-named beds or a receipt.
      Diagnostic command/log/intermediate files may remain in the unique failed directory.
      
      To consume a completed receipt with explicitly adopted narration, use the strict
      `--audio-mix-adoption` path in `explicit-audio-mix.md`. To deliver an already complete
      prepared bed without narration, a dedicated prepared-audio renderer will follow in a
      later release. Do not feed `prepared_bed.wav` through legacy source ducking or
      ambient BGM/loudness settings.
      
      ## 与旧入口的关系(中文摘要)
      
      原片完整解码一次再切样本,分别执行保留对白、低位原声和明确静音;渐变必须显式给定。
      已处理的音乐轨走 `frozen`,不能再次偏移、调增益或加渐变。若采用的原声底轨本身已包含
      完整音乐决定且不再叠加配乐,使用严格的 `score:{"kind":"none"}`;它生成真实全零 score,
      并保持 prepared 与 source 逐字节相同,不伪造静音音乐资产。
      
      这一步仅输出声音底轨和来源回执,不是最终视频。旧入口保留兼容行为;调用方须区分
      “底轨已验证”“配音已验证”和“完整混音已验证”三种状态,不能相互冒充。
      
    • subtitle-track.md 8.5 KB
      # Independent subtitle track contract (schema v1)
      
      `scripts/subtitles/track.py` loads a subtitle track whose cue times already use
      the final **output clock**. It validates the track against picture, edit, audio,
      and duration facts independently supplied by the caller. It does not align
      speech, remap source time, split text, repair cue boundaries, or prove that a
      subtitle is perceptually synchronized.
      
      ## Loader API
      
      ```python
      from fractions import Fraction
      from subtitles.track import load_subtitle_track
      
      loaded = load_subtitle_track(
          "subtitle_track.json",
          expected_picture_identity={
              "path": "/project/paired.mp4",
              "edit_plan": "/project/edit_plan.json",
          },
          expected_audio_identity={
              "selected_stream": 1,
              "sample_rate": 48000,
              "packet_count": 4700,
          },
          expected_duration_seconds=Fraction(duration_ts) * stream_time_base,
          reject_legacy_estimate=True,
      )
      metadata = loaded["metadata"]
      entries = loaded["entries"]
      ```
      
      The input may also be an in-memory mapping. `entries` is a list of dictionaries
      ready for the existing seconds-based subtitle renderer:
      
      ```json
      {
        "start": 1.0,
        "end": 3.0,
        "text": "example",
        "source": "narration",
        "source_ref": "narration:7",
        "timing_evidence": {
          "kind": "asr_boundary_calibrated",
          "evidence_refs": ["alignment-run:example"],
          "calibration": "asr_energy",
          "word_alignment": "none"
        }
      }
      ```
      
      The loader preserves cue text and boundaries. The only conversion is exact
      integer ticks to renderer-facing seconds. `metadata.timing_evidence_kinds`
      reports the labels present; it is not an aggregate precision verdict.
      
      ## Schema v1
      
      ```json
      {
        "schema_version": 1,
        "clock": {
          "kind": "output",
          "timebase": {"numerator": 1, "denominator": 30},
          "duration_ticks": 300
        },
        "overlap_policy": "forbid",
        "bindings": {
          "picture": {
            "path": "/project/paired.mp4",
            "edit_plan": "/project/edit_plan.json"
          },
          "audio": {
            "selected_stream": 1,
            "sample_rate": 48000,
            "packet_count": 4700
          }
        },
        "cues": [
          {
            "start_tick": 30,
            "end_tick": 90,
            "text": "example",
            "attribution": {"kind": "narration", "ref": "narration:7"},
            "timing_evidence": {
              "kind": "asr_boundary_calibrated",
              "evidence_refs": ["alignment-run:example"],
              "calibration": "asr_energy",
              "word_alignment": "none"
            }
          }
        ]
      }
      ```
      
      - `timebase` is rational seconds per tick. Cue intervals are half-open
        `[start_tick, end_tick)`.
      - `duration_ticks`, cue bounds, and stream indexes are nonnegative integers
        (booleans are not integers for this contract).
      - Cues must be ordered, non-overlapping, nonempty, and contained by the output
        duration. Adjacent half-open cues may touch.
      - Distinct tick bounds must remain a positive, finite interval after conversion
        to the renderer's float seconds. Tracks beyond that projection precision fail
        closed rather than becoming a zero-length rendered cue.
      - `overlap_policy` must be `forbid`. Schema v1 has no permissive overlap mode.
      - `attribution.kind` is `source` or `narration`; `ref` identifies the source
        utterance or narration item without changing its text.
      - Unknown fields and schema versions other than integer `1` are rejected, so a
        newer producer cannot be silently interpreted as v1. Legacy `sha256` /
        `edit_sha256` binding keys are the one exception: they are ignored.
      
      ## Independent binding checks
      
      The loader compares track declarations with the caller's current facts; it never
      treats a track's own binding as evidence that the track is fresh.
      
      - `picture.path` is mandatory and must resolve to the caller's current picture
        path. `edit_plan` is optional for inputs without a separately materialized
        edit plan; if the track contains it, the caller must supply the same path.
      - `audio.selected_stream`, `sample_rate` and `packet_count` describe the
        **actually adopted** audio stream, not merely a source filename or an intended
        mix manifest. Pass the facts probed from the current media; this module
        deliberately does not run ffprobe itself.
      - `expected_duration_seconds` is mandatory and checked against
        `duration_ticks * timebase`. Prefer `Fraction(duration_ts) * time_base` from
        the actual output stream to avoid decimal/container rounding. `int`, finite
        `float`, and finite `Decimal` are also accepted.
      
      ## Timing evidence labels
      
      Evidence labels describe how a cue boundary was obtained. They do not change
      the exact burn time represented by its ticks.
      
      | `kind` | Required calibration | Required word alignment | Evidence refs |
      |---|---|---|---|
      | `legacy_estimate` | `none` | `none` | optional |
      | `asr_boundary_calibrated` | `asr_energy` | `none` | required |
      | `word_timestamps` | `none` or `asr_energy` | `asr_words` | required |
      | `human_verified` | `human_boundary` | `none` or `human_words` | required |
      
      `asr_boundary_calibrated` is the label for a coarse-ASR boundary adjusted with
      ASR context and/or energy evidence. Energy onset is not proof of a phoneme, so
      this label cannot claim word alignment or human verification. Missing evidence
      references cannot claim any non-legacy kind. A strict caller may reject
      `legacy_estimate`; accepting another label still does not automatically call it
      strong alignment.
      
      Because cue intervals are half-open, a cue is not visible at any tick before its
      `start_tick`; the tick immediately preceding it belongs to whatever came before.
      
      ## Deliberate limits
      
      - Schema v1 has no source-to-cut/edit map and cannot map source-clock cues.
      - The module does not inspect media, select audio streams, or tolerate a
        self-declared binding without current caller facts.
      - Validation proves schema consistency and the requested bindings only. Actual
        rendered first/last subtitle frames and perceptual speech alignment require
        separate render/media review.
      
      ## Current assembly integration (bounded, not automatic alignment)
      
      Place `subtitle_track.json` in `work_dir` and select
      `assemble.py --audio-mode adopted-packet-copy`. A present track is a **complete
      replacement** of all generated narration/original-dialogue subtitles, not a
      partial patch and not merged with legacy subtitles. Carry every cue that should
      remain; an empty `cues` array intentionally removes all generated subtitles.
      Unknown patch/merge modes are rejected by schema v1. This does not detect words
      missing from the authored full track: acoustic/coverage review remains required.
      
      The assembly integration independently probes the chosen input stream for its
      sample rate and packet count, resolves the input path for the picture binding,
      and records the `{size, mtime_ns}` of the track, video and optional edit plan in
      `subtitle_track_validation.json`. It does not use the track's own declarations
      as current facts.
      
      Actual decoded frame PTS are read before projection. Integer cue ticks and the
      rational timebase are retained until each boundary is resolved to the first
      frame at or after it. `subtitle_track_validation.json` records original ticks,
      resolved frame indexes/PTS, quantization deltas, ASS thresholds, validation
      schema and projector version. SRT and timeline consume resolved frame seconds;
      ASS thresholds are chosen to switch on those same frames despite ASS's 10ms
      clock. A cue with no visible frame, or a boundary that ASS cannot distinguish,
      is rejected rather than silently dropped. The original author file is not
      rewritten. Legacy subtitles keep their previous rendering behavior.
      
      Each consumption re-checks the track/media/optional edit plan `{size, mtime_ns}`
      and the consumer duration; a rewritten input is stale and must be prepared
      again. Deleting the explicit track clears its previous validation record.
      An invalid/stale track never falls back to character-proportional timing.
      
      Current limits:
      - Integration is for output media starting at zero with an adopted AAC track,
        not raw-source-to-edited-output mapping or a newly mixed narration track.
      - The low-level preparation API accepts `edit_plan_path`, but the assembly CLI
        does not yet expose it. CLI tracks must omit `edit_plan`; that path binds the
        input media, not edit-plan ancestry.
      - The low-level policy `reject_legacy_estimate=True` is available. There is not
        yet a wired commercial-profile CLI gate. The current CLI preserves timing
        evidence labels and does not call every accepted cue precisely aligned.
      - The actual FFmpeg regression demonstrates frame timing with a synthetic cue,
        not phoneme alignment. `direct_listening` and `acoustic_alignment` remain
        `NOT_CHECKED`; evidence labels are declarations, not automatically verified
        human approval or proof that an external evidence reference is true.
      
  • scripts
    • adoption
      • audio_mix_binding.py 13.4 KB
        """Validate, render, and record an explicitly adopted prepared-bed plus narration mix."""
        
        import json
        from pathlib import Path
        import subprocess
        
        from assemble_constants import frame_clock_samples
        from adoption.frozen_audio import probe_audio_packets
        import adoption.narration_binding as narration_binding
        from pair_media import probe_picture, validate_pair_timing
        import source_score
        from adoption.strict_inputs import (
            read_json_bytes, require_fields, require_integer, require_local_path, require_number,
            run_logged, without_digests, write_json_atomic,
        )
        
        
        ARTIFACT = "audio_mix_binding"
        FILENAME = "audio_mix_binding.json"
        RATE = 48_000
        CHANNELS = 2
        CODEC = "pcm_f32le"
        
        
        def _picture_format(picture):
            samples = frame_clock_samples(picture["frame_count"], picture["fps"], RATE)
            return samples, {
                "fps": picture["fps"], "frame_count": picture["frame_count"],
                "duration": picture["duration"], "start": picture["start"],
                "decoder": picture["decoder"],
            }
        
        
        def load_adoption(path, *, input_video, narration_adoption_path, tts_segments):
            """Strict read-only preflight; no work artifacts are created here."""
            adoption_path, _, value = read_json_bytes(path, "audio mix adoption")
            value = without_digests(value, "audio mix adoption")
            require_fields(value, ["artifact", "schema_version", "prepared_receipt", "format", "segments",
                                   "master_gain_db"], "audio mix adoption")
            if value["artifact"] != "audio_mix_adoption" or type(value["schema_version"]) is not int \
                    or value["schema_version"] != 1:
                raise ValueError("unsupported audio_mix_adoption schema")
            input_video = require_local_path(input_video, "picture")
            picture = probe_picture(input_video)
            picture_samples, picture_summary = _picture_format(picture)
        
            narration_path = require_local_path(narration_adoption_path, "narration adoption")
            require_fields(value["format"], ["sample_rate", "channels", "total_samples"], "mix format")
            mix_format = value["format"]
            if mix_format != {"sample_rate": RATE, "channels": CHANNELS,
                              "total_samples": picture_samples}:
                raise ValueError("mix format differs from the actual picture sample clock")
            prepared_receipt = source_score.validate_prepared_receipt(
                value["prepared_receipt"], mix_format
            )
        
            if not isinstance(value["segments"], list) or len(value["segments"]) != len(tts_segments):
                raise ValueError("audio mix segments must exactly cover narration segments")
            normalized = []
            seen = set()
            for adopted, segment in zip(value["segments"], tts_segments):
                adopted = without_digests(adopted, "audio mix segment")
                require_fields(adopted, ["index", "output_start_sample", "gain"], "audio mix segment")
                index = require_integer(adopted["index"], "audio mix segment index")
                if index in seen or index != segment["index"]:
                    raise ValueError("audio mix segment index/order differs from narration")
                seen.add(index)
                normalized.append({
                    "index": index,
                    "output_start_sample": require_integer(
                        adopted["output_start_sample"], "output_start_sample"),
                    "gain": require_number(adopted["gain"], "narration gain", 0, 16),
                })
            return {
                "path": str(adoption_path),
                "picture": {"path": str(input_video), "clock": picture_summary},
                "picture_identity": picture,
                "prepared_receipt": prepared_receipt["reference"],
                "prepared": prepared_receipt["outputs"],
                "narration_adoption": {"path": str(narration_path)},
                "format": dict(mix_format), "segments": normalized,
                "master_gain_db": require_number(value["master_gain_db"], "master_gain_db", -24, 24),
                "runtime": None,
            }
        
        
        def _run(command, work_dir, label):
            run_logged(command, work_dir, label, prefix="explicit_")
        
        
        def _channels(path):
            result = subprocess.run(
                ["ffprobe", "-v", "error", "-select_streams", "a:0", "-show_entries",
                 "stream=channels", "-of", "default=nw=1:nk=1", str(path)],
                capture_output=True, text=True, timeout=600,
            )
            try:
                channels = int(result.stdout.strip())
            except ValueError as exc:
                raise ValueError("narration channel probe failed") from exc
            if result.returncode or channels not in {1, 2}:
                raise ValueError("explicit narration supports only mono or stereo input")
            return channels
        
        
        def render_explicit_mix(context, narration_context, tts_segments, work_dir):
            """Render complete direct 48 kHz placements, voice bus, premaster, and fixed master."""
            directory = Path(work_dir).resolve() / ".explicit_audio_mix"
            directory.mkdir(parents=True, exist_ok=False)
            by_index = {item["index"]: item for item in narration_context["segments"]}
            segment_memory = {item["index"]: item for item in tts_segments}
            rendered = []
            for position, adopted in enumerate(context["segments"]):
                source = by_index[adopted["index"]]["snapshot"]["path"]
                channels = _channels(source)
                placed = directory / f"placed_{position:04d}_{adopted['index']}.wav"
                if channels == 1:
                    channel_filter = (
                        "aformat=sample_fmts=flt,"
                        "pan=stereo|c0=0.7071067811865476*c0|c1=0.7071067811865476*c0,"
                        "aresample=48000:async=0:first_pts=0,aformat=sample_fmts=flt:"
                        "sample_rates=48000:channel_layouts=stereo"
                    )
                    matrix = "mono_equal_power"
                else:
                    channel_filter = (
                        "aresample=48000:async=0:first_pts=0,aformat=sample_fmts=flt:"
                        "sample_rates=48000:channel_layouts=stereo"
                    )
                    matrix = "stereo_identity"
                _run(["ffmpeg", "-nostdin", "-v", "error", "-n", "-i", str(source),
                      "-map", "0:a:0", "-af", channel_filter, "-c:a", CODEC, str(placed)],
                     directory, f"place_{position:04d}")
                facts = source_score._output_facts(placed)
                start = adopted["output_start_sample"]
                end = start + facts["pcm"]["samples"]
                if end > context["format"]["total_samples"]:
                    raise ValueError("complete narration segment does not fit the adopted mix clock")
                rendered.append({**adopted, "output_end_sample": end, "input_channels": channels,
                                 "channel_matrix": matrix, "placed": facts})
                memory = segment_memory[adopted["index"]]
                memory.update({
                    "placed_audio_path": str(placed),
                    "narration_conversion_path": str(placed),
                    "audio_duration": facts["pcm"]["samples"] / RATE,
                    "placed_audio_duration": facts["pcm"]["samples"] / RATE,
                    "actual_place_start": start / RATE, "actual_place_end": end / RATE,
                    "output_start_sample": start, "output_end_sample": end,
                    "adopted_gain": adopted["gain"],
                    "global_narration_speed": 1.0, "segment_tempo_factor": 1.0,
                    "effective_tempo": 1.0, "fit_status": "placed", "blocking": False,
                    "truncated": False, "truncate_reason": "none",
                })
            ordered = sorted(rendered, key=lambda item: item["output_start_sample"])
            if any(current["output_start_sample"] < previous["output_end_sample"]
                   for previous, current in zip(ordered, ordered[1:])):
                raise ValueError("explicit narration placements overlap")
        
            total = context["format"]["total_samples"]
            voice_bus = directory / "voice_bus.wav"
            command = ["ffmpeg", "-nostdin", "-v", "error", "-n"]
            for item in rendered:
                command += ["-i", item["placed"]["path"]]
            filters = [f"anullsrc=r={RATE}:cl=stereo,atrim=end_sample={total}[base]"]
            labels = ["[base]"]
            for position, item in enumerate(rendered):
                filters.append(
                    f"[{position}:a]volume={item['gain']:.17g},"
                    f"adelay={item['output_start_sample']}S:all=1[v{position}]"
                )
                labels.append(f"[v{position}]")
            filters.append("".join(labels) + f"amix=inputs={len(labels)}:duration=first:normalize=0,"
                           f"atrim=end_sample={total},aformat=sample_fmts=flt:sample_rates={RATE}:"
                           "channel_layouts=stereo[out]")
            command += ["-filter_complex", ";".join(filters), "-map", "[out]", "-c:a", CODEC,
                        str(voice_bus)]
            _run(command, directory, "voice_bus")
            voice_facts = source_score._output_facts(voice_bus)
            narration_binding.seal_render_inputs(narration_context, tts_segments, voice_bus)
            narration_context["sealed"]["narration_bus"]["consumption_status"] = \
                "CONSUMED_BY_EXPLICIT_MIX"
        
            prepared = context["prepared"]["prepared_bed.wav"]["path"]
            premaster = directory / "premaster.wav"
            _run(["ffmpeg", "-nostdin", "-v", "error", "-n", "-i", prepared, "-i", str(voice_bus),
                  "-filter_complex", f"[0:a][1:a]amix=inputs=2:duration=first:normalize=0,"
                  f"atrim=end_sample={total},aformat=sample_fmts=flt:sample_rates={RATE}:"
                  "channel_layouts=stereo[out]", "-map", "[out]", "-c:a", CODEC, str(premaster)],
                 directory, "premaster")
            master = directory / "master.wav"
            gain = 10 ** (context["master_gain_db"] / 20)
            _run(["ffmpeg", "-nostdin", "-v", "error", "-n", "-i", str(premaster),
                  "-af", f"volume={gain:.17g},aformat=sample_fmts=flt:sample_rates={RATE}:"
                  "channel_layouts=stereo", "-c:a", CODEC, str(master)], directory, "master")
            runtime = {
                "segments": rendered, "voice_bus": voice_facts,
                "premaster": source_score._output_facts(premaster),
                "master": source_score._output_facts(master), "master_gain_linear": gain,
            }
            if any(runtime[key]["pcm"]["samples"] != total
                   for key in ("voice_bus", "premaster", "master")):
                raise ValueError("explicit mix derivative sample count differs from picture clock")
            context["runtime"] = runtime
            return runtime
        
        
        def finalize_binding(context, narration_record, rendered_output, final_output, work_dir):
            """Probe the rendered candidate and write ``audio_mix_binding.json`` into work_dir."""
            if not context.get("runtime"):
                raise RuntimeError("explicit audio mix must be rendered before finalization")
            rendered = require_local_path(rendered_output, "rendered output")
            output_picture = probe_picture(rendered)
            input_clock = _picture_format(context["picture_identity"])[1]
            output_clock = _picture_format(output_picture)[1]
            for key in ("fps", "frame_count", "duration", "start"):
                if output_clock[key] != input_clock[key]:
                    raise ValueError("rendered output picture frame clock changed")
            packet_identity = (
                "EXACT" if output_picture == context["picture_identity"]
                else "REENCODED_CLOCK_MATCH"
            )
            audio = probe_audio_packets(rendered, 0)
            validate_pair_timing(output_picture, audio)
            decoded = Path(context["runtime"]["master"]["path"]).parent / "final_aac_decoded.wav"
            _run(["ffmpeg", "-nostdin", "-v", "error", "-n", "-i", str(rendered),
                  "-map", "0:a:0", "-af", "aformat=sample_fmts=flt:sample_rates=48000:"
                  "channel_layouts=stereo", "-c:a", CODEC, str(decoded)], decoded.parent,
                 "final_decode")
            context["runtime"]["final_decoded_pcm"] = source_score._output_facts(decoded)
            report = {
                "artifact": ARTIFACT, "schema_version": 1, "status": "FINALIZED",
                "adoption": {"path": context["path"]},
                "picture": context["picture"],
                "output_picture": {
                    **output_clock, "packet_identity": packet_identity,
                },
                "prepared_receipt": context["prepared_receipt"],
                "prepared": context["prepared"], "format": context["format"],
                "segments": context["runtime"]["segments"],
                "voice_bus": context["runtime"]["voice_bus"],
                "premaster": context["runtime"]["premaster"],
                "master": {**context["runtime"]["master"], "gain_db": context["master_gain_db"],
                           "gain_linear": context["runtime"]["master_gain_linear"]},
                "narration_input_binding": {"path": narration_record["path"], "status": "FINALIZED"},
                "final_output": {
                    "path": str(Path(final_output).resolve()),
                    "decoded_pcm": context["runtime"]["final_decoded_pcm"],
                    "audio_stream": {
                        "decoder": audio["decoder"], "packet_count": audio["packet_count"],
                        "payload_bytes": audio["payload_bytes"],
                        "start_time": audio["start_time"], "duration": audio["duration"],
                    },
                },
                "direct_listening": "NOT_CHECKED", "normal_speed_review": "NOT_CHECKED",
                "release_approved": False,
            }
            write_json_atomic(Path(work_dir).resolve() / FILENAME, report)
            return report
        
        
        def binding_record(work_dir):
            """The finalized mix binding's path and status; None when absent or malformed."""
            path = Path(work_dir) / FILENAME
            if not path.is_file():
                return None
            try:
                report = json.loads(path.read_text(encoding="utf-8"))
            except (UnicodeDecodeError, json.JSONDecodeError):
                return None
            if not isinstance(report, dict) or report.get("artifact") != ARTIFACT \
                    or report.get("schema_version") != 1 or report.get("status") != "FINALIZED":
                return None
            final = report.get("final_output")
            if not isinstance(final, dict) or not Path(str(final.get("path", ""))).is_file():
                return None
            narration = report.get("narration_input_binding")
            if not isinstance(narration, dict) or not Path(str(narration.get("path", ""))).is_file():
                return None
            return {"path": str(path.resolve()), "status": "FINALIZED"}
        
      • frozen_audio.py 7.6 KB
        """Probe and verify an adopted AAC stream without decoding or rewriting it."""
        
        import json
        from fractions import Fraction
        from pathlib import Path
        
        from lib import run_cmd
        
        
        def _fraction(value):
            return Fraction(str(value))
        
        
        def _packet_span(audio):
            """The stream duration proved by its packets: last packet end minus first packet start."""
            packets = audio["packets"]
            if not packets or packets[0]["pts"] is None or packets[-1]["pts"] is None \
                    or packets[-1]["duration"] is None:
                return None
            return Fraction(packets[-1]["pts"]) + Fraction(packets[-1]["duration"]) \
                - Fraction(packets[0]["pts"])
        
        
        def _rational_ticks(value, time_base):
            if value in (None, "N/A"):
                return None
            value = Fraction(int(value)) * Fraction(time_base)
            return f"{value.numerator}/{value.denominator}"
        
        
        def _probe(path, *args):
            result = run_cmd([
                "ffprobe", "-v", "error", *args, "-of", "json", str(path),
            ])
            if result.returncode != 0:
                raise RuntimeError(f"无法探测媒体 {path}: {result.stderr.strip()}")
            try:
                return json.loads(result.stdout)
            except (TypeError, ValueError) as exc:
                raise RuntimeError(f"ffprobe 返回无效 JSON: {path}") from exc
        
        
        def probe_audio_packets(path, audio_stream_index):
            """Return packet sizes and rational timestamps for one audio stream.
        
            ``audio_stream_index`` is the zero-based audio-stream ordinal accepted by
            ffmpeg's ``0:a:N`` selector, not the file-wide absolute stream index.
            """
            payload = _probe(
                path,
                "-select_streams", f"a:{audio_stream_index}",
                "-show_streams", "-show_packets",
                "-show_entries",
                "stream=index,codec_name,time_base,start_pts,start_time,duration_ts,duration,sample_rate,channels,"
                "channel_layout:packet=pts,dts,duration,size,side_data_list",
            )
            streams = payload.get("streams", [])
            if len(streams) != 1:
                raise RuntimeError(f"找不到音频流 a:{audio_stream_index}: {Path(path)}")
            stream = streams[0]
            time_base = stream.get("time_base")
            if not time_base:
                raise RuntimeError(f"音频流 a:{audio_stream_index} 缺少 time_base")
            codec = stream.get("codec_name")
            sample_rate = int(stream["sample_rate"]) if stream.get("sample_rate") else None
            channels = int(stream["channels"]) if stream.get("channels") else None
            channel_layout = stream.get("channel_layout")
            decoder = {
                "codec": codec,
                "sample_rate": sample_rate,
                "channels": channels,
                "channel_layout": channel_layout,
            }
            packets = []
            for packet in payload.get("packets", []):
                if packet.get("size") is None:
                    raise RuntimeError(f"音频流 a:{audio_stream_index} 的 packet 缺少 size")
                packets.append({
                    "size": int(packet["size"]),
                    "pts": _rational_ticks(packet.get("pts"), time_base),
                    "dts": _rational_ticks(packet.get("dts"), time_base),
                    "duration": _rational_ticks(packet.get("duration"), time_base),
                    "side_data_list": packet.get("side_data_list", []),
                })
            return {
                "selected_audio_stream_index": audio_stream_index,
                "absolute_stream_index": int(stream["index"]),
                "codec": codec,
                "time_base": time_base,
                "start_time": stream.get("start_time"),
                "duration": stream.get("duration"),
                "sample_rate": sample_rate,
                "channels": channels,
                "channel_layout": channel_layout,
                "decoder": decoder,
                "packet_count": len(packets),
                "payload_bytes": sum(packet["size"] for packet in packets),
                "packets": packets,
            }
        
        
        def _picture_interval(path):
            payload = _probe(
                path,
                "-select_streams", "v:0", "-show_streams",
                "-show_entries", "stream=time_base,start_pts,start_time,duration_ts,duration,avg_frame_rate",
            )
            streams = payload.get("streams", [])
            if len(streams) != 1:
                raise RuntimeError(f"找不到主画面流 v:0: {Path(path)}")
            stream = streams[0]
            start = _fraction(stream.get("start_time", "0"))
            if stream.get("duration") in (None, "N/A"):
                raise RuntimeError("主画面流缺少可验证时长")
            duration = _fraction(stream["duration"])
            frame_rate = Fraction(stream.get("avg_frame_rate", "0/1"))
            frame = Fraction(1, 1000) if frame_rate <= 0 else 1 / frame_rate
            return start, duration, frame
        
        
        def validate_adopted_source(path, audio_stream_index):
            """Fail unless the selected input is copyable AAC spanning the picture interval."""
            audio = probe_audio_packets(path, audio_stream_index)
            if audio["codec"] != "aac":
                raise RuntimeError(
                    f"adopted-packet-copy 当前只支持 MP4 AAC stream-copy;"
                    f"a:{audio_stream_index} 是 {audio['codec'] or 'unknown'}"
                )
            if not audio["packets"]:
                raise RuntimeError(f"音频流 a:{audio_stream_index} 没有 packet,不能采用")
            picture_start, picture_duration, frame_tolerance = _picture_interval(path)
            if audio["start_time"] in (None, "N/A") or audio["duration"] in (None, "N/A"):
                raise RuntimeError(f"音频流 a:{audio_stream_index} 缺少可验证起止时间")
            audio_start = _fraction(audio["start_time"])
            audio_duration = _fraction(audio["duration"])
            packet_durations = [
                _fraction(packet["duration"])
                for packet in audio["packets"] if packet["duration"] is not None
            ]
            tolerance = max([frame_tolerance, Fraction(1, 1000), *packet_durations])
            if (
                abs(audio_start - picture_start) > tolerance
                or abs(audio_duration - picture_duration) > tolerance
            ):
                raise RuntimeError(
                    "采用音频与画面时长/起点不兼容: "
                    f"picture={float(picture_start):.6f}+{float(picture_duration):.6f}s, "
                    f"audio={float(audio_start):.6f}+{float(audio_duration):.6f}s"
                )
            return audio
        
        
        def verify_adopted_audio(input_path, output_path, input_stream_index, output_stream_index=0):
            """Probe the rendered output and check packet count, payload bytes, duration and timing."""
            expected = probe_audio_packets(input_path, input_stream_index)
            actual = probe_audio_packets(output_path, output_stream_index)
            if expected["decoder"] != actual["decoder"]:
                raise RuntimeError("采用音频 decoder 参数已改变")
            if expected["packet_count"] != actual["packet_count"]:
                raise RuntimeError(
                    f"采用音频 packet count 已改变: {expected['packet_count']} -> {actual['packet_count']}"
                )
            if expected["payload_bytes"] != actual["payload_bytes"]:
                raise RuntimeError("采用音频 packet payload 总字节数不一致")
            if _packet_span(expected) != _packet_span(actual):
                raise RuntimeError("采用音频流时长已改变")
            expected_packet_core = [
                {key: value for key, value in packet.items() if key != "side_data_list"}
                for packet in expected["packets"]
            ]
            actual_packet_core = [
                {key: value for key, value in packet.items() if key != "side_data_list"}
                for packet in actual["packets"]
            ]
            if expected_packet_core != actual_packet_core:
                raise RuntimeError("采用音频 packet PTS/DTS/duration/size 不一致")
            if [packet["side_data_list"] for packet in expected["packets"]] != [
                packet["side_data_list"] for packet in actual["packets"]
            ]:
                raise RuntimeError("采用音频 packet side data 已改变")
            return {
                "verified": True,
                "selected_audio_stream_index": input_stream_index,
                "output_audio_stream_index": output_stream_index,
                "input": expected,
                "output": actual,
            }
        
      • narration_binding.py 14.5 KB
        """Record which narration inputs an assembly consumed: adoption, snapshot, placement, output."""
        
        import json
        import math
        from pathlib import Path
        import shutil
        import subprocess
        
        from adoption.frozen_audio import probe_audio_packets
        from adoption.strict_inputs import (
            read_json_bytes, require_fields, require_local_path, without_digests, write_json_atomic,
        )
        
        
        ARTIFACT = "narration_input_binding"
        FILENAME = "narration_input_binding.json"
        IDENTITY_STATUSES = frozenset({"UNADOPTED", "BOUND_TO_ADOPTION"})
        # The conservative default an adoption gets when it declares nothing stronger.
        TEMPO_POLICY = {
            "global_atempo": 1.0,
            "bounded_segment_fit": False,
            "segment_tempo_max": 1.0,
            "cumulative_tempo_max": 1.0,
            "cumulative_tempo_hard_max": 1.0,
        }
        TEMPO_NUMBERS = ("global_atempo", "segment_tempo_max", "cumulative_tempo_max",
                         "cumulative_tempo_hard_max")
        
        
        def validate_tempo_policy(value):
            """Accept any adoption-declared tempo policy whose shape and bounds hold.
        
            The adoption, not this module, decides how fast its own narration may be
            played; the module only refuses policies that are malformed or that would
            let a segment exceed the cumulative hard ceiling the policy itself declares.
            """
            require_fields(value, sorted(TEMPO_POLICY), "tempo_policy")
            policy = {}
            for key in TEMPO_NUMBERS:
                number = value[key]
                if type(number) not in (int, float) or not math.isfinite(number):
                    raise ValueError(f"tempo_policy {key} must be a finite number")
                policy[key] = float(number)
            if type(value["bounded_segment_fit"]) is not bool:
                raise ValueError("tempo_policy bounded_segment_fit must be a boolean")
            policy["bounded_segment_fit"] = value["bounded_segment_fit"]
            if policy["cumulative_tempo_max"] < 1.0:
                raise ValueError("tempo_policy cumulative_tempo_max must be at least 1.0")
            if policy["cumulative_tempo_hard_max"] < policy["cumulative_tempo_max"]:
                raise ValueError("tempo_policy cumulative_tempo_hard_max must not be below "
                                 "cumulative_tempo_max")
            if policy["segment_tempo_max"] < 1.0:
                raise ValueError("tempo_policy segment_tempo_max must be at least 1.0")
            if not 0 < policy["global_atempo"] <= policy["cumulative_tempo_hard_max"]:
                raise ValueError("tempo_policy global_atempo must be positive and within "
                                 "cumulative_tempo_hard_max")
            return policy
        
        
        def load_adoption(path, *, tts_meta_path, tts_segments):
            """Load a strict v1 adoption and check it against the current tts_meta segments."""
            if tts_meta_path is None:
                raise ValueError("narration adoption requires explicit tts_meta_path")
            adoption_path, _, adoption = read_json_bytes(path, "narration adoption")
            adoption = without_digests(adoption, "narration adoption")
            require_fields(adoption, ["artifact", "schema_version", "segments", "tempo_policy"],
                           "narration adoption")
            if adoption["artifact"] != "narration_adoption" or type(adoption["schema_version"]) is not int \
                    or adoption["schema_version"] != 1:
                raise ValueError("unsupported narration_adoption schema")
            tts_meta_path, _, tts_meta = read_json_bytes(tts_meta_path, "tts_meta")
            if not isinstance(tts_meta, dict) or not isinstance(tts_meta.get("segments"), list):
                raise ValueError("tts_meta requires a segments list")
            if tts_meta["segments"] != tts_segments:
                raise ValueError("in-memory narration segments differ from bound tts_meta")
            tempo_policy = validate_tempo_policy(adoption["tempo_policy"])
            if not isinstance(adoption["segments"], list) or len(adoption["segments"]) != len(tts_segments):
                raise ValueError("adoption segments must exactly cover tts_meta segments")
            normalized_segments = []
            for adopted, actual in zip(adoption["segments"], tts_segments):
                adopted = without_digests(adopted, "adoption segment")
                require_fields(adopted, ["index", "spoken_text", "requested_provider", "requested_voice"],
                               "adoption segment")
                if type(adopted["index"]) is not int or adopted["index"] != actual.get("index"):
                    raise ValueError("adoption segment index/order differs from tts_meta")
                spoken = actual.get("spoken_text", actual.get("narration"))
                if not isinstance(adopted["spoken_text"], str) or adopted["spoken_text"] != spoken:
                    raise ValueError("adoption spoken_text differs from tts_meta")
                for key in ("requested_provider", "requested_voice"):
                    if not isinstance(adopted[key], str) or not adopted[key]:
                        raise ValueError(f"adoption {key} must be a non-empty string")
                normalized_segments.append(dict(adopted))
            return {
                "path": str(adoption_path),
                "tts_meta": {"path": str(tts_meta_path)},
                "segments": normalized_segments, "tempo_policy": tempo_policy,
            }
        
        
        def _copy_snapshot(source, destination):
            """Copy seam kept small so tests can inject a post-preflight source mutation."""
            shutil.copyfile(source, destination)
        
        
        def prepare_binding(tts_segments, work_dir, *, narration_adoption_path=None,
                            tts_meta_path=None):
            """Validate first, then snapshot the adopted narration inputs into work_dir."""
            if not isinstance(tts_segments, list):
                raise ValueError("tts_segments must be a list")
            adoption = (
                load_adoption(narration_adoption_path, tts_meta_path=tts_meta_path,
                              tts_segments=tts_segments)
                if narration_adoption_path is not None else None
            )
            if not adoption:
                originals = [
                    {"index": segment["index"], "source": Path(segment["audio_path"]).resolve(),
                     "spoken_text": segment["spoken_text"]}
                    for segment in tts_segments
                ]
                return {
                    "identity_status": "UNADOPTED", "adoption": None, "tempo_policy": None,
                    "originals": originals, "active": False, "segments": [],
                }
            originals = []
            seen_indices = set()
            for position, segment in enumerate(tts_segments):
                if not isinstance(segment, dict) or type(segment.get("index")) is not int:
                    raise ValueError("each narration segment requires an integer index")
                if segment["index"] in seen_indices:
                    raise ValueError("narration segment indices must be unique")
                seen_indices.add(segment["index"])
                adopted = adoption["segments"][position]
                source = require_local_path(segment["audio_path"], "narration audio")
                originals.append({
                    "index": segment["index"], "source": source,
                    "spoken_text": segment["spoken_text"],
                    "requested_provider": adopted["requested_provider"],
                    "requested_voice": adopted["requested_voice"],
                })
            context = {
                "identity_status": "BOUND_TO_ADOPTION", "adoption": adoption,
                "tempo_policy": adoption["tempo_policy"],
                "originals": originals, "active": True, "segments": [], "sealed": None,
            }
            snapshot_dir = Path(work_dir).resolve() / ".narration_input_snapshots"
            snapshot_dir.mkdir(parents=True, exist_ok=False)
            for position, (segment, original) in enumerate(zip(tts_segments, originals)):
                suffix = original["source"].suffix or ".audio"
                snapshot = snapshot_dir / f"segment_{position:04d}_{segment['index']}{suffix}"
                _copy_snapshot(original["source"], snapshot)
                segment["audio_path"] = str(snapshot)
                segment["narration_input_original"] = {"path": str(original["source"])}
                segment["narration_input_snapshot"] = {"path": str(snapshot)}
                context["segments"].append({
                    "index": original["index"], "spoken_text": original["spoken_text"],
                    "requested_provider": original["requested_provider"],
                    "requested_voice": original["requested_voice"],
                    "original": dict(segment["narration_input_original"]),
                    "snapshot": dict(segment["narration_input_snapshot"]),
                })
            return context
        
        
        def _pcm(path):
            result = subprocess.run(
                ["ffprobe", "-v", "error", "-select_streams", "a:0", "-show_streams",
                 "-show_entries", "stream=codec_name,sample_fmt,sample_rate,channels,channel_layout",
                 "-of", "json", str(path)], capture_output=True, text=True, timeout=600,
            )
            if result.returncode:
                raise ValueError(f"audio probe failed: {result.stderr.strip()}")
            streams = json.loads(result.stdout).get("streams", [])
            if len(streams) != 1:
                raise ValueError("expected exactly one audio stream")
            stream = streams[0]
            return {key: stream.get(key) for key in (
                "codec_name", "sample_fmt", "sample_rate", "channels", "channel_layout"
            )}
        
        
        def _asset(path, *, pcm=False):
            path = require_local_path(path, "binding asset")
            value = {"path": str(path)}
            if pcm:
                value["pcm"] = _pcm(path)
            return value
        
        
        def seal_render_inputs(context, tts_segments, narration_wav):
            """Record every derived audio file that the final FFmpeg command will consume."""
            if not context.get("active"):
                return None
            by_index = {segment["index"]: segment for segment in tts_segments}
            sealed_segments = []
            for item in context["segments"]:
                segment = by_index[item["index"]]
                placed = segment.get("placed_audio_path")
                if not placed:
                    raise RuntimeError("active narration binding has no complete placed audio")
                conversion = segment.get("narration_conversion_path")
                conversion_asset = (
                    {"applied": True, **_asset(conversion, pcm=True)}
                    if conversion else {"applied": False}
                )
                sealed_segments.append({
                    "index": item["index"],
                    "conversion": conversion_asset,
                    "placed": _asset(placed, pcm=True),
                })
            context["sealed"] = {
                "segments": sealed_segments,
                "narration_bus": _asset(narration_wav, pcm=True),
            }
            return context["sealed"]
        
        
        def _active_report(context, rendered_output, final_output):
            if not context.get("sealed"):
                raise RuntimeError("active narration inputs must be sealed before finalization")
            sealed = {item["index"]: item for item in context["sealed"]["segments"]}
            segments = []
            for item in context["segments"]:
                segments.append({
                    **item,
                    "original": _asset(item["original"]["path"], pcm=True),
                    "snapshot": _asset(item["snapshot"]["path"], pcm=True),
                    "conversion": sealed[item["index"]]["conversion"],
                    "placed": sealed[item["index"]]["placed"],
                })
            rendered_path = require_local_path(rendered_output, "rendered output")
            final_audio = probe_audio_packets(rendered_path, 0)
            return {
                "artifact": ARTIFACT, "schema_version": 1, "status": "FINALIZED",
                "identity_status": context["identity_status"], "adoption": context["adoption"],
                "segments": segments, "narration_bus": context["sealed"]["narration_bus"],
                "final_output": {
                    "path": str(Path(final_output).resolve()),
                    "audio_stream": {
                        "decoder": final_audio["decoder"], "packet_count": final_audio["packet_count"],
                        "payload_bytes": final_audio["payload_bytes"],
                        "start_time": final_audio["start_time"], "duration": final_audio["duration"],
                    },
                },
                "voice_authentication": "NOT_CHECKED", "direct_listening": "NOT_CHECKED",
            }
        
        
        def finalize_binding(context, tts_segments, narration_wav, final_output, *, rendered_output=None):
            """Write ``narration_input_binding.json`` beside ``narration_wav`` and return it.
        
            ``rendered_output`` is the candidate file to probe when the published
            ``final_output`` path does not exist yet; it defaults to ``final_output``.
            """
            del tts_segments  # already sealed; the record describes what was consumed
            if not context.get("active"):
                report = {
                    "artifact": ARTIFACT, "schema_version": 1, "status": "FINALIZED",
                    "identity_status": "UNADOPTED", "adoption": None,
                    "segments": [
                        {
                            "index": original["index"], "spoken_text": original["spoken_text"],
                            "requested_provider": None, "requested_voice": None,
                            "original": {"path": str(original["source"])},
                            "snapshot": None, "conversion": {"applied": False},
                            "placed": None,
                        }
                        for original in context["originals"]
                    ],
                    "narration_bus": ({"path": str(Path(narration_wav).resolve())}
                                      if Path(narration_wav).is_file() else None),
                    "final_output": ({"path": str(Path(final_output).resolve())}
                                     if Path(final_output).is_file() else None),
                    "voice_authentication": "NOT_CHECKED", "direct_listening": "NOT_CHECKED",
                }
            else:
                report = _active_report(
                    context, final_output if rendered_output is None else rendered_output, final_output
                )
            write_json_atomic(Path(narration_wav).resolve().parent / FILENAME, report)
            return report
        
        
        def record_of(report, path):
            """The manifest/QC summary of one finalized binding report stored at ``path``."""
            return {"path": str(Path(path).resolve()),
                    "identity_status": report["identity_status"],
                    "tempo_policy": (
                        report["adoption"]["tempo_policy"]
                        if report["identity_status"] == "BOUND_TO_ADOPTION" else None
                    )}
        
        
        def binding_record(work_dir):
            """The finalized binding's path, identity status and tempo policy; None when absent/invalid."""
            path = Path(work_dir) / FILENAME
            if not path.is_file():
                return None
            try:
                value = json.loads(path.read_text(encoding="utf-8"))
            except (UnicodeDecodeError, json.JSONDecodeError):
                return None
            if (
                not isinstance(value, dict)
                or value.get("artifact") != ARTIFACT
                or type(value.get("schema_version")) is not int
                or value.get("schema_version") != 1
                or value.get("status") != "FINALIZED"
                or value.get("identity_status") not in IDENTITY_STATUSES
            ):
                return None
            if value["identity_status"] == "BOUND_TO_ADOPTION":
                adoption = value.get("adoption")
                if not isinstance(adoption, dict):
                    return None
                try:
                    validate_tempo_policy(adoption.get("tempo_policy"))
                except ValueError:
                    return None
            final_output = value.get("final_output")
            if not isinstance(final_output, dict) or not Path(str(final_output.get("path", ""))).is_file():
                return None
            return record_of(value, path)
        
      • strict_inputs.py 4.1 KB
        """Shared strict-input helpers for the explicit-input assemble modules.
        
        Every helper here fails closed with ``ValueError`` on malformed input and never
        guesses: a path must be a local file, a JSON document must parse, a rational must
        be canonical ``N/D``. ``run_logged`` and ``probe_json`` wrap the external tools so
        each caller logs the same evidence.
        """
        
        from fractions import Fraction
        import json
        import math
        from pathlib import Path
        import subprocess
        
        
        def require_fields(value, required, label):
            if not isinstance(value, dict) or set(value) != set(required):
                raise ValueError(f"{label} requires exactly fields {required}")
        
        
        def without_digests(value, label):
            """Drop the ``sha256`` / ``*_sha256`` keys older caller JSON declared; they are ignored."""
            if not isinstance(value, dict):
                raise ValueError(f"{label} must be an object")
            return {key: item for key, item in value.items()
                    if key != "sha256" and not key.endswith("_sha256")}
        
        
        def require_integer(value, label, minimum=0):
            if type(value) is not int or value < minimum:
                raise ValueError(f"{label} must be an integer >= {minimum}")
            return value
        
        
        def require_number(value, label, minimum, maximum):
            if type(value) not in (int, float) or not math.isfinite(value) \
                    or not minimum <= value <= maximum:
                raise ValueError(f"{label} must be finite in [{minimum},{maximum}]")
            return float(value)
        
        
        def require_local_path(path, label):
            if not isinstance(path, (str, Path)) or not str(path) or "://" in str(path):
                raise ValueError(f"{label} requires a local path")
            resolved = Path(path).resolve()
            if not resolved.is_file():
                raise ValueError(f"{label} is missing: {resolved}")
            return resolved
        
        
        def require_declared_path(value, label):
            """Resolve a caller ``{"path": ...}`` declaration to an existing local file."""
            if not isinstance(value, dict):
                raise ValueError(f"{label} must be an object with a path")
            return require_local_path(value.get("path"), label)
        
        
        def read_json_bytes(path, label):
            """Return (resolved_path, raw_bytes, parsed) for one local JSON document."""
            resolved = require_local_path(path, label)
            raw = resolved.read_bytes()
            try:
                return resolved, raw, json.loads(raw)
            except (UnicodeDecodeError, json.JSONDecodeError) as exc:
                raise ValueError(f"{label} is not valid JSON") from exc
        
        
        def write_json_atomic(path, value):
            path = Path(path)
            temporary = path.with_suffix(".writing.json")
            temporary.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
            temporary.replace(path)
        
        
        def canonical_fraction(value, label):
            if not isinstance(value, str):
                raise ValueError(f"{label} must be a canonical rational string")
            try:
                result = Fraction(value)
            except (ValueError, ZeroDivisionError) as exc:
                raise ValueError(f"invalid {label}") from exc
            if result <= 0 or value != f"{result.numerator}/{result.denominator}":
                raise ValueError(f"{label} must be a positive canonical N/D rational")
            return result
        
        
        def probe_json(path, *ffprobe_args):
            result = subprocess.run(
                ["ffprobe", "-v", "error", *ffprobe_args, "-of", "json", str(path)],
                capture_output=True, text=True, timeout=600,
            )
            if result.returncode or result.stderr.strip():
                raise ValueError(f"media probe failed for {path}: {result.stderr.strip()}")
            try:
                return json.loads(result.stdout)
            except json.JSONDecodeError as exc:
                raise ValueError(f"invalid media probe JSON for {path}") from exc
        
        
        def run_logged(command, directory, label, *, prefix="", timeout=3600):
            """Run one FFmpeg command, keeping ``<prefix><label>.command.json`` and ``.log``."""
            directory = Path(directory)
            name = f"{prefix}{label}"
            write_json_atomic(directory / f"{name}.command.json", command)
            result = subprocess.run(command, capture_output=True, text=True, timeout=timeout)
            (directory / f"{name}.log").write_text(result.stderr, encoding="utf-8")
            if result.returncode:
                raise RuntimeError(f"{name} FFmpeg failed; see {name}.log")
            return result
        
      • strict_publish.py 6 KB
        """Publish transaction for strict-adoption renders.
        
        An active narration binding renders to a hidden candidate file. Its binding (and,
        for an explicit mix, the audio mix binding) is written, QC gates the candidate,
        the media file is published, and QC runs once more against the published path.
        Any failure removes the candidate, the published file and the bindings written
        by this render, so a strict output never exists without its consumed-input record.
        """
        
        from pathlib import Path
        
        import assembly_contract
        import adoption.audio_mix_binding as audio_mix_binding
        import adoption.narration_binding as narration_binding
        
        
        def current_narration_binding(work_dir, audio_mode):
            """Read the narration record only for the explicit narration render path."""
            if audio_mode != "narration":
                return None
            return narration_binding.binding_record(work_dir)
        
        
        def current_audio_mix_binding(work_dir, audio_mode):
            if audio_mode != "narration":
                return None
            return audio_mix_binding.binding_record(work_dir)
        
        
        def publish_render(*, work_dir, binding, explicit_mix, tts_segments, narration_wav,
                           render_output, published_output, audio_mode, audio_operations,
                           adopted_audio, loudness_mode, loudnorm_measurement, visual_qc,
                           source_has_audio, video_duration, render_delivery, source_audio_status):
            """Gate the rendered candidate on assembly QC and publish it with its bindings.
        
            Returns the final output path. Inactive-binding renders only write their
            binding and QC; strict renders write bindings, publish, and re-run QC, or roll
            everything back.
            """
            active = bool(binding and binding["active"])
            work_dir = Path(work_dir)
            binding_path = work_dir / narration_binding.FILENAME
            mix_binding_path = work_dir / audio_mix_binding.FILENAME
            binding_written = False
            mix_binding_written = False
            try:
                if active:
                    report = narration_binding.finalize_binding(
                        binding, tts_segments,
                        (work_dir / "narration.wav" if explicit_mix is not None else narration_wav),
                        published_output, rendered_output=render_output,
                    )
                    binding_written = True
                    if not binding_path.is_file():
                        raise RuntimeError("narration binding 未写入")
                    narration_record = narration_binding.record_of(report, binding_path)
                    mix_record = None
                    if explicit_mix is not None:
                        audio_mix_binding.finalize_binding(
                            explicit_mix, narration_record, render_output, published_output, work_dir
                        )
                        mix_binding_written = True
                        mix_record = {"path": str(mix_binding_path.resolve()), "status": "FINALIZED"}
                elif binding:
                    narration_binding.finalize_binding(
                        binding, tts_segments, narration_wav, render_output
                    )
                    narration_record = current_narration_binding(work_dir, audio_mode)
                    mix_record = None
                else:
                    narration_record = current_narration_binding(work_dir, audio_mode)
                    mix_record = None
                assembly_qc = assembly_contract._build_assembly_qc(
                    tts_segments,
                    video_duration,
                    output_path=render_output,
                    source_has_audio=source_has_audio,
                    loudness_mode=loudness_mode,
                    loudnorm_measurement=loudnorm_measurement,
                    visual_qc=visual_qc,
                    audio_mode=audio_mode,
                    audio_operations=audio_operations,
                    adopted_audio=adopted_audio,
                    narration_input_binding=narration_record,
                    audio_mix_binding=mix_record,
                    source_audio_status=source_audio_status,
                    render_delivery=render_delivery,
                )
                if active and assembly_qc["blocking"]:
                    assembly_contract._write_assembly_qc(work_dir, assembly_qc)
                    codes = ", ".join(assembly_qc["blocking_codes"])
                    raise RuntimeError(f"身份约束渲染 QC 失败: {codes}")
                if active:
                    render_output.rename(published_output)
                    render_output = published_output
                    current_binding = current_narration_binding(work_dir, audio_mode)
                    if current_binding is None:
                        raise RuntimeError("已发布 narration binding 未通过终态检查")
                    current_mix_binding = current_audio_mix_binding(work_dir, audio_mode)
                    if explicit_mix is not None and current_mix_binding is None:
                        raise RuntimeError("已发布 audio mix binding 未通过终态检查")
                    assembly_qc = assembly_contract._build_assembly_qc(
                        tts_segments, video_duration, output_path=render_output,
                        source_has_audio=source_has_audio,
                        loudness_mode=loudness_mode, loudnorm_measurement=loudnorm_measurement,
                        visual_qc=visual_qc, audio_mode=audio_mode,
                        audio_operations=audio_operations, adopted_audio=adopted_audio,
                        narration_input_binding=current_binding,
                        audio_mix_binding=current_mix_binding,
                        source_audio_status=source_audio_status,
                        render_delivery=assembly_qc["delivery_qc"],
                    )
                    if assembly_qc["blocking"]:
                        assembly_qc["output"] = {
                            "path": str(published_output), "exists": False, "bytes": 0,
                        }
                        assembly_contract._write_assembly_qc(work_dir, assembly_qc)
                        codes = ", ".join(assembly_qc["blocking_codes"])
                        raise RuntimeError(f"身份约束渲染终态 QC 失败: {codes}")
                assembly_contract._write_assembly_qc(work_dir, assembly_qc)
            except Exception:
                if active:
                    render_output.unlink(missing_ok=True)
                    published_output.unlink(missing_ok=True)
                    if binding_written:
                        binding_path.unlink(missing_ok=True)
                    if mix_binding_written:
                        mix_binding_path.unlink(missing_ok=True)
                raise
            return render_output
        
      • __init__.py 94 B
        """Adoption family: narration/audio-mix binding + strict publish for adopted-source audio."""
        
    • jianying
      • builders.py 30.1 KB
        """Production material and segment builders for the JianYing exporter."""
        
        import json
        import os
        from copy import deepcopy
        
        from jianying.schema import scrub_platform_identity, us
        from jianying.templates import template
        from jianying.tracks import SEGMENT_RENDER_INDEX
        
        
        def timerange(start_us, dur_us):
            return {"start": int(start_us), "duration": int(dur_us)}
        
        
        def speed_material(new_id, speed=1.0):
            return {
                "id": new_id(),
                "speed": float(speed),
                "type": "speed",
                "mode": 0,
                "curve_speed": None,
            }
        
        
        def volume_keyframes(keyframes, seg_start_s, new_id):
            """Build one KFTypeVolume keyframe list from timeline-absolute points."""
            if not keyframes:
                return []
            kfs = []
            for kf in keyframes:
                kfs.append({
                    "curveType": "Line",
                    "graphID": "",
                    "left_control": {"x": 0.0, "y": 0.0},
                    "right_control": {"x": 0.0, "y": 0.0},
                    "id": new_id(),
                    "time_offset": max(0, us(kf["t"] - seg_start_s)),
                    "values": [round(float(kf["gain"]), 4)],
                })
            return [{
                "id": new_id(),
                "keyframe_list": kfs,
                "material_id": "",
                "property_type": "KFTypeVolume",
            }]
        
        
        def windowed_volume_keyframes(keyframes, seg_start_s, seg_end_s, default_gain, new_id):
            """Window timeline-absolute keyframes for one split/looped segment."""
            if not keyframes:
                return []
            start = float(seg_start_s)
            end = float(seg_end_s)
            default_gain = float(default_gain)
            ordered = sorted((
                {"t": float(kf["t"]), "gain": float(kf["gain"])}
                for kf in keyframes
                if "t" in kf and "gain" in kf
            ), key=lambda kf: kf["t"])
            if not ordered or end <= start:
                return []
        
            start_gain = default_gain
            for kf in ordered:
                if kf["t"] <= start:
                    start_gain = kf["gain"]
                else:
                    break
            inner = [kf for kf in ordered if start <= kf["t"] <= end]
            if not inner and abs(start_gain - default_gain) < 1e-4:
                return []
        
            selected = [{"t": start, "gain": start_gain}]
            for kf in inner:
                if abs(kf["t"] - start) < 1e-4:
                    selected[-1] = {"t": start, "gain": kf["gain"]}
                else:
                    selected.append(kf)
            if all(abs(kf["gain"] - default_gain) < 1e-4 for kf in selected):
                return []
            return volume_keyframes(selected, start, new_id)
        
        
        def clip_from_segment(segment):
            """Map optional authoring transforms (validated as objects by the timeline contract)."""
            scale = segment.get("scale", {})
            position = segment.get("position", {})
            flip = segment.get("flip", {})
            return {
                "alpha": round(float(segment.get("opacity", 1.0)), 4),
                "flip": {
                    "horizontal": bool(flip.get("horizontal", False)),
                    "vertical": bool(flip.get("vertical", False)),
                },
                "rotation": float(segment.get("rotation_degrees", 0.0)),
                "scale": {
                    "x": float(scale.get("x", 1.0)),
                    "y": float(scale.get("y", 1.0)),
                },
                # Timeline v2 uses JianYing's normalized, canvas-center, Y-up coordinates.
                "transform": {
                    "x": float(position.get("x", 0.0)),
                    "y": float(position.get("y", 0.0)),
                },
            }
        
        
        def base_segment(material_id, target_start_us, target_dur_us, volume, keyframes, new_id):
            segment = template("segment")
            segment.update({
                "id": new_id(),
                "material_id": material_id,
                "target_timerange": timerange(target_start_us, target_dur_us),
                "common_keyframes": keyframes,
                "track_render_index": SEGMENT_RENDER_INDEX,
                "render_index": SEGMENT_RENDER_INDEX,
                "volume": round(float(volume), 4),
            })
            return segment
        
        
        def audio_segment_piece(material_id, target_start_us, target_dur_us, source_start_us,
                                source_dur_us, volume, keyframes, new_id):
            seg = base_segment(material_id, target_start_us, target_dur_us, volume, keyframes, new_id)
            seg["source_timerange"] = timerange(source_start_us, source_dur_us)
            seg["extra_material_refs"] = []
            return seg
        
        
        def _hex_rgb(color):
            value = str(color or "#FFFFFF").lstrip("#")
            if len(value) == 8:
                value = value[:6]
            if len(value) != 6:
                raise ValueError(f"invalid text color: {color!r}")
            try:
                return [round(int(value[index:index + 2], 16) / 255.0, 6) for index in (0, 2, 4)]
            except ValueError as exc:
                raise ValueError(f"invalid text color: {color!r}") from exc
        
        
        def _text_style(style, start, end):
            fill_color = style.get("fill_color", "#FFFFFF")
            authored = template("text_style")["styles"][0]
            authored.update({
                "fill": {
                    "alpha": 1.0,
                    "content": {
                        "render_type": "solid",
                        "solid": {"alpha": 1.0, "color": _hex_rgb(fill_color)},
                    },
                },
                "range": [int(start), int(end)],
                "size": float(style.get("font_size", 8.0)),
                "bold": bool(style.get("bold", False)),
                "italic": bool(style.get("italic", False)),
                "underline": bool(style.get("underline", False)),
                "strokes": list(style.get("strokes", [])),
            })
            if style.get("font_path") or style.get("font_id"):
                authored["font"] = {
                    "id": str(style.get("font_id", "")),
                    "path": str(style.get("font_path", "")),
                }
            if style.get("stroke_color") and float(style.get("stroke_width", 0)) > 0:
                width = float(style["stroke_width"])
                authored["strokes"] = [{
                    "alpha": 1.0,
                    "content": {
                        "render_type": "solid",
                        "solid": {"alpha": 1.0, "color": _hex_rgb(style["stroke_color"])},
                    },
                    "width": round(0.00196 * (width ** 1.013), 6),
                }]
            if style.get("shadow_color"):
                authored["shadows"] = [{
                    "alpha": float(style.get("shadow_opacity", 90)) / 100.0,
                    "angle": int(style.get("shadow_angle", -45)),
                    "distance": int(style.get("shadow_width", 5)),
                    "feather": float(style.get("shadow_vague", 45)) / 100.0,
                    "content": {
                        "render_type": "solid",
                        "solid": {"alpha": 1.0, "color": _hex_rgb(style["shadow_color"])},
                    },
                }]
            if isinstance(style.get("effect_style"), dict):
                authored["effect_style"] = deepcopy(style["effect_style"])
            authored["use_letter_color"] = True
            return authored
        
        
        def rich_text_content(text, base_style=None, words=None, style_presets=None):
            """Build duo-video text styles using UTF-16 code-unit ranges."""
            base_style = dict(base_style or {})
            length = len(text.encode("utf-16-le")) // 2
            style_presets = style_presets or {}
            normalized_words = []
            boundaries = {0, length}
            for word in words or []:
                start = max(0, min(length, int(word.get("index", 0))))
                end = max(start, min(length, start + int(word.get("length", 0))))
                if end <= start:
                    continue
                word_style = dict(style_presets.get(str(word.get("style_id")), {}))
                word_style.update(word)
                normalized_words.append((start, end, word_style))
                boundaries.update((start, end))
        
            styles = []
            points = sorted(boundaries)
            for start, end in zip(points, points[1:]):
                style = dict(base_style)
                for word_start, word_end, word in normalized_words:
                    if word_start <= start and end <= word_end:
                        style.update({key: value for key, value in word.items() if key not in {"index", "length"}})
                styles.append(_text_style(style, start, end))
            content = template("text_style")
            content["text"] = text
            content["styles"] = styles
            return content
        
        
        RESOURCE_TRACKS = {
            "sound": "audios",
            "sticker": "stickers",
            "text_template": "text_templates",
            "video_effect": "video_effects",
            "face_effect": "video_effects",
        }
        
        
        def _resource_material(segment, kind, new_id):
            """The contract guarantees exactly one of material / resource_config (resolved package)."""
            config = segment.get("resource_config")
            if config is None:
                raw = deepcopy(segment["material"])
            else:
                raw = deepcopy(config["main_config"])
                resource_id = config.get("resource_id")
                if resource_id is not None:
                    raw.setdefault("resource_id", resource_id)
                if config.get("resources"):
                    raw["_bundle_resources"] = deepcopy(config["resources"])
                cover_img = config.get("cover_img")
                if kind == "sticker" and cover_img:
                    raw["icon_url"] = cover_img
                    raw["preview_cover_url"] = cover_img
            raw["id"] = new_id()
            return raw
        
        
        def build_resource_track(ctx, timeline_track):
            kind = timeline_track["kind"]
            materials_key = RESOURCE_TRACKS[kind]
            track_name = timeline_track.get("name", kind)
            for item in timeline_track["segments"]:
                ts, te = float(item["timeline_start"]), float(item["timeline_end"])
                duration_us = us(te - ts)
                authored_item = item
                package_name = item.get("resource_package")
                if package_name is not None:
                    package = ctx.resource_packages.get(str(package_name))
                    if not isinstance(package, dict):
                        raise ValueError(f"unknown JianYing resource package: {package_name}")
                    authored_item = dict(item)
                    authored_item["resource_config"] = package
                material = _resource_material(authored_item, kind, ctx.new_id)
                ctx.materials[materials_key].append(material)
                config = authored_item.get("resource_config") or {}
                if kind == "text_template":
                    subordinate_resources = deepcopy(material.get("_bundle_resources", []))
                    subordinate_texts = deepcopy(config.get("texts", []))
                    subordinate_effects = deepcopy(config.get("effects", []))
                    if subordinate_resources:
                        for subordinate in subordinate_texts + subordinate_effects:
                            if isinstance(subordinate, dict):
                                subordinate["_bundle_resources"] = deepcopy(subordinate_resources)
                    ctx.materials["texts"].extend(subordinate_texts)
                    ctx.materials["effects"].extend(subordinate_effects)
                seg = base_segment(material["id"], us(ts), duration_us, 1.0, [], ctx.new_id)
                speed_value = float(item.get("speed", 1.0))
                seg["source_timerange"] = timerange(0, round(duration_us * speed_value))
                seg["track_render_index"] = 0
                seg["extra_material_refs"] = []
                seg["speed"] = speed_value
                if abs(speed_value - 1.0) > 1e-9:
                    speed = speed_material(ctx.new_id, speed_value)
                    ctx.materials["speeds"].append(speed)
                    seg["extra_material_refs"].append(speed["id"])
                seg["clip"] = clip_from_segment(item)
                ctx.add_segment(kind, track_name, us(ts), duration_us, seg)
        
        
        def _attachment_material(spec, kind, new_id):
            material = deepcopy(spec)
            config = material.pop("main_config", None)
            resources = material.pop("resources", None)
            resource_id = material.pop("resource_id", None)
            if config is not None:
                if not isinstance(config, dict):
                    raise ValueError(f"video {kind}.main_config must be an object")
                merged = deepcopy(config)
                merged.update(material)
                material = merged
            if resource_id is not None:
                material.setdefault("resource_id", resource_id)
            if resources:
                material["_bundle_resources"] = deepcopy(resources)
            material["id"] = new_id()
            return material
        
        
        def _resolve_attachment_spec(ctx, spec, kind):
            if isinstance(spec, str):
                package = ctx.resource_packages.get(spec)
                if not isinstance(package, dict):
                    raise ValueError(f"unknown JianYing {kind} resource package: {spec}")
                return package
            return spec
        
        
        def apply_video_attachments(ctx, clip, segment):
            """Attach duo-video transition, mask, and LUT authoring data."""
            semantic_kind = "video"
            transition_spec = clip.get("transition")
            if transition_spec is not None:
                transition_spec = _resolve_attachment_spec(ctx, transition_spec, "transition")
                transition = _attachment_material(transition_spec, "transition", ctx.new_id)
                ctx.materials["transitions"].append(transition)
                segment["extra_material_refs"].append(transition["id"])
        
            mask_spec = clip.get("mask")
            if mask_spec is not None:
                mask_spec = _resolve_attachment_spec(ctx, mask_spec, "mask")
                mask = _attachment_material(mask_spec, "mask", ctx.new_id)
                ctx.materials["masks"].append(mask)
                ctx.materials["common_mask"].append(deepcopy(mask))
                segment["extra_material_refs"].append(mask["id"])
                semantic_kind = "mask"
        
            lut_spec = clip.get("lut")
            if lut_spec is not None:
                lut_spec = _resolve_attachment_spec(ctx, lut_spec, "lut")
                lut = _attachment_material(lut_spec, "lut", ctx.new_id)
                lut["value"] = float(lut.pop("strength", 100)) / 100.0
                lut.pop("skin_tone_correction", None)
                ctx.materials["effects"].append(lut)
                segment["extra_material_refs"].append(lut["id"])
                if lut_spec.get("skin_tone_correction") is not None:
                    lumi_hub_path = str(lut.get("lumi_hub_path") or "")
                    if not lumi_hub_path:
                        raise ValueError(
                            "LUT skin_tone_correction requires an offline effect "
                            "main_config with lumi_hub_path"
                        )
                    effect_path = lumi_hub_path.rsplit("/", 1)[0]
                    skin_tone = deepcopy(lut)
                    skin_tone["id"] = ctx.new_id()
                    skin_tone["type"] = "skin_tone_correction"
                    skin_tone["version"] = "v3"
                    skin_tone["value"] = float(lut_spec["skin_tone_correction"]) / 100.0
                    skin_tone["path"] = effect_path
                    skin_tone["lumi_hub_path"] = effect_path
                    ctx.materials["effects"].append(skin_tone)
                    segment["extra_material_refs"].append(skin_tone["id"])
            return semantic_kind
        
        
        def _track_object(ctx, name, track_type, segments, flag):
            return {
                "attribute": 0,
                "flag": flag,
                "id": ctx.new_id(),
                "is_default_name": True,
                "name": name,
                "segments": segments,
                "type": track_type,
            }
        
        
        def build_compound_video(ctx, clip, material, foreground_segment, track_name, ts, te):
            """Build duo-video's nested green-screen/compound draft structure."""
            background_spec = clip["green_background"]
            chroma_spec = _resolve_attachment_spec(ctx, clip["chroma"], "chroma")
        
            duration_us = us(te - ts)
            full_material_duration_us = material["duration"]
            outer_source_timerange = deepcopy(foreground_segment["source_timerange"])
            outer_target_timerange = deepcopy(foreground_segment["target_timerange"])
            background_path = background_spec["source_path"]
        
            background_id = ctx.new_id()
            background = template("video")
            background.update({
                "duration": full_material_duration_us,
                "height": int(background_spec.get("height") or ctx.height),
                "id": background_id,
                "material_name": os.path.basename(background_path),
                "path": background_path,
                "type": background_spec.get("type", "photo"),
                "width": int(background_spec.get("width") or ctx.width),
            })
        
            chroma = _attachment_material(chroma_spec, "chroma", ctx.new_id)
            foreground_segment = deepcopy(foreground_segment)
            foreground_segment["source_timerange"] = timerange(0, full_material_duration_us)
            foreground_segment["target_timerange"] = timerange(0, full_material_duration_us)
            foreground_segment["extra_material_refs"].append(chroma["id"])
        
            background_segment = base_segment(
                background_id, 0, full_material_duration_us, 1.0, [], ctx.new_id
            )
            background_segment["source_timerange"] = timerange(0, full_material_duration_us)
            background_segment["render_index"] = 1
            background_segment["track_render_index"] = 1
            background_segment["clip"] = clip_from_segment(background_spec)
        
            draft = template("draft")
            draft["id"] = ctx.new_id()
            draft["combination_id"] = ctx.new_id()
            nested = draft["draft"]
            nested["id"] = ctx.new_id()
            nested["canvas_config"] = {
                "width": ctx.width,
                "height": ctx.height,
                "ratio": "original",
            }
            nested["duration"] = full_material_duration_us
            nested["fps"] = float(ctx.fps)
            scrub_platform_identity(nested)
            nested["materials"]["videos"] = [material, background]
            nested["materials"]["chromas"] = [chroma]
            referenced = set(foreground_segment["extra_material_refs"])
            nested_speeds = [
                item for item in ctx.materials["speeds"] if item.get("id") in referenced
            ]
            if nested_speeds:
                nested["materials"]["speeds"] = nested_speeds
                ctx.materials["speeds"] = [
                    item for item in ctx.materials["speeds"] if item.get("id") not in referenced
                ]
            nested["tracks"] = [
                _track_object(ctx, "green_background", "video", [background_segment], 0),
                _track_object(ctx, "video", "video", [foreground_segment], 2),
            ]
            ctx.materials["drafts"].append(draft)
        
            combination_material = template("combination_video")
            combination_material.update({
                "duration": full_material_duration_us,
                "height": ctx.height,
                "id": ctx.new_id(),
                "width": ctx.width,
            })
            ctx.materials["videos"].append(combination_material)
        
            combination_segment = template("combination_segment")
            combination_segment.update({
                "id": ctx.new_id(),
                "material_id": combination_material["id"],
                "extra_material_refs": [draft["id"]],
                "source_timerange": outer_source_timerange,
                "target_timerange": outer_target_timerange,
            })
            semantic_kind = apply_video_attachments(ctx, clip, combination_segment)
            authored_track_name = "mask" if semantic_kind == "mask" else track_name
            ctx.add_segment(
                semantic_kind, authored_track_name, us(ts), duration_us, combination_segment
            )
        
        
        def build_video_track(ctx, timeline_track):
            track_name = timeline_track.get("name", "video")
            for clip in timeline_track["clips"]:
                ts, te = float(clip["timeline_start"]), float(clip["timeline_end"])
                ss, se = float(clip["source_start"]), float(clip["source_end"])
                path = clip["source_path"]
                speed_value = float(clip.get("speed", 1.0))
                reverse = bool(clip.get("reverse", False))
                if reverse:
                    # The contract allows omitting reverse_path so export_timeline_to_jianying
                    # can generate it; building directly from such a clip is an authoring error.
                    if "reverse_path" not in clip:
                        raise ValueError("reverse video clips require a local reverse_path")
                    path = clip["reverse_path"]
                src_dur_us, width, height = ctx.probe(path)
                if src_dur_us <= 0:
                    raise ValueError(f"JianYing video source has no probed duration: {path}")
                mat_id = ctx.new_id()
                material = template("video")
                material.update({
                    "duration": int(src_dur_us),
                    "height": height or ctx.height,
                    "id": mat_id,
                    "material_name": os.path.basename(path),
                    "path": path,
                    "width": width or ctx.width,
                })
                audio = clip.get("audio", {})
                keyframes = volume_keyframes(audio.get("volume_keyframes"), ts, ctx.new_id)
                volume = audio.get("base_gain", 1.0) if not keyframes else 1.0
                seg = base_segment(mat_id, us(ts), us(te - ts), volume, keyframes, ctx.new_id)
                source_start = ss
                if reverse:
                    source_start = max(0.0, (src_dur_us / 1_000_000) - se)
                seg["source_timerange"] = timerange(us(source_start), us(se - ss))
                seg["speed"] = speed_value
                seg["extra_material_refs"] = []
                if abs(speed_value - 1.0) > 1e-9:
                    speed = speed_material(ctx.new_id, speed_value)
                    ctx.materials["speeds"].append(speed)
                    seg["extra_material_refs"].append(speed["id"])
                seg["clip"] = clip_from_segment(clip)
                if clip.get("compound") or clip.get("green_background") or clip.get("chroma"):
                    build_compound_video(ctx, clip, material, seg, track_name, ts, te)
                    continue
                ctx.materials["videos"].append(material)
                semantic_kind = apply_video_attachments(ctx, clip, seg)
                authored_track_name = "mask" if semantic_kind == "mask" else track_name
                ctx.add_segment(semantic_kind, authored_track_name, us(ts), us(te - ts), seg)
        
        
        def build_audio_track(ctx, timeline_track):
            role = timeline_track.get("role", timeline_track.get("name", "audio"))
            track_name = timeline_track.get("name", role)
            for segment in timeline_track["segments"]:
                ts, te = float(segment["timeline_start"]), float(segment["timeline_end"])
                path = segment["source_path"]
                mat_dur_us, _width, _height = ctx.probe(path)
                if mat_dur_us <= 0:
                    raise ValueError(f"JianYing audio source has no probed duration: {path}")
                want_us = us(te - ts)
                speed_value = float(segment.get("speed", 1.0))
                place_us = want_us
                required_source_us = int(round(want_us * speed_value))
                if required_source_us > mat_dur_us:
                    if role == "bgm" and timeline_track.get("loop"):
                        place_us = want_us
                    else:
                        place_us = int(mat_dur_us / speed_value)
                    if role == "bgm" and not timeline_track.get("loop"):
                        ctx.note(
                            f"BGM 素材({mat_dur_us/1e6:.1f}s) 短于时间线({(te - ts):.1f}s),"
                            "剪映中未循环铺满(可在剪映里手动复制延长)"
                        )
                mat_id = ctx.new_id()
                material = template("audio")
                material.update({
                    "duration": int(mat_dur_us),
                    "id": mat_id,
                    "path": path,
                })
                ctx.materials["audios"].append(material)
                keyframes = volume_keyframes(segment.get("volume_keyframes"), ts, ctx.new_id)
                volume = segment.get("gain", 1.0) if not keyframes else 1.0
                if role == "bgm" and timeline_track.get("loop") and required_source_us > mat_dur_us:
                    cursor = 0
                    while cursor < want_us:
                        piece = min(int(mat_dur_us / speed_value), want_us - cursor)
                        if piece <= 0:
                            break
                        piece_start_s = ts + (cursor / 1_000_000)
                        piece_end_s = ts + ((cursor + piece) / 1_000_000)
                        piece_kfs = windowed_volume_keyframes(
                            segment.get("volume_keyframes"), piece_start_s, piece_end_s,
                            segment.get("gain", 1.0), ctx.new_id)
                        piece_volume = segment.get("gain", 1.0) if not piece_kfs else 1.0
                        piece_seg = audio_segment_piece(
                            mat_id, us(ts) + cursor, piece, 0, int(round(piece * speed_value)),
                            piece_volume, piece_kfs, ctx.new_id)
                        piece_seg["speed"] = speed_value
                        if abs(speed_value - 1.0) > 1e-9:
                            speed = speed_material(ctx.new_id, speed_value)
                            ctx.materials["speeds"].append(speed)
                            piece_seg["extra_material_refs"].append(speed["id"])
                        ctx.add_segment("audio", track_name, us(ts) + cursor, piece, piece_seg)
                        cursor += piece
                else:
                    audio_seg = audio_segment_piece(
                        mat_id, us(ts), place_us, 0, int(round(place_us * speed_value)),
                        volume, keyframes, ctx.new_id)
                    audio_seg["speed"] = speed_value
                    if abs(speed_value - 1.0) > 1e-9:
                        speed = speed_material(ctx.new_id, speed_value)
                        ctx.materials["speeds"].append(speed)
                        audio_seg["extra_material_refs"].append(speed["id"])
                    ctx.add_segment("audio", track_name, us(ts), place_us, audio_seg)
        
        
        def build_text_track(ctx, timeline_track):
            track_name = timeline_track.get("name", "text")
            track_kind = "subtitle" if track_name == "subtitle" else "text"
            for segment in timeline_track["segments"]:
                ts, te = float(segment["timeline_start"]), float(segment["timeline_end"])
                text = segment["text"]
                mat_id = ctx.new_id()
                authored_style = dict(ctx.style_presets.get(str(segment.get("style_id")), {}))
                authored_style.update(segment.get("style") or {})
                content = rich_text_content(
                    text, authored_style, segment.get("words"), ctx.style_presets
                )
                material = template("text")
                material.update({
                    "id": mat_id,
                    "content": json.dumps(content, ensure_ascii=False),
                    "type": "subtitle" if track_kind == "subtitle" else "text",
                    "alignment": int(authored_style.get("text_align", 1)),
                    "font_size": float(authored_style.get("font_size", 8.0)),
                    "text_color": authored_style.get("fill_color", "#FFFFFF"),
                    "line_spacing": float(authored_style.get("line_spacing", 0.02)),
                    "letter_spacing": float(authored_style.get("letter_spacing", 0.0)),
                    "check_flag": 15,
                })
                bundle_resources = []
                for content_style in content["styles"]:
                    font_path = content_style.get("font", {}).get("path")
                    if font_path and not str(font_path).startswith(("Resources/", "##_draftpath_placeholder_")):
                        bundle_resources.append({
                            "source_path": str(font_path),
                            "resource_kind": "fonts",
                        })
                    effect_path = content_style.get("effect_style", {}).get("path")
                    if effect_path and not str(effect_path).startswith(("Resources/", "##_draftpath_placeholder_")):
                        bundle_resources.append({
                            "source_path": str(effect_path),
                            "resource_kind": "effect",
                        })
                if bundle_resources:
                    material["_bundle_resources"] = bundle_resources
                if authored_style.get("background_color"):
                    material.update({
                        "background_color": authored_style["background_color"],
                        "background_alpha": float(authored_style.get("background_opacity", 100)) / 100.0,
                        "background_style": 1,
                        "background_height": float(authored_style.get("background_height", 14)) / 100.0,
                        "background_width": float(authored_style.get("background_width", 14)) / 100.0,
                        "background_horizontal_offset": float(authored_style.get("background_offset_x", 50)) * 0.02 - 1,
                        "background_vertical_offset": float(authored_style.get("background_offset_y", 50)) * 0.02 - 1,
                        "background_round_radius": float(authored_style.get("background_radius", 6)) / 100.0,
                        "check_flag": 31,
                    })
                if authored_style.get("stroke_color") and float(authored_style.get("stroke_width", 0)) > 0:
                    material["bold_width"] = round(
                        0.00196 * (float(authored_style["stroke_width"]) ** 1.013), 6
                    )
                    material["border_color"] = authored_style["stroke_color"]
                if authored_style.get("shadow_color"):
                    material.update({
                        "has_shadow": True,
                        "shadow_color": authored_style["shadow_color"],
                        "shadow_alpha": float(authored_style.get("shadow_opacity", 90)) / 100.0,
                        "shadow_angle": float(authored_style.get("shadow_angle", -45)),
                        "shadow_distance": float(authored_style.get("shadow_width", 5)),
                        "shadow_smoothing": float(authored_style.get("shadow_vague", 45)) / 100.0,
                    })
                if authored_style or segment.get("words"):
                    material["is_rich_text"] = True
                ctx.materials["texts"].append(material)
                duration_us = us(te - ts)
                seg = base_segment(mat_id, us(ts), duration_us, 1.0, [], ctx.new_id)
                seg["source_timerange"] = timerange(0, duration_us)
                seg["clip"] = clip_from_segment(segment)
                if track_kind == "subtitle" and "position" not in segment:
                    seg["clip"]["transform"]["y"] = -0.72
                ctx.add_segment(track_kind, track_name, us(ts), duration_us, seg)
        
        
        def build_image_track(ctx, timeline_track):
            """Build local image overlays as JianYing photo materials on video tracks."""
            track_name = timeline_track.get("name", "image")
            for segment in timeline_track["segments"]:
                ts, te = float(segment["timeline_start"]), float(segment["timeline_end"])
                duration_us = us(te - ts)
                path = segment["source_path"]
                _still_image_duration, width, height = ctx.probe(path)
                mat_id = ctx.new_id()
                material = template("video")
                material.update({
                    "duration": duration_us,
                    "height": height or ctx.height,
                    "id": mat_id,
                    "material_name": os.path.basename(path),
                    "path": path,
                    "type": "photo",
                    "width": width or ctx.width,
                })
                ctx.materials["videos"].append(material)
                seg = base_segment(mat_id, us(ts), duration_us, 1.0, [], ctx.new_id)
                speed_value = float(segment.get("speed", 1.0))
                seg["source_timerange"] = timerange(0, round(duration_us * speed_value))
                seg["clip"] = clip_from_segment(segment)
                seg["speed"] = speed_value
                if abs(speed_value - 1.0) > 1e-9:
                    speed = speed_material(ctx.new_id, speed_value)
                    ctx.materials["speeds"].append(speed)
                    seg["extra_material_refs"].append(speed["id"])
                semantic_kind = apply_video_attachments(ctx, segment, seg)
                authored_track_name = "mask" if semantic_kind == "mask" else track_name
                if semantic_kind == "video":
                    semantic_kind = "image"
                ctx.add_segment(semantic_kind, authored_track_name, us(ts), duration_us, seg)
        
        
        def build_timeline_track(ctx, timeline_track):
            kind = timeline_track["kind"]
            if kind in ("audio", "text") and not timeline_track["segments"]:
                return
            if kind == "video":
                build_video_track(ctx, timeline_track)
            elif kind == "audio":
                build_audio_track(ctx, timeline_track)
            elif kind == "text":
                build_text_track(ctx, timeline_track)
            elif kind == "image":
                build_image_track(ctx, timeline_track)
            else:  # normalize_timeline admits only RESOURCE_TRACKS kinds past this point
                build_resource_track(ctx, timeline_track)
        
      • model.py 3.4 KB
        """Thin internal model used at the JianYing adapter boundary."""
        
        from dataclasses import dataclass, field
        from collections.abc import Callable
        
        from jianying.schema import us
        from jianying.tracks import TrackAllocator
        
        
        ProbeFn = Callable[[str], tuple[int, int, int]]
        NewIdFn = Callable[[], str]
        
        
        @dataclass
        class DraftBuildContext:
            """Normalized draft build state and overlap-safe track collection."""
        
            width: int
            height: int
            fps: float
            total_us: int
            new_id: NewIdFn
            probe: ProbeFn
            resource_packages: dict = field(default_factory=dict)
            style_presets: dict = field(default_factory=dict)
            materials: dict[str, list] = field(
                default_factory=lambda: {
                    key: []
                    for key in (
                        "audios",
                        "chromas",
                        "common_mask",
                        "drafts",
                        "effects",
                        "masks",
                        "speeds",
                        "stickers",
                        "text_templates",
                        "texts",
                        "transitions",
                        "video_effects",
                        "videos",
                    )
                }
            )
            tracks: list[dict] = field(default_factory=list)
            notes: list[str] = field(default_factory=list)
            track_allocator: TrackAllocator = field(default_factory=TrackAllocator)
        
            @classmethod
            def from_timeline(cls, timeline, new_id, probe):
                """Build from a timeline already validated by jianying.timeline_contract."""
                canvas = timeline["canvas"]
                return cls(
                    width=canvas["width"],
                    height=canvas["height"],
                    fps=float(canvas["fps"]),
                    total_us=us(timeline["duration"]),
                    new_id=new_id,
                    probe=probe,
                    resource_packages=timeline.get("resource_packages", {}),
                    style_presets=timeline.get("style_presets", {}),
                )
        
            def add_segment(self, kind, base_name, start_us, duration_us, segment):
                allocated = self.track_allocator.allocate(kind, base_name, start_us, duration_us)
                track = next(
                    (
                        item for item in self.tracks
                        if item["_semantic_kind"] == allocated.kind
                        and item["name"] == allocated.name
                    ),
                    None,
                )
                if track is None:
                    track = {
                        "attribute": 0,
                        "flag": 0,
                        "id": self.new_id(),
                        "is_default_name": True,
                        "name": allocated.name,
                        "segments": [],
                        "type": allocated.track_type,
                        "_semantic_kind": allocated.kind,
                        "_layout_order": allocated.layout_order,
                    }
                    self.tracks.append(track)
                track["segments"].append(segment)
                return track
        
            def finalize_tracks(self):
                # Python's stable sort preserves the timeline's authored order inside
                # one semantic band (for example narration before BGM).
                self.tracks.sort(key=lambda item: item["_layout_order"])
                type_counts = {}
                for track in self.tracks:
                    count = type_counts.get(track["type"], 0)
                    track["flag"] = 0 if count == 0 else 2
                    type_counts[track["type"]] = count + 1
                    del track["_semantic_kind"]
                    del track["_layout_order"]
                return self.tracks
        
            def note(self, message):
                self.notes.append(message)
        
      • optional.py 1012 B
        """Failure-isolated optional JianYing export invoked after canonical rendering."""
        
        from pathlib import Path
        
        from lib import CONFIG, log
        
        def maybe_export_jianying(work_dir, out_dir, stem):
            """Lazy-import the optional 剪映 exporter and write a draft from timeline.json.
        
            The export is a documented fail-open sidecar: any failure is logged and never
            fails the already-rendered recap."""
            try:
                from export_jianying import export_timeline_to_jianying
                from timeline import load_timeline
                parent = out_dir or CONFIG["jianying_draft_dir"] or str(work_dir)
                draft_dir, notes = export_timeline_to_jianying(
                    load_timeline(Path(work_dir) / "timeline.json"), parent, draft_name=f"recap_{stem}",
                    bundle_media=CONFIG["jianying_bundle_media"])
                for n in notes:
                    log(f"  注意: {n}")
                log(f"剪映草稿已导出: {draft_dir}")
            except Exception as exc:
                log(f"  ⚠️ 剪映导出失败(不影响成片): {exc}")
        
      • schema.py 2.5 KB
        """Schema constants and skeleton factories for JianYing export.
        
        This module is intentionally data-oriented: it owns draft version metadata and
        the full `materials` parallel-array shape.
        """
        
        from jianying.templates import template
        
        # The full 剪映 materials object: ~45 parallel arrays. Only arrays backed by a
        # production builder are populated; retaining the full shape preserves compatibility.
        MATERIAL_KEYS = ["ai_translates", "audio_balances", "audio_effects", "audio_fades", "audio_track_indexes", "audios", "beats", "canvases", "chromas", "color_curves", "digital_humans", "drafts", "effects", "flowers", "green_screens", "handwrites", "hsl", "images", "log_color_wheels", "loudnesses", "manual_deformations", "masks", "common_mask", "material_animations", "material_colors", "multi_language_refs", "placeholders", "plugin_effects", "primary_color_wheels", "realtime_denoises", "shapes", "smart_crops", "smart_relights", "sound_channel_mappings", "speeds", "stickers", "tail_leaders", "text_templates", "texts", "time_marks", "transitions", "video_effects", "video_trackings", "videos", "vocal_beautifys", "vocal_separations"]
        
        
        def us(seconds):
            """Seconds (float) -> integer microseconds. The single seconds->µs boundary."""
            return int(round(float(seconds) * 1_000_000))
        
        
        def full_materials(filled):
            """Return a complete JianYing `materials` object with all known arrays."""
            out = {k: [] for k in MATERIAL_KEYS}
            out.update(filled)
            return out
        
        
        def scrub_platform_identity(project):
            """Remove hardware fingerprints carried by the pinned upstream templates."""
            for platform_key in ("last_modified_platform", "platform"):
                for identity_key in ("device_id", "hard_disk_id", "mac_address"):
                    project[platform_key][identity_key] = ""
            return project
        
        
        def draft_content_skeleton(draft_id, width, height, fps, total_us, materials, tracks):
            """Build the root `draft_content.json` / `draft_info.json` skeleton."""
            content = template("project")
            scrub_platform_identity(content)
            content["canvas_config"] = {"width": width, "height": height, "ratio": "original"}
            content["duration"] = int(total_us)
            content["fps"] = float(fps)
            content["id"] = draft_id
            content["materials"] = full_materials(materials)
            content["tracks"] = tracks
            return content
        
        
        def meta_info(draft_id, total_us):
            """Build the companion `draft_meta_info.json` skeleton."""
            meta = template("meta")
            meta["draft_id"] = draft_id
            meta["draft_timeline_materials_size_"] = 0
            meta["tm_duration"] = int(total_us)
            return meta
        
      • templates.py 1.2 KB
        """Load the pinned duo-video JianYing protocol templates."""
        
        import copy
        import json
        from functools import lru_cache
        from pathlib import Path
        
        
        _TEMPLATE_DIR = Path(__file__).resolve().parent.parent.parent / "references" / "jianying"
        _TEMPLATE_FILES = {
            "project": "empty_jy_project_info.json",
            "video": "empty_jy_material_video.json",
            "audio": "empty_yj_material_audio.json",
            "text": "empty_yj_material_text.json",
            "text_style": "empty_jy_text_styles.json",
            "segment": "empty_jy_segment.json",
            "draft": "empty_jy_draft.json",
            "combination_segment": "empty_jy_combination_segment.json",
            "combination_video": "empty_jy_combination_video_material.json",
            "meta": "empty_draft_meta_info.json",
            "meta_material": "empty_jy_meta_material_value.json",
        }
        
        
        @lru_cache(maxsize=None)
        def _read_template(name):
            try:
                filename = _TEMPLATE_FILES[name]
            except KeyError as exc:
                raise ValueError(f"unknown JianYing template: {name}") from exc
            with open(_TEMPLATE_DIR / filename, encoding="utf-8") as source:
                return json.load(source)
        
        
        def template(name):
            """Return a mutable deep copy of a pinned protocol template."""
            return copy.deepcopy(_read_template(name))
        
      • timeline_contract.py 9.4 KB
        """Timeline validation and migration at the JianYing adapter boundary."""
        
        import copy
        import math
        
        
        CURRENT_SCHEMA_VERSION = 2
        RESOURCE_TRACK_KINDS = {
            "face_effect", "sound", "sticker", "text_template", "video_effect",
        }
        
        def _error(path, expectation):
            raise ValueError(f"invalid timeline {path}: {expectation}")
        
        
        def _is_number(value):
            return isinstance(value, (int, float)) and not isinstance(value, bool) and math.isfinite(value)
        
        
        def _field_path(path, key):
            return f"{path}.{key}" if path else key
        
        
        def _require_number(container, key, path, *, minimum=None):
            field_path = _field_path(path, key)
            if key not in container or not _is_number(container[key]):
                _error(field_path, "must be a finite number")
            value = container[key]
            if minimum is not None and value < minimum:
                _error(field_path, f"must be >= {minimum}")
            return value
        
        
        def _require_string(container, key, path):
            value = container.get(key)
            if not isinstance(value, str) or not value:
                _error(_field_path(path, key), "must be a non-empty string")
            return value
        
        
        def _validate_span(item, path, start_key="timeline_start", end_key="timeline_end"):
            start = _require_number(item, start_key, path, minimum=0)
            end = _require_number(item, end_key, path, minimum=0)
            if end <= start:
                _error(f"{path}.{end_key}", f"must be greater than {start_key}")
        
        
        def _validate_transform(item, path):
            for field in ("scale", "position", "flip"):
                if field in item and not isinstance(item[field], dict):
                    _error(f"{path}.{field}", "must be an object")
        
        
        def _validate_resources(resources, path):
            if not isinstance(resources, list) or any(not isinstance(item, dict) for item in resources):
                _error(path, "must contain source_path objects")
            for index, item in enumerate(resources):
                _require_string(item, "source_path", f"{path}[{index}]")
        
        
        def _validate_video_clip(clip, path):
            if not isinstance(clip, dict):
                _error(path, "must be an object")
            _require_string(clip, "source_path", path)
            _validate_span(clip, path)
            _validate_span(clip, path, "source_start", "source_end")
            if "audio" in clip and not isinstance(clip["audio"], dict):
                _error(f"{path}.audio", "must be an object")
            if "speed" in clip:
                speed = _require_number(clip, "speed", path)
                if speed <= 0:
                    _error(f"{path}.speed", "must be greater than 0")
                source_duration = float(clip["source_end"]) - float(clip["source_start"])
                target_duration = float(clip["timeline_end"]) - float(clip["timeline_start"])
                expected_source_duration = target_duration * float(speed)
                if not math.isclose(source_duration, expected_source_duration, rel_tol=1e-6, abs_tol=1e-4):
                    _error(
                        path,
                        "source duration must equal target duration multiplied by speed "
                        f"({source_duration} != {target_duration} * {speed})",
                    )
            if "reverse" in clip and not isinstance(clip["reverse"], bool):
                _error(f"{path}.reverse", "must be a boolean")
            # A reversed clip may omit reverse_path: export_timeline_to_jianying generates it.
            if "reverse_path" in clip:
                _require_string(clip, "reverse_path", path)
            _validate_transform(clip, path)
            for field in ("transition", "mask", "lut", "chroma"):
                if field not in clip:
                    continue
                spec = clip[field]
                if isinstance(spec, dict):
                    if "resources" in spec:
                        _validate_resources(spec["resources"], f"{path}.{field}.resources")
                elif not isinstance(spec, str):
                    _error(f"{path}.{field}", "must be an object or resource-package name")
            if "compound" in clip and not isinstance(clip["compound"], bool):
                _error(f"{path}.compound", "must be a boolean")
            if clip.get("compound") or "green_background" in clip or "chroma" in clip:
                # Any one of these makes the clip a green-screen compound, which needs both.
                background = clip.get("green_background")
                if not isinstance(background, dict):
                    _error(f"{path}.green_background", "must be a local media object")
                _require_string(background, "source_path", f"{path}.green_background")
                _validate_transform(background, f"{path}.green_background")
                if "chroma" not in clip:
                    _error(f"{path}.chroma", "compound green-screen clips require a chroma object")
        
        
        def _validate_resource_config(config, path):
            if not isinstance(config, dict):
                _error(path, "must be an object")
            if not isinstance(config.get("main_config"), dict):
                _error(f"{path}.main_config", "must be an object")
            _validate_resources(config.get("resources", []), f"{path}.resources")
        
        
        def _validate_segment(segment, path, kind):
            if not isinstance(segment, dict):
                _error(path, "must be an object")
            _validate_span(segment, path)
            _validate_transform(segment, path)
            if "speed" in segment:
                speed = _require_number(segment, "speed", path)
                if speed <= 0:
                    _error(f"{path}.speed", "must be greater than 0")
            if kind in {"audio", "image"}:
                _require_string(segment, "source_path", path)
            elif kind == "text" and not isinstance(segment.get("text"), str):
                _error(f"{path}.text", "must be a string")
            elif kind in RESOURCE_TRACK_KINDS:
                sources = [key for key in ("material", "resource_config", "resource_package") if key in segment]
                if len(sources) != 1:
                    _error(path, "must define exactly one of material, resource_config, or resource_package")
                source = segment[sources[0]]
                if sources[0] == "resource_package":
                    if not isinstance(source, str) or not source:
                        _error(f"{path}.resource_package", "must be a non-empty string")
                elif sources[0] == "material" and not isinstance(source, dict):
                    _error(f"{path}.material", "must be an object")
                elif sources[0] == "resource_config":
                    _validate_resource_config(source, f"{path}.resource_config")
        
        
        def _validate_track(track, path):
            if not isinstance(track, dict):
                _error(path, "must be an object")
            kind = _require_string(track, "kind", path)
            if "name" in track and (not isinstance(track["name"], str) or not track["name"]):
                _error(f"{path}.name", "must be a non-empty string")
        
            if kind == "video":
                clips = track.get("clips")
                if not isinstance(clips, list):
                    _error(f"{path}.clips", "must be an array")
                for index, clip in enumerate(clips):
                    _validate_video_clip(clip, f"{path}.clips[{index}]")
                return
        
            if kind in {"audio", "image", "text"} | RESOURCE_TRACK_KINDS:
                segments = track.get("segments")
                if not isinstance(segments, list):
                    _error(f"{path}.segments", "must be an array")
                if kind == "audio":
                    if "role" in track and not isinstance(track["role"], str):
                        _error(f"{path}.role", "must be a string")
                    if "loop" in track and not isinstance(track["loop"], bool):
                        _error(f"{path}.loop", "must be a boolean")
                for index, segment in enumerate(segments):
                    _validate_segment(segment, f"{path}.segments[{index}]", kind)
                return
        
            _error(f"{path}.kind", f"unsupported track kind {kind!r}")
        
        
        def _validate_v2(timeline):
            canvas = timeline.get("canvas")
            if not isinstance(canvas, dict):
                _error("canvas", "must be an object")
            for dimension in ("width", "height"):
                value = canvas.get(dimension)
                if not isinstance(value, int) or isinstance(value, bool) or value <= 0:
                    _error(f"canvas.{dimension}", "must be a positive integer")
            fps = _require_number(canvas, "fps", "canvas")
            if fps <= 0:
                _error("canvas.fps", "must be greater than 0")
        
            _require_number(timeline, "duration", "", minimum=0)
            resource_packages = timeline.get("resource_packages", {})
            if not isinstance(resource_packages, dict):
                _error("resource_packages", "must be an object")
            for name, config in resource_packages.items():
                _validate_resource_config(config, f"resource_packages.{name}")
            if "style_presets" in timeline and not isinstance(timeline["style_presets"], dict):
                _error("style_presets", "must be an object")
            tracks = timeline.get("tracks")
            if not isinstance(tracks, list):
                _error("tracks", "must be an array")
            for index, track in enumerate(tracks):
                _validate_track(track, f"tracks[{index}]")
        
        
        def normalize_timeline(timeline):
            """Return a validated schema-v2 copy, migrating schema v1 when necessary."""
            if not isinstance(timeline, dict):
                _error("root", "must be an object")
            schema_version = timeline.get("schema_version")
            if not isinstance(schema_version, int) or isinstance(schema_version, bool):
                _error("schema_version", "must be integer 1 or 2")
            if schema_version not in {1, CURRENT_SCHEMA_VERSION}:
                raise ValueError(
                    f"unsupported timeline schema_version {schema_version}; "
                    f"supported versions are 1 and {CURRENT_SCHEMA_VERSION}"
                )
        
            normalized = copy.deepcopy(timeline)
            if schema_version == 1:
                # Schema v2 is an additive extension of v1 (local image tracks). The
                # migration therefore preserves all authored v1 fields and only advances
                # the version before applying the current contract.
                normalized["schema_version"] = CURRENT_SCHEMA_VERSION
            _validate_v2(normalized)
            return normalized
        
      • tracks.py 2.6 KB
        """Semantic track ordering and overlap-safe allocation for JianYing export.
        
        The order mirrors duo-video's authoring layout. It is used only to order track
        objects; JianYing segment ``render_index`` remains the schema default and must
        not be confused with this semantic layout value.
        """
        
        from dataclasses import dataclass
        
        
        @dataclass(frozen=True)
        class TrackBand:
            kind: str
            track_type: str
            layout_order: int
            description: str
        
        
        SEGMENT_RENDER_INDEX = 2
        
        
        TRACK_LAYOUT_BANDS = {
            "sound": TrackBand("sound", "audio", 10_000, "sound effects"),
            "audio": TrackBand("audio", "audio", 20_000, "narration, music, and general audio"),
            "green_screen": TrackBand("green_screen", "video", 30_000, "green-screen background"),
            "video": TrackBand("video", "video", 40_000, "base video"),
            "image": TrackBand("image", "video", 50_000, "image and photo overlays"),
            "mask": TrackBand("mask", "video", 60_000, "masks"),
            "effect": TrackBand("effect", "effect", 70_000, "video effects"),
            "video_effect": TrackBand("video_effect", "effect", 70_000, "video effects"),
            "face_effect": TrackBand("face_effect", "effect", 70_010, "face effects"),
            "sticker": TrackBand("sticker", "sticker", 80_000, "stickers"),
            "subtitle": TrackBand("subtitle", "text", 90_000, "subtitles"),
            "text": TrackBand("text", "text", 100_000, "plain text"),
            "text_template": TrackBand("text_template", "text", 110_000, "text templates"),
        }
        
        
        @dataclass(frozen=True)
        class AllocatedTrack:
            kind: str
            name: str
            track_type: str
            layout_order: int
        
        
        class TrackAllocator:
            """Allocate deterministic suffix tracks when same-name segments overlap.
        
            Intervals are half-open, so adjacent segments reuse a track while true
            overlap creates ``name-1``, ``name-2``, and so on.
            """
        
            def __init__(self):
                self._occupied = {}
        
            @staticmethod
            def _overlaps(start_us, duration_us, existing):
                end_us = int(start_us) + int(duration_us)
                return any(int(start_us) < old_end and end_us > old_start for old_start, old_end in existing)
        
            def allocate(self, kind, base_name, start_us, duration_us):
                band = TRACK_LAYOUT_BANDS[kind]
                suffix = 0
                while True:
                    name = base_name if suffix == 0 else f"{base_name}-{suffix}"
                    key = (kind, name)
                    occupied = self._occupied.setdefault(key, [])
                    if not self._overlaps(start_us, duration_us, occupied):
                        occupied.append((int(start_us), int(start_us) + int(duration_us)))
                        return AllocatedTrack(kind, name, band.track_type, band.layout_order + suffix)
                    suffix += 1
        
      • writer.py 14.9 KB
        """Safe writer and portable media bundler for JianYing draft folders."""
        
        import hashlib
        import json
        import os
        import shutil
        import tempfile
        import time
        import uuid
        import zipfile
        
        from jianying.templates import template
        
        
        DRAFT_PATH_PLACEHOLDER = "##_draftpath_placeholder_0E685133-18CE-45ED-8CB8-2904A212EC80_##"
        
        
        def validate_draft_name(draft_name):
            """Reject draft names that could escape or alias the requested parent dir."""
            if not isinstance(draft_name, str):
                raise TypeError("draft_name must be a string")
            if not draft_name or not draft_name.strip():
                raise ValueError("draft_name must not be empty")
            if os.path.isabs(draft_name):
                raise ValueError("draft_name must be a plain folder name, not an absolute path")
            if "/" in draft_name or "\\" in draft_name:
                raise ValueError("draft_name must not contain path separators")
            if draft_name in {".", ".."}:
                raise ValueError("draft_name must not be '.' or '..'")
        
        
        def _resource_kind(material, materials_key):
            if materials_key == "audios":
                return "audio", "music"
            if material.get("type") == "photo":
                return "image", "photo"
            return "video", "video"
        
        
        RESOURCE_DIRECTORY_BY_MATERIALS_KEY = {
            "chromas": "effect",
            "common_mask": "mask",
            "effects": "effect",
            "masks": "mask",
            "stickers": "sticker",
            "texts": "text",
            "text_templates": "text_template",
            "transitions": "transition",
            "video_effects": "effect",
        }
        
        
        def _unused_name(directory, basename, used):
            """Return a collision-free filename within one resource directory."""
            name = basename
            stem, ext = os.path.splitext(basename)
            suffix = 1
            while name in used or os.path.exists(os.path.join(directory, name)):
                name = f"{stem}_{suffix}{ext}"
                suffix += 1
            used.add(name)
            return name
        
        
        def _md5(path):
            digest = hashlib.md5(usedforsecurity=False)
            with open(path, "rb") as source:
                for chunk in iter(lambda: source.read(1024 * 1024), b""):
                    digest.update(chunk)
            return digest.hexdigest()
        
        
        def _meta_value(material, relative_path, metetype, copied_path, timestamp_ms):
            duration = int(material.get("duration") or 0)
            timestamp_s = timestamp_ms // 1000
            value = template("meta_material")
            value.update({
                "duration": duration,
                "height": int(material.get("height") or 0),
                # duo-video indexes imported local resources independently from the
                # draft-content material IDs.
                "id": str(uuid.uuid4()).upper(),
                "md5": _md5(copied_path),
                "metetype": metetype,
                "type": 0,
                "width": int(material.get("width") or 0),
                "create_time": timestamp_s,
                "extra_info": os.path.basename(relative_path),
                "file_Path": f"./{relative_path}",
                "import_time": timestamp_s,
                "import_time_ms": timestamp_ms,
                "item_source": 1,
                "roughcut_time_range": {"duration": duration, "start": 0},
                "sub_time_range": {"duration": -1, "start": -1},
            })
            return value
        
        
        def _material_sets(content):
            """Yield root and nested compound-draft material dictionaries."""
            materials = content["materials"]
            yield materials
            for draft in materials["drafts"]:
                yield from _material_sets(draft["draft"])
        
        
        def _replace_value(value, old, new):
            if isinstance(value, dict):
                for key, item in value.items():
                    value[key] = _replace_value(item, old, new)
            elif isinstance(value, list):
                for index, item in enumerate(value):
                    value[index] = _replace_value(item, old, new)
            elif isinstance(value, str):
                if value == old:
                    return new
            return value
        
        
        def _replace_material_resource_path(material, materials_key, old, new):
            """Rewrite one declared resource, including the known rich-text JSON field."""
            _replace_value(material, old, new)
            if materials_key != "texts" or not isinstance(material.get("content"), str):
                return
            content = json.loads(material["content"])
            _replace_value(content, old, new)
            material["content"] = json.dumps(content, ensure_ascii=False, separators=(",", ":"))
        
        
        def _safe_resource_target(target_path):
            normalized = os.path.normpath(str(target_path).replace("\\", "/"))
            if normalized in {"", "."} or os.path.isabs(normalized):
                raise ValueError(f"invalid JianYing resource target_path: {target_path}")
            if normalized == ".." or normalized.startswith("../"):
                raise ValueError(f"JianYing resource target_path escapes package: {target_path}")
            return normalized
        
        
        def _extract_zip(source, destination):
            os.makedirs(destination, exist_ok=False)
            destination_real = os.path.realpath(destination)
            with zipfile.ZipFile(source) as archive:
                for member in archive.infolist():
                    member_path = os.path.realpath(os.path.join(destination, member.filename))
                    if os.path.commonpath((destination_real, member_path)) != destination_real:
                        raise ValueError(f"unsafe path in JianYing resource archive: {member.filename}")
                archive.extractall(destination)
        
        
        def _copy_resource(source, resource_kind, resources_root, used, target_path=None):
            resource_dir = os.path.join(resources_root, resource_kind)
            os.makedirs(resource_dir, exist_ok=True)
            is_zip = os.path.isfile(source) and zipfile.is_zipfile(source)
            if target_path is None:
                basename = os.path.basename(source.rstrip(os.sep))
                if is_zip:
                    basename = os.path.splitext(basename)[0]
                relative_target = _unused_name(resource_dir, basename, used[resource_kind])
            else:
                relative_target = _safe_resource_target(target_path)
            copied_path = os.path.join(resource_dir, relative_target)
            copied_real = os.path.realpath(copied_path)
            if os.path.commonpath((os.path.realpath(resource_dir), copied_real)) != os.path.realpath(resource_dir):
                raise ValueError(f"JianYing resource target escapes package: {target_path}")
            os.makedirs(os.path.dirname(copied_path), exist_ok=True)
            if os.path.exists(copied_path):
                raise FileExistsError(f"duplicate JianYing resource target: {relative_target}")
            if is_zip:
                _extract_zip(source, copied_path)
            elif os.path.isdir(source):
                shutil.copytree(source, copied_path)
            else:
                shutil.copy2(source, copied_path)
            relative_path = f"Resources/local/{resource_kind}/{relative_target.replace(os.sep, '/')}"
            return copied_path, relative_path, f"{DRAFT_PATH_PLACEHOLDER}/{relative_path}"
        
        
        def _is_packaged_path(value):
            return str(value).startswith((DRAFT_PATH_PLACEHOLDER, "Resources/", "./Resources/"))
        
        
        def _descriptor(raw, default_kind, *, required):
            """`raw` is a contract-validated resource entry: an object with a non-empty source_path."""
            source = raw["source_path"]
            resource_kind = raw.get("resource_kind", default_kind)
            target_path = raw.get("target_path")
            if not isinstance(resource_kind, str) or resource_kind not in {
                "audio", "effect", "fonts", "image", "lut", "mask", "sticker",
                "text", "text_template", "transition", "video",
            }:
                raise ValueError(f"invalid JianYing resource kind: {resource_kind}")
            return {
                "source_path": source,
                "resource_kind": resource_kind,
                "target_path": target_path,
                "required": required,
            }
        
        
        def _material_resource_descriptors(material, default_kind):
            descriptors = [
                _descriptor(raw, default_kind, required=True)
                for raw in material.get("_bundle_resources", [])
            ]
            path = material.get("path")
            if isinstance(path, str) and path and not _is_packaged_path(path):
                if not any(item["source_path"] == path for item in descriptors):
                    descriptors.append({
                        "source_path": path,
                        "resource_kind": default_kind,
                        "target_path": None,
                        "required": False,
                    })
            return descriptors
        
        
        def bundle_media(content, meta, draft_dir):
            """Copy media into duo-video's Resources/local contract and index it.
        
            Materials that reference the same source file and resource kind share one
            copied file and one meta entry. Missing sources remain untouched and are
            reported to the caller so a non-portable reference is never disguised as a
            successfully bundled one.
            """
            resources_root = os.path.join(draft_dir, "Resources", "local")
            copied = {}
            resource_kinds = {
                "audio", "effect", "fonts", "image", "lut", "mask", "sticker",
                "text", "text_template", "transition", "video",
            }
            used = {kind: set() for kind in resource_kinds}
            meta_values = []
            notes = []
            timestamp_ms = int(time.time() * 1000)
        
            material_sets = list(_material_sets(content))
            for materials in material_sets:
                for materials_key in ("videos", "audios"):
                    for material in materials.get(materials_key, []):
                        src = material.get("path")
                        if not src:
                            continue
                        resource_kind, metetype = _resource_kind(material, materials_key)
                        source_key = (resource_kind, os.path.realpath(src))
                        existing = copied.get(source_key)
                        if existing is not None:
                            material["path"] = existing["draft_path"]
                            continue
                        if not os.path.isfile(src):
                            if not _is_packaged_path(src):
                                notes.append(f"素材缺失,未打包: {src}")
                            continue
        
                        copied_path, relative_path, draft_path = _copy_resource(
                            src, resource_kind, resources_root, used
                        )
                        material["path"] = draft_path
                        copied[source_key] = {"draft_path": draft_path}
                        meta_values.append(
                            _meta_value(material, relative_path, metetype, copied_path, timestamp_ms)
                        )
        
                for materials_key, resource_kind in RESOURCE_DIRECTORY_BY_MATERIALS_KEY.items():
                    for material in materials.get(materials_key, []):
                        descriptors = _material_resource_descriptors(
                            material, resource_kind
                        )
                        seen_descriptors = set()
                        for descriptor in descriptors:
                            src = descriptor["source_path"]
                            if _is_packaged_path(src):
                                continue
                            descriptor_key = (
                                descriptor["resource_kind"],
                                os.path.realpath(src),
                                descriptor["target_path"],
                            )
                            if descriptor_key in seen_descriptors:
                                continue
                            seen_descriptors.add(descriptor_key)
                            if not os.path.exists(src):
                                message = f"声明的剪映资源缺失: {src}"
                                if descriptor["required"]:
                                    raise ValueError(message)
                                notes.append(message)
                                continue
                            kind = descriptor["resource_kind"]
                            source_key = (kind, os.path.realpath(src), descriptor["target_path"])
                            existing = copied.get(source_key)
                            if existing is None:
                                _copied_path, _relative_path, draft_path = _copy_resource(
                                    src,
                                    kind,
                                    resources_root,
                                    used,
                                    target_path=descriptor["target_path"],
                                )
                                copied[source_key] = {"draft_path": draft_path}
                            else:
                                draft_path = existing["draft_path"]
                            _replace_material_resource_path(
                                material, materials_key, src, draft_path
                            )
        
            material_group = next(group for group in meta["draft_materials"] if group["type"] == 0)
            material_group["value"] = meta_values
            meta["draft_timeline_materials_size_"] = sum(
                os.path.getsize(os.path.join(draft_dir, value["file_Path"][2:]))
                for value in meta_values
            )
            return notes
        
        
        def strip_internal_resource_fields(content):
            for materials in _material_sets(content):
                for entries in materials.values():
                    for material in entries:
                        material.pop("_bundle_resources", None)
        
        
        def draft_dir_has_user_content(draft_dir):
            """Return True when writing here could overwrite an existing draft/material."""
            if not os.path.exists(draft_dir):
                return False
            try:
                return any(os.scandir(draft_dir))
            except OSError:
                return True
        
        
        def collision_safe_draft_dir(out_dir, draft_name):
            """Pick a fresh draft folder instead of overwriting an existing non-empty one."""
            validate_draft_name(draft_name)
            base = os.path.join(out_dir, draft_name)
            if not draft_dir_has_user_content(base):
                return base, draft_name
            idx = 2
            while True:
                candidate_name = f"{draft_name}_{idx}"
                candidate = os.path.join(out_dir, candidate_name)
                if not draft_dir_has_user_content(candidate):
                    return candidate, candidate_name
                idx += 1
        
        
        def write_draft(content, meta, notes, out_dir, draft_name, bundle_media_enabled=False):
            """Atomically write the three JianYing draft JSON files and optional bundle."""
            validate_draft_name(draft_name)
            out_dir = os.path.abspath(out_dir)
            os.makedirs(out_dir, exist_ok=True)
            draft_dir, actual_name = collision_safe_draft_dir(out_dir, draft_name)
            notes = list(notes)
            if actual_name != draft_name:
                notes.append(f"草稿目录已存在,改写为 {actual_name} 以避免覆盖")
        
            tmp_parent = tempfile.mkdtemp(prefix=f".{actual_name}.", dir=out_dir)
            tmp_dir = os.path.join(tmp_parent, actual_name)
            try:
                os.makedirs(tmp_dir, exist_ok=False)
                if bundle_media_enabled:
                    notes.extend(bundle_media(content, meta, tmp_dir))
                strip_internal_resource_fields(content)
                timestamp_ms = int(time.time() * 1000)
                meta["draft_name"] = actual_name
                meta["draft_fold_path"] = draft_dir
                meta["tm_draft_create"] = timestamp_ms
                meta["tm_draft_modified"] = timestamp_ms
                content["name"] = actual_name
                for fname in ("draft_content.json", "draft_info.json"):
                    with open(os.path.join(tmp_dir, fname), "w", encoding="utf-8") as f:
                        json.dump(content, f, ensure_ascii=False, indent=2)
                with open(os.path.join(tmp_dir, "draft_meta_info.json"), "w", encoding="utf-8") as f:
                    json.dump(meta, f, ensure_ascii=False, indent=2)
                if os.path.isdir(draft_dir) and not draft_dir_has_user_content(draft_dir):
                    os.rmdir(draft_dir)
                os.replace(tmp_dir, draft_dir)
            except Exception:
                shutil.rmtree(tmp_parent, ignore_errors=True)
                raise
            finally:
                if os.path.exists(tmp_parent):
                    shutil.rmtree(tmp_parent, ignore_errors=True)
            return draft_dir, notes
        
      • __init__.py 49 B
        """Jianying (剪映) draft export subpackage."""
        
    • subtitles
      • core.py 13.6 KB
        """Subtitle text shaping, timing, and measured-canvas geometry."""
        
        import os
        import re
        from decimal import Decimal
        
        from lib import CONFIG
        from assemble_constants import (
            SUBTITLE_STYLE_REF_H,
            SUBTITLE_STYLE_REF_W,
            _SUBTITLE_CLOSING_QUOTES,
            _SUBTITLE_TERMINAL_PUNCTUATION,
        )
        from media import _ratio_to_float
        
        def _seconds_to_srt_time(seconds):
            """Floor times to SRT milliseconds without float remainder artifacts.
        
            Coercing first keeps Fraction/Decimal/str inputs working and clamps a negative
            time to zero instead of emitting a negative-component SRT stamp.
            """
            seconds = max(0.0, float(seconds))
            total_ms = int(Decimal(str(seconds)) * 1000)
            h, remainder = divmod(total_ms, 3_600_000)
            m, remainder = divmod(remainder, 60_000)
            s, ms = divmod(remainder, 1000)
            return f"{h:02d}:{m:02d}:{s:02d},{ms:03d}"
        
        
        def _seconds_to_ass_time(seconds):
            """将秒数转为 ASS 时间格式 H:MM:SS.cc"""
            centiseconds = int(round(float(seconds) * 100))
            h = centiseconds // 360000
            centiseconds %= 360000
            m = centiseconds // 6000
            centiseconds %= 6000
            s = centiseconds // 100
            cs = centiseconds % 100
            return f"{h}:{m:02d}:{s:02d}.{cs:02d}"
        
        
        def _subtitle_style_config(canvas=None):
            """Return the internal default burn-in subtitle style.
        
            When ``canvas`` ({"width","height"}) is given AND the user has not pinned PlayRes via
            SUBTITLE_PLAY_RES_X/Y, the style is scaled to that canvas: PlayRes is set to the frame
            dimensions (so libass never stretches glyphs — the old hardcoded 1280x720 squished
            portrait text), horizontal metrics scale with width, vertical metrics with height, and
            the font is additionally capped so a full ``max_chars`` line fits the usable width. A
            16:9 source (or the 1280x720 default) reproduces the legacy values exactly.
            """
            style = {
                "font_name": CONFIG["subtitle_font_name"],
                "font_file": CONFIG["subtitle_font_file"],
                "font_size": CONFIG["subtitle_font_size"],
                "primary_color": CONFIG["subtitle_primary_color"],
                "outline_color": CONFIG["subtitle_outline_color"],
                "outline": CONFIG["subtitle_outline"],
                "shadow": CONFIG["subtitle_shadow"],
                "alignment": CONFIG["subtitle_alignment"],
                "margin_l": CONFIG["subtitle_margin_l"],
                "margin_r": CONFIG["subtitle_margin_r"],
                "margin_v": CONFIG["subtitle_margin_v"],
                "max_chars": CONFIG["subtitle_max_chars"],
                "play_res_x": CONFIG["subtitle_play_res_x"],
                "play_res_y": CONFIG["subtitle_play_res_y"],
            }
            pinned = "SUBTITLE_PLAY_RES_X" in os.environ or "SUBTITLE_PLAY_RES_Y" in os.environ
            if canvas is None or pinned:
                return style  # legacy / manually-pinned: unchanged
        
            cw, ch = canvas["width"], canvas["height"]
            base_font = float(style["font_size"])
            kx = cw / float(SUBTITLE_STYLE_REF_W)  # horizontal metrics ∝ width
            ky = ch / float(SUBTITLE_STYLE_REF_H)  # vertical metrics ∝ height
            margin_l = round(float(style["margin_l"]) * kx)
            margin_r = round(float(style["margin_r"]) * kx)
            margin_v = round(float(style["margin_v"]) * ky)
            # height-proportional size, then cap so a full line of CJK glyphs (≈1em wide) fits the
            # usable width — this is what keeps portrait text on-screen instead of overflowing.
            usable_w = max(1.0, cw - margin_l - margin_r)
            width_cap = usable_w / int(style["max_chars"])
            font_size = max(1, int(min(base_font * ky, width_cap)))  # floor so a full line never overflows
            font_scale = font_size / base_font
            style.update({
                "font_size": font_size,
                "outline": max(0, round(float(style["outline"]) * font_scale)),
                "shadow": max(0, round(float(style["shadow"]) * font_scale)),
                "margin_l": margin_l,
                "margin_r": margin_r,
                "margin_v": margin_v,
                "play_res_x": cw,
                "play_res_y": ch,
            })
            return style
        
        
        def _measured_subtitle_band(canvas):
            """Explicit (y_top, y_bot) display-frame subtitle band validated against the canvas, or None."""
            y_top = CONFIG["subtitle_y_top"]
            y_bot = CONFIG["subtitle_y_bot"]
            if y_top < 0 and y_bot < 0:
                return None
            sar_text = canvas["sample_aspect_ratio"]
            if abs(_ratio_to_float(sar_text, 0.0) - 1.0) >= 1e-9:
                raise ValueError(
                    f"字幕带坐标仅支持方形像素画布 (SAR 1:1);当前 SAR={sar_text}"
                )
            canvas_h = canvas["height"]
            if not 0 <= y_top < y_bot <= canvas_h:
                raise ValueError(
                    f"字幕带坐标无效: top={y_top}, bot={y_bot}, 画布高度={canvas_h};"
                    "必须满足 0 <= top < bot <= height"
                )
            return y_top, y_bot
        
        
        def _style_for_measured_subtitle_band(style, canvas):
            """Fit the ASS baseline and font into explicit auto-rotated display-frame Y coordinates."""
            style = dict(style)
            safe_area = _measured_subtitle_safe_area(style, canvas)
            if safe_area is None:
                return style
            alignment = style["alignment"]
            if alignment not in {1, 2, 3}:
                raise ValueError(
                    "measured subtitle coordinates require a bottom-aligned ASS style "
                    f"(SUBTITLE_ALIGNMENT 1/2/3); got {alignment}"
                )
            canvas_h = canvas["height"]
            scale_y = float(style["play_res_y"]) / canvas_h
            style["margin_v"] = max(0, round((canvas_h - CONFIG["subtitle_y_bot"]) * scale_y))
            current_font = int(style["font_size"])
            current_outline = float(style["outline"])
            current_shadow = float(style["shadow"])
            available_height = safe_area["height"]
            for candidate in range(current_font, 7, -1):
                scale = candidate / current_font
                outline = max(1 if current_outline > 0 else 0, round(current_outline * scale))
                shadow = max(0, round(current_shadow * scale))
                if candidate * 1.25 + outline * 2 + shadow <= available_height + 1e-6:
                    fitted_font = candidate
                    break
            else:
                # Keep the renderer's minimum readable size; visual QC will block because it cannot fit.
                fitted_font = min(current_font, 8)
            if fitted_font < current_font:
                scale = fitted_font / current_font
                style["font_size"] = fitted_font
                style["outline"] = max(
                    1 if current_outline > 0 else 0, round(current_outline * scale)
                )
                style["shadow"] = max(0, round(current_shadow * scale))
            return style
        
        
        def _measured_subtitle_safe_area(style, canvas):
            """Return the padded measured band in ASS PlayRes coordinates, or None."""
            band = _measured_subtitle_band(canvas)
            if band is None:
                return None
            y_top, y_bot = band
            canvas_h = canvas["height"]
            safe_top = max(0, y_top - CONFIG["subtitle_mask_padding"])
            # The ASS style remains bottom-anchored at the measured y_bot. Bottom mask padding hides
            # source glyph edges but is not usable subtitle layout space; only top padding can extend
            # the line box without moving its baseline below the measured band.
            play_x = int(style["play_res_x"])
            scale_y = int(style["play_res_y"]) / canvas_h
            margin_l = int(style["margin_l"])
            margin_r = int(style["margin_r"])
            return {
                "x": margin_l,
                "y": round(safe_top * scale_y),
                "width": max(1, play_x - margin_l - margin_r),
                "height": max(1, round((y_bot - safe_top) * scale_y)),
                "bottom_margin": max(0, round((canvas_h - y_bot) * scale_y)),
            }
        
        
        def _subtitle_display_text(text):
            """Return display-only subtitle text with trailing sentence punctuation removed.
        
            Narration/TTS source text stays untouched; this is applied only to SRT/ASS cue text.
            Closing quotes/brackets are preserved, so 「原声台词。」 renders as 「原声台词」.
            """
            text = text.strip()
            suffix = ""
            while text and text[-1] in _SUBTITLE_CLOSING_QUOTES:
                suffix = text[-1] + suffix
                text = text[:-1].rstrip()
            text = text.rstrip(_SUBTITLE_TERMINAL_PUNCTUATION).rstrip()
            return (text + suffix).strip()
        
        
        def _subtitle_chunk_weight(text):
            """Weight raw subtitle chunks for timing, independent of display punctuation cleanup."""
            return max(1, len(re.sub(r"\s+", "", text)))
        
        
        def _subtitle_entry_chunks(raw_chunks):
            """Pair raw chunks used for timing with their final display text.
        
            Timing remains based on the raw split topology. Terminal punctuation is stripped only
            on the emitted text, while quote-only suffix chunks are folded into the previous cue
            so a closing bracket never renders alone.
            """
            out = []
            for chunk in raw_chunks:
                display = _subtitle_display_text(chunk)
                if not display:
                    continue
                if all(ch in _SUBTITLE_CLOSING_QUOTES for ch in display):
                    if out:
                        out[-1]["text"] += display
                    continue
                out.append({"raw": chunk, "text": display})
            return out
        
        
        def _normalize_subtitle_text(text):
            """Normalize Chinese em-dashes in burned subtitle text: a run of one-or-more "—" (incl. "——")
            collapses to a single ",". Then collapse any resulting double commas (",,"→",") so the dash
            swap never leaves a doubled comma."""
            return re.sub(r",{2,}", ",", re.sub(r"—+", ",", text))
        
        
        def _split_subtitle_chunks(text, max_chars):
            """Split one narration block (often several sentences) into short display chunks.
        
            A block is synthesized as one continuous TTS utterance for fluent prosody, but showing the
            whole paragraph as a single subtitle would force a tall multi-line band and lag the picture.
            So we cut the block at punctuation into clauses, then greedily pack adjacent clauses into
            chunks of at most `max_chars` — each chunk renders as ONE readable line synced to its slice of
            the block's audio. Punctuation stays attached here for lossless splitting; the display layer
            strips terminal sentence marks per subtitle-cue style."""
            text = text.strip()
            if not text:
                return []
            breakers = ",。!?、;:…—,.!?;:"
            clauses, buf = [], ""
            for ch in text:
                buf += ch
                if ch in breakers:
                    clauses.append(buf)
                    buf = ""
            if buf.strip():
                clauses.append(buf)
            # Any single clause longer than max_chars is hard-wrapped so no chunk ever exceeds one line.
            # Balance those pieces instead of slicing exactly at max_chars: a 21-character clause must
            # not become a readable 20-character cue followed by a 1-character flash.
            sized = []
            for clause in clauses:
                if len(clause) <= max_chars:
                    sized.append(clause)
                else:
                    piece_count = (len(clause) + max_chars - 1) // max_chars
                    base, extra = divmod(len(clause), piece_count)
                    cursor = 0
                    for piece_index in range(piece_count):
                        width = base + (1 if piece_index < extra else 0)
                        sized.append(clause[cursor:cursor + width])
                        cursor += width
            chunks, cur = [], ""
            for clause in sized:
                sentence_closed = cur.rstrip().endswith(tuple(_SUBTITLE_TERMINAL_PUNCTUATION))
                if cur and (sentence_closed or len(cur) + len(clause) > max_chars):
                    chunks.append(cur)
                    cur = clause
                else:
                    cur += clause
            if cur.strip():
                chunks.append(cur)
            return [c.strip() for c in chunks if c.strip()]
        
        
        def _subtitle_entries(narration):
            """Collect subtitle entries from final TTS segment placement.
        
            Each placed segment is split into short one-line chunks and its played window
            [actual_place_start, actual_place_end] is distributed across them in proportion to character
            count — karaoke-style timing that keeps each line on screen only while it is roughly being
            spoken, instead of holding a whole paragraph for the segment's full duration. Segments that
            were not placed have a zero-width window and therefore produce no cue."""
            max_chars = CONFIG["subtitle_max_chars"]
            entries = []
            for seg in narration:
                text = seg["spoken_text"]
                start, end = float(seg["actual_place_start"]), float(seg["actual_place_end"])
                entries.extend(_distribute_chunks(_split_subtitle_chunks(text, max_chars), start, end))
            return entries
        
        
        def _distribute_chunks(chunks, start, end):
            """Distribute [start,end] across raw chunks while emitting display-clean text.
        
            Terminal subtitle punctuation is visual-only: it is stripped from final cue text,
            but the raw split chunks remain the timing topology. A slice too short to show on its
            own is folded into the previous line of the same block, so no chunk is ever dropped.
            """
            chunks = _subtitle_entry_chunks(chunks)
            if not chunks or end - start < 0.1:
                return []
            if len(chunks) == 1:
                return [{"start": start, "end": end, "text": chunks[0]["text"]}]
            total_chars = sum(_subtitle_chunk_weight(c["raw"]) for c in chunks)
            span = end - start
            out, cursor = [], start
            for i, chunk in enumerate(chunks):
                weight = _subtitle_chunk_weight(chunk["raw"])
                chunk_end = end if i == len(chunks) - 1 else cursor + span * (weight / total_chars)
                if out and chunk_end - cursor < 0.05:
                    out[-1]["text"] += chunk["text"]
                    out[-1]["end"] = chunk_end
                else:
                    out.append({"start": cursor, "end": chunk_end, "text": chunk["text"]})
                cursor = chunk_end
            return out
        
        
        def _bracketed_original_chunks(text, start, end, max_chars):
            """Split original dialogue into timed chunks wrapped in 「」 for visual distinction."""
            raw = text.strip()
            if raw.startswith("「") and raw.endswith("」"):
                raw = raw[1:-1].strip()
            chunks = _split_subtitle_chunks(raw, max_chars)
            if chunks:
                chunks[0] = "「" + chunks[0]
                chunks[-1] = chunks[-1] + "」"
            return _distribute_chunks(chunks, start, end)
        
      • render.py 3.5 KB
        """SRT/ASS serialization for narration and original-dialogue subtitles."""
        
        from source_subtitles import _combined_subtitle_entries
        from subtitles.core import (
            _normalize_subtitle_text,
            _seconds_to_ass_time,
            _seconds_to_srt_time,
            _style_for_measured_subtitle_band,
            _subtitle_style_config,
        )
        
        def _generate_srt(narration, work_dir, video_duration):
            """将解说脚本转为 SRT 字幕文件,使用实际音频放置时间;原声留白处补烧原声字幕。"""
            srt_lines = []
            # entries are already split into short one-line chunks, so no wrapping here.
            for idx, entry in enumerate(_combined_subtitle_entries(narration, work_dir, video_duration), start=1):
                srt_lines.append(str(idx))
                srt_lines.append(f"{_seconds_to_srt_time(entry['start'])} --> {_seconds_to_srt_time(entry['end'])}")
                srt_lines.append(entry["text"] if entry.get("_bound_track") else _normalize_subtitle_text(entry["text"]))
                srt_lines.append("")
            srt_path = work_dir / "subtitles.srt"
            srt_path.write_text("\n".join(srt_lines), encoding="utf-8")
            return srt_path
        
        
        def _escape_ass_text(text):
            """Escape user text for an ASS dialogue Text field."""
            return (
                text
                .replace("\\", "\\\\")
                .replace("{", "\\{")
                .replace("}", "\\}")
                .replace("\r\n", "\n")
                .replace("\r", "\n")
                .replace("\n", "\\N")
            )
        
        
        def _generate_ass(narration, work_dir, video_duration, canvas):
            """Generate an ASS subtitle file for readable hard-sub rendering, including the original
            dialogue during the original-audio gaps. canvas ({"width","height"}) scales the style to the
            real frame so portrait/竖屏 subtitles are not stretched."""
            style = _style_for_measured_subtitle_band(_subtitle_style_config(canvas), canvas)
            ass_lines = [
                "[Script Info]",
                "ScriptType: v4.00+",
                "WrapStyle: 0",
                "ScaledBorderAndShadow: yes",
                f"PlayResX: {int(style['play_res_x'])}",
                f"PlayResY: {int(style['play_res_y'])}",
                "",
                "[V4+ Styles]",
                (
                    "Format: Name, Fontname, Fontsize, PrimaryColour, SecondaryColour, "
                    "OutlineColour, BackColour, Bold, Italic, Underline, StrikeOut, "
                    "ScaleX, ScaleY, Spacing, Angle, BorderStyle, Outline, Shadow, "
                    "Alignment, MarginL, MarginR, MarginV, Encoding"
                ),
                (
                    "Style: Default,"
                    f"{style['font_name']},{style['font_size']},{style['primary_color']},&H000000FF,"
                    f"{style['outline_color']},&H64000000,0,0,0,0,100,100,0,0,1,"
                    f"{style['outline']},{style['shadow']},{style['alignment']},"
                    f"{style['margin_l']},{style['margin_r']},{style['margin_v']},1"
                ),
                "",
                "[Events]",
                "Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text",
            ]
            # entries are already split into short one-line chunks, so no wrapping here.
            for entry in _combined_subtitle_entries(narration, work_dir, video_duration):
                text = _escape_ass_text(entry["text"] if entry.get("_bound_track") else _normalize_subtitle_text(entry["text"]))
                ass_lines.append(
                    "Dialogue: 0,"
                    f"{entry.get('ass_start', _seconds_to_ass_time(entry['start']))},"
                    f"{entry.get('ass_end', _seconds_to_ass_time(entry['end']))},"
                    f"Default,,0,0,0,,{text}"
                )
        
            ass_path = work_dir / "subtitles.ass"
            ass_path.write_text("\n".join(ass_lines) + "\n", encoding="utf-8")
            return ass_path
        
      • track.py 12.6 KB
        """Strict loader for versioned, output-clock subtitle tracks.
        
        This module intentionally uses only the Python standard library so it can be
        loaded by the independently distributed ``video-assemble`` skill.
        """
        
        from __future__ import annotations
        
        import json
        import math
        from collections.abc import Mapping
        from decimal import Decimal
        from fractions import Fraction
        from pathlib import Path
        
        
        SCHEMA_VERSION = 1
        
        # Digest keys older tracks declared; they are ignored, never a reason to fail.
        _LEGACY_BINDING_KEYS = frozenset({"sha256", "edit_sha256"})
        _EVIDENCE_KINDS = frozenset(
            {
                "human_verified",
                "word_timestamps",
                "asr_boundary_calibrated",
                "legacy_estimate",
            }
        )
        _CALIBRATION_KINDS = frozenset({"none", "asr_energy", "human_boundary"})
        _WORD_ALIGNMENT_KINDS = frozenset({"none", "asr_words", "human_words"})
        
        
        class SubtitleTrackError(ValueError):
            """The subtitle track is malformed, stale, or incompatible with its caller."""
        
        
        def _fail(path, message):
            raise SubtitleTrackError(f"{path}: {message}")
        
        
        def _mapping(value, path):
            if not isinstance(value, Mapping):
                _fail(path, "must be an object")
            return value
        
        
        def _strict_fields(value, *, required, optional=(), path):
            value = _mapping(value, path)
            keys = set(value)
            missing = set(required) - keys
            unknown = keys - set(required) - set(optional)
            if missing:
                _fail(path, f"missing field(s): {', '.join(sorted(missing))}")
            if unknown:
                _fail(path, f"unknown field(s): {', '.join(sorted(unknown))}")
            return value
        
        
        def _integer(value, path, *, minimum=None):
            if isinstance(value, bool) or not isinstance(value, int):
                _fail(path, "must be an integer")
            if minimum is not None and value < minimum:
                _fail(path, f"must be >= {minimum}")
            return value
        
        
        def _nonempty_string(value, path):
            if not isinstance(value, str) or not value.strip():
                _fail(path, "must be a nonempty string")
            return value
        
        
        def _local_path(value, path):
            return str(Path(_nonempty_string(value, path)).resolve())
        
        
        def _seconds_fraction(value, path):
            if isinstance(value, bool):
                _fail(path, "duration must be finite and numeric")
            if isinstance(value, Fraction):
                result = value
            elif isinstance(value, Decimal):
                if not value.is_finite():
                    _fail(path, "duration must be finite")
                result = Fraction(value)
            elif isinstance(value, int):
                result = Fraction(value)
            elif isinstance(value, float):
                if not math.isfinite(value):
                    _fail(path, "duration must be finite")
                result = Fraction(str(value))
            else:
                _fail(path, "duration must be an int, float, Decimal, or Fraction")
            if result < 0:
                _fail(path, "duration must be nonnegative")
            return result
        
        
        def _load_document(path_or_mapping):
            if isinstance(path_or_mapping, Mapping):
                return path_or_mapping
            try:
                path = Path(path_or_mapping)
            except TypeError as exc:
                raise SubtitleTrackError("track must be a mapping or filesystem path") from exc
            try:
                value = json.loads(path.read_text(encoding="utf-8"))
            except (OSError, UnicodeError, json.JSONDecodeError) as exc:
                raise SubtitleTrackError(f"cannot read subtitle track {path}: {exc}") from exc
            return _mapping(value, "track")
        
        
        def _validate_picture_binding(binding, expected):
            binding = _strict_fields(
                binding, required={"path"}, optional={"edit_plan"} | _LEGACY_BINDING_KEYS,
                path="bindings.picture",
            )
            expected = _mapping(expected, "expected_picture_identity")
            picture_path = _local_path(binding["path"], "bindings.picture.path")
            expected_path = _local_path(expected.get("path"), "expected_picture_identity.path")
            if picture_path != expected_path:
                _fail("bindings.picture.path", "picture path does not match current picture")
        
            edit_plan = binding.get("edit_plan")
            if edit_plan is not None:
                edit_plan = _local_path(edit_plan, "bindings.picture.edit_plan")
                if "edit_plan" not in expected:
                    _fail("expected_picture_identity.edit_plan", "edit plan is required by track")
                expected_edit = _local_path(expected["edit_plan"], "expected_picture_identity.edit_plan")
                if edit_plan != expected_edit:
                    _fail("bindings.picture.edit_plan", "edit plan does not match current edit")
            return {"path": picture_path, **({"edit_plan": edit_plan} if edit_plan else {})}
        
        
        _AUDIO_BINDING_FACTS = ("selected_stream", "sample_rate", "packet_count")
        
        
        def _validate_audio_binding(binding, expected):
            binding = _strict_fields(
                binding, required=set(_AUDIO_BINDING_FACTS), optional=_LEGACY_BINDING_KEYS,
                path="bindings.audio",
            )
            expected = _mapping(expected, "expected_audio_identity")
            result = {}
            for key in _AUDIO_BINDING_FACTS:
                declared = _integer(binding[key], f"bindings.audio.{key}", minimum=0)
                actual = _integer(expected.get(key), f"expected_audio_identity.{key}", minimum=0)
                if declared != actual:
                    _fail(f"bindings.audio.{key}", f"{key} does not match adopted audio")
                result[key] = declared
            return result
        
        
        def _validate_evidence(value, cue_path):
            path = f"{cue_path}.timing_evidence"
            value = _strict_fields(
                value,
                required={"kind", "evidence_refs", "calibration", "word_alignment"},
                path=path,
            )
            kind = _nonempty_string(value["kind"], f"{path}.kind")
            calibration = _nonempty_string(value["calibration"], f"{path}.calibration")
            word_alignment = _nonempty_string(
                value["word_alignment"], f"{path}.word_alignment"
            )
            if kind not in _EVIDENCE_KINDS:
                _fail(path, f"timing_evidence kind is not supported: {kind!r}")
            if calibration not in _CALIBRATION_KINDS:
                _fail(path, f"timing_evidence calibration is not supported: {calibration!r}")
            if word_alignment not in _WORD_ALIGNMENT_KINDS:
                _fail(path, f"timing_evidence word_alignment is not supported: {word_alignment!r}")
            refs = value["evidence_refs"]
            if not isinstance(refs, list):
                _fail(path, "timing_evidence evidence_refs must be a list")
            refs = [_nonempty_string(ref, f"{path}.evidence_refs[{index}]") for index, ref in enumerate(refs)]
        
            valid_shape = {
                "legacy_estimate": calibration == "none" and word_alignment == "none",
                "asr_boundary_calibrated": calibration == "asr_energy" and word_alignment == "none",
                "word_timestamps": calibration in {"none", "asr_energy"} and word_alignment == "asr_words",
                "human_verified": calibration == "human_boundary" and word_alignment in {"none", "human_words"},
            }[kind]
            if not valid_shape:
                _fail(path, f"timing_evidence fields are inconsistent with kind {kind!r}")
            if kind != "legacy_estimate" and not refs:
                _fail(path, f"timing_evidence kind {kind!r} requires evidence_refs")
            return {
                "kind": kind,
                "evidence_refs": refs,
                "calibration": calibration,
                "word_alignment": word_alignment,
            }
        
        
        def load_subtitle_track(
            path_or_mapping,
            *,
            expected_picture_identity,
            expected_audio_identity,
            expected_duration_seconds,
            reject_legacy_estimate=False,
        ):
            """Validate schema v1 and return metadata plus second-based render entries.
        
            Identity arguments are current facts (paths, stream index, sample rate,
            packet count) supplied independently by the caller; declarations inside the
            track are compared against them, never accepted on their own. Cue boundaries
            and text are validated, not split or corrected.
            """
        
            if not isinstance(reject_legacy_estimate, bool):
                _fail("reject_legacy_estimate", "must be a boolean")
            track = _strict_fields(
                _load_document(path_or_mapping),
                required={"schema_version", "clock", "overlap_policy", "bindings", "cues"},
                path="track",
            )
            if type(track["schema_version"]) is not int or track["schema_version"] != SCHEMA_VERSION:
                _fail(
                    "schema_version",
                    f"only schema_version {SCHEMA_VERSION} is supported; got {track['schema_version']!r}",
                )
            if track["overlap_policy"] != "forbid":
                _fail("overlap_policy", "schema v1 requires fail-closed value 'forbid'")
        
            clock = _strict_fields(
                track["clock"],
                required={"kind", "timebase", "duration_ticks"},
                path="clock",
            )
            if clock["kind"] != "output":
                _fail("clock.kind", "schema v1 supports output clock only")
            timebase = _strict_fields(
                clock["timebase"], required={"numerator", "denominator"}, path="clock.timebase"
            )
            numerator = _integer(timebase["numerator"], "clock.timebase.numerator", minimum=1)
            denominator = _integer(timebase["denominator"], "clock.timebase.denominator", minimum=1)
            duration_ticks = _integer(clock["duration_ticks"], "clock.duration_ticks", minimum=0)
            tick_seconds = Fraction(numerator, denominator)
            duration = duration_ticks * tick_seconds
            expected_duration = _seconds_fraction(expected_duration_seconds, "expected_duration_seconds")
            if duration != expected_duration:
                _fail(
                    "clock.duration_ticks",
                    f"duration {duration} does not match independently supplied duration {expected_duration}",
                )
            try:
                duration_seconds = float(duration)
            except OverflowError:
                _fail("clock.duration_ticks", "duration must be representable as finite renderer seconds")
            if not math.isfinite(duration_seconds):
                _fail("clock.duration_ticks", "duration must be representable as finite renderer seconds")
        
            bindings = _strict_fields(
                track["bindings"], required={"picture", "audio"}, path="bindings"
            )
            picture = _validate_picture_binding(bindings["picture"], expected_picture_identity)
            audio = _validate_audio_binding(bindings["audio"], expected_audio_identity)
        
            cues = track["cues"]
            if not isinstance(cues, list):
                _fail("cues", "must be a list")
            entries = []
            kinds = set()
            previous_end = 0
            for index, raw_cue in enumerate(cues):
                cue_path = f"cues[{index}]"
                cue = _strict_fields(
                    raw_cue,
                    required={"start_tick", "end_tick", "text", "attribution", "timing_evidence"},
                    path=cue_path,
                )
                start_tick = _integer(cue["start_tick"], f"{cue_path}.start_tick", minimum=0)
                end_tick = _integer(cue["end_tick"], f"{cue_path}.end_tick", minimum=0)
                if end_tick <= start_tick:
                    _fail(cue_path, "half-open cue requires end_tick > start_tick")
                if end_tick > duration_ticks:
                    _fail(cue_path, "cue end_tick exceeds output duration")
                if start_tick < previous_end:
                    _fail(cue_path, "cue is out of order or overlaps the previous cue")
                previous_end = end_tick
                text = _nonempty_string(cue["text"], f"{cue_path}.text")
        
                attribution = _strict_fields(
                    cue["attribution"], required={"kind", "ref"}, path=f"{cue_path}.attribution"
                )
                source = _nonempty_string(attribution["kind"], f"{cue_path}.attribution.kind")
                if source not in {"source", "narration"}:
                    _fail(f"{cue_path}.attribution.kind", "must be 'source' or 'narration'")
                source_ref = _nonempty_string(attribution["ref"], f"{cue_path}.attribution.ref")
                evidence = _validate_evidence(cue["timing_evidence"], cue_path)
                if reject_legacy_estimate and evidence["kind"] == "legacy_estimate":
                    _fail(cue_path, "legacy_estimate is rejected by strict caller policy")
                start_seconds = float(start_tick * tick_seconds)
                end_seconds = float(end_tick * tick_seconds)
                if not (math.isfinite(start_seconds) and math.isfinite(end_seconds)):
                    _fail(cue_path, "cue float projection must be finite")
                if end_seconds <= start_seconds:
                    _fail(cue_path, "distinct cue ticks collapse in renderer float projection")
                kinds.add(evidence["kind"])
                entries.append(
                    {
                        "start": start_seconds,
                        "end": end_seconds,
                        "text": text,
                        "source": source,
                        "source_ref": source_ref,
                        "timing_evidence": evidence,
                    }
                )
        
            metadata = {
                "schema_version": SCHEMA_VERSION,
                "clock": {
                    "kind": "output",
                    "timebase": {"numerator": numerator, "denominator": denominator},
                    "duration_ticks": duration_ticks,
                    "duration_seconds": duration_seconds,
                },
                "overlap_policy": "forbid",
                "bindings": {"picture": picture, "audio": audio},
                "timing_evidence_kinds": sorted(kinds),
            }
            return {"metadata": metadata, "entries": entries}
        
      • track_binding.py 10.3 KB
        """Bind an explicit output-clock subtitle track to media before any consumer uses it.
        
        The first render integration deliberately supports adopted AAC only. A future
        newly mixed narration track needs its own final-mix identity, not this input's
        soundtrack. Bound cue timing is a declared decision, not proof of speech onset.
        """
        
        import bisect
        import json
        from fractions import Fraction
        from pathlib import Path
        
        from artifacts import file_identity
        from adoption.frozen_audio import probe_audio_packets
        from lib import run_cmd
        from subtitles.track import load_subtitle_track
        
        TRACK = 'subtitle_track.json'
        VALIDATION = 'subtitle_track_validation.json'
        VALIDATION_SCHEMA = 1
        PROJECTOR_VERSION = 1
        
        
        def current_bindings(video, selected_audio_stream=0, *, edit_plan_path=None):
            """Compute the binding facts independently of the supplied subtitle track.
        
            The picture binding is the resolved media path (plus the edit plan path when one
            exists); the audio binding is the selected stream ordinal with its sample rate and
            packet count.
            """
            packets = probe_audio_packets(video, selected_audio_stream)
            picture = {'path': str(Path(video).resolve())}
            if edit_plan_path is not None:
                picture['edit_plan'] = str(Path(edit_plan_path).resolve())
            return {'picture': picture,
                    'audio': {'selected_stream': selected_audio_stream,
                              'sample_rate': packets['sample_rate'],
                              'packet_count': packets['packet_count']}}
        
        
        def _picture_clock(video):
            result = run_cmd([
                'ffprobe', '-v', 'error', '-select_streams', 'v:0', '-show_streams', '-show_frames',
                '-show_entries', 'stream=time_base,start_pts,duration_ts:frame=pts', '-of', 'json', str(video),
            ])
            if result.returncode:
                raise ValueError(f'Cannot verify subtitle picture clock: {result.stderr}')
            data = json.loads(result.stdout)
            stream = data['streams'][0]
            timebase = Fraction(stream['time_base'])
            if int(stream['start_pts']) != 0:
                raise ValueError('Bound subtitle render requires an output-clock video starting at zero')
            duration = int(stream['duration_ts']) * timebase
            pts = [int(frame['pts']) * timebase for frame in data['frames']]
            if not pts or pts[0] != 0 or any(b <= a for a, b in zip(pts, pts[1:])) or pts[-1] >= duration:
                raise ValueError('Cannot verify picture frame clock: expected strictly increasing PTS within duration')
            return pts, duration
        
        
        def _project_boundary(target, frame_pts, duration):
            """Choose an ASS centisecond that changes on the first frame at/after a cue.
        
            ASS's 10ms clock is coarser than most frame clocks and can otherwise round a cue
            forward by one frame. Do not silently degrade if two relevant frames cannot be
            separated in that clock.
            libass receives presentation time in integer milliseconds from ffmpeg.
            """
            index = bisect.bisect_left(frame_pts, target)
            chosen = frame_pts[index] if index < len(frame_pts) else duration
            centiseconds = chosen * 100 // 1
            threshold_ms = int(centiseconds) * 10
            previous_ms = int(frame_pts[index - 1] * 1000) if index else -1
            if threshold_ms <= previous_ms:
                raise ValueError('ASS 10ms clock cannot represent this cue boundary; use a frame-capable renderer')
            ass_time = f'{centiseconds // 360000}:{centiseconds // 6000 % 60:02d}:{centiseconds // 100 % 60:02d}.{centiseconds % 100:02d}'
            return {'requested_time': str(target), 'frame_index': index, 'pts': str(chosen),
                    'delta': str(chosen - target), 'ass_time': ass_time}
        
        
        def prepare_subtitle_track(input_video, work_dir, video_duration, *, audio_mode,
                                   selected_audio_stream=0, edit_plan_path=None,
                                   reject_legacy_estimate=False):
            """Validate and project an explicit track before SRT/ASS/timeline/QC consumers.
        
            No track preserves legacy behavior. A present but invalid track fails instead
            of falling back to proportional timing. The persisted record is not approval.
            """
            work = Path(work_dir)
            path = work / TRACK
            validation = work / VALIDATION
            validation.unlink(missing_ok=True)
            if not path.exists():
                return None
            if audio_mode != 'adopted-packet-copy':
                raise ValueError('Explicit subtitle_track currently requires adopted-packet-copy; new mix is not bound')
            document = json.loads(path.read_bytes())
            bindings = current_bindings(input_video, selected_audio_stream, edit_plan_path=edit_plan_path)
            frame_pts, duration = _picture_clock(input_video)
            if abs(float(duration) - float(video_duration)) > 0.05:
                raise ValueError('Picture duration differs from assembly duration')
            loaded = load_subtitle_track(
                document, expected_picture_identity=bindings['picture'], expected_audio_identity=bindings['audio'],
                expected_duration_seconds=duration, reject_legacy_estimate=reject_legacy_estimate,
            )
            clock = document['clock']['timebase']
            tick = Fraction(clock['numerator'], clock['denominator'])
            for entry, cue in zip(loaded['entries'], document['cues']):
                start = _project_boundary(cue['start_tick'] * tick, frame_pts, duration)
                end = _project_boundary(cue['end_tick'] * tick, frame_pts, duration)
                if end['frame_index'] <= start['frame_index']:
                    raise ValueError('Subtitle cue has no visible frame in the actual picture clock')
                entry.update(_bound_track=True, ass_start=start['ass_time'], ass_end=end['ass_time'],
                             start=float(Fraction(start['pts'])), end=float(Fraction(end['pts'])))
                entry['frame_projection'] = {'requested_ticks': [cue['start_tick'], cue['end_tick']],
                                             'start': start, 'end': end}
            loaded['validation_schema'] = VALIDATION_SCHEMA
            loaded['projector_version'] = PROJECTOR_VERSION
            loaded['binding'] = {
                'input_video': str(Path(input_video).resolve()), 'identities': bindings,
                'inputs': {'track': file_identity(path), 'video': file_identity(input_video),
                           'edit_plan': file_identity(edit_plan_path) if edit_plan_path is not None else None},
                'duration': str(duration), 'frame_count': len(frame_pts),
                'edit_plan': str(Path(edit_plan_path).resolve()) if edit_plan_path is not None else None,
                'verification': 'media_binding_and_declared_frame_projection_only',
                'direct_listening': 'NOT_CHECKED', 'acoustic_alignment': 'NOT_CHECKED',
            }
            validation.write_text(json.dumps(loaded, ensure_ascii=False, indent=2) + '\n', encoding='utf-8')
            return loaded
        
        
        def _load_validation(work):
            path = Path(work) / VALIDATION
            if not path.exists():
                raise ValueError('Explicit subtitle track must be prepared against current media before use')
            loaded = json.loads(path.read_text(encoding='utf-8'))
            if (type(loaded.get('validation_schema')) is not int or loaded['validation_schema'] != VALIDATION_SCHEMA
                    or type(loaded.get('projector_version')) is not int or loaded['projector_version'] != PROJECTOR_VERSION):
                raise ValueError('Subtitle validation/projector version changed; prepare again')
            return loaded
        
        
        def bound_subtitle_entries(work_dir, video_duration):
            """Return prepared exact entries, or None only when no explicit track exists."""
            if work_dir is None:
                return None
            work = Path(work_dir)
            if not (work / TRACK).exists():
                return None
            loaded = _load_validation(work)
            record = loaded['binding']
            inputs = record['inputs']
            if file_identity(work / TRACK) != inputs['track']:
                raise ValueError('stale subtitle track: author file changed after prepare')
            video = Path(record['input_video'])
            if not video.is_file() or file_identity(video) != inputs['video']:
                raise ValueError('stale subtitle track: adopted media changed after prepare')
            if record['edit_plan'] and (not Path(record['edit_plan']).is_file()
                                        or file_identity(record['edit_plan']) != inputs['edit_plan']):
                raise ValueError('stale subtitle track: edit plan changed after prepare')
            if abs(float(Fraction(record['duration'])) - float(video_duration)) > 0.05:
                raise ValueError('stale subtitle track: consumer duration changed after prepare')
            return loaded['entries']
        
        
        def verify_rendered_picture(work_dir, output_path):
            """Check the actual rendered frame clock, not just pre-render declarations."""
            work = Path(work_dir)
            if not (work / TRACK).exists():
                return None
            path = work / VALIDATION
            loaded = json.loads(path.read_text(encoding='utf-8'))
            bound_subtitle_entries(work, float(Fraction(loaded['binding']['duration'])))
            loaded['rendered_picture'] = {'output': str(Path(output_path).resolve()),
                                          'frame_clock_verified': False}
            path.write_text(json.dumps(loaded, ensure_ascii=False, indent=2) + '\n', encoding='utf-8')
            if _picture_clock(output_path) != _picture_clock(loaded['binding']['input_video']):
                raise ValueError('Rendered frame clock changed; subtitle frame projection is no longer valid')
            loaded['rendered_picture'].update(frame_clock_verified=True,
                                              frame_count=loaded['binding']['frame_count'],
                                              identity=file_identity(output_path))
            path.write_text(json.dumps(loaded, ensure_ascii=False, indent=2) + '\n', encoding='utf-8')
            return loaded['rendered_picture']
        
        
        def manifest_subtitle_evidence(work_dir, input_video, output_path):
            """Only expose evidence for the actual input and verified output of this render."""
            work = Path(work_dir)
            if not (work / TRACK).exists():
                return None
            loaded = _load_validation(work)
            bound_subtitle_entries(work, float(Fraction(loaded['binding']['duration'])))
            rendered = loaded.get('rendered_picture', {})
            if rendered.get('frame_clock_verified') is not True:
                raise ValueError('Subtitle output frame clock has not been verified')
            if str(Path(input_video).resolve()) != loaded['binding']['input_video']:
                raise ValueError('stale subtitle manifest: input media differs')
            if not Path(output_path).is_file() or file_identity(output_path) != rendered.get('identity'):
                raise ValueError('stale subtitle manifest: actual output differs from verified render')
            return {'validation_path': str((work / VALIDATION).resolve()),
                    'binding': loaded['binding'], 'metadata': loaded['metadata'], 'rendered_picture': rendered}
        
      • __init__.py 56 B
        """Subtitle generation and track-binding subpackage."""
        
    • artifacts.py 1.4 KB
      """File identity and work-directory JSON helpers for video-assemble."""
      
      import json
      import os
      from pathlib import Path
      
      from lib import CONFIG
      
      
      def file_identity(path):
          """``{size, mtime_ns}`` of one input file: enough to notice a rewrite, cheap to take."""
          st = os.stat(os.fspath(path))
          return {"size": st.st_size, "mtime_ns": st.st_mtime_ns}
      
      
      def _artifact_identity(path):
          path = Path(path)
          return file_identity(path) if path.exists() else None
      
      
      def _explicit_source_video():
          """Return the cut-mode source video only when the caller opted in explicitly."""
          if not CONFIG["source_video_explicit"]:
              return ""
          return CONFIG["source_video"]
      
      
      def _source_video_identity():
          """``{path, size, mtime_ns}`` of the explicit cut-mode source video, else None."""
          source_video = _explicit_source_video()
          if not source_video:
              return None
          path = Path(source_video)
          return {"path": str(path.resolve()), **file_identity(path)}
      
      
      def _timeline_provenance_status(work_dir):
          data = _load_work_json(work_dir, "timeline.json")
          return None if data is None else data["provenance"]
      
      
      def _load_work_json(work_dir, name):
          """Parse a JSON artifact in work_dir; None only when the file does not exist."""
          path = Path(work_dir) / name
          if not path.exists():
              return None
          return json.loads(path.read_text(encoding="utf-8"))
      
    • assemble.py 29 KB
      """Canonical CLI and programmatic render entry for the self-contained video-assemble skill."""
      
      import json
      import os
      from pathlib import Path
      
      import artifacts
      import assemble_constants as constants
      import assembly_contract
      import assembly_settings
      import audio_mix
      import adoption.audio_mix_binding as audio_mix_binding
      import adoption.frozen_audio as frozen_audio
      import media
      import narration_audio
      import adoption.narration_binding as narration_binding
      import pair_media
      import render_preflight
      import adoption.strict_publish as strict_publish
      import subtitles.render as subtitle_render
      import subtitles.track_binding as subtitle_track_binding
      import timeline_emit
      import packaging
      import visual_render
      import lib
      
      __all__ = [
          "assemble_video",
          "main",
      ]
      
      
      AUDIO_MODES = ("narration", "source-mix", "adopted-packet-copy")
      
      
      _current_narration_binding = strict_publish.current_narration_binding
      _current_audio_mix_binding = strict_publish.current_audio_mix_binding
      
      
      def assemble_video(input_video, tts_segments, work_dir, output_path, *,
                         audio_mode="narration", audio_stream_index=0,
                         narration_adoption_path=None, tts_meta_path=None,
                         audio_mix_adoption_path=None):
          """组装最终视频"""
          if audio_mode not in AUDIO_MODES:
              raise RuntimeError(f"不支持的 audio_mode: {audio_mode}")
          if isinstance(audio_stream_index, bool) or not isinstance(audio_stream_index, int) or audio_stream_index < 0:
              raise RuntimeError("audio_stream_index 必须是非负整数")
          if audio_mode == "narration" and not tts_segments:
              raise RuntimeError("tts_meta.json 没有有效解说音频,已中止以避免生成无解说视频")
          if audio_mode == "narration" and audio_stream_index != 0:
              raise RuntimeError("narration 当前不支持非零 audio_stream_index")
          if audio_mode == "source-mix" and tts_segments:
              raise RuntimeError("source-mix 与 TTS 解说不兼容")
          if audio_mode != "narration" and tts_meta_path is not None:
              raise RuntimeError(f"tts_meta 与 audio_mode {audio_mode} 不兼容")
          bgm_path = lib.CONFIG["bgm_path"]
          has_bgm = bool(bgm_path) and os.path.exists(bgm_path)
          if audio_mode == "source-mix" and bgm_path and not has_bgm:
              raise RuntimeError(f"source-mix 声明的 BGM 文件不存在: {bgm_path}")
          if (
              audio_mode != "narration"
              and audio_stream_index != 0
              and lib.CONFIG["export_jianying"]
          ):
              raise RuntimeError("剪映导出当前不支持选择非零音频流")
          if audio_mode == "adopted-packet-copy":
              if tts_segments:
                  raise RuntimeError("adopted-packet-copy 与 TTS 解说不兼容")
              if bgm_path:
                  raise RuntimeError("adopted-packet-copy 与 BGM 混音不兼容")
          if audio_mode != "narration" and narration_adoption_path is not None:
              raise RuntimeError("narration adoption 仅适用于 narration audio_mode")
          if audio_mix_adoption_path is not None and (
              audio_mode != "narration" or narration_adoption_path is None or tts_meta_path is None
          ):
              raise RuntimeError("audio mix adoption 要求 narration 模式及显式 narration adoption/tts_meta")
      
          published_output = Path(output_path)
          if audio_mix_adoption_path is not None and published_output.exists():
              raise RuntimeError("显式音频混合要求新的 output_path,不能覆盖已有成片")
          explicit_mix = (
              audio_mix_binding.load_adoption(
                  audio_mix_adoption_path, input_video=input_video,
                  narration_adoption_path=narration_adoption_path, tts_segments=tts_segments,
              )
              if audio_mix_adoption_path is not None else None
          )
      
          binding = narration_binding.prepare_binding(
              tts_segments, work_dir,
              narration_adoption_path=narration_adoption_path,
              tts_meta_path=tts_meta_path,
          ) if audio_mode == "narration" else None
          render_output = published_output
          if binding and binding["active"]:
              if published_output.exists():
                  raise RuntimeError("身份约束渲染要求新的 output_path,不能覆盖已有成片")
              render_output = published_output.with_name(
                  f".{published_output.stem}.narration-rendering{published_output.suffix}"
              )
              render_output.unlink(missing_ok=True)
          output_path = render_output
          video_duration = lib.get_video_duration(input_video)
          canvas = media._probe_canvas(input_video)  # drives subtitle PlayRes/scale so 竖屏 text isn't stretched
          burn_subtitles = lib.CONFIG["burn_subtitles"]
          subtitle_track_binding.prepare_subtitle_track(
              input_video,
              work_dir,
              video_duration,
              audio_mode=audio_mode,
              selected_audio_stream=audio_stream_index,
          )
          adopted_source_audio = None
          if audio_mode == "adopted-packet-copy":
              adopted_source_audio = frozen_audio.validate_adopted_source(
                  input_video, audio_stream_index
              )
              if type(adopted_source_audio["sample_rate"]) is not int \
                      or adopted_source_audio["sample_rate"] <= 0:
                  raise RuntimeError("adopted audio sample rate must be a positive integer")
          elif audio_mode == "source-mix":
              frozen_audio.probe_audio_packets(input_video, audio_stream_index)
      
          # 解说整体提速(可选)后,将所有 TTS 片段按时间位置合成到与视频等长的音轨上
          narration_wav = None
          if audio_mode == "narration" and explicit_mix is not None:
              explicit_runtime = audio_mix_binding.render_explicit_mix(
                  explicit_mix, binding, tts_segments, work_dir
              )
              narration_wav = Path(explicit_runtime["voice_bus"]["path"])
          elif audio_mode == "narration":
              if binding["tempo_policy"]:
                  narration_audio._apply_narration_speed(
                      tts_segments, work_dir, tempo_policy=binding["tempo_policy"]
                  )
              else:
                  narration_audio._apply_narration_speed(tts_segments, work_dir)
              narration_wav = work_dir / "narration.wav"
              if binding["tempo_policy"]:
                  narration_audio._build_timed_narration(
                      tts_segments, narration_wav, video_duration, work_dir,
                      tempo_policy=binding["tempo_policy"],
                  )
                  if any(
                      segment.get("blocking") or segment.get("fit_status") == "no_safe_fit"
                      for segment in tts_segments
                  ):
                      raise RuntimeError(
                          "严格 narration adoption 存在 no_safe_fit,禁止提速或裁尾渲染"
                      )
              else:
                  narration_audio._build_timed_narration(
                      tts_segments, narration_wav, video_duration, work_dir
                  )
              handoffs = audio_mix._apply_source_sentence_handoffs(tts_segments, work_dir, video_duration)
              if handoffs:
                  lib.log(
                      "原声句末交接: "
                      + ", ".join(
                          f"{item['end']:.2f}s→{item.get('restore_at', item['end']):.2f}s({item['status']})"
                          for item in handoffs
                      )
                  )
              narration_binding.seal_render_inputs(binding, tts_segments, narration_wav)
      
          # 始终生成 SRT 字幕文件(原声留白处补烧原声字幕,传入成片时长以计算留白区间)
          srt_path = subtitle_render._generate_srt(tts_segments, work_dir, video_duration)
          lib.log(f"字幕文件: {srt_path}")
          ass_path = None
          if burn_subtitles:
              ass_path = subtitle_render._generate_ass(tts_segments, work_dir, video_duration, canvas)
              lib.log(f"压制字幕文件: {ass_path}")
      
          # 可选 BGM:作为一条独立音轨(input [2:a])混入,旁白处自动压低
          if explicit_mix is not None:
              lib.log("显式 adopted full-sound:忽略环境 BGM/duck/loudnorm/tempo 配置")
          elif bgm_path and not has_bgm:
              lib.log(f"  ⚠️ BGM 文件不存在,跳过: {bgm_path}")
          elif has_bgm:
              lib.log(f"BGM 铺底: {bgm_path} (音量 {lib.CONFIG['bgm_volume']},旁白时 {lib.CONFIG['bgm_ducking_volume']})")
      
          # 多轨时间线模型(timeline.json):canonical 渲染仍是 ffmpeg,此模型供检视/可选导出
          timeline_emit._emit_timeline(
              input_video, tts_segments, work_dir, video_duration, canvas, has_bgm,
              audio_mode=audio_mode, selected_audio_stream=audio_stream_index,
              explicit_audio_mix=(
                  {**explicit_mix, **explicit_mix["runtime"]} if explicit_mix is not None else None
              ),
          )
      
          overlay_filters, overlay_qc = visual_render._visual_overlay_filters(work_dir, canvas, video_duration)
          packaging_layers = packaging.load_packaging_layers(work_dir, canvas)
          mask_filter = visual_render._source_subtitle_mask_filter(canvas, work_dir, tts_segments, video_duration)
          visual_qc = visual_render._build_visual_qc(
              tts_segments,
              work_dir,
              video_duration,
              canvas,
              overlay_qc=overlay_qc,
              mask_filter=mask_filter,
          )
          visual_render._write_visual_qc(work_dir, visual_qc)
          assembly_qc_path = Path(work_dir) / constants.ASSEMBLY_QC
          assembly_qc_path.unlink(missing_ok=True)
          if visual_qc["blocking"]:
              codes = ", ".join(visual_qc["blocking_codes"])
              raise RuntimeError(f"视觉 QC 失败: {codes};详见 {Path(work_dir) / constants.VISUAL_QC}")
      
          # Select exactly one of three explicit audio paths. Only narration may synthesize
          # a missing original track; adopted copy never decodes, mixes, normalizes or trims.
          source_has_audio = media._has_audio_stream(input_video)
          adopted_audio = None
          loudnorm_measurement = None
          original_audio_input = []
          bgm_input = []
          filter_complex = None
          filter_args = []
          audio_input_args = []
          fc_script = None
          if audio_mode == "adopted-packet-copy":
              audio_map = f"0:a:{audio_stream_index}"
          elif audio_mode == "source-mix":
              audio_map = "[aoutln]"
              source_label = f"0:a:{audio_stream_index}"
              filter_complex = f"[{source_label}]volume={lib.CONFIG['idle_orig_volume']}[source]"
              if has_bgm:
                  bgm_input = ["-stream_loop", "-1", "-i", str(bgm_path)]
                  filter_complex += (
                      f";[1:a]volume={lib.CONFIG['bgm_volume']}[bgm]"
                      ";[source][bgm]amix=inputs=2:duration=first:dropout_transition=0[aout]"
                  )
              else:
                  filter_complex += ";[source]anull[aout]"
              final_ln = audio_mix.final_loudnorm_filter()
              filter_complex += f";[aout]{final_ln}[aoutln]"
              lib.log(f"source-mix 音频处理: source volume + {final_ln}")
          elif explicit_mix is not None:
              audio_map = "1:a:0"
              audio_input_args = ["-i", explicit_mix["runtime"]["master"]["path"]]
          else:
              # 混合原始音频 + 解说音频(+ 可选 BGM)
              if source_has_audio:
                  original_audio_label = "0:a"
                  bgm_audio_label = "2:a"
              else:
                  lib.log("源视频无音轨,使用静音原声音轨进行混音")
                  original_audio_input = [
                      "-f", "lavfi", "-t", str(video_duration),
                      "-i", "anullsrc=channel_layout=stereo:sample_rate=48000",
                  ]
                  original_audio_label = "2:a"
                  bgm_audio_label = "3:a"
              filter_complex = audio_mix._build_audio_filter_complex(
                  tts_segments,
                  has_bgm,
                  original_audio_label=original_audio_label,
                  bgm_audio_label=bgm_audio_label,
              )
              # BGM is input [2:a]; -stream_loop -1 loops it to cover the whole timeline (amix
              # duration=first + -t trim it back to the video length).
              bgm_input = ["-stream_loop", "-1", "-i", str(bgm_path)] if has_bgm else []
              # 末端整体响度归一:ducking 只管相对平衡,这一步统一成片绝对响度
              loudnorm_measurement = audio_mix._run_loudnorm_first_pass(
                  input_video,
                  narration_wav,
                  original_audio_input,
                  bgm_input,
                  filter_complex,
                  work_dir,
              )
              final_ln = audio_mix.final_loudnorm_filter(loudnorm_measurement)
              filter_complex += f";[aout]{final_ln}[aoutln]"
              lib.log(f"成片响度归一: {final_ln}")
              audio_map = "[aoutln]"
              audio_input_args = ["-i", str(narration_wav), *original_audio_input]
      
          # 对于超长 volume 表达式(多段解说),从脚本文件读取 filter_complex 避免命令行溢出
          if filter_complex is not None:
              if len(filter_complex.encode("utf-8")) > constants.FILTER_SCRIPT_THRESHOLD_BYTES:
                  fc_script = Path(work_dir) / ".filter_complex.txt"
                  fc_script.write_text(filter_complex, encoding="utf-8")
                  lib.log(f"使用 filter_complex 脚本文件 (表达式长度 {len(filter_complex.encode('utf-8'))} bytes)")
                  filter_args = lib.filter_file_args("filter_complex", fc_script)
              else:
                  filter_args = ["-filter_complex", filter_complex]
          cmd = [
              "ffmpeg", "-y",
              "-i", str(input_video),
              *audio_input_args,
              *bgm_input,
              *filter_args,
              # 0:v:0 (not 0:v): sources with attached cover art carry a second video stream.
              # -vf only ever applies to the first one, so mapping all of them makes ffmpeg
              # abort with "Could not write header (incorrect codec parameters ?)" and leave
              # an unreadable file — after the whole pipeline has already run.
              "-map", "0:v:0", "-map", audio_map,
          ]
      
          # Video filter chain: mask source subtitles first (drawbox), then burn our subtitles
          # on top. Either one forces a re-encode; with neither, the video stream is copied.
          crf = str(lib.CONFIG["output_crf"])  # env_int already clamps to >=0; keep 0 (lossless) intact
          preset = lib.CONFIG["output_preset"]
          max_h = lib.CONFIG["output_max_height"]
          vf_chain = []
          if mask_filter:
              vf_chain.append(mask_filter)
          vf_chain.extend(overlay_filters)
          if burn_subtitles:
              vf_chain.append(visual_render._subtitle_burn_filter(ass_path))
          # Downscale LAST so the mask + burned subtitles render at native resolution and are then
          # scaled down with the frame (crisp). The helper forces both dimensions even so an odd
          # OUTPUT_MAX_HEIGHT can't crash libx264; 'min(ih,H)' only ever shrinks the source.
          if max_h > 0:
              vf_chain.append(visual_render._output_downscale_filter(max_h))
          # yuv420p: 10-bit/4:2:2 sources re-encoded as-is play on desktop but fail on WeChat/
          # mobile/Safari; force 8-bit 4:2:0 so every recap is universally decodable. yuv420p also
          # needs EVEN width AND height, so normalize odd dims (4:2:2/4:4:4 permit them) before the
          # encode — otherwise libx264 aborts to a 0-byte file. The downscale helper already evens out.
          even = "scale=trunc(iw/2)*2:trunc(ih/2)*2"
          reencode = bool(vf_chain or packaging_layers) or lib.CONFIG["force_video_reencode"]
          notes = []
          video_filter_script = None
          if vf_chain or packaging_layers:
              if max_h <= 0:  # no downscale in the chain to force even dims
                  vf_chain.append(even)
              video_filter = packaging.compose_video_filter(
                  vf_chain, packaging_layers, mask_first=bool(mask_filter)
              )
              if len(video_filter.encode("utf-8")) > constants.FILTER_SCRIPT_THRESHOLD_BYTES:
                  video_filter_script = Path(work_dir) / ".video_filter.txt"
                  video_filter_script.write_text(video_filter, encoding="utf-8")
                  cmd += lib.filter_file_args("filter:v:0", video_filter_script)
                  lib.log(
                      "使用 video filter script "
                      f"(表达式长度 {len(video_filter.encode('utf-8'))} bytes)"
                  )
              else:
                  cmd += ["-vf", video_filter]
              cmd += ["-c:v", "libx264", "-preset", preset, "-crf", crf, "-pix_fmt", "yuv420p"]
              notes = ((["遮挡原字幕"] if mask_filter else [])
                       + ([f"包装图层×{len(packaging_layers)}"] if packaging_layers else [])
                       + ([f"视觉叠加×{len(overlay_filters)}"] if overlay_filters else [])
                       + (["压制解说字幕"] if burn_subtitles else [])
                       + ([f"缩放≤{max_h}p"] if max_h > 0 else []))
              lib.log(f"视频重编码: {' + '.join(notes)} (crf={crf}, preset={preset})")
          elif reencode:
              notes = ["force_video_reencode"]
              cmd += ["-vf", even, "-c:v", "libx264", "-preset", preset, "-crf", crf, "-pix_fmt", "yuv420p"]
          else:
              cmd += ["-c:v", "copy"]
      
          # +faststart relocates the moov atom to the front so web/social players can start
          # before the full file downloads; valid (and beneficial) on the copy path too.
          if audio_mode == "adopted-packet-copy":
              # No -t/-shortest: either would discard valid AAC priming or tail packets.
              cmd += ["-c:a", "copy", "-movie_timescale",
                      str(adopted_source_audio["sample_rate"]),
                      "-movflags", "+faststart", str(output_path)]
          elif explicit_mix is not None:
              cmd += ["-c:a", "aac", "-b:a", "192k", "-ar", "48000",
                      "-movie_timescale", "48000", "-movflags", "+faststart",
                      str(output_path)]
          else:
              cmd += ["-c:a", "aac", "-b:a", "192k", "-ar", "48000", "-movflags", "+faststart",
                      "-t", str(video_duration), str(output_path)]
          try:
              result = lib.run_cmd(cmd)
              if result.returncode != 0:
                  raise RuntimeError(f"视频组装失败: {result.stderr}")
          finally:
              # 清理临时 filter 脚本(无论 ffmpeg 是否成功)
              if fc_script is not None:
                  fc_script.unlink(missing_ok=True)
              if video_filter_script is not None:
                  video_filter_script.unlink(missing_ok=True)
      
          if audio_mode == "adopted-packet-copy":
              # Either check failing means the file at the final path is unverified: never leave it.
              try:
                  adopted_audio = frozen_audio.verify_adopted_audio(
                      input_video, output_path, audio_stream_index
                  )
                  pair_media.validate_aac_packet_interval(adopted_audio["output"])
              except (RuntimeError, ValueError):
                  output_path.unlink(missing_ok=True)
                  raise
          subtitle_track_binding.verify_rendered_picture(work_dir, output_path)
          audio_operations = {
              "narration": audio_mode == "narration",
              "source_mix": audio_mode == "source-mix",
              "bgm_mix": has_bgm and audio_mode != "adopted-packet-copy" and explicit_mix is None,
              "ducking": audio_mode == "narration" and explicit_mix is None,
              "loudness_normalization": (
                  audio_mode != "adopted-packet-copy" and explicit_mix is None
                  and lib.CONFIG["final_loudnorm"]
              ),
              "limiter": audio_mode != "adopted-packet-copy" and explicit_mix is None,
              "resample": audio_mode != "adopted-packet-copy",
              "tempo": audio_mode == "narration" and explicit_mix is None,
              "packet_copy": audio_mode == "adopted-packet-copy",
          }
          if explicit_mix is not None:
              audio_operations["explicit_audio_mix"] = True
          render_delivery = {
              "video_encode_passes": 1 if reencode else 0,
              "reencode_reason": notes,
              "audio_sample_rate": (
                  adopted_audio["output"]["sample_rate"] if adopted_audio else 48000
              ),
              "final_compat_notes": (
                  (["yuv420p"] if reencode else ["video_copy"])
                  + (["aac_packet_copy", "faststart"] if adopted_audio else ["aac_48000", "faststart"])
              ),
          }
          loudness_mode = (
              "not_run" if audio_mode == "adopted-packet-copy" else
              "fixed_master_gain_no_loudnorm" if explicit_mix is not None else None
          )
          source_audio_status = "prepared_bed_adopted" if explicit_mix is not None else None
          render_output = strict_publish.publish_render(
              work_dir=work_dir, binding=binding, explicit_mix=explicit_mix,
              tts_segments=tts_segments, narration_wav=narration_wav,
              render_output=render_output, published_output=published_output,
              audio_mode=audio_mode, audio_operations=audio_operations,
              adopted_audio=adopted_audio, loudness_mode=loudness_mode,
              loudnorm_measurement=loudnorm_measurement, visual_qc=visual_qc,
              source_has_audio=source_has_audio, video_duration=video_duration,
              render_delivery=render_delivery, source_audio_status=source_audio_status,
          )
          lib.log(f"最终视频: {render_output} ({render_output.stat().st_size / 1024 / 1024:.1f}MB)")
          return render_output
      
      
      def main():
          import argparse
          import shutil
          ap = argparse.ArgumentParser(
              description="video-assemble: mux narration audio over the video, duck the original, render subtitles.")
          ap.add_argument("video", help="source video (edited_source.mp4 in cut mode, else the original)")
          ap.add_argument("--work-dir", required=True)
          ap.add_argument("--tts-meta", default=None, help="tts_meta.json (default: <work-dir>/tts_meta.json)")
          ap.add_argument("--narration-adoption", default=None,
                          help="strict narration_adoption v1 bound to an explicit --tts-meta")
          ap.add_argument("--audio-mix-adoption", default=None,
                          help="strict audio_mix_adoption v1 for adopted prepared bed and narration")
          ap.add_argument(
              "--audio-mode", choices=AUDIO_MODES,
              default="narration", help="audio path (default: narration)",
          )
          ap.add_argument(
              "--audio-stream-index", type=int, default=0,
              help="zero-based input audio stream ordinal for source/adopted modes",
          )
          ap.add_argument("--recap-stem", default=None, help="final recap filename stem (default: video stem)")
          ap.add_argument("--output-dir", default=None)
          ap.add_argument("--burn-subtitles", action=argparse.BooleanOptionalAction, default=None,
                          help="burn narration subtitles into the video (default on; --no-burn-subtitles to disable)")
          ap.add_argument("--subtitle-y-top", type=int, default=None,
                          help="inclusive top of a measured subtitle band in display-frame pixels")
          ap.add_argument("--subtitle-y-bot", type=int, default=None,
                          help="exclusive bottom of a measured subtitle band in display-frame pixels")
          ap.add_argument("--source-video", default=None,
                          help="original source video (cut mode) so timeline.json / 剪映 export reference the real clips")
          ap.add_argument("--export-jianying", action="store_true",
                          help="also export an OPTIONAL 剪映/JianYing draft from timeline.json after rendering")
          ap.add_argument("--jianying-out", default=None, help="parent dir for the 剪映 draft (default: work-dir)")
          bundle_group = ap.add_mutually_exclusive_group()
          bundle_group.add_argument("--jianying-bundle-media", dest="jianying_bundle_media", action="store_true",
                                    help="copy media into the 剪映 draft folder (default on; portable/self-contained)")
          bundle_group.add_argument("--jianying-no-bundle-media", dest="jianying_bundle_media", action="store_false",
                                    help="do NOT copy media into the draft — reference in place (only if 剪映 can read those paths; macOS 剪映 usually cannot)")
          ap.set_defaults(jianying_bundle_media=None)
          args = ap.parse_args()
          work_dir = Path(args.work_dir)
          if args.burn_subtitles is not None:
              lib.CONFIG["burn_subtitles"] = args.burn_subtitles
          if (args.subtitle_y_top is None) != (args.subtitle_y_bot is None):
              ap.error("--subtitle-y-top and --subtitle-y-bot must be provided together")
          if args.subtitle_y_top is not None:
              if args.subtitle_y_top < 0 or args.subtitle_y_bot <= args.subtitle_y_top:
                  ap.error("subtitle Y coordinates must satisfy 0 <= top < bot")
              lib.CONFIG["subtitle_y_top"] = args.subtitle_y_top
              lib.CONFIG["subtitle_y_bot"] = args.subtitle_y_bot
              lib.CONFIG["mask_source_subtitles"] = True
              lib.CONFIG["source_subtitle_mask_policy"] = "opt_in"
              lib.CONFIG["source_subtitle_mask_policy_declared"] = True
              # A measured band is an explicit request to conceal the known source-caption
              # pixels.  The general 0.6 translucent look can leave white glyphs visible under
              # the generated subtitles; use an opaque mask unless the caller deliberately
              # chose a different opacity through the existing environment override.
              if "SUBTITLE_MASK_OPACITY" not in os.environ:
                  lib.CONFIG["subtitle_mask_opacity"] = 1.0
          if args.source_video:
              if not os.path.exists(args.source_video):
                  ap.error(f"--source-video does not exist: {args.source_video}")
              lib.CONFIG["source_video"] = args.source_video
              lib.CONFIG["source_video_explicit"] = True
          else:
              # SOURCE_VIDEO is an ambient env var in lib.CONFIG. Do not let a stale
              # shell value silently bind full-mode/direct timeline.json or JianYing
              # exports to an unrelated original; cut mode must pass --source-video.
              lib.CONFIG["source_video"] = ""
              lib.CONFIG["source_video_explicit"] = False
          if args.export_jianying:
              lib.CONFIG["export_jianying"] = True
          if args.jianying_bundle_media is not None:
              lib.CONFIG["jianying_bundle_media"] = args.jianying_bundle_media
          render_preflight._preflight_burn_subtitles()  # fail before the render if burn-in is on but ffmpeg lacks libass
          # Argument combinations are validated once, by assemble_video.
          tts_meta = Path(args.tts_meta) if args.tts_meta else None
          tts_segments = []
          if args.audio_mode == "narration":
              tts_meta = tts_meta or work_dir / "tts_meta.json"
              tts_segments = json.loads(tts_meta.read_text(encoding="utf-8"))["segments"]
          stem = args.recap_stem or Path(args.video).stem
          base = Path(args.output_dir) if args.output_dir else work_dir.parent
          final_output = assembly_contract._resolve_final_output(base, stem)
          if args.audio_mix_adoption is not None and final_output.exists():
              ap.error("explicit audio mix requires a new final delivery path")
          delivery_stage = None
          owned_alias = None
          output_path = work_dir / "output.mp4"
          try:
              assemble_video(
                  args.video, tts_segments, work_dir, output_path,
                  audio_mode=args.audio_mode, audio_stream_index=args.audio_stream_index,
                  narration_adoption_path=args.narration_adoption, tts_meta_path=tts_meta,
                  audio_mix_adoption_path=args.audio_mix_adoption,
              )
              assembly_qc = artifacts._load_work_json(work_dir, constants.ASSEMBLY_QC)
              if assembly_qc["blocking"]:
                  codes = ", ".join(assembly_qc["blocking_codes"])
                  raise SystemExit(
                      f"组装 QC 阻断交付: {codes};详见 {work_dir / constants.ASSEMBLY_QC}"
                  )
              base.mkdir(parents=True, exist_ok=True)
              if args.audio_mix_adoption is not None:
                  # Publish the delivery alias only when this process created it: stage a copy,
                  # hard-link it into place (fails if the alias appeared meanwhile), remember the
                  # inode, and roll back only an alias we own.
                  delivery_stage = final_output.with_name(
                      f".{final_output.name}.rendering-{os.getpid()}"
                  )
                  with output_path.open("rb") as source, delivery_stage.open("xb") as target:
                      shutil.copyfileobj(source, target)
                      target.flush()
                      os.fsync(target.fileno())
                  try:
                      os.link(delivery_stage, final_output)
                  except FileExistsError as exc:
                      raise RuntimeError("explicit audio mix delivery alias appeared during render") from exc
                  stat = final_output.stat()
                  owned_alias = (stat.st_dev, stat.st_ino)
                  delivery_stage.unlink()
                  delivery_stage = None
              else:
                  shutil.copy2(str(output_path), str(final_output))
              manifest = assembly_contract._assembly_manifest_payload(
                  args.video, tts_segments, work_dir, output_path,
                  tts_meta_path=tts_meta,
                  narration_input_binding=_current_narration_binding(work_dir, args.audio_mode),
                  audio_mix_binding=_current_audio_mix_binding(work_dir, args.audio_mode),
                  final_output=final_output,
                  settings_payload=assembly_settings.assembly_settings_payload,
                  audio_mode=args.audio_mode,
                  audio_stream_index=args.audio_stream_index,
              )
              assembly_contract._write_assembly_manifest(work_dir, manifest)
          except BaseException:
              if delivery_stage is not None:
                  delivery_stage.unlink(missing_ok=True)
              if owned_alias is not None and final_output.exists():
                  stat = final_output.stat()
                  if (stat.st_dev, stat.st_ino) == owned_alias:
                      final_output.unlink()
              raise
          lib.log(f"组装完成: {final_output}")
      
          # OPTIONAL, decoupled: export a 剪映 draft from the timeline (lazy import; never
          # required by the core render path).
          if lib.CONFIG["export_jianying"]:
              from jianying.optional import maybe_export_jianying
              maybe_export_jianying(work_dir, args.jianying_out, stem)
      
          print(json.dumps({"status": "assembled", "output": str(final_output), "work_dir": str(work_dir)},
                           ensure_ascii=False))
      
      
      if __name__ == "__main__":
          main()
      
    • assemble_constants.py 2 KB
      """Shared constants for the self-contained video-assemble skill."""
      
      from fractions import Fraction
      import math
      
      ASSEMBLY_MANIFEST = "assembly_manifest.json"
      ASSEMBLY_QC = "assembly_qc.json"
      VISUAL_QC = "visual_qc.json"
      VISUAL_OVERLAYS = "visual_overlays.json"
      SEGMENT_AUDIO_SCHEMA_VERSION = 1
      
      # The picture codecs every packet/frame clock proof in this skill accepts.
      SUPPORTED_PICTURE_CODECS = frozenset({"h264", "hevc"})
      
      # The one exact output audio clock every explicit-sound artifact in this skill uses.
      OUTPUT_SAMPLE_RATE = 48_000
      
      
      def frame_clock_samples(frame, fps, rate=OUTPUT_SAMPLE_RATE):
          """Project an exact frame boundary of a video clock onto the output sample clock.
      
          Broadcast rates such as 30000/1001 put frame boundaries between whole samples, so
          the exact Fraction position is rounded half-up once, here. Every sample bound in
          the skill - a segment edge and the whole-picture `total_samples` alike - is this
          single projection, so an integral clock is unchanged and a fractional one stays
          consistent across the explicit-sound tools.
          """
          exact = Fraction(int(frame) * rate, 1) / Fraction(fps)
          return math.floor(exact + Fraction(1, 2))
      
      
      FILTER_SCRIPT_THRESHOLD_BYTES = 8000
      
      # The default subtitle metrics were tuned in this reference canvas.
      SUBTITLE_STYLE_REF_W = 1280
      SUBTITLE_STYLE_REF_H = 720
      _SUBTITLE_TERMINAL_PUNCTUATION = "。!?!?…."
      _SUBTITLE_CLOSING_QUOTES = "」』”’))]】》〉\"'"
      
      _MIN_GAP_TO_SUBTITLE = 0.8
      _MIN_READABLE_SECONDS = 0.3
      _MIN_ASR_CLIP_OVERLAP = 0.05
      # timeline.py serializes interval bounds onto a 1e-4 second grid, flooring starts
      # and ceiling ends, so two bounds that were identical before serialization can come
      # back one grid step apart. Contiguity joins must tolerate that whole step.
      _TIMELINE_TIME_GRID_SECONDS = 1e-4
      _CLIP_CONTIGUITY_TOLERANCE = 1.5 * _TIMELINE_TIME_GRID_SECONDS
      _MAX_ORIGINAL_READ_CPS = 9.0
      _AUTO_ORIGINAL_READ_CPS = 6.0
      
      _SUPPORTED_VISUAL_OVERLAY_TYPES = {"top_title", "inline_label_or_callout"}
      
    • assembly_contract.py 13.6 KB
      """Assembly manifest/QC persistence and delivery contract helpers."""
      
      import json
      from fractions import Fraction
      import math
      import subprocess
      import wave
      from pathlib import Path
      
      from lib import CONFIG
      from assemble_constants import (
          ASSEMBLY_MANIFEST,
          ASSEMBLY_QC,
          SEGMENT_AUDIO_SCHEMA_VERSION,
      )
      from audio_mix import _loudness_mode
      from subtitles.track_binding import manifest_subtitle_evidence
      from artifacts import (
          _load_work_json,
          _source_video_identity,
          _timeline_provenance_status,
      )
      
      _AUDIO_QC_CODES = frozenset({
          "missing_narration", "skipped_segments", "no_safe_fit", "effective_tempo_exceeded",
          "empty_narration", "truncated_speech", "unsafe_source_handoff", "timeline_audio_mismatch",
      })
      
      
      def _assembly_manifest_payload(input_video, tts_segments, work_dir, output_path,
                                     tts_meta_path=None, final_output=None, *, settings_payload,
                                     audio_mode="narration", audio_stream_index=0,
                                     narration_input_binding=None, audio_mix_binding=None):
          """Slim render record. The orchestrator reads `final_output` to report the result;
          `source_video` stays None unless cut mode explicitly passed --source-video, proving a
          stale ambient SOURCE_VIDEO never leaked into a full-mode timeline / 剪映 export."""
          input_video = Path(input_video)
          output_path = Path(output_path)
          source_video_identity = _source_video_identity()
          qc_path = Path(work_dir) / ASSEMBLY_QC
          qc = _load_work_json(work_dir, ASSEMBLY_QC)  # always written by publish_render first
          settings = settings_payload(
              work_dir, audio_mode=audio_mode, audio_stream_index=audio_stream_index
          )
          payload = {
              "schema_version": 2,
              "input_video": str(input_video.resolve()),
              "source_video": source_video_identity["path"] if source_video_identity else None,
              "source_video_identity": source_video_identity,
              "tts_meta": str(Path(tts_meta_path).resolve()) if tts_meta_path else None,
              "tts_segments": len(tts_segments),
              "audio_mode": audio_mode,
              "selected_audio_stream_index": audio_stream_index,
              "assembly_settings": settings,
              "output_path": str(output_path.resolve()),
              "segment_audio_schema_version": SEGMENT_AUDIO_SCHEMA_VERSION,
              "qc_path": str(qc_path.resolve()),
              "qc_verdict": qc["verdict"],
              "qc_blocking_codes": qc["blocking_codes"],
              # The settings payload records the configured/fallback loudness policy; these QC
              # fields record what the just-finished render actually used after the loudnorm probe.
              "qc_loudness_mode": qc["loudness_mode"],
              "qc_loudnorm_measurement": qc["loudnorm_measurement"],
              "audio_operations": qc["audio_operations"],
              "adopted_audio": qc["adopted_audio"],
              "narration_input_binding": narration_input_binding,
              "audio_mix_binding": audio_mix_binding,
              "audio_segments": [
                  {
                      "index": seg["index"],
                      # Informational manifest fields: a tts_meta without the schema marker is v1,
                      # and the loudness measurements are None whenever voiceover skipped
                      # normalization (strict adoption fixtures in tests/orchestrator omit both).
                      "segment_audio_schema_version": seg.get(
                          "segment_audio_schema_version", SEGMENT_AUDIO_SCHEMA_VERSION
                      ),
                      "narration": seg["narration"],
                      "spoken_text": seg["spoken_text"],
                      "truncated": seg["truncated"],
                      "truncate_reason": seg["truncate_reason"],
                      "fit_status": seg["fit_status"],
                      "blocking": seg["blocking"],
                      "audio_duration": seg["audio_duration"],
                      "placed_audio_duration": seg["placed_audio_duration"],
                      "placed_audio_path": seg.get("placed_audio_path"),
                      "actual_place_start": seg.get("actual_place_start"),
                      "actual_place_end": seg.get("actual_place_end"),
                      "source_duck_end": seg.get("source_duck_end"),
                      "source_restore_at": seg.get("source_restore_at"),
                      "source_handoff_status": seg.get("source_handoff_status"),
                      "source_entry_status": seg.get("source_entry_status"),
                      "global_narration_speed": seg["global_narration_speed"],
                      "segment_tempo_factor": seg["segment_tempo_factor"],
                      "effective_tempo": seg["effective_tempo"],
                      "rms_dbfs_before": seg.get("rms_dbfs_before"),
                      "rms_dbfs_after": seg.get("rms_dbfs_after"),
                      "peak_after": seg.get("peak_after"),
                      "output_start_sample": seg.get("output_start_sample"),
                      "output_end_sample": seg.get("output_end_sample"),
                      "adopted_gain": seg.get("adopted_gain"),
                  }
                  for seg in tts_segments
              ],
          }
          if final_output is not None:
              payload["final_output"] = str(Path(final_output).resolve())
          provenance = _timeline_provenance_status(work_dir)
          if provenance:
              payload["timeline_provenance"] = provenance
          subtitle_evidence = manifest_subtitle_evidence(work_dir, input_video, output_path)
          if subtitle_evidence is not None:
              payload["subtitle_track"] = subtitle_evidence
          return payload
      
      
      def _write_assembly_manifest(work_dir, manifest):
          path = Path(work_dir) / ASSEMBLY_MANIFEST
          path.write_text(json.dumps(manifest, ensure_ascii=False, indent=2), encoding="utf-8")
          return path
      
      
      def _visual_qc_rollup(visual_qc):
          subtitles = visual_qc["subtitles"]
          overlays = visual_qc["overlays"]
          return {
              "present": True,
              "artifact": visual_qc["artifact"],
              "verdict": visual_qc["verdict"],
              "blocking": visual_qc["blocking"],
              "blocking_codes": list(visual_qc["blocking_codes"]),
              "summary": visual_qc["summary"],
              "geometry": visual_qc["geometry"],
              "subtitles": {
                  "entries": subtitles["entries"],
                  "multi_line": subtitles["multi_line"],
                  "overflow": subtitles["overflow"],
                  "safe_area": subtitles["safe_area"],
              },
              "mask": visual_qc["mask"],
              "overlays": {
                  "present": overlays["present"],
                  "rendered": overlays["rendered"],
                  "unsupported": overlays["unsupported"],
                  "overflow": overlays["overflow"],
              },
          }
      
      
      def _placed_audio_matches_timeline(seg):
          """True when the persisted per-beat WAV is exactly what the serialized timeline window plays."""
          placed_path = Path(seg["placed_audio_path"])
          if not placed_path.exists():
              return False
          try:
              with wave.open(str(placed_path), "rb") as placed_wav:
                  placed_duration = placed_wav.getnframes() / placed_wav.getframerate()
                  tolerance = 1.0 / placed_wav.getframerate()
          except wave.Error:
              result = subprocess.run(
                  ["ffprobe", "-v", "error", "-select_streams", "a:0", "-show_entries",
                   "stream=sample_rate,time_base,duration_ts", "-of", "json", str(placed_path)],
                  capture_output=True, text=True, timeout=600,
              )
              streams = json.loads(result.stdout).get("streams", []) if not result.returncode else []
              if len(streams) != 1:
                  return False
              stream = streams[0]
              try:
                  rate = int(stream["sample_rate"])
                  placed_duration = float(Fraction(stream["duration_ts"]) * Fraction(stream["time_base"]))
                  tolerance = 1.0 / rate
              except (KeyError, TypeError, ValueError, ZeroDivisionError):
                  return False
          timeline_start = math.floor(float(seg["actual_place_start"]) * 10_000 + 1e-9) / 10_000
          timeline_end = math.ceil(float(seg["actual_place_end"]) * 10_000 - 1e-9) / 10_000
          serialized_span = timeline_end - timeline_start
          return (
              abs(placed_duration - seg["placed_audio_duration"]) <= tolerance + 1e-9
              and serialized_span + 1e-9 >= placed_duration
          )
      
      
      def _build_assembly_qc(tts_segments, video_duration, *, audio_operations, render_delivery,
                             output_path=None, source_has_audio=None, loudness_mode=None,
                             loudnorm_measurement=None, visual_qc=None, audio_mode="narration",
                             adopted_audio=None,
                             narration_input_binding=None, audio_mix_binding=None,
                             source_audio_status=None):
          """Machine-readable assembly release gate.
      
          Visual facts are rolled up from visual_qc.json. Delivery/render facts live here
          (or render/delivery QC in future), never in visual_qc.json.
          """
          hard_max = CONFIG["narration_cumulative_tempo_hard_max"]
          segments = tts_segments
          no_safe = [
              s["index"] for s in segments
              if s["fit_status"] == "no_safe_fit"
              or s["truncate_reason"] in {"no_safe_boundary", "no_room"}
              or s["blocking"]
          ]
          skipped = [s["index"] for s in segments if s["fit_status"] == "skipped"]
          tempo_exceeded = []
          truncated = []
          handoff_failed = []
          timeline_audio_failed = []
          max_effective = 0.0
          for s in segments:
              eff = float(s["effective_tempo"])
              max_effective = max(max_effective, eff)
              if eff > hard_max + 1e-6:
                  tempo_exceeded.append(s["index"])
              if s["truncated"] or s["truncate_reason"] == "tail_trim_tolerance":
                  truncated.append(s["index"])
              if s.get("source_handoff_blocking", False):
                  handoff_failed.append(s["index"])
              if s["placed_audio_duration"] > 0 and not _placed_audio_matches_timeline(s):
                  timeline_audio_failed.append(s["index"])
      
          placed = [s["placed_audio_duration"] for s in segments]
          blocking_codes = []
          if audio_mode == "narration" and not segments:
              blocking_codes.append("missing_narration")
          if skipped:
              blocking_codes.append("skipped_segments")
          if no_safe:
              blocking_codes.append("no_safe_fit")
          if tempo_exceeded:
              blocking_codes.append("effective_tempo_exceeded")
          if truncated:
              blocking_codes.append("truncated_speech")
          if handoff_failed:
              blocking_codes.append("unsafe_source_handoff")
          if timeline_audio_failed:
              blocking_codes.append("timeline_audio_mismatch")
          if placed and max(placed) <= 0.0 and not no_safe:
              blocking_codes.append("empty_narration")
          visual_rollup = (
              _visual_qc_rollup(visual_qc)
              if visual_qc is not None
              else {"present": False, "verdict": "NOT_RUN", "blocking": False, "blocking_codes": [], "summary": {}}
          )
          if visual_rollup["blocking"]:
              blocking_codes.append("visual_qc_failed")
          if source_audio_status is not None:
              source_audio = source_audio_status
          elif source_has_audio is False:
              # Not blocking: assemble can synthesize a silent original track.
              source_audio = "synthetic_silence"
          elif source_has_audio is True:
              source_audio = "present"
          else:
              source_audio = "unknown"
      
          output = {}
          if output_path is not None:
              output_path = Path(output_path)
              output = {
                  "path": str(output_path),
                  "exists": output_path.exists(),
                  "bytes": output_path.stat().st_size if output_path.exists() else 0,
              }
              if output_path.exists() and output["bytes"] <= 0:
                  blocking_codes.append("empty_output")
      
          return {
              "schema_version": 1,
              "artifact": ASSEMBLY_QC,
              "verdict": "FAIL" if blocking_codes else "PASS",
              "blocking": bool(blocking_codes),
              "blocking_codes": blocking_codes,
              "duration": round(float(video_duration), 4),
              "audio_mode": audio_mode,
              "audio_operations": audio_operations,
              "adopted_audio": adopted_audio,
              "narration_input_binding": narration_input_binding,
              "audio_mix_binding": audio_mix_binding,
              "source_audio": source_audio,
              "loudness_mode": loudness_mode or _loudness_mode(loudnorm_measurement),
              "loudnorm_measurement": loudnorm_measurement,
              "release_gate": {
                  "verdict": "FAIL" if blocking_codes else "PASS",
                  "visual_qc": visual_rollup["verdict"],
                  "delivery_qc": "PASS",
                  "audio_qc": "FAIL" if _AUDIO_QC_CODES.intersection(blocking_codes) else "PASS",
              },
              "visual_qc": visual_rollup,
              "delivery_qc": {
                  "video_encode_passes": render_delivery["video_encode_passes"],
                  "reencode_reason": render_delivery["reencode_reason"],
                  "audio_sample_rate": render_delivery["audio_sample_rate"],
                  "final_compat_notes": render_delivery["final_compat_notes"],
              },
              "summary": {
                  "segments": len(segments),
                  "placed_segments": sum(1 for x in placed if x > 0.0),
                  "skipped_segments": skipped,
                  "no_safe_fit_segments": no_safe,
                  "tempo_exceeded_segments": tempo_exceeded,
                  "max_effective_tempo": round(max_effective, 4),
                  "truncated_segments": truncated,
                  "unsafe_source_handoff_segments": handoff_failed,
                  "timeline_audio_mismatch_segments": timeline_audio_failed,
              },
              "output": output,
          }
      
      
      def _write_assembly_qc(work_dir, qc):
          path = Path(work_dir) / ASSEMBLY_QC
          path.write_text(json.dumps(qc, ensure_ascii=False, indent=2), encoding="utf-8")
          return path
      
      
      def _resolve_final_output(base, stem):
          """The recap output is the stable human alias recap_<stem>.mp4, overwritten in place
          on every run so the iterate-on-narration loop always refreshes the same file."""
          return Path(base) / f"recap_{stem}.mp4"
      
    • assembly_settings.py 6.6 KB
      """Render-affecting settings payload recorded in the manifest and compared by resume logic."""
      
      from pathlib import Path
      
      from artifacts import _artifact_identity
      from assemble_constants import VISUAL_OVERLAYS
      from audio_mix import _loudness_mode, final_loudnorm_filter
      from lib import CONFIG
      from source_subtitles import _has_user_subtitles, _source_subtitle_mask_policy
      from subtitles.core import _subtitle_style_config
      from adoption.narration_binding import binding_record
      from adoption.audio_mix_binding import binding_record as audio_mix_binding_record
      from packaging import packaging_settings
      
      
      def assembly_settings_payload(work_dir=None, *, audio_mode="narration", audio_stream_index=0):
          """Settings that affect the rendered video, as a plain nested dict compared with ``==`` by
          pipeline resume logic. When work_dir is given, a user_subtitles presence flag and the
          ``{size, mtime_ns}`` identity of the overlay/subtitle-track inputs are included so dropping
          in or rewriting one of those files rebuilds the cached subtitles."""
          burn_subtitles = CONFIG["burn_subtitles"]
          mask_policy = _source_subtitle_mask_policy(work_dir)
          mask_source_subtitles = mask_policy["active"]
          overlay_identity = (
              _artifact_identity(Path(work_dir) / VISUAL_OVERLAYS) if work_dir is not None else None
          )
          subtitle_track_identity = (
              _artifact_identity(Path(work_dir) / "subtitle_track.json")
              if work_dir is not None else None
          )
          settings = {
              "user_subtitles": _has_user_subtitles(work_dir),
              "burn_subtitles": burn_subtitles,
              "force_video_reencode": CONFIG["force_video_reencode"],
              "encode": {
                  "output_crf": CONFIG["output_crf"],
                  "output_preset": CONFIG["output_preset"],
                  "output_max_height": CONFIG["output_max_height"],
              },
              "video_filters": {
                  "mask_source_subtitles": mask_source_subtitles,
                  "source_subtitle_mask_policy": mask_policy["policy"],
                  "source_subtitle_mask_policy_declared": mask_policy["declared"],
                  "source_subtitle_mask_policy_trigger": mask_policy["trigger"],
                  "source_subtitle_mask_ratio": (
                      CONFIG["source_subtitle_mask_ratio"] if mask_source_subtitles else None
                  ),
                  "source_subtitle_mask_timing": (
                      CONFIG["source_subtitle_mask_timing"] if mask_source_subtitles else None
                  ),
                  "subtitle_mask_opacity": (
                      CONFIG["subtitle_mask_opacity"] if mask_source_subtitles else None
                  ),
                  "subtitle_mask_padding": (
                      CONFIG["subtitle_mask_padding"] if mask_source_subtitles else None
                  ),
                  "subtitle_y_top": CONFIG["subtitle_y_top"],
                  "subtitle_y_bot": CONFIG["subtitle_y_bot"],
                  "visual_overlays": {
                      "artifact": VISUAL_OVERLAYS,
                      "present": overlay_identity is not None,
                      "identity": overlay_identity,
                  },
                  "packaging_layers": packaging_settings(work_dir),
              },
              "audio": {
                  "mode": audio_mode,
                  "selected_stream_index": audio_stream_index,
              },
          }
          explicit_mix = (
              audio_mix_binding_record(work_dir)
              if work_dir is not None and audio_mode == "narration" else None
          )
          if explicit_mix:
              settings["audio"]["path"] = "explicit_adopted_full_sound"
              settings["audio_mix_binding"] = explicit_mix
          if audio_mode == "narration":
              narration_binding = binding_record(work_dir) if work_dir else None
              settings["narration_input_binding"] = narration_binding
              adopted_tempo = (
                  narration_binding.get("tempo_policy") if narration_binding else None
              )
              settings["narration_timing"] = {
                  "delay_seconds": 0.0 if explicit_mix else CONFIG["narration_delay_seconds"],
                  "tail_pad_seconds": 0.0 if explicit_mix else CONFIG["narration_tail_pad_seconds"],
                  "fade_ms": 0 if explicit_mix else CONFIG["fade_ms"],
                  "narration_speed": (
                      1.0 if explicit_mix else
                      adopted_tempo["global_atempo"] if adopted_tempo else CONFIG["narration_speed"]
                  ),
                  "tempo_source": (
                      "explicit_audio_mix" if explicit_mix else
                      "adoption" if adopted_tempo else "configuration"
                  ),
                  "narration_cumulative_tempo_max": CONFIG["narration_cumulative_tempo_max"],
                  "tts_segment_tempo_max": (
                      adopted_tempo["segment_tempo_max"]
                      if explicit_mix and adopted_tempo else CONFIG["tts_segment_tempo_max"]
                  ),
              }
              if explicit_mix and adopted_tempo:
                  settings["narration_timing"]["narration_cumulative_tempo_max"] = \
                      adopted_tempo["cumulative_tempo_max"]
                  settings["narration_timing"]["narration_cumulative_tempo_hard_max"] = \
                      adopted_tempo["cumulative_tempo_hard_max"]
          if audio_mode in {"narration", "source-mix"} and not explicit_mix:
              # adopted-packet-copy never decodes or mixes, so mix settings cannot change it.
              settings["audio_mix"] = {
                  "ducking_mode": CONFIG["ducking_mode"],
                  "duck_fade_seconds": CONFIG["duck_fade_seconds"],
                  "duck_bridge_seconds": CONFIG["duck_bridge_seconds"],
                  "ducking_narr_weight": CONFIG["ducking_narr_weight"],
                  "ducking_orig_volume": CONFIG["ducking_orig_volume"],
                  "idle_orig_volume": CONFIG["idle_orig_volume"],
                  "speech_ducking_volume": CONFIG["speech_ducking_volume"],
                  "zone_ducking_volume": CONFIG["zone_ducking_volume"],
                  "ducking_threshold": CONFIG["ducking_threshold"],
                  "ducking_ratio": CONFIG["ducking_ratio"],
                  "ducking_attack": CONFIG["ducking_attack"],
                  "ducking_release": CONFIG["ducking_release"],
                  "ducking_level_sc": CONFIG["ducking_level_sc"],
                  "ducking_makeup": CONFIG["ducking_makeup"],
                  "final_loudnorm": final_loudnorm_filter(),
                  "loudness_mode": _loudness_mode(),
                  "bgm_path": CONFIG["bgm_path"],
                  "bgm_volume": CONFIG["bgm_volume"],
                  "bgm_ducking_volume": CONFIG["bgm_ducking_volume"],
              }
          if burn_subtitles:
              settings["subtitle_renderer"] = "ass"
              settings["subtitle_style"] = _subtitle_style_config()
          if subtitle_track_identity is not None:
              settings["subtitle_track"] = {
                  "artifact": "subtitle_track.json",
                  "identity": subtitle_track_identity,
              }
          return settings
      
    • audio_automation.py 6.7 KB
      """Shared audio automation semantics for ffmpeg render and editable timelines.
      
      This module is stdlib-only and is the single source for ducking window coalescing,
      pre-roll/hold/post-roll gain shape, timeline keyframes, and ffmpeg volume terms.
      """
      
      
      def _round_keyframe(t_s, gain):
          return {"t": round(float(t_s), 4), "gain": round(float(gain), 4)}
      
      
      def default_bridge(fade):
          return 2 * float(fade)
      
      
      def coalesce_duck_windows(windows, bridge):
          """Merge [(start, end, gain)] windows separated by gaps below `bridge`.
      
          A bridged mixed-level span uses the lowest gain across all members so neither
          ffmpeg nor timeline export swells louder in the bridged gap.
          """
          rel = sorted(
              ([float(s), float(e), float(g)] for s, e, g in windows if float(e) > float(s)),
              key=lambda w: w[0],
          )
          if not rel:
              return []
          bridge = float(bridge)
          merged = [rel[0][:]]
          for s, e, gain in rel[1:]:
              if s - merged[-1][1] < bridge:
                  merged[-1][1] = max(merged[-1][1], e)
                  merged[-1][2] = min(merged[-1][2], gain)
              else:
                  merged.append([s, e, gain])
          return merged
      
      
      def _t_minus(value):
          value = float(value)
          if value < 0:
              return f"t+{abs(value):.2f}"
          return f"t-{value:.2f}"
      
      
      def duck_ramp_expression(start, end, fade):
          """Return ffmpeg expression for the canonical duck shape.
      
          Semantics: ramp down during [start-fade, start], hold fully ducked on
          [start, end], and release during [end, end+fade].
          """
          start = float(start)
          end = float(end)
          fade = float(fade)
          if fade <= 0:
              return f"between(t,{start:.2f},{end:.2f})"
          ramp_start = start - fade
          ramp_end = end + fade
          return f"min(1,max(0,min({_t_minus(ramp_start)},{ramp_end:.2f}-t)/{fade:.2f}))"
      
      
      def ducking_expression(windows, idle, fade):
          """Build the ffmpeg volume expression for coalesced duck windows."""
          if not windows:
              return None
          idle = float(idle)
          terms = [
              f"+({float(level) - idle:.3f})*{duck_ramp_expression(s, e, fade)}"
              for s, e, level in windows
          ]
          return f"max(0,min(1,{idle}{''.join(terms)}))"
      
      
      def coalesce_release_duck_windows(windows, bridge):
          """Merge [(start, hold_end, gain, restore_at)] while preserving safe releases."""
          rel = sorted(
              ([float(s), float(e), float(g), max(float(e), float(r))] for s, e, g, r in windows if float(e) > float(s)),
              key=lambda row: row[0],
          )
          if not rel:
              return []
          bridge = float(bridge)
          merged = [rel[0][:]]
          for start, end, gain, restore in rel[1:]:
              if start - merged[-1][1] < bridge:
                  merged[-1][1] = max(merged[-1][1], end)
                  merged[-1][2] = min(merged[-1][2], gain)
                  merged[-1][3] = max(merged[-1][1], merged[-1][3], restore)
              else:
                  merged.append([start, end, gain, restore])
          return merged
      
      
      def release_ducking_expression(windows, idle, attack_fade, bridge=None):
          """FFmpeg gain expression with a fixed attack and a per-window safe release end."""
          attack = float(attack_fade)
          if bridge is None:
              bridge = default_bridge(attack)
          merged = coalesce_release_duck_windows(windows, bridge)
          if not merged:
              return None
          idle = float(idle)
          terms = []
          for start, hold_end, level, restore_at in merged:
              if attack > 0:
                  attack_term = f"({_t_minus(start - attack)})/{attack:.4f}"
              else:
                  attack_term = f"between(t,{start:.4f},{restore_at:.4f})"
              release = restore_at - hold_end
              if release > 1e-6:
                  release_term = f"({restore_at:.4f}-t)/{release:.4f}"
                  mask = f"min(1,max(0,min({attack_term},{release_term})))"
              else:
                  mask = f"between(t,{start:.4f},{hold_end:.4f})"
              terms.append(f"+({float(level) - idle:.3f})*{mask}")
          return f"max(0,min(1,{idle}{''.join(terms)}))"
      
      
      def release_ducking_keyframes(windows, idle, attack_fade, span_start, span_end, bridge=None):
          """Timeline keyframes matching `release_ducking_expression` exactly."""
          attack = float(attack_fade)
          if bridge is None:
              bridge = default_bridge(attack)
          span_start, span_end = float(span_start), float(span_end)
          normalized = [
              (max(span_start, start), min(span_end, end), gain, min(span_end, max(end, restore)))
              for start, end, gain, restore in ((float(s), float(e), float(g), float(r)) for s, e, g, r in windows)
              if end > span_start and start < span_end and end > start
          ]
          merged = coalesce_release_duck_windows(normalized, bridge)
          if not merged:
              return []
          points = [(span_start, float(idle))]
          for start, hold_end, level, restore_at in merged:
              points.extend([
                  (max(span_start, start - attack), float(idle)),
                  (start, level),
                  (hold_end, level),
                  (restore_at, float(idle)),
              ])
          points.append((span_end, float(idle)))
          points.sort(key=lambda point: point[0])
          out = []
          for when, gain in points:
              if out and abs(out[-1][0] - when) < 1e-4:
                  out[-1] = (when, min(out[-1][1], gain))
              else:
                  out.append((when, gain))
          return [_round_keyframe(when, gain) for when, gain in out]
      
      
      def variable_ducking_keyframes(windows, idle, fade, span_start, span_end, bridge=None):
          """Volume keyframes for per-window duck gains using canonical semantics."""
          fade = float(fade)
          if bridge is None:
              bridge = default_bridge(fade)
          span_start = float(span_start)
          span_end = float(span_end)
          rel = sorted(
              (max(span_start, float(w[0])), min(span_end, float(w[1])), float(w[2]))
              for w in windows
              if float(w[1]) > span_start and float(w[0]) < span_end and float(w[1]) > float(w[0])
          )
          merged = coalesce_duck_windows(rel, bridge)
          if not merged:
              return []
      
          pts = [(span_start, float(idle))]
          for s, e, level in merged:
              pts.append((max(span_start, s - fade), float(idle)))
              pts.append((s, level))
              pts.append((e, level))
              pts.append((min(span_end, e + fade), float(idle)))
          pts.append((span_end, float(idle)))
      
          pts.sort(key=lambda p: p[0])
          out = []
          for t, gain in pts:
              if out and abs(out[-1][0] - t) < 1e-4:
                  out[-1] = (t, min(out[-1][1], gain))
              else:
                  out.append((t, gain))
          return [_round_keyframe(t, gain) for t, gain in out]
      
      
      def fixed_ducking_keyframes(windows, idle, duck, fade, span_start, span_end, bridge=None):
          """Volume keyframes for fixed-gain duck windows using canonical semantics."""
          return variable_ducking_keyframes(
              [(float(s), float(e), float(duck)) for s, e in windows],
              idle,
              fade,
              span_start,
              span_end,
              bridge=bridge,
          )
      
    • audio_mix.py 18.5 KB
      """Loudness, source handoffs, ducking envelopes, and audio mix graphs."""
      
      import json
      import re
      from pathlib import Path
      
      from artifacts import _load_work_json
      from audio_automation import (
          coalesce_duck_windows,
          ducking_expression,
          release_ducking_expression,
      )
      from lib import CONFIG, filter_file_args, log, run_cmd
      
      def _limiter_filter():
          return f"alimiter=limit={CONFIG['final_limiter_peak']:.2f}:level=false"
      
      
      def _loudness_mode(measured=None):
          if not CONFIG["final_loudnorm"]:
              return "limiter_only"
          return "two_pass_linear" if measured else "equivalent"
      
      
      def final_loudnorm_filter(measured=None):
          """Final-mix loudness normalization/limiter filter from CONFIG.
      
          Ducking branches set only relative balance; this single stage owns the
          absolute output loudness so the recap is not left too quiet. When `measured`
          is supplied from a first loudnorm pass, ffmpeg runs the deterministic second
          pass; without it we still force the same target and peak limiter as a
          documented equivalent/fallback path.
          """
          if not CONFIG["final_loudnorm"]:
              return _limiter_filter()
          filt = (
              f"loudnorm=I={CONFIG['target_lufs']}"
              f":TP={CONFIG['target_true_peak']}"
              f":LRA={CONFIG['target_lra']}"
              f":linear=true"
          )
          if measured:
              for src, dst in (
                  ("input_i", "measured_I"),
                  ("input_tp", "measured_TP"),
                  ("input_lra", "measured_LRA"),
                  ("input_thresh", "measured_thresh"),
                  ("target_offset", "offset"),
              ):
                  if src in measured:
                      filt += f":{dst}={measured[src]}"
          filt += ":print_format=summary"
          return f"{filt},{_limiter_filter()}"
      
      
      def _parse_loudnorm_json(text):
          """Extract ffmpeg loudnorm JSON from stderr/stdout."""
          for match in reversed(list(re.finditer(r"\{[\s\S]*?\}", text))):
              try:
                  data = json.loads(match.group(0))
              except ValueError:
                  continue
              if {"input_i", "input_tp", "input_lra", "input_thresh", "target_offset"} <= set(data):
                  return data
          return None
      
      
      def _loudnorm_first_pass_filter():
          return (
              f"loudnorm=I={CONFIG['target_lufs']}"
              f":TP={CONFIG['target_true_peak']}"
              f":LRA={CONFIG['target_lra']}"
              f":print_format=json"
          )
      
      
      def _run_loudnorm_first_pass(input_video, narration_wav, original_audio_input,
                                   bgm_input, filter_complex, work_dir):
          """Measure the exact mixed audio graph before final render.
      
          Returns ffmpeg loudnorm JSON, or None when probing fails. The caller then
          falls back to the documented equivalent single-pass target+limiter filter.
          """
          if not CONFIG["final_loudnorm"]:
              return None
          probe_fc = f"{filter_complex};[aout]{_loudnorm_first_pass_filter()}[lnprobe]"
          probe_script = Path(work_dir) / ".filter_complex_loudnorm_probe.txt"
          probe_script.write_text(probe_fc, encoding="utf-8")
          cmd = [
              "ffmpeg", "-y",
              "-i", str(input_video),
              "-i", str(narration_wav),
              *original_audio_input,
              *bgm_input,
              *filter_file_args("filter_complex", probe_script),
              "-map", "[lnprobe]",
              "-f", "null", "-",
          ]
          try:
              result = run_cmd(cmd)
          finally:
              probe_script.unlink(missing_ok=True)
          if result.returncode != 0:
              log(f"  ⚠️ loudnorm 首遍测量失败,降级到目标滤镜+limiter: {result.stderr}")
              return None
          measured = _parse_loudnorm_json(result.stdout + "\n" + result.stderr)
          if not measured:
              log("  ⚠️ loudnorm 首遍未返回 JSON,降级到目标滤镜+limiter")
              return None
          return measured
      
      
      def _seg_place_window(seg):
          """A segment's actual placed (start, end) on the output timeline; zero-width when unplaced."""
          return seg["actual_place_start"], seg["actual_place_end"]
      
      
      def _load_sentence_handoff_anchors(work_dir):
          """Load high/medium sentence anchors and their measured pause windows."""
          work_dir = Path(work_dir)
          cut_mode = (work_dir / "edited_source.mp4").exists() or (
              work_dir / "clip_plan_validated.json"
          ).exists()
          artifact = "speech_boundary_anchors_output.json" if cut_mode else "speech_boundary_anchors.json"
          payload = _load_work_json(work_dir, artifact)
          if payload is None:
              return [], None, {"require_measured": cut_mode}
          if cut_mode:
              # Output-clock anchors are only trusted when they are at least as new as the cut plan.
              plan_path = work_dir / "clip_plan_validated.json"
              fresh = (
                  payload.get("schema_version") == 2
                  and payload.get("timeline") == "cut_output"
                  and plan_path.exists()
                  and (work_dir / artifact).stat().st_mtime_ns >= plan_path.stat().st_mtime_ns
              )
              if not fresh:
                  return [], None, {"require_measured": True}
              payload = {**payload, "require_measured": True}
          anchors = {}
          for item in payload["sentence_anchors"]:
              if item["confidence"] not in {"high", "medium"}:
                  continue
              when = float(item["time"])
              pause_start = float(item.get("pause_start", when - 0.12))
              row = {
                  "time": round(when, 4),
                  "pause_start": round(max(0.0, min(pause_start, when)), 4),
              }
              anchors[(row["time"], row["pause_start"])] = row
          return sorted(anchors.values(), key=lambda row: row["time"]), artifact, payload
      
      
      def _timed_rows(rows):
          return [{"start": float(row["start"]), "end": float(row["end"])} for row in rows]
      
      
      def _asr_segments(work_dir):
          """Cleaned ASR segments (asr_clean.json) when present, else raw asr_result.json; [] when absent."""
          clean = _load_work_json(work_dir, "asr_clean.json")
          if clean is not None:
              return clean["segments"]
          return _load_work_json(work_dir, "asr_result.json") or []
      
      
      def _handoff_speech_evidence(work_dir, payload):
          speech = _timed_rows(payload.get("speech_spans", []))
          quiet = _timed_rows(payload.get("quiet_windows", []))
          if payload.get("require_measured"):
              return speech, quiet
          if not speech:
              speech = _timed_rows(_asr_segments(work_dir))
          if not quiet:
              silence = _load_work_json(work_dir, "silence_periods.json") or []
              quiet = _timed_rows(row for row in silence if not row["has_speech"])
          return speech, quiet
      
      
      def _merged_handoff_intervals(start, end, rows):
          intervals = sorted(
              (max(start, row["start"]), min(end, row["end"]))
              for row in rows
              if row["end"] > start and row["start"] < end
          )
          merged = []
          for left, right in intervals:
              if right <= left:
                  continue
              if merged and left <= merged[-1][1]:
                  merged[-1] = (merged[-1][0], max(merged[-1][1], right))
              else:
                  merged.append((left, right))
          return merged
      
      
      def _speech_overlap_excluding_quiet(start, end, speech, quiet):
          speech_intervals = _merged_handoff_intervals(start, end, speech)
          quiet_intervals = _merged_handoff_intervals(start, end, quiet)
          overlap = sum(right - left for left, right in speech_intervals)
          for speech_left, speech_right in speech_intervals:
              overlap -= sum(
                  max(0.0, min(speech_right, quiet_right) - max(speech_left, quiet_left))
                  for quiet_left, quiet_right in quiet_intervals
              )
          return max(0.0, overlap)
      
      
      def _measured_speech_owned(
          start, end, speech, quiet, anchors, authored, require_measured=False
      ):
          duration = max(0.0, end - start)
          quiet_min = max(0.3, duration * CONFIG["quiet_overlap_min_ratio"])
          if speech:
              return _speech_overlap_excluding_quiet(start, end, speech, quiet) > 0.05
          quiet_overlap = sum(
              right - left for left, right in _merged_handoff_intervals(start, end, quiet)
          )
          if quiet and quiet_overlap >= quiet_min:
              return False
          return True if anchors or require_measured else bool(authored)
      
      
      def _entry_speech_owned(
          start, speech, quiet, anchors, authored, require_measured=False, tolerance=0.05
      ):
          if any(row["start"] - tolerance <= start <= row["end"] + tolerance for row in quiet):
              return False
          if any(row["start"] - tolerance <= start < row["end"] - tolerance for row in speech):
              return True
          if speech:
              return False
          return True if anchors or require_measured else bool(authored)
      
      
      def _work_has_source_speech(work_dir, speech_spans, require_measured):
          if speech_spans or require_measured:
              return True
          return any(item["text"].strip() for item in _asr_segments(work_dir))
      
      
      def _apply_source_sentence_handoffs(tts_segments, work_dir, video_duration):
          """Keep source audio ducked until a safe sentence boundary after narration.
      
          This does not move or trim narration. It only extends the ORIGINAL-audio duck
          envelope so returning the source track cannot reveal the middle of a sentence.
          """
          fade = CONFIG["duck_fade_seconds"]
          bridge = CONFIG["duck_bridge_seconds"]
          anchors, artifact, evidence_payload = _load_sentence_handoff_anchors(work_dir)
          speech_spans, quiet_windows = _handoff_speech_evidence(work_dir, evidence_payload)
          require_measured = evidence_payload.get("require_measured", False)
          placed = []
          for seg in tts_segments:
              start, end = _seg_place_window(seg)
              if end > start:
                  placed.append((start, end, seg))
          placed.sort(key=lambda item: (item[0], item[1]))
          if not placed:
              return []
      
          runs = []
          for start, end, seg in placed:
              if runs and start - runs[-1]["end"] <= bridge + 1e-6:
                  runs[-1]["end"] = max(runs[-1]["end"], end)
                  runs[-1]["segments"].append(seg)
              else:
                  runs.append({"start": start, "end": end, "segments": [seg]})
      
          source_has_speech = _work_has_source_speech(work_dir, speech_spans, require_measured)
          report = []
          for run in runs:
              ownership = []
              for seg in run["segments"]:
                  start, end = _seg_place_window(seg)
                  measured = _measured_speech_owned(
                      start,
                      end,
                      speech_spans,
                      quiet_windows,
                      anchors,
                      seg["overlaps_speech"],
                      require_measured=require_measured,
                  )
                  seg["overlaps_speech"] = measured
                  ownership.append(measured)
              first = run["segments"][0]
              entry_owned = _entry_speech_owned(
                  run["start"],
                  speech_spans,
                  quiet_windows,
                  anchors,
                  first["overlaps_speech"],
                  require_measured=require_measured,
              )
              speech_owned = entry_owned or any(ownership)
              if not speech_owned:
                  report.append({"start": run["start"], "end": run["end"], "status": "quiet_source"})
                  continue
              last = run["segments"][-1]
              start_safe = run["start"] <= 0.25 or any(
                  anchor["pause_start"] - 0.05 <= run["start"] <= anchor["time"] + 0.08
                  for anchor in anchors
              )
              if entry_owned and anchors and not start_safe:
                  first["source_handoff_blocking"] = True
                  first["source_entry_status"] = "unsafe_entry"
              elif not entry_owned:
                  first["source_entry_status"] = "quiet_source"
              else:
                  first["source_entry_status"] = "sentence_boundary" if anchors else "unverified"
      
              restore_anchor = next(
                  (anchor for anchor in anchors if anchor["time"] >= run["end"] - 0.01),
                  None,
              )
              if restore_anchor is not None:
                  # Hold the source low through its last spoken sample, then fit the release
                  # entirely inside the measured pause. Never begin the ramp `fade` seconds
                  # before the anchor when that would expose the final source phoneme.
                  duck_end = max(run["end"], restore_anchor["pause_start"])
                  restore_at = max(duck_end, restore_anchor["time"])
                  status = "sentence_boundary"
              elif anchors:
                  # No later complete source sentence: never expose a fragment at the tail.
                  restore_at = float(video_duration)
                  duck_end = float(video_duration)
                  status = "held_to_timeline_end"
              elif source_has_speech:
                  first["source_handoff_blocking"] = True
                  first["source_entry_status"] = "anchors_unavailable"
                  restore_at = run["end"] + fade
                  duck_end = run["end"]
                  status = "anchors_unavailable"
              else:
                  restore_at = run["end"] + fade
                  duck_end = run["end"]
                  status = "no_source_speech"
      
              last["source_duck_end"] = round(min(float(video_duration), duck_end), 4)
              last["source_restore_at"] = round(min(float(video_duration), restore_at), 4)
              last["source_handoff_status"] = status
              report.append({
                  "start": round(run["start"], 4),
                  "end": round(run["end"], 4),
                  "restore_at": last["source_restore_at"],
                  "status": status,
                  "anchor_artifact": artifact,
              })
          return report
      
      
      def _amix_tail(narr_vol, bgm_chain=""):
          """Mix the prepared original track [orig] (+ optional BGM bed) with the boosted
          narration [narr] into [aout]. bgm_chain, when given, defines [bgm] from input [2:a]."""
          narr = f"[1:a]volume={narr_vol},aresample=48000[narr];"
          if bgm_chain:
              return bgm_chain + narr + "[orig][bgm][narr]amix=inputs=3:duration=first:dropout_transition=0:normalize=0[aout]"
          return narr + "[orig][narr]amix=inputs=2:duration=first:dropout_transition=0:normalize=0[aout]"
      
      
      def _duck_envelope(tts_segments, idle, speech_vol, quiet_vol, fade, bridge):
          """Per-beat ducking automation for the ORIGINAL track.
      
          Uses the shared ducking contract: [start-fade,start] pre-roll ramp down,
          [start,end] held at the selected duck level, and [end,end+fade] release.
          Bridged spans use the most-ducked (lowest) level, matching timeline.json /
          JianYing keyframes. Returns a volume= expression, or None when no beat was
          placed (caller falls back to a constant).
          """
          windows = []
          for seg in tts_segments:
              start, narration_end = _seg_place_window(seg)
              if narration_end <= start:
                  continue
              hold_end = max(narration_end, seg.get("source_duck_end", narration_end))
              restore_at = max(hold_end, seg.get("source_restore_at", hold_end + fade))
              level = speech_vol if seg["overlaps_speech"] else quiet_vol
              windows.append((start, hold_end, level, restore_at))
          return release_ducking_expression(windows, idle, fade, bridge=bridge)
      
      
      def _bgm_envelope(tts_segments, base, duck, fade, bridge):
          """Per-beat ducking automation for the BGM track using the shared contract."""
          windows = [
              (start, end, duck)
              for start, end in map(_seg_place_window, tts_segments)
              if end > start
          ]
          return ducking_expression(coalesce_duck_windows(windows, bridge), base, fade)
      
      
      def _build_audio_filter_complex(
          tts_segments,
          has_bgm=False,
          *,
          original_audio_label="0:a",
          bgm_audio_label="2:a",
      ):
          """Compose the audio tracks into [aout], like a cut-software timeline.
      
          Tracks:
            - original (input [0:a], the video's own audio): ducked under each narration
              window by a per-beat volume envelope, but held up at `idle_orig_volume` in
              the gaps so the recap never drops to dead air between sentences.
            - bgm (input [2:a], optional): a looped music bed, gently ducked under narration.
            - narration (input [1:a]): the TTS, boosted and laid on top.
          CONFIG["ducking_mode"] (default "fixed") selects the original-track strategy:
          fixed = the gap-fill envelope above; sidechaincompress = auto-duck keyed off the
          narration; none = no ducking. Placement comes from actual_place_start/end.
          """
          ducking_mode = CONFIG["ducking_mode"]
          if ducking_mode == "sidechaincompress" and any(
              "source_duck_end" in seg and seg["source_duck_end"] > seg["actual_place_end"] + 1e-6
              for seg in tts_segments
          ):
              log("sidechaincompress 无法保持句末交接窗口,已回退 fixed ducking")
              ducking_mode = "fixed"
          narr_vol = CONFIG["ducking_narr_weight"]
          fade = CONFIG["duck_fade_seconds"]
          bridge = CONFIG["duck_bridge_seconds"]
          original_in = f"[{original_audio_label}]"
          bgm_in = f"[{bgm_audio_label}]"
      
          # BGM bed (input [2:a]): ducked under each narration window when present.
          bgm_chain = ""
          if has_bgm:
              base = CONFIG["bgm_volume"]
              bgm_expr = _bgm_envelope(tts_segments, base, CONFIG["bgm_ducking_volume"], fade, bridge)
              if bgm_expr:
                  bgm_chain = f"{bgm_in}volume='{bgm_expr}':eval=frame,aresample=48000[bgm];"
              else:
                  bgm_chain = f"{bgm_in}volume={base},aresample=48000[bgm];"
      
          if ducking_mode == "sidechaincompress":
              # The narration keys the compressor; split it so it can also be mixed in.
              head = (
                  f"{original_in}aresample=48000[o0];"
                  "[1:a]aresample=48000,asplit=2[sckey][scnarr];"
                  f"[o0][sckey]sidechaincompress="
                  f"threshold={CONFIG['ducking_threshold']}:ratio={CONFIG['ducking_ratio']}"
                  f":attack={CONFIG['ducking_attack']}:release={CONFIG['ducking_release']}"
                  f":knee=2.5:makeup={CONFIG['ducking_makeup']}:level_sc={CONFIG['ducking_level_sc']}[orig];"
              )
              narr = f"[scnarr]volume={narr_vol}[narr];"
              if bgm_chain:
                  return head + bgm_chain + narr + "[orig][bgm][narr]amix=inputs=3:duration=first:dropout_transition=0:normalize=0[aout]"
              return head + narr + "[orig][narr]amix=inputs=2:duration=first:dropout_transition=0:normalize=0[aout]"
      
          if ducking_mode == "none":
              return f"{original_in}aresample=48000[orig];" + _amix_tail(narr_vol, bgm_chain)
      
          # fixed (default): gap-fill ducking envelope on the original track.
          idle = CONFIG["idle_orig_volume"]
          speech_vol = CONFIG["speech_ducking_volume"]
          quiet_vol = CONFIG["zone_ducking_volume"]
          expr = _duck_envelope(tts_segments, idle, speech_vol, quiet_vol, fade, bridge)
          if expr:
              n_overlap = sum(1 for s in tts_segments if s["overlaps_speech"])
              n_quiet = len(tts_segments) - n_overlap
              log(f"gap-fill ducking: 间隙原声={idle}, 对白段={speech_vol}({n_overlap}), 安静段={quiet_vol}({n_quiet}), 桥接间隙<{bridge}s")
              orig = f"{original_in}volume='{expr}':eval=frame,aresample=48000[orig];"
          else:
              # No placement info at all: hold the original at a constant level.
              orig = f"{original_in}volume={CONFIG['ducking_orig_volume']},aresample=48000[orig];"
          return orig + _amix_tail(narr_vol, bgm_chain)
      
    • compose_foreground.py 14.1 KB
      #!/usr/bin/env python3
      """Compose caller-rendered RGBA pixels over a frozen H264/CFR/AAC base.
      
      This operation validates pixels and declared inputs. It does not generate or
      interpret titles, dialogue, brands, typography, or release approval.
      """
      
      import argparse
      from fractions import Fraction
      import json
      from pathlib import Path
      import subprocess
      
      from adoption.frozen_audio import probe_audio_packets, verify_adopted_audio
      from pair_media import probe_picture, validate_pair_timing
      from adoption.strict_inputs import (
          canonical_fraction, require_declared_path, require_fields, require_integer, run_logged,
          without_digests, write_json_atomic,
      )
      
      
      PATTERN = "frame_%06d.png"
      ENCODING = {
          "video_codec": "libx264", "preset": "fast", "crf": 18,
          "pixel_format": "yuv420p", "color": "bt709_tv", "audio_codec": "copy",
      }
      
      
      def _fields(value, required):
          value = without_digests(value, "compose plan")
          require_fields(value, required, "compose plan")
          return value
      
      
      def _local_file(value, label):
          _fields(value, ["path"])
          return {"path": str(require_declared_path(value, label))}
      
      
      def _sequence_paths(directory, pattern, start, end):
          if pattern != PATTERN:
              raise ValueError(f"Only the literal simple pattern {PATTERN!r} is supported")
          directory = Path(directory).resolve()
          if not directory.is_dir():
              raise FileNotFoundError(f"Sequence directory missing: {directory}")
          expected = [directory / (pattern % index) for index in range(start, end)]
          actual = sorted(directory.iterdir(), key=lambda item: item.name)
          if [item.name for item in actual] != [item.name for item in expected]:
              raise ValueError("Sequence requires exact contiguous files with no missing or extra entries")
          if not all(path.is_file() for path in expected):
              raise FileNotFoundError("Sequence frame missing or not a file")
          return directory, expected
      
      
      def _probe_image(path):
          result = subprocess.run(
              ["ffprobe", "-v", "error", "-select_streams", "v:0", "-show_streams",
               "-show_entries", "stream=codec_name,width,height,pix_fmt", "-of", "json", str(path)],
              capture_output=True, text=True, timeout=600,
          )
          if result.returncode or result.stderr.strip():
              raise ValueError(f"Image probe failed: {result.stderr.strip()}")
          streams = json.loads(result.stdout).get("streams", [])
          if len(streams) != 1 or streams[0].get("codec_name") != "png":
              raise ValueError("Foreground assets must be PNG images")
          return streams[0]
      
      
      def _validate_rgba(path, width, height):
          image = _probe_image(path)
          if (image.get("width"), image.get("height")) != (width, height):
              raise ValueError("PNG must exactly match the full output canvas")
          if image.get("pix_fmt") != "rgba":
              raise ValueError("PNG must contain an actual RGBA pixel format")
      
      
      def _sequence(value, required, *, local_count, width, height):
          value = _fields(value, required)
          directory = value["directory"]
          if not isinstance(directory, str) or not directory or "://" in directory:
              raise ValueError("Sequence requires a local directory")
          if value["pattern"] != PATTERN:
              raise ValueError(f"Only the literal simple pattern {PATTERN!r} is supported")
          resolved, paths = _sequence_paths(directory, value["pattern"], 0, local_count)
          for path in paths:
              _validate_rgba(path, width, height)
          return {**value, "directory": str(resolved), "frame_count": len(paths)}, paths
      
      
      def validate_endcard(value, *, foreground_end, total_frames, width, height):
          """Validate the exact endcard union and return normalized data and bound paths."""
          if not isinstance(value, dict) or value.get("kind") not in {"none", "still", "sequence"}:
              raise ValueError("Endcard kind must be none, still, or sequence")
          value = without_digests(value, "endcard")
          if value["kind"] == "none":
              _fields(value, ["kind"])
              if foreground_end != total_frames:
                  raise ValueError("No endcard requires foreground to cover the full frame clock")
              return {"kind": "none"}, []
          if value["kind"] == "still":
              required = ["kind", "path", "start_frame", "end_frame"]
          else:
              required = ["kind", "directory", "pattern", "start_frame", "end_frame"]
          _fields(value, required)
          end_start = require_integer(value["start_frame"], "endcard start_frame")
          end_end = require_integer(value["end_frame"], "endcard end_frame", minimum=1)
          if end_start >= end_end:
              raise ValueError("Endcard interval must be non-empty")
          if end_start != foreground_end or end_end != total_frames:
              raise ValueError("Foreground and endcard must exactly partition the base frame clock")
          if value["kind"] == "still":
              asset = _local_file({"path": value["path"]}, "endcard")
              _validate_rgba(asset["path"], width, height)
              return {**value, **asset}, [Path(asset["path"])]
          return _sequence(
              value, required, local_count=end_end - end_start, width=width, height=height,
          )
      
      
      def validate_plan(plan_path):
          plan_path = Path(plan_path).resolve()
          plan = json.loads(plan_path.read_bytes())
          _fields(plan, ["artifact", "schema_version", "base", "video", "foreground",
                         "endcard", "producer_receipt"])
          plan = {**plan, "video": _fields(plan["video"], ["fps", "width", "height", "total_frames"])}
          if plan["artifact"] != "foreground_compose_plan" or type(plan["schema_version"]) is not int \
                  or plan["schema_version"] != 1:
              raise ValueError("Unsupported foreground_compose_plan schema")
          base = _local_file(plan["base"], "base")
          receipt = _local_file(plan["producer_receipt"], "producer_receipt")
          fps = canonical_fraction(plan["video"]["fps"], "video fps")
          width = require_integer(plan["video"]["width"], "video width", minimum=1)
          height = require_integer(plan["video"]["height"], "video height", minimum=1)
          total = require_integer(plan["video"]["total_frames"], "video total_frames", minimum=1)
          _fields(plan["foreground"], ["directory", "pattern", "start_frame", "end_frame"])
          foreground_start = require_integer(plan["foreground"]["start_frame"], "foreground start_frame")
          foreground_end = require_integer(plan["foreground"]["end_frame"], "foreground end_frame", minimum=1)
          if foreground_start != 0 or foreground_end > total:
              raise ValueError("Foreground must start at frame 0 and not exceed the frame clock")
          foreground, foreground_paths = _sequence(
              plan["foreground"], ["directory", "pattern", "start_frame", "end_frame"],
              local_count=foreground_end, width=width, height=height,
          )
          normalized_endcard, endcard_paths = validate_endcard(
              plan["endcard"], foreground_end=foreground_end, total_frames=total,
              width=width, height=height,
          )
          picture = probe_picture(base["path"])
          decoder = picture["decoder"]
          required_color = {"codec_name": "h264", "pix_fmt": "yuv420p", "color_range": "tv",
                            "color_space": "bt709", "color_transfer": "bt709",
                            "color_primaries": "bt709"}
          if any(decoder.get(key) != value for key, value in required_color.items()):
              raise ValueError("Base must be H264 yuv420p with explicit BT.709 TV color metadata")
          if (decoder.get("width"), decoder.get("height"), picture["frame_count"],
                  Fraction(picture["fps"])) != (width, height, total, fps):
              raise ValueError("Declared fps/canvas/count does not match the actual base")
          audio = probe_audio_packets(base["path"], 0)
          validate_pair_timing(picture, audio)
          return {
              "plan_path": plan_path, "base": base,
              "video": {"fps": str(fps), "width": width, "height": height, "total_frames": total},
              "foreground": foreground, "foreground_paths": foreground_paths,
              "endcard": normalized_endcard, "endcard_paths": endcard_paths,
              "producer_receipt": receipt, "picture": picture, "audio": audio,
          }
      
      
      def _run_ffmpeg(command, directory):
          run_logged(command, directory, "compose", timeout=600)
      
      
      def _verify_output(path, validated):
          output_picture = probe_picture(path)
          expected = validated["picture"]
          for key in ("frame_pts", "frame_count", "fps", "duration", "start"):
              if output_picture[key] != expected[key]:
                  raise ValueError(f"Output picture {key} differs from the base")
          decoder_keys = ("codec_name", "width", "height", "pix_fmt", "sample_aspect_ratio",
                          "field_order", "color_range", "color_space", "color_transfer",
                          "color_primaries", "chroma_location")
          if {key: output_picture["decoder"].get(key) for key in decoder_keys} != \
                  {key: expected["decoder"].get(key) for key in decoder_keys}:
              raise ValueError("Output codec/canvas/color metadata differs from the base")
          proof = verify_adopted_audio(validated["base"]["path"], path, 0, 0)
          if validate_pair_timing(output_picture, proof["output"]) != \
                  validate_pair_timing(expected, validated["audio"]):
              raise ValueError("Output audio interval changed")
          result = subprocess.run(
              ["ffmpeg", "-v", "error", "-xerror", "-threads", "2", "-i", str(path),
               "-map", "0:v:0", "-map", "0:a:0", "-f", "null", "-"],
              capture_output=True, text=True, timeout=600,
          )
          if result.returncode or result.stderr.strip():
              raise ValueError("Foreground output full decode failed")
          return output_picture, proof
      
      
      def run_compose(plan_path, output_dir, *, plan_only=False):
          directory = Path(output_dir).resolve()
          directory.mkdir(parents=True, exist_ok=False)
          report_path = directory / "foreground_run.json"
          staged = directory / "foreground.rendering.mp4"
          output = directory / "foreground.mp4"
          report = {"artifact": "foreground_compose_run", "schema_version": 1,
                    "status": "PREPARING", "direct_listening": "NOT_CHECKED",
                    "normal_speed_review": "NOT_CHECKED", "release_approved": False,
                    "encoding": ENCODING}
          write_json_atomic(report_path, report)
          try:
              value = validate_plan(plan_path)
              report.update(
                  plan={"path": str(value["plan_path"])},
                  base=value["base"], video=value["video"], foreground=value["foreground"],
                  endcard=value["endcard"],
                  producer_receipt={**value["producer_receipt"],
                                    "semantic_validation": "DECLARED_NOT_CHECKED"},
              )
              if plan_only:
                  report["status"] = "PLANNED"
                  write_json_atomic(report_path, report)
                  return report
              fps = value["video"]["fps"]
              command = ["ffmpeg", "-nostdin", "-v", "error", "-n", "-copyts",
                         "-i", value["base"]["path"], "-framerate", fps, "-start_number", "0",
                         "-i", str(Path(value["foreground"]["directory"]) / PATTERN)]
              if value["endcard"]["kind"] == "none":
                  filters = (
                      "[0:v]format=rgb24[base];[1:v]format=rgba,setpts=PTS-STARTPTS[fg];"
                      "[base][fg]overlay=0:0:eof_action=pass:format=rgb,"
                      "scale=in_range=pc:out_range=tv:"
                      "in_color_matrix=bt709:out_color_matrix=bt709,format=yuv420p,"
                      f"trim=end_frame={value['video']['total_frames']}[outv]"
                  )
              elif value["endcard"]["kind"] == "still":
                  command += ["-loop", "1", "-framerate", fps, "-i", value["endcard"]["path"]]
              else:
                  command += ["-framerate", fps, "-start_number", "0", "-i",
                              str(Path(value["endcard"]["directory"]) / PATTERN)]
              if value["endcard"]["kind"] != "none":
                  end_start = value["endcard"]["start_frame"]
                  end_offset = Fraction(end_start, 1) / Fraction(fps)
                  filters = (
                      f"[0:v]format=rgb24[base];[1:v]format=rgba,setpts=PTS-STARTPTS[fg];"
                      f"[2:v]format=rgba,setpts=PTS-STARTPTS+{end_offset.numerator}/"
                      f"{end_offset.denominator}/TB[end];"
                      f"[base][fg]overlay=0:0:eof_action=pass:format=rgb:"
                      f"enable='lt(n,{end_start})'[body];"
                      f"[body][end]overlay=0:0:eof_action=repeat:format=rgb:"
                      f"enable='gte(n,{end_start})',scale=in_range=pc:out_range=tv:"
                      "in_color_matrix=bt709:out_color_matrix=bt709,format=yuv420p,"
                      f"trim=end_frame={value['video']['total_frames']}[outv]"
                  )
              command += ["-filter_complex", filters, "-map", "[outv]", "-map", "0:a:0",
                          "-c:v", ENCODING["video_codec"], "-preset", ENCODING["preset"],
                          "-crf", str(ENCODING["crf"]), "-threads", "2",
                          "-pix_fmt", ENCODING["pixel_format"],
                          "-color_range", "tv", "-colorspace", "bt709", "-color_primaries", "bt709",
                          "-color_trc", "bt709", "-c:a", "copy", "-movie_timescale",
                          str(value["audio"]["sample_rate"]), "-movflags", "+faststart",
                          str(staged)]
              _run_ffmpeg(command, directory)
              output_picture, audio_proof = _verify_output(staged, value)
              write_json_atomic(directory / "picture_identity.json", output_picture)
              write_json_atomic(directory / "adopted_audio_identity.json", audio_proof)
              report["output"] = {"path": str(output),
                                  "full_decode": "PASS", "frame_clock": "EXACT",
                                  "audio_packet_identity": "EXACT"}
              staged.rename(output)
              report["status"] = "FOREGROUND_RENDERED"
              write_json_atomic(report_path, report)
              return report
          except Exception as exc:
              staged.unlink(missing_ok=True)
              output.unlink(missing_ok=True)
              report.pop("output", None)
              report.update(status="FAILED", error=f"{type(exc).__name__}: {exc}")
              write_json_atomic(report_path, report)
              raise
      
      
      def main():
          parser = argparse.ArgumentParser(description=__doc__)
          parser.add_argument("plan")
          parser.add_argument("--output-dir", required=True)
          parser.add_argument("--plan-only", action="store_true")
          args = parser.parse_args()
          report = run_compose(args.plan, args.output_dir, plan_only=args.plan_only)
          print(json.dumps(report, ensure_ascii=False, indent=2))
      
      
      if __name__ == "__main__":
          main()
      
    • export_jianying.py 8.5 KB
      """Optional 剪映 / JianYing (CapCut) draft exporter — decoupled, stdlib + ffprobe only.
      
      Reads a backend-neutral `timeline.json` (see timeline.py) and writes a 剪映 draft
      folder (`draft_content.json` + `draft_info.json` + `draft_meta_info.json`) that the
      desktop app can open and the user can keep editing: video clips on the main track,
      the narration and BGM as their own audio tracks, the recap lines as a subtitle
      track, and the gap-fill ducking carried as native volume keyframes.
      
      The public entrypoints stay small (`us`, `build_draft`,
      `export_timeline_to_jianying`, CLI). Internally the exporter is split into schema/templates, a thin
      normalized build context, material/segment builders, track layout metadata, and
      a safe writer/bundler. This mirrors the useful schema boundaries from duo-video
      while ffmpeg remains the canonical renderer and JianYing export remains an
      optional sidecar.
      
      Schema and the draft skeleton are informed by the open-source pyJianYingDraft
      (© GuanYixuan, Apache-2.0) and capcut-mate (© Hommy, Apache-2.0). JSON protocol
      templates pinned from duo-video are vendored under its MIT license; builders
      are implemented locally and no upstream executable code, resource package,
      adapter binary, or credential is included. See ACKNOWLEDGEMENTS / 致谢.
      """
      
      import json
      import os
      import subprocess
      import tempfile
      import uuid
      
      from jianying.builders import build_timeline_track as _build_timeline_track
      from jianying.model import DraftBuildContext as _DraftBuildContext
      from jianying.schema import draft_content_skeleton as _draft_content_skeleton
      from jianying.schema import meta_info as _meta_info
      from jianying.schema import us
      from jianying.timeline_contract import normalize_timeline as _normalize_timeline
      from jianying.writer import write_draft as _write_draft
      
      __all__ = ["us", "build_draft", "export_timeline_to_jianying", "main"]
      
      
      def _default_id():
          return str(uuid.uuid4()).upper()
      
      
      def _probe_media(path):
          """Return (duration_us, width, height) via ffprobe.
      
          A probe failure raises: a draft that silently carries a zero or guessed media
          duration is worse than a failed optional export. Still images legitimately
          report no duration (0).
          """
          result = subprocess.run(
              [
                  "ffprobe",
                  "-v",
                  "error",
                  "-of",
                  "json",
                  "-show_entries",
                  "format=duration:stream=width,height,codec_type",
                  str(path),
              ],
              capture_output=True,
              text=True,
              timeout=30,
          )
          if result.returncode:
              raise RuntimeError(f"ffprobe failed for JianYing media {path}: {result.stderr.strip()}")
          data = json.loads(result.stdout)
          duration = data["format"].get("duration")
          width = height = 0
          for stream in data.get("streams", []):
              if stream.get("codec_type") == "video":
                  width, height = int(stream["width"]), int(stream["height"])
                  break
          return (us(float(duration)) if duration is not None else 0), width, height
      
      
      def build_draft(timeline, new_id=None, probe=None):
          """Build the 剪映 draft_content dict and companion meta from a timeline."""
          return _build_normalized_draft(_normalize_timeline(timeline), new_id, probe)
      
      
      def _build_normalized_draft(timeline, new_id=None, probe=None):
          """`timeline` has already passed jianying.timeline_contract.normalize_timeline."""
          new_id = new_id or _default_id
          probe = probe or _probe_media
          ctx = _DraftBuildContext.from_timeline(timeline, new_id, probe)
      
          for timeline_track in timeline["tracks"]:
              _build_timeline_track(ctx, timeline_track)
          ctx.finalize_tracks()
      
          draft_id = new_id()
          content = _draft_content_skeleton(
              draft_id,
              ctx.width,
              ctx.height,
              ctx.fps,
              ctx.total_us,
              ctx.materials,
              ctx.tracks,
          )
          meta = _meta_info(draft_id, ctx.total_us)
          return content, meta, ctx.notes
      
      
      def _generate_reversed_media(source_path, output_path):
          commands = [
              [
                  "ffmpeg",
                  "-y",
                  "-v",
                  "error",
                  "-i",
                  source_path,
                  "-vf",
                  "reverse",
                  "-af",
                  "areverse",
                  output_path,
              ],
              [
                  "ffmpeg",
                  "-y",
                  "-v",
                  "error",
                  "-i",
                  source_path,
                  "-vf",
                  "reverse",
                  "-an",
                  output_path,
              ],
          ]
          errors = []
          for command in commands:
              result = subprocess.run(command, capture_output=True, text=True)
              if result.returncode == 0 and os.path.isfile(output_path):
                  return
              errors.append(
                  (result.stderr or result.stdout or "unknown ffmpeg error").strip()
              )
          raise RuntimeError(
              f"failed to reverse JianYing source {source_path}: {'; '.join(errors)}"
          )
      
      
      def _prepare_reverse_sources(clips, temporary_dir):
          """Generate a reversed copy for each clip and record it as the clip's reverse_path."""
          generated = []
          for clip in clips:
              source_path = clip["source_path"]
              if not os.path.isfile(source_path):
                  raise ValueError(f"reverse source does not exist: {source_path}")
              output_path = os.path.join(
                  temporary_dir, f"reversed-{uuid.uuid4().hex}.mp4"
              )
              _generate_reversed_media(source_path, output_path)
              clip["reverse_path"] = output_path
              generated.append(source_path)
          return generated
      
      
      def export_timeline_to_jianying(
          timeline, out_dir, draft_name="recap", new_id=None, probe=None, bundle_media=True
      ):
          """Write a 剪映 draft folder under out_dir/draft_name. Returns (folder, notes).
      
          Referenced media is bundled by default so the draft is self-contained and
          portable. Pass bundle_media=False only when external absolute paths are
          intentionally required.
          """
          # Validate once at the boundary; everything below trusts the normalized copy.
          timeline = _normalize_timeline(timeline)
          reverse_clips = [
              clip
              for track in timeline["tracks"]
              if track["kind"] == "video"
              for clip in track["clips"]
              if clip.get("reverse") and "reverse_path" not in clip
          ]
          if reverse_clips and not bundle_media:
              raise ValueError("automatic reverse generation requires media bundling")
          if not reverse_clips:
              content, meta, notes = _build_normalized_draft(timeline, new_id=new_id, probe=probe)
              return _write_draft(
                  content,
                  meta,
                  notes,
                  out_dir,
                  draft_name,
                  bundle_media_enabled=bundle_media,
              )
      
          os.makedirs(out_dir, exist_ok=True)
          with tempfile.TemporaryDirectory(
              prefix="jianying-reverse-", dir=out_dir
          ) as temporary_dir:
              generated = _prepare_reverse_sources(reverse_clips, temporary_dir)
              content, meta, notes = _build_normalized_draft(timeline, new_id=new_id, probe=probe)
              notes.extend(f"已生成倒放素材: {source}" for source in generated)
              return _write_draft(
                  content,
                  meta,
                  notes,
                  out_dir,
                  draft_name,
                  bundle_media_enabled=True,
              )
      
      
      def main():
          import argparse
      
          ap = argparse.ArgumentParser(
              description="Export a timeline.json to a 剪映/JianYing draft folder."
          )
          ap.add_argument("timeline", help="path to timeline.json")
          ap.add_argument(
              "--out-dir", required=True, help="parent dir to create the draft folder in"
          )
          ap.add_argument("--name", default="recap", help="draft folder name")
          bundle_group = ap.add_mutually_exclusive_group()
          bundle_group.add_argument(
              "--bundle-media",
              dest="bundle_media",
              action="store_true",
              help="copy referenced media into the draft folder (default)",
          )
          bundle_group.add_argument(
              "--no-bundle-media",
              dest="bundle_media",
              action="store_false",
              help="keep external media paths instead of making a portable draft",
          )
          ap.set_defaults(bundle_media=True)
          args = ap.parse_args()
          with open(args.timeline, encoding="utf-8") as f:
              timeline = json.load(f)
          draft_dir, notes = export_timeline_to_jianying(
              timeline,
              args.out_dir,
              args.name,
              bundle_media=args.bundle_media,
          )
          for note in notes:
              print(f"  注意: {note}")
          print(
              json.dumps({"status": "exported", "draft_dir": draft_dir}, ensure_ascii=False)
          )
      
      
      if __name__ == "__main__":
          main()
      
    • lib.py 14 KB
      """Self-contained config + utilities for this skill (no cross-skill imports)."""
      import functools
      import math
      import os
      import shutil
      import subprocess
      import tempfile
      
      
      # ── 配置 ──────────────────────────────────────────────────────────────
      _EXISTING_CONFIG_REF = globals().get("CONFIG")
      
      
      def env_int(name, default, *, minimum=None):
          """Read an integer env var; an unset/empty value yields `default`, a malformed one is an error."""
          raw = os.environ.get(name, "")
          if raw == "":
              return default
          try:
              value = int(raw)
          except ValueError as exc:
              raise ValueError(f"环境变量 {name}={raw!r} 不是整数") from exc
          return value if minimum is None else max(minimum, value)
      
      
      def env_bool(name, default=False):
          """Read common boolean env var forms."""
          raw = os.environ.get(name, "")
          if raw == "":
              return default
          return raw.strip().lower() in {"1", "true", "yes", "y", "on"}
      
      
      def env_float(name, default, *, minimum=None):
          """Read a float env var; an unset/empty value yields `default`, a malformed one is an error."""
          raw = os.environ.get(name, "")
          if raw == "":
              return default
          try:
              value = float(raw)
          except ValueError as exc:
              raise ValueError(f"环境变量 {name}={raw!r} 不是数字") from exc
          if not math.isfinite(value):
              raise ValueError(f"环境变量 {name}={raw!r} 必须是有限数")
          return value if minimum is None else max(minimum, value)
      
      
      # Cross-language source: when the original audio is in a language the narration is NOT in
      # (e.g. a Japanese drama recapped in Chinese), the original speech bleeding under the narration
      # is just noise the viewer can't parse — it reads as 怪音. In that mode the original is ducked to
      # near-silent UNDER narration; it still plays full-volume in the original-audio gap blocks, where
      # a single language is fine. Explicit SPEECH_DUCKING_VOLUME / ZONE_DUCKING_VOLUME still override.
      _foreign_source_audio = env_bool("FOREIGN_SOURCE_AUDIO", False)
      _foreign_under_narration_volume = 0.05  # original volume under narration when source audio is foreign
      
      CONFIG = {
          "fade_ms": env_int("FADE_MS", 120, minimum=0),  # 每段 TTS 淡入淡出(ms);过大会让紧凑的句子一顿一顿,120ms 防爆音又不发闷
          "ducking_mode": "fixed",  # fixed | sidechaincompress | none
          "ducking_threshold": 0.15,
          "ducking_ratio": 3,
          "ducking_attack": 10,
          "ducking_release": 300,
          "ducking_level_sc": 2.0,
          "ducking_makeup": 1.2,
          "ducking_narr_weight": 1.5,
          "ducking_orig_volume": env_float("DUCKING_ORIG_VOLUME", 0.3, minimum=0.0),  # 解说时原声基准音量
          # Derived report of the FOREIGN_SOURCE_AUDIO knob this skill implements: it selects the
          # ducking volumes below. Declared so callers can see which policy is in effect.
          "foreign_source_audio": _foreign_source_audio,
          "zone_ducking_volume": env_float("ZONE_DUCKING_VOLUME",
              _foreign_under_narration_volume if _foreign_source_audio else 0.12, minimum=0.0),  # 解说时原声压低到的音量
          "idle_orig_volume": env_float("IDLE_ORIG_VOLUME", 1.0, minimum=0.0),  # 解说块之间的"原声块"音量:默认满音量(1.0),让精彩原声整段放出来,不被压低(用户要求解说成块、原声也成块)
          "duck_fade_seconds": env_float("DUCK_FADE_SECONDS", 0.3, minimum=0.0),  # 解说块/原声块切换的淡入淡出(秒),略放宽到 0.3 让满音量↔压低的过渡更顺
          "duck_bridge_seconds": env_float("DUCK_BRIDGE_SECONDS", 1.5, minimum=0.0),  # 仅把间隔小于此值的相邻解说窗口并成一段压低;超过则视为作者特意留的"原声块",原声放回满音量。默认 1.5s:解说块内部连续压低,块与块之间的留白放出满音量原声。该值只控制短间隔合并,不设定旁白/原声配额。调大→更连续铺底、原声块更少;调小→更碎
          "bgm_path": os.environ.get("BGM_PATH", "").strip(),  # 背景音乐文件(可选),留空则不加 BGM
          "source_video": os.environ.get("SOURCE_VIDEO", "").strip(),  # 剪辑模式下的原始视频(可选),用于时间线/剪映导出引用原片片段
          "source_video_explicit": False,  # 仅 assemble.py --source-video 显式传入时为 True;环境变量 SOURCE_VIDEO 不算显式
          "export_jianying": env_bool("EXPORT_JIANYING", False),  # 渲染后可选导出剪映草稿(默认关;与核心解耦)
          "jianying_draft_dir": os.environ.get("JIANYING_DRAFT_DIR", "").strip(),  # 剪映草稿输出父目录(留空=work_dir)
          "jianying_bundle_media": env_bool("JIANYING_BUNDLE_MEDIA", True),  # 默认开:macOS 剪映沙箱读不到外部路径,须把素材拷进草稿目录
          "bgm_volume": env_float("BGM_VOLUME", 0.18, minimum=0.0),  # BGM 铺底音量
          "bgm_ducking_volume": env_float("BGM_DUCKING_VOLUME", 0.10, minimum=0.0),  # 旁白时 BGM 压低到的音量
          "narration_speed": env_float("NARRATION_SPEED", 1.15, minimum=0.5),  # 解说整体提速(atempo),默认回到可懂区间;长片可设 1.0
          "narration_cumulative_tempo_max": env_float("NARRATION_CUMULATIVE_TEMPO_MAX", 1.35, minimum=1.0),  # TTS rate × 全局 atempo × 段内 atempo 的累计上限
          "narration_cumulative_tempo_hard_max": env_float("NARRATION_CUMULATIVE_TEMPO_HARD_MAX", 1.40, minimum=1.0),  # QC/阻断硬上限
          "tts_segment_tempo_max": env_float("TTS_SEGMENT_TEMPO_MAX", 1.20, minimum=1.0),  # 兼容旧段内 atempo 上限;实际会被累计预算收紧
          "mask_source_subtitles": env_bool("MASK_SOURCE_SUBTITLES", False),  # 遮挡原片烧录字幕;必须配合显式 SOURCE_SUBTITLE_MASK_POLICY
          "source_subtitle_mask_policy_declared": bool(os.environ.get("SOURCE_SUBTITLE_MASK_POLICY", "").strip()),
          "source_subtitle_mask_policy": (
              os.environ.get("SOURCE_SUBTITLE_MASK_POLICY", "").strip().lower()
              or "off"
          ),  # off | opt_in | safe | forced;MASK_SOURCE_SUBTITLES alone is legacy implicit and QC-blocking
          "source_subtitle_mask_ratio": env_float("SOURCE_SUBTITLE_MASK_RATIO", 0.14, minimum=0.0),  # 底部遮挡比例
          "source_subtitle_mask_timing": os.environ.get("SOURCE_SUBTITLE_MASK_TIMING", "narration").strip().lower(),  # all | narration;增强版默认仅解说时遮罩
          "subtitle_mask_opacity": min(1.0, env_float("SUBTITLE_MASK_OPACITY", 0.6, minimum=0.0)),  # 0=透明,1=全黑;增强版默认半透明
          "subtitle_mask_padding": env_int("SUBTITLE_MASK_PADDING", 4, minimum=0),
          "subtitle_y_top": env_int("SUBTITLE_Y_TOP", -1, minimum=-1),  # 自动旋转后的显示画布坐标;top/bot 同时有效时贴合原字幕带
          "subtitle_y_bot": env_int("SUBTITLE_Y_BOT", -1, minimum=-1),
          "narration_delay_seconds": env_float("NARRATION_DELAY_SECONDS", 0.0, minimum=0.0),  # 默认严格采用 Agent 写入的 start;旧项目可显式恢复延迟
          "narration_tighten": env_bool("NARRATION_TIGHTEN", True),  # 段落内把句子紧贴上一句实际收尾播放,句间间隔稳定≤tight_pause,杜绝"一句解说一段空白"的卡顿
          "narration_run_gap_seconds": env_float("NARRATION_RUN_GAP_SECONDS", 1.6, minimum=0.0),  # 作者留白超过此值=新段落(让精彩原声透出);小于则视为同一连续段落
          "narration_tight_pause_seconds": env_float("NARRATION_TIGHT_PAUSE_SECONDS", 0.35, minimum=0.0),  # 段落内句间固定间隔(秒)
          "narration_max_pull_seconds": env_float("NARRATION_MAX_PULL_SECONDS", 1.2, minimum=0.0),  # 收紧时一句最多比作者标注提前的秒数(漂移上限,越小越贴画面)
          "narration_tail_pad_seconds": 0.1,  # 解说尾部最少留白;短 slot 会自动压低 delay 避免截断
          "quiet_overlap_min_ratio": 0.8,  # 解说段至少多少比例落在安静窗口内才标记为非对白重叠
          "speech_ducking_volume": env_float("SPEECH_DUCKING_VOLUME",
              _foreign_under_narration_volume if _foreign_source_audio else 0.2, minimum=0.0),    # 解说与对白重叠时原声音量
          "burn_subtitles": env_bool("BURN_SUBTITLES", True),  # 烧录解说字幕(默认开;遮挡原字幕后需自带字幕,否则字幕区空白)
          "subtitle_original_in_gaps": env_bool("SUBTITLE_ORIGINAL_IN_GAPS", True),  # 原声留白处补烧原声台词字幕(来自 ASR)
          "force_video_reencode": env_bool("FORCE_VIDEO_REENCODE", False),  # 组装时重编码视频,修复部分容器时间戳问题
          # 成片压制(仅在重编码时生效:烧字幕/遮罩/缩放/FORCE_VIDEO_REENCODE 任一触发重编码)。
          "output_crf": env_int("OUTPUT_CRF", 18, minimum=0),          # x264 CRF;越大文件越小、画质越低(18≈视觉无损,23~26 体积更小)
          "output_preset": os.environ.get("OUTPUT_PRESET", "veryfast"),  # x264 preset;slow/slower 同 CRF 下体积更小但更慢
          "output_max_height": env_int("OUTPUT_MAX_HEIGHT", 0, minimum=0),  # >0 时把成片高度上限缩到该值(保持宽高比、偶数宽);0=不缩放
          # 成片末端整体响度归一(默认混音偏轻,归一后更接近常见短视频响度;样片约 -11.9,默认取更安全的 -14)
          "final_loudnorm": env_bool("FINAL_LOUDNORM", True),  # 组装末端做一次整体响度归一
          "target_lufs": env_float("TARGET_LUFS", -14.0),       # 目标综合响度 (LUFS)
          "target_true_peak": env_float("TARGET_TRUE_PEAK", -1.0),  # 目标真峰值 (dBTP)
          "target_lra": env_float("TARGET_LRA", 11.0),          # 目标响度范围 (LU)
          "final_limiter_peak": env_float("FINAL_LIMITER_PEAK", 0.98, minimum=0.1),  # loudnorm 后峰值保护 limiter
          "subtitle_font_name": os.environ.get("SUBTITLE_FONT_NAME", "Arial"),
          # 可选字体文件:ASS 烧录经 fontsdir 加载,画面文字经 drawtext fontfile 使用;family 名仍由 SUBTITLE_FONT_NAME 指定
          "subtitle_font_file": os.environ.get("SUBTITLE_FONT_FILE", "").strip(),
          "subtitle_font_size": env_int("SUBTITLE_FONT_SIZE", 42, minimum=8),
          "subtitle_primary_color": os.environ.get("SUBTITLE_PRIMARY_COLOR", "&H00FFFFFF"),
          "subtitle_outline_color": os.environ.get("SUBTITLE_OUTLINE_COLOR", "&H00000000"),
          "subtitle_outline": env_float("SUBTITLE_OUTLINE", 2.0, minimum=0.0),
          "subtitle_shadow": env_float("SUBTITLE_SHADOW", 1.0, minimum=0.0),
          "subtitle_margin_v": env_int("SUBTITLE_MARGIN_V", 48, minimum=0),
          "subtitle_margin_l": env_int("SUBTITLE_MARGIN_L", 40, minimum=0),
          "subtitle_margin_r": env_int("SUBTITLE_MARGIN_R", 40, minimum=0),
          "subtitle_alignment": env_int("SUBTITLE_ALIGNMENT", 2, minimum=1),
          "subtitle_max_chars": env_int("SUBTITLE_MAX_CHARS", 20, minimum=6),
          "subtitle_max_lines": env_int("SUBTITLE_MAX_LINES", 2, minimum=1),
          "subtitle_play_res_x": env_int("SUBTITLE_PLAY_RES_X", 1280, minimum=1),
          "subtitle_play_res_y": env_int("SUBTITLE_PLAY_RES_Y", 720, minimum=1),
      }
      if isinstance(_EXISTING_CONFIG_REF, dict):
          _EXISTING_CONFIG_REF.clear()
          _EXISTING_CONFIG_REF.update(CONFIG)
          CONFIG = _EXISTING_CONFIG_REF
      
      def narration_tempo_budget(tts_rate_offset=0.0):
          """Return the canonical tempo budget shared by voiceover and assemble.
      
          `effective_tempo` is the user-perceived cumulative compression:
          TTS rate × global narration atempo × per-segment atempo.  The segment atempo
          cap is therefore tightened by the configured global speed and TTS rate
          offset; callers must fail/shorten instead of time-trimming speech when the
          needed ratio exceeds `segment_tempo_max`.
          """
          global_speed = CONFIG["narration_speed"]
          rate_factor = max(0.01, 1.0 + float(tts_rate_offset))
          cumulative_max = CONFIG["narration_cumulative_tempo_max"]
          hard_max = max(cumulative_max, CONFIG["narration_cumulative_tempo_hard_max"])
          segment_tempo_max = max(1.0, min(
              CONFIG["tts_segment_tempo_max"], cumulative_max / (global_speed * rate_factor)
          ))
          return {
              "global_narration_speed": global_speed,
              "tts_rate_factor": rate_factor,
              "cumulative_tempo_max": cumulative_max,
              "cumulative_tempo_hard_max": hard_max,
              "segment_tempo_max": segment_tempo_max,
              "max_raw_duration_factor": global_speed * segment_tempo_max,
          }
      
      def log(msg):
          print(f"[video-recap] {msg}", flush=True)
      
      def run_cmd(cmd, **kwargs):
          """运行命令,返回 CompletedProcess"""
          display = " ".join(
              text if len(text) <= 240 else text[:237] + "..."
              for text in map(str, cmd)
          )
          log(f"运行: {display}")
          return subprocess.run(cmd, capture_output=True, text=True, **kwargs)
      
      
      # ffmpeg 7 added `-/option path` to read any option's value from a file; ffmpeg 9 removed the
      # older `-filter_complex_script` / `-filter_script` spellings, which are all ffmpeg <= 6 knows.
      _LEGACY_FILTER_FILE_OPTIONS = {
          "filter_complex": "-filter_complex_script",
          "filter:v:0": "-filter_script:v:0",
      }
      
      
      @functools.lru_cache(maxsize=None)
      def _ffmpeg_reads_option_files():
          """Whether the ffmpeg on PATH accepts `-/option path` (asked once per process)."""
          if shutil.which("ffmpeg") is None:
              return False
          with tempfile.TemporaryDirectory() as tmp:
              graph = os.path.join(tmp, "probe_filter.txt")
              with open(graph, "w", encoding="utf-8") as fh:
                  fh.write("null")
              result = subprocess.run(["ffmpeg", "-hide_banner", "-/filter_complex", graph],
                                      stdin=subprocess.DEVNULL, capture_output=True, text=True,
                                      timeout=20)
          return "Unrecognized option" not in result.stderr
      
      
      def filter_file_args(option, path):
          """ffmpeg arguments that load `option`'s filtergraph from `path`, spelled for this ffmpeg."""
          if _ffmpeg_reads_option_files():
              return [f"-/{option}", str(path)]
          return [_LEGACY_FILTER_FILE_OPTIONS[option], str(path)]
      
      
      def get_video_duration(video_path):
          """获取视频时长(秒)"""
          cmd = ["ffprobe", "-v", "quiet", "-show_entries", "format=duration",
                 "-of", "csv=p=0", str(video_path)]
          result = run_cmd(cmd)
          if result.returncode != 0:
              raise RuntimeError(f"ffprobe 无法读取时长 {video_path}: {result.stderr}")
          return float(result.stdout.strip())
      
    • media.py 8.2 KB
      """Media probing and source-clip provenance for video-assemble."""
      
      import json
      import os
      from pathlib import Path
      
      from artifacts import _explicit_source_video
      from lib import log, run_cmd
      
      def _load_cut_timeline_plan(work_dir):
          """The cut plan, preferring clip_plan_validated.json unless the raw plan is newer; None in full mode."""
          raw_plan_path = Path(work_dir) / "clip_plan.json"
          validated_plan_path = Path(work_dir) / "clip_plan_validated.json"
          if not validated_plan_path.exists():
              return json.loads(raw_plan_path.read_text(encoding="utf-8")) if raw_plan_path.exists() else None
          if not raw_plan_path.exists():
              return json.loads(validated_plan_path.read_text(encoding="utf-8"))
          if validated_plan_path.stat().st_mtime_ns >= raw_plan_path.stat().st_mtime_ns:
              return json.loads(validated_plan_path.read_text(encoding="utf-8"))
          return json.loads(raw_plan_path.read_text(encoding="utf-8"))
      
      
      def _plan_clip_spans(work_dir):
          """Cut-mode clip spans [{source_start, source_end, output_start, output_end, entry}], or None.
      
          clip_plan.json is either a bare list or {"clips": [...]}; a clip names its source range as
          source_start/source_end or start/end. Clips without explicit output_start/output_end are laid
          out back to back on the output timeline.
          """
          plan = _load_cut_timeline_plan(work_dir)
          if plan is None:
              return None
          entries = plan["clips"] if isinstance(plan, dict) else plan
          spans, cursor = [], 0.0
          for entry in entries:
              ss = float(entry.get("source_start", entry.get("start")))
              se = float(entry.get("source_end", entry.get("end")))
              if "output_start" in entry:
                  out_s, out_e = float(entry["output_start"]), float(entry["output_end"])
                  cursor = max(cursor, out_e)
              else:
                  out_s, out_e = cursor, cursor + (se - ss)
                  cursor = out_e
              spans.append({
                  "source_start": ss, "source_end": se,
                  "output_start": out_s, "output_end": out_e,
                  "entry": entry,
              })
          return spans
      
      
      def _ratio_to_float(value, default=1.0):
          """Parse an ffprobe ratio ("4:3", "16/9" or a bare number); unknown ratios yield `default`."""
          value = value.strip()
          if value in {"", "0:1", "0/1", "N/A"}:
              return default
          if ":" in value:
              num, den = value.split(":", 1)
          elif "/" in value:
              num, den = value.split("/", 1)
          else:
              return float(value)
          return float(num) / float(den) if float(den) else default
      
      
      def _fps_from_rate(value, default=30.0):
          """Parse an ffprobe frame rate ("30000/1001" or a bare number); a 0/0 rate yields `default`."""
          if "/" in value:
              num, den = value.split("/", 1)
              return round(float(num) / float(den), 3) if float(den) else default
          return round(float(value), 3)
      
      
      def _stream_rotation(stream):
          """Extract rotation from tags or side_data_list in ffprobe JSON."""
          for source in (stream.get("tags", {}).get("rotate"), stream.get("rotation")):
              if source not in (None, ""):
                  return int(round(float(source))) % 360
          for item in stream.get("side_data_list", []):
              if item.get("rotation") not in (None, ""):
                  return int(round(float(item["rotation"]))) % 360
          return 0
      
      
      def _canvas_from_stream(stream):
          storage_w = stream["width"]
          storage_h = stream["height"]
          fps = _fps_from_rate(stream["r_frame_rate"])
          sar_text = stream.get("sample_aspect_ratio", "1:1")
          dar_text = stream.get("display_aspect_ratio", "")
          sar = _ratio_to_float(sar_text, 1.0)
          rotation = _stream_rotation(stream)
      
          display_w = max(1, int(round(storage_w * sar)))
          display_h = max(1, storage_h)
          if dar_text and dar_text not in {"0:1", "N/A"}:
              dar = _ratio_to_float(dar_text, 0.0)
              # ffprobe sources are not consistent: some report DAR before rotation
              # (landscape value > 1 for a 90° stream), while some containers report the
              # already-rotated portrait DAR (< 1). Only apply DAR before swapping when it
              # describes the stored orientation.
              if dar > 0 and not (rotation in {90, 270} and dar < 1.0):
                  # Preserve height and adjust width. This keeps legacy square-pixel landscape
                  # byte-identical while honoring non-square pixel DAR metadata.
                  display_w = max(1, int(round(display_h * dar)))
          if rotation in {90, 270}:
              display_w, display_h = display_h, display_w
      
          return {
              "width": display_w,
              "height": display_h,
              "fps": fps,
              "storage_width": storage_w,
              "storage_height": storage_h,
              "rotation": rotation,
              "sample_aspect_ratio": sar_text,
              "display_aspect_ratio": dar_text or f"{display_w}:{display_h}",
          }
      
      
      def _probe_canvas(video_path):
          """Return rotation/SAR/DAR-aware canvas facts for a video.
      
          ``width``/``height`` are the display canvas used by subtitle/overlay geometry.
          For legacy square-pixel landscape sources, these remain the raw storage dimensions.
          """
          res = run_cmd([
              "ffprobe", "-v", "error", "-select_streams", "v:0",
              "-show_entries",
              "stream=width,height,r_frame_rate,avg_frame_rate,sample_aspect_ratio,display_aspect_ratio:stream_tags=rotate:stream_side_data=rotation",
              "-of", "json", str(video_path),
          ])
          if res.returncode != 0:
              raise RuntimeError(f"ffprobe 无法读取视频流 {video_path}: {res.stderr}")
          streams = json.loads(res.stdout)["streams"]
          if not streams:
              raise RuntimeError(f"{video_path} 没有视频流")
          return _canvas_from_stream(streams[0])
      
      
      def _has_audio_stream(video_path):
          """Return True when the input has an audio stream usable as [0:a]."""
          result = run_cmd([
              "ffprobe", "-v", "error", "-select_streams", "a:0",
              "-show_entries", "stream=index", "-of", "csv=p=0", str(video_path),
          ])
          return result.returncode == 0 and bool(result.stdout.strip())
      
      
      def _build_video_clips(input_video, work_dir, duration_s):
          """Video-track clips for the timeline.
      
          In cut mode each plan entry becomes a clip referencing the ORIGINAL source
          range. Multi-source validated plans carry per-clip source_path and do not
          require an explicit ambient --source-video. Without any declared source (full
          mode, or cut mode rendered without --source-video) the rendered input is one clip.
          """
          explicit_source_video = _explicit_source_video()
          spans = _plan_clip_spans(work_dir)
          multi_source = spans is not None and any(span["entry"].get("source_path") for span in spans)
          if spans is None or not (explicit_source_video or multi_source):
              return [{"source_path": str(input_video), "source_start": 0.0,
                       "source_end": float(duration_s), "timeline_start": 0.0,
                       "timeline_end": float(duration_s)}]
          clips = []
          for span in spans:
              entry = span["entry"]
              source_path = entry.get("source_path") or explicit_source_video
              timeline_start, timeline_end = span["output_start"], span["output_end"]
              if not source_path or not os.path.exists(source_path):
                  # Degrade ONLY this clip — point it at the rendered cut for its own output
                  # window — and keep real provenance for every present source, instead of
                  # collapsing the whole multi-source timeline.
                  log(f"  时间线: source_path 不存在,该片段降级为剪后成片片段: {source_path or '(unset)'}")
                  clips.append({"source_id": entry.get("source_id"),
                                "source_path": str(input_video),
                                "source_start": timeline_start,
                                "source_end": timeline_end,
                                "timeline_start": timeline_start,
                                "timeline_end": timeline_end,
                                "provenance_degraded": True,
                                "provenance_reason": f"missing_source_path:{source_path or 'unset'}"})
                  continue
              clips.append({"source_id": entry.get("source_id"),
                            "source_path": source_path,
                            "source_start": span["source_start"],
                            "source_end": span["source_end"],
                            "timeline_start": timeline_start,
                            "timeline_end": timeline_end})
          return clips
      
    • narration_audio.py 18.9 KB
      """Narration tempo fitting and sample-accurate timeline WAV placement."""
      
      import os
      import wave
      from pathlib import Path
      
      from lib import CONFIG, get_video_duration, log, narration_tempo_budget, run_cmd
      
      def _apply_narration_speed(
          tts_segments,
          work_dir,
          *,
          tempo_policy=None,
      ):
          """Globally speed up narration audio via atempo (CONFIG['narration_speed']).
      
          MiMo TTS reads a touch slowly for short-form recaps; a 1.1-1.2x bump makes it
          snappier without the chipmunk effect. Rewrites each segment's audio_path/duration
          to the sped copy so the rest of assembly is unchanged. No-op at speed 1.0.
          """
          speed = tempo_policy["global_atempo"] if tempo_policy else CONFIG["narration_speed"]
          if abs(speed - 1.0) <= 1e-3:
              return
          done = 0
          for seg in tts_segments:
              src = seg["audio_path"]
              if not os.path.exists(src):
                  continue  # reported as a skipped segment during placement
              out = str(Path(work_dir) / f"_spd_{seg['index']}.wav")
              res = run_cmd(["ffmpeg", "-y", "-i", src, "-filter:a", f"atempo={speed:.3f}",
                                    "-ar", "44100", "-ac", "1", "-acodec", "pcm_s16le", out])
              if res.returncode != 0:
                  raise RuntimeError(f"解说提速失败 {src}: {res.stderr}")
              seg["audio_path"] = out
              seg["narration_conversion_path"] = out
              seg["audio_duration"] = get_video_duration(out)
              done += 1
          log(f"解说整体提速: atempo={speed:.2f} ({done} 段)")
      
      
      def _adjust_tts_speed(
          audio_path,
          target_duration,
          tts_rate_offset=0.0,
          *,
          tempo_policy=None,
      ):
          """Fit overlong TTS with bounded atempo; never time-trim speech in assemble.
      
          Assemble has no word/sentence timestamps, so if bounded atempo cannot make the
          audio fit, it returns `fit_status=no_safe_fit` and leaves the original audio
          untouched for QC to block instead of guessing a spoken_text truncation.
          """
          audio_path = Path(audio_path)
          current_dur = get_video_duration(audio_path)
          budget = narration_tempo_budget(tts_rate_offset)
          if tempo_policy:
              budget.update({
                  "global_narration_speed": tempo_policy["global_atempo"],
                  "tts_rate_factor": 1.0,
                  "segment_tempo_max": tempo_policy["segment_tempo_max"],
                  "cumulative_tempo_max": tempo_policy["cumulative_tempo_max"],
                  "cumulative_tempo_hard_max": tempo_policy["cumulative_tempo_hard_max"],
              })
          meta = {
              "fit_status": "fits",
              "blocking": False,
              "tempo_factor": 1.0,
              "segment_tempo_factor": 1.0,
              "truncated": False,
              "truncate_reason": "none",
              "tts_rate_offset": float(tts_rate_offset),
              "audio_duration": current_dur,
              "placed_audio_duration": current_dur,
              "global_narration_speed": budget["global_narration_speed"],
              "effective_tempo": budget["global_narration_speed"] * budget["tts_rate_factor"],
              "cumulative_tempo_max": budget["cumulative_tempo_max"],
              "cumulative_tempo_hard_max": budget["cumulative_tempo_hard_max"],
          }
          if current_dur <= target_duration:
              return (str(audio_path), current_dur, meta)
      
          if tempo_policy and not tempo_policy["bounded_segment_fit"]:
              meta.update({
                  "fit_status": "no_safe_fit", "blocking": True,
                  "truncate_reason": "no_safe_boundary", "placed_audio_duration": 0.0,
                  "needed_tempo_factor": current_dur / target_duration,
              })
              return (str(audio_path), current_dur, meta)
      
          ratio = current_dur / target_duration
          effective_max = budget["segment_tempo_max"]
          if ratio > effective_max:
              meta.update({
                  "fit_status": "no_safe_fit",
                  "blocking": True,
                  "truncate_reason": "no_safe_boundary",
                  "placed_audio_duration": 0.0,
                  "needed_tempo_factor": ratio,
              })
              log(
                  f"  TTS 无安全放置: {current_dur:.1f}s 需 x{ratio:.2f},"
                  f"超过段内预算 x{effective_max:.2f}(assemble 不按时间硬切)"
              )
              return (str(audio_path), current_dur, meta)
      
          # 温和加速。给 atempo/容器时长舍入留出 0.2% 安全余量;宁可极轻微
          # 多加速,也不能在写入时间线时裁掉最后一个音节。
          tempo = min(ratio * 1.002, effective_max)
          adjusted_path = audio_path.with_name(f"{audio_path.stem}_adj{audio_path.suffix}")
          cmd = ["ffmpeg", "-y", "-i", str(audio_path),
                 "-filter:a", f"atempo={tempo:.6f}",
                 "-ar", "44100", "-ac", "1", str(adjusted_path)]
          result = run_cmd(cmd)
          if result.returncode != 0:
              raise RuntimeError(f"TTS 加速失败 {audio_path}: {result.stderr}")
          new_dur = get_video_duration(adjusted_path)
          if new_dur > target_duration + (1.0 / 44100.0):
              adjusted_path.unlink(missing_ok=True)
              meta.update({
                  "fit_status": "no_safe_fit",
                  "blocking": True,
                  "truncate_reason": "no_safe_boundary",
                  "placed_audio_duration": 0.0,
                  "needed_tempo_factor": new_dur / target_duration,
              })
              log(
                  f"  TTS 加速后仍超出安全窗口 {new_dur - target_duration:.3f}s;"
                  "禁止裁尾,交由 Agent 缩短/移动文本"
              )
              return (str(audio_path), current_dur, meta)
          meta.update({
              "fit_status": "tempo_adjusted",
              "tempo_factor": tempo,
              "segment_tempo_factor": tempo,
              "placed_audio_duration": new_dur,
              "effective_tempo": budget["global_narration_speed"] * budget["tts_rate_factor"] * tempo,
          })
          log(f"  TTS 温和加速: {current_dur:.1f}s → {new_dur:.1f}s (x{tempo:.2f})")
          return (str(adjusted_path), new_dur, meta)
      
      
      def _edge_quiet_samples(pcm16_mono, sample_count, *, from_start, threshold=260):
          """Count near-silent PCM16 samples at one edge of a mono buffer."""
          indices = range(sample_count) if from_start else range(sample_count - 1, -1, -1)
          quiet = 0
          for index in indices:
              offset = index * 2
              value = int.from_bytes(pcm16_mono[offset:offset + 2], "little", signed=True)
              if abs(value) > threshold:
                  break
              quiet += 1
          return quiet
      
      
      def _speech_safe_fade_lengths(pcm16_mono, sample_count, sample_rate, configured_ms):
          """Limit fades to edge silence so first/last syllables are never attenuated.
      
          When TTS has no measurable edge silence, retain only a 5ms anti-click ramp.
          """
          configured = min(int(max(0.0, float(configured_ms)) * sample_rate / 1000), sample_count // 4)
          if configured <= 0:
              return 0, 0
          anti_click = min(int(0.005 * sample_rate), configured)
          leading = _edge_quiet_samples(pcm16_mono, sample_count, from_start=True)
          trailing = _edge_quiet_samples(pcm16_mono, sample_count, from_start=False)
          return min(configured, max(anti_click, leading)), min(configured, max(anti_click, trailing))
      
      
      def _unplaced(seg, at, fit_status, reason, *, blocking=False):
          """Record a segment that contributes no audio: a zero-width window at `at` seconds."""
          seg["actual_place_start"] = at
          seg["actual_place_end"] = at
          seg["placed_audio_duration"] = 0.0
          seg["fit_status"] = fit_status
          seg["truncate_reason"] = reason
          seg["blocking"] = blocking
      
      
      def _build_timed_narration(
          tts_segments,
          output_wav,
          video_duration,
          work_dir,
          *,
          tempo_policy=None,
      ):
          """将 TTS 片段按时间轴放置到一条与视频等长的音轨上"""
          sample_rate = 44100
          total_samples = int(video_duration * sample_rate)
          buffer = bytearray(total_samples * 2)
          last_written_end = 0  # 追踪已写入位置,防止重叠
          prev_pause_samples = 0  # 前一段的 pause_after_ms,控制段间间隔
          skipped_count = 0  # 因 WAV 缺失或无安全放置而被跳过的段数
          placed_count = 0  # 真正写入音频的段数;防止"成功"生成全静音旁白
          no_safe_fit_count = 0  # 超预算但不能安全截断;交由 QC/manifest 阻断
          prev_authored_end = None  # 上一段作者标注的结束时间,用于判断"段落"边界
          run_gap = CONFIG["narration_run_gap_seconds"]   # 作者留白 > 此值 = 新段落
          tighten = CONFIG["narration_tighten"]
          tight_pause_samples = int(CONFIG["narration_tight_pause_seconds"] * sample_rate)
          # 漂移上限:收紧时一句最多比作者标注的时间提前 max_pull 秒,避免整段解说被全部压到前面、与画面脱节
          max_pull_samples = int(CONFIG["narration_max_pull_seconds"] * sample_rate)
          configured_delay = CONFIG["narration_delay_seconds"]
          tail_pad = CONFIG["narration_tail_pad_seconds"]
      
          for seg in tts_segments:
              wav_path = seg["audio_path"]
              pause_samples = int(seg["pause_after_ms"] * sample_rate / 1000)
              # 段落收紧:同一段落内(与上一句作者留白 <= run_gap)把这一句紧贴上一句的实际收尾播放,
              # 句间间隔固定为 tight_pause,不受 slot 内居中延迟 / TTS 时长波动影响。段落之间(作者特意留
              # 的大留白,让精彩原声透出)才放回原声。这样句间间隔稳定、不会出现"一句解说一段空白"。
              cur_authored_start = float(seg["start"])
              is_run_start = (placed_count == 0 or prev_authored_end is None
                              or cur_authored_start - prev_authored_end > run_gap)
              prev_authored_end = float(seg["end"])
      
              if not os.path.exists(wav_path):
                  _unplaced(seg, seg["start"], "skipped", "missing_wav")
                  prev_pause_samples = pause_samples
                  skipped_count += 1
                  continue
      
              original_wav_path = wav_path
              try:
                  with wave.open(wav_path, "rb") as wf_check:
                      needs_resample = (
                          wf_check.getnchannels(), wf_check.getsampwidth(), wf_check.getframerate()
                      ) != (1, 2, sample_rate)
              except (wave.Error, EOFError):
                  # Valid post-processed WAV may use IEEE float, which Python's wave
                  # reader does not support. FFmpeg performs the explicit PCM conversion.
                  needs_resample = True
      
              tts_rate_offset = seg["tts_rate_offset"]
              tts_dur = seg["audio_duration"]
      
              slot_duration = max(0.0, float(seg["end"]) - float(seg["start"]))
              max_delay = max(0.0, slot_duration - tts_dur - tail_pad)
              narration_delay = min(configured_delay, max_delay)
              start_sample = int((seg["start"] + narration_delay) * sample_rate)
              end_boundary = int(min(seg["end"], video_duration) * sample_rate)
      
              # 段间间隔:使用前一段的 pause_after_ms(来自 narration.json)
              min_start_with_pause = last_written_end + prev_pause_samples
              if tighten and not is_run_start:
                  # 段落内:紧贴上一句的实际收尾播放,句间间隔固定为 tight_pause(不被 slot 内居中延迟撑大),
                  # 但不早于"作者标注起始 - max_pull",防止整段被压到前面与画面脱节。
                  drift_floor = int(cur_authored_start * sample_rate) - max_pull_samples
                  actual_start = max(last_written_end + tight_pause_samples, drift_floor)
              else:
                  # 段落起点(或关闭收紧):尊重作者标注的起始 + 入场延迟,让画面/原声先立住
                  actual_start = max(start_sample, min_start_with_pause)
              actual_start = min(actual_start, end_boundary)  # 不超出 slot 边界
      
              # 根据实际可用空间决定是否加速
              available_samples = end_boundary - actual_start
              available_duration = max(available_samples / sample_rate, 0)
              if tts_dur > available_duration > 0:
                  if tempo_policy:
                      wav_path, _actual_dur, fit_meta = _adjust_tts_speed(
                          wav_path, available_duration, tts_rate_offset,
                          tempo_policy=tempo_policy,
                      )
                  else:
                      wav_path, _actual_dur, fit_meta = _adjust_tts_speed(
                          wav_path, available_duration, tts_rate_offset
                      )
                  seg.update({
                      "fit_status": fit_meta["fit_status"],
                      "segment_tempo_factor": fit_meta["segment_tempo_factor"],
                      "effective_tempo": fit_meta["effective_tempo"],
                      "global_narration_speed": fit_meta["global_narration_speed"],
                      "blocking": fit_meta["blocking"],
                  })
                  if fit_meta["fit_status"] == "no_safe_fit":
                      _unplaced(seg, actual_start / sample_rate, "no_safe_fit",
                                fit_meta["truncate_reason"], blocking=True)
                      prev_pause_samples = pause_samples
                      skipped_count += 1
                      no_safe_fit_count += 1
                      continue
              else:
                  budget = narration_tempo_budget(tts_rate_offset)
                  if tempo_policy:
                      budget.update({
                          "global_narration_speed": tempo_policy["global_atempo"],
                          "tts_rate_factor": 1.0,
                      })
                  seg.update({
                      "fit_status": "fits",
                      "segment_tempo_factor": 1.0,
                      "global_narration_speed": budget["global_narration_speed"],
                      "effective_tempo": budget["global_narration_speed"] * budget["tts_rate_factor"],
                      "blocking": False,
                  })
      
              # _adjust_tts_speed 输出固定 44100Hz mono 16bit,若文件被替换则无需 resample
              if wav_path != original_wav_path:
                  needs_resample = False
              if needs_resample:
                  tmp_path = str(Path(work_dir) / f"_rs_{seg['index']}.wav")
                  rs_result = run_cmd(["ffmpeg", "-y", "-i", wav_path,
                                              "-ar", str(sample_rate), "-ac", "1",
                                              "-acodec", "pcm_s16le", tmp_path])
                  if rs_result.returncode != 0:
                      log(f"  跳过: 重采样失败 {wav_path}: {rs_result.stderr}")
                      _unplaced(seg, seg["start"], "skipped", "resample_failed")
                      prev_pause_samples = pause_samples
                      skipped_count += 1
                      continue
                  wav_path = tmp_path
                  seg["narration_conversion_path"] = tmp_path
      
              with wave.open(wav_path, "rb") as wf:
                  wf_data = bytearray(wf.readframes(wf.getnframes()))
      
              # 按场景边界裁剪
              audio_samples = len(wf_data) // 2
              available = end_boundary - actual_start
              write_samples = audio_samples
      
              if write_samples <= 0 or available <= 0:
                  log(f"  跳过: {seg['start']:.1f}s-{seg['end']:.1f}s (无空间)")
                  _unplaced(seg, seg["start"], "no_safe_fit", "no_room", blocking=True)
                  prev_pause_samples = pause_samples
                  no_safe_fit_count += 1
                  continue
      
              if audio_samples > available:
                  # No tolerance-based trimming: even a few milliseconds may contain a
                  # consonant/vowel release. _adjust_tts_speed must produce a complete file
                  # that fits; otherwise block and ask the Agent to shorten/move the block.
                  over = (audio_samples - available) / sample_rate
                  log(f"  TTS 无安全放置: 段 {seg['index']} 超出可用窗口 {over:.3f}s;禁止裁尾,交由 QC 阻断")
                  _unplaced(seg, actual_start / sample_rate, "no_safe_fit", "no_safe_boundary", blocking=True)
                  prev_pause_samples = pause_samples
                  skipped_count += 1
                  no_safe_fit_count += 1
                  continue
      
              # 重叠检测:跳过与前段重叠的部分(在 fade 之前,避免截断后丢失 fade-in)
              if actual_start < last_written_end:
                  overlap_ms = (last_written_end - actual_start) * 1000 / sample_rate
                  if last_written_end >= actual_start + write_samples:
                      log(f"  跳过重叠段: {actual_start/sample_rate:.1f}s "
                             f"(与前段重叠 {overlap_ms:.0f}ms)")
                      _unplaced(seg, seg["start"], "no_safe_fit", "no_room", blocking=True)
                      prev_pause_samples = pause_samples
                      no_safe_fit_count += 1
                      continue
                  actual_start = last_written_end
                  available = end_boundary - actual_start
                  if write_samples > available:
                      log(f"  重叠 {overlap_ms:.0f}ms 后无安全完整窗口,跳过")
                      _unplaced(seg, actual_start / sample_rate, "no_safe_fit", "no_safe_boundary", blocking=True)
                      prev_pause_samples = pause_samples
                      skipped_count += 1
                      no_safe_fit_count += 1
                      continue
      
              # fade-in / fade-out(在 overlap 裁剪之后应用,确保正确的音频包络)
              fade_in_len, fade_out_len = _speech_safe_fade_lengths(
                  wf_data, write_samples, sample_rate, CONFIG["fade_ms"]
              )
              for i in range(fade_in_len):
                  gain = i / fade_in_len
                  s = i * 2
                  sample = int.from_bytes(wf_data[s:s+2], 'little', signed=True)
                  sample = int(sample * gain)
                  wf_data[s:s+2] = sample.to_bytes(2, 'little', signed=True)
              for i in range(fade_out_len):
                  gain = 1.0 - i / fade_out_len
                  s = (write_samples - 1 - i) * 2
                  if s < 0:
                      break
                  sample = int.from_bytes(wf_data[s:s+2], 'little', signed=True)
                  sample = int(sample * gain)
                  wf_data[s:s+2] = sample.to_bytes(2, 'little', signed=True)
      
              # Persist the exact complete per-beat PCM used by the canonical mix. Editable
              # exports must reference this file, not the longer pre-fit TTS input; otherwise
              # their timeline_end silently chops the final word even when ffmpeg is correct.
              placed_path = Path(work_dir) / f"_placed_{seg['index']:04d}.wav"
              with wave.open(str(placed_path), "wb") as placed_wav:
                  placed_wav.setnchannels(1)
                  placed_wav.setsampwidth(2)
                  placed_wav.setframerate(sample_rate)
                  placed_wav.writeframes(bytes(wf_data))
              seg["placed_audio_path"] = str(placed_path)
      
              buffer[actual_start * 2: actual_start * 2 + write_samples * 2] = wf_data
              seg["actual_place_start"] = actual_start / sample_rate
              seg["actual_place_end"] = (actual_start + write_samples) / sample_rate
              seg["placed_audio_duration"] = write_samples / sample_rate
              last_written_end = actual_start + write_samples
              prev_pause_samples = pause_samples
              placed_count += 1
      
          with wave.open(str(output_wav), "wb") as wf:
              wf.setnchannels(1)
              wf.setsampwidth(2)
              wf.setframerate(sample_rate)
              wf.writeframes(bytes(buffer))
      
          if tts_segments and placed_count == 0 and no_safe_fit_count == 0:
              output_wav.unlink(missing_ok=True)
              raise RuntimeError(
                  f"全部 {len(tts_segments)} 段解说均被跳过或未能写入"
                  f"(WAV 缺失或无可用时间;跳过 {skipped_count} 段),"
                  "已中止以避免生成无解说视频"
              )
      
          log(f"解说音轨: {video_duration:.1f}s, {len(tts_segments)} 段")
      
    • packaging.py 5.5 KB
      """Static packaging layers (frame / header / logo images) burned over the whole recap.
      
      The caller writes ``work_dir/packaging_layers.json`` (the orchestrator does so from a bound
      ``packaging`` template). Each layer is a local image scaled into a rect on a declared
      canvas. ffmpeg reads the images with ``movie=`` sources, so the render stays a single-input
      video filter; ``timeline.json`` gets matching image segments for editable export.
      """
      
      import json
      from pathlib import Path
      
      import lib
      from artifacts import file_identity
      from visual_render import _escape_subtitle_filter_path
      
      PACKAGING_LAYERS = "packaging_layers.json"
      _IMAGE_EXTS = {".png", ".jpg", ".jpeg", ".webp"}
      
      
      def load_packaging_layers(work_dir, canvas):
          """Validated layers for this canvas, or [] when the work_dir declares none."""
          path = Path(work_dir) / PACKAGING_LAYERS
          if not path.exists():
              return []
          plan = json.loads(path.read_text(encoding="utf-8"))
          declared = plan["canvas"]
          if (declared["width"], declared["height"]) != (canvas["width"], canvas["height"]):
              raise RuntimeError(
                  f"{PACKAGING_LAYERS} 按 {declared['width']}x{declared['height']} 设计,"
                  f"成片画布是 {canvas['width']}x{canvas['height']}"
              )
          layers = []
          for layer in plan["layers"]:
              image = Path(layer["path"]).expanduser().resolve()
              rect = layer["rect"]
              if not image.is_file() or image.suffix.lower() not in _IMAGE_EXTS:
                  raise RuntimeError(f"包装图层 {layer['name']} 的图片不可用: {image}")
              if (min(rect["x"], rect["y"]) < 0 or rect["width"] <= 0 or rect["height"] <= 0
                      or rect["x"] + rect["width"] > canvas["width"]
                      or rect["y"] + rect["height"] > canvas["height"]):
                  raise RuntimeError(f"包装图层 {layer['name']} 超出画布")
              layers.append({"name": layer["name"], "path": str(image), "rect": dict(rect),
                             "image_size": _image_size(image)})
          return layers
      
      
      def _image_size(image):
          """Pixel size of a layer image; the editor fits by it, the render stretches to the rect."""
          res = lib.run_cmd(["ffprobe", "-v", "error", "-select_streams", "v:0",
                             "-show_entries", "stream=width,height", "-of", "json", str(image)])
          try:
              stream = json.loads(res.stdout)["streams"][0]
              return {"width": int(stream["width"]), "height": int(stream["height"])}
          except (ValueError, KeyError, IndexError):
              raise RuntimeError(f"无法读取包装图层尺寸: {image}") from None
      
      
      def compose_video_filter(chain, layers, *, mask_first):
          """Join the existing filter chain, inserting packaging layers after the source mask.
      
          Order: source-subtitle mask → packaging layers → text overlays / burned subtitles →
          scaling. Without layers this is the plain comma-joined chain.
          """
          if not layers:
              return ",".join(chain)
          head, tail = (chain[:1], chain[1:]) if mask_first else ([], list(chain))
          graph = [
              f"movie=filename='{_escape_subtitle_filter_path(layer['path'])}',"
              f"scale={layer['rect']['width']}:{layer['rect']['height']},format=rgba[pk{i}]"
              for i, layer in enumerate(layers)
          ]
          current = "in"
          if head:
              graph.append(f"[in]{head[0]}[pm]")
              current = "pm"
          for i, layer in enumerate(layers):
              out = "out" if i == len(layers) - 1 and not tail else f"po{i}"
              graph.append(f"[{current}][pk{i}]overlay=x={layer['rect']['x']}:y={layer['rect']['y']}[{out}]")
              current = out
          if tail:
              graph.append(f"[{current}]{','.join(tail)}[out]")
          return ";".join(graph)
      
      
      def timeline_image_segments(layers, canvas, duration_s):
          """Full-length image segments in the timeline's center-origin, Y-up transform.
      
          The editor shows an image fitted inside the canvas (keeping its own aspect) at scale 1;
          the render stretches it to the rect, so x and y scales are derived separately.
          """
          width, height = canvas["width"], canvas["height"]
          segments = []
          for layer in layers:
              rect, size = layer["rect"], layer["image_size"]
              fit = min(width / size["width"], height / size["height"])
              segments.append({
                  "source_path": layer["path"],
                  "timeline_start": 0.0,
                  "timeline_end": duration_s,
                  "scale": {"x": round(rect["width"] / (size["width"] * fit), 6),
                            "y": round(rect["height"] / (size["height"] * fit), 6)},
                  "position": {
                      "x": round((rect["x"] + rect["width"] / 2 - width / 2) / (width / 2), 6),
                      "y": round((height / 2 - rect["y"] - rect["height"] / 2) / (height / 2), 6),
                  },
              })
          return segments
      
      
      def packaging_settings(work_dir):
          """What the manifest records: the plan file and each layer image's identity."""
          path = Path(work_dir) / PACKAGING_LAYERS if work_dir is not None else None
          if path is None or not path.exists():
              return {"artifact": PACKAGING_LAYERS, "present": False, "layers": []}
          plan = json.loads(path.read_text(encoding="utf-8"))
          layers = []
          for layer in plan["layers"]:
              image = Path(layer["path"]).expanduser().resolve()
              layers.append({"name": layer["name"], "path": str(image), "rect": layer["rect"],
                             **(file_identity(image) if image.is_file() else {})})
          return {"artifact": PACKAGING_LAYERS, "present": True, "identity": file_identity(path),
                  "template": plan.get("template"), "layers": layers}
      
    • pair_media.py 10.9 KB
      #!/usr/bin/env python3
      """Pair explicit picture and adopted audio assets without re-encoding either stream.
      
      This is asset pairing, not a claim of source reconstruction, correct editorial
      selection, speech alignment, or release approval. Existing assemble then consumes
      paired.mp4; its subtitle binding must be newly computed on that container/a:0.
      """
      
      import argparse
      from fractions import Fraction
      import json
      from pathlib import Path
      import subprocess
      
      from assemble_constants import SUPPORTED_PICTURE_CODECS
      from adoption.frozen_audio import probe_audio_packets, verify_adopted_audio
      from adoption.strict_inputs import (
          probe_json, require_declared_path, require_fields, require_integer, run_logged,
          without_digests, write_json_atomic,
      )
      
      
      _probe = probe_json
      
      
      def _asset(value, audio=False):
          value = without_digests(value, 'asset')
          require_fields(value, ['path', 'selected_stream'] if audio else ['path'], 'asset')
          path = require_declared_path(value, 'asset')
          if audio:
              return {'path': str(path),
                      'selected_stream': require_integer(value['selected_stream'], 'Audio selected_stream')}
          return {'path': str(path)}
      
      
      def _time(value):
          if value in (None, 'N/A'):
              raise ValueError('Missing media timing')
          try:
              return Fraction(str(value))
          except (ValueError, ZeroDivisionError) as exc:
              raise ValueError('Invalid media timing') from exc
      
      
      def probe_picture(path):
          """Actual full CFR presentation clock plus codec-order packet sizes and timestamps."""
          data = _probe(path, '-select_streams', 'v:0', '-show_streams', '-show_packets',
                        '-show_format', '-show_entries',
                        'format=format_name:stream=codec_name,profile,level,width,height,pix_fmt,'
                        'sample_aspect_ratio,field_order,color_range,color_space,color_transfer,'
                        'color_primaries,chroma_location,time_base,start_pts,duration_ts,avg_frame_rate'
                        ':packet=pts,dts,duration,size,side_data_list')
          streams = data.get('streams', [])
          if len(streams) != 1 or streams[0].get('codec_name') not in SUPPORTED_PICTURE_CODECS:
              raise ValueError('Pairing requires one selected H264/HEVC picture stream')
          if 'mp4' not in data.get('format', {}).get('format_name', '').split(','):
              raise ValueError('Pairing currently requires MP4-family picture container')
          v = streams[0]
          fps, tb = _time(v.get('avg_frame_rate')), _time(v.get('time_base'))
          if not 1 <= fps <= 120 or tb <= 0 or v.get('start_pts') != 0:
              raise ValueError('Picture requires positive CFR and a known zero start')
          duration = _time(v.get('duration_ts')) * tb
          frames = _probe(path, '-select_streams', 'v:0', '-show_frames',
                          '-show_entries', 'frame=pts')['frames']
          pts = [_time(frame.get('pts')) * tb for frame in frames]
          if not pts or any(t != Fraction(i, fps) for i, t in enumerate(pts)) or duration != len(pts) / fps:
              raise ValueError('Picture requires complete zero-origin CFR frame clock and exact duration')
          decoder_keys = ['codec_name', 'profile', 'level', 'width', 'height', 'pix_fmt',
                          'sample_aspect_ratio', 'field_order', 'color_range', 'color_space',
                          'color_transfer', 'color_primaries', 'chroma_location']
          packets = []
          for packet in data.get('packets', []):
              if packet.get('size') is None:
                  raise ValueError('Picture packet missing size')
              ticks = {key: str(_time(packet.get(key)) * tb) for key in ['pts', 'dts', 'duration']}
              if _time(ticks['duration']) <= 0:
                  raise ValueError('Picture packet has invalid duration')
              packets.append({**ticks, 'size': int(packet['size']),
                              'side_data_list': packet.get('side_data_list', [])})
          if len(packets) != len(pts):
              raise ValueError('Picture packet/frame counts disagree')
          return {'decoder': {key: v.get(key) for key in decoder_keys}, 'packets': packets,
                  'frame_pts': [str(t) for t in pts], 'frame_count': len(pts), 'fps': str(fps),
                  'duration': str(duration), 'start': '0'}
      
      
      def validate_aac_packet_interval(audio):
          """Validate exact AAC packet continuity against the stream-header interval."""
          if audio['codec'] != 'aac':
              raise ValueError('Pairing requires AAC adopted audio')
          packets = audio['packets']
          rate = audio['sample_rate']
          if type(rate) is not int or rate <= 0 or len(packets) < 2:
              raise ValueError('AAC packet timing is incomplete')
          nominal = _time(packets[0]['duration'])
          # Accept known AAC frame sample counts, never an arbitrary huge duration.
          if nominal * rate not in {960, 1024, 1920, 2048}:
              raise ValueError('Unsupported AAC nominal packet duration')
          for i, packet in enumerate(packets):
              d = _time(packet['duration'])
              if d <= 0 or d > nominal or (i < len(packets) - 1 and d != nominal):
                  raise ValueError('Nonuniform or invalid AAC packet duration')
              p, t = _time(packet['pts']), _time(packet['dts'])
              if p != t:
                  raise ValueError('Unsupported AAC PTS/DTS difference')
              if i and p != _time(packets[i-1]['pts']) + _time(packets[i-1]['duration']):
                  raise ValueError('AAC packet clock has gaps or overlaps')
          audio_start, audio_duration = _time(audio['start_time']), _time(audio['duration'])
          if audio_duration <= 0:
              raise ValueError('Invalid audio interval')
          audio_end = audio_start + audio_duration
          first_pts = _time(packets[0]['pts'])
          last_end = _time(packets[-1]['pts']) + _time(packets[-1]['duration'])
          if (not 0 <= audio_start - first_pts <= nominal
                  or abs(last_end - audio_end) > Fraction(1, rate)):
              raise ValueError('AAC packet clock does not match stream interval')
          return audio_start, audio_end, nominal
      
      
      def validate_pair_timing(picture, audio):
          """Narrow AAC/CFR compatibility; not a perceptual synchronization verdict."""
          audio_start, audio_end, nominal = validate_aac_packet_interval(audio)
          tolerance = max(1 / Fraction(picture['fps']), nominal)
          end_delta = audio_end - Fraction(picture['duration'])
          if abs(audio_start) > tolerance or abs(end_delta) > tolerance:
              raise ValueError('Picture/audio interval mismatch; no implicit trim, padding, offset or retime')
          return {'start_delta': str(audio_start),
                  'end_delta': str(end_delta),
                  'tolerance': str(tolerance), 'nominal_aac_packet': str(nominal)}
      
      
      def _run_mux(command, directory):
          run_logged(command, directory, 'mux', timeout=600)
      
      
      def _verify_output(path):
          data = _probe(path, '-show_streams')
          if [stream.get('codec_type') for stream in data['streams']] != ['video', 'audio']:
              raise ValueError('Paired output must contain only v:0 and a:0')
          result = subprocess.run(['ffmpeg', '-v', 'error', '-xerror', '-threads', '2', '-i', str(path),
                                   '-map', '0:v:0', '-map', '0:a:0', '-f', 'null', '-'],
                                  capture_output=True, text=True, timeout=600)
          if result.returncode or result.stderr.strip():
              raise ValueError('Paired output full decode failed')
      
      
      def run_pair(plan_path, output_dir, *, plan_only=False):
          plan_path, directory = Path(plan_path).resolve(), Path(output_dir).resolve()
          directory.mkdir(parents=True, exist_ok=False)
          report_path = directory / 'pair_run.json'
          staged, output = directory / 'paired.rendering.mp4', directory / 'paired.mp4'
          report = {'artifact': 'media_pair_run', 'schema_version': 1, 'status': 'PREPARING',
                    'direct_listening': 'NOT_CHECKED', 'normal_speed_review': 'NOT_CHECKED',
                    'release_approved': False}
          write_json_atomic(report_path, report)
          try:
              plan = json.loads(plan_path.read_bytes())
              require_fields(plan, ['artifact', 'schema_version', 'picture', 'audio'], 'media_pair plan')
              if plan['artifact'] != 'media_pair' or type(plan['schema_version']) is not int or plan['schema_version'] != 1:
                  raise ValueError('Unsupported media_pair schema')
              picture, audio = _asset(plan['picture']), _asset(plan['audio'], audio=True)
              report.update(plan={'path': str(plan_path)}, inputs={'picture': picture, 'audio': audio})
              video_facts = probe_picture(picture['path'])
              audio_facts = probe_audio_packets(audio['path'], audio['selected_stream'])
              report['timing'] = validate_pair_timing(video_facts, audio_facts)
              report['picture'] = {k: v for k, v in video_facts.items() if k not in ['packets', 'frame_pts']}
              report['audio'] = {'input_stream': audio['selected_stream'], 'output_stream': 0,
                                 'packet_count': audio_facts['packet_count']}
              if plan_only:
                  report['status'] = 'PLANNED'
                  write_json_atomic(report_path, report)
                  return report
              command = ['ffmpeg', '-nostdin', '-v', 'error', '-n', '-copyts',
                         '-i', picture['path'], '-i', audio['path'], '-map', '0:v:0',
                         '-map', f"1:a:{audio['selected_stream']}", '-c', 'copy',
                         '-movie_timescale', str(audio_facts['sample_rate']),
                         '-movflags', '+faststart', str(staged)]
              _run_mux(command, directory)
              if probe_picture(staged) != video_facts:
                  raise ValueError('Paired picture packets, decoder, geometry/color or full frame clock changed')
              proof = verify_adopted_audio(audio['path'], staged, audio['selected_stream'], 0)
              output_timing = validate_pair_timing(video_facts, proof['output'])
              if output_timing != report['timing']:
                  raise ValueError('Output audio presentation interval changed')
              report['output_timing'] = output_timing
              _verify_output(staged)
              write_json_atomic(directory / 'picture_identity.json', video_facts)
              write_json_atomic(directory / 'adopted_audio_identity.json', proof)
              report['output'] = {'path': str(output), 'full_decode': 'PASS',
                                  'picture_identity': 'EXACT', 'audio_packet_identity': 'EXACT'}
              staged.rename(output)
              report['status'] = 'PAIR_RENDERED'
              write_json_atomic(report_path, report)
              return report
          except Exception as exc:
              # A partial unique run is evidence, never a final/current asset. Keep logs.
              staged.unlink(missing_ok=True)
              output.unlink(missing_ok=True)
              report.pop('output', None)
              report.update(status='FAILED', error=f'{type(exc).__name__}: {exc}')
              write_json_atomic(report_path, report)
              raise
      
      
      def main():
          parser = argparse.ArgumentParser(description=__doc__)
          parser.add_argument('plan')
          parser.add_argument('--output-dir', required=True)
          parser.add_argument('--plan-only', action='store_true')
          args = parser.parse_args()
          report = run_pair(args.plan, args.output_dir, plan_only=args.plan_only)
          print(json.dumps(report, ensure_ascii=False, indent=2))
      
      
      if __name__ == '__main__':
          main()
      
    • render_preflight.py 1.3 KB
      """Local ffmpeg capability preflight for subtitle burn-in."""
      
      import shutil
      import subprocess
      
      from lib import CONFIG
      
      def _ffmpeg_filters():
          """Return ffmpeg's compiled-in filter names (caller has confirmed ffmpeg exists)."""
          result = subprocess.run(["ffmpeg", "-hide_banner", "-filters"],
                                  text=True, capture_output=True, timeout=20)
          if result.returncode != 0:
              return set()
          filters = set()
          for line in result.stdout.splitlines():
              parts = line.split()
              if len(parts) >= 2 and parts[0] and parts[0][0] in ".TSCAPN|":
                  filters.add(parts[1])
          return filters
      
      
      def _preflight_burn_subtitles():
          """Fail before the (re-encoding) render when burn-in is on but ffmpeg lacks the libass
          `subtitles` filter. Only fires when ffmpeg EXISTS but can't burn — an absent ffmpeg fails
          the render regardless."""
          if not CONFIG["burn_subtitles"]:
              return
          if shutil.which("ffmpeg") is None:
              return
          if "subtitles" not in _ffmpeg_filters():
              raise SystemExit(
                  "字幕烧录已开启,但当前 ffmpeg 不支持 subtitles/libass 滤镜,渲染会在最后一步失败。\n"
                  "  解决:安装带 libass 的 ffmpeg,或加 --no-burn-subtitles 关闭烧录(仍输出 .srt 外挂字幕)。")
      
    • source_score.py 23.7 KB
      #!/usr/bin/env python3
      """Prepare exact source-audio and continuous-score beds from a strict sample plan."""
      
      import argparse
      from fractions import Fraction
      import json
      import os
      import math
      from pathlib import Path
      import re
      import shutil
      import subprocess
      
      from assemble_constants import SUPPORTED_PICTURE_CODECS, frame_clock_samples
      from adoption.strict_inputs import (
          canonical_fraction, probe_json, read_json_bytes, require_declared_path, require_fields,
          require_integer, require_local_path, require_number, run_logged, without_digests,
          write_json_atomic,
      )
      
      
      RATE = 48_000
      CHANNELS = 2
      CODEC = "pcm_f32le"
      SOURCE_ROLES = {"protected_original", "mixed_original_under_narration"}
      FADE_SHAPES = {"linear", "half_cosine"}
      
      
      def _probe_cfr(path):
          data = probe_json(
              path, "-select_streams", "v:0", "-show_streams", "-show_packets",
              "-show_entries",
              "stream=codec_name,time_base,start_pts,start_time,duration_ts,duration,"
              "avg_frame_rate,r_frame_rate:packet=pts,duration",
          )
          streams = data.get("streams", [])
          if len(streams) != 1:
              raise ValueError("source requires exactly one selected v:0 clock")
          stream = streams[0]
          codec = stream.get("codec_name")
          if codec not in SUPPORTED_PICTURE_CODECS:
              raise ValueError("v1 packet/frame CFR proof supports H264 and HEVC picture sources only")
          average = canonical_fraction(stream.get("avg_frame_rate"), "actual average frame rate")
          real = canonical_fraction(stream.get("r_frame_rate"), "actual real frame rate")
          if average != real:
              raise ValueError("source must be same-speed CFR")
          start_time = Fraction(stream.get("start_time", "0"))
          if start_time != 0:
              raise ValueError("source video frame clock must start at zero")
          time_base = Fraction(stream["time_base"])
          packets = data.get("packets", [])
          try:
              pts = sorted(Fraction(int(packet["pts"])) * time_base for packet in packets)
              durations = [Fraction(int(packet["duration"])) * time_base for packet in packets]
          except (KeyError, TypeError, ValueError, ZeroDivisionError) as exc:
              raise ValueError("H26x packet/frame clock proof is incomplete") from exc
          frame_duration = 1 / average
          count = len(pts)
          if not pts or any(value != index * frame_duration for index, value in enumerate(pts)):
              raise ValueError("source picture packets do not prove a zero-origin CFR frame clock")
          if any(duration != frame_duration for duration in durations):
              raise ValueError("source picture packet durations are not one CFR frame")
          return {"fps": str(average), "frame_count": count,
                  "time_base": stream.get("time_base"), "start_time": "0",
                  "codec_name": codec, "clock_proof": "H26X_ONE_PACKET_PER_FRAME_PTS"}
      
      
      def _probe_audio(path, ordinal):
          data = probe_json(
              path, "-select_streams", f"a:{ordinal}", "-show_streams",
              "-show_entries", "stream=index,codec_name,sample_fmt,sample_rate,channels,"
              "channel_layout,time_base,start_pts,start_time,duration_ts,duration",
          )
          streams = data.get("streams", [])
          if len(streams) != 1:
              raise ValueError(f"audio stream a:{ordinal} is missing or ambiguous")
          stream = streams[0]
          return {key: stream.get(key) for key in (
              "index", "codec_name", "sample_fmt", "sample_rate", "channels",
              "channel_layout", "time_base", "start_pts", "start_time", "duration_ts", "duration"
          )}
      
      
      def _probe_pcm(path):
          stream = _probe_audio(path, 0)
          if stream["codec_name"] not in {"pcm_s16le", "pcm_s24le", "pcm_f32le"}:
              raise ValueError("canonical/frozen WAV requires PCM16, PCM24, or PCM float")
          if stream["sample_rate"] != str(RATE) or stream["channels"] != CHANNELS:
              raise ValueError("canonical/frozen WAV must be 48 kHz stereo")
          time_base = Fraction(stream["time_base"])
          samples = int(Fraction(stream["duration_ts"]) * time_base * RATE)
          if Fraction(stream["duration_ts"]) * time_base * RATE != samples:
              raise ValueError("WAV duration is not an integral 48 kHz sample count")
          return {"codec_name": stream["codec_name"], "sample_fmt": stream["sample_fmt"],
                  "sample_rate": RATE, "channels": CHANNELS, "samples": samples}
      
      
      def _validate_fades(value, duration, curves, label):
          fade_in = require_integer(value["fade_in_samples"], f"{label} fade_in_samples")
          fade_out = require_integer(value["fade_out_samples"], f"{label} fade_out_samples")
          if fade_in + fade_out > duration:
              raise ValueError(f"{label} fades overlap")
          if fade_in == 1 or fade_out == 1:
              raise ValueError(f"{label} fades require zero or at least two samples")
          if value["fade_shape"] not in curves:
              raise ValueError(f"unsupported {label} fade_shape")
          return fade_in, fade_out
      
      
      def load_plan(plan_path):
          plan_path, _, plan = read_json_bytes(plan_path, "source score plan")
          require_fields(plan, ["artifact", "schema_version", "output", "source_segments",
                         "source_silence", "score"], "source score plan")
          if plan["artifact"] != "source_score_plan" or type(plan["schema_version"]) is not int \
                  or plan["schema_version"] != 1:
              raise ValueError("unsupported source_score_plan schema")
          require_fields(plan["output"], ["sample_rate", "channels", "total_samples"], "output")
          if plan["output"]["sample_rate"] != RATE or plan["output"]["channels"] != CHANNELS:
              raise ValueError("v1 output is fixed at 48 kHz stereo")
          total = require_integer(plan["output"]["total_samples"], "total_samples", 1)
          if not isinstance(plan["source_segments"], list) or not isinstance(plan["source_silence"], list):
              raise ValueError("source segments and explicit silence must be lists")
          segments = []
          asset_cache = {}
          picture_cache = {}
          ids = set()
          segment_fields = [
              "id", "path", "audio_stream", "source_fps", "source_start_frame",
              "source_end_frame", "output_start_sample", "gain", "fade_in_samples",
              "fade_out_samples", "fade_shape", "role",
          ]
          for value in plan["source_segments"]:
              value = without_digests(value, "source segment")
              require_fields(value, segment_fields, "source segment")
              if not isinstance(value["id"], str) or not value["id"] or value["id"] in ids:
                  raise ValueError("source segment id must be unique and non-empty")
              ids.add(value["id"])
              ordinal = require_integer(value["audio_stream"], "audio_stream")
              path = require_declared_path(value, "source")
              fps = canonical_fraction(value["source_fps"], "source_fps")
              start = require_integer(value["source_start_frame"], "source_start_frame")
              end = require_integer(value["source_end_frame"], "source_end_frame", 1)
              if end <= start:
                  raise ValueError("source frame interval must be non-empty")
              key = (str(path), ordinal)
              if key not in asset_cache:
                  if str(path) not in picture_cache:
                      picture_cache[str(path)] = _probe_cfr(path)
                  picture = picture_cache[str(path)]
                  audio = _probe_audio(path, ordinal)
                  asset_cache[key] = {"path": str(path), "audio_stream": ordinal,
                                      "picture": picture, "audio": audio}
              facts = asset_cache[key]
              if facts["picture"]["fps"] != str(fps) or end > facts["picture"]["frame_count"]:
                  raise ValueError("declared source frame clock/range differs from actual CFR source")
              source_start_sample = frame_clock_samples(start, fps, RATE)
              source_end_sample = frame_clock_samples(end, fps, RATE)
              duration = source_end_sample - source_start_sample
              if duration <= 0:
                  raise ValueError("source frame interval is shorter than one 48 kHz sample")
              fade_in, fade_out = _validate_fades(value, duration, {"linear"}, "source")
              output_start = require_integer(value["output_start_sample"], "output_start_sample")
              output_end = output_start + duration
              if output_end > total or value["role"] not in SOURCE_ROLES:
                  raise ValueError("source output range or role is unsupported")
              segments.append({**value, "path": str(path),
                               "gain": require_number(value["gain"], "source gain", 0, 16),
                               "source_start_sample": source_start_sample,
                               "source_end_sample": source_end_sample,
                               "output_end_sample": output_end, "fade_in_samples": fade_in,
                               "fade_out_samples": fade_out, "asset_key": key})
          silence = []
          for value in plan["source_silence"]:
              require_fields(value, ["output_start_sample", "output_end_sample", "role"], "source silence")
              start = require_integer(value["output_start_sample"], "silence output_start_sample")
              end = require_integer(value["output_end_sample"], "silence output_end_sample", 1)
              if value["role"] != "silence" or not start < end <= total:
                  raise ValueError("invalid explicit source silence range")
              silence.append(dict(value))
          coverage = sorted(
              [(item["output_start_sample"], item["output_end_sample"]) for item in segments]
              + [(item["output_start_sample"], item["output_end_sample"]) for item in silence]
          )
          cursor = 0
          for start, end in coverage:
              if start != cursor:
                  raise ValueError("source bed requires exact nonoverlapping coverage with explicit silence")
              cursor = end
          if cursor != total:
              raise ValueError("source bed has an implicit tail gap")
          score = plan["score"]
          if not isinstance(score, dict) or score.get("kind") not in {"raw", "frozen", "none"}:
              raise ValueError("score kind must be raw, frozen, or none")
          score = without_digests(score, "score")
          if score["kind"] == "raw":
              require_fields(score, ["kind", "path", "audio_stream", "source_offset_sample",
                              "gain", "fade_in_samples", "fade_out_samples", "fade_shape"], "raw score")
              score_path = require_declared_path(score, "raw score")
              ordinal = require_integer(score["audio_stream"], "score audio_stream")
              offset = require_integer(score["source_offset_sample"], "score source_offset_sample")
              fade_in, fade_out = _validate_fades(score, total, FADE_SHAPES, "score")
              score = {**score, "path": str(score_path),
                       "audio_stream": ordinal, "source_offset_sample": offset,
                       "gain": require_number(score["gain"], "score gain", 0, 16),
                       "fade_in_samples": fade_in, "fade_out_samples": fade_out,
                       "audio": _probe_audio(score_path, ordinal)}
          elif score["kind"] == "frozen":
              require_fields(score, ["kind", "path", "audio_stream"], "frozen score")
              score_path = require_declared_path(score, "frozen score")
              ordinal = require_integer(score["audio_stream"], "score audio_stream")
              if ordinal != 0:
                  raise ValueError("frozen WAV supports only a:0")
              pcm = _probe_pcm(score_path)
              if pcm["samples"] != total:
                  raise ValueError("frozen score must exactly match total_samples")
              score = {**score, "path": str(score_path), "pcm": pcm}
          else:
              require_fields(score, ["kind"], "none score")
          return {"plan_path": plan_path,
                  "output": {**plan["output"], "codec": CODEC}, "source_segments": segments,
                  "source_silence": silence, "source_assets": list(asset_cache.values()),
                  "score": score}
      
      
      _run = run_logged
      
      
      def _decode_command(path, ordinal, output):
          return ["ffmpeg", "-nostdin", "-v", "error", "-n", "-copyts", "-i", str(path),
                  "-map", f"0:a:{ordinal}", "-af",
                  "aresample=48000:async=0:first_pts=0,aformat=sample_fmts=flt:"
                  "sample_rates=48000:channel_layouts=stereo", "-c:a", CODEC, str(output)]
      
      
      def _fade_filters(duration, fade_in, fade_out, curve):
          factors = []
          if fade_in:
              if curve == "linear":
                  factors.append(f"min(1\\,n/{fade_in - 1})")
              else:
                  position = f"max(0\\,min(1\\,n/{fade_in - 1}))"
                  factors.append(f"0.5-0.5*cos(PI*{position})")
          if fade_out:
              remaining = f"({duration - 1}-n)"
              if curve == "linear":
                  factors.append(f"max(0\\,min(1\\,{remaining}/{fade_out - 1}))")
              else:
                  position = f"max(0\\,min(1\\,{remaining}/{fade_out - 1}))"
                  factors.append(
                      f"0.5-0.5*cos(PI*{position})"
                  )
          if not factors:
              return []
          gain = "*".join(f"({factor})" for factor in factors)
          return [f"aeval=val(0)*({gain})|val(1)*({gain}):c=same"]
      
      
      def _render_source(plan, decoded, output, directory):
          total = plan["output"]["total_samples"]
          command = ["ffmpeg", "-nostdin", "-v", "error", "-n"]
          assets = {tuple(asset_key): index for index, asset_key in enumerate(decoded)}
          for asset_key in decoded:
              command += ["-i", str(decoded[asset_key])]
          filters = [f"anullsrc=r={RATE}:cl=stereo,atrim=end_sample={total}[base]"]
          inputs = ["[base]"]
          for index, segment in enumerate(plan["source_segments"]):
              duration = segment["source_end_sample"] - segment["source_start_sample"]
              chain = [f"[{assets[tuple(segment['asset_key'])]}:a]atrim="
                       f"start_sample={segment['source_start_sample']}:"
                       f"end_sample={segment['source_end_sample']}", "asetpts=PTS-STARTPTS",
                       f"volume={segment['gain']:.17g}"]
              chain += _fade_filters(duration, segment["fade_in_samples"],
                                     segment["fade_out_samples"], "linear")
              chain.append(f"adelay={segment['output_start_sample']}S:all=1[src{index}]")
              filters.append(",".join(chain))
              inputs.append(f"[src{index}]")
          filters.append("".join(inputs) + f"amix=inputs={len(inputs)}:duration=first:normalize=0,"
                         f"atrim=end_sample={total},aformat=sample_fmts=flt:"
                         f"sample_rates={RATE}:channel_layouts=stereo[out]")
          command += ["-filter_complex", ";".join(filters), "-map", "[out]", "-c:a", CODEC,
                      str(output)]
          _run(command, directory, "source")
      
      
      def _render_score(plan, decoded_score, output, directory):
          total = plan["output"]["total_samples"]
          score = plan["score"]
          if score["kind"] == "none":
              command = [
                  "ffmpeg", "-nostdin", "-v", "error", "-n", "-f", "lavfi", "-i",
                  f"anullsrc=r={RATE}:cl=stereo", "-af", f"atrim=end_sample={total}",
                  "-c:a", CODEC, str(output),
              ]
          elif score["kind"] == "frozen":
              command = ["ffmpeg", "-nostdin", "-v", "error", "-n", "-i", score["path"],
                         "-map", "0:a:0", "-af", "aformat=sample_fmts=flt:sample_rates=48000:"
                         "channel_layouts=stereo", "-c:a", CODEC, str(output)]
          else:
              chain = [f"atrim=start_sample={score['source_offset_sample']}:"
                       f"end_sample={score['source_offset_sample'] + total}",
                       "asetpts=PTS-STARTPTS", f"volume={score['gain']:.17g}"]
              chain += _fade_filters(total, score["fade_in_samples"], score["fade_out_samples"],
                                     score["fade_shape"])
              chain += [f"atrim=end_sample={total}", "aformat=sample_fmts=flt:sample_rates=48000:"
                        "channel_layouts=stereo"]
              command = ["ffmpeg", "-nostdin", "-v", "error", "-n", "-i", str(decoded_score),
                         "-af", ",".join(chain), "-c:a", CODEC, str(output)]
          _run(command, directory, "score")
      
      
      def _astats(path):
          result = subprocess.run(
              ["ffmpeg", "-nostdin", "-v", "info", "-i", str(path), "-af",
               "astats=metadata=0:reset=0", "-f", "null", "-"],
              capture_output=True, text=True, timeout=3600,
          )
          if result.returncode:
              raise ValueError("output astats decode failed")
          peaks = re.findall(r"Peak level dB:\s*(-?inf|[-+0-9.]+)", result.stderr, re.I)
          nan_counts = re.findall(r"Number of NaNs:\s*(\d+)", result.stderr)
          inf_counts = re.findall(r"Number of Infs:\s*(\d+)", result.stderr)
          if not peaks:
              raise ValueError("output peak statistics unavailable")
          peak_db = max(float(value) if value.lower() != "-inf" else -math.inf for value in peaks)
          finite = all(int(value) == 0 for value in nan_counts + inf_counts)
          if not finite:
              raise ValueError("output contains non-finite PCM samples")
          return {"finite": True, "peak": 0.0 if peak_db == -math.inf else 10 ** (peak_db / 20)}
      
      
      def _output_facts(path):
          """PCM format, sample count, size and peak/finiteness of one 48 kHz stereo float WAV."""
          return {"path": str(path), "bytes": os.stat(path).st_size, "pcm": _probe_pcm(path),
                  **_astats(path)}
      
      
      def validate_prepared_receipt(reference, expected_format):
          """Validate one completed prepared-bed receipt and its three current PCM stems."""
          require_fields(without_digests(reference, "prepared receipt reference"), ["path"],
                         "prepared receipt reference")
          require_fields(expected_format, ["sample_rate", "channels", "total_samples"],
                  "expected prepared format")
          if expected_format["sample_rate"] != RATE or expected_format["channels"] != CHANNELS:
              raise ValueError("prepared format must be 48 kHz stereo")
          require_integer(expected_format["total_samples"], "prepared total_samples", 1)
          receipt_path = require_declared_path(reference, "prepared receipt")
          receipt = read_json_bytes(receipt_path, "prepared receipt")[2]
          if not isinstance(receipt, dict) or receipt.get("artifact") != "prepared_bed_receipt" \
                  or receipt.get("schema_version") != 1 or receipt.get("status") != "PREPARED":
              raise ValueError("prepared receipt is not a completed v1 artifact")
          required_format = {**expected_format, "codec": CODEC}
          if receipt.get("format") != required_format:
              raise ValueError("prepared receipt format differs from expected picture clock")
          outputs = receipt.get("outputs")
          if not isinstance(outputs, dict) or set(outputs) != {
              "source_bed.wav", "score_bed.wav", "prepared_bed.wav"
          }:
              raise ValueError("prepared receipt requires exactly three named bed outputs")
          prepared = {}
          for name, declared in outputs.items():
              if not isinstance(declared, dict):
                  raise ValueError(f"prepared receipt {name} identity is invalid")
              path = require_local_path(declared.get("path"), name)
              stream = _probe_audio(path, 0)
              if stream["codec_name"] != CODEC or stream["sample_rate"] != str(RATE) or \
                      stream["channels"] != CHANNELS or \
                      (stream["start_time"] not in (None, "N/A") and
                       Fraction(stream["start_time"]) != 0):
                  raise ValueError(f"{name} must be zero-origin 48 kHz stereo float PCM")
              actual = _output_facts(path)
              if actual["pcm"]["codec_name"] != CODEC \
                      or actual["pcm"]["samples"] != expected_format["total_samples"]:
                  raise ValueError(f"{name} does not match the required full PCM format")
              prepared[name] = actual
          return {
              "reference": {"path": str(receipt_path)},
              "format": dict(expected_format), "outputs": prepared, "receipt": receipt,
          }
      
      
      def prepare_source_score(plan_path, output_dir):
          directory = Path(output_dir).resolve()
          directory.mkdir(parents=True, exist_ok=False)
          finals = {name: directory / name for name in
                    ("source_bed.wav", "score_bed.wav", "prepared_bed.wav")}
          staged = {name: directory / f".{Path(name).stem}.rendering.wav" for name in finals}
          receipt_path = directory / "prepared_bed_receipt.json"
          try:
              plan = load_plan(plan_path)
              decoded = {}
              for index, asset in enumerate(plan["source_assets"]):
                  key = (asset["path"], asset["audio_stream"])
                  path = directory / f".source_{index:03d}.decoded.wav"
                  _run(_decode_command(asset["path"], asset["audio_stream"], path),
                       directory, f"decode_source_{index:03d}")
                  decoded[key] = path
                  asset["canonical_pcm"] = _output_facts(path)
                  required = max(
                      segment["source_end_sample"] for segment in plan["source_segments"]
                      if tuple(segment["asset_key"]) == key
                  )
                  if asset["canonical_pcm"]["pcm"]["samples"] < required:
                      raise ValueError("decoded source audio is too short for a selected frame range")
              score_decode = None
              if plan["score"]["kind"] == "raw":
                  score_decode = directory / ".score.decoded.wav"
                  _run(_decode_command(plan["score"]["path"], plan["score"]["audio_stream"],
                                       score_decode), directory, "decode_score")
                  plan["score"]["canonical_decode"] = _output_facts(score_decode)
                  if plan["score"]["canonical_decode"]["pcm"]["samples"] < \
                          plan["score"]["source_offset_sample"] + plan["output"]["total_samples"]:
                      raise ValueError("raw score is too short for continuous offset window")
              _render_source(plan, decoded, staged["source_bed.wav"], directory)
              _render_score(plan, score_decode, staged["score_bed.wav"], directory)
              rendered = {
                  "source_bed.wav": _output_facts(staged["source_bed.wav"]),
                  "score_bed.wav": _output_facts(staged["score_bed.wav"]),
              }
              if plan["score"]["kind"] == "none":
                  shutil.copyfile(staged["source_bed.wav"], staged["prepared_bed.wav"])
              else:
                  command = [
                      "ffmpeg", "-nostdin", "-v", "error", "-n",
                      "-i", str(staged["source_bed.wav"]), "-i", str(staged["score_bed.wav"]),
                      "-filter_complex", f"[0:a][1:a]amix=inputs=2:duration=first:normalize=0,"
                      f"atrim=end_sample={plan['output']['total_samples']},aformat=sample_fmts=flt:"
                      "sample_rates=48000:channel_layouts=stereo[out]", "-map", "[out]",
                      "-c:a", CODEC, str(staged["prepared_bed.wav"]),
                  ]
                  _run(command, directory, "prepare")
              outputs = {**rendered, "prepared_bed.wav": _output_facts(staged["prepared_bed.wav"])}
              if any(value["pcm"]["samples"] != plan["output"]["total_samples"]
                     for value in outputs.values()):
                  raise ValueError("prepared bed sample counts differ from plan")
              if plan["score"]["kind"] == "none" and (
                      outputs["prepared_bed.wav"]["bytes"] != outputs["source_bed.wav"]["bytes"]
                      or outputs["prepared_bed.wav"]["pcm"] != outputs["source_bed.wav"]["pcm"]):
                  raise ValueError("none score must preserve the source bed exactly")
              outputs["prepared_bed.wav"]["headroom_policy"] = "FLOAT_PRESERVED_NO_MASTER"
              receipt = {
                  "artifact": "prepared_bed_receipt", "schema_version": 1, "status": "PREPARED",
                  "plan": {"path": str(plan["plan_path"])},
                  "format": plan["output"], "source_assets": plan["source_assets"],
                  "source_segments": [{key: value for key, value in item.items() if key != "asset_key"}
                                      for item in plan["source_segments"]],
                  "source_silence": plan["source_silence"], "score": plan["score"],
                  "outputs": outputs, "direct_listening": "NOT_CHECKED", "release_approved": False,
              }
              for name, path in staged.items():
                  path.rename(finals[name])
                  receipt["outputs"][name]["path"] = str(finals[name])
              write_json_atomic(receipt_path, receipt)
              return receipt
          except Exception:
              for path in [*staged.values(), *finals.values(), receipt_path]:
                  path.unlink(missing_ok=True)
              raise
      
      
      def main():
          parser = argparse.ArgumentParser(description=__doc__)
          parser.add_argument("plan")
          parser.add_argument("--output-dir", required=True)
          args = parser.parse_args()
          print(json.dumps(prepare_source_score(args.plan, args.output_dir), ensure_ascii=False))
      
      
      if __name__ == "__main__":
          main()
      
    • source_subtitles.py 20 KB
      """Original-dialogue subtitle loading, source mapping, and gap placement."""
      
      import json
      import re
      from pathlib import Path
      
      from artifacts import _load_work_json
      from audio_mix import _seg_place_window
      from lib import CONFIG
      from media import _plan_clip_spans
      from assemble_constants import (
          _AUTO_ORIGINAL_READ_CPS,
          _CLIP_CONTIGUITY_TOLERANCE,
          _MAX_ORIGINAL_READ_CPS,
          _MIN_ASR_CLIP_OVERLAP,
          _MIN_GAP_TO_SUBTITLE,
          _MIN_READABLE_SECONDS,
          _SUBTITLE_CLOSING_QUOTES,
      )
      from subtitles.track_binding import bound_subtitle_entries
      from subtitles.core import (
          _bracketed_original_chunks,
          _subtitle_entries,
      )
      
      def _has_user_subtitles(work_dir):
          """True when the user dropped a bring-your-own original-subtitle file into work_dir."""
          return work_dir is not None and any(
              (Path(work_dir) / name).exists()
              for name in ("user_subtitles.json", "user_subtitles.srt", "user_subtitles.ass")
          )
      
      
      def _source_subtitle_mask_policy(work_dir=None):
          """Explicit source-subtitle mask policy and trigger facts for visual QC/cache keys.
      
          Older builds treated ``MASK_SOURCE_SUBTITLES=True`` as an ambient default black
          band. The visual contract now requires an explicit policy, so a bare truthy
          legacy flag is represented as ``legacy_implicit`` and blocks the visual gate
          instead of silently masking picture information.
          """
          burn = CONFIG["burn_subtitles"]
          raw_policy = CONFIG["source_subtitle_mask_policy"]
          legacy_flag = CONFIG["mask_source_subtitles"]
          allowed = {"off", "opt_in", "safe", "forced"}
          declared = CONFIG["source_subtitle_mask_policy_declared"] or raw_policy in {"opt_in", "safe", "forced"}
          implicit = False
          if legacy_flag and not declared:
              raw_policy = "legacy_implicit"
              implicit = True
          elif raw_policy not in allowed:
              implicit = True
          user_subtitles = _has_user_subtitles(work_dir)
          active = False
          trigger = "policy_off"
          reason = "source subtitle masking disabled by explicit policy"
          if raw_policy == "off":
              active = False
          elif raw_policy in {"opt_in", "forced"}:
              active = burn and legacy_flag
              trigger = "burn_subtitles_and_legacy_mask_flag"
              reason = "explicit policy permits masking only with burned recap subtitles"
          elif raw_policy == "safe":
              active = burn and (legacy_flag or user_subtitles)
              trigger = "safe_policy_with_burned_subtitles"
              reason = "safe policy masks only when recap subtitles are burned and an original-subtitle source is declared"
          else:
              active = False
              trigger = "implicit_or_invalid_policy"
              reason = "mask_source_subtitles requires explicit SOURCE_SUBTITLE_MASK_POLICY"
          if not burn and active:
              active = False
              trigger = "burn_subtitles_disabled"
              reason = "mask-only black band is forbidden without burned recap subtitles"
          return {
              "policy": raw_policy,
              "declared": bool(declared and raw_policy in allowed),
              "active": bool(active),
              "scope": (
                  "measured_source_subtitle_band"
                  if active and 0 <= CONFIG["subtitle_y_top"] < CONFIG["subtitle_y_bot"]
                  else ("bottom_source_subtitle_band" if active else "none")
              ),
              "trigger": trigger,
              "reason": reason,
              "burn_subtitles": burn,
              "legacy_mask_flag": legacy_flag,
              "user_subtitles_present": user_subtitles,
              "blocking": implicit,
          }
      
      
      def _load_original_asr(work_dir):
          """The original speech transcription (asr_result.json), SOURCE-time [{start,end,text}]; [] when absent."""
          data = _load_work_json(work_dir, "asr_result.json")
          if data is None:
              return []
          return [{"start": float(s["start"]), "end": float(s["end"]), "text": s["text"]} for s in data]
      
      
      def _load_agent_original_subtitles(work_dir):
          """Agent-calibrated original-dialogue subtitles (original_subtitles.json): OUTPUT-time
          [{start,end,text}] the writer authors alongside narration.json — the corrected, gap-aligned
          transcript of what is ACTUALLY said in each original-audio gap (ASR errors/names fixed).
          None when absent or empty (then assemble falls back to a conservative auto-ASR mapping)."""
          data = _load_work_json(work_dir, "original_subtitles.json")
          if not data:
              return None
          return [{"start": float(s["start"]), "end": float(s["end"]), "text": s["text"]} for s in data]
      
      
      def _user_subtitle_entries(rows, source):
          """Validate user-authored {start,end,text} rows; a malformed row is an error, not a skip."""
          out = []
          for index, row in enumerate(rows, start=1):
              try:
                  start, end, text = float(row["start"]), float(row["end"]), row["text"].strip()
              except (KeyError, TypeError, ValueError, AttributeError) as exc:
                  raise ValueError(f"{source} 第 {index} 条缺少或无法解析 start/end/text: {row!r}") from exc
              if end <= start or not text:
                  raise ValueError(f"{source} 第 {index} 条无效(需要 end > start 且 text 非空): {row!r}")
              out.append({"start": start, "end": end, "text": text})
          return out
      
      
      def _parse_srt_timestamp(value):
          """Parse an SRT 'HH:MM:SS,mmm' (or ASS 'H:MM:SS.cc') timestamp into seconds."""
          m = re.match(r"\s*(\d+):(\d{1,2}):(\d{1,2})[.,](\d{1,3})\s*$", value)
          if not m:
              raise ValueError(f"无法解析字幕时间戳: {value!r}")
          h, mm, ss, frac = m.groups()
          return int(h) * 3600 + int(mm) * 60 + int(ss) + int(frac) / (10 ** len(frac))
      
      
      def _parse_srt_text(text, source):
          """Minimal SRT parser → [{start,end,text}]. Blank lines and missing indices are tolerated;
          a block without a parseable timing line is an error. Cues with no text are dropped."""
          segs = []
          for block in re.split(r"\n\s*\n", text.replace("\r\n", "\n").replace("\r", "\n")):
              lines = [ln for ln in block.split("\n") if ln.strip()]
              if not lines:
                  continue
              if lines[0].strip().isdigit():
                  lines = lines[1:]
              if not lines or "-->" not in lines[0]:
                  raise ValueError(f"{source}: 字幕块缺少时间行: {block.strip()!r}")
              start_text, end_text = lines[0].split("-->", 1)
              start, end = _parse_srt_timestamp(start_text), _parse_srt_timestamp(end_text)
              if end <= start:
                  raise ValueError(f"{source}: 字幕结束时间必须晚于开始时间: {lines[0].strip()!r}")
              body = " ".join(lines[1:]).strip()
              if body:
                  segs.append({"start": start, "end": end, "text": body})
          return segs
      
      
      def _parse_ass_text(text, source):
          """Minimal ASS Dialogue parser → [{start,end,text}] (Start, End are fields 2 and 3)."""
          segs = []
          for line in text.replace("\r\n", "\n").replace("\r", "\n").split("\n"):
              if not line.startswith("Dialogue:"):
                  continue
              fields = line[len("Dialogue:"):].split(",", 9)
              if len(fields) < 10:
                  raise ValueError(f"{source}: Dialogue 行字段不足: {line!r}")
              start, end = _parse_srt_timestamp(fields[1]), _parse_srt_timestamp(fields[2])
              if end <= start:
                  raise ValueError(f"{source}: 字幕结束时间必须晚于开始时间: {line!r}")
              body = re.sub(r"\{[^}]*\}", "", fields[9]).replace("\\N", " ").replace("\\n", " ").strip()
              if body:
                  segs.append({"start": start, "end": end, "text": body})
          return segs
      
      
      def _load_user_original_subtitles(work_dir):
          """User-supplied original-dialogue subtitles, the highest-priority source (above the agent file).
      
          Accepts (first existing wins):
            - user_subtitles.json: a bare list [{start,end,text}] (treated as OUTPUT-time, used verbatim),
              OR a wrapper {"timeline":"source"|"output", "lines":[...]} — "source" is remapped to OUTPUT
              via the cut clip spans, "output" (default) is used directly.
            - user_subtitles.srt / user_subtitles.ass: parsed minimally and defaulted to SOURCE-time,
              so they are remapped to OUTPUT via the cut clip spans.
          Returns OUTPUT-time [{start,end,text}], or None when no user file exists. A malformed
          file raises: the user asked for these subtitles, so silently falling back is wrong."""
          work = Path(work_dir)
          json_path = work / "user_subtitles.json"
          if json_path.exists():
              try:
                  data = json.loads(json_path.read_text(encoding="utf-8"))
              except ValueError as exc:
                  raise ValueError(f"{json_path} 不是合法 JSON: {exc}") from exc
              if isinstance(data, dict):
                  timeline = data.get("timeline", "output")
                  if timeline not in {"source", "output"}:
                      raise ValueError(f"{json_path}: timeline 必须是 source 或 output,当前为 {timeline!r}")
                  rows = data["lines"]
              else:
                  timeline, rows = "output", data
              segs = _user_subtitle_entries(rows, json_path.name)
              if timeline == "source":
                  segs = _map_asr_to_output(segs, _plan_clip_spans(work))
              return segs
      
          for name in ("user_subtitles.srt", "user_subtitles.ass"):
              path = work / name
              if not path.exists():
                  continue
              parser = _parse_ass_text if name.endswith(".ass") else _parse_srt_text
              segs = parser(path.read_text(encoding="utf-8"), name)
              # .srt/.ass default to SOURCE-time → remap onto the output timeline (identity in full mode).
              return _map_asr_to_output(segs, _plan_clip_spans(work))
      
          return None
      
      
      def _map_asr_to_output(asr_segs, clip_spans):
          """Map SOURCE-time ASR segments onto the OUTPUT timeline. Full mode (clip_spans None) is
          identity. Keep one utterance across source/output-continuous picture cuts;
          real source deletions, output gaps and repeated playback remain separate."""
          if clip_spans is None:
              return [dict(s) for s in asr_segs]
          out = []
          for seg in asr_segs:
              fragments = []
              for c in clip_spans:
                  ov_s, ov_e = max(seg["start"], c["source_start"]), min(seg["end"], c["source_end"])
                  if ov_e <= ov_s:
                      continue
                  start = c["output_start"] + (ov_s - c["source_start"])
                  end = c["output_start"] + (ov_e - c["source_start"])
                  entry = c.get("entry", {})
                  source_key = (c.get("source_id") or entry.get("source_id"),
                                c.get("source_path") or entry.get("source_path"))
                  if (fragments and abs(fragments[-1]["source_end"] - ov_s) < _CLIP_CONTIGUITY_TOLERANCE
                          and abs(fragments[-1]["end"] - start) < _CLIP_CONTIGUITY_TOLERANCE
                          and fragments[-1]["source_key"] == source_key):
                      fragments[-1].update(end=end, source_end=ov_e)
                  else:
                      fragments.append({"start": start, "end": end, "source_end": ov_e,
                                        "source_key": source_key})
              # Filter after joining: individually tiny pieces may be one readable phrase.
              out.extend({"start": f["start"], "end": f["end"], "text": seg["text"]}
                         for f in fragments if f["end"] - f["start"] > _MIN_ASR_CLIP_OVERLAP)
          return out
      
      
      def _narration_gap_windows(tts_segments, video_duration, min_gap=_MIN_GAP_TO_SUBTITLE):
          """OUTPUT-timeline stretches with NO narration (the original-audio blocks): the complement of
          the merged narration placement windows within [0, video_duration], keeping gaps >= min_gap."""
          placed = sorted(map(_seg_place_window, tts_segments), key=lambda w: w[0])
          merged = []
          for s, e in placed:
              if e - s <= 0:
                  continue
              if merged and s <= merged[-1][1]:
                  merged[-1][1] = max(merged[-1][1], e)
              else:
                  merged.append([s, e])
          gaps, cursor = [], 0.0
          for s, e in merged:
              if s - cursor >= min_gap:
                  gaps.append((cursor, s))
              cursor = max(cursor, e)
          if video_duration - cursor >= min_gap:
              gaps.append((cursor, float(video_duration)))
          return gaps
      
      
      def _original_gap_subtitle_entries(tts_segments, work_dir, video_duration):
          """Subtitle entries for the ORIGINAL dialogue during the original-audio blocks (narration
          gaps), so the band is not blank while the original speaks. Off unless we are burning and
          subtitle_original_in_gaps is set; no-op when there is no ASR. Cut mode remaps ASR to output."""
          # Fill the gaps when either (a) we are masking the source's own burned-in subs (so the band is
          # blank without us), or (b) the user supplied their own subtitle file — a clear signal they want
          # the original dialogue shown, e.g. a clean/foreign source with mask OFF (no burned subs to
          # double). Without a user file we keep the mask requirement so we don't double the source's own
          # visible subs. subtitle_original_in_gaps is the explicit override either way.
          if not (CONFIG["burn_subtitles"]
                  and CONFIG["subtitle_original_in_gaps"]
                  and (_source_subtitle_mask_covers_gaps(work_dir) or _has_user_subtitles(work_dir))):
              return []
          gaps = _narration_gap_windows(tts_segments, video_duration)
          if not gaps:
              return []
      
          # Source ladder (highest priority first): user-supplied file → agent-calibrated transcript →
          # conservative auto-ASR mapping. The user file and the agent file are time-precise (their spans
          # are the real on-screen windows), so they take the interval-clip "precise" path; raw ASR is
          # coarse and stays on the midpoint+over-render-guard fallback path.
          max_chars = CONFIG["subtitle_max_chars"]
          user = _load_user_original_subtitles(work_dir)
          if user is not None:
              return _precise_gap_entries(user, gaps, max_chars)
          agent = _load_agent_original_subtitles(work_dir)
          if agent is not None:
              return _precise_gap_entries(agent, gaps, max_chars)
          asr = _load_original_asr(work_dir)
          if not asr:
              return []
          return _fallback_gap_entries(_map_asr_to_output(asr, _plan_clip_spans(work_dir)), gaps, max_chars)
      
      
      def _precise_gap_entries(candidates, gaps, max_chars):
          """Precise path for time-accurate sources (user / agent-calibrated): interval-CLIP each line
          across the gap boundaries it overlaps, emitting one sub-entry per overlapped gap (clipped to
          that gap). A line straddling two gaps is split, not snapped to one or dropped; only sub-fragments
          shorter than _MIN_READABLE_SECONDS are dropped. No over-render guard (the source is trusted)."""
          entries = []
          for seg in candidates:
              text = seg["text"]
              seg_start, seg_end = float(seg["start"]), float(seg["end"])
              overlaps = [
                  (max(seg_start, gs), min(seg_end, ge))
                  for gs, ge in gaps
                  if min(seg_end, ge) - max(seg_start, gs) >= _MIN_READABLE_SECONDS
              ]
              if not overlaps:
                  continue
              if len(overlaps) == 1:
                  # the common case (a line authored within one gap): show it whole in that gap
                  cs, ce = overlaps[0]
                  entries.extend(_bracketed_original_chunks(text, cs, ce, max_chars))
                  continue
              # the line straddles a narration block: show each gap only ITS portion of the text
              # (proportional to the time the line overlaps that gap) instead of the whole line twice.
              seg_dur = seg_end - seg_start
              n = len(text)
              for cs, ce in overlaps:
                  lo = max(0, int(round((cs - seg_start) / seg_dur * n)))
                  hi = min(n, int(round((ce - seg_start) / seg_dur * n)))
                  piece = text[lo:hi].strip()
                  if piece:
                      entries.extend(_bracketed_original_chunks(piece, cs, ce, max_chars))
          return entries
      
      
      def _split_sentences_keep_delims(text):
          """Split on terminal CJK sentence marks 。!? keeping each delimiter with its sentence. A
          fragment that is only closing quotes/brackets (e.g. a trailing 」 after a 。 inside a quote) is
          re-attached to the previous sentence so quoted speech is never split off into a bare 」."""
          parts = [p.strip() for p in re.split(r"(?<=[。!?])", text) if p.strip()]
          merged = []
          for part in parts:
              if merged and all(ch in _SUBTITLE_CLOSING_QUOTES for ch in part):
                  merged[-1] += part
              else:
                  merged.append(part)
          return merged
      
      
      def _fallback_gap_entries(candidates, gaps, max_chars):
          """Coarse-ASR fallback. Each coarse-ASR line spans a whole window with no per-sentence onset, so
          it is split into WHOLE sentences (never mid-word); each sentence is assigned to the gap its
          char-proportional midpoint lands in, and within a gap the assigned sentences are packed
          SEQUENTIALLY from the first one's estimated onset at a comfortable read rate — so two lines in
          one gap never overlap or scatter to char-proportional tail slots — capped at the gap end. An
          over-dense gap front-truncates (shown) rather than dropping to blank."""
          # 1) split each coarse line into WHOLE sentences (never mid-word) and assign each to the gap
          #    its char-proportional midpoint lands in (the only "which gap" signal coarse ASR gives).
          buckets = {}  # gap_index -> [(estimated_onset, sentence_text)]
          for seg in candidates:
              for text in _split_sentences_keep_delims(seg["text"]):
                  sub = _sentence_subspan(seg, text)
                  mid = (sub["start"] + sub["end"]) / 2.0
                  gi = next((i for i, (gs, ge) in enumerate(gaps) if gs <= mid < ge), None)
                  if gi is None:
                      continue
                  buckets.setdefault(gi, []).append((sub["start"], text))
          # 2) within each gap, pack the assigned sentences SEQUENTIALLY from the gap onset at a
          #    comfortable read rate. Anchoring to the gap onset (vs each sentence's char-proportional
          #    tail position) stops a line heard early from being shoved to the END of its window — the
          #    coarse-ASR lag. Full mode keeps the real ASR onset (the first sentence's own start); an
          #    over-dense gap front-truncates (shown) rather than dropping to blank.
          entries = []
          for gi, items in buckets.items():
              gs, ge = gaps[gi]
              items.sort(key=lambda it: it[0])
              # start at the first assigned sentence's estimated onset (clamped into the gap), then pack
              # the rest sequentially so they never overlap or scatter to char-proportional tail slots.
              cursor = min(ge, max(gs, min(start for start, _ in items)))
              for _, text in items:
                  if cursor >= ge - _MIN_READABLE_SECONDS:
                      break
                  ce2 = min(ge, cursor + max(_MIN_READABLE_SECONDS, len(text) / _AUTO_ORIGINAL_READ_CPS))
                  if ce2 - cursor < _MIN_READABLE_SECONDS:
                      break
                  max_len = int((ce2 - cursor) * _MAX_ORIGINAL_READ_CPS)
                  entries.extend(_bracketed_original_chunks(text[:max_len], cursor, ce2, max_chars))
                  cursor = ce2
          return entries
      
      
      def _sentence_subspan(seg, sentence):
          """The slice of seg's [start,end] window that this sentence occupies, by character proportion.
          Single-sentence lines return the whole span unchanged."""
          full = seg["text"].strip()
          idx = full.find(sentence)
          if not full or sentence == full or idx < 0:
              return {"start": seg["start"], "end": seg["end"]}
          span = seg["end"] - seg["start"]
          s = seg["start"] + span * (idx / len(full))
          e = seg["start"] + span * ((idx + len(sentence)) / len(full))
          return {"start": s, "end": e}
      
      
      def _combined_subtitle_entries(narration, work_dir, video_duration):
          """Narration subtitle entries plus original-dialogue entries in the gaps, sorted by start.
          Original entries are confined to narration gaps, so they never overlap narration entries."""
          bound = bound_subtitle_entries(work_dir, video_duration)
          if bound is not None:
              return bound
          entries = _subtitle_entries(narration)
          entries.extend(_original_gap_subtitle_entries(narration, work_dir, video_duration))
          entries.sort(key=lambda x: (x["start"], x["end"]))
          return entries
      
      
      def _source_subtitle_mask_covers_gaps(work_dir=None):
          """Whether the effective source mask hides hardcoded subtitles outside narration."""
          if not _source_subtitle_mask_policy(work_dir)["active"]:
              return False
          return CONFIG["subtitle_mask_opacity"] >= 1.0 - 1e-9 and CONFIG["source_subtitle_mask_timing"] == "all"
      
    • timeline.py 9.2 KB
      """Multi-track timeline model for the recap (backend-neutral, stdlib only).
      
      A `Timeline` is a small, serializable representation of the finished recap as a
      set of tracks — exactly like a cut-tool project:
      
        - one **video** track: the source clip(s), each carrying its own *original
          audio* with a per-clip volume automation (the ducking: a continuous low bed
          under narration, held across short inter-sentence gaps, back up only at the
          lead-in/out and genuine long gaps);
        - one **narration** audio track: the placed TTS beats;
        - an optional **bgm** audio track: a looped music bed with its own ducking;
        - one **subtitle** (text) track: the narration lines.
        - optional **image** tracks: local photo overlays with normalized center-origin,
          Y-up transforms for editable JianYing export.
      
      The canonical ducking semantics live in `audio_automation.py`; ffmpeg
      (`assemble.py`) and this timeline model both derive their automation from that
      shared source. This model is emitted as `timeline.json` and consumed by the
      *optional* 剪映 exporter. The model itself knows nothing about ffmpeg or 剪映 —
      times are plain seconds and volumes are plain gains, so any backend can read it.
      """
      
      import json
      import math
      from copy import deepcopy
      
      from audio_automation import fixed_ducking_keyframes as ducking_keyframes
      from audio_automation import release_ducking_keyframes
      
      SCHEMA_VERSION = 2
      
      
      def _ceil_time(value, digits=4):
          """Round an interval end outward so serialization can never shorten media."""
          scale = 10 ** digits
          return math.ceil((float(value) * scale) - 1e-9) / scale
      
      
      def _floor_time(value, digits=4):
          """Round an interval start outward so serialization cannot clip source samples."""
          scale = 10 ** digits
          return math.floor((float(value) * scale) + 1e-9) / scale
      
      
      def build_timeline(canvas, duration_s, video_clips, narration_segments,
                         bgm=None, ducking=None, subtitle_segments=None,
                         image_segments=(), resource_packages=None,
                         style_presets=None, extra_tracks=()):
          """Assemble a Timeline dict from resolved placement data.
      
          canvas: {"width", "height", "fps"}
          duration_s: total output length (seconds)
          video_clips: ordered [{"source_path", "source_start", "source_end",
                       "timeline_start", "timeline_end"}] (cut mode: one per clip;
                       full mode: a single clip spanning the whole video).
          narration_segments: placed beats [{"source_path", "timeline_start",
                       "timeline_end", "text", "overlaps_speech", "gain"}]; a zero-width
                       beat (an unplaced segment) is skipped.
          subtitle_segments: optional display-ready text cues [{"text", "timeline_start",
                       "timeline_end"}]. When present, this is authoritative for the
                       subtitle/text track; narration segment text remains raw editor metadata.
          bgm: optional {"source_path", "volume", "ducking_volume", "fade"}.
          ducking: {"idle", "speech", "quiet", "fade", "bridge"} for the original-audio
                   automation; None disables original ducking (flat original). `bridge` holds
                   the duck across inter-beat gaps shorter than it.
          image_segments: optional v2 local image overlays [{"source_path", "timeline_start",
                       "timeline_end", ...authoring extensions}], passed through as authored.
          """
          placed = [
              s for s in narration_segments
              if float(s["timeline_end"]) > float(s["timeline_start"])
          ]
          windows = [(float(s["timeline_start"]), float(s["timeline_end"])) for s in placed]
          duck_windows = []
          if ducking is not None:
              for s in placed:
                  end = float(s["timeline_end"])
                  hold_end = max(end, float(s.get("source_duck_end", end)))
                  restore_at = max(
                      hold_end, float(s.get("source_restore_at", hold_end + float(ducking["fade"])))
                  )
                  level = float(ducking["speech" if s["overlaps_speech"] else "quiet"])
                  duck_windows.append((float(s["timeline_start"]), hold_end, level, restore_at))
      
          # --- video track: each clip carries its original audio + ducking automation
          video_clip_objs = []
          for c in video_clips:
              ts, te = float(c["timeline_start"]), float(c["timeline_end"])
              audio = {"role": "original", "volume_keyframes": []}
              if ducking is not None:
                  audio["volume_keyframes"] = release_ducking_keyframes(
                      duck_windows, ducking["idle"], ducking["fade"], ts, te,
                      bridge=ducking["bridge"])
                  audio["base_gain"] = round(float(ducking["idle"]), 4)
              else:
                  audio["base_gain"] = 1.0
              video_clip = {
                  "source_path": c["source_path"],
                  "source_start": round(float(c["source_start"]), 4),
                  "source_end": round(float(c["source_end"]), 4),
                  "timeline_start": round(ts, 4),
                  "timeline_end": round(te, 4),
                  "audio": audio,
              }
              for key in (
                  "chroma", "compound", "flip", "green_background", "lut", "mask",
                  "opacity", "position", "reverse", "reverse_path", "rotation_degrees",
                  "scale", "speed", "transition",
              ):
                  if key in c:
                      video_clip[key] = deepcopy(c[key])
              video_clip_objs.append(video_clip)
      
          tracks = [{"kind": "video", "name": "video", "clips": video_clip_objs}]
      
          # --- narration track
          narr_segs = []
          for s in placed:
              narration = {
                  "source_path": s["source_path"],
                  "timeline_start": _floor_time(s["timeline_start"], 4),
                  "timeline_end": _ceil_time(s["timeline_end"], 4),
                  "gain": round(float(s["gain"]), 4),
                  "text": s["text"],
                  "overlaps_speech": bool(s["overlaps_speech"]),
              }
              for key in ("source_duck_end", "source_restore_at", "source_handoff_status", "source_entry_status"):
                  if key in s:
                      narration[key] = deepcopy(s[key])
              if "speed" in s:
                  narration["speed"] = float(s["speed"])
              narr_segs.append(narration)
          if narr_segs:
              tracks.append({"kind": "audio", "name": "narration", "role": "narration",
                             "segments": narr_segs})
      
          # --- bgm track (optional, looped, ducked under narration)
          if bgm:
              base = float(bgm["volume"])
              duck = float(bgm["ducking_volume"])
              fade = float(bgm["fade"])
              kfs = ducking_keyframes(windows, base, duck, fade, 0.0, duration_s,
                                      bridge=ducking["bridge"] if ducking else None)
              tracks.append({
                  "kind": "audio", "name": "bgm", "role": "bgm", "loop": True,
                  "segments": [{
                      "source_path": bgm["source_path"],
                      "timeline_start": 0.0,
                      "timeline_end": round(float(duration_s), 4),
                      "gain": round(base, 4),
                      "volume_keyframes": kfs,
                  }],
              })
      
          # --- subtitle (text) track: empty cues carry nothing to display
          text_source = subtitle_segments if subtitle_segments is not None else narration_segments
          text_segs = []
          for s in text_source:
              if not s.get("text"):
                  continue
              ts, te = float(s["timeline_start"]), float(s["timeline_end"])
              if te <= ts:
                  continue
              text_segment = {
                  "text": s["text"],
                  "timeline_start": round(ts, 4),
                  "timeline_end": round(te, 4),
              }
              for key in ("flip", "opacity", "position", "rotation_degrees", "scale", "style", "style_id", "words"):
                  if key in s:
                      text_segment[key] = deepcopy(s[key])
              text_segs.append(text_segment)
          if text_segs:
              tracks.append({"kind": "text", "name": "subtitle", "segments": text_segs})
      
          # --- local image overlays (optional, timeline schema v2). Transform fields are
          # optional authoring extensions; the JianYing exporter validates and defaults them.
          images = []
          for segment in image_segments:
              image = {
                  "source_path": segment["source_path"],
                  "timeline_start": round(float(segment["timeline_start"]), 4),
                  "timeline_end": round(float(segment["timeline_end"]), 4),
              }
              for key in ("flip", "lut", "mask", "opacity", "position", "rotation_degrees",
                          "scale", "speed", "transition"):
                  if key in segment:
                      image[key] = deepcopy(segment[key])
              images.append(image)
          if images:
              tracks.append({"kind": "image", "name": "image", "segments": images})
      
          tracks.extend(deepcopy(track) for track in extra_tracks)
      
          timeline = {
              "schema_version": SCHEMA_VERSION,
              "canvas": {"width": int(canvas["width"]), "height": int(canvas["height"]),
                         "fps": float(canvas["fps"])},
              "duration": round(float(duration_s), 4),
              "tracks": tracks,
          }
          if resource_packages:
              timeline["resource_packages"] = deepcopy(resource_packages)
          if style_presets:
              timeline["style_presets"] = deepcopy(style_presets)
          return timeline
      
      
      def save_timeline(timeline, path):
          with open(path, "w", encoding="utf-8") as f:
              json.dump(timeline, f, ensure_ascii=False, indent=2)
          return path
      
      
      def load_timeline(path):
          with open(path, encoding="utf-8") as f:
              return json.load(f)
      
    • timeline_emit.py 7.7 KB
      """Backend-neutral timeline emission for the video-assemble skill."""
      
      from pathlib import Path
      
      from audio_mix import _seg_place_window
      from lib import CONFIG, log
      from media import _build_video_clips
      from source_subtitles import _combined_subtitle_entries
      from timeline import build_timeline, save_timeline
      import packaging
      
      def _timeline_subtitle_segments(tts_segments, work_dir, duration_s):
          """Display-ready subtitle cues for timeline/export text tracks.
      
          The narration audio track keeps raw semantic text for editor reference; this
          payload mirrors SRT/ASS display policy, including terminal-punctuation cleanup
          and original-dialogue gap subtitles when configured.
          """
          return [
              {
                  "text": entry["text"],
                  "timeline_start": float(entry["start"]),
                  "timeline_end": float(entry["end"]),
              }
              for entry in _combined_subtitle_entries(tts_segments, work_dir, duration_s)
          ]
      
      
      def _emit_timeline(input_video, tts_segments, work_dir, duration_s, canvas, has_bgm, *,
                         audio_mode="narration", selected_audio_stream=0,
                         explicit_audio_mix=None):
          """Build and persist the backend-neutral multi-track timeline.json."""
          if audio_mode == "narration" and explicit_audio_mix is None:
              video_clips = _build_video_clips(input_video, work_dir, duration_s)
          else:
              # Non-narration sound comes from the actual current picture input as one
              # complete interval. Re-expanding an old cut plan would substitute different
              # source sound and make the optional editor project misrepresent the render.
              video_clips = [{
                  "source_path": str(Path(input_video)),
                  "source_start": 0.0,
                  "source_end": float(duration_s),
                  "timeline_start": 0.0,
                  "timeline_end": float(duration_s),
              }]
          # Mix segments are 1:1 with the UNFILTERED narration list, so they must be looked
          # up by their own index; a skipped beat would otherwise shift every later gain and
          # sample bound onto the wrong segment.
          mix_by_index = (
              {item["index"]: item for item in explicit_audio_mix["segments"]}
              if explicit_audio_mix is not None else {}
          )
          narration_segments = []
          placed_indices = []
          for seg in tts_segments:
              s, e = _seg_place_window(seg)
              if e <= s:
                  continue
              narration_item = {
                  # JianYing must consume the exact WAV written into narration.wav. In
                  # particular, a tempo-adjusted beat cannot reference its longer pre-fit
                  # source or the editor will trim its final words at timeline_end.
                  "source_path": seg["placed_audio_path"],
                  "timeline_start": s, "timeline_end": e,
                  "text": seg["narration"],
                  "overlaps_speech": seg["overlaps_speech"],
                  "gain": (
                      mix_by_index[seg["index"]]["gain"]
                      if explicit_audio_mix is not None else 1.0
                  ),
              }
              for key in ("source_duck_end", "source_restore_at", "source_handoff_status", "source_entry_status"):
                  if key in seg:
                      narration_item[key] = seg[key]
              narration_segments.append(narration_item)
              placed_indices.append(seg["index"])
          fade = CONFIG["duck_fade_seconds"]
          bgm = None
          if has_bgm and explicit_audio_mix is None:
              bgm = {"source_path": CONFIG["bgm_path"],
                     "volume": CONFIG["bgm_volume"],
                     "ducking_volume": CONFIG["bgm_ducking_volume"],
                     "fade": fade}
          # carry ducking automation whenever ducking is on at all; even under sidechain
          # mode the draft gets editable volume keyframes (ffmpeg stays the canonical mix)
          ducking = None
          if audio_mode == "narration" and explicit_audio_mix is None \
                  and CONFIG["ducking_mode"] != "none":
              ducking = {"idle": CONFIG["idle_orig_volume"],
                         "speech": CONFIG["speech_ducking_volume"],
                         "quiet": CONFIG["zone_ducking_volume"],
                         "fade": fade,
                         "bridge": CONFIG["duck_bridge_seconds"]}
          subtitle_segments = _timeline_subtitle_segments(tts_segments, work_dir, duration_s)
          timeline = build_timeline(canvas, duration_s, video_clips,
                                    narration_segments, bgm=bgm, ducking=ducking,
                                    subtitle_segments=subtitle_segments,
                                    image_segments=packaging.timeline_image_segments(
                                        packaging.load_packaging_layers(work_dir, canvas),
                                        canvas, duration_s))
          if explicit_audio_mix is not None:
              for clip in timeline["tracks"][0]["clips"]:
                  clip["audio"] = {
                      "role": "picture_audio_not_consumed", "base_gain": 0.0,
                      "volume_keyframes": [],
                  }
              narration_track = next(
                  (track for track in timeline["tracks"] if track.get("name") == "narration"), None
              )
              if narration_track:
                  for segment, index in zip(narration_track["segments"], placed_indices):
                      adopted = mix_by_index[index]
                      segment.update({
                          "gain": adopted["gain"],
                          "output_start_sample": adopted["output_start_sample"],
                          "output_end_sample": adopted["output_end_sample"],
                          "sample_rate": 48_000,
                      })
              timeline["tracks"].append({
                  "kind": "audio", "name": "prepared_bed", "role": "prepared_bed",
                  "segments": [{
                      "source_path": explicit_audio_mix["prepared"]["prepared_bed.wav"]["path"],
                      "timeline_start": 0.0, "timeline_end": float(duration_s), "gain": 1.0,
                  }],
              })
              timeline["audio_delivery"] = {
                  "mode": "explicit_adopted_full_sound", "sample_rate": 48_000,
                  "total_samples": explicit_audio_mix["format"]["total_samples"],
                  "master_gain_db": explicit_audio_mix["master_gain_db"],
                  "canonical_renderer": "ffmpeg_explicit_mix",
                  "reconstructable": False,
                  "reconstructable_reason": (
                      "optional editor export does not implement the adopted 48 kHz "
                      "prepared-bed, per-segment gain, and fixed-master chain"
                  ),
              }
          if audio_mode != "narration":
              source_gain = 1.0 if audio_mode == "adopted-packet-copy" else CONFIG["idle_orig_volume"]
              for clip in timeline["tracks"][0]["clips"]:
                  clip["audio"].update({
                      "base_gain": round(float(source_gain), 4),
                      "selected_stream": selected_audio_stream,
                      "mode": audio_mode,
                  })
              timeline["audio_delivery"] = {
                  "mode": audio_mode,
                  "selected_stream": selected_audio_stream,
                  "packet_frozen": audio_mode == "adopted-packet-copy",
                  "reconstructable": (
                      audio_mode == "adopted-packet-copy" and selected_audio_stream == 0
                  ),
                  "canonical_renderer": "stream_copy" if audio_mode == "adopted-packet-copy" else "ffmpeg_mix",
              }
          degraded = [
              {"source_path": clip["source_path"], "reason": clip["provenance_reason"]}
              for clip in video_clips
              if clip.get("provenance_degraded")
          ]
          if degraded:
              timeline["provenance"] = {"degraded": True, "degraded_clips": degraded}
              log(f"  ⚠️ 时间线 provenance 降级: {degraded[0]['reason']} ({len(degraded)} clip)")
          else:
              timeline["provenance"] = {"degraded": False}
          out = Path(work_dir) / "timeline.json"
          save_timeline(timeline, out)
          log(f"时间线模型: {out} ({len(timeline['tracks'])} 轨)")
          return timeline
      
    • visual_render.py 16.8 KB
      """Visual overlays, subtitle layout QC, masking, and video filter helpers."""
      
      import json
      import re
      from pathlib import Path
      
      from assemble_constants import (
          SUBTITLE_STYLE_REF_H,
          VISUAL_OVERLAYS,
          VISUAL_QC,
          _SUPPORTED_VISUAL_OVERLAY_TYPES,
      )
      from audio_automation import coalesce_duck_windows
      from audio_mix import _seg_place_window
      from lib import CONFIG
      from source_subtitles import (
          _combined_subtitle_entries,
          _original_gap_subtitle_entries,
          _source_subtitle_mask_policy,
      )
      from subtitles.core import (
          _measured_subtitle_band,
          _measured_subtitle_safe_area,
          _normalize_subtitle_text,
          _style_for_measured_subtitle_band,
          _subtitle_style_config,
      )
      
      _OVERFLOW_KINDS = {
          "max_lines_exceeded": "line_count",
          "safe_width_exceeded": "line_width",
          "safe_height_exceeded": "safe_area",
      }
      
      
      def _visual_text_units(text):
          """Approximate visual text width in em units for deterministic geometry QC."""
          units = 0.0
          for ch in text:
              if ch.isspace():
                  units += 0.35
              elif ord(ch) < 128:
                  units += 0.56
              else:
                  units += 1.0
          return units
      
      
      def _subtitle_layout_qc(entries, style, safe_area=None):
          """Machine-check subtitle safe-area/multiline/overflow facts for visual_qc.json."""
          play_x = int(style["play_res_x"])
          play_y = int(style["play_res_y"])
          margin_l = int(style["margin_l"])
          margin_r = int(style["margin_r"])
          margin_v = int(style["margin_v"])
          font_size = float(style["font_size"])
          max_lines = CONFIG["subtitle_max_lines"]
          if safe_area is None:
              safe_area = {
                  "x": margin_l,
                  "y": margin_v,
                  "width": max(1, play_x - margin_l - margin_r),
                  "height": max(1, play_y - 2 * margin_v),
                  "bottom_margin": margin_v,
              }
          usable_w = float(safe_area["width"])
          line_h = font_size * 1.25
          overflow_entries = []
          violations = []
          multi_line_entries = []
          max_observed_lines = 0
          entry_facts = []
          for i, entry in enumerate(entries):
              raw_text = _normalize_subtitle_text(entry["text"])
              lines = [ln for ln in re.split(r"(?:\\N|\n)+", raw_text) if ln != ""] or [""]
              line_count = len(lines)
              max_observed_lines = max(max_observed_lines, line_count)
              max_w = max(_visual_text_units(line) * font_size for line in lines)
              band_h = line_count * line_h + float(style["outline"]) * 2 + float(style["shadow"])
              overflow_reasons = []
              if line_count > max_lines:
                  overflow_reasons.append("max_lines_exceeded")
              if max_w > usable_w + 1e-6:
                  overflow_reasons.append("safe_width_exceeded")
              if band_h > safe_area["height"] + 1e-6:
                  overflow_reasons.append("safe_height_exceeded")
              fact = {
                  "index": i,
                  "start": round(float(entry["start"]), 3),
                  "end": round(float(entry["end"]), 3),
                  "line_count": line_count,
                  "max_line_width": round(max_w, 2),
                  "safe_width": round(usable_w, 2),
                  "band_height": round(band_h, 2),
                  "overflow": bool(overflow_reasons),
                  "overflow_reasons": overflow_reasons,
              }
              entry_facts.append(fact)
              if line_count > 1:
                  multi_line_entries.append(i)
              if overflow_reasons:
                  overflow_entries.append(fact)
                  violations.extend(
                      {"index": i, "kind": _OVERFLOW_KINDS[reason], "reason": reason}
                      for reason in overflow_reasons
                  )
          return {
              "enabled": CONFIG["burn_subtitles"],
              "renderer": "ass" if CONFIG["burn_subtitles"] else "sidecar_srt",
              "style": {
                  "font_size": int(font_size),
                  "max_chars": int(style["max_chars"]),
                  "max_lines": max_lines,
                  "play_res_x": play_x,
                  "play_res_y": play_y,
                  "alignment": int(style["alignment"]),
                  "margin_l": margin_l,
                  "margin_r": margin_r,
                  "margin_v": margin_v,
              },
              "safe_area": safe_area,
              "entries": len(entry_facts),
              "max_lines": max_observed_lines,
              "max_observed_lines": max_observed_lines,
              "multi_line": bool(multi_line_entries),
              "multi_line_entries": multi_line_entries,
              "overflow": bool(overflow_entries),
              "overflow_entries": overflow_entries,
              "violations": violations,
              "entry_facts": entry_facts,
          }
      
      
      def _load_visual_overlays(work_dir):
          """Return (overlays, source) from the canonical visual_overlays.json handoff."""
          path = Path(work_dir) / VISUAL_OVERLAYS
          if not path.exists():
              return [], {"present": False, "path": str(path)}
          try:
              data = json.loads(path.read_text(encoding="utf-8"))
          except json.JSONDecodeError as exc:
              raise ValueError(f"{path}: JSON 无效") from exc
          if (
              not isinstance(data, dict)
              or type(data.get("schema_version")) is not int
              or data["schema_version"] != 1
              or not isinstance(data.get("overlays"), list)
              or not all(isinstance(item, dict) for item in data["overlays"])
          ):
              raise ValueError(f"{path}: visual_overlays.json schema 无效")
          source = {
              "present": True,
              "path": str(path),
              "schema_version": 1,
          }
          return data["overlays"], source
      
      
      def _escape_drawtext_text(text):
          return (
              text
              .replace("\\", "\\\\")
              .replace(":", "\\:")
              .replace("'", "\\'")
              .replace("%", "\\%")
              .replace("\n", "\\n")
          )
      
      
      def _overlay_time_window(overlay, video_duration):
          start = float(overlay.get("start", 0.0))
          end = float(overlay.get("end", video_duration))
          return start, max(start, end)
      
      
      def _overlay_bbox(overlay, canvas, *, default_y):
          width = canvas["width"]
          height = canvas["height"]
          text = overlay["text"]
          font_size = int(overlay.get("font_size", max(18, round(height * 0.045))))
          lines = [ln for ln in text.splitlines() if ln.strip()] or [text]
          max_w = max(_visual_text_units(ln) * font_size for ln in lines)
          text_h = len(lines) * font_size * 1.25
          if overlay["type"] == "top_title":
              x = max(0.0, (width - max_w) / 2)
              y = float(overlay.get("y", default_y))
          else:
              # Fractions of the canvas in [0, 1] are normalized coordinates; larger values are pixels.
              x = float(overlay.get("x", 0.08))
              y = float(overlay.get("y", 0.25))
              if 0.0 <= x <= 1.0:
                  x *= width
              if 0.0 <= y <= 1.0:
                  y *= height
          return {
              "x": round(x, 2),
              "y": round(y, 2),
              "width": round(max_w, 2),
              "height": round(text_h, 2),
              "font_size": font_size,
              "line_count": len(lines),
              "overflow": x < 0 or y < 0 or x + max_w > width or y + text_h > height,
          }
      
      
      def _visual_overlay_filters(work_dir, canvas, video_duration):
          """Render the first-release canonical visual_overlays.json contract.
      
          Only two semantic renderers are supported: top_title and inline_label_or_callout.
          Unsupported types are QC-blocking and deliberately do not silently render.
          """
          overlays, source = _load_visual_overlays(work_dir)
          default_top_y = max(24, round(canvas["height"] * 0.05))
          filters = []
          facts = []
          unsupported = []
          overflow = []
          for idx, overlay in enumerate(overlays):
              typ = overlay.get("type")
              text = overlay.get("text", "").strip()
              if typ not in _SUPPORTED_VISUAL_OVERLAY_TYPES:
                  unsupported.append({"index": idx, "type": typ, "reason": "unsupported_overlay_type"})
                  continue
              if not text:
                  unsupported.append({"index": idx, "type": typ, "reason": "missing_text"})
                  continue
              start, end = _overlay_time_window(overlay, video_duration)
              bbox = _overlay_bbox(overlay, canvas, default_y=default_top_y)
              if bbox["overflow"]:
                  overflow.append({"index": idx, "type": typ, "bbox": bbox})
              font_size = bbox["font_size"]
              safe_text = _escape_drawtext_text(text)
              enable = f"between(t\\,{start:.3f}\\,{end:.3f})"
              if typ == "top_title":
                  filt = (
                      "drawtext="
                      f"{_drawtext_font_option()}text='{safe_text}':x=(w-text_w)/2:y={int(bbox['y'])}:"
                      f"fontsize={font_size}:fontcolor=white:borderw=2:bordercolor=black@0.85:"
                      f"box=1:boxcolor=black@0.35:boxborderw=12:enable='{enable}'"
                  )
              else:
                  filt = (
                      "drawtext="
                      f"{_drawtext_font_option()}text='{safe_text}':x={int(bbox['x'])}:y={int(bbox['y'])}:"
                      f"fontsize={font_size}:fontcolor=white:borderw=2:bordercolor=black@0.85:"
                      f"box=1:boxcolor=black@0.45:boxborderw=8:enable='{enable}'"
                  )
              filters.append(filt)
              facts.append({
                  "index": idx,
                  "type": typ,
                  "text_chars": len(text),
                  "start": round(start, 3),
                  "end": round(end, 3),
                  "bbox": bbox,
              })
          qc = {
              "source": source,
              "supported_types": sorted(_SUPPORTED_VISUAL_OVERLAY_TYPES),
              "present": source["present"],
              "count": len(overlays),
              "rendered": len(facts),
              "facts": facts,
              "unsupported": unsupported,
              "overflow": overflow,
          }
          return filters, qc
      
      
      def _build_visual_qc(tts_segments, work_dir, video_duration, canvas, *, overlay_qc=None, mask_filter=None):
          entries = _combined_subtitle_entries(tts_segments, work_dir, video_duration)
          style = _style_for_measured_subtitle_band(_subtitle_style_config(canvas), canvas)
          subtitle_layout = _subtitle_layout_qc(
              entries, style, safe_area=_measured_subtitle_safe_area(style, canvas)
          )
          mask = _source_subtitle_mask_policy(work_dir)
          mask.update({
              "ratio": min(0.5, CONFIG["source_subtitle_mask_ratio"]) if mask["active"] else None,
              "filter": "drawbox" if mask_filter else None,
              "opacity": CONFIG["subtitle_mask_opacity"],
              "timing": CONFIG["source_subtitle_mask_timing"],
              "subtitle_y_top": CONFIG["subtitle_y_top"],
              "subtitle_y_bot": CONFIG["subtitle_y_bot"],
          })
          if overlay_qc is None:
              overlay_qc = _visual_overlay_filters(work_dir, canvas, video_duration)[1]
          blocking_codes = []
          if mask["blocking"]:
              blocking_codes.append("mask_policy_not_explicit")
          if subtitle_layout["overflow"]:
              blocking_codes.append("subtitle_overflow")
          if overlay_qc["unsupported"]:
              blocking_codes.append("unsupported_visual_overlay")
          if overlay_qc["overflow"]:
              blocking_codes.append("visual_overlay_overflow")
          return {
              "schema_version": 1,
              "artifact": VISUAL_QC,
              "verdict": "FAIL" if blocking_codes else "PASS",
              "blocking": bool(blocking_codes),
              "blocking_codes": blocking_codes,
              "geometry": {
                  "canvas": {
                      "width": canvas["width"],
                      "height": canvas["height"],
                      "fps": canvas["fps"],
                  },
                  "storage": {
                      "width": canvas["storage_width"],
                      "height": canvas["storage_height"],
                  },
                  "rotation": canvas["rotation"],
                  "sample_aspect_ratio": canvas["sample_aspect_ratio"],
                  "display_aspect_ratio": canvas["display_aspect_ratio"],
              },
              "subtitles": subtitle_layout,
              "mask": mask,
              "overlays": overlay_qc,
              "summary": {
                  "subtitle_entries": subtitle_layout["entries"],
                  "subtitle_overflow": subtitle_layout["overflow"],
                  "subtitle_multi_line": subtitle_layout["multi_line"],
                  "mask_policy": mask["policy"],
                  "mask_active": mask["active"],
                  "overlay_rendered": overlay_qc["rendered"],
                  "overlay_unsupported": len(overlay_qc["unsupported"]),
              },
          }
      
      
      def _write_visual_qc(work_dir, qc):
          path = Path(work_dir) / VISUAL_QC
          path.write_text(json.dumps(qc, ensure_ascii=False, indent=2), encoding="utf-8")
          return path
      
      
      def _escape_subtitle_filter_path(path):
          """Escape a path for ffmpeg subtitle/ass video filter arguments."""
          text = str(path).replace("\\", "/")
          for raw, escaped in (
              ("\\", "\\\\"),
              (":", "\\:"),
              ("'", "\\'"),
              (",", "\\,"),
              ("[", "\\["),
              ("]", "\\]"),
          ):
              text = text.replace(raw, escaped)
          return text
      
      
      def _subtitle_burn_filter(subtitle_path):
          """Build the ffmpeg video filter used for hard-sub rendering.
      
          With SUBTITLE_FONT_FILE set, libass also loads the fonts in that file's directory so
          the ASS style's family name resolves to the declared file instead of a system font.
          """
          filt = f"subtitles=filename='{_escape_subtitle_filter_path(subtitle_path)}'"
          font_file = CONFIG["subtitle_font_file"]
          if font_file:
              fonts_dir = Path(font_file).expanduser().resolve().parent
              filt += f":fontsdir='{_escape_subtitle_filter_path(fonts_dir)}'"
          return filt
      
      
      def _drawtext_font_option():
          font_file = CONFIG["subtitle_font_file"]
          if not font_file:
              return ""
          return f"fontfile='{_escape_subtitle_filter_path(Path(font_file).expanduser().resolve())}':"
      
      
      def _output_downscale_filter(max_h):
          """Lanczos downscale that forces BOTH output dimensions even (libx264/yuv420p need it).
      
          -2 keeps the aspect ratio with an even width; 2*trunc(min(ih,H)/2) caps the height at H
          yet forces it even, so an odd OUTPUT_MAX_HEIGHT (e.g. 721) cannot produce an odd height
          that makes libx264 abort with an empty output file. 'min(ih,H)' only ever shrinks.
          """
          return f"scale=-2:'2*trunc(min(ih,{max_h})/2)':flags=lanczos"
      
      
      def _source_subtitle_mask_filter(canvas, work_dir, tts_segments, video_duration):
          """Return source-subtitle drawbox filters, optionally scoped to narration windows.
      
          Many source videos (e.g. 庆余年) ship hardcoded subtitles; without this the recap
          shows the original subs AND our narration subs stacked. Once masking is explicitly enabled,
          the enhanced default is a measured, translucent narration-only band; opacity and timing
          remain configurable.
          """
          policy = _source_subtitle_mask_policy(work_dir)
          if not policy["active"]:
              return None
          opacity = CONFIG["subtitle_mask_opacity"]
          timing = CONFIG["source_subtitle_mask_timing"]
          if timing not in {"all", "narration"}:
              raise ValueError(f"SOURCE_SUBTITLE_MASK_TIMING 必须是 all 或 narration,当前为 {timing!r}")
      
          band = _measured_subtitle_band(canvas)
          if band is not None:
              y_top, y_bot = band
              padding = CONFIG["subtitle_mask_padding"]
              mask_top = max(0, y_top - padding)
              mask_bot = min(canvas["height"], y_bot + padding)
              geometry = f"x=0:y={mask_top}:w=iw:h={mask_bot - mask_top}"
          else:
              # Our subtitle cues are one line. Keep the mask large enough for that line and its
              # margin, but never regress to the old two-line bar that hid ~23% of the image.
              style = _subtitle_style_config(canvas)
              play_res_y = float(style["play_res_y"])
              line_h = float(style["font_size"]) * 1.25
              pad = 10.0 * play_res_y / SUBTITLE_STYLE_REF_H
              sub_band = (float(style["margin_v"]) + line_h + pad) / play_res_y
              ratio = min(0.5, max(CONFIG["source_subtitle_mask_ratio"], sub_band))
              geometry = f"x=0:y=ih-ih*{ratio:.3f}:w=iw:h=ih*{ratio:.3f}"
      
          base = f"drawbox={geometry}:color=black@{opacity:.2f}:t=fill"
          filters = []
          if opacity > 0:
              if timing == "all":
                  filters.append(base)
              else:
                  windows = [
                      (start, end, 0.0)
                      for start, end in map(_seg_place_window, tts_segments)
                      if end > start
                  ]
                  # Avoid overlapping drawboxes: stacking two 60%-black masks would darken the
                  # overlap to 84%. Coalescing also keeps long filter chains smaller.
                  filters.extend(
                      f"{base}:enable='between(t,{start:.3f},{end:.3f})'"
                      for start, end, _ in coalesce_duck_windows(windows, bridge=0.001)
                  )
      
          # A translucent mask deliberately leaves the source glyphs visible. Whenever we burn a
          # replacement original-dialogue subtitle into a gap, cover that exact window opaquely first;
          # otherwise the source hard-sub and replacement text are stacked on top of each other.
          if not (timing == "all" and opacity >= 1.0 - 1e-9):
              replacement_windows = [
                  (entry["start"], entry["end"], 0.0)
                  for entry in _original_gap_subtitle_entries(tts_segments, work_dir, video_duration)
              ]
              opaque = f"drawbox={geometry}:color=black@1.00:t=fill"
              filters.extend(
                  f"{opaque}:enable='between(t,{start:.3f},{end:.3f})'"
                  for start, end, _ in coalesce_duck_windows(replacement_windows, bridge=0.001)
              )
          return ",".join(filters) if filters else None
      
  • SKILL.md 8.7 KB
    ---
    name: video-assemble
    user-invocable: false
    description: >
     合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录,
     最后做响度标准化。作为最终合成阶段使用。输入源视频、tts_meta.json 与旁白位置;
     输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。
    ---
    
    ## 1. 定位
    
    本技能负责最终合成:
    
    1. 把各段旁白音频放到视频时间线上。
    2. 在旁白窗口内压低原声,支持 fixed / sidechain / zone 模式。
    3. 根据旁白位置生成 `subtitles.srt`;默认同时生成并烧录 `subtitles.ass`,`--no-burn-subtitles` 可关闭。
    4. 可选把最终响度标准化到目标 LUFS。
    
    ## 2. 声音收尾契约
    
    合成阶段只实现创作决定,不凭空制造决定。Agent 在写旁白位置前,已在 `visual_audio_board.json` 为每个 beat 指定 `audio_owner`:
    
    - `original_dialogue`
    - `action_sound`
    - `ambience` / `music`
    - `silence`
    - `narration`
    
    因此,旁白间隙是主动选择,不是必须填满的空白。不要为了“更满”而加入通用 BGM、压住必须听见的台词或消除有意义的沉默。
    
    当前渲染器不解析 `visual_audio_board.json`;Agent 通过旁白时间、`overlaps_speech`、原声留白与现有混音参数落实这些决定。
    
    ## 3. 输入契约
    
    - `<video>`:源视频;cut 模式下为 `edited_source.mp4`。
    - `work_dir/tts_meta.json`:默认 `narration` 模式必需;配音阶段写出的 `{segments: [...]}`。每段包含 `audio_path`、时间、`pause_after_ms`、`overlaps_speech` 和用于混音/字幕的位置。显式 `source-mix` / `adopted-packet-copy` 模式不读取它。
    - 已采用的配音使用显式 `--tts-meta` 和 `--narration-adoption`:后者由调用方独立确认文字、请求的引擎/声线和速度策略,不能从待消费元数据自动“批准”出来。完整格式与记录边界见 `references/narration-adoption.md`。
    - 已采用的完整声音底轨与逐段配音可再传 `--audio-mix-adoption`;严格格式、48 kHz 声道矩阵和双 binding 事务见 `references/explicit-audio-mix.md`。
    
    下面的 `scripts/...` 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。
    
    ## 4. 运行命令
    
    ```bash
    python3 scripts/assemble.py <video> --work-dir <work_dir> \
      [--audio-mode narration|source-mix|adopted-packet-copy] [--audio-stream-index <N>] \
      [--tts-meta <tts_meta.json> --narration-adoption <narration_adoption.json>] \
      [--audio-mix-adoption <audio_mix_adoption.json>] \
      [--recap-stem <name>] [--output-dir <dir>] [--no-burn-subtitles] \
      [--subtitle-y-top <inclusive-y> --subtitle-y-bot <exclusive-y>] \
      [--source-video <orig.mp4>] [--export-jianying [--jianying-out <dir>]]
    ```
    
    ## 5. 输出契约
    
    - `recap_<stem>.mp4`:稳定的最终输出别名;每次运行覆盖更新。
    - `work_dir/output.mp4`:工作目录内成片。
    - `subtitles.srt`:旁白字幕;烧录时另有 `subtitles.ass`。
    - `timeline.json`:后端无关的多轨模型,包含视频、原声、旁白、BGM、字幕和 ducking 自动化。
    - `_placed_*.wav`:实际写入主混音的完整逐段旁白 PCM;时间线与剪映只引用这些文件。
    - `narration_input_binding.json`:旁白输入、转换、实际放置、旁白总轨和最终音轨的消费记录(路径、PCM 参数、packet 计数)。区分未采用与已绑定采用决定两种状态;不等于声线鉴定或听审。
    - `audio_mix_binding.json`:显式完整声音分支消费的画面时钟、底轨、48 kHz 配音放置、premaster、固定 master gain、最终 PCM/AAC 事实与 narration binding 路径的记录。
    - `assembly_manifest.json`:输入来源、cut 来源标识(路径、大小、mtime)、渲染设置与最终输出路径。
    - `assembly_qc.json`:旁白完整性、原声句末交接、时间线素材时长与交付质量的发布门禁。
    - 剪映草稿目录:仅 `--export-jianying` 时生成,包含 `draft_content.json`、`draft_info.json` 与 `draft_meta_info.json`。
    
    ## 6. 合成规则
    
    - 音频模式的处理与冻结语义见 `references/audio-modes.md`。默认仍为 `narration`;另外两种模式必须显式选择。
    - `--audio-mix-adoption` 只与显式 `--tts-meta`、`--narration-adoption` 同时使用;它保留 `narration` 模式名,但跳过旧速度/适配、原声 handoff、环境 BGM、duck、loudnorm 和 limiter。
    - 音频按轨道混合:原声、可选 BGM 与旁白各自独立。
    - 旁白不做任何容差裁尾;温和加速后仍放不下即 `no_safe_fit`。每段 `_placed_*.wav`
      必须与序列化后的时间线区间等长或更短,否则 `timeline_audio_mismatch` 阻断。
    - 已采用配音的 v1 合同只支持原速、禁止段内适速;不能让环境默认 1.15 倍速或旧缓存覆盖它。放不下就修订安排,不裁尾。严格运行使用新工作目录与新输出路径;输入/实际混音来源变动或 QC 失败时,不发布候选成片。没有采用文件的旧入口仍是兼容模式,不自动获得同等证据。
    - 原声在旁白结束后保持压低到下一可靠句末的 `pause_start`,只在实测停顿内渐强,
      于 `source_restore_at` 回满;无后续锚点时保持压低到时间线末端,而不是放出半句。
    - `--export-jianying` / `EXPORT_JIANYING=1` 可把 `timeline.json` 导出为可编辑草稿。cut 模式应传 `--source-video <orig>`,让草稿引用真实原片区间。
    - 剪映导出默认把视频、音频与图片复制到 `Resources/local/{video,audio,image}`,保持草稿可搬迁;`--jianying-no-bundle-media` 只适合原路径始终可访问的情况。
    - 重叠覆盖物会拆到编号轨道;非空目标目录不会覆盖,而会创建编号兄弟目录。
    - 常速、倒放、变换、富文本、转场、蒙版、LUT、绿幕复合草稿及显式特效轨道通过 timeline v2 扩展表达。需要素材包的功能只接受调用方合法提供的离线资源。
    - 剪映草稿引用未烧录的源视频,因此原片硬字幕仍会保留,必要时在剪映内另行遮罩。
    - 字幕外观可用 `SUBTITLE_FONT_SIZE`、`SUBTITLE_MARGIN_V`、`SUBTITLE_MAX_CHARS` 等控制。
    - `SUBTITLE_Y_TOP/BOT` 把 ASS 基线放到测得的原片字幕区域,坐标为半开 `[top, bot)`;显式遮罩策略下默认 `SUBTITLE_MASK_OPACITY=0.6`,`SOURCE_SUBTITLE_MASK_TIMING=narration`。
    - 原声在旁白间隙回到 `IDLE_ORIG_VOLUME`,旁白下压到 `SPEECH_DUCKING_VOLUME`;`DUCK_FADE_SECONDS` 控制过渡。还可配置 `DUCKING_MODE`、`ZONE_DUCKING_VOLUME`、`FINAL_LOUDNORM` 与 `TARGET_LUFS`。
    - 可通过 `BGM_PATH` 指定 BGM;它会循环到成片长度,并按 `BGM_VOLUME` / `BGM_DUCKING_VOLUME` 混音。不要在没有创作依据时设置通用 BGM。
    - 烧录字幕需要带 `subtitles` / libass 的 ffmpeg;合成阶段会预检并在缺失时明确失败。
    - 原声留白中的对白字幕优先读取 Agent 校对的 `original_subtitles.json`;否则保守映射 ASR。只有遮罩覆盖留白或用户字幕明确要求替换时才烧录原声对白,并用 `「」` 与旁白区分。
    
    ### 按原片区间准备声音,而不是整体压低旧成片
    
    已有多段原声取舍和独立 BGM 决定时,先用 `references/source-score.md` 的独立
    `source_score.py` 从原片声音流按精确帧区间重建原声轨、音乐轨及两者之和;它只输出
    声音底轨和来源回执。要与逐段已采用配音合成,再由调用方提供 `references/explicit-audio-mix.md`
    的严格 adoption;不要将底轨塞入旧入口自动 duck,也不要从含旧解说的成片取整条声音冒充干净原声。
    
    ## 7. 字幕与可选包装
    
    先锁定画面、剪点、旁白和混音,再投入字幕动画或边框包装;字幕样式不能掩盖叙事、剪点或声音问题。
    普通交付优先使用现有 ASS 路径;只有用户需要逐 cue 排版、动画或透明图层时,才用项目级代码渲染器,
    并按 `references/foreground-compose.md` 把它生成的 RGBA 序列叠到锁定母版。包装顺序与样帧抽检清单见
    `references/packaging.md`。
    
    ## 8. 能力边界
    
    - 不生成旁白文字,不合成 TTS,不重新转写视频。
    - 字幕烧录默认开启;关闭时不会重编码绘制字幕区域。
    
    显式输出轴字幕轨的独立合同、完整替换语义和当前边界见 `references/subtitle-track.md`。
    
    画面回原片重建后,若需保留另一文件中的已采用完整混音,先按
    `references/pair-media.md` 显式配对独立画面与音轨。配对只复制流,不补字幕或片名卡;
    后续字幕轨必须重新绑定配对后的容器与 `a:0`,不能继续沿用旧版本身份。
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related