video-assemble
合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录, 最后做响度标准化。作为最终合成阶段使用。输入源视频、tts_meta.json 与旁白位置; 输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。
Install
npx skills add https://github.com/zenstory-ai/oh-story-dsh/tree/main/packages/knowledge/video-recap/skills/video-assemble
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install zenstory-ai-oh-story-dsh@llmmart
git clone https://github.com/zenstory-ai/oh-story-dsh.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole zenstory-ai/oh-story-dsh collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
1. 定位
本技能负责最终合成:
- 把各段旁白音频放到视频时间线上。
- 在旁白窗口内压低原声,支持 fixed / sidechain / zone 模式。
- 根据旁白位置生成
subtitles.srt;默认同时生成并烧录subtitles.ass,--no-burn-subtitles可关闭。 - 可选把最终响度标准化到目标 LUFS。
2. 声音收尾契约
合成阶段只实现创作决定,不凭空制造决定。Agent 在写旁白位置前,已在 visual_audio_board.json 为每个 beat 指定 audio_owner:
original_dialogueaction_soundambience/musicsilencenarration
因此,旁白间隙是主动选择,不是必须填满的空白。不要为了“更满”而加入通用 BGM、压住必须听见的台词或消除有意义的沉默。
当前渲染器不解析 visual_audio_board.json;Agent 通过旁白时间、overlaps_speech、原声留白与现有混音参数落实这些决定。
3. 输入契约
<video>:源视频;cut 模式下为edited_source.mp4。work_dir/tts_meta.json:默认narration模式必需;配音阶段写出的{segments: [...]}。每段包含audio_path、时间、pause_after_ms、overlaps_speech和用于混音/字幕的位置。显式source-mix/adopted-packet-copy模式不读取它。- 已采用的配音使用显式
--tts-meta和--narration-adoption:后者由调用方独立确认文字、请求的引擎/声线和速度策略,不能从待消费元数据自动“批准”出来。完整格式与记录边界见references/narration-adoption.md。 - 已采用的完整声音底轨与逐段配音可再传
--audio-mix-adoption;严格格式、48 kHz 声道矩阵和双 binding 事务见references/explicit-audio-mix.md。
下面的 scripts/... 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。
4. 运行命令
python3 scripts/assemble.py <video> --work-dir <work_dir> \
[--audio-mode narration|source-mix|adopted-packet-copy] [--audio-stream-index <N>] \
[--tts-meta <tts_meta.json> --narration-adoption <narration_adoption.json>] \
[--audio-mix-adoption <audio_mix_adoption.json>] \
[--recap-stem <name>] [--output-dir <dir>] [--no-burn-subtitles] \
[--subtitle-y-top <inclusive-y> --subtitle-y-bot <exclusive-y>] \
[--source-video <orig.mp4>] [--export-jianying [--jianying-out <dir>]]
5. 输出契约
recap_<stem>.mp4:稳定的最终输出别名;每次运行覆盖更新。work_dir/output.mp4:工作目录内成片。subtitles.srt:旁白字幕;烧录时另有subtitles.ass。timeline.json:后端无关的多轨模型,包含视频、原声、旁白、BGM、字幕和 ducking 自动化。_placed_*.wav:实际写入主混音的完整逐段旁白 PCM;时间线与剪映只引用这些文件。narration_input_binding.json:旁白输入、转换、实际放置、旁白总轨和最终音轨的消费记录(路径、PCM 参数、packet 计数)。区分未采用与已绑定采用决定两种状态;不等于声线鉴定或听审。audio_mix_binding.json:显式完整声音分支消费的画面时钟、底轨、48 kHz 配音放置、premaster、固定 master gain、最终 PCM/AAC 事实与 narration binding 路径的记录。assembly_manifest.json:输入来源、cut 来源标识(路径、大小、mtime)、渲染设置与最终输出路径。assembly_qc.json:旁白完整性、原声句末交接、时间线素材时长与交付质量的发布门禁。- 剪映草稿目录:仅
--export-jianying时生成,包含draft_content.json、draft_info.json与draft_meta_info.json。
6. 合成规则
- 音频模式的处理与冻结语义见
references/audio-modes.md。默认仍为narration;另外两种模式必须显式选择。 --audio-mix-adoption只与显式--tts-meta、--narration-adoption同时使用;它保留narration模式名,但跳过旧速度/适配、原声 handoff、环境 BGM、duck、loudnorm 和 limiter。- 音频按轨道混合:原声、可选 BGM 与旁白各自独立。
- 旁白不做任何容差裁尾;温和加速后仍放不下即
no_safe_fit。每段_placed_*.wav必须与序列化后的时间线区间等长或更短,否则timeline_audio_mismatch阻断。 - 已采用配音的 v1 合同只支持原速、禁止段内适速;不能让环境默认 1.15 倍速或旧缓存覆盖它。放不下就修订安排,不裁尾。严格运行使用新工作目录与新输出路径;输入/实际混音来源变动或 QC 失败时,不发布候选成片。没有采用文件的旧入口仍是兼容模式,不自动获得同等证据。
- 原声在旁白结束后保持压低到下一可靠句末的
pause_start,只在实测停顿内渐强, 于source_restore_at回满;无后续锚点时保持压低到时间线末端,而不是放出半句。 --export-jianying/EXPORT_JIANYING=1可把timeline.json导出为可编辑草稿。cut 模式应传--source-video <orig>,让草稿引用真实原片区间。- 剪映导出默认把视频、音频与图片复制到
Resources/local/{video,audio,image},保持草稿可搬迁;--jianying-no-bundle-media只适合原路径始终可访问的情况。 - 重叠覆盖物会拆到编号轨道;非空目标目录不会覆盖,而会创建编号兄弟目录。
- 常速、倒放、变换、富文本、转场、蒙版、LUT、绿幕复合草稿及显式特效轨道通过 timeline v2 扩展表达。需要素材包的功能只接受调用方合法提供的离线资源。
- 剪映草稿引用未烧录的源视频,因此原片硬字幕仍会保留,必要时在剪映内另行遮罩。
- 字幕外观可用
SUBTITLE_FONT_SIZE、SUBTITLE_MARGIN_V、SUBTITLE_MAX_CHARS等控制。 SUBTITLE_Y_TOP/BOT把 ASS 基线放到测得的原片字幕区域,坐标为半开[top, bot);显式遮罩策略下默认SUBTITLE_MASK_OPACITY=0.6,SOURCE_SUBTITLE_MASK_TIMING=narration。- 原声在旁白间隙回到
IDLE_ORIG_VOLUME,旁白下压到SPEECH_DUCKING_VOLUME;DUCK_FADE_SECONDS控制过渡。还可配置DUCKING_MODE、ZONE_DUCKING_VOLUME、FINAL_LOUDNORM与TARGET_LUFS。 - 可通过
BGM_PATH指定 BGM;它会循环到成片长度,并按BGM_VOLUME/BGM_DUCKING_VOLUME混音。不要在没有创作依据时设置通用 BGM。 - 烧录字幕需要带
subtitles/ libass 的 ffmpeg;合成阶段会预检并在缺失时明确失败。 - 原声留白中的对白字幕优先读取 Agent 校对的
original_subtitles.json;否则保守映射 ASR。只有遮罩覆盖留白或用户字幕明确要求替换时才烧录原声对白,并用「」与旁白区分。
按原片区间准备声音,而不是整体压低旧成片
已有多段原声取舍和独立 BGM 决定时,先用 references/source-score.md 的独立
source_score.py 从原片声音流按精确帧区间重建原声轨、音乐轨及两者之和;它只输出
声音底轨和来源回执。要与逐段已采用配音合成,再由调用方提供 references/explicit-audio-mix.md
的严格 adoption;不要将底轨塞入旧入口自动 duck,也不要从含旧解说的成片取整条声音冒充干净原声。
7. 字幕与可选包装
先锁定画面、剪点、旁白和混音,再投入字幕动画或边框包装;字幕样式不能掩盖叙事、剪点或声音问题。
普通交付优先使用现有 ASS 路径;只有用户需要逐 cue 排版、动画或透明图层时,才用项目级代码渲染器,
并按 references/foreground-compose.md 把它生成的 RGBA 序列叠到锁定母版。包装顺序与样帧抽检清单见
references/packaging.md。
8. 能力边界
- 不生成旁白文字,不合成 TTS,不重新转写视频。
- 字幕烧录默认开启;关闭时不会重编码绘制字幕区域。
显式输出轴字幕轨的独立合同、完整替换语义和当前边界见 references/subtitle-track.md。
画面回原片重建后,若需保留另一文件中的已采用完整混音,先按
references/pair-media.md 显式配对独立画面与音轨。配对只复制流,不补字幕或片名卡;
后续字幕轨必须重新绑定配对后的容器与 a:0,不能继续沿用旧版本身份。
Files (oh-story-dsh)
-
references
-
jianying
-
empty_draft_meta_info.json 1.7 KB
{ "cloud_package_completed_time": "", "draft_cloud_capcut_purchase_info": "", "draft_cloud_last_action_download": false, "draft_cloud_materials": [], "draft_cloud_purchase_info": "", "draft_cloud_template_id": "", "draft_cloud_tutorial_info": "", "draft_cloud_videocut_purchase_info": "", "draft_cover": "", "draft_deeplink_url": "", "draft_enterprise_info": { "draft_enterprise_extra": "", "draft_enterprise_id": "", "draft_enterprise_name": "", "enterprise_material": [] }, "draft_fold_path": "", "draft_id": "", "draft_is_ai_packaging_used": false, "draft_is_ai_shorts": false, "draft_is_article_video_draft": false, "draft_is_from_deeplink": "false", "draft_is_invisible": false, "draft_materials": [ { "type": 0, "value": [] }, { "type": 1, "value": [] }, { "type": 2, "value": [] }, { "type": 3, "value": [] }, { "type": 6, "value": [] }, { "type": 7, "value": [] }, { "type": 8, "value": [] } ], "draft_materials_copied_info": [], "draft_name": "", "draft_new_version": "", "draft_removable_storage_device": "", "draft_root_path": "", "draft_segment_extra_info": [], "draft_timeline_materials_size_": 9734196, "draft_type": "", "tm_draft_cloud_completed": "", "tm_draft_cloud_modified": 0, "tm_draft_create": 1765203775977, "tm_draft_modified": 1765203775977, "tm_draft_removed": 0, "tm_duration": 249100000 } -
empty_jy_combination_segment.json 1.6 KB
{ "is_tone_modify": false, "enable_adjust_mask": false, "keyframe_refs": [], "enable_adjust": true, "enable_video_mask": true, "enable_hsl": false, "speed": 1, "clip": { "rotation": 0, "transform": { "y": 0, "x": 0 }, "flip": { "vertical": false, "horizontal": false }, "scale": { "y": 1, "x": 1 }, "alpha": 1 }, "intensifies_audio": false, "uniform_scale": { "on": true, "value": 1 }, "responsive_layout": { "vertical_pos_layout": 0, "target_follow": "", "enable": false, "horizontal_pos_layout": 0, "size_layout": 0 }, "enable_hsl_curves": true, "desc": "", "state": 0, "enable_lut": true, "reverse": false, "target_timerange": { "start": 0, "duration": 0 }, "cartoon": false, "enable_color_correct_adjust": false, "last_nonzero_volume": 1, "enable_smart_color_adjust": false, "source": "segmentsourcenormal", "render_timerange": { "start": 0, "duration": 0 }, "render_index": 0, "group_id": "", "enable_color_match_adjust": false, "enable_color_curves": true, "raw_segment_id": "", "volume": 1, "visible": true, "track_render_index": 0, "track_attribute": 0, "is_loop": false, "id": "", "enable_color_wheels": true, "extra_material_refs": [], "template_id": "", "template_scene": "default", "material_id": "", "common_keyframes": [], "digital_human_template_group_id": "", "color_correct_alg_result": "", "hdr_settings": { "mode": 1, "intensity": 1, "nits": 1000 }, "is_placeholder": false, "source_timerange": { "start": 0, "duration": 0 } } -
empty_jy_combination_video_material.json 2.2 KB
{ "category_name": "", "aigc_type": "none", "duration": 0, "category_id": "", "beauty_face_auto_preset_infos": [], "beauty_body_preset_id": "", "picture_from": "none", "cartoon_path": "", "intensifies_audio_path": "", "crop_ratio": "free", "source_platform": 0, "local_material_from": "", "media_path": "", "is_unified_beauty_mode": false, "width": 0, "video_algorithm": { "algorithms": [], "path": "", "story_video_modify_video_config": { "is_overwrite_last_video": false, "tracker_task_id": "", "task_id": "" }, "ai_background_configs": [], "ai_in_painting_config": [], "gameplay_configs": [] }, "type": "video", "height": 0, "material_id": "", "has_audio": true, "is_copyright": true, "source": 0, "intensifies_path": "", "crop_scale": 1, "live_photo_cover_path": "", "reverse_path": "", "reverse_intensifies_path": "", "picture_set_category_id": "", "crop": { "upper_right_x": 1, "lower_right_y": 1, "upper_right_y": 0, "upper_left_y": 0, "lower_left_y": 1, "lower_right_x": 1, "lower_left_x": 0, "upper_left_x": 0 }, "request_id": "", "local_material_id": "", "picture_set_category_name": "", "path": "", "is_text_edit_overdub": false, "check_flag": 62978047, "origin_material_id": "", "material_url": "", "beauty_face_auto_preset": { "preset_id": "", "scene": "", "name": "", "rate_map": "" }, "is_ai_generate_content": false, "extra_type_option": 2, "id": "", "aigc_item_id": "", "material_name": "复合片段1", "beauty_face_preset_infos": [], "team_id": "", "local_id": "", "matting": { "expansion": 0, "feather": 0, "flag": 0, "reverse": false, "strokes": [], "custom_matting_id": "", "has_use_quick_eraser": false, "has_use_quick_brush": false, "path": "", "interactiveTime": [], "enable_matting_stroke": false }, "formula_id": "", "has_sound_separated": false, "live_photo_timestamp": -1, "stable": { "stable_level": 0, "time_range": { "start": 0, "duration": 0 }, "matrix_path": "" }, "aigc_history_id": "" } -
empty_jy_draft.json 5.4 KB
{ "draft_cover_path": "", "type": "combination", "combination_id": "", "draft_config_path": "", "draft": { "version": 360000, "new_version": "111.0.0", "source": "default", "update_time": 0, "keyframe_graph_list": [], "duration": 86600000, "uneven_animation_template_info": { "order": "", "sub_template_info_list": [], "content": "", "composition": "" }, "tracks": [], "canvas_config": { "height": 1080, "ratio": "original", "width": 1920 }, "relationships": [], "last_modified_platform": { "app_version": "", "os_version": "", "os": "", "hard_disk_id": "", "device_id": "", "mac_address": "", "app_id": 0, "app_source": "" }, "color_space": -1, "static_cover_image_path": "", "render_index_track_mode_on": true, "keyframes": { "handwrites": [], "videos": [], "texts": [], "audios": [], "adjusts": [], "stickers": [], "filters": [], "effects": [] }, "platform": { "app_version": "9.6.0", "os_version": "15.1", "os": "mac", "hard_disk_id": "9d5cea4f22458a4e59d643d07162d324", "device_id": "0b02b9bd41947815545481af8a5bde46", "mac_address": "f8e784e2995a6aed158ef584d6245be8", "app_id": 3704, "app_source": "lv" }, "smart_ads_info": { "page_from": "", "routine": "", "draft_url": "" }, "path": "", "id": "AFCFCFAC-468E-46B6-92DC-8F2C66411CA3", "name": "", "lyrics_effects": [], "fps": 30, "config": { "use_float_render": false, "adjust_max_index": 1, "multi_language_list": [], "multi_language_main": "none", "lyrics_taskinfo": [], "lyrics_recognition_id": "", "system_font_list": [], "sticker_max_index": 1, "multi_language_mode": "none", "material_save_mode": 0, "attachment_info": [], "video_mute": false, "maintrack_adsorb": false, "multi_language_current": "none", "subtitle_sync": true, "subtitle_recognition_id": "", "combination_max_index": 1, "extract_audio_last_index": 1, "subtitle_taskinfo": [], "record_audio_last_index": 1, "lyrics_sync": true, "original_sound_last_index": 1 }, "create_time": 0, "draft_type": "video", "function_assistant_info": { "auto_adjust": false, "auto_caption": false, "enhande_voice": false, "auto_adjust_fixed": false, "fixed_rec_applied": false, "retouch_segid_list": [], "retouch": false, "enhande_voice_fixed": false, "enhance_quality_segid_list": [], "deflicker_segid_list": [], "auto_caption_template_id": "", "color_correction": false, "caption_opt": false, "audio_noise_segid_list": [], "auto_caption_segid_list": [], "normalize_loudness_fixed": false, "normalize_loudness_segid_list": [], "eye_correction": false, "enhance_quality_fixed": false, "color_correction_fixed_value": 50, "normalize_loudness_audio_denoise_segid_list": [], "video_noise_segid_list": [], "eye_correction_segid_list": [], "retouch_fixed": false, "auto_adjust_segid_list": [], "enhance_voice_segid_list": [], "smooth_slow_motion_fixed": false, "caption_opt_segid_list": [], "color_correction_segid_list": [], "smooth_slow_motion": false, "smart_segid_list": [], "smart_rec_applied": false, "color_correction_fixed": false, "fps": { "den": 1, "num": 0 }, "auto_adjust_fixed_value": 50, "normalize_loudness": false, "enhance_quality": false }, "materials": { "audio_fades": [], "images": [], "flowers": [], "smart_crops": [], "common_mask": [], "time_marks": [], "shapes": [], "vocal_beautifys": [], "plugin_effects": [], "audios": [], "smart_relights": [], "video_radius": [], "vocal_separations": [], "video_effects": [], "video_shadows": [], "videos": [], "manual_deformations": [], "video_strokes": [], "audio_effects": [], "realtime_denoises": [], "handwrites": [], "texts": [], "text_templates": [], "hsl_curves": [], "tail_leaders": [], "manual_beautys": [], "material_colors": [], "multi_language_refs": [], "material_animations": [], "chromas": [], "speeds": [], "audio_pannings": [], "sound_channel_mappings": [], "ai_translates": [], "video_trackings": [], "primary_color_wheels": [], "placeholders": [], "stickers": [], "placeholder_infos": [], "beats": [], "audio_pitch_shifts": [], "color_curves": [], "green_screens": [], "drafts": [], "digital_humans": [], "effects": [], "digital_human_model_dressing": [], "audio_balances": [], "hsl": [], "log_color_wheels": [], "canvases": [], "loudnesses": [], "transitions": [], "audio_track_indexes": [] }, "is_drop_frame_timecode": false, "free_render_index_mode_on": false }, "category_id": "", "precompile_combination": false, "combination_type": "none", "formula_id": "", "name": "", "id": "", "draft_file_path": "", "category_name": "" } -
empty_jy_material_video.json 1.8 KB
{ "aigc_history_id": "", "aigc_item_id": "", "aigc_type": "none", "audio_fade": null, "cartoon_path": "", "category_id": "", "category_name": "", "check_flag": 63487, "crop": { "lower_left_x": 0.0, "lower_left_y": 1.0, "lower_right_x": 1.0, "lower_right_y": 1.0, "upper_left_x": 0.0, "upper_left_y": 0.0, "upper_right_x": 1.0, "upper_right_y": 0.0 }, "crop_ratio": "free", "crop_scale": 1.0, "duration": 0, "extra_type_option": 0, "formula_id": "", "freeze": null, "has_audio": true, "height": 1920, "id": "", "intensifies_audio_path": "", "intensifies_path": "", "is_ai_generate_content": false, "is_copyright": false, "is_text_edit_overdub": false, "is_unified_beauty_mode": false, "local_id": "", "local_material_id": "", "material_id": "", "material_name": "", "material_url": "", "matting": { "flag": 0, "has_use_quick_brush": false, "has_use_quick_eraser": false, "interactiveTime": [], "path": "", "strokes": [] }, "media_path": "", "object_locked": null, "origin_material_id": "", "path": "", "picture_from": "none", "picture_set_category_id": "", "picture_set_category_name": "", "request_id": "", "reverse_intensifies_path": "", "reverse_path": "", "smart_motion": null, "source": 0, "source_platform": 0, "stable": { "matrix_path": "", "stable_level": 0, "time_range": { "duration": 0, "start": 0 } }, "team_id": "", "type": "video", "video_algorithm": { "algorithms": [], "complement_frame_config": null, "deflicker": null, "gameplay_configs": [], "motion_blur_config": null, "noise_reduction": null, "path": "", "quality_enhance": null, "time_range": null }, "width": 1080 } -
empty_jy_meta_material_value.json 358 B
{ "duration": 0, "height": 0, "id": "", "md5": "", "metetype": "", "type": 0, "width": 1080, "create_time": 0, "extra_info": "", "file_Path": "", "import_time": 0, "import_time_ms": 0, "item_source": 1, "roughcut_time_range": { "duration": 0, "start": 0 }, "sub_time_range": { "duration": -1, "start": -1 } } -
empty_jy_project_info.json 3.6 KB
{ "canvas_config": { "height": 1080, "ratio": "original", "width": 1920 }, "color_space": 0, "config": { "adjust_max_index": 1, "attachment_info": [], "combination_max_index": 1, "export_range": null, "extract_audio_last_index": 1, "lyrics_recognition_id": "", "lyrics_sync": true, "lyrics_taskinfo": [], "maintrack_adsorb": true, "material_save_mode": 0, "multi_language_current": "none", "multi_language_list": [], "multi_language_main": "none", "multi_language_mode": "none", "original_sound_last_index": 1, "record_audio_last_index": 1, "sticker_max_index": 1, "subtitle_keywords_config": null, "subtitle_recognition_id": "", "subtitle_sync": true, "subtitle_taskinfo": [], "system_font_list": [], "video_mute": false, "zoom_info_params": null }, "cover": null, "create_time": 0, "duration": 0, "extra_info": null, "fps": 30.0, "free_render_index_mode_on": false, "group_container": null, "id": "", "keyframe_graph_list": [], "keyframes": { "adjusts": [], "audios": [], "effects": [], "filters": [], "handwrites": [], "stickers": [], "texts": [], "videos": [] }, "last_modified_platform": { "app_id": 3704, "app_source": "lv", "app_version": "5.9.5-beta1", "device_id": "0b02b9bd41947815545481af8a5bde46", "hard_disk_id": "9d5cea4f22458a4e59d643d07162d324", "mac_address": "046f91694f9a5ed2c112de702af7131b", "os": "mac", "os_version": "15.1" }, "materials": { "ai_translates": [], "audio_balances": [], "audio_effects": [], "audio_fades": [], "audio_track_indexes": [], "audios": [], "beats": [], "canvases": [], "chromas": [], "color_curves": [], "digital_humans": [], "drafts": [], "effects": [], "flowers": [], "green_screens": [], "handwrites": [], "hsl": [], "images": [], "log_color_wheels": [], "loudnesses": [], "manual_deformations": [], "masks": [], "common_mask": [], "material_animations": [], "material_colors": [], "multi_language_refs": [], "placeholders": [], "plugin_effects": [], "primary_color_wheels": [], "realtime_denoises": [], "shapes": [], "smart_crops": [], "smart_relights": [], "sound_channel_mappings": [], "speeds": [], "stickers": [], "tail_leaders": [], "text_templates": [], "texts": [], "time_marks": [], "transitions": [], "video_effects": [], "video_trackings": [], "videos": [], "vocal_beautifys": [], "vocal_separations": [] }, "mutable_config": null, "name": "", "new_version": "111.0.0", "platform": { "app_id": 3704, "app_source": "lv", "app_version": "5.9.5-beta1", "device_id": "0b02b9bd41947815545481af8a5bde46", "hard_disk_id": "46f90855e144229994b84640e2c31384", "mac_address": "046f91694f9a5ed2c112de702af7131b", "os": "mac", "os_version": "15.1" }, "relationships": [], "render_index_track_mode_on": false, "retouch_cover": null, "source": "default", "static_cover_image_path": "", "time_marks": null, "tracks": [], "update_time": 0, "version": 360000 } -
empty_jy_segment.json 1.3 KB
{ "caption_info": null, "cartoon": false, "clip": { "alpha": 1.0, "flip": { "horizontal": false, "vertical": false }, "rotation": 0.0, "scale": { "x": 1.0, "y": 1.0 }, "transform": { "x": 0.0, "y": 0.0 } }, "common_keyframes": [], "enable_adjust": true, "enable_color_correct_adjust": false, "enable_color_curves": true, "enable_color_match_adjust": false, "enable_color_wheels": true, "enable_lut": true, "enable_smart_color_adjust": false, "extra_material_refs": [], "group_id": "", "hdr_settings": { "intensity": 1.0, "mode": 1, "nits": 1000 }, "id": "", "intensifies_audio": false, "is_placeholder": false, "is_tone_modify": false, "keyframe_refs": [], "last_nonzero_volume": 1.0, "material_id": "", "render_index": 0, "responsive_layout": { "enable": false, "horizontal_pos_layout": 0, "size_layout": 0, "target_follow": "", "vertical_pos_layout": 0 }, "reverse": false, "source_timerange": { "duration": 0, "start": 0 }, "speed": 1.0, "target_timerange": { "duration": 0, "start": 0 }, "template_id": "", "template_scene": "default", "track_attribute": 0, "track_render_index": 0, "uniform_scale": { "on": true, "value": 1.0 }, "visible": true, "volume": 1.0 } -
empty_jy_text_styles.json 440 B
{ "styles": [ { "fill": { "alpha": 1.0, "content": { "render_type": "solid", "solid": { "alpha": 1.0, "color": [ 1.0, 1.0, 1.0 ] } } }, "font": { "id": "", "path": "" }, "range": [ 0, 4 ], "size": 15.0 } ], "text": "" } -
empty_yj_material_audio.json 865 B
{ "aigc_history_id": "", "aigc_item_id": "", "app_id": 0, "category_id": "", "category_name": "", "check_flag": 1, "copyright_limit_type": "none", "duration": 0, "effect_id": "", "formula_id": "", "id": "", "intensifies_path": "", "is_ai_clone_tone": false, "is_text_edit_overdub": false, "is_ugc": false, "local_material_id": "", "music_id": "", "name": "", "path": "", "query": "", "request_id": "", "resource_id": "", "search_id": "", "source_from": "", "source_platform": 0, "team_id": "", "text_id": "", "tone_category_id": "", "tone_category_name": "", "tone_effect_id": "", "tone_effect_name": "", "tone_platform": "", "tone_second_category_id": "", "tone_second_category_name": "", "tone_speaker": "", "tone_type": "", "type": "extract_music", "video_id": "", "wave_points": [] } -
empty_yj_material_text.json 2.4 KB
{ "add_type": 0, "alignment": 1, "background_alpha": 1.0, "background_color": "", "background_height": 0.0, "background_horizontal_offset": 0.0, "background_round_radius": 0.0, "background_style": 0, "background_vertical_offset": 0.0, "background_width": 0.0, "base_content": "", "bold_width": 0.0, "border_alpha": 1.0, "border_color": "", "border_width": 0.08, "caption_template_info": { "category_id": "", "category_name": "", "effect_id": "", "is_new": false, "path": "", "request_id": "", "resource_id": "", "resource_name": "", "source_platform": 0 }, "check_flag": 7, "combo_info": { "text_templates": [] }, "content": "", "fixed_height": -1.0, "fixed_width": -1.0, "font_category_id": "", "font_category_name": "", "font_id": "", "font_name": "", "font_path": "", "font_resource_id": "", "font_size": 15.0, "font_source_platform": 0, "font_team_id": "", "font_title": "none", "font_url": "", "fonts": [], "force_apply_line_max_width": false, "global_alpha": 1.0, "group_id": "", "has_shadow": false, "id": "", "initial_scale": 1.0, "inner_padding": -1.0, "is_rich_text": false, "italic_degree": 0, "ktv_color": "", "language": "", "layer_weight": 1, "letter_spacing": 0.0, "line_feed": 1, "line_max_width": 0.82, "line_spacing": 0.02, "multi_language_current": "none", "name": "", "original_size": [], "preset_category": "", "preset_category_id": "", "preset_has_set_alignment": false, "preset_id": "", "preset_index": 0, "preset_name": "", "recognize_task_id": "", "recognize_type": 0, "relevance_segment": [], "shadow_alpha": 0.9, "shadow_angle": -45.0, "shadow_color": "", "shadow_distance": 5.0, "shadow_point": { "x": 0.0, "y": 0.0 }, "shadow_smoothing": 0.45, "shape_clip_x": false, "shape_clip_y": false, "source_from": "", "style_name": "", "sub_type": 0, "subtitle_keywords": null, "subtitle_template_original_fontsize": 0.0, "text_alpha": 1.0, "text_color": "#FFFFFF", "text_curve": null, "text_preset_resource_id": "", "text_size": 30, "text_to_audio_ids": [], "tts_auto_update": false, "type": "text", "typesetting": 0, "underline": false, "underline_offset": 0.22, "underline_width": 0.05, "use_effect_default_color": true, "words": { "end_time": [], "start_time": [], "text": [] } } -
LICENSE.duo-video 1 KB · in bundle
-
SOURCE.md 728 B
# JianYing protocol templates These JSON protocol templates are pinned to [`duoec/duo-video`](https://github.com/duoec/duo-video) commit `ef4eb46c823910553f901649f2f13fd7575e748f`, under its MIT license. They are data/schema baselines, not executable upstream code. Runtime builders deep-copy the templates and replace authored values such as IDs, paths, timings, canvas dimensions, and resource configuration. The required copyright and permission notice is preserved in `LICENSE.duo-video` in this directory. The exporter never uses duo-video's embedded example credentials. Resource-ID features accept an offline `material` or `resource_config` payload so official resource packages can be supplied legally by the caller.
-
-
audio-modes.md 3.5 KB
# Assembly audio modes `assemble.py` keeps `narration` as its API and CLI default. Non-narration behavior is opt-in with `--audio-mode`; `--audio-stream-index N` is the zero-based audio ordinal used by FFmpeg's `0:a:N` selector. ## `narration` - Requires non-empty `tts_meta.json` segments, as before. - Builds/places narration WAV, performs source ducking and optional BGM mix, then applies the configured final loudness/limiter stage and AAC encoding. - Missing source audio may use the existing synthetic-silence fallback. - With the additional strict `--audio-mix-adoption`, narration instead consumes an adopted prepared bed plus complete direct-to-48-kHz narration placements. This is still narration mode, but it bypasses ambient BGM, ducking, speed/fit, loudnorm, and limiter operations. See `explicit-audio-mix.md`. ## `source-mix` - Does not read or require `tts_meta.json` and never creates narration audio. - Uses the selected input audio stream, applies `IDLE_ORIG_VOLUME`, mixes a declared `BGM_PATH` when present, and runs the configured final loudness/limiter stage before AAC encoding. - A declared but missing `BGM_PATH` is an error in this new mode (the legacy narration-mode warn-and-skip behavior remains unchanged). - This is a processed mix, not frozen audio. `assembly_qc.json` reports the operations that actually ran. ## `adopted-packet-copy` - First implementation supports a selected AAC stream from `<video>` only. - Rejects TTS segments, explicit `--tts-meta`, any configured `BGM_PATH`, a missing/wrong stream, unsupported codec, or audio/picture interval mismatch. - Maps that stream with `-c:a copy`. It does not build narration, mix, duck, normalize, limit, resample, change tempo, or add silence. - It deliberately omits `-t` and `-shortest`, preserving AAC priming and tail packets. After rendering, the output is probed and compared with the input: decoder parameters, packet count, total payload bytes, the packet-clock span, and every packet's size and PTS/DTS/duration converted to rational time must match. Packet side data also must match, including AAC skip-sample/ discard-padding values and their reason fields. - QC/manifest evidence records selected stream, absolute stream index, codec, time base, sample rate, channels/layout, packet count, payload bytes, packet details, and actual operation flags. The adopted-copy settings payload excludes narration/mix/loudness defaults, so irrelevant ambient settings do not invalidate a frozen-audio render. Video filters and video re-encoding remain allowed; packet verification must still pass afterward. These modes do not create a commercial quality profile, automatic speech alignment, listening approval, or release approval. An explicit subtitle track is only a version/media-bound timing declaration; its evidence labels and `NOT_CHECKED` acoustic/listening status remain authoritative. Matching packets and decoder parameters show the encoded audio was copied; they do not prove perceptual quality or that a person listened to the result. `timeline.json` represents either non-narration mode as one complete clip from the current input video and records the selected stream. Adopted copy uses gain 1 with no automation. Source-mix records processed/non-frozen delivery and is not claimed to be reconstructable by an editor. The optional JianYing exporter currently supports only selected stream 0; requesting export with another stream fails instead of silently substituting the default stream. -
explicit-audio-mix.md 3.7 KB
# Explicit adopted full-sound mix Use this path only when the picture, a completed `prepared_bed_receipt`, an exact narration adoption, narration placements/gains, and one fixed master gain have already been independently selected. It is still `audio_mode=narration`, but it bypasses the legacy narration timing, ducking, ambient BGM, loudness-normalization, and limiter path. ```bash python3 scripts/assemble.py picture.mp4 --work-dir NEW_WORK \ --tts-meta /local/tts_meta.json \ --narration-adoption /local/narration_adoption.json \ --audio-mix-adoption /local/audio_mix_adoption.json ``` All strict output paths, including the CLI delivery alias, must be new. The CLI copies to a hidden delivery stage and creates the alias with an exclusive atomic link, so a concurrent or existing file is never overwritten. ## Adoption schema v1 ```json { "artifact": "audio_mix_adoption", "schema_version": 1, "prepared_receipt": {"path": "/local/prepared_bed_receipt.json"}, "format": {"sample_rate": 48000, "channels": 2, "total_samples": 1856000}, "segments": [ {"index": 0, "output_start_sample": 366000, "gain": 0.4251421093940735} ], "master_gain_db": 0.75 } ``` Unknown or missing fields fail; legacy `*_sha256` keys are ignored. The format must equal the actual zero-origin picture frame clock at 48 kHz and the prepared receipt. The receipt, all three float PCM beds (format, sample count, finite PCM), the picture, the narration adoption file and the ordered segment indices are checked before narration snapshots are written. ## Exact audio operations Each narration snapshot is decoded directly and completely to `pcm_f32le`, 48 kHz stereo. Version 1 uses one fixed channel matrix, recorded per segment as `mono_equal_power` or `stereo_identity`: - mono is panned to left and right at `1/sqrt(2)` per channel; - stereo preserves independent left and right channels at unit gain; - inputs with more than two channels are rejected. There is no intermediate 44.1 kHz mono placement, speed change, fade, trim, tail pad, normalization, or automatic fit. Every complete converted WAV must fit its adopted integer `output_start_sample`; placements may not overlap. Adopted per-segment gains form a float `voice_bus.wav`. The producer's `prepared_bed.wav` and voice bus form a float premaster, then the sole adopted `master_gain_db` forms the float master consumed by the final AAC encode. The picture's old audio is never a mixer input. ## Records and publication `narration_input_binding.json` records the actual direct 48 kHz placed files and voice bus, marking the bus `CONSUMED_BY_EXPLICIT_MIX`. `audio_mix_binding.json` records the mix adoption path, picture path and clock, prepared receipt path and stems, converted placements (path, channel matrix, PCM facts), voice bus, premaster, master, the narration binding path, the final decoded PCM facts, the final AAC decoder/packet count/payload bytes, and the actual output picture decoder/frame clock. The output picture must preserve fps, frame count, zero start, and duration. A video-copy path reports `packet_identity: EXACT` when decoder parameters and packet sizes/timestamps are unchanged; an allowed visual re-encode reports `REENCODED_CLOCK_MATCH`. Timeline, settings, QC, and manifest identify the explicit path rather than reporting ambient ducking/BGM/loudnorm operations. The candidate render is written to a hidden file. Both bindings are written, QC must pass, the media is published, and a second QC runs against the published path. Any render, binding, or QC failure removes the candidate, final-named media and the bindings written by this run instead of overwriting an older success. Diagnostic/staging audio may remain in the unique work directory. -
foreground-compose.md 5.1 KB
# Caller-rendered foreground composition `compose_foreground.py` performs one narrow operation: it overlays a caller-rendered, full-canvas RGBA PNG sequence on an already locked H264/CFR/AAC base, optionally replaces a declared tail with an explicit endcard asset, and copies `a:0` without decoding or rewriting it. The PNG pixels are the source schema. This is not a CSS, font, text, subtitle, logo, animation, or plugin renderer. ## Command ```bash python3 scripts/compose_foreground.py foreground_plan.json \ --output-dir a-new-directory [--plan-only] ``` The output directory must not exist. A normal run publishes `foreground.mp4` only after the staged media's frame clock, canvas/color metadata, AAC decoder/packets/PTS/ side data, and a full audio/video decode pass verification. A failure retains `foreground_run.json` and available logs but removes staged/final media. `--plan-only` validates the complete plan and inputs and writes only `foreground_run.json`; it does not render media or update any current pointer. ## Strict plan schema v1 ```json { "artifact": "foreground_compose_plan", "schema_version": 1, "base": {"path": "/local/base.mp4"}, "video": {"fps": "30/1", "width": 1280, "height": 720, "total_frames": 300}, "foreground": { "directory": "/local/foreground_sequence", "pattern": "frame_%06d.png", "start_frame": 0, "end_frame": 270 }, "endcard": { "kind": "sequence", "directory": "/local/endcard_sequence", "pattern": "frame_%06d.png", "start_frame": 270, "end_frame": 300 }, "producer_receipt": {"path": "/local/producer_receipt.json"} } ``` `endcard` is an exact tagged union. The alternative static form is: ```json { "kind": "still", "path": "/local/endcard.png", "start_frame": 270, "end_frame": 300 } ``` When the foreground itself covers the entire declared frame clock, the only accepted no-tail form is exactly: ```json {"kind": "none"} ``` It accepts no additional fields. It is invalid when the foreground reserves any tail; there is no implicit filler, repeated last frame, or generated endcard. Use `still` only when a genuinely static full-canvas endcard is intended. A static PNG does not reconstruct a delivered fade or other animated endcard; supply the caller-rendered `sequence` form for that case. Both sequences use **local zero-based filenames** regardless of their output range. Thus the endcard example contains `frame_000000.png` through `frame_000029.png`, mapped to output frames `[270,300)`. The foreground must start at output frame zero. Its end must equal the endcard start, and the endcard end must equal the independently probed base frame count. With `kind: "none"`, the foreground end must instead equal that full base frame count. Bounds are half-open. The only accepted filename pattern is the literal `frame_%06d.png`. The directory must contain exactly the expected contiguous entries—no missing or extra files. Every actual image must probe as PNG, `rgba`, and the exact declared full canvas. Repeated files and symlinks are allowed so transparent holds need not duplicate storage. The run report records each sequence's directory and validated `frame_count`. Legacy `sha256` / `ordered_sha256` keys in a plan are ignored. ## Base and output invariants The first version accepts one selected `v:0` H264, zero-origin CFR, YUV420P, BT.709/TV-range base and adopted AAC `a:0`. Declared fps, canvas, and frame count are checked independently against actual decoded frame PTS. Picture and audio intervals must satisfy the narrow `pair_media` timing contract. Composition explicitly overlays in RGB, converts to BT.709 limited-range YUV420P, limits the final video filter to the declared half-open frame clock, and re-encodes only video with fixed `libx264 -preset fast -crf 18` settings. A no-tail run uses only the base and foreground inputs; it does not synthesize or pad a tail. The filter EOF prevents still-image input loops without imposing an output-level video frame cap that could stop accepted trailing AAC packets from being copied. The run report records that encoder, preset, CRF, pixel format, color contract, and audio-copy mode explicitly. H.264 CRF encoding is lossy: preserved transparent regions are visually checked against tolerances, not claimed pixel-identical to decoded base pixels. The compositor does not use `-r`, `-shortest`, or `-t`. Audio is mapped from base `a:0` with `-c:a copy`; output verification requires matching decoder parameters, packet count, payload bytes, PTS/DTS/duration/size, and side data. ## Provenance and review boundary The producer receipt is a caller artifact that must exist when the plan is read. Its declarations may describe roles or content decisions such as title/note/brand, `dialogues=[]`, or hidden markers. Core records those declarations as `DECLARED_NOT_CHECKED`; it does not infer semantic truth from pixels or claim that dialogue/subtitle duplication is absent. A separate producer/reviewer must establish that evidence. Every run reports `direct_listening` and `normal_speed_review` as `NOT_CHECKED` and `release_approved` as false. Pixel validation and deterministic rendering are not commercial-quality or release approval. -
narration-adoption.md 4.7 KB
# Narration adoption and the consumed-input record Narration assembly accepts legacy `tts_meta.json`, but a current file by itself is not an adoption decision. Use an explicit adoption document when the selected spoken text, requested provider/voice, and tempo policy must be bound to the actual final mix. ```bash python3 scripts/assemble.py input.mp4 --work-dir work \ --tts-meta /local/tts_meta.json \ --narration-adoption /local/narration_adoption.json ``` `--narration-adoption` is narration-only and requires an explicitly supplied `--tts-meta`; it never adopts an ambient work-directory file automatically. ## Strict v1 schema ```json { "artifact": "narration_adoption", "schema_version": 1, "segments": [ { "index": 0, "spoken_text": "exact selected words", "requested_provider": "caller-selected-provider", "requested_voice": "caller-selected-voice" } ], "tempo_policy": { "global_atempo": 1.0, "bounded_segment_fit": false, "segment_tempo_max": 1.0, "cumulative_tempo_max": 1.0, "cumulative_tempo_hard_max": 1.0 } } ``` That policy is the conservative default. An adoption may declare its own values and they are honoured: `global_atempo` must be positive and no greater than `cumulative_tempo_hard_max`, `segment_tempo_max` and `cumulative_tempo_max` must be at least `1.0`, `cumulative_tempo_hard_max` must not be below `cumulative_tempo_max`, and `bounded_segment_fit` must be a boolean. Malformed or out-of-bounds policies fail; ambient defaults such as a global 1.15 speed never override an adoption. With `bounded_segment_fit` false, an adopted segment that does not fit its authored window blocks without time trimming or bounded fit. The adoption's ordered segment indices and spoken text must match both the segment list in the supplied `tts_meta.json` and the in-memory list passed to `assemble_video`; those two lists must be equal to each other. Unknown fields fail; legacy `tts_meta_sha256` / `processed_wav_sha256` keys are ignored. The consumer never creates an adoption from metadata that it is about to consume. ## Validation and snapshots Every adopted `audio_path` must exist before any media is probed or rendered. The adopted inputs are then copied to per-run snapshots under `work/.narration_input_snapshots/`; the original files are never modified, and the render reads only the snapshots. Every conversion, placed WAV, and the completed narration bus is recorded with its path and probed PCM facts before the final FFmpeg command runs. On the legacy narration-mix path, Python's standard WAV reader cannot open every valid post-processed WAV encoding. Noncanonical input, including `pcm_f32le`, 48 kHz, or stereo WAV, is explicitly decoded by the existing FFmpeg executable to 44.1 kHz mono PCM16 before placement. The conversion path and actual PCM facts are recorded; the consumer does not claim that converted samples equal the original samples. The explicit full-sound path described in `explicit-audio-mix.md` deliberately does not use that 44.1 kHz mono conversion. It consumes the snapshots directly as complete 48 kHz float placements, preserving native stereo channels. An adopted render writes to a hidden candidate path. The binding is written, QC runs against the candidate, the candidate is renamed to the final path, and QC runs once more against the published file. A blocking QC removes the candidate and the binding instead of publishing; the run records the final output as absent rather than claiming a missing file was published. Non-final work or diagnostic artifacts may remain for investigation. ## Binding report After successful assembly, `work/narration_input_binding.json` records: - original input path and snapshot path; - any explicit conversion path and PCM parameters; - each complete placed WAV path and PCM parameters; - `narration.wav` path and PCM parameters; - the final output path and its encoded audio-stream facts (decoder parameters, packet count, payload bytes, start time, duration); - the adoption path, its `tts_meta` path, and the exact tempo policy when supplied. The report uses one of two identity statuses: - `UNADOPTED`: no adoption was supplied; the legacy inputs were consumed as given; - `BOUND_TO_ADOPTION`: the adoption's selection and tempo policy were carried through the final mix. `BOUND_TO_ADOPTION` records what was consumed, not provider truth, acoustic speaker identity, direct listening, naturalness, or release approval; voice authentication and direct listening stay `NOT_CHECKED`. QC, manifest, and settings records reference a binding only while its recorded final output path still exists; source-mix and adopted-packet-copy modes do not reuse narration binding evidence. -
packaging.md 2.8 KB
# 字幕与包装的推荐顺序 先锁定画面、剪点、旁白和混音,再投入字幕动画或边框包装。字幕样式不能掩盖叙事、剪点或声音问题。 普通交付优先使用现有 ASS 路径。只有用户需要更精细的逐 cue 排版、动画或透明图层时,才使用 Remotion 或其他代码渲染器作为**项目级可选实现**,不要把特定框架、字体、颜色或黄字写成核心依赖。 ## 推荐顺序 1. 输出无包装的锁定母版,确认音画内容不再变化。 2. 从实际 TTS / 时间线生成 captions;TTS 块保持连续思路,字幕可以按阅读宽度拆 cue。 3. 先抽检开头、亮背景、暗背景、人物近景和长字幕样帧,确定字号、安全区、描边、阴影及是否需要底板。 4. 渲染完整透明字幕层,再 overlay 到锁定母版;合成后确认音轨未被意外改写。 5. 完整播放实际最终文件,并复查字幕遮脸、跳字、断行、首尾帧和边界处残影。 已有项目级渲染器能输出精确包装时,使用 `foreground-compose.md` 将它生成的 RGBA 序列叠到锁定母版, 而不是用通用白字黑框近似品牌样式。该操作保留实际帧钟和 AAC 包,不生成字体或文案;只有通过验证的 新文件才写入新目录。片名卡有渐显或动画时须提供完整序列,不能冻结最后一张图代替。 ## 静态包装图层 `packaging_layers.json` 包框、标题条、角标 logo 这类整片不动的图片,可以直接在合成时叠加,不必先渲染成逐帧序列。 在 `work_dir` 写 `packaging_layers.json`: ```json { "canvas": {"width": 1080, "height": 1920}, "layers": [ {"name": "frame", "path": "/abs/frame.png", "rect": {"x": 0, "y": 0, "width": 1080, "height": 1920}} ], "template": {"id": "brand-frame", "version": 2} } ``` - `canvas` 必须等于成片画布,否则合成前报错;`rect` 必须在画布内,图片缩放到 `rect` 大小(透明区域保持透明)。 - 叠加顺序:遮原字幕 → 包装图层(按数组顺序)→ 画面文字 → 解说字幕 → 缩放。 - `timeline.json` 同时得到对应的 image 轨,剪映草稿里的包框可单独编辑。 - `assembly_manifest.json` 的 `video_filters.packaging_layers` 记录每张图片的路径与 `{size, mtime_ns}`。 - 编排器按项目绑定的 `packaging` 模板写出这个文件时会带 `written_by` 标记; 手写的文件不会被它删除。有渐显或动画的包装仍走 `foreground-compose.md`。 ## 原则 包装价值来自稳定、可读、与内容一致的排版,不来自效果数量。先建立统一字体、颜色、描边/阴影和轻量动效; 底板、边框、花字与音效只有解决具体可读性或叙事任务时才加入。一个样帧好看不代表全片成立。 -
pair-media.md 4.1 KB
# Pair rebuilt picture with a retained adopted soundtrack When the picture is rebuilt/reframed but the adopted mix must not change, use `pair_media.py` before `assemble.py --audio-mode adopted-packet-copy`. This is a real two-input stream copy, not re-encoding an old finished video's picture. It does not generate narration, align speech, design packaging, or create an end card. It does not decide which inputs the editor intended to adopt. ```json { "artifact": "media_pair", "schema_version": 1, "picture": {"path": "/project/picture.mp4"}, "audio": {"path": "/project/adopted.m4a", "selected_stream": 0} } ``` Both paths must be explicit local files. `selected_stream` is the audio ordinal (`a:N`), not the absolute stream index. The audio donor may be an old MP4 with unrelated video: only the selected audio is used. The picture's audio is ignored. Both files must exist when the plan is read. The plan is strict: unknown fields do not silently request unsupported retime, gain, trimming or offset operations (legacy `sha256` keys are ignored). ```bash python3 scripts/pair_media.py pair.json --output-dir new-pair-directory --plan-only # A PLANNED directory is not reusable as a rendered candidate: choose a NEW one. python3 scripts/pair_media.py pair.json --output-dir new-render-directory ``` The output directory must not already exist. No existing input, final asset, current pointer, or user-approved version is overwritten. `--plan-only` checks actual inputs and timing but creates no video and invokes no mux. A rendered run provides: - `paired.mp4`: selected picture and audio, both compressed-stream copied. - `pair_run.json`: `PLANNED`, `PAIR_RENDERED`, or `FAILED`, explicit input and plan/output paths, frame count, timing tolerance and output stream numbering. - `picture_identity.json`: decoder, geometry/color, complete actual frame clock and compressed packet size/timing/side-data facts. - `adopted_audio_identity.json`: donor and output AAC packet/decoder facts. - mux command/log for reconstruction and diagnosis. The first implementation deliberately accepts H264 or HEVC in an MP4-family container, complete zero-origin CFR picture (1–120 fps), and contiguous AAC packets with known nominal sample duration. No VFR, retime, offset, format conversion, padding, trimming, `-shortest`, normalization or gain is inferred. Packet priming and skip metadata are retained. The packet clock must also agree with the declared audio stream interval (at most one nominal priming packet before the start; the packet end matches the stream end within one sample). This is checked again on the actual muxed output. Picture/audio starts and ends must differ by no more than the larger of one picture frame or one nominal AAC packet. The interval check is compatibility, **not perceptual synchronization or acoustic alignment**. A tiny container-tail tolerance does not authorize cutting a word. Muxing goes to a staging file. The output's video decoder/packet sizes and timestamps/full frame clock/geometry/color and adopted AAC packets must match the inputs, the output must contain only `v:0,a:0`, and full decode must pass before final publication. Failure leaves a FAILED record and logs, not a final `paired.mp4`. Rebuilding source geometry is certified by its own upstream render/map evidence, not by the existence of a successfully paired container. ## Subtitle integration: bind AFTER pairing The output audio ordinal is always **0**, even when the donor used `a:1`. Create the complete output-clock `subtitle_track.json` against `subtitles.track_binding.current_bindings(paired_video, 0)`, then run the existing adopted assembly path. Do not reuse a binding computed against the picture-only file or the donor's old stream ordinal. A present stale subtitle track fails rather than falling back to estimated timings. See [subtitle-track.md](subtitle-track.md). Pairing preserves the input picture, including any explicit black tail; it does not turn that black tail into a branded end card. Listening, normal-speed review, editorial intent and release approval remain separate from this mechanical proof. -
source-score.md 6.7 KB
# Source and score bed producer `source_score.py` is a standalone producer for explicitly authored source-audio intervals and one continuous score playhead. It does not mix narration, inspect old masters, infer dialogue, choose music, retime sound, normalize loudness, limit peaks, or publish a final video. ```bash python3 scripts/source_score.py source_score_plan.json --output-dir a-new-directory ``` The output directory must not exist. Success creates: - `source_bed.wav` — reordered/pre-gained source intervals plus explicit silence; - `score_bed.wav` — one continuous raw-score window or an adopted frozen stem; - `prepared_bed.wav` — source plus score, without master gain/limiting; - `prepared_bed_receipt.json` — input paths, clocks, sample ranges, PCM facts and QC facts. All three beds are `pcm_f32le`, 48 kHz, stereo. Float output avoids repeated integer quantization and preserves values outside `[-1,1]`; the receipt records actual peak, finite-sample status, and `FLOAT_PRESERVED_NO_MASTER`. A later explicitly adopted consumer owns final master gain and delivery encoding. ## Strict plan schema v1 ```json { "artifact": "source_score_plan", "schema_version": 1, "output": {"sample_rate": 48000, "channels": 2, "total_samples": 1920000}, "source_segments": [ { "id": "protected-dialogue-1", "path": "/local/source.mov", "audio_stream": 0, "source_fps": "24/1", "source_start_frame": 120, "source_end_frame": 180, "output_start_sample": 240000, "gain": 1.0, "fade_in_samples": 0, "fade_out_samples": 96, "fade_shape": "linear", "role": "protected_original" } ], "source_silence": [ {"output_start_sample": 0, "output_end_sample": 240000, "role": "silence"}, {"output_start_sample": 360000, "output_end_sample": 1920000, "role": "silence"} ], "score": { "kind": "raw", "path": "/local/score.wav", "audio_stream": 0, "source_offset_sample": 384000, "gain": 0.18, "fade_in_samples": 33600, "fade_out_samples": 120000, "fade_shape": "half_cosine" } } ``` Every source range is half-open in CFR picture frames. Version 1 accepts H.264 and HEVC picture sources whose complete packet PTS/duration grid proves one zero-origin packet per same-speed frame. Frame boundaries are projected onto the 48 kHz clock by rounding the exact position half-up once, so `48000/fps` need not be an integer and a broadcast rate such as 30000/1001 is accepted; every derived bound comes from those rounded positions. Other codecs, incomplete packet timing, VFR, retime fields, nominal time seeks, and unknown fields are rejected rather than called verified (legacy `sha256` keys are ignored). Each unique selected source stream is decoded from its actual PTS clock to canonical 48 kHz stereo exactly once. All authored intervals are then cut from that canonical sample array—never independently decoded or `-ss`-seeked per edit. The source segment duration is `(source_end_frame-source_start_frame)*48000/fps`. Its output end is derived, not declared. Source segments and explicit silence ranges must form an exact, nonoverlapping partition of `[0,total_samples)`. There are no implicit gaps. Supported roles are `protected_original` and `mixed_original_under_narration`; their gains and labels are retained unchanged. Source fades are explicitly `linear`. A fade of `N>=2` samples uses inclusive endpoints: fade-in gain is `i/(N-1)`, and fade-out gain is `(N-1-i)/(N-1)`. Thus the first/last samples are exactly zero and the opposite endpoints exactly one. Zero means no fade; a one-sample fade is rejected as ambiguous. Fade windows may not overlap within a segment. ## Score modes `raw` decodes the selected score stream once, advances one playhead from `source_offset_sample`, and takes exactly `total_samples`. The input must be long enough; there is no implicit loop. Gain and one pair of whole-score fades are applied once, so score playback never resets at picture/source cuts. `linear` and `half_cosine` fades use inclusive endpoints. When the adopted source bed is already the complete soundtrack and no additional music is selected, use the exact no-score form: ```json {"kind":"none"} ``` It accepts no other fields. `score_bed.wav` is real all-zero float PCM at the exact output length, while `prepared_bed.wav` is an exact canonical PCM copy of `source_bed.wav`; the receipt records `kind:none` rather than inventing a music asset. An adopted score that already contains its chosen offset/gain/fades uses the exact alternative schema: ```json {"kind":"frozen","path":"/local/frozen.wav","audio_stream":0} ``` Frozen input must be PCM16, PCM24, or PCM float, 48 kHz stereo, and exactly `total_samples`. PCM16/24 values convert exactly to float canonical representation. No offset, gain, fade, loop, normalization or other processing fields are accepted. The receipt records the frozen file path and its probed PCM facts. ## Receipt and failure boundary `prepared_bed_receipt` schema version 1 records the plan path, each input path, selected stream, actual CFR/audio clocks, canonical decoded PCM facts, resolved input/output sample ranges, gains, fades, roles, and each output's path, byte size, format, sample count, finite status and peak. Every declared input must exist and probe as declared before work starts. A selected source range must fit the actual canonical decoded samples; silence padding can never conceal a short source. WAVs remain hidden staging artifacts until every output validates. FFmpeg failure, invalid ranges, non-finite PCM, or an existing target directory cannot publish final-named beds or a receipt. Diagnostic command/log/intermediate files may remain in the unique failed directory. To consume a completed receipt with explicitly adopted narration, use the strict `--audio-mix-adoption` path in `explicit-audio-mix.md`. To deliver an already complete prepared bed without narration, a dedicated prepared-audio renderer will follow in a later release. Do not feed `prepared_bed.wav` through legacy source ducking or ambient BGM/loudness settings. ## 与旧入口的关系(中文摘要) 原片完整解码一次再切样本,分别执行保留对白、低位原声和明确静音;渐变必须显式给定。 已处理的音乐轨走 `frozen`,不能再次偏移、调增益或加渐变。若采用的原声底轨本身已包含 完整音乐决定且不再叠加配乐,使用严格的 `score:{"kind":"none"}`;它生成真实全零 score, 并保持 prepared 与 source 逐字节相同,不伪造静音音乐资产。 这一步仅输出声音底轨和来源回执,不是最终视频。旧入口保留兼容行为;调用方须区分 “底轨已验证”“配音已验证”和“完整混音已验证”三种状态,不能相互冒充。 -
subtitle-track.md 8.5 KB
# Independent subtitle track contract (schema v1) `scripts/subtitles/track.py` loads a subtitle track whose cue times already use the final **output clock**. It validates the track against picture, edit, audio, and duration facts independently supplied by the caller. It does not align speech, remap source time, split text, repair cue boundaries, or prove that a subtitle is perceptually synchronized. ## Loader API ```python from fractions import Fraction from subtitles.track import load_subtitle_track loaded = load_subtitle_track( "subtitle_track.json", expected_picture_identity={ "path": "/project/paired.mp4", "edit_plan": "/project/edit_plan.json", }, expected_audio_identity={ "selected_stream": 1, "sample_rate": 48000, "packet_count": 4700, }, expected_duration_seconds=Fraction(duration_ts) * stream_time_base, reject_legacy_estimate=True, ) metadata = loaded["metadata"] entries = loaded["entries"] ``` The input may also be an in-memory mapping. `entries` is a list of dictionaries ready for the existing seconds-based subtitle renderer: ```json { "start": 1.0, "end": 3.0, "text": "example", "source": "narration", "source_ref": "narration:7", "timing_evidence": { "kind": "asr_boundary_calibrated", "evidence_refs": ["alignment-run:example"], "calibration": "asr_energy", "word_alignment": "none" } } ``` The loader preserves cue text and boundaries. The only conversion is exact integer ticks to renderer-facing seconds. `metadata.timing_evidence_kinds` reports the labels present; it is not an aggregate precision verdict. ## Schema v1 ```json { "schema_version": 1, "clock": { "kind": "output", "timebase": {"numerator": 1, "denominator": 30}, "duration_ticks": 300 }, "overlap_policy": "forbid", "bindings": { "picture": { "path": "/project/paired.mp4", "edit_plan": "/project/edit_plan.json" }, "audio": { "selected_stream": 1, "sample_rate": 48000, "packet_count": 4700 } }, "cues": [ { "start_tick": 30, "end_tick": 90, "text": "example", "attribution": {"kind": "narration", "ref": "narration:7"}, "timing_evidence": { "kind": "asr_boundary_calibrated", "evidence_refs": ["alignment-run:example"], "calibration": "asr_energy", "word_alignment": "none" } } ] } ``` - `timebase` is rational seconds per tick. Cue intervals are half-open `[start_tick, end_tick)`. - `duration_ticks`, cue bounds, and stream indexes are nonnegative integers (booleans are not integers for this contract). - Cues must be ordered, non-overlapping, nonempty, and contained by the output duration. Adjacent half-open cues may touch. - Distinct tick bounds must remain a positive, finite interval after conversion to the renderer's float seconds. Tracks beyond that projection precision fail closed rather than becoming a zero-length rendered cue. - `overlap_policy` must be `forbid`. Schema v1 has no permissive overlap mode. - `attribution.kind` is `source` or `narration`; `ref` identifies the source utterance or narration item without changing its text. - Unknown fields and schema versions other than integer `1` are rejected, so a newer producer cannot be silently interpreted as v1. Legacy `sha256` / `edit_sha256` binding keys are the one exception: they are ignored. ## Independent binding checks The loader compares track declarations with the caller's current facts; it never treats a track's own binding as evidence that the track is fresh. - `picture.path` is mandatory and must resolve to the caller's current picture path. `edit_plan` is optional for inputs without a separately materialized edit plan; if the track contains it, the caller must supply the same path. - `audio.selected_stream`, `sample_rate` and `packet_count` describe the **actually adopted** audio stream, not merely a source filename or an intended mix manifest. Pass the facts probed from the current media; this module deliberately does not run ffprobe itself. - `expected_duration_seconds` is mandatory and checked against `duration_ticks * timebase`. Prefer `Fraction(duration_ts) * time_base` from the actual output stream to avoid decimal/container rounding. `int`, finite `float`, and finite `Decimal` are also accepted. ## Timing evidence labels Evidence labels describe how a cue boundary was obtained. They do not change the exact burn time represented by its ticks. | `kind` | Required calibration | Required word alignment | Evidence refs | |---|---|---|---| | `legacy_estimate` | `none` | `none` | optional | | `asr_boundary_calibrated` | `asr_energy` | `none` | required | | `word_timestamps` | `none` or `asr_energy` | `asr_words` | required | | `human_verified` | `human_boundary` | `none` or `human_words` | required | `asr_boundary_calibrated` is the label for a coarse-ASR boundary adjusted with ASR context and/or energy evidence. Energy onset is not proof of a phoneme, so this label cannot claim word alignment or human verification. Missing evidence references cannot claim any non-legacy kind. A strict caller may reject `legacy_estimate`; accepting another label still does not automatically call it strong alignment. Because cue intervals are half-open, a cue is not visible at any tick before its `start_tick`; the tick immediately preceding it belongs to whatever came before. ## Deliberate limits - Schema v1 has no source-to-cut/edit map and cannot map source-clock cues. - The module does not inspect media, select audio streams, or tolerate a self-declared binding without current caller facts. - Validation proves schema consistency and the requested bindings only. Actual rendered first/last subtitle frames and perceptual speech alignment require separate render/media review. ## Current assembly integration (bounded, not automatic alignment) Place `subtitle_track.json` in `work_dir` and select `assemble.py --audio-mode adopted-packet-copy`. A present track is a **complete replacement** of all generated narration/original-dialogue subtitles, not a partial patch and not merged with legacy subtitles. Carry every cue that should remain; an empty `cues` array intentionally removes all generated subtitles. Unknown patch/merge modes are rejected by schema v1. This does not detect words missing from the authored full track: acoustic/coverage review remains required. The assembly integration independently probes the chosen input stream for its sample rate and packet count, resolves the input path for the picture binding, and records the `{size, mtime_ns}` of the track, video and optional edit plan in `subtitle_track_validation.json`. It does not use the track's own declarations as current facts. Actual decoded frame PTS are read before projection. Integer cue ticks and the rational timebase are retained until each boundary is resolved to the first frame at or after it. `subtitle_track_validation.json` records original ticks, resolved frame indexes/PTS, quantization deltas, ASS thresholds, validation schema and projector version. SRT and timeline consume resolved frame seconds; ASS thresholds are chosen to switch on those same frames despite ASS's 10ms clock. A cue with no visible frame, or a boundary that ASS cannot distinguish, is rejected rather than silently dropped. The original author file is not rewritten. Legacy subtitles keep their previous rendering behavior. Each consumption re-checks the track/media/optional edit plan `{size, mtime_ns}` and the consumer duration; a rewritten input is stale and must be prepared again. Deleting the explicit track clears its previous validation record. An invalid/stale track never falls back to character-proportional timing. Current limits: - Integration is for output media starting at zero with an adopted AAC track, not raw-source-to-edited-output mapping or a newly mixed narration track. - The low-level preparation API accepts `edit_plan_path`, but the assembly CLI does not yet expose it. CLI tracks must omit `edit_plan`; that path binds the input media, not edit-plan ancestry. - The low-level policy `reject_legacy_estimate=True` is available. There is not yet a wired commercial-profile CLI gate. The current CLI preserves timing evidence labels and does not call every accepted cue precisely aligned. - The actual FFmpeg regression demonstrates frame timing with a synthetic cue, not phoneme alignment. `direct_listening` and `acoustic_alignment` remain `NOT_CHECKED`; evidence labels are declarations, not automatically verified human approval or proof that an external evidence reference is true.
-
-
scripts
-
adoption
-
audio_mix_binding.py 13.4 KB
"""Validate, render, and record an explicitly adopted prepared-bed plus narration mix.""" import json from pathlib import Path import subprocess from assemble_constants import frame_clock_samples from adoption.frozen_audio import probe_audio_packets import adoption.narration_binding as narration_binding from pair_media import probe_picture, validate_pair_timing import source_score from adoption.strict_inputs import ( read_json_bytes, require_fields, require_integer, require_local_path, require_number, run_logged, without_digests, write_json_atomic, ) ARTIFACT = "audio_mix_binding" FILENAME = "audio_mix_binding.json" RATE = 48_000 CHANNELS = 2 CODEC = "pcm_f32le" def _picture_format(picture): samples = frame_clock_samples(picture["frame_count"], picture["fps"], RATE) return samples, { "fps": picture["fps"], "frame_count": picture["frame_count"], "duration": picture["duration"], "start": picture["start"], "decoder": picture["decoder"], } def load_adoption(path, *, input_video, narration_adoption_path, tts_segments): """Strict read-only preflight; no work artifacts are created here.""" adoption_path, _, value = read_json_bytes(path, "audio mix adoption") value = without_digests(value, "audio mix adoption") require_fields(value, ["artifact", "schema_version", "prepared_receipt", "format", "segments", "master_gain_db"], "audio mix adoption") if value["artifact"] != "audio_mix_adoption" or type(value["schema_version"]) is not int \ or value["schema_version"] != 1: raise ValueError("unsupported audio_mix_adoption schema") input_video = require_local_path(input_video, "picture") picture = probe_picture(input_video) picture_samples, picture_summary = _picture_format(picture) narration_path = require_local_path(narration_adoption_path, "narration adoption") require_fields(value["format"], ["sample_rate", "channels", "total_samples"], "mix format") mix_format = value["format"] if mix_format != {"sample_rate": RATE, "channels": CHANNELS, "total_samples": picture_samples}: raise ValueError("mix format differs from the actual picture sample clock") prepared_receipt = source_score.validate_prepared_receipt( value["prepared_receipt"], mix_format ) if not isinstance(value["segments"], list) or len(value["segments"]) != len(tts_segments): raise ValueError("audio mix segments must exactly cover narration segments") normalized = [] seen = set() for adopted, segment in zip(value["segments"], tts_segments): adopted = without_digests(adopted, "audio mix segment") require_fields(adopted, ["index", "output_start_sample", "gain"], "audio mix segment") index = require_integer(adopted["index"], "audio mix segment index") if index in seen or index != segment["index"]: raise ValueError("audio mix segment index/order differs from narration") seen.add(index) normalized.append({ "index": index, "output_start_sample": require_integer( adopted["output_start_sample"], "output_start_sample"), "gain": require_number(adopted["gain"], "narration gain", 0, 16), }) return { "path": str(adoption_path), "picture": {"path": str(input_video), "clock": picture_summary}, "picture_identity": picture, "prepared_receipt": prepared_receipt["reference"], "prepared": prepared_receipt["outputs"], "narration_adoption": {"path": str(narration_path)}, "format": dict(mix_format), "segments": normalized, "master_gain_db": require_number(value["master_gain_db"], "master_gain_db", -24, 24), "runtime": None, } def _run(command, work_dir, label): run_logged(command, work_dir, label, prefix="explicit_") def _channels(path): result = subprocess.run( ["ffprobe", "-v", "error", "-select_streams", "a:0", "-show_entries", "stream=channels", "-of", "default=nw=1:nk=1", str(path)], capture_output=True, text=True, timeout=600, ) try: channels = int(result.stdout.strip()) except ValueError as exc: raise ValueError("narration channel probe failed") from exc if result.returncode or channels not in {1, 2}: raise ValueError("explicit narration supports only mono or stereo input") return channels def render_explicit_mix(context, narration_context, tts_segments, work_dir): """Render complete direct 48 kHz placements, voice bus, premaster, and fixed master.""" directory = Path(work_dir).resolve() / ".explicit_audio_mix" directory.mkdir(parents=True, exist_ok=False) by_index = {item["index"]: item for item in narration_context["segments"]} segment_memory = {item["index"]: item for item in tts_segments} rendered = [] for position, adopted in enumerate(context["segments"]): source = by_index[adopted["index"]]["snapshot"]["path"] channels = _channels(source) placed = directory / f"placed_{position:04d}_{adopted['index']}.wav" if channels == 1: channel_filter = ( "aformat=sample_fmts=flt," "pan=stereo|c0=0.7071067811865476*c0|c1=0.7071067811865476*c0," "aresample=48000:async=0:first_pts=0,aformat=sample_fmts=flt:" "sample_rates=48000:channel_layouts=stereo" ) matrix = "mono_equal_power" else: channel_filter = ( "aresample=48000:async=0:first_pts=0,aformat=sample_fmts=flt:" "sample_rates=48000:channel_layouts=stereo" ) matrix = "stereo_identity" _run(["ffmpeg", "-nostdin", "-v", "error", "-n", "-i", str(source), "-map", "0:a:0", "-af", channel_filter, "-c:a", CODEC, str(placed)], directory, f"place_{position:04d}") facts = source_score._output_facts(placed) start = adopted["output_start_sample"] end = start + facts["pcm"]["samples"] if end > context["format"]["total_samples"]: raise ValueError("complete narration segment does not fit the adopted mix clock") rendered.append({**adopted, "output_end_sample": end, "input_channels": channels, "channel_matrix": matrix, "placed": facts}) memory = segment_memory[adopted["index"]] memory.update({ "placed_audio_path": str(placed), "narration_conversion_path": str(placed), "audio_duration": facts["pcm"]["samples"] / RATE, "placed_audio_duration": facts["pcm"]["samples"] / RATE, "actual_place_start": start / RATE, "actual_place_end": end / RATE, "output_start_sample": start, "output_end_sample": end, "adopted_gain": adopted["gain"], "global_narration_speed": 1.0, "segment_tempo_factor": 1.0, "effective_tempo": 1.0, "fit_status": "placed", "blocking": False, "truncated": False, "truncate_reason": "none", }) ordered = sorted(rendered, key=lambda item: item["output_start_sample"]) if any(current["output_start_sample"] < previous["output_end_sample"] for previous, current in zip(ordered, ordered[1:])): raise ValueError("explicit narration placements overlap") total = context["format"]["total_samples"] voice_bus = directory / "voice_bus.wav" command = ["ffmpeg", "-nostdin", "-v", "error", "-n"] for item in rendered: command += ["-i", item["placed"]["path"]] filters = [f"anullsrc=r={RATE}:cl=stereo,atrim=end_sample={total}[base]"] labels = ["[base]"] for position, item in enumerate(rendered): filters.append( f"[{position}:a]volume={item['gain']:.17g}," f"adelay={item['output_start_sample']}S:all=1[v{position}]" ) labels.append(f"[v{position}]") filters.append("".join(labels) + f"amix=inputs={len(labels)}:duration=first:normalize=0," f"atrim=end_sample={total},aformat=sample_fmts=flt:sample_rates={RATE}:" "channel_layouts=stereo[out]") command += ["-filter_complex", ";".join(filters), "-map", "[out]", "-c:a", CODEC, str(voice_bus)] _run(command, directory, "voice_bus") voice_facts = source_score._output_facts(voice_bus) narration_binding.seal_render_inputs(narration_context, tts_segments, voice_bus) narration_context["sealed"]["narration_bus"]["consumption_status"] = \ "CONSUMED_BY_EXPLICIT_MIX" prepared = context["prepared"]["prepared_bed.wav"]["path"] premaster = directory / "premaster.wav" _run(["ffmpeg", "-nostdin", "-v", "error", "-n", "-i", prepared, "-i", str(voice_bus), "-filter_complex", f"[0:a][1:a]amix=inputs=2:duration=first:normalize=0," f"atrim=end_sample={total},aformat=sample_fmts=flt:sample_rates={RATE}:" "channel_layouts=stereo[out]", "-map", "[out]", "-c:a", CODEC, str(premaster)], directory, "premaster") master = directory / "master.wav" gain = 10 ** (context["master_gain_db"] / 20) _run(["ffmpeg", "-nostdin", "-v", "error", "-n", "-i", str(premaster), "-af", f"volume={gain:.17g},aformat=sample_fmts=flt:sample_rates={RATE}:" "channel_layouts=stereo", "-c:a", CODEC, str(master)], directory, "master") runtime = { "segments": rendered, "voice_bus": voice_facts, "premaster": source_score._output_facts(premaster), "master": source_score._output_facts(master), "master_gain_linear": gain, } if any(runtime[key]["pcm"]["samples"] != total for key in ("voice_bus", "premaster", "master")): raise ValueError("explicit mix derivative sample count differs from picture clock") context["runtime"] = runtime return runtime def finalize_binding(context, narration_record, rendered_output, final_output, work_dir): """Probe the rendered candidate and write ``audio_mix_binding.json`` into work_dir.""" if not context.get("runtime"): raise RuntimeError("explicit audio mix must be rendered before finalization") rendered = require_local_path(rendered_output, "rendered output") output_picture = probe_picture(rendered) input_clock = _picture_format(context["picture_identity"])[1] output_clock = _picture_format(output_picture)[1] for key in ("fps", "frame_count", "duration", "start"): if output_clock[key] != input_clock[key]: raise ValueError("rendered output picture frame clock changed") packet_identity = ( "EXACT" if output_picture == context["picture_identity"] else "REENCODED_CLOCK_MATCH" ) audio = probe_audio_packets(rendered, 0) validate_pair_timing(output_picture, audio) decoded = Path(context["runtime"]["master"]["path"]).parent / "final_aac_decoded.wav" _run(["ffmpeg", "-nostdin", "-v", "error", "-n", "-i", str(rendered), "-map", "0:a:0", "-af", "aformat=sample_fmts=flt:sample_rates=48000:" "channel_layouts=stereo", "-c:a", CODEC, str(decoded)], decoded.parent, "final_decode") context["runtime"]["final_decoded_pcm"] = source_score._output_facts(decoded) report = { "artifact": ARTIFACT, "schema_version": 1, "status": "FINALIZED", "adoption": {"path": context["path"]}, "picture": context["picture"], "output_picture": { **output_clock, "packet_identity": packet_identity, }, "prepared_receipt": context["prepared_receipt"], "prepared": context["prepared"], "format": context["format"], "segments": context["runtime"]["segments"], "voice_bus": context["runtime"]["voice_bus"], "premaster": context["runtime"]["premaster"], "master": {**context["runtime"]["master"], "gain_db": context["master_gain_db"], "gain_linear": context["runtime"]["master_gain_linear"]}, "narration_input_binding": {"path": narration_record["path"], "status": "FINALIZED"}, "final_output": { "path": str(Path(final_output).resolve()), "decoded_pcm": context["runtime"]["final_decoded_pcm"], "audio_stream": { "decoder": audio["decoder"], "packet_count": audio["packet_count"], "payload_bytes": audio["payload_bytes"], "start_time": audio["start_time"], "duration": audio["duration"], }, }, "direct_listening": "NOT_CHECKED", "normal_speed_review": "NOT_CHECKED", "release_approved": False, } write_json_atomic(Path(work_dir).resolve() / FILENAME, report) return report def binding_record(work_dir): """The finalized mix binding's path and status; None when absent or malformed.""" path = Path(work_dir) / FILENAME if not path.is_file(): return None try: report = json.loads(path.read_text(encoding="utf-8")) except (UnicodeDecodeError, json.JSONDecodeError): return None if not isinstance(report, dict) or report.get("artifact") != ARTIFACT \ or report.get("schema_version") != 1 or report.get("status") != "FINALIZED": return None final = report.get("final_output") if not isinstance(final, dict) or not Path(str(final.get("path", ""))).is_file(): return None narration = report.get("narration_input_binding") if not isinstance(narration, dict) or not Path(str(narration.get("path", ""))).is_file(): return None return {"path": str(path.resolve()), "status": "FINALIZED"} -
frozen_audio.py 7.6 KB
"""Probe and verify an adopted AAC stream without decoding or rewriting it.""" import json from fractions import Fraction from pathlib import Path from lib import run_cmd def _fraction(value): return Fraction(str(value)) def _packet_span(audio): """The stream duration proved by its packets: last packet end minus first packet start.""" packets = audio["packets"] if not packets or packets[0]["pts"] is None or packets[-1]["pts"] is None \ or packets[-1]["duration"] is None: return None return Fraction(packets[-1]["pts"]) + Fraction(packets[-1]["duration"]) \ - Fraction(packets[0]["pts"]) def _rational_ticks(value, time_base): if value in (None, "N/A"): return None value = Fraction(int(value)) * Fraction(time_base) return f"{value.numerator}/{value.denominator}" def _probe(path, *args): result = run_cmd([ "ffprobe", "-v", "error", *args, "-of", "json", str(path), ]) if result.returncode != 0: raise RuntimeError(f"无法探测媒体 {path}: {result.stderr.strip()}") try: return json.loads(result.stdout) except (TypeError, ValueError) as exc: raise RuntimeError(f"ffprobe 返回无效 JSON: {path}") from exc def probe_audio_packets(path, audio_stream_index): """Return packet sizes and rational timestamps for one audio stream. ``audio_stream_index`` is the zero-based audio-stream ordinal accepted by ffmpeg's ``0:a:N`` selector, not the file-wide absolute stream index. """ payload = _probe( path, "-select_streams", f"a:{audio_stream_index}", "-show_streams", "-show_packets", "-show_entries", "stream=index,codec_name,time_base,start_pts,start_time,duration_ts,duration,sample_rate,channels," "channel_layout:packet=pts,dts,duration,size,side_data_list", ) streams = payload.get("streams", []) if len(streams) != 1: raise RuntimeError(f"找不到音频流 a:{audio_stream_index}: {Path(path)}") stream = streams[0] time_base = stream.get("time_base") if not time_base: raise RuntimeError(f"音频流 a:{audio_stream_index} 缺少 time_base") codec = stream.get("codec_name") sample_rate = int(stream["sample_rate"]) if stream.get("sample_rate") else None channels = int(stream["channels"]) if stream.get("channels") else None channel_layout = stream.get("channel_layout") decoder = { "codec": codec, "sample_rate": sample_rate, "channels": channels, "channel_layout": channel_layout, } packets = [] for packet in payload.get("packets", []): if packet.get("size") is None: raise RuntimeError(f"音频流 a:{audio_stream_index} 的 packet 缺少 size") packets.append({ "size": int(packet["size"]), "pts": _rational_ticks(packet.get("pts"), time_base), "dts": _rational_ticks(packet.get("dts"), time_base), "duration": _rational_ticks(packet.get("duration"), time_base), "side_data_list": packet.get("side_data_list", []), }) return { "selected_audio_stream_index": audio_stream_index, "absolute_stream_index": int(stream["index"]), "codec": codec, "time_base": time_base, "start_time": stream.get("start_time"), "duration": stream.get("duration"), "sample_rate": sample_rate, "channels": channels, "channel_layout": channel_layout, "decoder": decoder, "packet_count": len(packets), "payload_bytes": sum(packet["size"] for packet in packets), "packets": packets, } def _picture_interval(path): payload = _probe( path, "-select_streams", "v:0", "-show_streams", "-show_entries", "stream=time_base,start_pts,start_time,duration_ts,duration,avg_frame_rate", ) streams = payload.get("streams", []) if len(streams) != 1: raise RuntimeError(f"找不到主画面流 v:0: {Path(path)}") stream = streams[0] start = _fraction(stream.get("start_time", "0")) if stream.get("duration") in (None, "N/A"): raise RuntimeError("主画面流缺少可验证时长") duration = _fraction(stream["duration"]) frame_rate = Fraction(stream.get("avg_frame_rate", "0/1")) frame = Fraction(1, 1000) if frame_rate <= 0 else 1 / frame_rate return start, duration, frame def validate_adopted_source(path, audio_stream_index): """Fail unless the selected input is copyable AAC spanning the picture interval.""" audio = probe_audio_packets(path, audio_stream_index) if audio["codec"] != "aac": raise RuntimeError( f"adopted-packet-copy 当前只支持 MP4 AAC stream-copy;" f"a:{audio_stream_index} 是 {audio['codec'] or 'unknown'}" ) if not audio["packets"]: raise RuntimeError(f"音频流 a:{audio_stream_index} 没有 packet,不能采用") picture_start, picture_duration, frame_tolerance = _picture_interval(path) if audio["start_time"] in (None, "N/A") or audio["duration"] in (None, "N/A"): raise RuntimeError(f"音频流 a:{audio_stream_index} 缺少可验证起止时间") audio_start = _fraction(audio["start_time"]) audio_duration = _fraction(audio["duration"]) packet_durations = [ _fraction(packet["duration"]) for packet in audio["packets"] if packet["duration"] is not None ] tolerance = max([frame_tolerance, Fraction(1, 1000), *packet_durations]) if ( abs(audio_start - picture_start) > tolerance or abs(audio_duration - picture_duration) > tolerance ): raise RuntimeError( "采用音频与画面时长/起点不兼容: " f"picture={float(picture_start):.6f}+{float(picture_duration):.6f}s, " f"audio={float(audio_start):.6f}+{float(audio_duration):.6f}s" ) return audio def verify_adopted_audio(input_path, output_path, input_stream_index, output_stream_index=0): """Probe the rendered output and check packet count, payload bytes, duration and timing.""" expected = probe_audio_packets(input_path, input_stream_index) actual = probe_audio_packets(output_path, output_stream_index) if expected["decoder"] != actual["decoder"]: raise RuntimeError("采用音频 decoder 参数已改变") if expected["packet_count"] != actual["packet_count"]: raise RuntimeError( f"采用音频 packet count 已改变: {expected['packet_count']} -> {actual['packet_count']}" ) if expected["payload_bytes"] != actual["payload_bytes"]: raise RuntimeError("采用音频 packet payload 总字节数不一致") if _packet_span(expected) != _packet_span(actual): raise RuntimeError("采用音频流时长已改变") expected_packet_core = [ {key: value for key, value in packet.items() if key != "side_data_list"} for packet in expected["packets"] ] actual_packet_core = [ {key: value for key, value in packet.items() if key != "side_data_list"} for packet in actual["packets"] ] if expected_packet_core != actual_packet_core: raise RuntimeError("采用音频 packet PTS/DTS/duration/size 不一致") if [packet["side_data_list"] for packet in expected["packets"]] != [ packet["side_data_list"] for packet in actual["packets"] ]: raise RuntimeError("采用音频 packet side data 已改变") return { "verified": True, "selected_audio_stream_index": input_stream_index, "output_audio_stream_index": output_stream_index, "input": expected, "output": actual, } -
narration_binding.py 14.5 KB
"""Record which narration inputs an assembly consumed: adoption, snapshot, placement, output.""" import json import math from pathlib import Path import shutil import subprocess from adoption.frozen_audio import probe_audio_packets from adoption.strict_inputs import ( read_json_bytes, require_fields, require_local_path, without_digests, write_json_atomic, ) ARTIFACT = "narration_input_binding" FILENAME = "narration_input_binding.json" IDENTITY_STATUSES = frozenset({"UNADOPTED", "BOUND_TO_ADOPTION"}) # The conservative default an adoption gets when it declares nothing stronger. TEMPO_POLICY = { "global_atempo": 1.0, "bounded_segment_fit": False, "segment_tempo_max": 1.0, "cumulative_tempo_max": 1.0, "cumulative_tempo_hard_max": 1.0, } TEMPO_NUMBERS = ("global_atempo", "segment_tempo_max", "cumulative_tempo_max", "cumulative_tempo_hard_max") def validate_tempo_policy(value): """Accept any adoption-declared tempo policy whose shape and bounds hold. The adoption, not this module, decides how fast its own narration may be played; the module only refuses policies that are malformed or that would let a segment exceed the cumulative hard ceiling the policy itself declares. """ require_fields(value, sorted(TEMPO_POLICY), "tempo_policy") policy = {} for key in TEMPO_NUMBERS: number = value[key] if type(number) not in (int, float) or not math.isfinite(number): raise ValueError(f"tempo_policy {key} must be a finite number") policy[key] = float(number) if type(value["bounded_segment_fit"]) is not bool: raise ValueError("tempo_policy bounded_segment_fit must be a boolean") policy["bounded_segment_fit"] = value["bounded_segment_fit"] if policy["cumulative_tempo_max"] < 1.0: raise ValueError("tempo_policy cumulative_tempo_max must be at least 1.0") if policy["cumulative_tempo_hard_max"] < policy["cumulative_tempo_max"]: raise ValueError("tempo_policy cumulative_tempo_hard_max must not be below " "cumulative_tempo_max") if policy["segment_tempo_max"] < 1.0: raise ValueError("tempo_policy segment_tempo_max must be at least 1.0") if not 0 < policy["global_atempo"] <= policy["cumulative_tempo_hard_max"]: raise ValueError("tempo_policy global_atempo must be positive and within " "cumulative_tempo_hard_max") return policy def load_adoption(path, *, tts_meta_path, tts_segments): """Load a strict v1 adoption and check it against the current tts_meta segments.""" if tts_meta_path is None: raise ValueError("narration adoption requires explicit tts_meta_path") adoption_path, _, adoption = read_json_bytes(path, "narration adoption") adoption = without_digests(adoption, "narration adoption") require_fields(adoption, ["artifact", "schema_version", "segments", "tempo_policy"], "narration adoption") if adoption["artifact"] != "narration_adoption" or type(adoption["schema_version"]) is not int \ or adoption["schema_version"] != 1: raise ValueError("unsupported narration_adoption schema") tts_meta_path, _, tts_meta = read_json_bytes(tts_meta_path, "tts_meta") if not isinstance(tts_meta, dict) or not isinstance(tts_meta.get("segments"), list): raise ValueError("tts_meta requires a segments list") if tts_meta["segments"] != tts_segments: raise ValueError("in-memory narration segments differ from bound tts_meta") tempo_policy = validate_tempo_policy(adoption["tempo_policy"]) if not isinstance(adoption["segments"], list) or len(adoption["segments"]) != len(tts_segments): raise ValueError("adoption segments must exactly cover tts_meta segments") normalized_segments = [] for adopted, actual in zip(adoption["segments"], tts_segments): adopted = without_digests(adopted, "adoption segment") require_fields(adopted, ["index", "spoken_text", "requested_provider", "requested_voice"], "adoption segment") if type(adopted["index"]) is not int or adopted["index"] != actual.get("index"): raise ValueError("adoption segment index/order differs from tts_meta") spoken = actual.get("spoken_text", actual.get("narration")) if not isinstance(adopted["spoken_text"], str) or adopted["spoken_text"] != spoken: raise ValueError("adoption spoken_text differs from tts_meta") for key in ("requested_provider", "requested_voice"): if not isinstance(adopted[key], str) or not adopted[key]: raise ValueError(f"adoption {key} must be a non-empty string") normalized_segments.append(dict(adopted)) return { "path": str(adoption_path), "tts_meta": {"path": str(tts_meta_path)}, "segments": normalized_segments, "tempo_policy": tempo_policy, } def _copy_snapshot(source, destination): """Copy seam kept small so tests can inject a post-preflight source mutation.""" shutil.copyfile(source, destination) def prepare_binding(tts_segments, work_dir, *, narration_adoption_path=None, tts_meta_path=None): """Validate first, then snapshot the adopted narration inputs into work_dir.""" if not isinstance(tts_segments, list): raise ValueError("tts_segments must be a list") adoption = ( load_adoption(narration_adoption_path, tts_meta_path=tts_meta_path, tts_segments=tts_segments) if narration_adoption_path is not None else None ) if not adoption: originals = [ {"index": segment["index"], "source": Path(segment["audio_path"]).resolve(), "spoken_text": segment["spoken_text"]} for segment in tts_segments ] return { "identity_status": "UNADOPTED", "adoption": None, "tempo_policy": None, "originals": originals, "active": False, "segments": [], } originals = [] seen_indices = set() for position, segment in enumerate(tts_segments): if not isinstance(segment, dict) or type(segment.get("index")) is not int: raise ValueError("each narration segment requires an integer index") if segment["index"] in seen_indices: raise ValueError("narration segment indices must be unique") seen_indices.add(segment["index"]) adopted = adoption["segments"][position] source = require_local_path(segment["audio_path"], "narration audio") originals.append({ "index": segment["index"], "source": source, "spoken_text": segment["spoken_text"], "requested_provider": adopted["requested_provider"], "requested_voice": adopted["requested_voice"], }) context = { "identity_status": "BOUND_TO_ADOPTION", "adoption": adoption, "tempo_policy": adoption["tempo_policy"], "originals": originals, "active": True, "segments": [], "sealed": None, } snapshot_dir = Path(work_dir).resolve() / ".narration_input_snapshots" snapshot_dir.mkdir(parents=True, exist_ok=False) for position, (segment, original) in enumerate(zip(tts_segments, originals)): suffix = original["source"].suffix or ".audio" snapshot = snapshot_dir / f"segment_{position:04d}_{segment['index']}{suffix}" _copy_snapshot(original["source"], snapshot) segment["audio_path"] = str(snapshot) segment["narration_input_original"] = {"path": str(original["source"])} segment["narration_input_snapshot"] = {"path": str(snapshot)} context["segments"].append({ "index": original["index"], "spoken_text": original["spoken_text"], "requested_provider": original["requested_provider"], "requested_voice": original["requested_voice"], "original": dict(segment["narration_input_original"]), "snapshot": dict(segment["narration_input_snapshot"]), }) return context def _pcm(path): result = subprocess.run( ["ffprobe", "-v", "error", "-select_streams", "a:0", "-show_streams", "-show_entries", "stream=codec_name,sample_fmt,sample_rate,channels,channel_layout", "-of", "json", str(path)], capture_output=True, text=True, timeout=600, ) if result.returncode: raise ValueError(f"audio probe failed: {result.stderr.strip()}") streams = json.loads(result.stdout).get("streams", []) if len(streams) != 1: raise ValueError("expected exactly one audio stream") stream = streams[0] return {key: stream.get(key) for key in ( "codec_name", "sample_fmt", "sample_rate", "channels", "channel_layout" )} def _asset(path, *, pcm=False): path = require_local_path(path, "binding asset") value = {"path": str(path)} if pcm: value["pcm"] = _pcm(path) return value def seal_render_inputs(context, tts_segments, narration_wav): """Record every derived audio file that the final FFmpeg command will consume.""" if not context.get("active"): return None by_index = {segment["index"]: segment for segment in tts_segments} sealed_segments = [] for item in context["segments"]: segment = by_index[item["index"]] placed = segment.get("placed_audio_path") if not placed: raise RuntimeError("active narration binding has no complete placed audio") conversion = segment.get("narration_conversion_path") conversion_asset = ( {"applied": True, **_asset(conversion, pcm=True)} if conversion else {"applied": False} ) sealed_segments.append({ "index": item["index"], "conversion": conversion_asset, "placed": _asset(placed, pcm=True), }) context["sealed"] = { "segments": sealed_segments, "narration_bus": _asset(narration_wav, pcm=True), } return context["sealed"] def _active_report(context, rendered_output, final_output): if not context.get("sealed"): raise RuntimeError("active narration inputs must be sealed before finalization") sealed = {item["index"]: item for item in context["sealed"]["segments"]} segments = [] for item in context["segments"]: segments.append({ **item, "original": _asset(item["original"]["path"], pcm=True), "snapshot": _asset(item["snapshot"]["path"], pcm=True), "conversion": sealed[item["index"]]["conversion"], "placed": sealed[item["index"]]["placed"], }) rendered_path = require_local_path(rendered_output, "rendered output") final_audio = probe_audio_packets(rendered_path, 0) return { "artifact": ARTIFACT, "schema_version": 1, "status": "FINALIZED", "identity_status": context["identity_status"], "adoption": context["adoption"], "segments": segments, "narration_bus": context["sealed"]["narration_bus"], "final_output": { "path": str(Path(final_output).resolve()), "audio_stream": { "decoder": final_audio["decoder"], "packet_count": final_audio["packet_count"], "payload_bytes": final_audio["payload_bytes"], "start_time": final_audio["start_time"], "duration": final_audio["duration"], }, }, "voice_authentication": "NOT_CHECKED", "direct_listening": "NOT_CHECKED", } def finalize_binding(context, tts_segments, narration_wav, final_output, *, rendered_output=None): """Write ``narration_input_binding.json`` beside ``narration_wav`` and return it. ``rendered_output`` is the candidate file to probe when the published ``final_output`` path does not exist yet; it defaults to ``final_output``. """ del tts_segments # already sealed; the record describes what was consumed if not context.get("active"): report = { "artifact": ARTIFACT, "schema_version": 1, "status": "FINALIZED", "identity_status": "UNADOPTED", "adoption": None, "segments": [ { "index": original["index"], "spoken_text": original["spoken_text"], "requested_provider": None, "requested_voice": None, "original": {"path": str(original["source"])}, "snapshot": None, "conversion": {"applied": False}, "placed": None, } for original in context["originals"] ], "narration_bus": ({"path": str(Path(narration_wav).resolve())} if Path(narration_wav).is_file() else None), "final_output": ({"path": str(Path(final_output).resolve())} if Path(final_output).is_file() else None), "voice_authentication": "NOT_CHECKED", "direct_listening": "NOT_CHECKED", } else: report = _active_report( context, final_output if rendered_output is None else rendered_output, final_output ) write_json_atomic(Path(narration_wav).resolve().parent / FILENAME, report) return report def record_of(report, path): """The manifest/QC summary of one finalized binding report stored at ``path``.""" return {"path": str(Path(path).resolve()), "identity_status": report["identity_status"], "tempo_policy": ( report["adoption"]["tempo_policy"] if report["identity_status"] == "BOUND_TO_ADOPTION" else None )} def binding_record(work_dir): """The finalized binding's path, identity status and tempo policy; None when absent/invalid.""" path = Path(work_dir) / FILENAME if not path.is_file(): return None try: value = json.loads(path.read_text(encoding="utf-8")) except (UnicodeDecodeError, json.JSONDecodeError): return None if ( not isinstance(value, dict) or value.get("artifact") != ARTIFACT or type(value.get("schema_version")) is not int or value.get("schema_version") != 1 or value.get("status") != "FINALIZED" or value.get("identity_status") not in IDENTITY_STATUSES ): return None if value["identity_status"] == "BOUND_TO_ADOPTION": adoption = value.get("adoption") if not isinstance(adoption, dict): return None try: validate_tempo_policy(adoption.get("tempo_policy")) except ValueError: return None final_output = value.get("final_output") if not isinstance(final_output, dict) or not Path(str(final_output.get("path", ""))).is_file(): return None return record_of(value, path) -
strict_inputs.py 4.1 KB
"""Shared strict-input helpers for the explicit-input assemble modules. Every helper here fails closed with ``ValueError`` on malformed input and never guesses: a path must be a local file, a JSON document must parse, a rational must be canonical ``N/D``. ``run_logged`` and ``probe_json`` wrap the external tools so each caller logs the same evidence. """ from fractions import Fraction import json import math from pathlib import Path import subprocess def require_fields(value, required, label): if not isinstance(value, dict) or set(value) != set(required): raise ValueError(f"{label} requires exactly fields {required}") def without_digests(value, label): """Drop the ``sha256`` / ``*_sha256`` keys older caller JSON declared; they are ignored.""" if not isinstance(value, dict): raise ValueError(f"{label} must be an object") return {key: item for key, item in value.items() if key != "sha256" and not key.endswith("_sha256")} def require_integer(value, label, minimum=0): if type(value) is not int or value < minimum: raise ValueError(f"{label} must be an integer >= {minimum}") return value def require_number(value, label, minimum, maximum): if type(value) not in (int, float) or not math.isfinite(value) \ or not minimum <= value <= maximum: raise ValueError(f"{label} must be finite in [{minimum},{maximum}]") return float(value) def require_local_path(path, label): if not isinstance(path, (str, Path)) or not str(path) or "://" in str(path): raise ValueError(f"{label} requires a local path") resolved = Path(path).resolve() if not resolved.is_file(): raise ValueError(f"{label} is missing: {resolved}") return resolved def require_declared_path(value, label): """Resolve a caller ``{"path": ...}`` declaration to an existing local file.""" if not isinstance(value, dict): raise ValueError(f"{label} must be an object with a path") return require_local_path(value.get("path"), label) def read_json_bytes(path, label): """Return (resolved_path, raw_bytes, parsed) for one local JSON document.""" resolved = require_local_path(path, label) raw = resolved.read_bytes() try: return resolved, raw, json.loads(raw) except (UnicodeDecodeError, json.JSONDecodeError) as exc: raise ValueError(f"{label} is not valid JSON") from exc def write_json_atomic(path, value): path = Path(path) temporary = path.with_suffix(".writing.json") temporary.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8") temporary.replace(path) def canonical_fraction(value, label): if not isinstance(value, str): raise ValueError(f"{label} must be a canonical rational string") try: result = Fraction(value) except (ValueError, ZeroDivisionError) as exc: raise ValueError(f"invalid {label}") from exc if result <= 0 or value != f"{result.numerator}/{result.denominator}": raise ValueError(f"{label} must be a positive canonical N/D rational") return result def probe_json(path, *ffprobe_args): result = subprocess.run( ["ffprobe", "-v", "error", *ffprobe_args, "-of", "json", str(path)], capture_output=True, text=True, timeout=600, ) if result.returncode or result.stderr.strip(): raise ValueError(f"media probe failed for {path}: {result.stderr.strip()}") try: return json.loads(result.stdout) except json.JSONDecodeError as exc: raise ValueError(f"invalid media probe JSON for {path}") from exc def run_logged(command, directory, label, *, prefix="", timeout=3600): """Run one FFmpeg command, keeping ``<prefix><label>.command.json`` and ``.log``.""" directory = Path(directory) name = f"{prefix}{label}" write_json_atomic(directory / f"{name}.command.json", command) result = subprocess.run(command, capture_output=True, text=True, timeout=timeout) (directory / f"{name}.log").write_text(result.stderr, encoding="utf-8") if result.returncode: raise RuntimeError(f"{name} FFmpeg failed; see {name}.log") return result -
strict_publish.py 6 KB
"""Publish transaction for strict-adoption renders. An active narration binding renders to a hidden candidate file. Its binding (and, for an explicit mix, the audio mix binding) is written, QC gates the candidate, the media file is published, and QC runs once more against the published path. Any failure removes the candidate, the published file and the bindings written by this render, so a strict output never exists without its consumed-input record. """ from pathlib import Path import assembly_contract import adoption.audio_mix_binding as audio_mix_binding import adoption.narration_binding as narration_binding def current_narration_binding(work_dir, audio_mode): """Read the narration record only for the explicit narration render path.""" if audio_mode != "narration": return None return narration_binding.binding_record(work_dir) def current_audio_mix_binding(work_dir, audio_mode): if audio_mode != "narration": return None return audio_mix_binding.binding_record(work_dir) def publish_render(*, work_dir, binding, explicit_mix, tts_segments, narration_wav, render_output, published_output, audio_mode, audio_operations, adopted_audio, loudness_mode, loudnorm_measurement, visual_qc, source_has_audio, video_duration, render_delivery, source_audio_status): """Gate the rendered candidate on assembly QC and publish it with its bindings. Returns the final output path. Inactive-binding renders only write their binding and QC; strict renders write bindings, publish, and re-run QC, or roll everything back. """ active = bool(binding and binding["active"]) work_dir = Path(work_dir) binding_path = work_dir / narration_binding.FILENAME mix_binding_path = work_dir / audio_mix_binding.FILENAME binding_written = False mix_binding_written = False try: if active: report = narration_binding.finalize_binding( binding, tts_segments, (work_dir / "narration.wav" if explicit_mix is not None else narration_wav), published_output, rendered_output=render_output, ) binding_written = True if not binding_path.is_file(): raise RuntimeError("narration binding 未写入") narration_record = narration_binding.record_of(report, binding_path) mix_record = None if explicit_mix is not None: audio_mix_binding.finalize_binding( explicit_mix, narration_record, render_output, published_output, work_dir ) mix_binding_written = True mix_record = {"path": str(mix_binding_path.resolve()), "status": "FINALIZED"} elif binding: narration_binding.finalize_binding( binding, tts_segments, narration_wav, render_output ) narration_record = current_narration_binding(work_dir, audio_mode) mix_record = None else: narration_record = current_narration_binding(work_dir, audio_mode) mix_record = None assembly_qc = assembly_contract._build_assembly_qc( tts_segments, video_duration, output_path=render_output, source_has_audio=source_has_audio, loudness_mode=loudness_mode, loudnorm_measurement=loudnorm_measurement, visual_qc=visual_qc, audio_mode=audio_mode, audio_operations=audio_operations, adopted_audio=adopted_audio, narration_input_binding=narration_record, audio_mix_binding=mix_record, source_audio_status=source_audio_status, render_delivery=render_delivery, ) if active and assembly_qc["blocking"]: assembly_contract._write_assembly_qc(work_dir, assembly_qc) codes = ", ".join(assembly_qc["blocking_codes"]) raise RuntimeError(f"身份约束渲染 QC 失败: {codes}") if active: render_output.rename(published_output) render_output = published_output current_binding = current_narration_binding(work_dir, audio_mode) if current_binding is None: raise RuntimeError("已发布 narration binding 未通过终态检查") current_mix_binding = current_audio_mix_binding(work_dir, audio_mode) if explicit_mix is not None and current_mix_binding is None: raise RuntimeError("已发布 audio mix binding 未通过终态检查") assembly_qc = assembly_contract._build_assembly_qc( tts_segments, video_duration, output_path=render_output, source_has_audio=source_has_audio, loudness_mode=loudness_mode, loudnorm_measurement=loudnorm_measurement, visual_qc=visual_qc, audio_mode=audio_mode, audio_operations=audio_operations, adopted_audio=adopted_audio, narration_input_binding=current_binding, audio_mix_binding=current_mix_binding, source_audio_status=source_audio_status, render_delivery=assembly_qc["delivery_qc"], ) if assembly_qc["blocking"]: assembly_qc["output"] = { "path": str(published_output), "exists": False, "bytes": 0, } assembly_contract._write_assembly_qc(work_dir, assembly_qc) codes = ", ".join(assembly_qc["blocking_codes"]) raise RuntimeError(f"身份约束渲染终态 QC 失败: {codes}") assembly_contract._write_assembly_qc(work_dir, assembly_qc) except Exception: if active: render_output.unlink(missing_ok=True) published_output.unlink(missing_ok=True) if binding_written: binding_path.unlink(missing_ok=True) if mix_binding_written: mix_binding_path.unlink(missing_ok=True) raise return render_output -
__init__.py 94 B
"""Adoption family: narration/audio-mix binding + strict publish for adopted-source audio."""
-
-
jianying
-
builders.py 30.1 KB
"""Production material and segment builders for the JianYing exporter.""" import json import os from copy import deepcopy from jianying.schema import scrub_platform_identity, us from jianying.templates import template from jianying.tracks import SEGMENT_RENDER_INDEX def timerange(start_us, dur_us): return {"start": int(start_us), "duration": int(dur_us)} def speed_material(new_id, speed=1.0): return { "id": new_id(), "speed": float(speed), "type": "speed", "mode": 0, "curve_speed": None, } def volume_keyframes(keyframes, seg_start_s, new_id): """Build one KFTypeVolume keyframe list from timeline-absolute points.""" if not keyframes: return [] kfs = [] for kf in keyframes: kfs.append({ "curveType": "Line", "graphID": "", "left_control": {"x": 0.0, "y": 0.0}, "right_control": {"x": 0.0, "y": 0.0}, "id": new_id(), "time_offset": max(0, us(kf["t"] - seg_start_s)), "values": [round(float(kf["gain"]), 4)], }) return [{ "id": new_id(), "keyframe_list": kfs, "material_id": "", "property_type": "KFTypeVolume", }] def windowed_volume_keyframes(keyframes, seg_start_s, seg_end_s, default_gain, new_id): """Window timeline-absolute keyframes for one split/looped segment.""" if not keyframes: return [] start = float(seg_start_s) end = float(seg_end_s) default_gain = float(default_gain) ordered = sorted(( {"t": float(kf["t"]), "gain": float(kf["gain"])} for kf in keyframes if "t" in kf and "gain" in kf ), key=lambda kf: kf["t"]) if not ordered or end <= start: return [] start_gain = default_gain for kf in ordered: if kf["t"] <= start: start_gain = kf["gain"] else: break inner = [kf for kf in ordered if start <= kf["t"] <= end] if not inner and abs(start_gain - default_gain) < 1e-4: return [] selected = [{"t": start, "gain": start_gain}] for kf in inner: if abs(kf["t"] - start) < 1e-4: selected[-1] = {"t": start, "gain": kf["gain"]} else: selected.append(kf) if all(abs(kf["gain"] - default_gain) < 1e-4 for kf in selected): return [] return volume_keyframes(selected, start, new_id) def clip_from_segment(segment): """Map optional authoring transforms (validated as objects by the timeline contract).""" scale = segment.get("scale", {}) position = segment.get("position", {}) flip = segment.get("flip", {}) return { "alpha": round(float(segment.get("opacity", 1.0)), 4), "flip": { "horizontal": bool(flip.get("horizontal", False)), "vertical": bool(flip.get("vertical", False)), }, "rotation": float(segment.get("rotation_degrees", 0.0)), "scale": { "x": float(scale.get("x", 1.0)), "y": float(scale.get("y", 1.0)), }, # Timeline v2 uses JianYing's normalized, canvas-center, Y-up coordinates. "transform": { "x": float(position.get("x", 0.0)), "y": float(position.get("y", 0.0)), }, } def base_segment(material_id, target_start_us, target_dur_us, volume, keyframes, new_id): segment = template("segment") segment.update({ "id": new_id(), "material_id": material_id, "target_timerange": timerange(target_start_us, target_dur_us), "common_keyframes": keyframes, "track_render_index": SEGMENT_RENDER_INDEX, "render_index": SEGMENT_RENDER_INDEX, "volume": round(float(volume), 4), }) return segment def audio_segment_piece(material_id, target_start_us, target_dur_us, source_start_us, source_dur_us, volume, keyframes, new_id): seg = base_segment(material_id, target_start_us, target_dur_us, volume, keyframes, new_id) seg["source_timerange"] = timerange(source_start_us, source_dur_us) seg["extra_material_refs"] = [] return seg def _hex_rgb(color): value = str(color or "#FFFFFF").lstrip("#") if len(value) == 8: value = value[:6] if len(value) != 6: raise ValueError(f"invalid text color: {color!r}") try: return [round(int(value[index:index + 2], 16) / 255.0, 6) for index in (0, 2, 4)] except ValueError as exc: raise ValueError(f"invalid text color: {color!r}") from exc def _text_style(style, start, end): fill_color = style.get("fill_color", "#FFFFFF") authored = template("text_style")["styles"][0] authored.update({ "fill": { "alpha": 1.0, "content": { "render_type": "solid", "solid": {"alpha": 1.0, "color": _hex_rgb(fill_color)}, }, }, "range": [int(start), int(end)], "size": float(style.get("font_size", 8.0)), "bold": bool(style.get("bold", False)), "italic": bool(style.get("italic", False)), "underline": bool(style.get("underline", False)), "strokes": list(style.get("strokes", [])), }) if style.get("font_path") or style.get("font_id"): authored["font"] = { "id": str(style.get("font_id", "")), "path": str(style.get("font_path", "")), } if style.get("stroke_color") and float(style.get("stroke_width", 0)) > 0: width = float(style["stroke_width"]) authored["strokes"] = [{ "alpha": 1.0, "content": { "render_type": "solid", "solid": {"alpha": 1.0, "color": _hex_rgb(style["stroke_color"])}, }, "width": round(0.00196 * (width ** 1.013), 6), }] if style.get("shadow_color"): authored["shadows"] = [{ "alpha": float(style.get("shadow_opacity", 90)) / 100.0, "angle": int(style.get("shadow_angle", -45)), "distance": int(style.get("shadow_width", 5)), "feather": float(style.get("shadow_vague", 45)) / 100.0, "content": { "render_type": "solid", "solid": {"alpha": 1.0, "color": _hex_rgb(style["shadow_color"])}, }, }] if isinstance(style.get("effect_style"), dict): authored["effect_style"] = deepcopy(style["effect_style"]) authored["use_letter_color"] = True return authored def rich_text_content(text, base_style=None, words=None, style_presets=None): """Build duo-video text styles using UTF-16 code-unit ranges.""" base_style = dict(base_style or {}) length = len(text.encode("utf-16-le")) // 2 style_presets = style_presets or {} normalized_words = [] boundaries = {0, length} for word in words or []: start = max(0, min(length, int(word.get("index", 0)))) end = max(start, min(length, start + int(word.get("length", 0)))) if end <= start: continue word_style = dict(style_presets.get(str(word.get("style_id")), {})) word_style.update(word) normalized_words.append((start, end, word_style)) boundaries.update((start, end)) styles = [] points = sorted(boundaries) for start, end in zip(points, points[1:]): style = dict(base_style) for word_start, word_end, word in normalized_words: if word_start <= start and end <= word_end: style.update({key: value for key, value in word.items() if key not in {"index", "length"}}) styles.append(_text_style(style, start, end)) content = template("text_style") content["text"] = text content["styles"] = styles return content RESOURCE_TRACKS = { "sound": "audios", "sticker": "stickers", "text_template": "text_templates", "video_effect": "video_effects", "face_effect": "video_effects", } def _resource_material(segment, kind, new_id): """The contract guarantees exactly one of material / resource_config (resolved package).""" config = segment.get("resource_config") if config is None: raw = deepcopy(segment["material"]) else: raw = deepcopy(config["main_config"]) resource_id = config.get("resource_id") if resource_id is not None: raw.setdefault("resource_id", resource_id) if config.get("resources"): raw["_bundle_resources"] = deepcopy(config["resources"]) cover_img = config.get("cover_img") if kind == "sticker" and cover_img: raw["icon_url"] = cover_img raw["preview_cover_url"] = cover_img raw["id"] = new_id() return raw def build_resource_track(ctx, timeline_track): kind = timeline_track["kind"] materials_key = RESOURCE_TRACKS[kind] track_name = timeline_track.get("name", kind) for item in timeline_track["segments"]: ts, te = float(item["timeline_start"]), float(item["timeline_end"]) duration_us = us(te - ts) authored_item = item package_name = item.get("resource_package") if package_name is not None: package = ctx.resource_packages.get(str(package_name)) if not isinstance(package, dict): raise ValueError(f"unknown JianYing resource package: {package_name}") authored_item = dict(item) authored_item["resource_config"] = package material = _resource_material(authored_item, kind, ctx.new_id) ctx.materials[materials_key].append(material) config = authored_item.get("resource_config") or {} if kind == "text_template": subordinate_resources = deepcopy(material.get("_bundle_resources", [])) subordinate_texts = deepcopy(config.get("texts", [])) subordinate_effects = deepcopy(config.get("effects", [])) if subordinate_resources: for subordinate in subordinate_texts + subordinate_effects: if isinstance(subordinate, dict): subordinate["_bundle_resources"] = deepcopy(subordinate_resources) ctx.materials["texts"].extend(subordinate_texts) ctx.materials["effects"].extend(subordinate_effects) seg = base_segment(material["id"], us(ts), duration_us, 1.0, [], ctx.new_id) speed_value = float(item.get("speed", 1.0)) seg["source_timerange"] = timerange(0, round(duration_us * speed_value)) seg["track_render_index"] = 0 seg["extra_material_refs"] = [] seg["speed"] = speed_value if abs(speed_value - 1.0) > 1e-9: speed = speed_material(ctx.new_id, speed_value) ctx.materials["speeds"].append(speed) seg["extra_material_refs"].append(speed["id"]) seg["clip"] = clip_from_segment(item) ctx.add_segment(kind, track_name, us(ts), duration_us, seg) def _attachment_material(spec, kind, new_id): material = deepcopy(spec) config = material.pop("main_config", None) resources = material.pop("resources", None) resource_id = material.pop("resource_id", None) if config is not None: if not isinstance(config, dict): raise ValueError(f"video {kind}.main_config must be an object") merged = deepcopy(config) merged.update(material) material = merged if resource_id is not None: material.setdefault("resource_id", resource_id) if resources: material["_bundle_resources"] = deepcopy(resources) material["id"] = new_id() return material def _resolve_attachment_spec(ctx, spec, kind): if isinstance(spec, str): package = ctx.resource_packages.get(spec) if not isinstance(package, dict): raise ValueError(f"unknown JianYing {kind} resource package: {spec}") return package return spec def apply_video_attachments(ctx, clip, segment): """Attach duo-video transition, mask, and LUT authoring data.""" semantic_kind = "video" transition_spec = clip.get("transition") if transition_spec is not None: transition_spec = _resolve_attachment_spec(ctx, transition_spec, "transition") transition = _attachment_material(transition_spec, "transition", ctx.new_id) ctx.materials["transitions"].append(transition) segment["extra_material_refs"].append(transition["id"]) mask_spec = clip.get("mask") if mask_spec is not None: mask_spec = _resolve_attachment_spec(ctx, mask_spec, "mask") mask = _attachment_material(mask_spec, "mask", ctx.new_id) ctx.materials["masks"].append(mask) ctx.materials["common_mask"].append(deepcopy(mask)) segment["extra_material_refs"].append(mask["id"]) semantic_kind = "mask" lut_spec = clip.get("lut") if lut_spec is not None: lut_spec = _resolve_attachment_spec(ctx, lut_spec, "lut") lut = _attachment_material(lut_spec, "lut", ctx.new_id) lut["value"] = float(lut.pop("strength", 100)) / 100.0 lut.pop("skin_tone_correction", None) ctx.materials["effects"].append(lut) segment["extra_material_refs"].append(lut["id"]) if lut_spec.get("skin_tone_correction") is not None: lumi_hub_path = str(lut.get("lumi_hub_path") or "") if not lumi_hub_path: raise ValueError( "LUT skin_tone_correction requires an offline effect " "main_config with lumi_hub_path" ) effect_path = lumi_hub_path.rsplit("/", 1)[0] skin_tone = deepcopy(lut) skin_tone["id"] = ctx.new_id() skin_tone["type"] = "skin_tone_correction" skin_tone["version"] = "v3" skin_tone["value"] = float(lut_spec["skin_tone_correction"]) / 100.0 skin_tone["path"] = effect_path skin_tone["lumi_hub_path"] = effect_path ctx.materials["effects"].append(skin_tone) segment["extra_material_refs"].append(skin_tone["id"]) return semantic_kind def _track_object(ctx, name, track_type, segments, flag): return { "attribute": 0, "flag": flag, "id": ctx.new_id(), "is_default_name": True, "name": name, "segments": segments, "type": track_type, } def build_compound_video(ctx, clip, material, foreground_segment, track_name, ts, te): """Build duo-video's nested green-screen/compound draft structure.""" background_spec = clip["green_background"] chroma_spec = _resolve_attachment_spec(ctx, clip["chroma"], "chroma") duration_us = us(te - ts) full_material_duration_us = material["duration"] outer_source_timerange = deepcopy(foreground_segment["source_timerange"]) outer_target_timerange = deepcopy(foreground_segment["target_timerange"]) background_path = background_spec["source_path"] background_id = ctx.new_id() background = template("video") background.update({ "duration": full_material_duration_us, "height": int(background_spec.get("height") or ctx.height), "id": background_id, "material_name": os.path.basename(background_path), "path": background_path, "type": background_spec.get("type", "photo"), "width": int(background_spec.get("width") or ctx.width), }) chroma = _attachment_material(chroma_spec, "chroma", ctx.new_id) foreground_segment = deepcopy(foreground_segment) foreground_segment["source_timerange"] = timerange(0, full_material_duration_us) foreground_segment["target_timerange"] = timerange(0, full_material_duration_us) foreground_segment["extra_material_refs"].append(chroma["id"]) background_segment = base_segment( background_id, 0, full_material_duration_us, 1.0, [], ctx.new_id ) background_segment["source_timerange"] = timerange(0, full_material_duration_us) background_segment["render_index"] = 1 background_segment["track_render_index"] = 1 background_segment["clip"] = clip_from_segment(background_spec) draft = template("draft") draft["id"] = ctx.new_id() draft["combination_id"] = ctx.new_id() nested = draft["draft"] nested["id"] = ctx.new_id() nested["canvas_config"] = { "width": ctx.width, "height": ctx.height, "ratio": "original", } nested["duration"] = full_material_duration_us nested["fps"] = float(ctx.fps) scrub_platform_identity(nested) nested["materials"]["videos"] = [material, background] nested["materials"]["chromas"] = [chroma] referenced = set(foreground_segment["extra_material_refs"]) nested_speeds = [ item for item in ctx.materials["speeds"] if item.get("id") in referenced ] if nested_speeds: nested["materials"]["speeds"] = nested_speeds ctx.materials["speeds"] = [ item for item in ctx.materials["speeds"] if item.get("id") not in referenced ] nested["tracks"] = [ _track_object(ctx, "green_background", "video", [background_segment], 0), _track_object(ctx, "video", "video", [foreground_segment], 2), ] ctx.materials["drafts"].append(draft) combination_material = template("combination_video") combination_material.update({ "duration": full_material_duration_us, "height": ctx.height, "id": ctx.new_id(), "width": ctx.width, }) ctx.materials["videos"].append(combination_material) combination_segment = template("combination_segment") combination_segment.update({ "id": ctx.new_id(), "material_id": combination_material["id"], "extra_material_refs": [draft["id"]], "source_timerange": outer_source_timerange, "target_timerange": outer_target_timerange, }) semantic_kind = apply_video_attachments(ctx, clip, combination_segment) authored_track_name = "mask" if semantic_kind == "mask" else track_name ctx.add_segment( semantic_kind, authored_track_name, us(ts), duration_us, combination_segment ) def build_video_track(ctx, timeline_track): track_name = timeline_track.get("name", "video") for clip in timeline_track["clips"]: ts, te = float(clip["timeline_start"]), float(clip["timeline_end"]) ss, se = float(clip["source_start"]), float(clip["source_end"]) path = clip["source_path"] speed_value = float(clip.get("speed", 1.0)) reverse = bool(clip.get("reverse", False)) if reverse: # The contract allows omitting reverse_path so export_timeline_to_jianying # can generate it; building directly from such a clip is an authoring error. if "reverse_path" not in clip: raise ValueError("reverse video clips require a local reverse_path") path = clip["reverse_path"] src_dur_us, width, height = ctx.probe(path) if src_dur_us <= 0: raise ValueError(f"JianYing video source has no probed duration: {path}") mat_id = ctx.new_id() material = template("video") material.update({ "duration": int(src_dur_us), "height": height or ctx.height, "id": mat_id, "material_name": os.path.basename(path), "path": path, "width": width or ctx.width, }) audio = clip.get("audio", {}) keyframes = volume_keyframes(audio.get("volume_keyframes"), ts, ctx.new_id) volume = audio.get("base_gain", 1.0) if not keyframes else 1.0 seg = base_segment(mat_id, us(ts), us(te - ts), volume, keyframes, ctx.new_id) source_start = ss if reverse: source_start = max(0.0, (src_dur_us / 1_000_000) - se) seg["source_timerange"] = timerange(us(source_start), us(se - ss)) seg["speed"] = speed_value seg["extra_material_refs"] = [] if abs(speed_value - 1.0) > 1e-9: speed = speed_material(ctx.new_id, speed_value) ctx.materials["speeds"].append(speed) seg["extra_material_refs"].append(speed["id"]) seg["clip"] = clip_from_segment(clip) if clip.get("compound") or clip.get("green_background") or clip.get("chroma"): build_compound_video(ctx, clip, material, seg, track_name, ts, te) continue ctx.materials["videos"].append(material) semantic_kind = apply_video_attachments(ctx, clip, seg) authored_track_name = "mask" if semantic_kind == "mask" else track_name ctx.add_segment(semantic_kind, authored_track_name, us(ts), us(te - ts), seg) def build_audio_track(ctx, timeline_track): role = timeline_track.get("role", timeline_track.get("name", "audio")) track_name = timeline_track.get("name", role) for segment in timeline_track["segments"]: ts, te = float(segment["timeline_start"]), float(segment["timeline_end"]) path = segment["source_path"] mat_dur_us, _width, _height = ctx.probe(path) if mat_dur_us <= 0: raise ValueError(f"JianYing audio source has no probed duration: {path}") want_us = us(te - ts) speed_value = float(segment.get("speed", 1.0)) place_us = want_us required_source_us = int(round(want_us * speed_value)) if required_source_us > mat_dur_us: if role == "bgm" and timeline_track.get("loop"): place_us = want_us else: place_us = int(mat_dur_us / speed_value) if role == "bgm" and not timeline_track.get("loop"): ctx.note( f"BGM 素材({mat_dur_us/1e6:.1f}s) 短于时间线({(te - ts):.1f}s)," "剪映中未循环铺满(可在剪映里手动复制延长)" ) mat_id = ctx.new_id() material = template("audio") material.update({ "duration": int(mat_dur_us), "id": mat_id, "path": path, }) ctx.materials["audios"].append(material) keyframes = volume_keyframes(segment.get("volume_keyframes"), ts, ctx.new_id) volume = segment.get("gain", 1.0) if not keyframes else 1.0 if role == "bgm" and timeline_track.get("loop") and required_source_us > mat_dur_us: cursor = 0 while cursor < want_us: piece = min(int(mat_dur_us / speed_value), want_us - cursor) if piece <= 0: break piece_start_s = ts + (cursor / 1_000_000) piece_end_s = ts + ((cursor + piece) / 1_000_000) piece_kfs = windowed_volume_keyframes( segment.get("volume_keyframes"), piece_start_s, piece_end_s, segment.get("gain", 1.0), ctx.new_id) piece_volume = segment.get("gain", 1.0) if not piece_kfs else 1.0 piece_seg = audio_segment_piece( mat_id, us(ts) + cursor, piece, 0, int(round(piece * speed_value)), piece_volume, piece_kfs, ctx.new_id) piece_seg["speed"] = speed_value if abs(speed_value - 1.0) > 1e-9: speed = speed_material(ctx.new_id, speed_value) ctx.materials["speeds"].append(speed) piece_seg["extra_material_refs"].append(speed["id"]) ctx.add_segment("audio", track_name, us(ts) + cursor, piece, piece_seg) cursor += piece else: audio_seg = audio_segment_piece( mat_id, us(ts), place_us, 0, int(round(place_us * speed_value)), volume, keyframes, ctx.new_id) audio_seg["speed"] = speed_value if abs(speed_value - 1.0) > 1e-9: speed = speed_material(ctx.new_id, speed_value) ctx.materials["speeds"].append(speed) audio_seg["extra_material_refs"].append(speed["id"]) ctx.add_segment("audio", track_name, us(ts), place_us, audio_seg) def build_text_track(ctx, timeline_track): track_name = timeline_track.get("name", "text") track_kind = "subtitle" if track_name == "subtitle" else "text" for segment in timeline_track["segments"]: ts, te = float(segment["timeline_start"]), float(segment["timeline_end"]) text = segment["text"] mat_id = ctx.new_id() authored_style = dict(ctx.style_presets.get(str(segment.get("style_id")), {})) authored_style.update(segment.get("style") or {}) content = rich_text_content( text, authored_style, segment.get("words"), ctx.style_presets ) material = template("text") material.update({ "id": mat_id, "content": json.dumps(content, ensure_ascii=False), "type": "subtitle" if track_kind == "subtitle" else "text", "alignment": int(authored_style.get("text_align", 1)), "font_size": float(authored_style.get("font_size", 8.0)), "text_color": authored_style.get("fill_color", "#FFFFFF"), "line_spacing": float(authored_style.get("line_spacing", 0.02)), "letter_spacing": float(authored_style.get("letter_spacing", 0.0)), "check_flag": 15, }) bundle_resources = [] for content_style in content["styles"]: font_path = content_style.get("font", {}).get("path") if font_path and not str(font_path).startswith(("Resources/", "##_draftpath_placeholder_")): bundle_resources.append({ "source_path": str(font_path), "resource_kind": "fonts", }) effect_path = content_style.get("effect_style", {}).get("path") if effect_path and not str(effect_path).startswith(("Resources/", "##_draftpath_placeholder_")): bundle_resources.append({ "source_path": str(effect_path), "resource_kind": "effect", }) if bundle_resources: material["_bundle_resources"] = bundle_resources if authored_style.get("background_color"): material.update({ "background_color": authored_style["background_color"], "background_alpha": float(authored_style.get("background_opacity", 100)) / 100.0, "background_style": 1, "background_height": float(authored_style.get("background_height", 14)) / 100.0, "background_width": float(authored_style.get("background_width", 14)) / 100.0, "background_horizontal_offset": float(authored_style.get("background_offset_x", 50)) * 0.02 - 1, "background_vertical_offset": float(authored_style.get("background_offset_y", 50)) * 0.02 - 1, "background_round_radius": float(authored_style.get("background_radius", 6)) / 100.0, "check_flag": 31, }) if authored_style.get("stroke_color") and float(authored_style.get("stroke_width", 0)) > 0: material["bold_width"] = round( 0.00196 * (float(authored_style["stroke_width"]) ** 1.013), 6 ) material["border_color"] = authored_style["stroke_color"] if authored_style.get("shadow_color"): material.update({ "has_shadow": True, "shadow_color": authored_style["shadow_color"], "shadow_alpha": float(authored_style.get("shadow_opacity", 90)) / 100.0, "shadow_angle": float(authored_style.get("shadow_angle", -45)), "shadow_distance": float(authored_style.get("shadow_width", 5)), "shadow_smoothing": float(authored_style.get("shadow_vague", 45)) / 100.0, }) if authored_style or segment.get("words"): material["is_rich_text"] = True ctx.materials["texts"].append(material) duration_us = us(te - ts) seg = base_segment(mat_id, us(ts), duration_us, 1.0, [], ctx.new_id) seg["source_timerange"] = timerange(0, duration_us) seg["clip"] = clip_from_segment(segment) if track_kind == "subtitle" and "position" not in segment: seg["clip"]["transform"]["y"] = -0.72 ctx.add_segment(track_kind, track_name, us(ts), duration_us, seg) def build_image_track(ctx, timeline_track): """Build local image overlays as JianYing photo materials on video tracks.""" track_name = timeline_track.get("name", "image") for segment in timeline_track["segments"]: ts, te = float(segment["timeline_start"]), float(segment["timeline_end"]) duration_us = us(te - ts) path = segment["source_path"] _still_image_duration, width, height = ctx.probe(path) mat_id = ctx.new_id() material = template("video") material.update({ "duration": duration_us, "height": height or ctx.height, "id": mat_id, "material_name": os.path.basename(path), "path": path, "type": "photo", "width": width or ctx.width, }) ctx.materials["videos"].append(material) seg = base_segment(mat_id, us(ts), duration_us, 1.0, [], ctx.new_id) speed_value = float(segment.get("speed", 1.0)) seg["source_timerange"] = timerange(0, round(duration_us * speed_value)) seg["clip"] = clip_from_segment(segment) seg["speed"] = speed_value if abs(speed_value - 1.0) > 1e-9: speed = speed_material(ctx.new_id, speed_value) ctx.materials["speeds"].append(speed) seg["extra_material_refs"].append(speed["id"]) semantic_kind = apply_video_attachments(ctx, segment, seg) authored_track_name = "mask" if semantic_kind == "mask" else track_name if semantic_kind == "video": semantic_kind = "image" ctx.add_segment(semantic_kind, authored_track_name, us(ts), duration_us, seg) def build_timeline_track(ctx, timeline_track): kind = timeline_track["kind"] if kind in ("audio", "text") and not timeline_track["segments"]: return if kind == "video": build_video_track(ctx, timeline_track) elif kind == "audio": build_audio_track(ctx, timeline_track) elif kind == "text": build_text_track(ctx, timeline_track) elif kind == "image": build_image_track(ctx, timeline_track) else: # normalize_timeline admits only RESOURCE_TRACKS kinds past this point build_resource_track(ctx, timeline_track) -
model.py 3.4 KB
"""Thin internal model used at the JianYing adapter boundary.""" from dataclasses import dataclass, field from collections.abc import Callable from jianying.schema import us from jianying.tracks import TrackAllocator ProbeFn = Callable[[str], tuple[int, int, int]] NewIdFn = Callable[[], str] @dataclass class DraftBuildContext: """Normalized draft build state and overlap-safe track collection.""" width: int height: int fps: float total_us: int new_id: NewIdFn probe: ProbeFn resource_packages: dict = field(default_factory=dict) style_presets: dict = field(default_factory=dict) materials: dict[str, list] = field( default_factory=lambda: { key: [] for key in ( "audios", "chromas", "common_mask", "drafts", "effects", "masks", "speeds", "stickers", "text_templates", "texts", "transitions", "video_effects", "videos", ) } ) tracks: list[dict] = field(default_factory=list) notes: list[str] = field(default_factory=list) track_allocator: TrackAllocator = field(default_factory=TrackAllocator) @classmethod def from_timeline(cls, timeline, new_id, probe): """Build from a timeline already validated by jianying.timeline_contract.""" canvas = timeline["canvas"] return cls( width=canvas["width"], height=canvas["height"], fps=float(canvas["fps"]), total_us=us(timeline["duration"]), new_id=new_id, probe=probe, resource_packages=timeline.get("resource_packages", {}), style_presets=timeline.get("style_presets", {}), ) def add_segment(self, kind, base_name, start_us, duration_us, segment): allocated = self.track_allocator.allocate(kind, base_name, start_us, duration_us) track = next( ( item for item in self.tracks if item["_semantic_kind"] == allocated.kind and item["name"] == allocated.name ), None, ) if track is None: track = { "attribute": 0, "flag": 0, "id": self.new_id(), "is_default_name": True, "name": allocated.name, "segments": [], "type": allocated.track_type, "_semantic_kind": allocated.kind, "_layout_order": allocated.layout_order, } self.tracks.append(track) track["segments"].append(segment) return track def finalize_tracks(self): # Python's stable sort preserves the timeline's authored order inside # one semantic band (for example narration before BGM). self.tracks.sort(key=lambda item: item["_layout_order"]) type_counts = {} for track in self.tracks: count = type_counts.get(track["type"], 0) track["flag"] = 0 if count == 0 else 2 type_counts[track["type"]] = count + 1 del track["_semantic_kind"] del track["_layout_order"] return self.tracks def note(self, message): self.notes.append(message) -
optional.py 1012 B
"""Failure-isolated optional JianYing export invoked after canonical rendering.""" from pathlib import Path from lib import CONFIG, log def maybe_export_jianying(work_dir, out_dir, stem): """Lazy-import the optional 剪映 exporter and write a draft from timeline.json. The export is a documented fail-open sidecar: any failure is logged and never fails the already-rendered recap.""" try: from export_jianying import export_timeline_to_jianying from timeline import load_timeline parent = out_dir or CONFIG["jianying_draft_dir"] or str(work_dir) draft_dir, notes = export_timeline_to_jianying( load_timeline(Path(work_dir) / "timeline.json"), parent, draft_name=f"recap_{stem}", bundle_media=CONFIG["jianying_bundle_media"]) for n in notes: log(f" 注意: {n}") log(f"剪映草稿已导出: {draft_dir}") except Exception as exc: log(f" ⚠️ 剪映导出失败(不影响成片): {exc}") -
schema.py 2.5 KB
"""Schema constants and skeleton factories for JianYing export. This module is intentionally data-oriented: it owns draft version metadata and the full `materials` parallel-array shape. """ from jianying.templates import template # The full 剪映 materials object: ~45 parallel arrays. Only arrays backed by a # production builder are populated; retaining the full shape preserves compatibility. MATERIAL_KEYS = ["ai_translates", "audio_balances", "audio_effects", "audio_fades", "audio_track_indexes", "audios", "beats", "canvases", "chromas", "color_curves", "digital_humans", "drafts", "effects", "flowers", "green_screens", "handwrites", "hsl", "images", "log_color_wheels", "loudnesses", "manual_deformations", "masks", "common_mask", "material_animations", "material_colors", "multi_language_refs", "placeholders", "plugin_effects", "primary_color_wheels", "realtime_denoises", "shapes", "smart_crops", "smart_relights", "sound_channel_mappings", "speeds", "stickers", "tail_leaders", "text_templates", "texts", "time_marks", "transitions", "video_effects", "video_trackings", "videos", "vocal_beautifys", "vocal_separations"] def us(seconds): """Seconds (float) -> integer microseconds. The single seconds->µs boundary.""" return int(round(float(seconds) * 1_000_000)) def full_materials(filled): """Return a complete JianYing `materials` object with all known arrays.""" out = {k: [] for k in MATERIAL_KEYS} out.update(filled) return out def scrub_platform_identity(project): """Remove hardware fingerprints carried by the pinned upstream templates.""" for platform_key in ("last_modified_platform", "platform"): for identity_key in ("device_id", "hard_disk_id", "mac_address"): project[platform_key][identity_key] = "" return project def draft_content_skeleton(draft_id, width, height, fps, total_us, materials, tracks): """Build the root `draft_content.json` / `draft_info.json` skeleton.""" content = template("project") scrub_platform_identity(content) content["canvas_config"] = {"width": width, "height": height, "ratio": "original"} content["duration"] = int(total_us) content["fps"] = float(fps) content["id"] = draft_id content["materials"] = full_materials(materials) content["tracks"] = tracks return content def meta_info(draft_id, total_us): """Build the companion `draft_meta_info.json` skeleton.""" meta = template("meta") meta["draft_id"] = draft_id meta["draft_timeline_materials_size_"] = 0 meta["tm_duration"] = int(total_us) return meta -
templates.py 1.2 KB
"""Load the pinned duo-video JianYing protocol templates.""" import copy import json from functools import lru_cache from pathlib import Path _TEMPLATE_DIR = Path(__file__).resolve().parent.parent.parent / "references" / "jianying" _TEMPLATE_FILES = { "project": "empty_jy_project_info.json", "video": "empty_jy_material_video.json", "audio": "empty_yj_material_audio.json", "text": "empty_yj_material_text.json", "text_style": "empty_jy_text_styles.json", "segment": "empty_jy_segment.json", "draft": "empty_jy_draft.json", "combination_segment": "empty_jy_combination_segment.json", "combination_video": "empty_jy_combination_video_material.json", "meta": "empty_draft_meta_info.json", "meta_material": "empty_jy_meta_material_value.json", } @lru_cache(maxsize=None) def _read_template(name): try: filename = _TEMPLATE_FILES[name] except KeyError as exc: raise ValueError(f"unknown JianYing template: {name}") from exc with open(_TEMPLATE_DIR / filename, encoding="utf-8") as source: return json.load(source) def template(name): """Return a mutable deep copy of a pinned protocol template.""" return copy.deepcopy(_read_template(name)) -
timeline_contract.py 9.4 KB
"""Timeline validation and migration at the JianYing adapter boundary.""" import copy import math CURRENT_SCHEMA_VERSION = 2 RESOURCE_TRACK_KINDS = { "face_effect", "sound", "sticker", "text_template", "video_effect", } def _error(path, expectation): raise ValueError(f"invalid timeline {path}: {expectation}") def _is_number(value): return isinstance(value, (int, float)) and not isinstance(value, bool) and math.isfinite(value) def _field_path(path, key): return f"{path}.{key}" if path else key def _require_number(container, key, path, *, minimum=None): field_path = _field_path(path, key) if key not in container or not _is_number(container[key]): _error(field_path, "must be a finite number") value = container[key] if minimum is not None and value < minimum: _error(field_path, f"must be >= {minimum}") return value def _require_string(container, key, path): value = container.get(key) if not isinstance(value, str) or not value: _error(_field_path(path, key), "must be a non-empty string") return value def _validate_span(item, path, start_key="timeline_start", end_key="timeline_end"): start = _require_number(item, start_key, path, minimum=0) end = _require_number(item, end_key, path, minimum=0) if end <= start: _error(f"{path}.{end_key}", f"must be greater than {start_key}") def _validate_transform(item, path): for field in ("scale", "position", "flip"): if field in item and not isinstance(item[field], dict): _error(f"{path}.{field}", "must be an object") def _validate_resources(resources, path): if not isinstance(resources, list) or any(not isinstance(item, dict) for item in resources): _error(path, "must contain source_path objects") for index, item in enumerate(resources): _require_string(item, "source_path", f"{path}[{index}]") def _validate_video_clip(clip, path): if not isinstance(clip, dict): _error(path, "must be an object") _require_string(clip, "source_path", path) _validate_span(clip, path) _validate_span(clip, path, "source_start", "source_end") if "audio" in clip and not isinstance(clip["audio"], dict): _error(f"{path}.audio", "must be an object") if "speed" in clip: speed = _require_number(clip, "speed", path) if speed <= 0: _error(f"{path}.speed", "must be greater than 0") source_duration = float(clip["source_end"]) - float(clip["source_start"]) target_duration = float(clip["timeline_end"]) - float(clip["timeline_start"]) expected_source_duration = target_duration * float(speed) if not math.isclose(source_duration, expected_source_duration, rel_tol=1e-6, abs_tol=1e-4): _error( path, "source duration must equal target duration multiplied by speed " f"({source_duration} != {target_duration} * {speed})", ) if "reverse" in clip and not isinstance(clip["reverse"], bool): _error(f"{path}.reverse", "must be a boolean") # A reversed clip may omit reverse_path: export_timeline_to_jianying generates it. if "reverse_path" in clip: _require_string(clip, "reverse_path", path) _validate_transform(clip, path) for field in ("transition", "mask", "lut", "chroma"): if field not in clip: continue spec = clip[field] if isinstance(spec, dict): if "resources" in spec: _validate_resources(spec["resources"], f"{path}.{field}.resources") elif not isinstance(spec, str): _error(f"{path}.{field}", "must be an object or resource-package name") if "compound" in clip and not isinstance(clip["compound"], bool): _error(f"{path}.compound", "must be a boolean") if clip.get("compound") or "green_background" in clip or "chroma" in clip: # Any one of these makes the clip a green-screen compound, which needs both. background = clip.get("green_background") if not isinstance(background, dict): _error(f"{path}.green_background", "must be a local media object") _require_string(background, "source_path", f"{path}.green_background") _validate_transform(background, f"{path}.green_background") if "chroma" not in clip: _error(f"{path}.chroma", "compound green-screen clips require a chroma object") def _validate_resource_config(config, path): if not isinstance(config, dict): _error(path, "must be an object") if not isinstance(config.get("main_config"), dict): _error(f"{path}.main_config", "must be an object") _validate_resources(config.get("resources", []), f"{path}.resources") def _validate_segment(segment, path, kind): if not isinstance(segment, dict): _error(path, "must be an object") _validate_span(segment, path) _validate_transform(segment, path) if "speed" in segment: speed = _require_number(segment, "speed", path) if speed <= 0: _error(f"{path}.speed", "must be greater than 0") if kind in {"audio", "image"}: _require_string(segment, "source_path", path) elif kind == "text" and not isinstance(segment.get("text"), str): _error(f"{path}.text", "must be a string") elif kind in RESOURCE_TRACK_KINDS: sources = [key for key in ("material", "resource_config", "resource_package") if key in segment] if len(sources) != 1: _error(path, "must define exactly one of material, resource_config, or resource_package") source = segment[sources[0]] if sources[0] == "resource_package": if not isinstance(source, str) or not source: _error(f"{path}.resource_package", "must be a non-empty string") elif sources[0] == "material" and not isinstance(source, dict): _error(f"{path}.material", "must be an object") elif sources[0] == "resource_config": _validate_resource_config(source, f"{path}.resource_config") def _validate_track(track, path): if not isinstance(track, dict): _error(path, "must be an object") kind = _require_string(track, "kind", path) if "name" in track and (not isinstance(track["name"], str) or not track["name"]): _error(f"{path}.name", "must be a non-empty string") if kind == "video": clips = track.get("clips") if not isinstance(clips, list): _error(f"{path}.clips", "must be an array") for index, clip in enumerate(clips): _validate_video_clip(clip, f"{path}.clips[{index}]") return if kind in {"audio", "image", "text"} | RESOURCE_TRACK_KINDS: segments = track.get("segments") if not isinstance(segments, list): _error(f"{path}.segments", "must be an array") if kind == "audio": if "role" in track and not isinstance(track["role"], str): _error(f"{path}.role", "must be a string") if "loop" in track and not isinstance(track["loop"], bool): _error(f"{path}.loop", "must be a boolean") for index, segment in enumerate(segments): _validate_segment(segment, f"{path}.segments[{index}]", kind) return _error(f"{path}.kind", f"unsupported track kind {kind!r}") def _validate_v2(timeline): canvas = timeline.get("canvas") if not isinstance(canvas, dict): _error("canvas", "must be an object") for dimension in ("width", "height"): value = canvas.get(dimension) if not isinstance(value, int) or isinstance(value, bool) or value <= 0: _error(f"canvas.{dimension}", "must be a positive integer") fps = _require_number(canvas, "fps", "canvas") if fps <= 0: _error("canvas.fps", "must be greater than 0") _require_number(timeline, "duration", "", minimum=0) resource_packages = timeline.get("resource_packages", {}) if not isinstance(resource_packages, dict): _error("resource_packages", "must be an object") for name, config in resource_packages.items(): _validate_resource_config(config, f"resource_packages.{name}") if "style_presets" in timeline and not isinstance(timeline["style_presets"], dict): _error("style_presets", "must be an object") tracks = timeline.get("tracks") if not isinstance(tracks, list): _error("tracks", "must be an array") for index, track in enumerate(tracks): _validate_track(track, f"tracks[{index}]") def normalize_timeline(timeline): """Return a validated schema-v2 copy, migrating schema v1 when necessary.""" if not isinstance(timeline, dict): _error("root", "must be an object") schema_version = timeline.get("schema_version") if not isinstance(schema_version, int) or isinstance(schema_version, bool): _error("schema_version", "must be integer 1 or 2") if schema_version not in {1, CURRENT_SCHEMA_VERSION}: raise ValueError( f"unsupported timeline schema_version {schema_version}; " f"supported versions are 1 and {CURRENT_SCHEMA_VERSION}" ) normalized = copy.deepcopy(timeline) if schema_version == 1: # Schema v2 is an additive extension of v1 (local image tracks). The # migration therefore preserves all authored v1 fields and only advances # the version before applying the current contract. normalized["schema_version"] = CURRENT_SCHEMA_VERSION _validate_v2(normalized) return normalized -
tracks.py 2.6 KB
"""Semantic track ordering and overlap-safe allocation for JianYing export. The order mirrors duo-video's authoring layout. It is used only to order track objects; JianYing segment ``render_index`` remains the schema default and must not be confused with this semantic layout value. """ from dataclasses import dataclass @dataclass(frozen=True) class TrackBand: kind: str track_type: str layout_order: int description: str SEGMENT_RENDER_INDEX = 2 TRACK_LAYOUT_BANDS = { "sound": TrackBand("sound", "audio", 10_000, "sound effects"), "audio": TrackBand("audio", "audio", 20_000, "narration, music, and general audio"), "green_screen": TrackBand("green_screen", "video", 30_000, "green-screen background"), "video": TrackBand("video", "video", 40_000, "base video"), "image": TrackBand("image", "video", 50_000, "image and photo overlays"), "mask": TrackBand("mask", "video", 60_000, "masks"), "effect": TrackBand("effect", "effect", 70_000, "video effects"), "video_effect": TrackBand("video_effect", "effect", 70_000, "video effects"), "face_effect": TrackBand("face_effect", "effect", 70_010, "face effects"), "sticker": TrackBand("sticker", "sticker", 80_000, "stickers"), "subtitle": TrackBand("subtitle", "text", 90_000, "subtitles"), "text": TrackBand("text", "text", 100_000, "plain text"), "text_template": TrackBand("text_template", "text", 110_000, "text templates"), } @dataclass(frozen=True) class AllocatedTrack: kind: str name: str track_type: str layout_order: int class TrackAllocator: """Allocate deterministic suffix tracks when same-name segments overlap. Intervals are half-open, so adjacent segments reuse a track while true overlap creates ``name-1``, ``name-2``, and so on. """ def __init__(self): self._occupied = {} @staticmethod def _overlaps(start_us, duration_us, existing): end_us = int(start_us) + int(duration_us) return any(int(start_us) < old_end and end_us > old_start for old_start, old_end in existing) def allocate(self, kind, base_name, start_us, duration_us): band = TRACK_LAYOUT_BANDS[kind] suffix = 0 while True: name = base_name if suffix == 0 else f"{base_name}-{suffix}" key = (kind, name) occupied = self._occupied.setdefault(key, []) if not self._overlaps(start_us, duration_us, occupied): occupied.append((int(start_us), int(start_us) + int(duration_us))) return AllocatedTrack(kind, name, band.track_type, band.layout_order + suffix) suffix += 1 -
writer.py 14.9 KB
"""Safe writer and portable media bundler for JianYing draft folders.""" import hashlib import json import os import shutil import tempfile import time import uuid import zipfile from jianying.templates import template DRAFT_PATH_PLACEHOLDER = "##_draftpath_placeholder_0E685133-18CE-45ED-8CB8-2904A212EC80_##" def validate_draft_name(draft_name): """Reject draft names that could escape or alias the requested parent dir.""" if not isinstance(draft_name, str): raise TypeError("draft_name must be a string") if not draft_name or not draft_name.strip(): raise ValueError("draft_name must not be empty") if os.path.isabs(draft_name): raise ValueError("draft_name must be a plain folder name, not an absolute path") if "/" in draft_name or "\\" in draft_name: raise ValueError("draft_name must not contain path separators") if draft_name in {".", ".."}: raise ValueError("draft_name must not be '.' or '..'") def _resource_kind(material, materials_key): if materials_key == "audios": return "audio", "music" if material.get("type") == "photo": return "image", "photo" return "video", "video" RESOURCE_DIRECTORY_BY_MATERIALS_KEY = { "chromas": "effect", "common_mask": "mask", "effects": "effect", "masks": "mask", "stickers": "sticker", "texts": "text", "text_templates": "text_template", "transitions": "transition", "video_effects": "effect", } def _unused_name(directory, basename, used): """Return a collision-free filename within one resource directory.""" name = basename stem, ext = os.path.splitext(basename) suffix = 1 while name in used or os.path.exists(os.path.join(directory, name)): name = f"{stem}_{suffix}{ext}" suffix += 1 used.add(name) return name def _md5(path): digest = hashlib.md5(usedforsecurity=False) with open(path, "rb") as source: for chunk in iter(lambda: source.read(1024 * 1024), b""): digest.update(chunk) return digest.hexdigest() def _meta_value(material, relative_path, metetype, copied_path, timestamp_ms): duration = int(material.get("duration") or 0) timestamp_s = timestamp_ms // 1000 value = template("meta_material") value.update({ "duration": duration, "height": int(material.get("height") or 0), # duo-video indexes imported local resources independently from the # draft-content material IDs. "id": str(uuid.uuid4()).upper(), "md5": _md5(copied_path), "metetype": metetype, "type": 0, "width": int(material.get("width") or 0), "create_time": timestamp_s, "extra_info": os.path.basename(relative_path), "file_Path": f"./{relative_path}", "import_time": timestamp_s, "import_time_ms": timestamp_ms, "item_source": 1, "roughcut_time_range": {"duration": duration, "start": 0}, "sub_time_range": {"duration": -1, "start": -1}, }) return value def _material_sets(content): """Yield root and nested compound-draft material dictionaries.""" materials = content["materials"] yield materials for draft in materials["drafts"]: yield from _material_sets(draft["draft"]) def _replace_value(value, old, new): if isinstance(value, dict): for key, item in value.items(): value[key] = _replace_value(item, old, new) elif isinstance(value, list): for index, item in enumerate(value): value[index] = _replace_value(item, old, new) elif isinstance(value, str): if value == old: return new return value def _replace_material_resource_path(material, materials_key, old, new): """Rewrite one declared resource, including the known rich-text JSON field.""" _replace_value(material, old, new) if materials_key != "texts" or not isinstance(material.get("content"), str): return content = json.loads(material["content"]) _replace_value(content, old, new) material["content"] = json.dumps(content, ensure_ascii=False, separators=(",", ":")) def _safe_resource_target(target_path): normalized = os.path.normpath(str(target_path).replace("\\", "/")) if normalized in {"", "."} or os.path.isabs(normalized): raise ValueError(f"invalid JianYing resource target_path: {target_path}") if normalized == ".." or normalized.startswith("../"): raise ValueError(f"JianYing resource target_path escapes package: {target_path}") return normalized def _extract_zip(source, destination): os.makedirs(destination, exist_ok=False) destination_real = os.path.realpath(destination) with zipfile.ZipFile(source) as archive: for member in archive.infolist(): member_path = os.path.realpath(os.path.join(destination, member.filename)) if os.path.commonpath((destination_real, member_path)) != destination_real: raise ValueError(f"unsafe path in JianYing resource archive: {member.filename}") archive.extractall(destination) def _copy_resource(source, resource_kind, resources_root, used, target_path=None): resource_dir = os.path.join(resources_root, resource_kind) os.makedirs(resource_dir, exist_ok=True) is_zip = os.path.isfile(source) and zipfile.is_zipfile(source) if target_path is None: basename = os.path.basename(source.rstrip(os.sep)) if is_zip: basename = os.path.splitext(basename)[0] relative_target = _unused_name(resource_dir, basename, used[resource_kind]) else: relative_target = _safe_resource_target(target_path) copied_path = os.path.join(resource_dir, relative_target) copied_real = os.path.realpath(copied_path) if os.path.commonpath((os.path.realpath(resource_dir), copied_real)) != os.path.realpath(resource_dir): raise ValueError(f"JianYing resource target escapes package: {target_path}") os.makedirs(os.path.dirname(copied_path), exist_ok=True) if os.path.exists(copied_path): raise FileExistsError(f"duplicate JianYing resource target: {relative_target}") if is_zip: _extract_zip(source, copied_path) elif os.path.isdir(source): shutil.copytree(source, copied_path) else: shutil.copy2(source, copied_path) relative_path = f"Resources/local/{resource_kind}/{relative_target.replace(os.sep, '/')}" return copied_path, relative_path, f"{DRAFT_PATH_PLACEHOLDER}/{relative_path}" def _is_packaged_path(value): return str(value).startswith((DRAFT_PATH_PLACEHOLDER, "Resources/", "./Resources/")) def _descriptor(raw, default_kind, *, required): """`raw` is a contract-validated resource entry: an object with a non-empty source_path.""" source = raw["source_path"] resource_kind = raw.get("resource_kind", default_kind) target_path = raw.get("target_path") if not isinstance(resource_kind, str) or resource_kind not in { "audio", "effect", "fonts", "image", "lut", "mask", "sticker", "text", "text_template", "transition", "video", }: raise ValueError(f"invalid JianYing resource kind: {resource_kind}") return { "source_path": source, "resource_kind": resource_kind, "target_path": target_path, "required": required, } def _material_resource_descriptors(material, default_kind): descriptors = [ _descriptor(raw, default_kind, required=True) for raw in material.get("_bundle_resources", []) ] path = material.get("path") if isinstance(path, str) and path and not _is_packaged_path(path): if not any(item["source_path"] == path for item in descriptors): descriptors.append({ "source_path": path, "resource_kind": default_kind, "target_path": None, "required": False, }) return descriptors def bundle_media(content, meta, draft_dir): """Copy media into duo-video's Resources/local contract and index it. Materials that reference the same source file and resource kind share one copied file and one meta entry. Missing sources remain untouched and are reported to the caller so a non-portable reference is never disguised as a successfully bundled one. """ resources_root = os.path.join(draft_dir, "Resources", "local") copied = {} resource_kinds = { "audio", "effect", "fonts", "image", "lut", "mask", "sticker", "text", "text_template", "transition", "video", } used = {kind: set() for kind in resource_kinds} meta_values = [] notes = [] timestamp_ms = int(time.time() * 1000) material_sets = list(_material_sets(content)) for materials in material_sets: for materials_key in ("videos", "audios"): for material in materials.get(materials_key, []): src = material.get("path") if not src: continue resource_kind, metetype = _resource_kind(material, materials_key) source_key = (resource_kind, os.path.realpath(src)) existing = copied.get(source_key) if existing is not None: material["path"] = existing["draft_path"] continue if not os.path.isfile(src): if not _is_packaged_path(src): notes.append(f"素材缺失,未打包: {src}") continue copied_path, relative_path, draft_path = _copy_resource( src, resource_kind, resources_root, used ) material["path"] = draft_path copied[source_key] = {"draft_path": draft_path} meta_values.append( _meta_value(material, relative_path, metetype, copied_path, timestamp_ms) ) for materials_key, resource_kind in RESOURCE_DIRECTORY_BY_MATERIALS_KEY.items(): for material in materials.get(materials_key, []): descriptors = _material_resource_descriptors( material, resource_kind ) seen_descriptors = set() for descriptor in descriptors: src = descriptor["source_path"] if _is_packaged_path(src): continue descriptor_key = ( descriptor["resource_kind"], os.path.realpath(src), descriptor["target_path"], ) if descriptor_key in seen_descriptors: continue seen_descriptors.add(descriptor_key) if not os.path.exists(src): message = f"声明的剪映资源缺失: {src}" if descriptor["required"]: raise ValueError(message) notes.append(message) continue kind = descriptor["resource_kind"] source_key = (kind, os.path.realpath(src), descriptor["target_path"]) existing = copied.get(source_key) if existing is None: _copied_path, _relative_path, draft_path = _copy_resource( src, kind, resources_root, used, target_path=descriptor["target_path"], ) copied[source_key] = {"draft_path": draft_path} else: draft_path = existing["draft_path"] _replace_material_resource_path( material, materials_key, src, draft_path ) material_group = next(group for group in meta["draft_materials"] if group["type"] == 0) material_group["value"] = meta_values meta["draft_timeline_materials_size_"] = sum( os.path.getsize(os.path.join(draft_dir, value["file_Path"][2:])) for value in meta_values ) return notes def strip_internal_resource_fields(content): for materials in _material_sets(content): for entries in materials.values(): for material in entries: material.pop("_bundle_resources", None) def draft_dir_has_user_content(draft_dir): """Return True when writing here could overwrite an existing draft/material.""" if not os.path.exists(draft_dir): return False try: return any(os.scandir(draft_dir)) except OSError: return True def collision_safe_draft_dir(out_dir, draft_name): """Pick a fresh draft folder instead of overwriting an existing non-empty one.""" validate_draft_name(draft_name) base = os.path.join(out_dir, draft_name) if not draft_dir_has_user_content(base): return base, draft_name idx = 2 while True: candidate_name = f"{draft_name}_{idx}" candidate = os.path.join(out_dir, candidate_name) if not draft_dir_has_user_content(candidate): return candidate, candidate_name idx += 1 def write_draft(content, meta, notes, out_dir, draft_name, bundle_media_enabled=False): """Atomically write the three JianYing draft JSON files and optional bundle.""" validate_draft_name(draft_name) out_dir = os.path.abspath(out_dir) os.makedirs(out_dir, exist_ok=True) draft_dir, actual_name = collision_safe_draft_dir(out_dir, draft_name) notes = list(notes) if actual_name != draft_name: notes.append(f"草稿目录已存在,改写为 {actual_name} 以避免覆盖") tmp_parent = tempfile.mkdtemp(prefix=f".{actual_name}.", dir=out_dir) tmp_dir = os.path.join(tmp_parent, actual_name) try: os.makedirs(tmp_dir, exist_ok=False) if bundle_media_enabled: notes.extend(bundle_media(content, meta, tmp_dir)) strip_internal_resource_fields(content) timestamp_ms = int(time.time() * 1000) meta["draft_name"] = actual_name meta["draft_fold_path"] = draft_dir meta["tm_draft_create"] = timestamp_ms meta["tm_draft_modified"] = timestamp_ms content["name"] = actual_name for fname in ("draft_content.json", "draft_info.json"): with open(os.path.join(tmp_dir, fname), "w", encoding="utf-8") as f: json.dump(content, f, ensure_ascii=False, indent=2) with open(os.path.join(tmp_dir, "draft_meta_info.json"), "w", encoding="utf-8") as f: json.dump(meta, f, ensure_ascii=False, indent=2) if os.path.isdir(draft_dir) and not draft_dir_has_user_content(draft_dir): os.rmdir(draft_dir) os.replace(tmp_dir, draft_dir) except Exception: shutil.rmtree(tmp_parent, ignore_errors=True) raise finally: if os.path.exists(tmp_parent): shutil.rmtree(tmp_parent, ignore_errors=True) return draft_dir, notes -
__init__.py 49 B
"""Jianying (剪映) draft export subpackage."""
-
-
subtitles
-
core.py 13.6 KB
"""Subtitle text shaping, timing, and measured-canvas geometry.""" import os import re from decimal import Decimal from lib import CONFIG from assemble_constants import ( SUBTITLE_STYLE_REF_H, SUBTITLE_STYLE_REF_W, _SUBTITLE_CLOSING_QUOTES, _SUBTITLE_TERMINAL_PUNCTUATION, ) from media import _ratio_to_float def _seconds_to_srt_time(seconds): """Floor times to SRT milliseconds without float remainder artifacts. Coercing first keeps Fraction/Decimal/str inputs working and clamps a negative time to zero instead of emitting a negative-component SRT stamp. """ seconds = max(0.0, float(seconds)) total_ms = int(Decimal(str(seconds)) * 1000) h, remainder = divmod(total_ms, 3_600_000) m, remainder = divmod(remainder, 60_000) s, ms = divmod(remainder, 1000) return f"{h:02d}:{m:02d}:{s:02d},{ms:03d}" def _seconds_to_ass_time(seconds): """将秒数转为 ASS 时间格式 H:MM:SS.cc""" centiseconds = int(round(float(seconds) * 100)) h = centiseconds // 360000 centiseconds %= 360000 m = centiseconds // 6000 centiseconds %= 6000 s = centiseconds // 100 cs = centiseconds % 100 return f"{h}:{m:02d}:{s:02d}.{cs:02d}" def _subtitle_style_config(canvas=None): """Return the internal default burn-in subtitle style. When ``canvas`` ({"width","height"}) is given AND the user has not pinned PlayRes via SUBTITLE_PLAY_RES_X/Y, the style is scaled to that canvas: PlayRes is set to the frame dimensions (so libass never stretches glyphs — the old hardcoded 1280x720 squished portrait text), horizontal metrics scale with width, vertical metrics with height, and the font is additionally capped so a full ``max_chars`` line fits the usable width. A 16:9 source (or the 1280x720 default) reproduces the legacy values exactly. """ style = { "font_name": CONFIG["subtitle_font_name"], "font_file": CONFIG["subtitle_font_file"], "font_size": CONFIG["subtitle_font_size"], "primary_color": CONFIG["subtitle_primary_color"], "outline_color": CONFIG["subtitle_outline_color"], "outline": CONFIG["subtitle_outline"], "shadow": CONFIG["subtitle_shadow"], "alignment": CONFIG["subtitle_alignment"], "margin_l": CONFIG["subtitle_margin_l"], "margin_r": CONFIG["subtitle_margin_r"], "margin_v": CONFIG["subtitle_margin_v"], "max_chars": CONFIG["subtitle_max_chars"], "play_res_x": CONFIG["subtitle_play_res_x"], "play_res_y": CONFIG["subtitle_play_res_y"], } pinned = "SUBTITLE_PLAY_RES_X" in os.environ or "SUBTITLE_PLAY_RES_Y" in os.environ if canvas is None or pinned: return style # legacy / manually-pinned: unchanged cw, ch = canvas["width"], canvas["height"] base_font = float(style["font_size"]) kx = cw / float(SUBTITLE_STYLE_REF_W) # horizontal metrics ∝ width ky = ch / float(SUBTITLE_STYLE_REF_H) # vertical metrics ∝ height margin_l = round(float(style["margin_l"]) * kx) margin_r = round(float(style["margin_r"]) * kx) margin_v = round(float(style["margin_v"]) * ky) # height-proportional size, then cap so a full line of CJK glyphs (≈1em wide) fits the # usable width — this is what keeps portrait text on-screen instead of overflowing. usable_w = max(1.0, cw - margin_l - margin_r) width_cap = usable_w / int(style["max_chars"]) font_size = max(1, int(min(base_font * ky, width_cap))) # floor so a full line never overflows font_scale = font_size / base_font style.update({ "font_size": font_size, "outline": max(0, round(float(style["outline"]) * font_scale)), "shadow": max(0, round(float(style["shadow"]) * font_scale)), "margin_l": margin_l, "margin_r": margin_r, "margin_v": margin_v, "play_res_x": cw, "play_res_y": ch, }) return style def _measured_subtitle_band(canvas): """Explicit (y_top, y_bot) display-frame subtitle band validated against the canvas, or None.""" y_top = CONFIG["subtitle_y_top"] y_bot = CONFIG["subtitle_y_bot"] if y_top < 0 and y_bot < 0: return None sar_text = canvas["sample_aspect_ratio"] if abs(_ratio_to_float(sar_text, 0.0) - 1.0) >= 1e-9: raise ValueError( f"字幕带坐标仅支持方形像素画布 (SAR 1:1);当前 SAR={sar_text}" ) canvas_h = canvas["height"] if not 0 <= y_top < y_bot <= canvas_h: raise ValueError( f"字幕带坐标无效: top={y_top}, bot={y_bot}, 画布高度={canvas_h};" "必须满足 0 <= top < bot <= height" ) return y_top, y_bot def _style_for_measured_subtitle_band(style, canvas): """Fit the ASS baseline and font into explicit auto-rotated display-frame Y coordinates.""" style = dict(style) safe_area = _measured_subtitle_safe_area(style, canvas) if safe_area is None: return style alignment = style["alignment"] if alignment not in {1, 2, 3}: raise ValueError( "measured subtitle coordinates require a bottom-aligned ASS style " f"(SUBTITLE_ALIGNMENT 1/2/3); got {alignment}" ) canvas_h = canvas["height"] scale_y = float(style["play_res_y"]) / canvas_h style["margin_v"] = max(0, round((canvas_h - CONFIG["subtitle_y_bot"]) * scale_y)) current_font = int(style["font_size"]) current_outline = float(style["outline"]) current_shadow = float(style["shadow"]) available_height = safe_area["height"] for candidate in range(current_font, 7, -1): scale = candidate / current_font outline = max(1 if current_outline > 0 else 0, round(current_outline * scale)) shadow = max(0, round(current_shadow * scale)) if candidate * 1.25 + outline * 2 + shadow <= available_height + 1e-6: fitted_font = candidate break else: # Keep the renderer's minimum readable size; visual QC will block because it cannot fit. fitted_font = min(current_font, 8) if fitted_font < current_font: scale = fitted_font / current_font style["font_size"] = fitted_font style["outline"] = max( 1 if current_outline > 0 else 0, round(current_outline * scale) ) style["shadow"] = max(0, round(current_shadow * scale)) return style def _measured_subtitle_safe_area(style, canvas): """Return the padded measured band in ASS PlayRes coordinates, or None.""" band = _measured_subtitle_band(canvas) if band is None: return None y_top, y_bot = band canvas_h = canvas["height"] safe_top = max(0, y_top - CONFIG["subtitle_mask_padding"]) # The ASS style remains bottom-anchored at the measured y_bot. Bottom mask padding hides # source glyph edges but is not usable subtitle layout space; only top padding can extend # the line box without moving its baseline below the measured band. play_x = int(style["play_res_x"]) scale_y = int(style["play_res_y"]) / canvas_h margin_l = int(style["margin_l"]) margin_r = int(style["margin_r"]) return { "x": margin_l, "y": round(safe_top * scale_y), "width": max(1, play_x - margin_l - margin_r), "height": max(1, round((y_bot - safe_top) * scale_y)), "bottom_margin": max(0, round((canvas_h - y_bot) * scale_y)), } def _subtitle_display_text(text): """Return display-only subtitle text with trailing sentence punctuation removed. Narration/TTS source text stays untouched; this is applied only to SRT/ASS cue text. Closing quotes/brackets are preserved, so 「原声台词。」 renders as 「原声台词」. """ text = text.strip() suffix = "" while text and text[-1] in _SUBTITLE_CLOSING_QUOTES: suffix = text[-1] + suffix text = text[:-1].rstrip() text = text.rstrip(_SUBTITLE_TERMINAL_PUNCTUATION).rstrip() return (text + suffix).strip() def _subtitle_chunk_weight(text): """Weight raw subtitle chunks for timing, independent of display punctuation cleanup.""" return max(1, len(re.sub(r"\s+", "", text))) def _subtitle_entry_chunks(raw_chunks): """Pair raw chunks used for timing with their final display text. Timing remains based on the raw split topology. Terminal punctuation is stripped only on the emitted text, while quote-only suffix chunks are folded into the previous cue so a closing bracket never renders alone. """ out = [] for chunk in raw_chunks: display = _subtitle_display_text(chunk) if not display: continue if all(ch in _SUBTITLE_CLOSING_QUOTES for ch in display): if out: out[-1]["text"] += display continue out.append({"raw": chunk, "text": display}) return out def _normalize_subtitle_text(text): """Normalize Chinese em-dashes in burned subtitle text: a run of one-or-more "—" (incl. "——") collapses to a single ",". Then collapse any resulting double commas (",,"→",") so the dash swap never leaves a doubled comma.""" return re.sub(r",{2,}", ",", re.sub(r"—+", ",", text)) def _split_subtitle_chunks(text, max_chars): """Split one narration block (often several sentences) into short display chunks. A block is synthesized as one continuous TTS utterance for fluent prosody, but showing the whole paragraph as a single subtitle would force a tall multi-line band and lag the picture. So we cut the block at punctuation into clauses, then greedily pack adjacent clauses into chunks of at most `max_chars` — each chunk renders as ONE readable line synced to its slice of the block's audio. Punctuation stays attached here for lossless splitting; the display layer strips terminal sentence marks per subtitle-cue style.""" text = text.strip() if not text: return [] breakers = ",。!?、;:…—,.!?;:" clauses, buf = [], "" for ch in text: buf += ch if ch in breakers: clauses.append(buf) buf = "" if buf.strip(): clauses.append(buf) # Any single clause longer than max_chars is hard-wrapped so no chunk ever exceeds one line. # Balance those pieces instead of slicing exactly at max_chars: a 21-character clause must # not become a readable 20-character cue followed by a 1-character flash. sized = [] for clause in clauses: if len(clause) <= max_chars: sized.append(clause) else: piece_count = (len(clause) + max_chars - 1) // max_chars base, extra = divmod(len(clause), piece_count) cursor = 0 for piece_index in range(piece_count): width = base + (1 if piece_index < extra else 0) sized.append(clause[cursor:cursor + width]) cursor += width chunks, cur = [], "" for clause in sized: sentence_closed = cur.rstrip().endswith(tuple(_SUBTITLE_TERMINAL_PUNCTUATION)) if cur and (sentence_closed or len(cur) + len(clause) > max_chars): chunks.append(cur) cur = clause else: cur += clause if cur.strip(): chunks.append(cur) return [c.strip() for c in chunks if c.strip()] def _subtitle_entries(narration): """Collect subtitle entries from final TTS segment placement. Each placed segment is split into short one-line chunks and its played window [actual_place_start, actual_place_end] is distributed across them in proportion to character count — karaoke-style timing that keeps each line on screen only while it is roughly being spoken, instead of holding a whole paragraph for the segment's full duration. Segments that were not placed have a zero-width window and therefore produce no cue.""" max_chars = CONFIG["subtitle_max_chars"] entries = [] for seg in narration: text = seg["spoken_text"] start, end = float(seg["actual_place_start"]), float(seg["actual_place_end"]) entries.extend(_distribute_chunks(_split_subtitle_chunks(text, max_chars), start, end)) return entries def _distribute_chunks(chunks, start, end): """Distribute [start,end] across raw chunks while emitting display-clean text. Terminal subtitle punctuation is visual-only: it is stripped from final cue text, but the raw split chunks remain the timing topology. A slice too short to show on its own is folded into the previous line of the same block, so no chunk is ever dropped. """ chunks = _subtitle_entry_chunks(chunks) if not chunks or end - start < 0.1: return [] if len(chunks) == 1: return [{"start": start, "end": end, "text": chunks[0]["text"]}] total_chars = sum(_subtitle_chunk_weight(c["raw"]) for c in chunks) span = end - start out, cursor = [], start for i, chunk in enumerate(chunks): weight = _subtitle_chunk_weight(chunk["raw"]) chunk_end = end if i == len(chunks) - 1 else cursor + span * (weight / total_chars) if out and chunk_end - cursor < 0.05: out[-1]["text"] += chunk["text"] out[-1]["end"] = chunk_end else: out.append({"start": cursor, "end": chunk_end, "text": chunk["text"]}) cursor = chunk_end return out def _bracketed_original_chunks(text, start, end, max_chars): """Split original dialogue into timed chunks wrapped in 「」 for visual distinction.""" raw = text.strip() if raw.startswith("「") and raw.endswith("」"): raw = raw[1:-1].strip() chunks = _split_subtitle_chunks(raw, max_chars) if chunks: chunks[0] = "「" + chunks[0] chunks[-1] = chunks[-1] + "」" return _distribute_chunks(chunks, start, end) -
render.py 3.5 KB
"""SRT/ASS serialization for narration and original-dialogue subtitles.""" from source_subtitles import _combined_subtitle_entries from subtitles.core import ( _normalize_subtitle_text, _seconds_to_ass_time, _seconds_to_srt_time, _style_for_measured_subtitle_band, _subtitle_style_config, ) def _generate_srt(narration, work_dir, video_duration): """将解说脚本转为 SRT 字幕文件,使用实际音频放置时间;原声留白处补烧原声字幕。""" srt_lines = [] # entries are already split into short one-line chunks, so no wrapping here. for idx, entry in enumerate(_combined_subtitle_entries(narration, work_dir, video_duration), start=1): srt_lines.append(str(idx)) srt_lines.append(f"{_seconds_to_srt_time(entry['start'])} --> {_seconds_to_srt_time(entry['end'])}") srt_lines.append(entry["text"] if entry.get("_bound_track") else _normalize_subtitle_text(entry["text"])) srt_lines.append("") srt_path = work_dir / "subtitles.srt" srt_path.write_text("\n".join(srt_lines), encoding="utf-8") return srt_path def _escape_ass_text(text): """Escape user text for an ASS dialogue Text field.""" return ( text .replace("\\", "\\\\") .replace("{", "\\{") .replace("}", "\\}") .replace("\r\n", "\n") .replace("\r", "\n") .replace("\n", "\\N") ) def _generate_ass(narration, work_dir, video_duration, canvas): """Generate an ASS subtitle file for readable hard-sub rendering, including the original dialogue during the original-audio gaps. canvas ({"width","height"}) scales the style to the real frame so portrait/竖屏 subtitles are not stretched.""" style = _style_for_measured_subtitle_band(_subtitle_style_config(canvas), canvas) ass_lines = [ "[Script Info]", "ScriptType: v4.00+", "WrapStyle: 0", "ScaledBorderAndShadow: yes", f"PlayResX: {int(style['play_res_x'])}", f"PlayResY: {int(style['play_res_y'])}", "", "[V4+ Styles]", ( "Format: Name, Fontname, Fontsize, PrimaryColour, SecondaryColour, " "OutlineColour, BackColour, Bold, Italic, Underline, StrikeOut, " "ScaleX, ScaleY, Spacing, Angle, BorderStyle, Outline, Shadow, " "Alignment, MarginL, MarginR, MarginV, Encoding" ), ( "Style: Default," f"{style['font_name']},{style['font_size']},{style['primary_color']},&H000000FF," f"{style['outline_color']},&H64000000,0,0,0,0,100,100,0,0,1," f"{style['outline']},{style['shadow']},{style['alignment']}," f"{style['margin_l']},{style['margin_r']},{style['margin_v']},1" ), "", "[Events]", "Format: Layer, Start, End, Style, Name, MarginL, MarginR, MarginV, Effect, Text", ] # entries are already split into short one-line chunks, so no wrapping here. for entry in _combined_subtitle_entries(narration, work_dir, video_duration): text = _escape_ass_text(entry["text"] if entry.get("_bound_track") else _normalize_subtitle_text(entry["text"])) ass_lines.append( "Dialogue: 0," f"{entry.get('ass_start', _seconds_to_ass_time(entry['start']))}," f"{entry.get('ass_end', _seconds_to_ass_time(entry['end']))}," f"Default,,0,0,0,,{text}" ) ass_path = work_dir / "subtitles.ass" ass_path.write_text("\n".join(ass_lines) + "\n", encoding="utf-8") return ass_path -
track.py 12.6 KB
"""Strict loader for versioned, output-clock subtitle tracks. This module intentionally uses only the Python standard library so it can be loaded by the independently distributed ``video-assemble`` skill. """ from __future__ import annotations import json import math from collections.abc import Mapping from decimal import Decimal from fractions import Fraction from pathlib import Path SCHEMA_VERSION = 1 # Digest keys older tracks declared; they are ignored, never a reason to fail. _LEGACY_BINDING_KEYS = frozenset({"sha256", "edit_sha256"}) _EVIDENCE_KINDS = frozenset( { "human_verified", "word_timestamps", "asr_boundary_calibrated", "legacy_estimate", } ) _CALIBRATION_KINDS = frozenset({"none", "asr_energy", "human_boundary"}) _WORD_ALIGNMENT_KINDS = frozenset({"none", "asr_words", "human_words"}) class SubtitleTrackError(ValueError): """The subtitle track is malformed, stale, or incompatible with its caller.""" def _fail(path, message): raise SubtitleTrackError(f"{path}: {message}") def _mapping(value, path): if not isinstance(value, Mapping): _fail(path, "must be an object") return value def _strict_fields(value, *, required, optional=(), path): value = _mapping(value, path) keys = set(value) missing = set(required) - keys unknown = keys - set(required) - set(optional) if missing: _fail(path, f"missing field(s): {', '.join(sorted(missing))}") if unknown: _fail(path, f"unknown field(s): {', '.join(sorted(unknown))}") return value def _integer(value, path, *, minimum=None): if isinstance(value, bool) or not isinstance(value, int): _fail(path, "must be an integer") if minimum is not None and value < minimum: _fail(path, f"must be >= {minimum}") return value def _nonempty_string(value, path): if not isinstance(value, str) or not value.strip(): _fail(path, "must be a nonempty string") return value def _local_path(value, path): return str(Path(_nonempty_string(value, path)).resolve()) def _seconds_fraction(value, path): if isinstance(value, bool): _fail(path, "duration must be finite and numeric") if isinstance(value, Fraction): result = value elif isinstance(value, Decimal): if not value.is_finite(): _fail(path, "duration must be finite") result = Fraction(value) elif isinstance(value, int): result = Fraction(value) elif isinstance(value, float): if not math.isfinite(value): _fail(path, "duration must be finite") result = Fraction(str(value)) else: _fail(path, "duration must be an int, float, Decimal, or Fraction") if result < 0: _fail(path, "duration must be nonnegative") return result def _load_document(path_or_mapping): if isinstance(path_or_mapping, Mapping): return path_or_mapping try: path = Path(path_or_mapping) except TypeError as exc: raise SubtitleTrackError("track must be a mapping or filesystem path") from exc try: value = json.loads(path.read_text(encoding="utf-8")) except (OSError, UnicodeError, json.JSONDecodeError) as exc: raise SubtitleTrackError(f"cannot read subtitle track {path}: {exc}") from exc return _mapping(value, "track") def _validate_picture_binding(binding, expected): binding = _strict_fields( binding, required={"path"}, optional={"edit_plan"} | _LEGACY_BINDING_KEYS, path="bindings.picture", ) expected = _mapping(expected, "expected_picture_identity") picture_path = _local_path(binding["path"], "bindings.picture.path") expected_path = _local_path(expected.get("path"), "expected_picture_identity.path") if picture_path != expected_path: _fail("bindings.picture.path", "picture path does not match current picture") edit_plan = binding.get("edit_plan") if edit_plan is not None: edit_plan = _local_path(edit_plan, "bindings.picture.edit_plan") if "edit_plan" not in expected: _fail("expected_picture_identity.edit_plan", "edit plan is required by track") expected_edit = _local_path(expected["edit_plan"], "expected_picture_identity.edit_plan") if edit_plan != expected_edit: _fail("bindings.picture.edit_plan", "edit plan does not match current edit") return {"path": picture_path, **({"edit_plan": edit_plan} if edit_plan else {})} _AUDIO_BINDING_FACTS = ("selected_stream", "sample_rate", "packet_count") def _validate_audio_binding(binding, expected): binding = _strict_fields( binding, required=set(_AUDIO_BINDING_FACTS), optional=_LEGACY_BINDING_KEYS, path="bindings.audio", ) expected = _mapping(expected, "expected_audio_identity") result = {} for key in _AUDIO_BINDING_FACTS: declared = _integer(binding[key], f"bindings.audio.{key}", minimum=0) actual = _integer(expected.get(key), f"expected_audio_identity.{key}", minimum=0) if declared != actual: _fail(f"bindings.audio.{key}", f"{key} does not match adopted audio") result[key] = declared return result def _validate_evidence(value, cue_path): path = f"{cue_path}.timing_evidence" value = _strict_fields( value, required={"kind", "evidence_refs", "calibration", "word_alignment"}, path=path, ) kind = _nonempty_string(value["kind"], f"{path}.kind") calibration = _nonempty_string(value["calibration"], f"{path}.calibration") word_alignment = _nonempty_string( value["word_alignment"], f"{path}.word_alignment" ) if kind not in _EVIDENCE_KINDS: _fail(path, f"timing_evidence kind is not supported: {kind!r}") if calibration not in _CALIBRATION_KINDS: _fail(path, f"timing_evidence calibration is not supported: {calibration!r}") if word_alignment not in _WORD_ALIGNMENT_KINDS: _fail(path, f"timing_evidence word_alignment is not supported: {word_alignment!r}") refs = value["evidence_refs"] if not isinstance(refs, list): _fail(path, "timing_evidence evidence_refs must be a list") refs = [_nonempty_string(ref, f"{path}.evidence_refs[{index}]") for index, ref in enumerate(refs)] valid_shape = { "legacy_estimate": calibration == "none" and word_alignment == "none", "asr_boundary_calibrated": calibration == "asr_energy" and word_alignment == "none", "word_timestamps": calibration in {"none", "asr_energy"} and word_alignment == "asr_words", "human_verified": calibration == "human_boundary" and word_alignment in {"none", "human_words"}, }[kind] if not valid_shape: _fail(path, f"timing_evidence fields are inconsistent with kind {kind!r}") if kind != "legacy_estimate" and not refs: _fail(path, f"timing_evidence kind {kind!r} requires evidence_refs") return { "kind": kind, "evidence_refs": refs, "calibration": calibration, "word_alignment": word_alignment, } def load_subtitle_track( path_or_mapping, *, expected_picture_identity, expected_audio_identity, expected_duration_seconds, reject_legacy_estimate=False, ): """Validate schema v1 and return metadata plus second-based render entries. Identity arguments are current facts (paths, stream index, sample rate, packet count) supplied independently by the caller; declarations inside the track are compared against them, never accepted on their own. Cue boundaries and text are validated, not split or corrected. """ if not isinstance(reject_legacy_estimate, bool): _fail("reject_legacy_estimate", "must be a boolean") track = _strict_fields( _load_document(path_or_mapping), required={"schema_version", "clock", "overlap_policy", "bindings", "cues"}, path="track", ) if type(track["schema_version"]) is not int or track["schema_version"] != SCHEMA_VERSION: _fail( "schema_version", f"only schema_version {SCHEMA_VERSION} is supported; got {track['schema_version']!r}", ) if track["overlap_policy"] != "forbid": _fail("overlap_policy", "schema v1 requires fail-closed value 'forbid'") clock = _strict_fields( track["clock"], required={"kind", "timebase", "duration_ticks"}, path="clock", ) if clock["kind"] != "output": _fail("clock.kind", "schema v1 supports output clock only") timebase = _strict_fields( clock["timebase"], required={"numerator", "denominator"}, path="clock.timebase" ) numerator = _integer(timebase["numerator"], "clock.timebase.numerator", minimum=1) denominator = _integer(timebase["denominator"], "clock.timebase.denominator", minimum=1) duration_ticks = _integer(clock["duration_ticks"], "clock.duration_ticks", minimum=0) tick_seconds = Fraction(numerator, denominator) duration = duration_ticks * tick_seconds expected_duration = _seconds_fraction(expected_duration_seconds, "expected_duration_seconds") if duration != expected_duration: _fail( "clock.duration_ticks", f"duration {duration} does not match independently supplied duration {expected_duration}", ) try: duration_seconds = float(duration) except OverflowError: _fail("clock.duration_ticks", "duration must be representable as finite renderer seconds") if not math.isfinite(duration_seconds): _fail("clock.duration_ticks", "duration must be representable as finite renderer seconds") bindings = _strict_fields( track["bindings"], required={"picture", "audio"}, path="bindings" ) picture = _validate_picture_binding(bindings["picture"], expected_picture_identity) audio = _validate_audio_binding(bindings["audio"], expected_audio_identity) cues = track["cues"] if not isinstance(cues, list): _fail("cues", "must be a list") entries = [] kinds = set() previous_end = 0 for index, raw_cue in enumerate(cues): cue_path = f"cues[{index}]" cue = _strict_fields( raw_cue, required={"start_tick", "end_tick", "text", "attribution", "timing_evidence"}, path=cue_path, ) start_tick = _integer(cue["start_tick"], f"{cue_path}.start_tick", minimum=0) end_tick = _integer(cue["end_tick"], f"{cue_path}.end_tick", minimum=0) if end_tick <= start_tick: _fail(cue_path, "half-open cue requires end_tick > start_tick") if end_tick > duration_ticks: _fail(cue_path, "cue end_tick exceeds output duration") if start_tick < previous_end: _fail(cue_path, "cue is out of order or overlaps the previous cue") previous_end = end_tick text = _nonempty_string(cue["text"], f"{cue_path}.text") attribution = _strict_fields( cue["attribution"], required={"kind", "ref"}, path=f"{cue_path}.attribution" ) source = _nonempty_string(attribution["kind"], f"{cue_path}.attribution.kind") if source not in {"source", "narration"}: _fail(f"{cue_path}.attribution.kind", "must be 'source' or 'narration'") source_ref = _nonempty_string(attribution["ref"], f"{cue_path}.attribution.ref") evidence = _validate_evidence(cue["timing_evidence"], cue_path) if reject_legacy_estimate and evidence["kind"] == "legacy_estimate": _fail(cue_path, "legacy_estimate is rejected by strict caller policy") start_seconds = float(start_tick * tick_seconds) end_seconds = float(end_tick * tick_seconds) if not (math.isfinite(start_seconds) and math.isfinite(end_seconds)): _fail(cue_path, "cue float projection must be finite") if end_seconds <= start_seconds: _fail(cue_path, "distinct cue ticks collapse in renderer float projection") kinds.add(evidence["kind"]) entries.append( { "start": start_seconds, "end": end_seconds, "text": text, "source": source, "source_ref": source_ref, "timing_evidence": evidence, } ) metadata = { "schema_version": SCHEMA_VERSION, "clock": { "kind": "output", "timebase": {"numerator": numerator, "denominator": denominator}, "duration_ticks": duration_ticks, "duration_seconds": duration_seconds, }, "overlap_policy": "forbid", "bindings": {"picture": picture, "audio": audio}, "timing_evidence_kinds": sorted(kinds), } return {"metadata": metadata, "entries": entries} -
track_binding.py 10.3 KB
"""Bind an explicit output-clock subtitle track to media before any consumer uses it. The first render integration deliberately supports adopted AAC only. A future newly mixed narration track needs its own final-mix identity, not this input's soundtrack. Bound cue timing is a declared decision, not proof of speech onset. """ import bisect import json from fractions import Fraction from pathlib import Path from artifacts import file_identity from adoption.frozen_audio import probe_audio_packets from lib import run_cmd from subtitles.track import load_subtitle_track TRACK = 'subtitle_track.json' VALIDATION = 'subtitle_track_validation.json' VALIDATION_SCHEMA = 1 PROJECTOR_VERSION = 1 def current_bindings(video, selected_audio_stream=0, *, edit_plan_path=None): """Compute the binding facts independently of the supplied subtitle track. The picture binding is the resolved media path (plus the edit plan path when one exists); the audio binding is the selected stream ordinal with its sample rate and packet count. """ packets = probe_audio_packets(video, selected_audio_stream) picture = {'path': str(Path(video).resolve())} if edit_plan_path is not None: picture['edit_plan'] = str(Path(edit_plan_path).resolve()) return {'picture': picture, 'audio': {'selected_stream': selected_audio_stream, 'sample_rate': packets['sample_rate'], 'packet_count': packets['packet_count']}} def _picture_clock(video): result = run_cmd([ 'ffprobe', '-v', 'error', '-select_streams', 'v:0', '-show_streams', '-show_frames', '-show_entries', 'stream=time_base,start_pts,duration_ts:frame=pts', '-of', 'json', str(video), ]) if result.returncode: raise ValueError(f'Cannot verify subtitle picture clock: {result.stderr}') data = json.loads(result.stdout) stream = data['streams'][0] timebase = Fraction(stream['time_base']) if int(stream['start_pts']) != 0: raise ValueError('Bound subtitle render requires an output-clock video starting at zero') duration = int(stream['duration_ts']) * timebase pts = [int(frame['pts']) * timebase for frame in data['frames']] if not pts or pts[0] != 0 or any(b <= a for a, b in zip(pts, pts[1:])) or pts[-1] >= duration: raise ValueError('Cannot verify picture frame clock: expected strictly increasing PTS within duration') return pts, duration def _project_boundary(target, frame_pts, duration): """Choose an ASS centisecond that changes on the first frame at/after a cue. ASS's 10ms clock is coarser than most frame clocks and can otherwise round a cue forward by one frame. Do not silently degrade if two relevant frames cannot be separated in that clock. libass receives presentation time in integer milliseconds from ffmpeg. """ index = bisect.bisect_left(frame_pts, target) chosen = frame_pts[index] if index < len(frame_pts) else duration centiseconds = chosen * 100 // 1 threshold_ms = int(centiseconds) * 10 previous_ms = int(frame_pts[index - 1] * 1000) if index else -1 if threshold_ms <= previous_ms: raise ValueError('ASS 10ms clock cannot represent this cue boundary; use a frame-capable renderer') ass_time = f'{centiseconds // 360000}:{centiseconds // 6000 % 60:02d}:{centiseconds // 100 % 60:02d}.{centiseconds % 100:02d}' return {'requested_time': str(target), 'frame_index': index, 'pts': str(chosen), 'delta': str(chosen - target), 'ass_time': ass_time} def prepare_subtitle_track(input_video, work_dir, video_duration, *, audio_mode, selected_audio_stream=0, edit_plan_path=None, reject_legacy_estimate=False): """Validate and project an explicit track before SRT/ASS/timeline/QC consumers. No track preserves legacy behavior. A present but invalid track fails instead of falling back to proportional timing. The persisted record is not approval. """ work = Path(work_dir) path = work / TRACK validation = work / VALIDATION validation.unlink(missing_ok=True) if not path.exists(): return None if audio_mode != 'adopted-packet-copy': raise ValueError('Explicit subtitle_track currently requires adopted-packet-copy; new mix is not bound') document = json.loads(path.read_bytes()) bindings = current_bindings(input_video, selected_audio_stream, edit_plan_path=edit_plan_path) frame_pts, duration = _picture_clock(input_video) if abs(float(duration) - float(video_duration)) > 0.05: raise ValueError('Picture duration differs from assembly duration') loaded = load_subtitle_track( document, expected_picture_identity=bindings['picture'], expected_audio_identity=bindings['audio'], expected_duration_seconds=duration, reject_legacy_estimate=reject_legacy_estimate, ) clock = document['clock']['timebase'] tick = Fraction(clock['numerator'], clock['denominator']) for entry, cue in zip(loaded['entries'], document['cues']): start = _project_boundary(cue['start_tick'] * tick, frame_pts, duration) end = _project_boundary(cue['end_tick'] * tick, frame_pts, duration) if end['frame_index'] <= start['frame_index']: raise ValueError('Subtitle cue has no visible frame in the actual picture clock') entry.update(_bound_track=True, ass_start=start['ass_time'], ass_end=end['ass_time'], start=float(Fraction(start['pts'])), end=float(Fraction(end['pts']))) entry['frame_projection'] = {'requested_ticks': [cue['start_tick'], cue['end_tick']], 'start': start, 'end': end} loaded['validation_schema'] = VALIDATION_SCHEMA loaded['projector_version'] = PROJECTOR_VERSION loaded['binding'] = { 'input_video': str(Path(input_video).resolve()), 'identities': bindings, 'inputs': {'track': file_identity(path), 'video': file_identity(input_video), 'edit_plan': file_identity(edit_plan_path) if edit_plan_path is not None else None}, 'duration': str(duration), 'frame_count': len(frame_pts), 'edit_plan': str(Path(edit_plan_path).resolve()) if edit_plan_path is not None else None, 'verification': 'media_binding_and_declared_frame_projection_only', 'direct_listening': 'NOT_CHECKED', 'acoustic_alignment': 'NOT_CHECKED', } validation.write_text(json.dumps(loaded, ensure_ascii=False, indent=2) + '\n', encoding='utf-8') return loaded def _load_validation(work): path = Path(work) / VALIDATION if not path.exists(): raise ValueError('Explicit subtitle track must be prepared against current media before use') loaded = json.loads(path.read_text(encoding='utf-8')) if (type(loaded.get('validation_schema')) is not int or loaded['validation_schema'] != VALIDATION_SCHEMA or type(loaded.get('projector_version')) is not int or loaded['projector_version'] != PROJECTOR_VERSION): raise ValueError('Subtitle validation/projector version changed; prepare again') return loaded def bound_subtitle_entries(work_dir, video_duration): """Return prepared exact entries, or None only when no explicit track exists.""" if work_dir is None: return None work = Path(work_dir) if not (work / TRACK).exists(): return None loaded = _load_validation(work) record = loaded['binding'] inputs = record['inputs'] if file_identity(work / TRACK) != inputs['track']: raise ValueError('stale subtitle track: author file changed after prepare') video = Path(record['input_video']) if not video.is_file() or file_identity(video) != inputs['video']: raise ValueError('stale subtitle track: adopted media changed after prepare') if record['edit_plan'] and (not Path(record['edit_plan']).is_file() or file_identity(record['edit_plan']) != inputs['edit_plan']): raise ValueError('stale subtitle track: edit plan changed after prepare') if abs(float(Fraction(record['duration'])) - float(video_duration)) > 0.05: raise ValueError('stale subtitle track: consumer duration changed after prepare') return loaded['entries'] def verify_rendered_picture(work_dir, output_path): """Check the actual rendered frame clock, not just pre-render declarations.""" work = Path(work_dir) if not (work / TRACK).exists(): return None path = work / VALIDATION loaded = json.loads(path.read_text(encoding='utf-8')) bound_subtitle_entries(work, float(Fraction(loaded['binding']['duration']))) loaded['rendered_picture'] = {'output': str(Path(output_path).resolve()), 'frame_clock_verified': False} path.write_text(json.dumps(loaded, ensure_ascii=False, indent=2) + '\n', encoding='utf-8') if _picture_clock(output_path) != _picture_clock(loaded['binding']['input_video']): raise ValueError('Rendered frame clock changed; subtitle frame projection is no longer valid') loaded['rendered_picture'].update(frame_clock_verified=True, frame_count=loaded['binding']['frame_count'], identity=file_identity(output_path)) path.write_text(json.dumps(loaded, ensure_ascii=False, indent=2) + '\n', encoding='utf-8') return loaded['rendered_picture'] def manifest_subtitle_evidence(work_dir, input_video, output_path): """Only expose evidence for the actual input and verified output of this render.""" work = Path(work_dir) if not (work / TRACK).exists(): return None loaded = _load_validation(work) bound_subtitle_entries(work, float(Fraction(loaded['binding']['duration']))) rendered = loaded.get('rendered_picture', {}) if rendered.get('frame_clock_verified') is not True: raise ValueError('Subtitle output frame clock has not been verified') if str(Path(input_video).resolve()) != loaded['binding']['input_video']: raise ValueError('stale subtitle manifest: input media differs') if not Path(output_path).is_file() or file_identity(output_path) != rendered.get('identity'): raise ValueError('stale subtitle manifest: actual output differs from verified render') return {'validation_path': str((work / VALIDATION).resolve()), 'binding': loaded['binding'], 'metadata': loaded['metadata'], 'rendered_picture': rendered} -
__init__.py 56 B
"""Subtitle generation and track-binding subpackage."""
-
-
artifacts.py 1.4 KB
"""File identity and work-directory JSON helpers for video-assemble.""" import json import os from pathlib import Path from lib import CONFIG def file_identity(path): """``{size, mtime_ns}`` of one input file: enough to notice a rewrite, cheap to take.""" st = os.stat(os.fspath(path)) return {"size": st.st_size, "mtime_ns": st.st_mtime_ns} def _artifact_identity(path): path = Path(path) return file_identity(path) if path.exists() else None def _explicit_source_video(): """Return the cut-mode source video only when the caller opted in explicitly.""" if not CONFIG["source_video_explicit"]: return "" return CONFIG["source_video"] def _source_video_identity(): """``{path, size, mtime_ns}`` of the explicit cut-mode source video, else None.""" source_video = _explicit_source_video() if not source_video: return None path = Path(source_video) return {"path": str(path.resolve()), **file_identity(path)} def _timeline_provenance_status(work_dir): data = _load_work_json(work_dir, "timeline.json") return None if data is None else data["provenance"] def _load_work_json(work_dir, name): """Parse a JSON artifact in work_dir; None only when the file does not exist.""" path = Path(work_dir) / name if not path.exists(): return None return json.loads(path.read_text(encoding="utf-8")) -
assemble.py 29 KB
"""Canonical CLI and programmatic render entry for the self-contained video-assemble skill.""" import json import os from pathlib import Path import artifacts import assemble_constants as constants import assembly_contract import assembly_settings import audio_mix import adoption.audio_mix_binding as audio_mix_binding import adoption.frozen_audio as frozen_audio import media import narration_audio import adoption.narration_binding as narration_binding import pair_media import render_preflight import adoption.strict_publish as strict_publish import subtitles.render as subtitle_render import subtitles.track_binding as subtitle_track_binding import timeline_emit import packaging import visual_render import lib __all__ = [ "assemble_video", "main", ] AUDIO_MODES = ("narration", "source-mix", "adopted-packet-copy") _current_narration_binding = strict_publish.current_narration_binding _current_audio_mix_binding = strict_publish.current_audio_mix_binding def assemble_video(input_video, tts_segments, work_dir, output_path, *, audio_mode="narration", audio_stream_index=0, narration_adoption_path=None, tts_meta_path=None, audio_mix_adoption_path=None): """组装最终视频""" if audio_mode not in AUDIO_MODES: raise RuntimeError(f"不支持的 audio_mode: {audio_mode}") if isinstance(audio_stream_index, bool) or not isinstance(audio_stream_index, int) or audio_stream_index < 0: raise RuntimeError("audio_stream_index 必须是非负整数") if audio_mode == "narration" and not tts_segments: raise RuntimeError("tts_meta.json 没有有效解说音频,已中止以避免生成无解说视频") if audio_mode == "narration" and audio_stream_index != 0: raise RuntimeError("narration 当前不支持非零 audio_stream_index") if audio_mode == "source-mix" and tts_segments: raise RuntimeError("source-mix 与 TTS 解说不兼容") if audio_mode != "narration" and tts_meta_path is not None: raise RuntimeError(f"tts_meta 与 audio_mode {audio_mode} 不兼容") bgm_path = lib.CONFIG["bgm_path"] has_bgm = bool(bgm_path) and os.path.exists(bgm_path) if audio_mode == "source-mix" and bgm_path and not has_bgm: raise RuntimeError(f"source-mix 声明的 BGM 文件不存在: {bgm_path}") if ( audio_mode != "narration" and audio_stream_index != 0 and lib.CONFIG["export_jianying"] ): raise RuntimeError("剪映导出当前不支持选择非零音频流") if audio_mode == "adopted-packet-copy": if tts_segments: raise RuntimeError("adopted-packet-copy 与 TTS 解说不兼容") if bgm_path: raise RuntimeError("adopted-packet-copy 与 BGM 混音不兼容") if audio_mode != "narration" and narration_adoption_path is not None: raise RuntimeError("narration adoption 仅适用于 narration audio_mode") if audio_mix_adoption_path is not None and ( audio_mode != "narration" or narration_adoption_path is None or tts_meta_path is None ): raise RuntimeError("audio mix adoption 要求 narration 模式及显式 narration adoption/tts_meta") published_output = Path(output_path) if audio_mix_adoption_path is not None and published_output.exists(): raise RuntimeError("显式音频混合要求新的 output_path,不能覆盖已有成片") explicit_mix = ( audio_mix_binding.load_adoption( audio_mix_adoption_path, input_video=input_video, narration_adoption_path=narration_adoption_path, tts_segments=tts_segments, ) if audio_mix_adoption_path is not None else None ) binding = narration_binding.prepare_binding( tts_segments, work_dir, narration_adoption_path=narration_adoption_path, tts_meta_path=tts_meta_path, ) if audio_mode == "narration" else None render_output = published_output if binding and binding["active"]: if published_output.exists(): raise RuntimeError("身份约束渲染要求新的 output_path,不能覆盖已有成片") render_output = published_output.with_name( f".{published_output.stem}.narration-rendering{published_output.suffix}" ) render_output.unlink(missing_ok=True) output_path = render_output video_duration = lib.get_video_duration(input_video) canvas = media._probe_canvas(input_video) # drives subtitle PlayRes/scale so 竖屏 text isn't stretched burn_subtitles = lib.CONFIG["burn_subtitles"] subtitle_track_binding.prepare_subtitle_track( input_video, work_dir, video_duration, audio_mode=audio_mode, selected_audio_stream=audio_stream_index, ) adopted_source_audio = None if audio_mode == "adopted-packet-copy": adopted_source_audio = frozen_audio.validate_adopted_source( input_video, audio_stream_index ) if type(adopted_source_audio["sample_rate"]) is not int \ or adopted_source_audio["sample_rate"] <= 0: raise RuntimeError("adopted audio sample rate must be a positive integer") elif audio_mode == "source-mix": frozen_audio.probe_audio_packets(input_video, audio_stream_index) # 解说整体提速(可选)后,将所有 TTS 片段按时间位置合成到与视频等长的音轨上 narration_wav = None if audio_mode == "narration" and explicit_mix is not None: explicit_runtime = audio_mix_binding.render_explicit_mix( explicit_mix, binding, tts_segments, work_dir ) narration_wav = Path(explicit_runtime["voice_bus"]["path"]) elif audio_mode == "narration": if binding["tempo_policy"]: narration_audio._apply_narration_speed( tts_segments, work_dir, tempo_policy=binding["tempo_policy"] ) else: narration_audio._apply_narration_speed(tts_segments, work_dir) narration_wav = work_dir / "narration.wav" if binding["tempo_policy"]: narration_audio._build_timed_narration( tts_segments, narration_wav, video_duration, work_dir, tempo_policy=binding["tempo_policy"], ) if any( segment.get("blocking") or segment.get("fit_status") == "no_safe_fit" for segment in tts_segments ): raise RuntimeError( "严格 narration adoption 存在 no_safe_fit,禁止提速或裁尾渲染" ) else: narration_audio._build_timed_narration( tts_segments, narration_wav, video_duration, work_dir ) handoffs = audio_mix._apply_source_sentence_handoffs(tts_segments, work_dir, video_duration) if handoffs: lib.log( "原声句末交接: " + ", ".join( f"{item['end']:.2f}s→{item.get('restore_at', item['end']):.2f}s({item['status']})" for item in handoffs ) ) narration_binding.seal_render_inputs(binding, tts_segments, narration_wav) # 始终生成 SRT 字幕文件(原声留白处补烧原声字幕,传入成片时长以计算留白区间) srt_path = subtitle_render._generate_srt(tts_segments, work_dir, video_duration) lib.log(f"字幕文件: {srt_path}") ass_path = None if burn_subtitles: ass_path = subtitle_render._generate_ass(tts_segments, work_dir, video_duration, canvas) lib.log(f"压制字幕文件: {ass_path}") # 可选 BGM:作为一条独立音轨(input [2:a])混入,旁白处自动压低 if explicit_mix is not None: lib.log("显式 adopted full-sound:忽略环境 BGM/duck/loudnorm/tempo 配置") elif bgm_path and not has_bgm: lib.log(f" ⚠️ BGM 文件不存在,跳过: {bgm_path}") elif has_bgm: lib.log(f"BGM 铺底: {bgm_path} (音量 {lib.CONFIG['bgm_volume']},旁白时 {lib.CONFIG['bgm_ducking_volume']})") # 多轨时间线模型(timeline.json):canonical 渲染仍是 ffmpeg,此模型供检视/可选导出 timeline_emit._emit_timeline( input_video, tts_segments, work_dir, video_duration, canvas, has_bgm, audio_mode=audio_mode, selected_audio_stream=audio_stream_index, explicit_audio_mix=( {**explicit_mix, **explicit_mix["runtime"]} if explicit_mix is not None else None ), ) overlay_filters, overlay_qc = visual_render._visual_overlay_filters(work_dir, canvas, video_duration) packaging_layers = packaging.load_packaging_layers(work_dir, canvas) mask_filter = visual_render._source_subtitle_mask_filter(canvas, work_dir, tts_segments, video_duration) visual_qc = visual_render._build_visual_qc( tts_segments, work_dir, video_duration, canvas, overlay_qc=overlay_qc, mask_filter=mask_filter, ) visual_render._write_visual_qc(work_dir, visual_qc) assembly_qc_path = Path(work_dir) / constants.ASSEMBLY_QC assembly_qc_path.unlink(missing_ok=True) if visual_qc["blocking"]: codes = ", ".join(visual_qc["blocking_codes"]) raise RuntimeError(f"视觉 QC 失败: {codes};详见 {Path(work_dir) / constants.VISUAL_QC}") # Select exactly one of three explicit audio paths. Only narration may synthesize # a missing original track; adopted copy never decodes, mixes, normalizes or trims. source_has_audio = media._has_audio_stream(input_video) adopted_audio = None loudnorm_measurement = None original_audio_input = [] bgm_input = [] filter_complex = None filter_args = [] audio_input_args = [] fc_script = None if audio_mode == "adopted-packet-copy": audio_map = f"0:a:{audio_stream_index}" elif audio_mode == "source-mix": audio_map = "[aoutln]" source_label = f"0:a:{audio_stream_index}" filter_complex = f"[{source_label}]volume={lib.CONFIG['idle_orig_volume']}[source]" if has_bgm: bgm_input = ["-stream_loop", "-1", "-i", str(bgm_path)] filter_complex += ( f";[1:a]volume={lib.CONFIG['bgm_volume']}[bgm]" ";[source][bgm]amix=inputs=2:duration=first:dropout_transition=0[aout]" ) else: filter_complex += ";[source]anull[aout]" final_ln = audio_mix.final_loudnorm_filter() filter_complex += f";[aout]{final_ln}[aoutln]" lib.log(f"source-mix 音频处理: source volume + {final_ln}") elif explicit_mix is not None: audio_map = "1:a:0" audio_input_args = ["-i", explicit_mix["runtime"]["master"]["path"]] else: # 混合原始音频 + 解说音频(+ 可选 BGM) if source_has_audio: original_audio_label = "0:a" bgm_audio_label = "2:a" else: lib.log("源视频无音轨,使用静音原声音轨进行混音") original_audio_input = [ "-f", "lavfi", "-t", str(video_duration), "-i", "anullsrc=channel_layout=stereo:sample_rate=48000", ] original_audio_label = "2:a" bgm_audio_label = "3:a" filter_complex = audio_mix._build_audio_filter_complex( tts_segments, has_bgm, original_audio_label=original_audio_label, bgm_audio_label=bgm_audio_label, ) # BGM is input [2:a]; -stream_loop -1 loops it to cover the whole timeline (amix # duration=first + -t trim it back to the video length). bgm_input = ["-stream_loop", "-1", "-i", str(bgm_path)] if has_bgm else [] # 末端整体响度归一:ducking 只管相对平衡,这一步统一成片绝对响度 loudnorm_measurement = audio_mix._run_loudnorm_first_pass( input_video, narration_wav, original_audio_input, bgm_input, filter_complex, work_dir, ) final_ln = audio_mix.final_loudnorm_filter(loudnorm_measurement) filter_complex += f";[aout]{final_ln}[aoutln]" lib.log(f"成片响度归一: {final_ln}") audio_map = "[aoutln]" audio_input_args = ["-i", str(narration_wav), *original_audio_input] # 对于超长 volume 表达式(多段解说),从脚本文件读取 filter_complex 避免命令行溢出 if filter_complex is not None: if len(filter_complex.encode("utf-8")) > constants.FILTER_SCRIPT_THRESHOLD_BYTES: fc_script = Path(work_dir) / ".filter_complex.txt" fc_script.write_text(filter_complex, encoding="utf-8") lib.log(f"使用 filter_complex 脚本文件 (表达式长度 {len(filter_complex.encode('utf-8'))} bytes)") filter_args = lib.filter_file_args("filter_complex", fc_script) else: filter_args = ["-filter_complex", filter_complex] cmd = [ "ffmpeg", "-y", "-i", str(input_video), *audio_input_args, *bgm_input, *filter_args, # 0:v:0 (not 0:v): sources with attached cover art carry a second video stream. # -vf only ever applies to the first one, so mapping all of them makes ffmpeg # abort with "Could not write header (incorrect codec parameters ?)" and leave # an unreadable file — after the whole pipeline has already run. "-map", "0:v:0", "-map", audio_map, ] # Video filter chain: mask source subtitles first (drawbox), then burn our subtitles # on top. Either one forces a re-encode; with neither, the video stream is copied. crf = str(lib.CONFIG["output_crf"]) # env_int already clamps to >=0; keep 0 (lossless) intact preset = lib.CONFIG["output_preset"] max_h = lib.CONFIG["output_max_height"] vf_chain = [] if mask_filter: vf_chain.append(mask_filter) vf_chain.extend(overlay_filters) if burn_subtitles: vf_chain.append(visual_render._subtitle_burn_filter(ass_path)) # Downscale LAST so the mask + burned subtitles render at native resolution and are then # scaled down with the frame (crisp). The helper forces both dimensions even so an odd # OUTPUT_MAX_HEIGHT can't crash libx264; 'min(ih,H)' only ever shrinks the source. if max_h > 0: vf_chain.append(visual_render._output_downscale_filter(max_h)) # yuv420p: 10-bit/4:2:2 sources re-encoded as-is play on desktop but fail on WeChat/ # mobile/Safari; force 8-bit 4:2:0 so every recap is universally decodable. yuv420p also # needs EVEN width AND height, so normalize odd dims (4:2:2/4:4:4 permit them) before the # encode — otherwise libx264 aborts to a 0-byte file. The downscale helper already evens out. even = "scale=trunc(iw/2)*2:trunc(ih/2)*2" reencode = bool(vf_chain or packaging_layers) or lib.CONFIG["force_video_reencode"] notes = [] video_filter_script = None if vf_chain or packaging_layers: if max_h <= 0: # no downscale in the chain to force even dims vf_chain.append(even) video_filter = packaging.compose_video_filter( vf_chain, packaging_layers, mask_first=bool(mask_filter) ) if len(video_filter.encode("utf-8")) > constants.FILTER_SCRIPT_THRESHOLD_BYTES: video_filter_script = Path(work_dir) / ".video_filter.txt" video_filter_script.write_text(video_filter, encoding="utf-8") cmd += lib.filter_file_args("filter:v:0", video_filter_script) lib.log( "使用 video filter script " f"(表达式长度 {len(video_filter.encode('utf-8'))} bytes)" ) else: cmd += ["-vf", video_filter] cmd += ["-c:v", "libx264", "-preset", preset, "-crf", crf, "-pix_fmt", "yuv420p"] notes = ((["遮挡原字幕"] if mask_filter else []) + ([f"包装图层×{len(packaging_layers)}"] if packaging_layers else []) + ([f"视觉叠加×{len(overlay_filters)}"] if overlay_filters else []) + (["压制解说字幕"] if burn_subtitles else []) + ([f"缩放≤{max_h}p"] if max_h > 0 else [])) lib.log(f"视频重编码: {' + '.join(notes)} (crf={crf}, preset={preset})") elif reencode: notes = ["force_video_reencode"] cmd += ["-vf", even, "-c:v", "libx264", "-preset", preset, "-crf", crf, "-pix_fmt", "yuv420p"] else: cmd += ["-c:v", "copy"] # +faststart relocates the moov atom to the front so web/social players can start # before the full file downloads; valid (and beneficial) on the copy path too. if audio_mode == "adopted-packet-copy": # No -t/-shortest: either would discard valid AAC priming or tail packets. cmd += ["-c:a", "copy", "-movie_timescale", str(adopted_source_audio["sample_rate"]), "-movflags", "+faststart", str(output_path)] elif explicit_mix is not None: cmd += ["-c:a", "aac", "-b:a", "192k", "-ar", "48000", "-movie_timescale", "48000", "-movflags", "+faststart", str(output_path)] else: cmd += ["-c:a", "aac", "-b:a", "192k", "-ar", "48000", "-movflags", "+faststart", "-t", str(video_duration), str(output_path)] try: result = lib.run_cmd(cmd) if result.returncode != 0: raise RuntimeError(f"视频组装失败: {result.stderr}") finally: # 清理临时 filter 脚本(无论 ffmpeg 是否成功) if fc_script is not None: fc_script.unlink(missing_ok=True) if video_filter_script is not None: video_filter_script.unlink(missing_ok=True) if audio_mode == "adopted-packet-copy": # Either check failing means the file at the final path is unverified: never leave it. try: adopted_audio = frozen_audio.verify_adopted_audio( input_video, output_path, audio_stream_index ) pair_media.validate_aac_packet_interval(adopted_audio["output"]) except (RuntimeError, ValueError): output_path.unlink(missing_ok=True) raise subtitle_track_binding.verify_rendered_picture(work_dir, output_path) audio_operations = { "narration": audio_mode == "narration", "source_mix": audio_mode == "source-mix", "bgm_mix": has_bgm and audio_mode != "adopted-packet-copy" and explicit_mix is None, "ducking": audio_mode == "narration" and explicit_mix is None, "loudness_normalization": ( audio_mode != "adopted-packet-copy" and explicit_mix is None and lib.CONFIG["final_loudnorm"] ), "limiter": audio_mode != "adopted-packet-copy" and explicit_mix is None, "resample": audio_mode != "adopted-packet-copy", "tempo": audio_mode == "narration" and explicit_mix is None, "packet_copy": audio_mode == "adopted-packet-copy", } if explicit_mix is not None: audio_operations["explicit_audio_mix"] = True render_delivery = { "video_encode_passes": 1 if reencode else 0, "reencode_reason": notes, "audio_sample_rate": ( adopted_audio["output"]["sample_rate"] if adopted_audio else 48000 ), "final_compat_notes": ( (["yuv420p"] if reencode else ["video_copy"]) + (["aac_packet_copy", "faststart"] if adopted_audio else ["aac_48000", "faststart"]) ), } loudness_mode = ( "not_run" if audio_mode == "adopted-packet-copy" else "fixed_master_gain_no_loudnorm" if explicit_mix is not None else None ) source_audio_status = "prepared_bed_adopted" if explicit_mix is not None else None render_output = strict_publish.publish_render( work_dir=work_dir, binding=binding, explicit_mix=explicit_mix, tts_segments=tts_segments, narration_wav=narration_wav, render_output=render_output, published_output=published_output, audio_mode=audio_mode, audio_operations=audio_operations, adopted_audio=adopted_audio, loudness_mode=loudness_mode, loudnorm_measurement=loudnorm_measurement, visual_qc=visual_qc, source_has_audio=source_has_audio, video_duration=video_duration, render_delivery=render_delivery, source_audio_status=source_audio_status, ) lib.log(f"最终视频: {render_output} ({render_output.stat().st_size / 1024 / 1024:.1f}MB)") return render_output def main(): import argparse import shutil ap = argparse.ArgumentParser( description="video-assemble: mux narration audio over the video, duck the original, render subtitles.") ap.add_argument("video", help="source video (edited_source.mp4 in cut mode, else the original)") ap.add_argument("--work-dir", required=True) ap.add_argument("--tts-meta", default=None, help="tts_meta.json (default: <work-dir>/tts_meta.json)") ap.add_argument("--narration-adoption", default=None, help="strict narration_adoption v1 bound to an explicit --tts-meta") ap.add_argument("--audio-mix-adoption", default=None, help="strict audio_mix_adoption v1 for adopted prepared bed and narration") ap.add_argument( "--audio-mode", choices=AUDIO_MODES, default="narration", help="audio path (default: narration)", ) ap.add_argument( "--audio-stream-index", type=int, default=0, help="zero-based input audio stream ordinal for source/adopted modes", ) ap.add_argument("--recap-stem", default=None, help="final recap filename stem (default: video stem)") ap.add_argument("--output-dir", default=None) ap.add_argument("--burn-subtitles", action=argparse.BooleanOptionalAction, default=None, help="burn narration subtitles into the video (default on; --no-burn-subtitles to disable)") ap.add_argument("--subtitle-y-top", type=int, default=None, help="inclusive top of a measured subtitle band in display-frame pixels") ap.add_argument("--subtitle-y-bot", type=int, default=None, help="exclusive bottom of a measured subtitle band in display-frame pixels") ap.add_argument("--source-video", default=None, help="original source video (cut mode) so timeline.json / 剪映 export reference the real clips") ap.add_argument("--export-jianying", action="store_true", help="also export an OPTIONAL 剪映/JianYing draft from timeline.json after rendering") ap.add_argument("--jianying-out", default=None, help="parent dir for the 剪映 draft (default: work-dir)") bundle_group = ap.add_mutually_exclusive_group() bundle_group.add_argument("--jianying-bundle-media", dest="jianying_bundle_media", action="store_true", help="copy media into the 剪映 draft folder (default on; portable/self-contained)") bundle_group.add_argument("--jianying-no-bundle-media", dest="jianying_bundle_media", action="store_false", help="do NOT copy media into the draft — reference in place (only if 剪映 can read those paths; macOS 剪映 usually cannot)") ap.set_defaults(jianying_bundle_media=None) args = ap.parse_args() work_dir = Path(args.work_dir) if args.burn_subtitles is not None: lib.CONFIG["burn_subtitles"] = args.burn_subtitles if (args.subtitle_y_top is None) != (args.subtitle_y_bot is None): ap.error("--subtitle-y-top and --subtitle-y-bot must be provided together") if args.subtitle_y_top is not None: if args.subtitle_y_top < 0 or args.subtitle_y_bot <= args.subtitle_y_top: ap.error("subtitle Y coordinates must satisfy 0 <= top < bot") lib.CONFIG["subtitle_y_top"] = args.subtitle_y_top lib.CONFIG["subtitle_y_bot"] = args.subtitle_y_bot lib.CONFIG["mask_source_subtitles"] = True lib.CONFIG["source_subtitle_mask_policy"] = "opt_in" lib.CONFIG["source_subtitle_mask_policy_declared"] = True # A measured band is an explicit request to conceal the known source-caption # pixels. The general 0.6 translucent look can leave white glyphs visible under # the generated subtitles; use an opaque mask unless the caller deliberately # chose a different opacity through the existing environment override. if "SUBTITLE_MASK_OPACITY" not in os.environ: lib.CONFIG["subtitle_mask_opacity"] = 1.0 if args.source_video: if not os.path.exists(args.source_video): ap.error(f"--source-video does not exist: {args.source_video}") lib.CONFIG["source_video"] = args.source_video lib.CONFIG["source_video_explicit"] = True else: # SOURCE_VIDEO is an ambient env var in lib.CONFIG. Do not let a stale # shell value silently bind full-mode/direct timeline.json or JianYing # exports to an unrelated original; cut mode must pass --source-video. lib.CONFIG["source_video"] = "" lib.CONFIG["source_video_explicit"] = False if args.export_jianying: lib.CONFIG["export_jianying"] = True if args.jianying_bundle_media is not None: lib.CONFIG["jianying_bundle_media"] = args.jianying_bundle_media render_preflight._preflight_burn_subtitles() # fail before the render if burn-in is on but ffmpeg lacks libass # Argument combinations are validated once, by assemble_video. tts_meta = Path(args.tts_meta) if args.tts_meta else None tts_segments = [] if args.audio_mode == "narration": tts_meta = tts_meta or work_dir / "tts_meta.json" tts_segments = json.loads(tts_meta.read_text(encoding="utf-8"))["segments"] stem = args.recap_stem or Path(args.video).stem base = Path(args.output_dir) if args.output_dir else work_dir.parent final_output = assembly_contract._resolve_final_output(base, stem) if args.audio_mix_adoption is not None and final_output.exists(): ap.error("explicit audio mix requires a new final delivery path") delivery_stage = None owned_alias = None output_path = work_dir / "output.mp4" try: assemble_video( args.video, tts_segments, work_dir, output_path, audio_mode=args.audio_mode, audio_stream_index=args.audio_stream_index, narration_adoption_path=args.narration_adoption, tts_meta_path=tts_meta, audio_mix_adoption_path=args.audio_mix_adoption, ) assembly_qc = artifacts._load_work_json(work_dir, constants.ASSEMBLY_QC) if assembly_qc["blocking"]: codes = ", ".join(assembly_qc["blocking_codes"]) raise SystemExit( f"组装 QC 阻断交付: {codes};详见 {work_dir / constants.ASSEMBLY_QC}" ) base.mkdir(parents=True, exist_ok=True) if args.audio_mix_adoption is not None: # Publish the delivery alias only when this process created it: stage a copy, # hard-link it into place (fails if the alias appeared meanwhile), remember the # inode, and roll back only an alias we own. delivery_stage = final_output.with_name( f".{final_output.name}.rendering-{os.getpid()}" ) with output_path.open("rb") as source, delivery_stage.open("xb") as target: shutil.copyfileobj(source, target) target.flush() os.fsync(target.fileno()) try: os.link(delivery_stage, final_output) except FileExistsError as exc: raise RuntimeError("explicit audio mix delivery alias appeared during render") from exc stat = final_output.stat() owned_alias = (stat.st_dev, stat.st_ino) delivery_stage.unlink() delivery_stage = None else: shutil.copy2(str(output_path), str(final_output)) manifest = assembly_contract._assembly_manifest_payload( args.video, tts_segments, work_dir, output_path, tts_meta_path=tts_meta, narration_input_binding=_current_narration_binding(work_dir, args.audio_mode), audio_mix_binding=_current_audio_mix_binding(work_dir, args.audio_mode), final_output=final_output, settings_payload=assembly_settings.assembly_settings_payload, audio_mode=args.audio_mode, audio_stream_index=args.audio_stream_index, ) assembly_contract._write_assembly_manifest(work_dir, manifest) except BaseException: if delivery_stage is not None: delivery_stage.unlink(missing_ok=True) if owned_alias is not None and final_output.exists(): stat = final_output.stat() if (stat.st_dev, stat.st_ino) == owned_alias: final_output.unlink() raise lib.log(f"组装完成: {final_output}") # OPTIONAL, decoupled: export a 剪映 draft from the timeline (lazy import; never # required by the core render path). if lib.CONFIG["export_jianying"]: from jianying.optional import maybe_export_jianying maybe_export_jianying(work_dir, args.jianying_out, stem) print(json.dumps({"status": "assembled", "output": str(final_output), "work_dir": str(work_dir)}, ensure_ascii=False)) if __name__ == "__main__": main() -
assemble_constants.py 2 KB
"""Shared constants for the self-contained video-assemble skill.""" from fractions import Fraction import math ASSEMBLY_MANIFEST = "assembly_manifest.json" ASSEMBLY_QC = "assembly_qc.json" VISUAL_QC = "visual_qc.json" VISUAL_OVERLAYS = "visual_overlays.json" SEGMENT_AUDIO_SCHEMA_VERSION = 1 # The picture codecs every packet/frame clock proof in this skill accepts. SUPPORTED_PICTURE_CODECS = frozenset({"h264", "hevc"}) # The one exact output audio clock every explicit-sound artifact in this skill uses. OUTPUT_SAMPLE_RATE = 48_000 def frame_clock_samples(frame, fps, rate=OUTPUT_SAMPLE_RATE): """Project an exact frame boundary of a video clock onto the output sample clock. Broadcast rates such as 30000/1001 put frame boundaries between whole samples, so the exact Fraction position is rounded half-up once, here. Every sample bound in the skill - a segment edge and the whole-picture `total_samples` alike - is this single projection, so an integral clock is unchanged and a fractional one stays consistent across the explicit-sound tools. """ exact = Fraction(int(frame) * rate, 1) / Fraction(fps) return math.floor(exact + Fraction(1, 2)) FILTER_SCRIPT_THRESHOLD_BYTES = 8000 # The default subtitle metrics were tuned in this reference canvas. SUBTITLE_STYLE_REF_W = 1280 SUBTITLE_STYLE_REF_H = 720 _SUBTITLE_TERMINAL_PUNCTUATION = "。!?!?…." _SUBTITLE_CLOSING_QUOTES = "」』”’))]】》〉\"'" _MIN_GAP_TO_SUBTITLE = 0.8 _MIN_READABLE_SECONDS = 0.3 _MIN_ASR_CLIP_OVERLAP = 0.05 # timeline.py serializes interval bounds onto a 1e-4 second grid, flooring starts # and ceiling ends, so two bounds that were identical before serialization can come # back one grid step apart. Contiguity joins must tolerate that whole step. _TIMELINE_TIME_GRID_SECONDS = 1e-4 _CLIP_CONTIGUITY_TOLERANCE = 1.5 * _TIMELINE_TIME_GRID_SECONDS _MAX_ORIGINAL_READ_CPS = 9.0 _AUTO_ORIGINAL_READ_CPS = 6.0 _SUPPORTED_VISUAL_OVERLAY_TYPES = {"top_title", "inline_label_or_callout"} -
assembly_contract.py 13.6 KB
"""Assembly manifest/QC persistence and delivery contract helpers.""" import json from fractions import Fraction import math import subprocess import wave from pathlib import Path from lib import CONFIG from assemble_constants import ( ASSEMBLY_MANIFEST, ASSEMBLY_QC, SEGMENT_AUDIO_SCHEMA_VERSION, ) from audio_mix import _loudness_mode from subtitles.track_binding import manifest_subtitle_evidence from artifacts import ( _load_work_json, _source_video_identity, _timeline_provenance_status, ) _AUDIO_QC_CODES = frozenset({ "missing_narration", "skipped_segments", "no_safe_fit", "effective_tempo_exceeded", "empty_narration", "truncated_speech", "unsafe_source_handoff", "timeline_audio_mismatch", }) def _assembly_manifest_payload(input_video, tts_segments, work_dir, output_path, tts_meta_path=None, final_output=None, *, settings_payload, audio_mode="narration", audio_stream_index=0, narration_input_binding=None, audio_mix_binding=None): """Slim render record. The orchestrator reads `final_output` to report the result; `source_video` stays None unless cut mode explicitly passed --source-video, proving a stale ambient SOURCE_VIDEO never leaked into a full-mode timeline / 剪映 export.""" input_video = Path(input_video) output_path = Path(output_path) source_video_identity = _source_video_identity() qc_path = Path(work_dir) / ASSEMBLY_QC qc = _load_work_json(work_dir, ASSEMBLY_QC) # always written by publish_render first settings = settings_payload( work_dir, audio_mode=audio_mode, audio_stream_index=audio_stream_index ) payload = { "schema_version": 2, "input_video": str(input_video.resolve()), "source_video": source_video_identity["path"] if source_video_identity else None, "source_video_identity": source_video_identity, "tts_meta": str(Path(tts_meta_path).resolve()) if tts_meta_path else None, "tts_segments": len(tts_segments), "audio_mode": audio_mode, "selected_audio_stream_index": audio_stream_index, "assembly_settings": settings, "output_path": str(output_path.resolve()), "segment_audio_schema_version": SEGMENT_AUDIO_SCHEMA_VERSION, "qc_path": str(qc_path.resolve()), "qc_verdict": qc["verdict"], "qc_blocking_codes": qc["blocking_codes"], # The settings payload records the configured/fallback loudness policy; these QC # fields record what the just-finished render actually used after the loudnorm probe. "qc_loudness_mode": qc["loudness_mode"], "qc_loudnorm_measurement": qc["loudnorm_measurement"], "audio_operations": qc["audio_operations"], "adopted_audio": qc["adopted_audio"], "narration_input_binding": narration_input_binding, "audio_mix_binding": audio_mix_binding, "audio_segments": [ { "index": seg["index"], # Informational manifest fields: a tts_meta without the schema marker is v1, # and the loudness measurements are None whenever voiceover skipped # normalization (strict adoption fixtures in tests/orchestrator omit both). "segment_audio_schema_version": seg.get( "segment_audio_schema_version", SEGMENT_AUDIO_SCHEMA_VERSION ), "narration": seg["narration"], "spoken_text": seg["spoken_text"], "truncated": seg["truncated"], "truncate_reason": seg["truncate_reason"], "fit_status": seg["fit_status"], "blocking": seg["blocking"], "audio_duration": seg["audio_duration"], "placed_audio_duration": seg["placed_audio_duration"], "placed_audio_path": seg.get("placed_audio_path"), "actual_place_start": seg.get("actual_place_start"), "actual_place_end": seg.get("actual_place_end"), "source_duck_end": seg.get("source_duck_end"), "source_restore_at": seg.get("source_restore_at"), "source_handoff_status": seg.get("source_handoff_status"), "source_entry_status": seg.get("source_entry_status"), "global_narration_speed": seg["global_narration_speed"], "segment_tempo_factor": seg["segment_tempo_factor"], "effective_tempo": seg["effective_tempo"], "rms_dbfs_before": seg.get("rms_dbfs_before"), "rms_dbfs_after": seg.get("rms_dbfs_after"), "peak_after": seg.get("peak_after"), "output_start_sample": seg.get("output_start_sample"), "output_end_sample": seg.get("output_end_sample"), "adopted_gain": seg.get("adopted_gain"), } for seg in tts_segments ], } if final_output is not None: payload["final_output"] = str(Path(final_output).resolve()) provenance = _timeline_provenance_status(work_dir) if provenance: payload["timeline_provenance"] = provenance subtitle_evidence = manifest_subtitle_evidence(work_dir, input_video, output_path) if subtitle_evidence is not None: payload["subtitle_track"] = subtitle_evidence return payload def _write_assembly_manifest(work_dir, manifest): path = Path(work_dir) / ASSEMBLY_MANIFEST path.write_text(json.dumps(manifest, ensure_ascii=False, indent=2), encoding="utf-8") return path def _visual_qc_rollup(visual_qc): subtitles = visual_qc["subtitles"] overlays = visual_qc["overlays"] return { "present": True, "artifact": visual_qc["artifact"], "verdict": visual_qc["verdict"], "blocking": visual_qc["blocking"], "blocking_codes": list(visual_qc["blocking_codes"]), "summary": visual_qc["summary"], "geometry": visual_qc["geometry"], "subtitles": { "entries": subtitles["entries"], "multi_line": subtitles["multi_line"], "overflow": subtitles["overflow"], "safe_area": subtitles["safe_area"], }, "mask": visual_qc["mask"], "overlays": { "present": overlays["present"], "rendered": overlays["rendered"], "unsupported": overlays["unsupported"], "overflow": overlays["overflow"], }, } def _placed_audio_matches_timeline(seg): """True when the persisted per-beat WAV is exactly what the serialized timeline window plays.""" placed_path = Path(seg["placed_audio_path"]) if not placed_path.exists(): return False try: with wave.open(str(placed_path), "rb") as placed_wav: placed_duration = placed_wav.getnframes() / placed_wav.getframerate() tolerance = 1.0 / placed_wav.getframerate() except wave.Error: result = subprocess.run( ["ffprobe", "-v", "error", "-select_streams", "a:0", "-show_entries", "stream=sample_rate,time_base,duration_ts", "-of", "json", str(placed_path)], capture_output=True, text=True, timeout=600, ) streams = json.loads(result.stdout).get("streams", []) if not result.returncode else [] if len(streams) != 1: return False stream = streams[0] try: rate = int(stream["sample_rate"]) placed_duration = float(Fraction(stream["duration_ts"]) * Fraction(stream["time_base"])) tolerance = 1.0 / rate except (KeyError, TypeError, ValueError, ZeroDivisionError): return False timeline_start = math.floor(float(seg["actual_place_start"]) * 10_000 + 1e-9) / 10_000 timeline_end = math.ceil(float(seg["actual_place_end"]) * 10_000 - 1e-9) / 10_000 serialized_span = timeline_end - timeline_start return ( abs(placed_duration - seg["placed_audio_duration"]) <= tolerance + 1e-9 and serialized_span + 1e-9 >= placed_duration ) def _build_assembly_qc(tts_segments, video_duration, *, audio_operations, render_delivery, output_path=None, source_has_audio=None, loudness_mode=None, loudnorm_measurement=None, visual_qc=None, audio_mode="narration", adopted_audio=None, narration_input_binding=None, audio_mix_binding=None, source_audio_status=None): """Machine-readable assembly release gate. Visual facts are rolled up from visual_qc.json. Delivery/render facts live here (or render/delivery QC in future), never in visual_qc.json. """ hard_max = CONFIG["narration_cumulative_tempo_hard_max"] segments = tts_segments no_safe = [ s["index"] for s in segments if s["fit_status"] == "no_safe_fit" or s["truncate_reason"] in {"no_safe_boundary", "no_room"} or s["blocking"] ] skipped = [s["index"] for s in segments if s["fit_status"] == "skipped"] tempo_exceeded = [] truncated = [] handoff_failed = [] timeline_audio_failed = [] max_effective = 0.0 for s in segments: eff = float(s["effective_tempo"]) max_effective = max(max_effective, eff) if eff > hard_max + 1e-6: tempo_exceeded.append(s["index"]) if s["truncated"] or s["truncate_reason"] == "tail_trim_tolerance": truncated.append(s["index"]) if s.get("source_handoff_blocking", False): handoff_failed.append(s["index"]) if s["placed_audio_duration"] > 0 and not _placed_audio_matches_timeline(s): timeline_audio_failed.append(s["index"]) placed = [s["placed_audio_duration"] for s in segments] blocking_codes = [] if audio_mode == "narration" and not segments: blocking_codes.append("missing_narration") if skipped: blocking_codes.append("skipped_segments") if no_safe: blocking_codes.append("no_safe_fit") if tempo_exceeded: blocking_codes.append("effective_tempo_exceeded") if truncated: blocking_codes.append("truncated_speech") if handoff_failed: blocking_codes.append("unsafe_source_handoff") if timeline_audio_failed: blocking_codes.append("timeline_audio_mismatch") if placed and max(placed) <= 0.0 and not no_safe: blocking_codes.append("empty_narration") visual_rollup = ( _visual_qc_rollup(visual_qc) if visual_qc is not None else {"present": False, "verdict": "NOT_RUN", "blocking": False, "blocking_codes": [], "summary": {}} ) if visual_rollup["blocking"]: blocking_codes.append("visual_qc_failed") if source_audio_status is not None: source_audio = source_audio_status elif source_has_audio is False: # Not blocking: assemble can synthesize a silent original track. source_audio = "synthetic_silence" elif source_has_audio is True: source_audio = "present" else: source_audio = "unknown" output = {} if output_path is not None: output_path = Path(output_path) output = { "path": str(output_path), "exists": output_path.exists(), "bytes": output_path.stat().st_size if output_path.exists() else 0, } if output_path.exists() and output["bytes"] <= 0: blocking_codes.append("empty_output") return { "schema_version": 1, "artifact": ASSEMBLY_QC, "verdict": "FAIL" if blocking_codes else "PASS", "blocking": bool(blocking_codes), "blocking_codes": blocking_codes, "duration": round(float(video_duration), 4), "audio_mode": audio_mode, "audio_operations": audio_operations, "adopted_audio": adopted_audio, "narration_input_binding": narration_input_binding, "audio_mix_binding": audio_mix_binding, "source_audio": source_audio, "loudness_mode": loudness_mode or _loudness_mode(loudnorm_measurement), "loudnorm_measurement": loudnorm_measurement, "release_gate": { "verdict": "FAIL" if blocking_codes else "PASS", "visual_qc": visual_rollup["verdict"], "delivery_qc": "PASS", "audio_qc": "FAIL" if _AUDIO_QC_CODES.intersection(blocking_codes) else "PASS", }, "visual_qc": visual_rollup, "delivery_qc": { "video_encode_passes": render_delivery["video_encode_passes"], "reencode_reason": render_delivery["reencode_reason"], "audio_sample_rate": render_delivery["audio_sample_rate"], "final_compat_notes": render_delivery["final_compat_notes"], }, "summary": { "segments": len(segments), "placed_segments": sum(1 for x in placed if x > 0.0), "skipped_segments": skipped, "no_safe_fit_segments": no_safe, "tempo_exceeded_segments": tempo_exceeded, "max_effective_tempo": round(max_effective, 4), "truncated_segments": truncated, "unsafe_source_handoff_segments": handoff_failed, "timeline_audio_mismatch_segments": timeline_audio_failed, }, "output": output, } def _write_assembly_qc(work_dir, qc): path = Path(work_dir) / ASSEMBLY_QC path.write_text(json.dumps(qc, ensure_ascii=False, indent=2), encoding="utf-8") return path def _resolve_final_output(base, stem): """The recap output is the stable human alias recap_<stem>.mp4, overwritten in place on every run so the iterate-on-narration loop always refreshes the same file.""" return Path(base) / f"recap_{stem}.mp4" -
assembly_settings.py 6.6 KB
"""Render-affecting settings payload recorded in the manifest and compared by resume logic.""" from pathlib import Path from artifacts import _artifact_identity from assemble_constants import VISUAL_OVERLAYS from audio_mix import _loudness_mode, final_loudnorm_filter from lib import CONFIG from source_subtitles import _has_user_subtitles, _source_subtitle_mask_policy from subtitles.core import _subtitle_style_config from adoption.narration_binding import binding_record from adoption.audio_mix_binding import binding_record as audio_mix_binding_record from packaging import packaging_settings def assembly_settings_payload(work_dir=None, *, audio_mode="narration", audio_stream_index=0): """Settings that affect the rendered video, as a plain nested dict compared with ``==`` by pipeline resume logic. When work_dir is given, a user_subtitles presence flag and the ``{size, mtime_ns}`` identity of the overlay/subtitle-track inputs are included so dropping in or rewriting one of those files rebuilds the cached subtitles.""" burn_subtitles = CONFIG["burn_subtitles"] mask_policy = _source_subtitle_mask_policy(work_dir) mask_source_subtitles = mask_policy["active"] overlay_identity = ( _artifact_identity(Path(work_dir) / VISUAL_OVERLAYS) if work_dir is not None else None ) subtitle_track_identity = ( _artifact_identity(Path(work_dir) / "subtitle_track.json") if work_dir is not None else None ) settings = { "user_subtitles": _has_user_subtitles(work_dir), "burn_subtitles": burn_subtitles, "force_video_reencode": CONFIG["force_video_reencode"], "encode": { "output_crf": CONFIG["output_crf"], "output_preset": CONFIG["output_preset"], "output_max_height": CONFIG["output_max_height"], }, "video_filters": { "mask_source_subtitles": mask_source_subtitles, "source_subtitle_mask_policy": mask_policy["policy"], "source_subtitle_mask_policy_declared": mask_policy["declared"], "source_subtitle_mask_policy_trigger": mask_policy["trigger"], "source_subtitle_mask_ratio": ( CONFIG["source_subtitle_mask_ratio"] if mask_source_subtitles else None ), "source_subtitle_mask_timing": ( CONFIG["source_subtitle_mask_timing"] if mask_source_subtitles else None ), "subtitle_mask_opacity": ( CONFIG["subtitle_mask_opacity"] if mask_source_subtitles else None ), "subtitle_mask_padding": ( CONFIG["subtitle_mask_padding"] if mask_source_subtitles else None ), "subtitle_y_top": CONFIG["subtitle_y_top"], "subtitle_y_bot": CONFIG["subtitle_y_bot"], "visual_overlays": { "artifact": VISUAL_OVERLAYS, "present": overlay_identity is not None, "identity": overlay_identity, }, "packaging_layers": packaging_settings(work_dir), }, "audio": { "mode": audio_mode, "selected_stream_index": audio_stream_index, }, } explicit_mix = ( audio_mix_binding_record(work_dir) if work_dir is not None and audio_mode == "narration" else None ) if explicit_mix: settings["audio"]["path"] = "explicit_adopted_full_sound" settings["audio_mix_binding"] = explicit_mix if audio_mode == "narration": narration_binding = binding_record(work_dir) if work_dir else None settings["narration_input_binding"] = narration_binding adopted_tempo = ( narration_binding.get("tempo_policy") if narration_binding else None ) settings["narration_timing"] = { "delay_seconds": 0.0 if explicit_mix else CONFIG["narration_delay_seconds"], "tail_pad_seconds": 0.0 if explicit_mix else CONFIG["narration_tail_pad_seconds"], "fade_ms": 0 if explicit_mix else CONFIG["fade_ms"], "narration_speed": ( 1.0 if explicit_mix else adopted_tempo["global_atempo"] if adopted_tempo else CONFIG["narration_speed"] ), "tempo_source": ( "explicit_audio_mix" if explicit_mix else "adoption" if adopted_tempo else "configuration" ), "narration_cumulative_tempo_max": CONFIG["narration_cumulative_tempo_max"], "tts_segment_tempo_max": ( adopted_tempo["segment_tempo_max"] if explicit_mix and adopted_tempo else CONFIG["tts_segment_tempo_max"] ), } if explicit_mix and adopted_tempo: settings["narration_timing"]["narration_cumulative_tempo_max"] = \ adopted_tempo["cumulative_tempo_max"] settings["narration_timing"]["narration_cumulative_tempo_hard_max"] = \ adopted_tempo["cumulative_tempo_hard_max"] if audio_mode in {"narration", "source-mix"} and not explicit_mix: # adopted-packet-copy never decodes or mixes, so mix settings cannot change it. settings["audio_mix"] = { "ducking_mode": CONFIG["ducking_mode"], "duck_fade_seconds": CONFIG["duck_fade_seconds"], "duck_bridge_seconds": CONFIG["duck_bridge_seconds"], "ducking_narr_weight": CONFIG["ducking_narr_weight"], "ducking_orig_volume": CONFIG["ducking_orig_volume"], "idle_orig_volume": CONFIG["idle_orig_volume"], "speech_ducking_volume": CONFIG["speech_ducking_volume"], "zone_ducking_volume": CONFIG["zone_ducking_volume"], "ducking_threshold": CONFIG["ducking_threshold"], "ducking_ratio": CONFIG["ducking_ratio"], "ducking_attack": CONFIG["ducking_attack"], "ducking_release": CONFIG["ducking_release"], "ducking_level_sc": CONFIG["ducking_level_sc"], "ducking_makeup": CONFIG["ducking_makeup"], "final_loudnorm": final_loudnorm_filter(), "loudness_mode": _loudness_mode(), "bgm_path": CONFIG["bgm_path"], "bgm_volume": CONFIG["bgm_volume"], "bgm_ducking_volume": CONFIG["bgm_ducking_volume"], } if burn_subtitles: settings["subtitle_renderer"] = "ass" settings["subtitle_style"] = _subtitle_style_config() if subtitle_track_identity is not None: settings["subtitle_track"] = { "artifact": "subtitle_track.json", "identity": subtitle_track_identity, } return settings -
audio_automation.py 6.7 KB
"""Shared audio automation semantics for ffmpeg render and editable timelines. This module is stdlib-only and is the single source for ducking window coalescing, pre-roll/hold/post-roll gain shape, timeline keyframes, and ffmpeg volume terms. """ def _round_keyframe(t_s, gain): return {"t": round(float(t_s), 4), "gain": round(float(gain), 4)} def default_bridge(fade): return 2 * float(fade) def coalesce_duck_windows(windows, bridge): """Merge [(start, end, gain)] windows separated by gaps below `bridge`. A bridged mixed-level span uses the lowest gain across all members so neither ffmpeg nor timeline export swells louder in the bridged gap. """ rel = sorted( ([float(s), float(e), float(g)] for s, e, g in windows if float(e) > float(s)), key=lambda w: w[0], ) if not rel: return [] bridge = float(bridge) merged = [rel[0][:]] for s, e, gain in rel[1:]: if s - merged[-1][1] < bridge: merged[-1][1] = max(merged[-1][1], e) merged[-1][2] = min(merged[-1][2], gain) else: merged.append([s, e, gain]) return merged def _t_minus(value): value = float(value) if value < 0: return f"t+{abs(value):.2f}" return f"t-{value:.2f}" def duck_ramp_expression(start, end, fade): """Return ffmpeg expression for the canonical duck shape. Semantics: ramp down during [start-fade, start], hold fully ducked on [start, end], and release during [end, end+fade]. """ start = float(start) end = float(end) fade = float(fade) if fade <= 0: return f"between(t,{start:.2f},{end:.2f})" ramp_start = start - fade ramp_end = end + fade return f"min(1,max(0,min({_t_minus(ramp_start)},{ramp_end:.2f}-t)/{fade:.2f}))" def ducking_expression(windows, idle, fade): """Build the ffmpeg volume expression for coalesced duck windows.""" if not windows: return None idle = float(idle) terms = [ f"+({float(level) - idle:.3f})*{duck_ramp_expression(s, e, fade)}" for s, e, level in windows ] return f"max(0,min(1,{idle}{''.join(terms)}))" def coalesce_release_duck_windows(windows, bridge): """Merge [(start, hold_end, gain, restore_at)] while preserving safe releases.""" rel = sorted( ([float(s), float(e), float(g), max(float(e), float(r))] for s, e, g, r in windows if float(e) > float(s)), key=lambda row: row[0], ) if not rel: return [] bridge = float(bridge) merged = [rel[0][:]] for start, end, gain, restore in rel[1:]: if start - merged[-1][1] < bridge: merged[-1][1] = max(merged[-1][1], end) merged[-1][2] = min(merged[-1][2], gain) merged[-1][3] = max(merged[-1][1], merged[-1][3], restore) else: merged.append([start, end, gain, restore]) return merged def release_ducking_expression(windows, idle, attack_fade, bridge=None): """FFmpeg gain expression with a fixed attack and a per-window safe release end.""" attack = float(attack_fade) if bridge is None: bridge = default_bridge(attack) merged = coalesce_release_duck_windows(windows, bridge) if not merged: return None idle = float(idle) terms = [] for start, hold_end, level, restore_at in merged: if attack > 0: attack_term = f"({_t_minus(start - attack)})/{attack:.4f}" else: attack_term = f"between(t,{start:.4f},{restore_at:.4f})" release = restore_at - hold_end if release > 1e-6: release_term = f"({restore_at:.4f}-t)/{release:.4f}" mask = f"min(1,max(0,min({attack_term},{release_term})))" else: mask = f"between(t,{start:.4f},{hold_end:.4f})" terms.append(f"+({float(level) - idle:.3f})*{mask}") return f"max(0,min(1,{idle}{''.join(terms)}))" def release_ducking_keyframes(windows, idle, attack_fade, span_start, span_end, bridge=None): """Timeline keyframes matching `release_ducking_expression` exactly.""" attack = float(attack_fade) if bridge is None: bridge = default_bridge(attack) span_start, span_end = float(span_start), float(span_end) normalized = [ (max(span_start, start), min(span_end, end), gain, min(span_end, max(end, restore))) for start, end, gain, restore in ((float(s), float(e), float(g), float(r)) for s, e, g, r in windows) if end > span_start and start < span_end and end > start ] merged = coalesce_release_duck_windows(normalized, bridge) if not merged: return [] points = [(span_start, float(idle))] for start, hold_end, level, restore_at in merged: points.extend([ (max(span_start, start - attack), float(idle)), (start, level), (hold_end, level), (restore_at, float(idle)), ]) points.append((span_end, float(idle))) points.sort(key=lambda point: point[0]) out = [] for when, gain in points: if out and abs(out[-1][0] - when) < 1e-4: out[-1] = (when, min(out[-1][1], gain)) else: out.append((when, gain)) return [_round_keyframe(when, gain) for when, gain in out] def variable_ducking_keyframes(windows, idle, fade, span_start, span_end, bridge=None): """Volume keyframes for per-window duck gains using canonical semantics.""" fade = float(fade) if bridge is None: bridge = default_bridge(fade) span_start = float(span_start) span_end = float(span_end) rel = sorted( (max(span_start, float(w[0])), min(span_end, float(w[1])), float(w[2])) for w in windows if float(w[1]) > span_start and float(w[0]) < span_end and float(w[1]) > float(w[0]) ) merged = coalesce_duck_windows(rel, bridge) if not merged: return [] pts = [(span_start, float(idle))] for s, e, level in merged: pts.append((max(span_start, s - fade), float(idle))) pts.append((s, level)) pts.append((e, level)) pts.append((min(span_end, e + fade), float(idle))) pts.append((span_end, float(idle))) pts.sort(key=lambda p: p[0]) out = [] for t, gain in pts: if out and abs(out[-1][0] - t) < 1e-4: out[-1] = (t, min(out[-1][1], gain)) else: out.append((t, gain)) return [_round_keyframe(t, gain) for t, gain in out] def fixed_ducking_keyframes(windows, idle, duck, fade, span_start, span_end, bridge=None): """Volume keyframes for fixed-gain duck windows using canonical semantics.""" return variable_ducking_keyframes( [(float(s), float(e), float(duck)) for s, e in windows], idle, fade, span_start, span_end, bridge=bridge, ) -
audio_mix.py 18.5 KB
"""Loudness, source handoffs, ducking envelopes, and audio mix graphs.""" import json import re from pathlib import Path from artifacts import _load_work_json from audio_automation import ( coalesce_duck_windows, ducking_expression, release_ducking_expression, ) from lib import CONFIG, filter_file_args, log, run_cmd def _limiter_filter(): return f"alimiter=limit={CONFIG['final_limiter_peak']:.2f}:level=false" def _loudness_mode(measured=None): if not CONFIG["final_loudnorm"]: return "limiter_only" return "two_pass_linear" if measured else "equivalent" def final_loudnorm_filter(measured=None): """Final-mix loudness normalization/limiter filter from CONFIG. Ducking branches set only relative balance; this single stage owns the absolute output loudness so the recap is not left too quiet. When `measured` is supplied from a first loudnorm pass, ffmpeg runs the deterministic second pass; without it we still force the same target and peak limiter as a documented equivalent/fallback path. """ if not CONFIG["final_loudnorm"]: return _limiter_filter() filt = ( f"loudnorm=I={CONFIG['target_lufs']}" f":TP={CONFIG['target_true_peak']}" f":LRA={CONFIG['target_lra']}" f":linear=true" ) if measured: for src, dst in ( ("input_i", "measured_I"), ("input_tp", "measured_TP"), ("input_lra", "measured_LRA"), ("input_thresh", "measured_thresh"), ("target_offset", "offset"), ): if src in measured: filt += f":{dst}={measured[src]}" filt += ":print_format=summary" return f"{filt},{_limiter_filter()}" def _parse_loudnorm_json(text): """Extract ffmpeg loudnorm JSON from stderr/stdout.""" for match in reversed(list(re.finditer(r"\{[\s\S]*?\}", text))): try: data = json.loads(match.group(0)) except ValueError: continue if {"input_i", "input_tp", "input_lra", "input_thresh", "target_offset"} <= set(data): return data return None def _loudnorm_first_pass_filter(): return ( f"loudnorm=I={CONFIG['target_lufs']}" f":TP={CONFIG['target_true_peak']}" f":LRA={CONFIG['target_lra']}" f":print_format=json" ) def _run_loudnorm_first_pass(input_video, narration_wav, original_audio_input, bgm_input, filter_complex, work_dir): """Measure the exact mixed audio graph before final render. Returns ffmpeg loudnorm JSON, or None when probing fails. The caller then falls back to the documented equivalent single-pass target+limiter filter. """ if not CONFIG["final_loudnorm"]: return None probe_fc = f"{filter_complex};[aout]{_loudnorm_first_pass_filter()}[lnprobe]" probe_script = Path(work_dir) / ".filter_complex_loudnorm_probe.txt" probe_script.write_text(probe_fc, encoding="utf-8") cmd = [ "ffmpeg", "-y", "-i", str(input_video), "-i", str(narration_wav), *original_audio_input, *bgm_input, *filter_file_args("filter_complex", probe_script), "-map", "[lnprobe]", "-f", "null", "-", ] try: result = run_cmd(cmd) finally: probe_script.unlink(missing_ok=True) if result.returncode != 0: log(f" ⚠️ loudnorm 首遍测量失败,降级到目标滤镜+limiter: {result.stderr}") return None measured = _parse_loudnorm_json(result.stdout + "\n" + result.stderr) if not measured: log(" ⚠️ loudnorm 首遍未返回 JSON,降级到目标滤镜+limiter") return None return measured def _seg_place_window(seg): """A segment's actual placed (start, end) on the output timeline; zero-width when unplaced.""" return seg["actual_place_start"], seg["actual_place_end"] def _load_sentence_handoff_anchors(work_dir): """Load high/medium sentence anchors and their measured pause windows.""" work_dir = Path(work_dir) cut_mode = (work_dir / "edited_source.mp4").exists() or ( work_dir / "clip_plan_validated.json" ).exists() artifact = "speech_boundary_anchors_output.json" if cut_mode else "speech_boundary_anchors.json" payload = _load_work_json(work_dir, artifact) if payload is None: return [], None, {"require_measured": cut_mode} if cut_mode: # Output-clock anchors are only trusted when they are at least as new as the cut plan. plan_path = work_dir / "clip_plan_validated.json" fresh = ( payload.get("schema_version") == 2 and payload.get("timeline") == "cut_output" and plan_path.exists() and (work_dir / artifact).stat().st_mtime_ns >= plan_path.stat().st_mtime_ns ) if not fresh: return [], None, {"require_measured": True} payload = {**payload, "require_measured": True} anchors = {} for item in payload["sentence_anchors"]: if item["confidence"] not in {"high", "medium"}: continue when = float(item["time"]) pause_start = float(item.get("pause_start", when - 0.12)) row = { "time": round(when, 4), "pause_start": round(max(0.0, min(pause_start, when)), 4), } anchors[(row["time"], row["pause_start"])] = row return sorted(anchors.values(), key=lambda row: row["time"]), artifact, payload def _timed_rows(rows): return [{"start": float(row["start"]), "end": float(row["end"])} for row in rows] def _asr_segments(work_dir): """Cleaned ASR segments (asr_clean.json) when present, else raw asr_result.json; [] when absent.""" clean = _load_work_json(work_dir, "asr_clean.json") if clean is not None: return clean["segments"] return _load_work_json(work_dir, "asr_result.json") or [] def _handoff_speech_evidence(work_dir, payload): speech = _timed_rows(payload.get("speech_spans", [])) quiet = _timed_rows(payload.get("quiet_windows", [])) if payload.get("require_measured"): return speech, quiet if not speech: speech = _timed_rows(_asr_segments(work_dir)) if not quiet: silence = _load_work_json(work_dir, "silence_periods.json") or [] quiet = _timed_rows(row for row in silence if not row["has_speech"]) return speech, quiet def _merged_handoff_intervals(start, end, rows): intervals = sorted( (max(start, row["start"]), min(end, row["end"])) for row in rows if row["end"] > start and row["start"] < end ) merged = [] for left, right in intervals: if right <= left: continue if merged and left <= merged[-1][1]: merged[-1] = (merged[-1][0], max(merged[-1][1], right)) else: merged.append((left, right)) return merged def _speech_overlap_excluding_quiet(start, end, speech, quiet): speech_intervals = _merged_handoff_intervals(start, end, speech) quiet_intervals = _merged_handoff_intervals(start, end, quiet) overlap = sum(right - left for left, right in speech_intervals) for speech_left, speech_right in speech_intervals: overlap -= sum( max(0.0, min(speech_right, quiet_right) - max(speech_left, quiet_left)) for quiet_left, quiet_right in quiet_intervals ) return max(0.0, overlap) def _measured_speech_owned( start, end, speech, quiet, anchors, authored, require_measured=False ): duration = max(0.0, end - start) quiet_min = max(0.3, duration * CONFIG["quiet_overlap_min_ratio"]) if speech: return _speech_overlap_excluding_quiet(start, end, speech, quiet) > 0.05 quiet_overlap = sum( right - left for left, right in _merged_handoff_intervals(start, end, quiet) ) if quiet and quiet_overlap >= quiet_min: return False return True if anchors or require_measured else bool(authored) def _entry_speech_owned( start, speech, quiet, anchors, authored, require_measured=False, tolerance=0.05 ): if any(row["start"] - tolerance <= start <= row["end"] + tolerance for row in quiet): return False if any(row["start"] - tolerance <= start < row["end"] - tolerance for row in speech): return True if speech: return False return True if anchors or require_measured else bool(authored) def _work_has_source_speech(work_dir, speech_spans, require_measured): if speech_spans or require_measured: return True return any(item["text"].strip() for item in _asr_segments(work_dir)) def _apply_source_sentence_handoffs(tts_segments, work_dir, video_duration): """Keep source audio ducked until a safe sentence boundary after narration. This does not move or trim narration. It only extends the ORIGINAL-audio duck envelope so returning the source track cannot reveal the middle of a sentence. """ fade = CONFIG["duck_fade_seconds"] bridge = CONFIG["duck_bridge_seconds"] anchors, artifact, evidence_payload = _load_sentence_handoff_anchors(work_dir) speech_spans, quiet_windows = _handoff_speech_evidence(work_dir, evidence_payload) require_measured = evidence_payload.get("require_measured", False) placed = [] for seg in tts_segments: start, end = _seg_place_window(seg) if end > start: placed.append((start, end, seg)) placed.sort(key=lambda item: (item[0], item[1])) if not placed: return [] runs = [] for start, end, seg in placed: if runs and start - runs[-1]["end"] <= bridge + 1e-6: runs[-1]["end"] = max(runs[-1]["end"], end) runs[-1]["segments"].append(seg) else: runs.append({"start": start, "end": end, "segments": [seg]}) source_has_speech = _work_has_source_speech(work_dir, speech_spans, require_measured) report = [] for run in runs: ownership = [] for seg in run["segments"]: start, end = _seg_place_window(seg) measured = _measured_speech_owned( start, end, speech_spans, quiet_windows, anchors, seg["overlaps_speech"], require_measured=require_measured, ) seg["overlaps_speech"] = measured ownership.append(measured) first = run["segments"][0] entry_owned = _entry_speech_owned( run["start"], speech_spans, quiet_windows, anchors, first["overlaps_speech"], require_measured=require_measured, ) speech_owned = entry_owned or any(ownership) if not speech_owned: report.append({"start": run["start"], "end": run["end"], "status": "quiet_source"}) continue last = run["segments"][-1] start_safe = run["start"] <= 0.25 or any( anchor["pause_start"] - 0.05 <= run["start"] <= anchor["time"] + 0.08 for anchor in anchors ) if entry_owned and anchors and not start_safe: first["source_handoff_blocking"] = True first["source_entry_status"] = "unsafe_entry" elif not entry_owned: first["source_entry_status"] = "quiet_source" else: first["source_entry_status"] = "sentence_boundary" if anchors else "unverified" restore_anchor = next( (anchor for anchor in anchors if anchor["time"] >= run["end"] - 0.01), None, ) if restore_anchor is not None: # Hold the source low through its last spoken sample, then fit the release # entirely inside the measured pause. Never begin the ramp `fade` seconds # before the anchor when that would expose the final source phoneme. duck_end = max(run["end"], restore_anchor["pause_start"]) restore_at = max(duck_end, restore_anchor["time"]) status = "sentence_boundary" elif anchors: # No later complete source sentence: never expose a fragment at the tail. restore_at = float(video_duration) duck_end = float(video_duration) status = "held_to_timeline_end" elif source_has_speech: first["source_handoff_blocking"] = True first["source_entry_status"] = "anchors_unavailable" restore_at = run["end"] + fade duck_end = run["end"] status = "anchors_unavailable" else: restore_at = run["end"] + fade duck_end = run["end"] status = "no_source_speech" last["source_duck_end"] = round(min(float(video_duration), duck_end), 4) last["source_restore_at"] = round(min(float(video_duration), restore_at), 4) last["source_handoff_status"] = status report.append({ "start": round(run["start"], 4), "end": round(run["end"], 4), "restore_at": last["source_restore_at"], "status": status, "anchor_artifact": artifact, }) return report def _amix_tail(narr_vol, bgm_chain=""): """Mix the prepared original track [orig] (+ optional BGM bed) with the boosted narration [narr] into [aout]. bgm_chain, when given, defines [bgm] from input [2:a].""" narr = f"[1:a]volume={narr_vol},aresample=48000[narr];" if bgm_chain: return bgm_chain + narr + "[orig][bgm][narr]amix=inputs=3:duration=first:dropout_transition=0:normalize=0[aout]" return narr + "[orig][narr]amix=inputs=2:duration=first:dropout_transition=0:normalize=0[aout]" def _duck_envelope(tts_segments, idle, speech_vol, quiet_vol, fade, bridge): """Per-beat ducking automation for the ORIGINAL track. Uses the shared ducking contract: [start-fade,start] pre-roll ramp down, [start,end] held at the selected duck level, and [end,end+fade] release. Bridged spans use the most-ducked (lowest) level, matching timeline.json / JianYing keyframes. Returns a volume= expression, or None when no beat was placed (caller falls back to a constant). """ windows = [] for seg in tts_segments: start, narration_end = _seg_place_window(seg) if narration_end <= start: continue hold_end = max(narration_end, seg.get("source_duck_end", narration_end)) restore_at = max(hold_end, seg.get("source_restore_at", hold_end + fade)) level = speech_vol if seg["overlaps_speech"] else quiet_vol windows.append((start, hold_end, level, restore_at)) return release_ducking_expression(windows, idle, fade, bridge=bridge) def _bgm_envelope(tts_segments, base, duck, fade, bridge): """Per-beat ducking automation for the BGM track using the shared contract.""" windows = [ (start, end, duck) for start, end in map(_seg_place_window, tts_segments) if end > start ] return ducking_expression(coalesce_duck_windows(windows, bridge), base, fade) def _build_audio_filter_complex( tts_segments, has_bgm=False, *, original_audio_label="0:a", bgm_audio_label="2:a", ): """Compose the audio tracks into [aout], like a cut-software timeline. Tracks: - original (input [0:a], the video's own audio): ducked under each narration window by a per-beat volume envelope, but held up at `idle_orig_volume` in the gaps so the recap never drops to dead air between sentences. - bgm (input [2:a], optional): a looped music bed, gently ducked under narration. - narration (input [1:a]): the TTS, boosted and laid on top. CONFIG["ducking_mode"] (default "fixed") selects the original-track strategy: fixed = the gap-fill envelope above; sidechaincompress = auto-duck keyed off the narration; none = no ducking. Placement comes from actual_place_start/end. """ ducking_mode = CONFIG["ducking_mode"] if ducking_mode == "sidechaincompress" and any( "source_duck_end" in seg and seg["source_duck_end"] > seg["actual_place_end"] + 1e-6 for seg in tts_segments ): log("sidechaincompress 无法保持句末交接窗口,已回退 fixed ducking") ducking_mode = "fixed" narr_vol = CONFIG["ducking_narr_weight"] fade = CONFIG["duck_fade_seconds"] bridge = CONFIG["duck_bridge_seconds"] original_in = f"[{original_audio_label}]" bgm_in = f"[{bgm_audio_label}]" # BGM bed (input [2:a]): ducked under each narration window when present. bgm_chain = "" if has_bgm: base = CONFIG["bgm_volume"] bgm_expr = _bgm_envelope(tts_segments, base, CONFIG["bgm_ducking_volume"], fade, bridge) if bgm_expr: bgm_chain = f"{bgm_in}volume='{bgm_expr}':eval=frame,aresample=48000[bgm];" else: bgm_chain = f"{bgm_in}volume={base},aresample=48000[bgm];" if ducking_mode == "sidechaincompress": # The narration keys the compressor; split it so it can also be mixed in. head = ( f"{original_in}aresample=48000[o0];" "[1:a]aresample=48000,asplit=2[sckey][scnarr];" f"[o0][sckey]sidechaincompress=" f"threshold={CONFIG['ducking_threshold']}:ratio={CONFIG['ducking_ratio']}" f":attack={CONFIG['ducking_attack']}:release={CONFIG['ducking_release']}" f":knee=2.5:makeup={CONFIG['ducking_makeup']}:level_sc={CONFIG['ducking_level_sc']}[orig];" ) narr = f"[scnarr]volume={narr_vol}[narr];" if bgm_chain: return head + bgm_chain + narr + "[orig][bgm][narr]amix=inputs=3:duration=first:dropout_transition=0:normalize=0[aout]" return head + narr + "[orig][narr]amix=inputs=2:duration=first:dropout_transition=0:normalize=0[aout]" if ducking_mode == "none": return f"{original_in}aresample=48000[orig];" + _amix_tail(narr_vol, bgm_chain) # fixed (default): gap-fill ducking envelope on the original track. idle = CONFIG["idle_orig_volume"] speech_vol = CONFIG["speech_ducking_volume"] quiet_vol = CONFIG["zone_ducking_volume"] expr = _duck_envelope(tts_segments, idle, speech_vol, quiet_vol, fade, bridge) if expr: n_overlap = sum(1 for s in tts_segments if s["overlaps_speech"]) n_quiet = len(tts_segments) - n_overlap log(f"gap-fill ducking: 间隙原声={idle}, 对白段={speech_vol}({n_overlap}), 安静段={quiet_vol}({n_quiet}), 桥接间隙<{bridge}s") orig = f"{original_in}volume='{expr}':eval=frame,aresample=48000[orig];" else: # No placement info at all: hold the original at a constant level. orig = f"{original_in}volume={CONFIG['ducking_orig_volume']},aresample=48000[orig];" return orig + _amix_tail(narr_vol, bgm_chain) -
compose_foreground.py 14.1 KB
#!/usr/bin/env python3 """Compose caller-rendered RGBA pixels over a frozen H264/CFR/AAC base. This operation validates pixels and declared inputs. It does not generate or interpret titles, dialogue, brands, typography, or release approval. """ import argparse from fractions import Fraction import json from pathlib import Path import subprocess from adoption.frozen_audio import probe_audio_packets, verify_adopted_audio from pair_media import probe_picture, validate_pair_timing from adoption.strict_inputs import ( canonical_fraction, require_declared_path, require_fields, require_integer, run_logged, without_digests, write_json_atomic, ) PATTERN = "frame_%06d.png" ENCODING = { "video_codec": "libx264", "preset": "fast", "crf": 18, "pixel_format": "yuv420p", "color": "bt709_tv", "audio_codec": "copy", } def _fields(value, required): value = without_digests(value, "compose plan") require_fields(value, required, "compose plan") return value def _local_file(value, label): _fields(value, ["path"]) return {"path": str(require_declared_path(value, label))} def _sequence_paths(directory, pattern, start, end): if pattern != PATTERN: raise ValueError(f"Only the literal simple pattern {PATTERN!r} is supported") directory = Path(directory).resolve() if not directory.is_dir(): raise FileNotFoundError(f"Sequence directory missing: {directory}") expected = [directory / (pattern % index) for index in range(start, end)] actual = sorted(directory.iterdir(), key=lambda item: item.name) if [item.name for item in actual] != [item.name for item in expected]: raise ValueError("Sequence requires exact contiguous files with no missing or extra entries") if not all(path.is_file() for path in expected): raise FileNotFoundError("Sequence frame missing or not a file") return directory, expected def _probe_image(path): result = subprocess.run( ["ffprobe", "-v", "error", "-select_streams", "v:0", "-show_streams", "-show_entries", "stream=codec_name,width,height,pix_fmt", "-of", "json", str(path)], capture_output=True, text=True, timeout=600, ) if result.returncode or result.stderr.strip(): raise ValueError(f"Image probe failed: {result.stderr.strip()}") streams = json.loads(result.stdout).get("streams", []) if len(streams) != 1 or streams[0].get("codec_name") != "png": raise ValueError("Foreground assets must be PNG images") return streams[0] def _validate_rgba(path, width, height): image = _probe_image(path) if (image.get("width"), image.get("height")) != (width, height): raise ValueError("PNG must exactly match the full output canvas") if image.get("pix_fmt") != "rgba": raise ValueError("PNG must contain an actual RGBA pixel format") def _sequence(value, required, *, local_count, width, height): value = _fields(value, required) directory = value["directory"] if not isinstance(directory, str) or not directory or "://" in directory: raise ValueError("Sequence requires a local directory") if value["pattern"] != PATTERN: raise ValueError(f"Only the literal simple pattern {PATTERN!r} is supported") resolved, paths = _sequence_paths(directory, value["pattern"], 0, local_count) for path in paths: _validate_rgba(path, width, height) return {**value, "directory": str(resolved), "frame_count": len(paths)}, paths def validate_endcard(value, *, foreground_end, total_frames, width, height): """Validate the exact endcard union and return normalized data and bound paths.""" if not isinstance(value, dict) or value.get("kind") not in {"none", "still", "sequence"}: raise ValueError("Endcard kind must be none, still, or sequence") value = without_digests(value, "endcard") if value["kind"] == "none": _fields(value, ["kind"]) if foreground_end != total_frames: raise ValueError("No endcard requires foreground to cover the full frame clock") return {"kind": "none"}, [] if value["kind"] == "still": required = ["kind", "path", "start_frame", "end_frame"] else: required = ["kind", "directory", "pattern", "start_frame", "end_frame"] _fields(value, required) end_start = require_integer(value["start_frame"], "endcard start_frame") end_end = require_integer(value["end_frame"], "endcard end_frame", minimum=1) if end_start >= end_end: raise ValueError("Endcard interval must be non-empty") if end_start != foreground_end or end_end != total_frames: raise ValueError("Foreground and endcard must exactly partition the base frame clock") if value["kind"] == "still": asset = _local_file({"path": value["path"]}, "endcard") _validate_rgba(asset["path"], width, height) return {**value, **asset}, [Path(asset["path"])] return _sequence( value, required, local_count=end_end - end_start, width=width, height=height, ) def validate_plan(plan_path): plan_path = Path(plan_path).resolve() plan = json.loads(plan_path.read_bytes()) _fields(plan, ["artifact", "schema_version", "base", "video", "foreground", "endcard", "producer_receipt"]) plan = {**plan, "video": _fields(plan["video"], ["fps", "width", "height", "total_frames"])} if plan["artifact"] != "foreground_compose_plan" or type(plan["schema_version"]) is not int \ or plan["schema_version"] != 1: raise ValueError("Unsupported foreground_compose_plan schema") base = _local_file(plan["base"], "base") receipt = _local_file(plan["producer_receipt"], "producer_receipt") fps = canonical_fraction(plan["video"]["fps"], "video fps") width = require_integer(plan["video"]["width"], "video width", minimum=1) height = require_integer(plan["video"]["height"], "video height", minimum=1) total = require_integer(plan["video"]["total_frames"], "video total_frames", minimum=1) _fields(plan["foreground"], ["directory", "pattern", "start_frame", "end_frame"]) foreground_start = require_integer(plan["foreground"]["start_frame"], "foreground start_frame") foreground_end = require_integer(plan["foreground"]["end_frame"], "foreground end_frame", minimum=1) if foreground_start != 0 or foreground_end > total: raise ValueError("Foreground must start at frame 0 and not exceed the frame clock") foreground, foreground_paths = _sequence( plan["foreground"], ["directory", "pattern", "start_frame", "end_frame"], local_count=foreground_end, width=width, height=height, ) normalized_endcard, endcard_paths = validate_endcard( plan["endcard"], foreground_end=foreground_end, total_frames=total, width=width, height=height, ) picture = probe_picture(base["path"]) decoder = picture["decoder"] required_color = {"codec_name": "h264", "pix_fmt": "yuv420p", "color_range": "tv", "color_space": "bt709", "color_transfer": "bt709", "color_primaries": "bt709"} if any(decoder.get(key) != value for key, value in required_color.items()): raise ValueError("Base must be H264 yuv420p with explicit BT.709 TV color metadata") if (decoder.get("width"), decoder.get("height"), picture["frame_count"], Fraction(picture["fps"])) != (width, height, total, fps): raise ValueError("Declared fps/canvas/count does not match the actual base") audio = probe_audio_packets(base["path"], 0) validate_pair_timing(picture, audio) return { "plan_path": plan_path, "base": base, "video": {"fps": str(fps), "width": width, "height": height, "total_frames": total}, "foreground": foreground, "foreground_paths": foreground_paths, "endcard": normalized_endcard, "endcard_paths": endcard_paths, "producer_receipt": receipt, "picture": picture, "audio": audio, } def _run_ffmpeg(command, directory): run_logged(command, directory, "compose", timeout=600) def _verify_output(path, validated): output_picture = probe_picture(path) expected = validated["picture"] for key in ("frame_pts", "frame_count", "fps", "duration", "start"): if output_picture[key] != expected[key]: raise ValueError(f"Output picture {key} differs from the base") decoder_keys = ("codec_name", "width", "height", "pix_fmt", "sample_aspect_ratio", "field_order", "color_range", "color_space", "color_transfer", "color_primaries", "chroma_location") if {key: output_picture["decoder"].get(key) for key in decoder_keys} != \ {key: expected["decoder"].get(key) for key in decoder_keys}: raise ValueError("Output codec/canvas/color metadata differs from the base") proof = verify_adopted_audio(validated["base"]["path"], path, 0, 0) if validate_pair_timing(output_picture, proof["output"]) != \ validate_pair_timing(expected, validated["audio"]): raise ValueError("Output audio interval changed") result = subprocess.run( ["ffmpeg", "-v", "error", "-xerror", "-threads", "2", "-i", str(path), "-map", "0:v:0", "-map", "0:a:0", "-f", "null", "-"], capture_output=True, text=True, timeout=600, ) if result.returncode or result.stderr.strip(): raise ValueError("Foreground output full decode failed") return output_picture, proof def run_compose(plan_path, output_dir, *, plan_only=False): directory = Path(output_dir).resolve() directory.mkdir(parents=True, exist_ok=False) report_path = directory / "foreground_run.json" staged = directory / "foreground.rendering.mp4" output = directory / "foreground.mp4" report = {"artifact": "foreground_compose_run", "schema_version": 1, "status": "PREPARING", "direct_listening": "NOT_CHECKED", "normal_speed_review": "NOT_CHECKED", "release_approved": False, "encoding": ENCODING} write_json_atomic(report_path, report) try: value = validate_plan(plan_path) report.update( plan={"path": str(value["plan_path"])}, base=value["base"], video=value["video"], foreground=value["foreground"], endcard=value["endcard"], producer_receipt={**value["producer_receipt"], "semantic_validation": "DECLARED_NOT_CHECKED"}, ) if plan_only: report["status"] = "PLANNED" write_json_atomic(report_path, report) return report fps = value["video"]["fps"] command = ["ffmpeg", "-nostdin", "-v", "error", "-n", "-copyts", "-i", value["base"]["path"], "-framerate", fps, "-start_number", "0", "-i", str(Path(value["foreground"]["directory"]) / PATTERN)] if value["endcard"]["kind"] == "none": filters = ( "[0:v]format=rgb24[base];[1:v]format=rgba,setpts=PTS-STARTPTS[fg];" "[base][fg]overlay=0:0:eof_action=pass:format=rgb," "scale=in_range=pc:out_range=tv:" "in_color_matrix=bt709:out_color_matrix=bt709,format=yuv420p," f"trim=end_frame={value['video']['total_frames']}[outv]" ) elif value["endcard"]["kind"] == "still": command += ["-loop", "1", "-framerate", fps, "-i", value["endcard"]["path"]] else: command += ["-framerate", fps, "-start_number", "0", "-i", str(Path(value["endcard"]["directory"]) / PATTERN)] if value["endcard"]["kind"] != "none": end_start = value["endcard"]["start_frame"] end_offset = Fraction(end_start, 1) / Fraction(fps) filters = ( f"[0:v]format=rgb24[base];[1:v]format=rgba,setpts=PTS-STARTPTS[fg];" f"[2:v]format=rgba,setpts=PTS-STARTPTS+{end_offset.numerator}/" f"{end_offset.denominator}/TB[end];" f"[base][fg]overlay=0:0:eof_action=pass:format=rgb:" f"enable='lt(n,{end_start})'[body];" f"[body][end]overlay=0:0:eof_action=repeat:format=rgb:" f"enable='gte(n,{end_start})',scale=in_range=pc:out_range=tv:" "in_color_matrix=bt709:out_color_matrix=bt709,format=yuv420p," f"trim=end_frame={value['video']['total_frames']}[outv]" ) command += ["-filter_complex", filters, "-map", "[outv]", "-map", "0:a:0", "-c:v", ENCODING["video_codec"], "-preset", ENCODING["preset"], "-crf", str(ENCODING["crf"]), "-threads", "2", "-pix_fmt", ENCODING["pixel_format"], "-color_range", "tv", "-colorspace", "bt709", "-color_primaries", "bt709", "-color_trc", "bt709", "-c:a", "copy", "-movie_timescale", str(value["audio"]["sample_rate"]), "-movflags", "+faststart", str(staged)] _run_ffmpeg(command, directory) output_picture, audio_proof = _verify_output(staged, value) write_json_atomic(directory / "picture_identity.json", output_picture) write_json_atomic(directory / "adopted_audio_identity.json", audio_proof) report["output"] = {"path": str(output), "full_decode": "PASS", "frame_clock": "EXACT", "audio_packet_identity": "EXACT"} staged.rename(output) report["status"] = "FOREGROUND_RENDERED" write_json_atomic(report_path, report) return report except Exception as exc: staged.unlink(missing_ok=True) output.unlink(missing_ok=True) report.pop("output", None) report.update(status="FAILED", error=f"{type(exc).__name__}: {exc}") write_json_atomic(report_path, report) raise def main(): parser = argparse.ArgumentParser(description=__doc__) parser.add_argument("plan") parser.add_argument("--output-dir", required=True) parser.add_argument("--plan-only", action="store_true") args = parser.parse_args() report = run_compose(args.plan, args.output_dir, plan_only=args.plan_only) print(json.dumps(report, ensure_ascii=False, indent=2)) if __name__ == "__main__": main() -
export_jianying.py 8.5 KB
"""Optional 剪映 / JianYing (CapCut) draft exporter — decoupled, stdlib + ffprobe only. Reads a backend-neutral `timeline.json` (see timeline.py) and writes a 剪映 draft folder (`draft_content.json` + `draft_info.json` + `draft_meta_info.json`) that the desktop app can open and the user can keep editing: video clips on the main track, the narration and BGM as their own audio tracks, the recap lines as a subtitle track, and the gap-fill ducking carried as native volume keyframes. The public entrypoints stay small (`us`, `build_draft`, `export_timeline_to_jianying`, CLI). Internally the exporter is split into schema/templates, a thin normalized build context, material/segment builders, track layout metadata, and a safe writer/bundler. This mirrors the useful schema boundaries from duo-video while ffmpeg remains the canonical renderer and JianYing export remains an optional sidecar. Schema and the draft skeleton are informed by the open-source pyJianYingDraft (© GuanYixuan, Apache-2.0) and capcut-mate (© Hommy, Apache-2.0). JSON protocol templates pinned from duo-video are vendored under its MIT license; builders are implemented locally and no upstream executable code, resource package, adapter binary, or credential is included. See ACKNOWLEDGEMENTS / 致谢. """ import json import os import subprocess import tempfile import uuid from jianying.builders import build_timeline_track as _build_timeline_track from jianying.model import DraftBuildContext as _DraftBuildContext from jianying.schema import draft_content_skeleton as _draft_content_skeleton from jianying.schema import meta_info as _meta_info from jianying.schema import us from jianying.timeline_contract import normalize_timeline as _normalize_timeline from jianying.writer import write_draft as _write_draft __all__ = ["us", "build_draft", "export_timeline_to_jianying", "main"] def _default_id(): return str(uuid.uuid4()).upper() def _probe_media(path): """Return (duration_us, width, height) via ffprobe. A probe failure raises: a draft that silently carries a zero or guessed media duration is worse than a failed optional export. Still images legitimately report no duration (0). """ result = subprocess.run( [ "ffprobe", "-v", "error", "-of", "json", "-show_entries", "format=duration:stream=width,height,codec_type", str(path), ], capture_output=True, text=True, timeout=30, ) if result.returncode: raise RuntimeError(f"ffprobe failed for JianYing media {path}: {result.stderr.strip()}") data = json.loads(result.stdout) duration = data["format"].get("duration") width = height = 0 for stream in data.get("streams", []): if stream.get("codec_type") == "video": width, height = int(stream["width"]), int(stream["height"]) break return (us(float(duration)) if duration is not None else 0), width, height def build_draft(timeline, new_id=None, probe=None): """Build the 剪映 draft_content dict and companion meta from a timeline.""" return _build_normalized_draft(_normalize_timeline(timeline), new_id, probe) def _build_normalized_draft(timeline, new_id=None, probe=None): """`timeline` has already passed jianying.timeline_contract.normalize_timeline.""" new_id = new_id or _default_id probe = probe or _probe_media ctx = _DraftBuildContext.from_timeline(timeline, new_id, probe) for timeline_track in timeline["tracks"]: _build_timeline_track(ctx, timeline_track) ctx.finalize_tracks() draft_id = new_id() content = _draft_content_skeleton( draft_id, ctx.width, ctx.height, ctx.fps, ctx.total_us, ctx.materials, ctx.tracks, ) meta = _meta_info(draft_id, ctx.total_us) return content, meta, ctx.notes def _generate_reversed_media(source_path, output_path): commands = [ [ "ffmpeg", "-y", "-v", "error", "-i", source_path, "-vf", "reverse", "-af", "areverse", output_path, ], [ "ffmpeg", "-y", "-v", "error", "-i", source_path, "-vf", "reverse", "-an", output_path, ], ] errors = [] for command in commands: result = subprocess.run(command, capture_output=True, text=True) if result.returncode == 0 and os.path.isfile(output_path): return errors.append( (result.stderr or result.stdout or "unknown ffmpeg error").strip() ) raise RuntimeError( f"failed to reverse JianYing source {source_path}: {'; '.join(errors)}" ) def _prepare_reverse_sources(clips, temporary_dir): """Generate a reversed copy for each clip and record it as the clip's reverse_path.""" generated = [] for clip in clips: source_path = clip["source_path"] if not os.path.isfile(source_path): raise ValueError(f"reverse source does not exist: {source_path}") output_path = os.path.join( temporary_dir, f"reversed-{uuid.uuid4().hex}.mp4" ) _generate_reversed_media(source_path, output_path) clip["reverse_path"] = output_path generated.append(source_path) return generated def export_timeline_to_jianying( timeline, out_dir, draft_name="recap", new_id=None, probe=None, bundle_media=True ): """Write a 剪映 draft folder under out_dir/draft_name. Returns (folder, notes). Referenced media is bundled by default so the draft is self-contained and portable. Pass bundle_media=False only when external absolute paths are intentionally required. """ # Validate once at the boundary; everything below trusts the normalized copy. timeline = _normalize_timeline(timeline) reverse_clips = [ clip for track in timeline["tracks"] if track["kind"] == "video" for clip in track["clips"] if clip.get("reverse") and "reverse_path" not in clip ] if reverse_clips and not bundle_media: raise ValueError("automatic reverse generation requires media bundling") if not reverse_clips: content, meta, notes = _build_normalized_draft(timeline, new_id=new_id, probe=probe) return _write_draft( content, meta, notes, out_dir, draft_name, bundle_media_enabled=bundle_media, ) os.makedirs(out_dir, exist_ok=True) with tempfile.TemporaryDirectory( prefix="jianying-reverse-", dir=out_dir ) as temporary_dir: generated = _prepare_reverse_sources(reverse_clips, temporary_dir) content, meta, notes = _build_normalized_draft(timeline, new_id=new_id, probe=probe) notes.extend(f"已生成倒放素材: {source}" for source in generated) return _write_draft( content, meta, notes, out_dir, draft_name, bundle_media_enabled=True, ) def main(): import argparse ap = argparse.ArgumentParser( description="Export a timeline.json to a 剪映/JianYing draft folder." ) ap.add_argument("timeline", help="path to timeline.json") ap.add_argument( "--out-dir", required=True, help="parent dir to create the draft folder in" ) ap.add_argument("--name", default="recap", help="draft folder name") bundle_group = ap.add_mutually_exclusive_group() bundle_group.add_argument( "--bundle-media", dest="bundle_media", action="store_true", help="copy referenced media into the draft folder (default)", ) bundle_group.add_argument( "--no-bundle-media", dest="bundle_media", action="store_false", help="keep external media paths instead of making a portable draft", ) ap.set_defaults(bundle_media=True) args = ap.parse_args() with open(args.timeline, encoding="utf-8") as f: timeline = json.load(f) draft_dir, notes = export_timeline_to_jianying( timeline, args.out_dir, args.name, bundle_media=args.bundle_media, ) for note in notes: print(f" 注意: {note}") print( json.dumps({"status": "exported", "draft_dir": draft_dir}, ensure_ascii=False) ) if __name__ == "__main__": main() -
lib.py 14 KB
"""Self-contained config + utilities for this skill (no cross-skill imports).""" import functools import math import os import shutil import subprocess import tempfile # ── 配置 ────────────────────────────────────────────────────────────── _EXISTING_CONFIG_REF = globals().get("CONFIG") def env_int(name, default, *, minimum=None): """Read an integer env var; an unset/empty value yields `default`, a malformed one is an error.""" raw = os.environ.get(name, "") if raw == "": return default try: value = int(raw) except ValueError as exc: raise ValueError(f"环境变量 {name}={raw!r} 不是整数") from exc return value if minimum is None else max(minimum, value) def env_bool(name, default=False): """Read common boolean env var forms.""" raw = os.environ.get(name, "") if raw == "": return default return raw.strip().lower() in {"1", "true", "yes", "y", "on"} def env_float(name, default, *, minimum=None): """Read a float env var; an unset/empty value yields `default`, a malformed one is an error.""" raw = os.environ.get(name, "") if raw == "": return default try: value = float(raw) except ValueError as exc: raise ValueError(f"环境变量 {name}={raw!r} 不是数字") from exc if not math.isfinite(value): raise ValueError(f"环境变量 {name}={raw!r} 必须是有限数") return value if minimum is None else max(minimum, value) # Cross-language source: when the original audio is in a language the narration is NOT in # (e.g. a Japanese drama recapped in Chinese), the original speech bleeding under the narration # is just noise the viewer can't parse — it reads as 怪音. In that mode the original is ducked to # near-silent UNDER narration; it still plays full-volume in the original-audio gap blocks, where # a single language is fine. Explicit SPEECH_DUCKING_VOLUME / ZONE_DUCKING_VOLUME still override. _foreign_source_audio = env_bool("FOREIGN_SOURCE_AUDIO", False) _foreign_under_narration_volume = 0.05 # original volume under narration when source audio is foreign CONFIG = { "fade_ms": env_int("FADE_MS", 120, minimum=0), # 每段 TTS 淡入淡出(ms);过大会让紧凑的句子一顿一顿,120ms 防爆音又不发闷 "ducking_mode": "fixed", # fixed | sidechaincompress | none "ducking_threshold": 0.15, "ducking_ratio": 3, "ducking_attack": 10, "ducking_release": 300, "ducking_level_sc": 2.0, "ducking_makeup": 1.2, "ducking_narr_weight": 1.5, "ducking_orig_volume": env_float("DUCKING_ORIG_VOLUME", 0.3, minimum=0.0), # 解说时原声基准音量 # Derived report of the FOREIGN_SOURCE_AUDIO knob this skill implements: it selects the # ducking volumes below. Declared so callers can see which policy is in effect. "foreign_source_audio": _foreign_source_audio, "zone_ducking_volume": env_float("ZONE_DUCKING_VOLUME", _foreign_under_narration_volume if _foreign_source_audio else 0.12, minimum=0.0), # 解说时原声压低到的音量 "idle_orig_volume": env_float("IDLE_ORIG_VOLUME", 1.0, minimum=0.0), # 解说块之间的"原声块"音量:默认满音量(1.0),让精彩原声整段放出来,不被压低(用户要求解说成块、原声也成块) "duck_fade_seconds": env_float("DUCK_FADE_SECONDS", 0.3, minimum=0.0), # 解说块/原声块切换的淡入淡出(秒),略放宽到 0.3 让满音量↔压低的过渡更顺 "duck_bridge_seconds": env_float("DUCK_BRIDGE_SECONDS", 1.5, minimum=0.0), # 仅把间隔小于此值的相邻解说窗口并成一段压低;超过则视为作者特意留的"原声块",原声放回满音量。默认 1.5s:解说块内部连续压低,块与块之间的留白放出满音量原声。该值只控制短间隔合并,不设定旁白/原声配额。调大→更连续铺底、原声块更少;调小→更碎 "bgm_path": os.environ.get("BGM_PATH", "").strip(), # 背景音乐文件(可选),留空则不加 BGM "source_video": os.environ.get("SOURCE_VIDEO", "").strip(), # 剪辑模式下的原始视频(可选),用于时间线/剪映导出引用原片片段 "source_video_explicit": False, # 仅 assemble.py --source-video 显式传入时为 True;环境变量 SOURCE_VIDEO 不算显式 "export_jianying": env_bool("EXPORT_JIANYING", False), # 渲染后可选导出剪映草稿(默认关;与核心解耦) "jianying_draft_dir": os.environ.get("JIANYING_DRAFT_DIR", "").strip(), # 剪映草稿输出父目录(留空=work_dir) "jianying_bundle_media": env_bool("JIANYING_BUNDLE_MEDIA", True), # 默认开:macOS 剪映沙箱读不到外部路径,须把素材拷进草稿目录 "bgm_volume": env_float("BGM_VOLUME", 0.18, minimum=0.0), # BGM 铺底音量 "bgm_ducking_volume": env_float("BGM_DUCKING_VOLUME", 0.10, minimum=0.0), # 旁白时 BGM 压低到的音量 "narration_speed": env_float("NARRATION_SPEED", 1.15, minimum=0.5), # 解说整体提速(atempo),默认回到可懂区间;长片可设 1.0 "narration_cumulative_tempo_max": env_float("NARRATION_CUMULATIVE_TEMPO_MAX", 1.35, minimum=1.0), # TTS rate × 全局 atempo × 段内 atempo 的累计上限 "narration_cumulative_tempo_hard_max": env_float("NARRATION_CUMULATIVE_TEMPO_HARD_MAX", 1.40, minimum=1.0), # QC/阻断硬上限 "tts_segment_tempo_max": env_float("TTS_SEGMENT_TEMPO_MAX", 1.20, minimum=1.0), # 兼容旧段内 atempo 上限;实际会被累计预算收紧 "mask_source_subtitles": env_bool("MASK_SOURCE_SUBTITLES", False), # 遮挡原片烧录字幕;必须配合显式 SOURCE_SUBTITLE_MASK_POLICY "source_subtitle_mask_policy_declared": bool(os.environ.get("SOURCE_SUBTITLE_MASK_POLICY", "").strip()), "source_subtitle_mask_policy": ( os.environ.get("SOURCE_SUBTITLE_MASK_POLICY", "").strip().lower() or "off" ), # off | opt_in | safe | forced;MASK_SOURCE_SUBTITLES alone is legacy implicit and QC-blocking "source_subtitle_mask_ratio": env_float("SOURCE_SUBTITLE_MASK_RATIO", 0.14, minimum=0.0), # 底部遮挡比例 "source_subtitle_mask_timing": os.environ.get("SOURCE_SUBTITLE_MASK_TIMING", "narration").strip().lower(), # all | narration;增强版默认仅解说时遮罩 "subtitle_mask_opacity": min(1.0, env_float("SUBTITLE_MASK_OPACITY", 0.6, minimum=0.0)), # 0=透明,1=全黑;增强版默认半透明 "subtitle_mask_padding": env_int("SUBTITLE_MASK_PADDING", 4, minimum=0), "subtitle_y_top": env_int("SUBTITLE_Y_TOP", -1, minimum=-1), # 自动旋转后的显示画布坐标;top/bot 同时有效时贴合原字幕带 "subtitle_y_bot": env_int("SUBTITLE_Y_BOT", -1, minimum=-1), "narration_delay_seconds": env_float("NARRATION_DELAY_SECONDS", 0.0, minimum=0.0), # 默认严格采用 Agent 写入的 start;旧项目可显式恢复延迟 "narration_tighten": env_bool("NARRATION_TIGHTEN", True), # 段落内把句子紧贴上一句实际收尾播放,句间间隔稳定≤tight_pause,杜绝"一句解说一段空白"的卡顿 "narration_run_gap_seconds": env_float("NARRATION_RUN_GAP_SECONDS", 1.6, minimum=0.0), # 作者留白超过此值=新段落(让精彩原声透出);小于则视为同一连续段落 "narration_tight_pause_seconds": env_float("NARRATION_TIGHT_PAUSE_SECONDS", 0.35, minimum=0.0), # 段落内句间固定间隔(秒) "narration_max_pull_seconds": env_float("NARRATION_MAX_PULL_SECONDS", 1.2, minimum=0.0), # 收紧时一句最多比作者标注提前的秒数(漂移上限,越小越贴画面) "narration_tail_pad_seconds": 0.1, # 解说尾部最少留白;短 slot 会自动压低 delay 避免截断 "quiet_overlap_min_ratio": 0.8, # 解说段至少多少比例落在安静窗口内才标记为非对白重叠 "speech_ducking_volume": env_float("SPEECH_DUCKING_VOLUME", _foreign_under_narration_volume if _foreign_source_audio else 0.2, minimum=0.0), # 解说与对白重叠时原声音量 "burn_subtitles": env_bool("BURN_SUBTITLES", True), # 烧录解说字幕(默认开;遮挡原字幕后需自带字幕,否则字幕区空白) "subtitle_original_in_gaps": env_bool("SUBTITLE_ORIGINAL_IN_GAPS", True), # 原声留白处补烧原声台词字幕(来自 ASR) "force_video_reencode": env_bool("FORCE_VIDEO_REENCODE", False), # 组装时重编码视频,修复部分容器时间戳问题 # 成片压制(仅在重编码时生效:烧字幕/遮罩/缩放/FORCE_VIDEO_REENCODE 任一触发重编码)。 "output_crf": env_int("OUTPUT_CRF", 18, minimum=0), # x264 CRF;越大文件越小、画质越低(18≈视觉无损,23~26 体积更小) "output_preset": os.environ.get("OUTPUT_PRESET", "veryfast"), # x264 preset;slow/slower 同 CRF 下体积更小但更慢 "output_max_height": env_int("OUTPUT_MAX_HEIGHT", 0, minimum=0), # >0 时把成片高度上限缩到该值(保持宽高比、偶数宽);0=不缩放 # 成片末端整体响度归一(默认混音偏轻,归一后更接近常见短视频响度;样片约 -11.9,默认取更安全的 -14) "final_loudnorm": env_bool("FINAL_LOUDNORM", True), # 组装末端做一次整体响度归一 "target_lufs": env_float("TARGET_LUFS", -14.0), # 目标综合响度 (LUFS) "target_true_peak": env_float("TARGET_TRUE_PEAK", -1.0), # 目标真峰值 (dBTP) "target_lra": env_float("TARGET_LRA", 11.0), # 目标响度范围 (LU) "final_limiter_peak": env_float("FINAL_LIMITER_PEAK", 0.98, minimum=0.1), # loudnorm 后峰值保护 limiter "subtitle_font_name": os.environ.get("SUBTITLE_FONT_NAME", "Arial"), # 可选字体文件:ASS 烧录经 fontsdir 加载,画面文字经 drawtext fontfile 使用;family 名仍由 SUBTITLE_FONT_NAME 指定 "subtitle_font_file": os.environ.get("SUBTITLE_FONT_FILE", "").strip(), "subtitle_font_size": env_int("SUBTITLE_FONT_SIZE", 42, minimum=8), "subtitle_primary_color": os.environ.get("SUBTITLE_PRIMARY_COLOR", "&H00FFFFFF"), "subtitle_outline_color": os.environ.get("SUBTITLE_OUTLINE_COLOR", "&H00000000"), "subtitle_outline": env_float("SUBTITLE_OUTLINE", 2.0, minimum=0.0), "subtitle_shadow": env_float("SUBTITLE_SHADOW", 1.0, minimum=0.0), "subtitle_margin_v": env_int("SUBTITLE_MARGIN_V", 48, minimum=0), "subtitle_margin_l": env_int("SUBTITLE_MARGIN_L", 40, minimum=0), "subtitle_margin_r": env_int("SUBTITLE_MARGIN_R", 40, minimum=0), "subtitle_alignment": env_int("SUBTITLE_ALIGNMENT", 2, minimum=1), "subtitle_max_chars": env_int("SUBTITLE_MAX_CHARS", 20, minimum=6), "subtitle_max_lines": env_int("SUBTITLE_MAX_LINES", 2, minimum=1), "subtitle_play_res_x": env_int("SUBTITLE_PLAY_RES_X", 1280, minimum=1), "subtitle_play_res_y": env_int("SUBTITLE_PLAY_RES_Y", 720, minimum=1), } if isinstance(_EXISTING_CONFIG_REF, dict): _EXISTING_CONFIG_REF.clear() _EXISTING_CONFIG_REF.update(CONFIG) CONFIG = _EXISTING_CONFIG_REF def narration_tempo_budget(tts_rate_offset=0.0): """Return the canonical tempo budget shared by voiceover and assemble. `effective_tempo` is the user-perceived cumulative compression: TTS rate × global narration atempo × per-segment atempo. The segment atempo cap is therefore tightened by the configured global speed and TTS rate offset; callers must fail/shorten instead of time-trimming speech when the needed ratio exceeds `segment_tempo_max`. """ global_speed = CONFIG["narration_speed"] rate_factor = max(0.01, 1.0 + float(tts_rate_offset)) cumulative_max = CONFIG["narration_cumulative_tempo_max"] hard_max = max(cumulative_max, CONFIG["narration_cumulative_tempo_hard_max"]) segment_tempo_max = max(1.0, min( CONFIG["tts_segment_tempo_max"], cumulative_max / (global_speed * rate_factor) )) return { "global_narration_speed": global_speed, "tts_rate_factor": rate_factor, "cumulative_tempo_max": cumulative_max, "cumulative_tempo_hard_max": hard_max, "segment_tempo_max": segment_tempo_max, "max_raw_duration_factor": global_speed * segment_tempo_max, } def log(msg): print(f"[video-recap] {msg}", flush=True) def run_cmd(cmd, **kwargs): """运行命令,返回 CompletedProcess""" display = " ".join( text if len(text) <= 240 else text[:237] + "..." for text in map(str, cmd) ) log(f"运行: {display}") return subprocess.run(cmd, capture_output=True, text=True, **kwargs) # ffmpeg 7 added `-/option path` to read any option's value from a file; ffmpeg 9 removed the # older `-filter_complex_script` / `-filter_script` spellings, which are all ffmpeg <= 6 knows. _LEGACY_FILTER_FILE_OPTIONS = { "filter_complex": "-filter_complex_script", "filter:v:0": "-filter_script:v:0", } @functools.lru_cache(maxsize=None) def _ffmpeg_reads_option_files(): """Whether the ffmpeg on PATH accepts `-/option path` (asked once per process).""" if shutil.which("ffmpeg") is None: return False with tempfile.TemporaryDirectory() as tmp: graph = os.path.join(tmp, "probe_filter.txt") with open(graph, "w", encoding="utf-8") as fh: fh.write("null") result = subprocess.run(["ffmpeg", "-hide_banner", "-/filter_complex", graph], stdin=subprocess.DEVNULL, capture_output=True, text=True, timeout=20) return "Unrecognized option" not in result.stderr def filter_file_args(option, path): """ffmpeg arguments that load `option`'s filtergraph from `path`, spelled for this ffmpeg.""" if _ffmpeg_reads_option_files(): return [f"-/{option}", str(path)] return [_LEGACY_FILTER_FILE_OPTIONS[option], str(path)] def get_video_duration(video_path): """获取视频时长(秒)""" cmd = ["ffprobe", "-v", "quiet", "-show_entries", "format=duration", "-of", "csv=p=0", str(video_path)] result = run_cmd(cmd) if result.returncode != 0: raise RuntimeError(f"ffprobe 无法读取时长 {video_path}: {result.stderr}") return float(result.stdout.strip()) -
media.py 8.2 KB
"""Media probing and source-clip provenance for video-assemble.""" import json import os from pathlib import Path from artifacts import _explicit_source_video from lib import log, run_cmd def _load_cut_timeline_plan(work_dir): """The cut plan, preferring clip_plan_validated.json unless the raw plan is newer; None in full mode.""" raw_plan_path = Path(work_dir) / "clip_plan.json" validated_plan_path = Path(work_dir) / "clip_plan_validated.json" if not validated_plan_path.exists(): return json.loads(raw_plan_path.read_text(encoding="utf-8")) if raw_plan_path.exists() else None if not raw_plan_path.exists(): return json.loads(validated_plan_path.read_text(encoding="utf-8")) if validated_plan_path.stat().st_mtime_ns >= raw_plan_path.stat().st_mtime_ns: return json.loads(validated_plan_path.read_text(encoding="utf-8")) return json.loads(raw_plan_path.read_text(encoding="utf-8")) def _plan_clip_spans(work_dir): """Cut-mode clip spans [{source_start, source_end, output_start, output_end, entry}], or None. clip_plan.json is either a bare list or {"clips": [...]}; a clip names its source range as source_start/source_end or start/end. Clips without explicit output_start/output_end are laid out back to back on the output timeline. """ plan = _load_cut_timeline_plan(work_dir) if plan is None: return None entries = plan["clips"] if isinstance(plan, dict) else plan spans, cursor = [], 0.0 for entry in entries: ss = float(entry.get("source_start", entry.get("start"))) se = float(entry.get("source_end", entry.get("end"))) if "output_start" in entry: out_s, out_e = float(entry["output_start"]), float(entry["output_end"]) cursor = max(cursor, out_e) else: out_s, out_e = cursor, cursor + (se - ss) cursor = out_e spans.append({ "source_start": ss, "source_end": se, "output_start": out_s, "output_end": out_e, "entry": entry, }) return spans def _ratio_to_float(value, default=1.0): """Parse an ffprobe ratio ("4:3", "16/9" or a bare number); unknown ratios yield `default`.""" value = value.strip() if value in {"", "0:1", "0/1", "N/A"}: return default if ":" in value: num, den = value.split(":", 1) elif "/" in value: num, den = value.split("/", 1) else: return float(value) return float(num) / float(den) if float(den) else default def _fps_from_rate(value, default=30.0): """Parse an ffprobe frame rate ("30000/1001" or a bare number); a 0/0 rate yields `default`.""" if "/" in value: num, den = value.split("/", 1) return round(float(num) / float(den), 3) if float(den) else default return round(float(value), 3) def _stream_rotation(stream): """Extract rotation from tags or side_data_list in ffprobe JSON.""" for source in (stream.get("tags", {}).get("rotate"), stream.get("rotation")): if source not in (None, ""): return int(round(float(source))) % 360 for item in stream.get("side_data_list", []): if item.get("rotation") not in (None, ""): return int(round(float(item["rotation"]))) % 360 return 0 def _canvas_from_stream(stream): storage_w = stream["width"] storage_h = stream["height"] fps = _fps_from_rate(stream["r_frame_rate"]) sar_text = stream.get("sample_aspect_ratio", "1:1") dar_text = stream.get("display_aspect_ratio", "") sar = _ratio_to_float(sar_text, 1.0) rotation = _stream_rotation(stream) display_w = max(1, int(round(storage_w * sar))) display_h = max(1, storage_h) if dar_text and dar_text not in {"0:1", "N/A"}: dar = _ratio_to_float(dar_text, 0.0) # ffprobe sources are not consistent: some report DAR before rotation # (landscape value > 1 for a 90° stream), while some containers report the # already-rotated portrait DAR (< 1). Only apply DAR before swapping when it # describes the stored orientation. if dar > 0 and not (rotation in {90, 270} and dar < 1.0): # Preserve height and adjust width. This keeps legacy square-pixel landscape # byte-identical while honoring non-square pixel DAR metadata. display_w = max(1, int(round(display_h * dar))) if rotation in {90, 270}: display_w, display_h = display_h, display_w return { "width": display_w, "height": display_h, "fps": fps, "storage_width": storage_w, "storage_height": storage_h, "rotation": rotation, "sample_aspect_ratio": sar_text, "display_aspect_ratio": dar_text or f"{display_w}:{display_h}", } def _probe_canvas(video_path): """Return rotation/SAR/DAR-aware canvas facts for a video. ``width``/``height`` are the display canvas used by subtitle/overlay geometry. For legacy square-pixel landscape sources, these remain the raw storage dimensions. """ res = run_cmd([ "ffprobe", "-v", "error", "-select_streams", "v:0", "-show_entries", "stream=width,height,r_frame_rate,avg_frame_rate,sample_aspect_ratio,display_aspect_ratio:stream_tags=rotate:stream_side_data=rotation", "-of", "json", str(video_path), ]) if res.returncode != 0: raise RuntimeError(f"ffprobe 无法读取视频流 {video_path}: {res.stderr}") streams = json.loads(res.stdout)["streams"] if not streams: raise RuntimeError(f"{video_path} 没有视频流") return _canvas_from_stream(streams[0]) def _has_audio_stream(video_path): """Return True when the input has an audio stream usable as [0:a].""" result = run_cmd([ "ffprobe", "-v", "error", "-select_streams", "a:0", "-show_entries", "stream=index", "-of", "csv=p=0", str(video_path), ]) return result.returncode == 0 and bool(result.stdout.strip()) def _build_video_clips(input_video, work_dir, duration_s): """Video-track clips for the timeline. In cut mode each plan entry becomes a clip referencing the ORIGINAL source range. Multi-source validated plans carry per-clip source_path and do not require an explicit ambient --source-video. Without any declared source (full mode, or cut mode rendered without --source-video) the rendered input is one clip. """ explicit_source_video = _explicit_source_video() spans = _plan_clip_spans(work_dir) multi_source = spans is not None and any(span["entry"].get("source_path") for span in spans) if spans is None or not (explicit_source_video or multi_source): return [{"source_path": str(input_video), "source_start": 0.0, "source_end": float(duration_s), "timeline_start": 0.0, "timeline_end": float(duration_s)}] clips = [] for span in spans: entry = span["entry"] source_path = entry.get("source_path") or explicit_source_video timeline_start, timeline_end = span["output_start"], span["output_end"] if not source_path or not os.path.exists(source_path): # Degrade ONLY this clip — point it at the rendered cut for its own output # window — and keep real provenance for every present source, instead of # collapsing the whole multi-source timeline. log(f" 时间线: source_path 不存在,该片段降级为剪后成片片段: {source_path or '(unset)'}") clips.append({"source_id": entry.get("source_id"), "source_path": str(input_video), "source_start": timeline_start, "source_end": timeline_end, "timeline_start": timeline_start, "timeline_end": timeline_end, "provenance_degraded": True, "provenance_reason": f"missing_source_path:{source_path or 'unset'}"}) continue clips.append({"source_id": entry.get("source_id"), "source_path": source_path, "source_start": span["source_start"], "source_end": span["source_end"], "timeline_start": timeline_start, "timeline_end": timeline_end}) return clips -
narration_audio.py 18.9 KB
"""Narration tempo fitting and sample-accurate timeline WAV placement.""" import os import wave from pathlib import Path from lib import CONFIG, get_video_duration, log, narration_tempo_budget, run_cmd def _apply_narration_speed( tts_segments, work_dir, *, tempo_policy=None, ): """Globally speed up narration audio via atempo (CONFIG['narration_speed']). MiMo TTS reads a touch slowly for short-form recaps; a 1.1-1.2x bump makes it snappier without the chipmunk effect. Rewrites each segment's audio_path/duration to the sped copy so the rest of assembly is unchanged. No-op at speed 1.0. """ speed = tempo_policy["global_atempo"] if tempo_policy else CONFIG["narration_speed"] if abs(speed - 1.0) <= 1e-3: return done = 0 for seg in tts_segments: src = seg["audio_path"] if not os.path.exists(src): continue # reported as a skipped segment during placement out = str(Path(work_dir) / f"_spd_{seg['index']}.wav") res = run_cmd(["ffmpeg", "-y", "-i", src, "-filter:a", f"atempo={speed:.3f}", "-ar", "44100", "-ac", "1", "-acodec", "pcm_s16le", out]) if res.returncode != 0: raise RuntimeError(f"解说提速失败 {src}: {res.stderr}") seg["audio_path"] = out seg["narration_conversion_path"] = out seg["audio_duration"] = get_video_duration(out) done += 1 log(f"解说整体提速: atempo={speed:.2f} ({done} 段)") def _adjust_tts_speed( audio_path, target_duration, tts_rate_offset=0.0, *, tempo_policy=None, ): """Fit overlong TTS with bounded atempo; never time-trim speech in assemble. Assemble has no word/sentence timestamps, so if bounded atempo cannot make the audio fit, it returns `fit_status=no_safe_fit` and leaves the original audio untouched for QC to block instead of guessing a spoken_text truncation. """ audio_path = Path(audio_path) current_dur = get_video_duration(audio_path) budget = narration_tempo_budget(tts_rate_offset) if tempo_policy: budget.update({ "global_narration_speed": tempo_policy["global_atempo"], "tts_rate_factor": 1.0, "segment_tempo_max": tempo_policy["segment_tempo_max"], "cumulative_tempo_max": tempo_policy["cumulative_tempo_max"], "cumulative_tempo_hard_max": tempo_policy["cumulative_tempo_hard_max"], }) meta = { "fit_status": "fits", "blocking": False, "tempo_factor": 1.0, "segment_tempo_factor": 1.0, "truncated": False, "truncate_reason": "none", "tts_rate_offset": float(tts_rate_offset), "audio_duration": current_dur, "placed_audio_duration": current_dur, "global_narration_speed": budget["global_narration_speed"], "effective_tempo": budget["global_narration_speed"] * budget["tts_rate_factor"], "cumulative_tempo_max": budget["cumulative_tempo_max"], "cumulative_tempo_hard_max": budget["cumulative_tempo_hard_max"], } if current_dur <= target_duration: return (str(audio_path), current_dur, meta) if tempo_policy and not tempo_policy["bounded_segment_fit"]: meta.update({ "fit_status": "no_safe_fit", "blocking": True, "truncate_reason": "no_safe_boundary", "placed_audio_duration": 0.0, "needed_tempo_factor": current_dur / target_duration, }) return (str(audio_path), current_dur, meta) ratio = current_dur / target_duration effective_max = budget["segment_tempo_max"] if ratio > effective_max: meta.update({ "fit_status": "no_safe_fit", "blocking": True, "truncate_reason": "no_safe_boundary", "placed_audio_duration": 0.0, "needed_tempo_factor": ratio, }) log( f" TTS 无安全放置: {current_dur:.1f}s 需 x{ratio:.2f}," f"超过段内预算 x{effective_max:.2f}(assemble 不按时间硬切)" ) return (str(audio_path), current_dur, meta) # 温和加速。给 atempo/容器时长舍入留出 0.2% 安全余量;宁可极轻微 # 多加速,也不能在写入时间线时裁掉最后一个音节。 tempo = min(ratio * 1.002, effective_max) adjusted_path = audio_path.with_name(f"{audio_path.stem}_adj{audio_path.suffix}") cmd = ["ffmpeg", "-y", "-i", str(audio_path), "-filter:a", f"atempo={tempo:.6f}", "-ar", "44100", "-ac", "1", str(adjusted_path)] result = run_cmd(cmd) if result.returncode != 0: raise RuntimeError(f"TTS 加速失败 {audio_path}: {result.stderr}") new_dur = get_video_duration(adjusted_path) if new_dur > target_duration + (1.0 / 44100.0): adjusted_path.unlink(missing_ok=True) meta.update({ "fit_status": "no_safe_fit", "blocking": True, "truncate_reason": "no_safe_boundary", "placed_audio_duration": 0.0, "needed_tempo_factor": new_dur / target_duration, }) log( f" TTS 加速后仍超出安全窗口 {new_dur - target_duration:.3f}s;" "禁止裁尾,交由 Agent 缩短/移动文本" ) return (str(audio_path), current_dur, meta) meta.update({ "fit_status": "tempo_adjusted", "tempo_factor": tempo, "segment_tempo_factor": tempo, "placed_audio_duration": new_dur, "effective_tempo": budget["global_narration_speed"] * budget["tts_rate_factor"] * tempo, }) log(f" TTS 温和加速: {current_dur:.1f}s → {new_dur:.1f}s (x{tempo:.2f})") return (str(adjusted_path), new_dur, meta) def _edge_quiet_samples(pcm16_mono, sample_count, *, from_start, threshold=260): """Count near-silent PCM16 samples at one edge of a mono buffer.""" indices = range(sample_count) if from_start else range(sample_count - 1, -1, -1) quiet = 0 for index in indices: offset = index * 2 value = int.from_bytes(pcm16_mono[offset:offset + 2], "little", signed=True) if abs(value) > threshold: break quiet += 1 return quiet def _speech_safe_fade_lengths(pcm16_mono, sample_count, sample_rate, configured_ms): """Limit fades to edge silence so first/last syllables are never attenuated. When TTS has no measurable edge silence, retain only a 5ms anti-click ramp. """ configured = min(int(max(0.0, float(configured_ms)) * sample_rate / 1000), sample_count // 4) if configured <= 0: return 0, 0 anti_click = min(int(0.005 * sample_rate), configured) leading = _edge_quiet_samples(pcm16_mono, sample_count, from_start=True) trailing = _edge_quiet_samples(pcm16_mono, sample_count, from_start=False) return min(configured, max(anti_click, leading)), min(configured, max(anti_click, trailing)) def _unplaced(seg, at, fit_status, reason, *, blocking=False): """Record a segment that contributes no audio: a zero-width window at `at` seconds.""" seg["actual_place_start"] = at seg["actual_place_end"] = at seg["placed_audio_duration"] = 0.0 seg["fit_status"] = fit_status seg["truncate_reason"] = reason seg["blocking"] = blocking def _build_timed_narration( tts_segments, output_wav, video_duration, work_dir, *, tempo_policy=None, ): """将 TTS 片段按时间轴放置到一条与视频等长的音轨上""" sample_rate = 44100 total_samples = int(video_duration * sample_rate) buffer = bytearray(total_samples * 2) last_written_end = 0 # 追踪已写入位置,防止重叠 prev_pause_samples = 0 # 前一段的 pause_after_ms,控制段间间隔 skipped_count = 0 # 因 WAV 缺失或无安全放置而被跳过的段数 placed_count = 0 # 真正写入音频的段数;防止"成功"生成全静音旁白 no_safe_fit_count = 0 # 超预算但不能安全截断;交由 QC/manifest 阻断 prev_authored_end = None # 上一段作者标注的结束时间,用于判断"段落"边界 run_gap = CONFIG["narration_run_gap_seconds"] # 作者留白 > 此值 = 新段落 tighten = CONFIG["narration_tighten"] tight_pause_samples = int(CONFIG["narration_tight_pause_seconds"] * sample_rate) # 漂移上限:收紧时一句最多比作者标注的时间提前 max_pull 秒,避免整段解说被全部压到前面、与画面脱节 max_pull_samples = int(CONFIG["narration_max_pull_seconds"] * sample_rate) configured_delay = CONFIG["narration_delay_seconds"] tail_pad = CONFIG["narration_tail_pad_seconds"] for seg in tts_segments: wav_path = seg["audio_path"] pause_samples = int(seg["pause_after_ms"] * sample_rate / 1000) # 段落收紧:同一段落内(与上一句作者留白 <= run_gap)把这一句紧贴上一句的实际收尾播放, # 句间间隔固定为 tight_pause,不受 slot 内居中延迟 / TTS 时长波动影响。段落之间(作者特意留 # 的大留白,让精彩原声透出)才放回原声。这样句间间隔稳定、不会出现"一句解说一段空白"。 cur_authored_start = float(seg["start"]) is_run_start = (placed_count == 0 or prev_authored_end is None or cur_authored_start - prev_authored_end > run_gap) prev_authored_end = float(seg["end"]) if not os.path.exists(wav_path): _unplaced(seg, seg["start"], "skipped", "missing_wav") prev_pause_samples = pause_samples skipped_count += 1 continue original_wav_path = wav_path try: with wave.open(wav_path, "rb") as wf_check: needs_resample = ( wf_check.getnchannels(), wf_check.getsampwidth(), wf_check.getframerate() ) != (1, 2, sample_rate) except (wave.Error, EOFError): # Valid post-processed WAV may use IEEE float, which Python's wave # reader does not support. FFmpeg performs the explicit PCM conversion. needs_resample = True tts_rate_offset = seg["tts_rate_offset"] tts_dur = seg["audio_duration"] slot_duration = max(0.0, float(seg["end"]) - float(seg["start"])) max_delay = max(0.0, slot_duration - tts_dur - tail_pad) narration_delay = min(configured_delay, max_delay) start_sample = int((seg["start"] + narration_delay) * sample_rate) end_boundary = int(min(seg["end"], video_duration) * sample_rate) # 段间间隔:使用前一段的 pause_after_ms(来自 narration.json) min_start_with_pause = last_written_end + prev_pause_samples if tighten and not is_run_start: # 段落内:紧贴上一句的实际收尾播放,句间间隔固定为 tight_pause(不被 slot 内居中延迟撑大), # 但不早于"作者标注起始 - max_pull",防止整段被压到前面与画面脱节。 drift_floor = int(cur_authored_start * sample_rate) - max_pull_samples actual_start = max(last_written_end + tight_pause_samples, drift_floor) else: # 段落起点(或关闭收紧):尊重作者标注的起始 + 入场延迟,让画面/原声先立住 actual_start = max(start_sample, min_start_with_pause) actual_start = min(actual_start, end_boundary) # 不超出 slot 边界 # 根据实际可用空间决定是否加速 available_samples = end_boundary - actual_start available_duration = max(available_samples / sample_rate, 0) if tts_dur > available_duration > 0: if tempo_policy: wav_path, _actual_dur, fit_meta = _adjust_tts_speed( wav_path, available_duration, tts_rate_offset, tempo_policy=tempo_policy, ) else: wav_path, _actual_dur, fit_meta = _adjust_tts_speed( wav_path, available_duration, tts_rate_offset ) seg.update({ "fit_status": fit_meta["fit_status"], "segment_tempo_factor": fit_meta["segment_tempo_factor"], "effective_tempo": fit_meta["effective_tempo"], "global_narration_speed": fit_meta["global_narration_speed"], "blocking": fit_meta["blocking"], }) if fit_meta["fit_status"] == "no_safe_fit": _unplaced(seg, actual_start / sample_rate, "no_safe_fit", fit_meta["truncate_reason"], blocking=True) prev_pause_samples = pause_samples skipped_count += 1 no_safe_fit_count += 1 continue else: budget = narration_tempo_budget(tts_rate_offset) if tempo_policy: budget.update({ "global_narration_speed": tempo_policy["global_atempo"], "tts_rate_factor": 1.0, }) seg.update({ "fit_status": "fits", "segment_tempo_factor": 1.0, "global_narration_speed": budget["global_narration_speed"], "effective_tempo": budget["global_narration_speed"] * budget["tts_rate_factor"], "blocking": False, }) # _adjust_tts_speed 输出固定 44100Hz mono 16bit,若文件被替换则无需 resample if wav_path != original_wav_path: needs_resample = False if needs_resample: tmp_path = str(Path(work_dir) / f"_rs_{seg['index']}.wav") rs_result = run_cmd(["ffmpeg", "-y", "-i", wav_path, "-ar", str(sample_rate), "-ac", "1", "-acodec", "pcm_s16le", tmp_path]) if rs_result.returncode != 0: log(f" 跳过: 重采样失败 {wav_path}: {rs_result.stderr}") _unplaced(seg, seg["start"], "skipped", "resample_failed") prev_pause_samples = pause_samples skipped_count += 1 continue wav_path = tmp_path seg["narration_conversion_path"] = tmp_path with wave.open(wav_path, "rb") as wf: wf_data = bytearray(wf.readframes(wf.getnframes())) # 按场景边界裁剪 audio_samples = len(wf_data) // 2 available = end_boundary - actual_start write_samples = audio_samples if write_samples <= 0 or available <= 0: log(f" 跳过: {seg['start']:.1f}s-{seg['end']:.1f}s (无空间)") _unplaced(seg, seg["start"], "no_safe_fit", "no_room", blocking=True) prev_pause_samples = pause_samples no_safe_fit_count += 1 continue if audio_samples > available: # No tolerance-based trimming: even a few milliseconds may contain a # consonant/vowel release. _adjust_tts_speed must produce a complete file # that fits; otherwise block and ask the Agent to shorten/move the block. over = (audio_samples - available) / sample_rate log(f" TTS 无安全放置: 段 {seg['index']} 超出可用窗口 {over:.3f}s;禁止裁尾,交由 QC 阻断") _unplaced(seg, actual_start / sample_rate, "no_safe_fit", "no_safe_boundary", blocking=True) prev_pause_samples = pause_samples skipped_count += 1 no_safe_fit_count += 1 continue # 重叠检测:跳过与前段重叠的部分(在 fade 之前,避免截断后丢失 fade-in) if actual_start < last_written_end: overlap_ms = (last_written_end - actual_start) * 1000 / sample_rate if last_written_end >= actual_start + write_samples: log(f" 跳过重叠段: {actual_start/sample_rate:.1f}s " f"(与前段重叠 {overlap_ms:.0f}ms)") _unplaced(seg, seg["start"], "no_safe_fit", "no_room", blocking=True) prev_pause_samples = pause_samples no_safe_fit_count += 1 continue actual_start = last_written_end available = end_boundary - actual_start if write_samples > available: log(f" 重叠 {overlap_ms:.0f}ms 后无安全完整窗口,跳过") _unplaced(seg, actual_start / sample_rate, "no_safe_fit", "no_safe_boundary", blocking=True) prev_pause_samples = pause_samples skipped_count += 1 no_safe_fit_count += 1 continue # fade-in / fade-out(在 overlap 裁剪之后应用,确保正确的音频包络) fade_in_len, fade_out_len = _speech_safe_fade_lengths( wf_data, write_samples, sample_rate, CONFIG["fade_ms"] ) for i in range(fade_in_len): gain = i / fade_in_len s = i * 2 sample = int.from_bytes(wf_data[s:s+2], 'little', signed=True) sample = int(sample * gain) wf_data[s:s+2] = sample.to_bytes(2, 'little', signed=True) for i in range(fade_out_len): gain = 1.0 - i / fade_out_len s = (write_samples - 1 - i) * 2 if s < 0: break sample = int.from_bytes(wf_data[s:s+2], 'little', signed=True) sample = int(sample * gain) wf_data[s:s+2] = sample.to_bytes(2, 'little', signed=True) # Persist the exact complete per-beat PCM used by the canonical mix. Editable # exports must reference this file, not the longer pre-fit TTS input; otherwise # their timeline_end silently chops the final word even when ffmpeg is correct. placed_path = Path(work_dir) / f"_placed_{seg['index']:04d}.wav" with wave.open(str(placed_path), "wb") as placed_wav: placed_wav.setnchannels(1) placed_wav.setsampwidth(2) placed_wav.setframerate(sample_rate) placed_wav.writeframes(bytes(wf_data)) seg["placed_audio_path"] = str(placed_path) buffer[actual_start * 2: actual_start * 2 + write_samples * 2] = wf_data seg["actual_place_start"] = actual_start / sample_rate seg["actual_place_end"] = (actual_start + write_samples) / sample_rate seg["placed_audio_duration"] = write_samples / sample_rate last_written_end = actual_start + write_samples prev_pause_samples = pause_samples placed_count += 1 with wave.open(str(output_wav), "wb") as wf: wf.setnchannels(1) wf.setsampwidth(2) wf.setframerate(sample_rate) wf.writeframes(bytes(buffer)) if tts_segments and placed_count == 0 and no_safe_fit_count == 0: output_wav.unlink(missing_ok=True) raise RuntimeError( f"全部 {len(tts_segments)} 段解说均被跳过或未能写入" f"(WAV 缺失或无可用时间;跳过 {skipped_count} 段)," "已中止以避免生成无解说视频" ) log(f"解说音轨: {video_duration:.1f}s, {len(tts_segments)} 段") -
packaging.py 5.5 KB
"""Static packaging layers (frame / header / logo images) burned over the whole recap. The caller writes ``work_dir/packaging_layers.json`` (the orchestrator does so from a bound ``packaging`` template). Each layer is a local image scaled into a rect on a declared canvas. ffmpeg reads the images with ``movie=`` sources, so the render stays a single-input video filter; ``timeline.json`` gets matching image segments for editable export. """ import json from pathlib import Path import lib from artifacts import file_identity from visual_render import _escape_subtitle_filter_path PACKAGING_LAYERS = "packaging_layers.json" _IMAGE_EXTS = {".png", ".jpg", ".jpeg", ".webp"} def load_packaging_layers(work_dir, canvas): """Validated layers for this canvas, or [] when the work_dir declares none.""" path = Path(work_dir) / PACKAGING_LAYERS if not path.exists(): return [] plan = json.loads(path.read_text(encoding="utf-8")) declared = plan["canvas"] if (declared["width"], declared["height"]) != (canvas["width"], canvas["height"]): raise RuntimeError( f"{PACKAGING_LAYERS} 按 {declared['width']}x{declared['height']} 设计," f"成片画布是 {canvas['width']}x{canvas['height']}" ) layers = [] for layer in plan["layers"]: image = Path(layer["path"]).expanduser().resolve() rect = layer["rect"] if not image.is_file() or image.suffix.lower() not in _IMAGE_EXTS: raise RuntimeError(f"包装图层 {layer['name']} 的图片不可用: {image}") if (min(rect["x"], rect["y"]) < 0 or rect["width"] <= 0 or rect["height"] <= 0 or rect["x"] + rect["width"] > canvas["width"] or rect["y"] + rect["height"] > canvas["height"]): raise RuntimeError(f"包装图层 {layer['name']} 超出画布") layers.append({"name": layer["name"], "path": str(image), "rect": dict(rect), "image_size": _image_size(image)}) return layers def _image_size(image): """Pixel size of a layer image; the editor fits by it, the render stretches to the rect.""" res = lib.run_cmd(["ffprobe", "-v", "error", "-select_streams", "v:0", "-show_entries", "stream=width,height", "-of", "json", str(image)]) try: stream = json.loads(res.stdout)["streams"][0] return {"width": int(stream["width"]), "height": int(stream["height"])} except (ValueError, KeyError, IndexError): raise RuntimeError(f"无法读取包装图层尺寸: {image}") from None def compose_video_filter(chain, layers, *, mask_first): """Join the existing filter chain, inserting packaging layers after the source mask. Order: source-subtitle mask → packaging layers → text overlays / burned subtitles → scaling. Without layers this is the plain comma-joined chain. """ if not layers: return ",".join(chain) head, tail = (chain[:1], chain[1:]) if mask_first else ([], list(chain)) graph = [ f"movie=filename='{_escape_subtitle_filter_path(layer['path'])}'," f"scale={layer['rect']['width']}:{layer['rect']['height']},format=rgba[pk{i}]" for i, layer in enumerate(layers) ] current = "in" if head: graph.append(f"[in]{head[0]}[pm]") current = "pm" for i, layer in enumerate(layers): out = "out" if i == len(layers) - 1 and not tail else f"po{i}" graph.append(f"[{current}][pk{i}]overlay=x={layer['rect']['x']}:y={layer['rect']['y']}[{out}]") current = out if tail: graph.append(f"[{current}]{','.join(tail)}[out]") return ";".join(graph) def timeline_image_segments(layers, canvas, duration_s): """Full-length image segments in the timeline's center-origin, Y-up transform. The editor shows an image fitted inside the canvas (keeping its own aspect) at scale 1; the render stretches it to the rect, so x and y scales are derived separately. """ width, height = canvas["width"], canvas["height"] segments = [] for layer in layers: rect, size = layer["rect"], layer["image_size"] fit = min(width / size["width"], height / size["height"]) segments.append({ "source_path": layer["path"], "timeline_start": 0.0, "timeline_end": duration_s, "scale": {"x": round(rect["width"] / (size["width"] * fit), 6), "y": round(rect["height"] / (size["height"] * fit), 6)}, "position": { "x": round((rect["x"] + rect["width"] / 2 - width / 2) / (width / 2), 6), "y": round((height / 2 - rect["y"] - rect["height"] / 2) / (height / 2), 6), }, }) return segments def packaging_settings(work_dir): """What the manifest records: the plan file and each layer image's identity.""" path = Path(work_dir) / PACKAGING_LAYERS if work_dir is not None else None if path is None or not path.exists(): return {"artifact": PACKAGING_LAYERS, "present": False, "layers": []} plan = json.loads(path.read_text(encoding="utf-8")) layers = [] for layer in plan["layers"]: image = Path(layer["path"]).expanduser().resolve() layers.append({"name": layer["name"], "path": str(image), "rect": layer["rect"], **(file_identity(image) if image.is_file() else {})}) return {"artifact": PACKAGING_LAYERS, "present": True, "identity": file_identity(path), "template": plan.get("template"), "layers": layers} -
pair_media.py 10.9 KB
#!/usr/bin/env python3 """Pair explicit picture and adopted audio assets without re-encoding either stream. This is asset pairing, not a claim of source reconstruction, correct editorial selection, speech alignment, or release approval. Existing assemble then consumes paired.mp4; its subtitle binding must be newly computed on that container/a:0. """ import argparse from fractions import Fraction import json from pathlib import Path import subprocess from assemble_constants import SUPPORTED_PICTURE_CODECS from adoption.frozen_audio import probe_audio_packets, verify_adopted_audio from adoption.strict_inputs import ( probe_json, require_declared_path, require_fields, require_integer, run_logged, without_digests, write_json_atomic, ) _probe = probe_json def _asset(value, audio=False): value = without_digests(value, 'asset') require_fields(value, ['path', 'selected_stream'] if audio else ['path'], 'asset') path = require_declared_path(value, 'asset') if audio: return {'path': str(path), 'selected_stream': require_integer(value['selected_stream'], 'Audio selected_stream')} return {'path': str(path)} def _time(value): if value in (None, 'N/A'): raise ValueError('Missing media timing') try: return Fraction(str(value)) except (ValueError, ZeroDivisionError) as exc: raise ValueError('Invalid media timing') from exc def probe_picture(path): """Actual full CFR presentation clock plus codec-order packet sizes and timestamps.""" data = _probe(path, '-select_streams', 'v:0', '-show_streams', '-show_packets', '-show_format', '-show_entries', 'format=format_name:stream=codec_name,profile,level,width,height,pix_fmt,' 'sample_aspect_ratio,field_order,color_range,color_space,color_transfer,' 'color_primaries,chroma_location,time_base,start_pts,duration_ts,avg_frame_rate' ':packet=pts,dts,duration,size,side_data_list') streams = data.get('streams', []) if len(streams) != 1 or streams[0].get('codec_name') not in SUPPORTED_PICTURE_CODECS: raise ValueError('Pairing requires one selected H264/HEVC picture stream') if 'mp4' not in data.get('format', {}).get('format_name', '').split(','): raise ValueError('Pairing currently requires MP4-family picture container') v = streams[0] fps, tb = _time(v.get('avg_frame_rate')), _time(v.get('time_base')) if not 1 <= fps <= 120 or tb <= 0 or v.get('start_pts') != 0: raise ValueError('Picture requires positive CFR and a known zero start') duration = _time(v.get('duration_ts')) * tb frames = _probe(path, '-select_streams', 'v:0', '-show_frames', '-show_entries', 'frame=pts')['frames'] pts = [_time(frame.get('pts')) * tb for frame in frames] if not pts or any(t != Fraction(i, fps) for i, t in enumerate(pts)) or duration != len(pts) / fps: raise ValueError('Picture requires complete zero-origin CFR frame clock and exact duration') decoder_keys = ['codec_name', 'profile', 'level', 'width', 'height', 'pix_fmt', 'sample_aspect_ratio', 'field_order', 'color_range', 'color_space', 'color_transfer', 'color_primaries', 'chroma_location'] packets = [] for packet in data.get('packets', []): if packet.get('size') is None: raise ValueError('Picture packet missing size') ticks = {key: str(_time(packet.get(key)) * tb) for key in ['pts', 'dts', 'duration']} if _time(ticks['duration']) <= 0: raise ValueError('Picture packet has invalid duration') packets.append({**ticks, 'size': int(packet['size']), 'side_data_list': packet.get('side_data_list', [])}) if len(packets) != len(pts): raise ValueError('Picture packet/frame counts disagree') return {'decoder': {key: v.get(key) for key in decoder_keys}, 'packets': packets, 'frame_pts': [str(t) for t in pts], 'frame_count': len(pts), 'fps': str(fps), 'duration': str(duration), 'start': '0'} def validate_aac_packet_interval(audio): """Validate exact AAC packet continuity against the stream-header interval.""" if audio['codec'] != 'aac': raise ValueError('Pairing requires AAC adopted audio') packets = audio['packets'] rate = audio['sample_rate'] if type(rate) is not int or rate <= 0 or len(packets) < 2: raise ValueError('AAC packet timing is incomplete') nominal = _time(packets[0]['duration']) # Accept known AAC frame sample counts, never an arbitrary huge duration. if nominal * rate not in {960, 1024, 1920, 2048}: raise ValueError('Unsupported AAC nominal packet duration') for i, packet in enumerate(packets): d = _time(packet['duration']) if d <= 0 or d > nominal or (i < len(packets) - 1 and d != nominal): raise ValueError('Nonuniform or invalid AAC packet duration') p, t = _time(packet['pts']), _time(packet['dts']) if p != t: raise ValueError('Unsupported AAC PTS/DTS difference') if i and p != _time(packets[i-1]['pts']) + _time(packets[i-1]['duration']): raise ValueError('AAC packet clock has gaps or overlaps') audio_start, audio_duration = _time(audio['start_time']), _time(audio['duration']) if audio_duration <= 0: raise ValueError('Invalid audio interval') audio_end = audio_start + audio_duration first_pts = _time(packets[0]['pts']) last_end = _time(packets[-1]['pts']) + _time(packets[-1]['duration']) if (not 0 <= audio_start - first_pts <= nominal or abs(last_end - audio_end) > Fraction(1, rate)): raise ValueError('AAC packet clock does not match stream interval') return audio_start, audio_end, nominal def validate_pair_timing(picture, audio): """Narrow AAC/CFR compatibility; not a perceptual synchronization verdict.""" audio_start, audio_end, nominal = validate_aac_packet_interval(audio) tolerance = max(1 / Fraction(picture['fps']), nominal) end_delta = audio_end - Fraction(picture['duration']) if abs(audio_start) > tolerance or abs(end_delta) > tolerance: raise ValueError('Picture/audio interval mismatch; no implicit trim, padding, offset or retime') return {'start_delta': str(audio_start), 'end_delta': str(end_delta), 'tolerance': str(tolerance), 'nominal_aac_packet': str(nominal)} def _run_mux(command, directory): run_logged(command, directory, 'mux', timeout=600) def _verify_output(path): data = _probe(path, '-show_streams') if [stream.get('codec_type') for stream in data['streams']] != ['video', 'audio']: raise ValueError('Paired output must contain only v:0 and a:0') result = subprocess.run(['ffmpeg', '-v', 'error', '-xerror', '-threads', '2', '-i', str(path), '-map', '0:v:0', '-map', '0:a:0', '-f', 'null', '-'], capture_output=True, text=True, timeout=600) if result.returncode or result.stderr.strip(): raise ValueError('Paired output full decode failed') def run_pair(plan_path, output_dir, *, plan_only=False): plan_path, directory = Path(plan_path).resolve(), Path(output_dir).resolve() directory.mkdir(parents=True, exist_ok=False) report_path = directory / 'pair_run.json' staged, output = directory / 'paired.rendering.mp4', directory / 'paired.mp4' report = {'artifact': 'media_pair_run', 'schema_version': 1, 'status': 'PREPARING', 'direct_listening': 'NOT_CHECKED', 'normal_speed_review': 'NOT_CHECKED', 'release_approved': False} write_json_atomic(report_path, report) try: plan = json.loads(plan_path.read_bytes()) require_fields(plan, ['artifact', 'schema_version', 'picture', 'audio'], 'media_pair plan') if plan['artifact'] != 'media_pair' or type(plan['schema_version']) is not int or plan['schema_version'] != 1: raise ValueError('Unsupported media_pair schema') picture, audio = _asset(plan['picture']), _asset(plan['audio'], audio=True) report.update(plan={'path': str(plan_path)}, inputs={'picture': picture, 'audio': audio}) video_facts = probe_picture(picture['path']) audio_facts = probe_audio_packets(audio['path'], audio['selected_stream']) report['timing'] = validate_pair_timing(video_facts, audio_facts) report['picture'] = {k: v for k, v in video_facts.items() if k not in ['packets', 'frame_pts']} report['audio'] = {'input_stream': audio['selected_stream'], 'output_stream': 0, 'packet_count': audio_facts['packet_count']} if plan_only: report['status'] = 'PLANNED' write_json_atomic(report_path, report) return report command = ['ffmpeg', '-nostdin', '-v', 'error', '-n', '-copyts', '-i', picture['path'], '-i', audio['path'], '-map', '0:v:0', '-map', f"1:a:{audio['selected_stream']}", '-c', 'copy', '-movie_timescale', str(audio_facts['sample_rate']), '-movflags', '+faststart', str(staged)] _run_mux(command, directory) if probe_picture(staged) != video_facts: raise ValueError('Paired picture packets, decoder, geometry/color or full frame clock changed') proof = verify_adopted_audio(audio['path'], staged, audio['selected_stream'], 0) output_timing = validate_pair_timing(video_facts, proof['output']) if output_timing != report['timing']: raise ValueError('Output audio presentation interval changed') report['output_timing'] = output_timing _verify_output(staged) write_json_atomic(directory / 'picture_identity.json', video_facts) write_json_atomic(directory / 'adopted_audio_identity.json', proof) report['output'] = {'path': str(output), 'full_decode': 'PASS', 'picture_identity': 'EXACT', 'audio_packet_identity': 'EXACT'} staged.rename(output) report['status'] = 'PAIR_RENDERED' write_json_atomic(report_path, report) return report except Exception as exc: # A partial unique run is evidence, never a final/current asset. Keep logs. staged.unlink(missing_ok=True) output.unlink(missing_ok=True) report.pop('output', None) report.update(status='FAILED', error=f'{type(exc).__name__}: {exc}') write_json_atomic(report_path, report) raise def main(): parser = argparse.ArgumentParser(description=__doc__) parser.add_argument('plan') parser.add_argument('--output-dir', required=True) parser.add_argument('--plan-only', action='store_true') args = parser.parse_args() report = run_pair(args.plan, args.output_dir, plan_only=args.plan_only) print(json.dumps(report, ensure_ascii=False, indent=2)) if __name__ == '__main__': main() -
render_preflight.py 1.3 KB
"""Local ffmpeg capability preflight for subtitle burn-in.""" import shutil import subprocess from lib import CONFIG def _ffmpeg_filters(): """Return ffmpeg's compiled-in filter names (caller has confirmed ffmpeg exists).""" result = subprocess.run(["ffmpeg", "-hide_banner", "-filters"], text=True, capture_output=True, timeout=20) if result.returncode != 0: return set() filters = set() for line in result.stdout.splitlines(): parts = line.split() if len(parts) >= 2 and parts[0] and parts[0][0] in ".TSCAPN|": filters.add(parts[1]) return filters def _preflight_burn_subtitles(): """Fail before the (re-encoding) render when burn-in is on but ffmpeg lacks the libass `subtitles` filter. Only fires when ffmpeg EXISTS but can't burn — an absent ffmpeg fails the render regardless.""" if not CONFIG["burn_subtitles"]: return if shutil.which("ffmpeg") is None: return if "subtitles" not in _ffmpeg_filters(): raise SystemExit( "字幕烧录已开启,但当前 ffmpeg 不支持 subtitles/libass 滤镜,渲染会在最后一步失败。\n" " 解决:安装带 libass 的 ffmpeg,或加 --no-burn-subtitles 关闭烧录(仍输出 .srt 外挂字幕)。") -
source_score.py 23.7 KB
#!/usr/bin/env python3 """Prepare exact source-audio and continuous-score beds from a strict sample plan.""" import argparse from fractions import Fraction import json import os import math from pathlib import Path import re import shutil import subprocess from assemble_constants import SUPPORTED_PICTURE_CODECS, frame_clock_samples from adoption.strict_inputs import ( canonical_fraction, probe_json, read_json_bytes, require_declared_path, require_fields, require_integer, require_local_path, require_number, run_logged, without_digests, write_json_atomic, ) RATE = 48_000 CHANNELS = 2 CODEC = "pcm_f32le" SOURCE_ROLES = {"protected_original", "mixed_original_under_narration"} FADE_SHAPES = {"linear", "half_cosine"} def _probe_cfr(path): data = probe_json( path, "-select_streams", "v:0", "-show_streams", "-show_packets", "-show_entries", "stream=codec_name,time_base,start_pts,start_time,duration_ts,duration," "avg_frame_rate,r_frame_rate:packet=pts,duration", ) streams = data.get("streams", []) if len(streams) != 1: raise ValueError("source requires exactly one selected v:0 clock") stream = streams[0] codec = stream.get("codec_name") if codec not in SUPPORTED_PICTURE_CODECS: raise ValueError("v1 packet/frame CFR proof supports H264 and HEVC picture sources only") average = canonical_fraction(stream.get("avg_frame_rate"), "actual average frame rate") real = canonical_fraction(stream.get("r_frame_rate"), "actual real frame rate") if average != real: raise ValueError("source must be same-speed CFR") start_time = Fraction(stream.get("start_time", "0")) if start_time != 0: raise ValueError("source video frame clock must start at zero") time_base = Fraction(stream["time_base"]) packets = data.get("packets", []) try: pts = sorted(Fraction(int(packet["pts"])) * time_base for packet in packets) durations = [Fraction(int(packet["duration"])) * time_base for packet in packets] except (KeyError, TypeError, ValueError, ZeroDivisionError) as exc: raise ValueError("H26x packet/frame clock proof is incomplete") from exc frame_duration = 1 / average count = len(pts) if not pts or any(value != index * frame_duration for index, value in enumerate(pts)): raise ValueError("source picture packets do not prove a zero-origin CFR frame clock") if any(duration != frame_duration for duration in durations): raise ValueError("source picture packet durations are not one CFR frame") return {"fps": str(average), "frame_count": count, "time_base": stream.get("time_base"), "start_time": "0", "codec_name": codec, "clock_proof": "H26X_ONE_PACKET_PER_FRAME_PTS"} def _probe_audio(path, ordinal): data = probe_json( path, "-select_streams", f"a:{ordinal}", "-show_streams", "-show_entries", "stream=index,codec_name,sample_fmt,sample_rate,channels," "channel_layout,time_base,start_pts,start_time,duration_ts,duration", ) streams = data.get("streams", []) if len(streams) != 1: raise ValueError(f"audio stream a:{ordinal} is missing or ambiguous") stream = streams[0] return {key: stream.get(key) for key in ( "index", "codec_name", "sample_fmt", "sample_rate", "channels", "channel_layout", "time_base", "start_pts", "start_time", "duration_ts", "duration" )} def _probe_pcm(path): stream = _probe_audio(path, 0) if stream["codec_name"] not in {"pcm_s16le", "pcm_s24le", "pcm_f32le"}: raise ValueError("canonical/frozen WAV requires PCM16, PCM24, or PCM float") if stream["sample_rate"] != str(RATE) or stream["channels"] != CHANNELS: raise ValueError("canonical/frozen WAV must be 48 kHz stereo") time_base = Fraction(stream["time_base"]) samples = int(Fraction(stream["duration_ts"]) * time_base * RATE) if Fraction(stream["duration_ts"]) * time_base * RATE != samples: raise ValueError("WAV duration is not an integral 48 kHz sample count") return {"codec_name": stream["codec_name"], "sample_fmt": stream["sample_fmt"], "sample_rate": RATE, "channels": CHANNELS, "samples": samples} def _validate_fades(value, duration, curves, label): fade_in = require_integer(value["fade_in_samples"], f"{label} fade_in_samples") fade_out = require_integer(value["fade_out_samples"], f"{label} fade_out_samples") if fade_in + fade_out > duration: raise ValueError(f"{label} fades overlap") if fade_in == 1 or fade_out == 1: raise ValueError(f"{label} fades require zero or at least two samples") if value["fade_shape"] not in curves: raise ValueError(f"unsupported {label} fade_shape") return fade_in, fade_out def load_plan(plan_path): plan_path, _, plan = read_json_bytes(plan_path, "source score plan") require_fields(plan, ["artifact", "schema_version", "output", "source_segments", "source_silence", "score"], "source score plan") if plan["artifact"] != "source_score_plan" or type(plan["schema_version"]) is not int \ or plan["schema_version"] != 1: raise ValueError("unsupported source_score_plan schema") require_fields(plan["output"], ["sample_rate", "channels", "total_samples"], "output") if plan["output"]["sample_rate"] != RATE or plan["output"]["channels"] != CHANNELS: raise ValueError("v1 output is fixed at 48 kHz stereo") total = require_integer(plan["output"]["total_samples"], "total_samples", 1) if not isinstance(plan["source_segments"], list) or not isinstance(plan["source_silence"], list): raise ValueError("source segments and explicit silence must be lists") segments = [] asset_cache = {} picture_cache = {} ids = set() segment_fields = [ "id", "path", "audio_stream", "source_fps", "source_start_frame", "source_end_frame", "output_start_sample", "gain", "fade_in_samples", "fade_out_samples", "fade_shape", "role", ] for value in plan["source_segments"]: value = without_digests(value, "source segment") require_fields(value, segment_fields, "source segment") if not isinstance(value["id"], str) or not value["id"] or value["id"] in ids: raise ValueError("source segment id must be unique and non-empty") ids.add(value["id"]) ordinal = require_integer(value["audio_stream"], "audio_stream") path = require_declared_path(value, "source") fps = canonical_fraction(value["source_fps"], "source_fps") start = require_integer(value["source_start_frame"], "source_start_frame") end = require_integer(value["source_end_frame"], "source_end_frame", 1) if end <= start: raise ValueError("source frame interval must be non-empty") key = (str(path), ordinal) if key not in asset_cache: if str(path) not in picture_cache: picture_cache[str(path)] = _probe_cfr(path) picture = picture_cache[str(path)] audio = _probe_audio(path, ordinal) asset_cache[key] = {"path": str(path), "audio_stream": ordinal, "picture": picture, "audio": audio} facts = asset_cache[key] if facts["picture"]["fps"] != str(fps) or end > facts["picture"]["frame_count"]: raise ValueError("declared source frame clock/range differs from actual CFR source") source_start_sample = frame_clock_samples(start, fps, RATE) source_end_sample = frame_clock_samples(end, fps, RATE) duration = source_end_sample - source_start_sample if duration <= 0: raise ValueError("source frame interval is shorter than one 48 kHz sample") fade_in, fade_out = _validate_fades(value, duration, {"linear"}, "source") output_start = require_integer(value["output_start_sample"], "output_start_sample") output_end = output_start + duration if output_end > total or value["role"] not in SOURCE_ROLES: raise ValueError("source output range or role is unsupported") segments.append({**value, "path": str(path), "gain": require_number(value["gain"], "source gain", 0, 16), "source_start_sample": source_start_sample, "source_end_sample": source_end_sample, "output_end_sample": output_end, "fade_in_samples": fade_in, "fade_out_samples": fade_out, "asset_key": key}) silence = [] for value in plan["source_silence"]: require_fields(value, ["output_start_sample", "output_end_sample", "role"], "source silence") start = require_integer(value["output_start_sample"], "silence output_start_sample") end = require_integer(value["output_end_sample"], "silence output_end_sample", 1) if value["role"] != "silence" or not start < end <= total: raise ValueError("invalid explicit source silence range") silence.append(dict(value)) coverage = sorted( [(item["output_start_sample"], item["output_end_sample"]) for item in segments] + [(item["output_start_sample"], item["output_end_sample"]) for item in silence] ) cursor = 0 for start, end in coverage: if start != cursor: raise ValueError("source bed requires exact nonoverlapping coverage with explicit silence") cursor = end if cursor != total: raise ValueError("source bed has an implicit tail gap") score = plan["score"] if not isinstance(score, dict) or score.get("kind") not in {"raw", "frozen", "none"}: raise ValueError("score kind must be raw, frozen, or none") score = without_digests(score, "score") if score["kind"] == "raw": require_fields(score, ["kind", "path", "audio_stream", "source_offset_sample", "gain", "fade_in_samples", "fade_out_samples", "fade_shape"], "raw score") score_path = require_declared_path(score, "raw score") ordinal = require_integer(score["audio_stream"], "score audio_stream") offset = require_integer(score["source_offset_sample"], "score source_offset_sample") fade_in, fade_out = _validate_fades(score, total, FADE_SHAPES, "score") score = {**score, "path": str(score_path), "audio_stream": ordinal, "source_offset_sample": offset, "gain": require_number(score["gain"], "score gain", 0, 16), "fade_in_samples": fade_in, "fade_out_samples": fade_out, "audio": _probe_audio(score_path, ordinal)} elif score["kind"] == "frozen": require_fields(score, ["kind", "path", "audio_stream"], "frozen score") score_path = require_declared_path(score, "frozen score") ordinal = require_integer(score["audio_stream"], "score audio_stream") if ordinal != 0: raise ValueError("frozen WAV supports only a:0") pcm = _probe_pcm(score_path) if pcm["samples"] != total: raise ValueError("frozen score must exactly match total_samples") score = {**score, "path": str(score_path), "pcm": pcm} else: require_fields(score, ["kind"], "none score") return {"plan_path": plan_path, "output": {**plan["output"], "codec": CODEC}, "source_segments": segments, "source_silence": silence, "source_assets": list(asset_cache.values()), "score": score} _run = run_logged def _decode_command(path, ordinal, output): return ["ffmpeg", "-nostdin", "-v", "error", "-n", "-copyts", "-i", str(path), "-map", f"0:a:{ordinal}", "-af", "aresample=48000:async=0:first_pts=0,aformat=sample_fmts=flt:" "sample_rates=48000:channel_layouts=stereo", "-c:a", CODEC, str(output)] def _fade_filters(duration, fade_in, fade_out, curve): factors = [] if fade_in: if curve == "linear": factors.append(f"min(1\\,n/{fade_in - 1})") else: position = f"max(0\\,min(1\\,n/{fade_in - 1}))" factors.append(f"0.5-0.5*cos(PI*{position})") if fade_out: remaining = f"({duration - 1}-n)" if curve == "linear": factors.append(f"max(0\\,min(1\\,{remaining}/{fade_out - 1}))") else: position = f"max(0\\,min(1\\,{remaining}/{fade_out - 1}))" factors.append( f"0.5-0.5*cos(PI*{position})" ) if not factors: return [] gain = "*".join(f"({factor})" for factor in factors) return [f"aeval=val(0)*({gain})|val(1)*({gain}):c=same"] def _render_source(plan, decoded, output, directory): total = plan["output"]["total_samples"] command = ["ffmpeg", "-nostdin", "-v", "error", "-n"] assets = {tuple(asset_key): index for index, asset_key in enumerate(decoded)} for asset_key in decoded: command += ["-i", str(decoded[asset_key])] filters = [f"anullsrc=r={RATE}:cl=stereo,atrim=end_sample={total}[base]"] inputs = ["[base]"] for index, segment in enumerate(plan["source_segments"]): duration = segment["source_end_sample"] - segment["source_start_sample"] chain = [f"[{assets[tuple(segment['asset_key'])]}:a]atrim=" f"start_sample={segment['source_start_sample']}:" f"end_sample={segment['source_end_sample']}", "asetpts=PTS-STARTPTS", f"volume={segment['gain']:.17g}"] chain += _fade_filters(duration, segment["fade_in_samples"], segment["fade_out_samples"], "linear") chain.append(f"adelay={segment['output_start_sample']}S:all=1[src{index}]") filters.append(",".join(chain)) inputs.append(f"[src{index}]") filters.append("".join(inputs) + f"amix=inputs={len(inputs)}:duration=first:normalize=0," f"atrim=end_sample={total},aformat=sample_fmts=flt:" f"sample_rates={RATE}:channel_layouts=stereo[out]") command += ["-filter_complex", ";".join(filters), "-map", "[out]", "-c:a", CODEC, str(output)] _run(command, directory, "source") def _render_score(plan, decoded_score, output, directory): total = plan["output"]["total_samples"] score = plan["score"] if score["kind"] == "none": command = [ "ffmpeg", "-nostdin", "-v", "error", "-n", "-f", "lavfi", "-i", f"anullsrc=r={RATE}:cl=stereo", "-af", f"atrim=end_sample={total}", "-c:a", CODEC, str(output), ] elif score["kind"] == "frozen": command = ["ffmpeg", "-nostdin", "-v", "error", "-n", "-i", score["path"], "-map", "0:a:0", "-af", "aformat=sample_fmts=flt:sample_rates=48000:" "channel_layouts=stereo", "-c:a", CODEC, str(output)] else: chain = [f"atrim=start_sample={score['source_offset_sample']}:" f"end_sample={score['source_offset_sample'] + total}", "asetpts=PTS-STARTPTS", f"volume={score['gain']:.17g}"] chain += _fade_filters(total, score["fade_in_samples"], score["fade_out_samples"], score["fade_shape"]) chain += [f"atrim=end_sample={total}", "aformat=sample_fmts=flt:sample_rates=48000:" "channel_layouts=stereo"] command = ["ffmpeg", "-nostdin", "-v", "error", "-n", "-i", str(decoded_score), "-af", ",".join(chain), "-c:a", CODEC, str(output)] _run(command, directory, "score") def _astats(path): result = subprocess.run( ["ffmpeg", "-nostdin", "-v", "info", "-i", str(path), "-af", "astats=metadata=0:reset=0", "-f", "null", "-"], capture_output=True, text=True, timeout=3600, ) if result.returncode: raise ValueError("output astats decode failed") peaks = re.findall(r"Peak level dB:\s*(-?inf|[-+0-9.]+)", result.stderr, re.I) nan_counts = re.findall(r"Number of NaNs:\s*(\d+)", result.stderr) inf_counts = re.findall(r"Number of Infs:\s*(\d+)", result.stderr) if not peaks: raise ValueError("output peak statistics unavailable") peak_db = max(float(value) if value.lower() != "-inf" else -math.inf for value in peaks) finite = all(int(value) == 0 for value in nan_counts + inf_counts) if not finite: raise ValueError("output contains non-finite PCM samples") return {"finite": True, "peak": 0.0 if peak_db == -math.inf else 10 ** (peak_db / 20)} def _output_facts(path): """PCM format, sample count, size and peak/finiteness of one 48 kHz stereo float WAV.""" return {"path": str(path), "bytes": os.stat(path).st_size, "pcm": _probe_pcm(path), **_astats(path)} def validate_prepared_receipt(reference, expected_format): """Validate one completed prepared-bed receipt and its three current PCM stems.""" require_fields(without_digests(reference, "prepared receipt reference"), ["path"], "prepared receipt reference") require_fields(expected_format, ["sample_rate", "channels", "total_samples"], "expected prepared format") if expected_format["sample_rate"] != RATE or expected_format["channels"] != CHANNELS: raise ValueError("prepared format must be 48 kHz stereo") require_integer(expected_format["total_samples"], "prepared total_samples", 1) receipt_path = require_declared_path(reference, "prepared receipt") receipt = read_json_bytes(receipt_path, "prepared receipt")[2] if not isinstance(receipt, dict) or receipt.get("artifact") != "prepared_bed_receipt" \ or receipt.get("schema_version") != 1 or receipt.get("status") != "PREPARED": raise ValueError("prepared receipt is not a completed v1 artifact") required_format = {**expected_format, "codec": CODEC} if receipt.get("format") != required_format: raise ValueError("prepared receipt format differs from expected picture clock") outputs = receipt.get("outputs") if not isinstance(outputs, dict) or set(outputs) != { "source_bed.wav", "score_bed.wav", "prepared_bed.wav" }: raise ValueError("prepared receipt requires exactly three named bed outputs") prepared = {} for name, declared in outputs.items(): if not isinstance(declared, dict): raise ValueError(f"prepared receipt {name} identity is invalid") path = require_local_path(declared.get("path"), name) stream = _probe_audio(path, 0) if stream["codec_name"] != CODEC or stream["sample_rate"] != str(RATE) or \ stream["channels"] != CHANNELS or \ (stream["start_time"] not in (None, "N/A") and Fraction(stream["start_time"]) != 0): raise ValueError(f"{name} must be zero-origin 48 kHz stereo float PCM") actual = _output_facts(path) if actual["pcm"]["codec_name"] != CODEC \ or actual["pcm"]["samples"] != expected_format["total_samples"]: raise ValueError(f"{name} does not match the required full PCM format") prepared[name] = actual return { "reference": {"path": str(receipt_path)}, "format": dict(expected_format), "outputs": prepared, "receipt": receipt, } def prepare_source_score(plan_path, output_dir): directory = Path(output_dir).resolve() directory.mkdir(parents=True, exist_ok=False) finals = {name: directory / name for name in ("source_bed.wav", "score_bed.wav", "prepared_bed.wav")} staged = {name: directory / f".{Path(name).stem}.rendering.wav" for name in finals} receipt_path = directory / "prepared_bed_receipt.json" try: plan = load_plan(plan_path) decoded = {} for index, asset in enumerate(plan["source_assets"]): key = (asset["path"], asset["audio_stream"]) path = directory / f".source_{index:03d}.decoded.wav" _run(_decode_command(asset["path"], asset["audio_stream"], path), directory, f"decode_source_{index:03d}") decoded[key] = path asset["canonical_pcm"] = _output_facts(path) required = max( segment["source_end_sample"] for segment in plan["source_segments"] if tuple(segment["asset_key"]) == key ) if asset["canonical_pcm"]["pcm"]["samples"] < required: raise ValueError("decoded source audio is too short for a selected frame range") score_decode = None if plan["score"]["kind"] == "raw": score_decode = directory / ".score.decoded.wav" _run(_decode_command(plan["score"]["path"], plan["score"]["audio_stream"], score_decode), directory, "decode_score") plan["score"]["canonical_decode"] = _output_facts(score_decode) if plan["score"]["canonical_decode"]["pcm"]["samples"] < \ plan["score"]["source_offset_sample"] + plan["output"]["total_samples"]: raise ValueError("raw score is too short for continuous offset window") _render_source(plan, decoded, staged["source_bed.wav"], directory) _render_score(plan, score_decode, staged["score_bed.wav"], directory) rendered = { "source_bed.wav": _output_facts(staged["source_bed.wav"]), "score_bed.wav": _output_facts(staged["score_bed.wav"]), } if plan["score"]["kind"] == "none": shutil.copyfile(staged["source_bed.wav"], staged["prepared_bed.wav"]) else: command = [ "ffmpeg", "-nostdin", "-v", "error", "-n", "-i", str(staged["source_bed.wav"]), "-i", str(staged["score_bed.wav"]), "-filter_complex", f"[0:a][1:a]amix=inputs=2:duration=first:normalize=0," f"atrim=end_sample={plan['output']['total_samples']},aformat=sample_fmts=flt:" "sample_rates=48000:channel_layouts=stereo[out]", "-map", "[out]", "-c:a", CODEC, str(staged["prepared_bed.wav"]), ] _run(command, directory, "prepare") outputs = {**rendered, "prepared_bed.wav": _output_facts(staged["prepared_bed.wav"])} if any(value["pcm"]["samples"] != plan["output"]["total_samples"] for value in outputs.values()): raise ValueError("prepared bed sample counts differ from plan") if plan["score"]["kind"] == "none" and ( outputs["prepared_bed.wav"]["bytes"] != outputs["source_bed.wav"]["bytes"] or outputs["prepared_bed.wav"]["pcm"] != outputs["source_bed.wav"]["pcm"]): raise ValueError("none score must preserve the source bed exactly") outputs["prepared_bed.wav"]["headroom_policy"] = "FLOAT_PRESERVED_NO_MASTER" receipt = { "artifact": "prepared_bed_receipt", "schema_version": 1, "status": "PREPARED", "plan": {"path": str(plan["plan_path"])}, "format": plan["output"], "source_assets": plan["source_assets"], "source_segments": [{key: value for key, value in item.items() if key != "asset_key"} for item in plan["source_segments"]], "source_silence": plan["source_silence"], "score": plan["score"], "outputs": outputs, "direct_listening": "NOT_CHECKED", "release_approved": False, } for name, path in staged.items(): path.rename(finals[name]) receipt["outputs"][name]["path"] = str(finals[name]) write_json_atomic(receipt_path, receipt) return receipt except Exception: for path in [*staged.values(), *finals.values(), receipt_path]: path.unlink(missing_ok=True) raise def main(): parser = argparse.ArgumentParser(description=__doc__) parser.add_argument("plan") parser.add_argument("--output-dir", required=True) args = parser.parse_args() print(json.dumps(prepare_source_score(args.plan, args.output_dir), ensure_ascii=False)) if __name__ == "__main__": main() -
source_subtitles.py 20 KB
"""Original-dialogue subtitle loading, source mapping, and gap placement.""" import json import re from pathlib import Path from artifacts import _load_work_json from audio_mix import _seg_place_window from lib import CONFIG from media import _plan_clip_spans from assemble_constants import ( _AUTO_ORIGINAL_READ_CPS, _CLIP_CONTIGUITY_TOLERANCE, _MAX_ORIGINAL_READ_CPS, _MIN_ASR_CLIP_OVERLAP, _MIN_GAP_TO_SUBTITLE, _MIN_READABLE_SECONDS, _SUBTITLE_CLOSING_QUOTES, ) from subtitles.track_binding import bound_subtitle_entries from subtitles.core import ( _bracketed_original_chunks, _subtitle_entries, ) def _has_user_subtitles(work_dir): """True when the user dropped a bring-your-own original-subtitle file into work_dir.""" return work_dir is not None and any( (Path(work_dir) / name).exists() for name in ("user_subtitles.json", "user_subtitles.srt", "user_subtitles.ass") ) def _source_subtitle_mask_policy(work_dir=None): """Explicit source-subtitle mask policy and trigger facts for visual QC/cache keys. Older builds treated ``MASK_SOURCE_SUBTITLES=True`` as an ambient default black band. The visual contract now requires an explicit policy, so a bare truthy legacy flag is represented as ``legacy_implicit`` and blocks the visual gate instead of silently masking picture information. """ burn = CONFIG["burn_subtitles"] raw_policy = CONFIG["source_subtitle_mask_policy"] legacy_flag = CONFIG["mask_source_subtitles"] allowed = {"off", "opt_in", "safe", "forced"} declared = CONFIG["source_subtitle_mask_policy_declared"] or raw_policy in {"opt_in", "safe", "forced"} implicit = False if legacy_flag and not declared: raw_policy = "legacy_implicit" implicit = True elif raw_policy not in allowed: implicit = True user_subtitles = _has_user_subtitles(work_dir) active = False trigger = "policy_off" reason = "source subtitle masking disabled by explicit policy" if raw_policy == "off": active = False elif raw_policy in {"opt_in", "forced"}: active = burn and legacy_flag trigger = "burn_subtitles_and_legacy_mask_flag" reason = "explicit policy permits masking only with burned recap subtitles" elif raw_policy == "safe": active = burn and (legacy_flag or user_subtitles) trigger = "safe_policy_with_burned_subtitles" reason = "safe policy masks only when recap subtitles are burned and an original-subtitle source is declared" else: active = False trigger = "implicit_or_invalid_policy" reason = "mask_source_subtitles requires explicit SOURCE_SUBTITLE_MASK_POLICY" if not burn and active: active = False trigger = "burn_subtitles_disabled" reason = "mask-only black band is forbidden without burned recap subtitles" return { "policy": raw_policy, "declared": bool(declared and raw_policy in allowed), "active": bool(active), "scope": ( "measured_source_subtitle_band" if active and 0 <= CONFIG["subtitle_y_top"] < CONFIG["subtitle_y_bot"] else ("bottom_source_subtitle_band" if active else "none") ), "trigger": trigger, "reason": reason, "burn_subtitles": burn, "legacy_mask_flag": legacy_flag, "user_subtitles_present": user_subtitles, "blocking": implicit, } def _load_original_asr(work_dir): """The original speech transcription (asr_result.json), SOURCE-time [{start,end,text}]; [] when absent.""" data = _load_work_json(work_dir, "asr_result.json") if data is None: return [] return [{"start": float(s["start"]), "end": float(s["end"]), "text": s["text"]} for s in data] def _load_agent_original_subtitles(work_dir): """Agent-calibrated original-dialogue subtitles (original_subtitles.json): OUTPUT-time [{start,end,text}] the writer authors alongside narration.json — the corrected, gap-aligned transcript of what is ACTUALLY said in each original-audio gap (ASR errors/names fixed). None when absent or empty (then assemble falls back to a conservative auto-ASR mapping).""" data = _load_work_json(work_dir, "original_subtitles.json") if not data: return None return [{"start": float(s["start"]), "end": float(s["end"]), "text": s["text"]} for s in data] def _user_subtitle_entries(rows, source): """Validate user-authored {start,end,text} rows; a malformed row is an error, not a skip.""" out = [] for index, row in enumerate(rows, start=1): try: start, end, text = float(row["start"]), float(row["end"]), row["text"].strip() except (KeyError, TypeError, ValueError, AttributeError) as exc: raise ValueError(f"{source} 第 {index} 条缺少或无法解析 start/end/text: {row!r}") from exc if end <= start or not text: raise ValueError(f"{source} 第 {index} 条无效(需要 end > start 且 text 非空): {row!r}") out.append({"start": start, "end": end, "text": text}) return out def _parse_srt_timestamp(value): """Parse an SRT 'HH:MM:SS,mmm' (or ASS 'H:MM:SS.cc') timestamp into seconds.""" m = re.match(r"\s*(\d+):(\d{1,2}):(\d{1,2})[.,](\d{1,3})\s*$", value) if not m: raise ValueError(f"无法解析字幕时间戳: {value!r}") h, mm, ss, frac = m.groups() return int(h) * 3600 + int(mm) * 60 + int(ss) + int(frac) / (10 ** len(frac)) def _parse_srt_text(text, source): """Minimal SRT parser → [{start,end,text}]. Blank lines and missing indices are tolerated; a block without a parseable timing line is an error. Cues with no text are dropped.""" segs = [] for block in re.split(r"\n\s*\n", text.replace("\r\n", "\n").replace("\r", "\n")): lines = [ln for ln in block.split("\n") if ln.strip()] if not lines: continue if lines[0].strip().isdigit(): lines = lines[1:] if not lines or "-->" not in lines[0]: raise ValueError(f"{source}: 字幕块缺少时间行: {block.strip()!r}") start_text, end_text = lines[0].split("-->", 1) start, end = _parse_srt_timestamp(start_text), _parse_srt_timestamp(end_text) if end <= start: raise ValueError(f"{source}: 字幕结束时间必须晚于开始时间: {lines[0].strip()!r}") body = " ".join(lines[1:]).strip() if body: segs.append({"start": start, "end": end, "text": body}) return segs def _parse_ass_text(text, source): """Minimal ASS Dialogue parser → [{start,end,text}] (Start, End are fields 2 and 3).""" segs = [] for line in text.replace("\r\n", "\n").replace("\r", "\n").split("\n"): if not line.startswith("Dialogue:"): continue fields = line[len("Dialogue:"):].split(",", 9) if len(fields) < 10: raise ValueError(f"{source}: Dialogue 行字段不足: {line!r}") start, end = _parse_srt_timestamp(fields[1]), _parse_srt_timestamp(fields[2]) if end <= start: raise ValueError(f"{source}: 字幕结束时间必须晚于开始时间: {line!r}") body = re.sub(r"\{[^}]*\}", "", fields[9]).replace("\\N", " ").replace("\\n", " ").strip() if body: segs.append({"start": start, "end": end, "text": body}) return segs def _load_user_original_subtitles(work_dir): """User-supplied original-dialogue subtitles, the highest-priority source (above the agent file). Accepts (first existing wins): - user_subtitles.json: a bare list [{start,end,text}] (treated as OUTPUT-time, used verbatim), OR a wrapper {"timeline":"source"|"output", "lines":[...]} — "source" is remapped to OUTPUT via the cut clip spans, "output" (default) is used directly. - user_subtitles.srt / user_subtitles.ass: parsed minimally and defaulted to SOURCE-time, so they are remapped to OUTPUT via the cut clip spans. Returns OUTPUT-time [{start,end,text}], or None when no user file exists. A malformed file raises: the user asked for these subtitles, so silently falling back is wrong.""" work = Path(work_dir) json_path = work / "user_subtitles.json" if json_path.exists(): try: data = json.loads(json_path.read_text(encoding="utf-8")) except ValueError as exc: raise ValueError(f"{json_path} 不是合法 JSON: {exc}") from exc if isinstance(data, dict): timeline = data.get("timeline", "output") if timeline not in {"source", "output"}: raise ValueError(f"{json_path}: timeline 必须是 source 或 output,当前为 {timeline!r}") rows = data["lines"] else: timeline, rows = "output", data segs = _user_subtitle_entries(rows, json_path.name) if timeline == "source": segs = _map_asr_to_output(segs, _plan_clip_spans(work)) return segs for name in ("user_subtitles.srt", "user_subtitles.ass"): path = work / name if not path.exists(): continue parser = _parse_ass_text if name.endswith(".ass") else _parse_srt_text segs = parser(path.read_text(encoding="utf-8"), name) # .srt/.ass default to SOURCE-time → remap onto the output timeline (identity in full mode). return _map_asr_to_output(segs, _plan_clip_spans(work)) return None def _map_asr_to_output(asr_segs, clip_spans): """Map SOURCE-time ASR segments onto the OUTPUT timeline. Full mode (clip_spans None) is identity. Keep one utterance across source/output-continuous picture cuts; real source deletions, output gaps and repeated playback remain separate.""" if clip_spans is None: return [dict(s) for s in asr_segs] out = [] for seg in asr_segs: fragments = [] for c in clip_spans: ov_s, ov_e = max(seg["start"], c["source_start"]), min(seg["end"], c["source_end"]) if ov_e <= ov_s: continue start = c["output_start"] + (ov_s - c["source_start"]) end = c["output_start"] + (ov_e - c["source_start"]) entry = c.get("entry", {}) source_key = (c.get("source_id") or entry.get("source_id"), c.get("source_path") or entry.get("source_path")) if (fragments and abs(fragments[-1]["source_end"] - ov_s) < _CLIP_CONTIGUITY_TOLERANCE and abs(fragments[-1]["end"] - start) < _CLIP_CONTIGUITY_TOLERANCE and fragments[-1]["source_key"] == source_key): fragments[-1].update(end=end, source_end=ov_e) else: fragments.append({"start": start, "end": end, "source_end": ov_e, "source_key": source_key}) # Filter after joining: individually tiny pieces may be one readable phrase. out.extend({"start": f["start"], "end": f["end"], "text": seg["text"]} for f in fragments if f["end"] - f["start"] > _MIN_ASR_CLIP_OVERLAP) return out def _narration_gap_windows(tts_segments, video_duration, min_gap=_MIN_GAP_TO_SUBTITLE): """OUTPUT-timeline stretches with NO narration (the original-audio blocks): the complement of the merged narration placement windows within [0, video_duration], keeping gaps >= min_gap.""" placed = sorted(map(_seg_place_window, tts_segments), key=lambda w: w[0]) merged = [] for s, e in placed: if e - s <= 0: continue if merged and s <= merged[-1][1]: merged[-1][1] = max(merged[-1][1], e) else: merged.append([s, e]) gaps, cursor = [], 0.0 for s, e in merged: if s - cursor >= min_gap: gaps.append((cursor, s)) cursor = max(cursor, e) if video_duration - cursor >= min_gap: gaps.append((cursor, float(video_duration))) return gaps def _original_gap_subtitle_entries(tts_segments, work_dir, video_duration): """Subtitle entries for the ORIGINAL dialogue during the original-audio blocks (narration gaps), so the band is not blank while the original speaks. Off unless we are burning and subtitle_original_in_gaps is set; no-op when there is no ASR. Cut mode remaps ASR to output.""" # Fill the gaps when either (a) we are masking the source's own burned-in subs (so the band is # blank without us), or (b) the user supplied their own subtitle file — a clear signal they want # the original dialogue shown, e.g. a clean/foreign source with mask OFF (no burned subs to # double). Without a user file we keep the mask requirement so we don't double the source's own # visible subs. subtitle_original_in_gaps is the explicit override either way. if not (CONFIG["burn_subtitles"] and CONFIG["subtitle_original_in_gaps"] and (_source_subtitle_mask_covers_gaps(work_dir) or _has_user_subtitles(work_dir))): return [] gaps = _narration_gap_windows(tts_segments, video_duration) if not gaps: return [] # Source ladder (highest priority first): user-supplied file → agent-calibrated transcript → # conservative auto-ASR mapping. The user file and the agent file are time-precise (their spans # are the real on-screen windows), so they take the interval-clip "precise" path; raw ASR is # coarse and stays on the midpoint+over-render-guard fallback path. max_chars = CONFIG["subtitle_max_chars"] user = _load_user_original_subtitles(work_dir) if user is not None: return _precise_gap_entries(user, gaps, max_chars) agent = _load_agent_original_subtitles(work_dir) if agent is not None: return _precise_gap_entries(agent, gaps, max_chars) asr = _load_original_asr(work_dir) if not asr: return [] return _fallback_gap_entries(_map_asr_to_output(asr, _plan_clip_spans(work_dir)), gaps, max_chars) def _precise_gap_entries(candidates, gaps, max_chars): """Precise path for time-accurate sources (user / agent-calibrated): interval-CLIP each line across the gap boundaries it overlaps, emitting one sub-entry per overlapped gap (clipped to that gap). A line straddling two gaps is split, not snapped to one or dropped; only sub-fragments shorter than _MIN_READABLE_SECONDS are dropped. No over-render guard (the source is trusted).""" entries = [] for seg in candidates: text = seg["text"] seg_start, seg_end = float(seg["start"]), float(seg["end"]) overlaps = [ (max(seg_start, gs), min(seg_end, ge)) for gs, ge in gaps if min(seg_end, ge) - max(seg_start, gs) >= _MIN_READABLE_SECONDS ] if not overlaps: continue if len(overlaps) == 1: # the common case (a line authored within one gap): show it whole in that gap cs, ce = overlaps[0] entries.extend(_bracketed_original_chunks(text, cs, ce, max_chars)) continue # the line straddles a narration block: show each gap only ITS portion of the text # (proportional to the time the line overlaps that gap) instead of the whole line twice. seg_dur = seg_end - seg_start n = len(text) for cs, ce in overlaps: lo = max(0, int(round((cs - seg_start) / seg_dur * n))) hi = min(n, int(round((ce - seg_start) / seg_dur * n))) piece = text[lo:hi].strip() if piece: entries.extend(_bracketed_original_chunks(piece, cs, ce, max_chars)) return entries def _split_sentences_keep_delims(text): """Split on terminal CJK sentence marks 。!? keeping each delimiter with its sentence. A fragment that is only closing quotes/brackets (e.g. a trailing 」 after a 。 inside a quote) is re-attached to the previous sentence so quoted speech is never split off into a bare 」.""" parts = [p.strip() for p in re.split(r"(?<=[。!?])", text) if p.strip()] merged = [] for part in parts: if merged and all(ch in _SUBTITLE_CLOSING_QUOTES for ch in part): merged[-1] += part else: merged.append(part) return merged def _fallback_gap_entries(candidates, gaps, max_chars): """Coarse-ASR fallback. Each coarse-ASR line spans a whole window with no per-sentence onset, so it is split into WHOLE sentences (never mid-word); each sentence is assigned to the gap its char-proportional midpoint lands in, and within a gap the assigned sentences are packed SEQUENTIALLY from the first one's estimated onset at a comfortable read rate — so two lines in one gap never overlap or scatter to char-proportional tail slots — capped at the gap end. An over-dense gap front-truncates (shown) rather than dropping to blank.""" # 1) split each coarse line into WHOLE sentences (never mid-word) and assign each to the gap # its char-proportional midpoint lands in (the only "which gap" signal coarse ASR gives). buckets = {} # gap_index -> [(estimated_onset, sentence_text)] for seg in candidates: for text in _split_sentences_keep_delims(seg["text"]): sub = _sentence_subspan(seg, text) mid = (sub["start"] + sub["end"]) / 2.0 gi = next((i for i, (gs, ge) in enumerate(gaps) if gs <= mid < ge), None) if gi is None: continue buckets.setdefault(gi, []).append((sub["start"], text)) # 2) within each gap, pack the assigned sentences SEQUENTIALLY from the gap onset at a # comfortable read rate. Anchoring to the gap onset (vs each sentence's char-proportional # tail position) stops a line heard early from being shoved to the END of its window — the # coarse-ASR lag. Full mode keeps the real ASR onset (the first sentence's own start); an # over-dense gap front-truncates (shown) rather than dropping to blank. entries = [] for gi, items in buckets.items(): gs, ge = gaps[gi] items.sort(key=lambda it: it[0]) # start at the first assigned sentence's estimated onset (clamped into the gap), then pack # the rest sequentially so they never overlap or scatter to char-proportional tail slots. cursor = min(ge, max(gs, min(start for start, _ in items))) for _, text in items: if cursor >= ge - _MIN_READABLE_SECONDS: break ce2 = min(ge, cursor + max(_MIN_READABLE_SECONDS, len(text) / _AUTO_ORIGINAL_READ_CPS)) if ce2 - cursor < _MIN_READABLE_SECONDS: break max_len = int((ce2 - cursor) * _MAX_ORIGINAL_READ_CPS) entries.extend(_bracketed_original_chunks(text[:max_len], cursor, ce2, max_chars)) cursor = ce2 return entries def _sentence_subspan(seg, sentence): """The slice of seg's [start,end] window that this sentence occupies, by character proportion. Single-sentence lines return the whole span unchanged.""" full = seg["text"].strip() idx = full.find(sentence) if not full or sentence == full or idx < 0: return {"start": seg["start"], "end": seg["end"]} span = seg["end"] - seg["start"] s = seg["start"] + span * (idx / len(full)) e = seg["start"] + span * ((idx + len(sentence)) / len(full)) return {"start": s, "end": e} def _combined_subtitle_entries(narration, work_dir, video_duration): """Narration subtitle entries plus original-dialogue entries in the gaps, sorted by start. Original entries are confined to narration gaps, so they never overlap narration entries.""" bound = bound_subtitle_entries(work_dir, video_duration) if bound is not None: return bound entries = _subtitle_entries(narration) entries.extend(_original_gap_subtitle_entries(narration, work_dir, video_duration)) entries.sort(key=lambda x: (x["start"], x["end"])) return entries def _source_subtitle_mask_covers_gaps(work_dir=None): """Whether the effective source mask hides hardcoded subtitles outside narration.""" if not _source_subtitle_mask_policy(work_dir)["active"]: return False return CONFIG["subtitle_mask_opacity"] >= 1.0 - 1e-9 and CONFIG["source_subtitle_mask_timing"] == "all" -
timeline.py 9.2 KB
"""Multi-track timeline model for the recap (backend-neutral, stdlib only). A `Timeline` is a small, serializable representation of the finished recap as a set of tracks — exactly like a cut-tool project: - one **video** track: the source clip(s), each carrying its own *original audio* with a per-clip volume automation (the ducking: a continuous low bed under narration, held across short inter-sentence gaps, back up only at the lead-in/out and genuine long gaps); - one **narration** audio track: the placed TTS beats; - an optional **bgm** audio track: a looped music bed with its own ducking; - one **subtitle** (text) track: the narration lines. - optional **image** tracks: local photo overlays with normalized center-origin, Y-up transforms for editable JianYing export. The canonical ducking semantics live in `audio_automation.py`; ffmpeg (`assemble.py`) and this timeline model both derive their automation from that shared source. This model is emitted as `timeline.json` and consumed by the *optional* 剪映 exporter. The model itself knows nothing about ffmpeg or 剪映 — times are plain seconds and volumes are plain gains, so any backend can read it. """ import json import math from copy import deepcopy from audio_automation import fixed_ducking_keyframes as ducking_keyframes from audio_automation import release_ducking_keyframes SCHEMA_VERSION = 2 def _ceil_time(value, digits=4): """Round an interval end outward so serialization can never shorten media.""" scale = 10 ** digits return math.ceil((float(value) * scale) - 1e-9) / scale def _floor_time(value, digits=4): """Round an interval start outward so serialization cannot clip source samples.""" scale = 10 ** digits return math.floor((float(value) * scale) + 1e-9) / scale def build_timeline(canvas, duration_s, video_clips, narration_segments, bgm=None, ducking=None, subtitle_segments=None, image_segments=(), resource_packages=None, style_presets=None, extra_tracks=()): """Assemble a Timeline dict from resolved placement data. canvas: {"width", "height", "fps"} duration_s: total output length (seconds) video_clips: ordered [{"source_path", "source_start", "source_end", "timeline_start", "timeline_end"}] (cut mode: one per clip; full mode: a single clip spanning the whole video). narration_segments: placed beats [{"source_path", "timeline_start", "timeline_end", "text", "overlaps_speech", "gain"}]; a zero-width beat (an unplaced segment) is skipped. subtitle_segments: optional display-ready text cues [{"text", "timeline_start", "timeline_end"}]. When present, this is authoritative for the subtitle/text track; narration segment text remains raw editor metadata. bgm: optional {"source_path", "volume", "ducking_volume", "fade"}. ducking: {"idle", "speech", "quiet", "fade", "bridge"} for the original-audio automation; None disables original ducking (flat original). `bridge` holds the duck across inter-beat gaps shorter than it. image_segments: optional v2 local image overlays [{"source_path", "timeline_start", "timeline_end", ...authoring extensions}], passed through as authored. """ placed = [ s for s in narration_segments if float(s["timeline_end"]) > float(s["timeline_start"]) ] windows = [(float(s["timeline_start"]), float(s["timeline_end"])) for s in placed] duck_windows = [] if ducking is not None: for s in placed: end = float(s["timeline_end"]) hold_end = max(end, float(s.get("source_duck_end", end))) restore_at = max( hold_end, float(s.get("source_restore_at", hold_end + float(ducking["fade"]))) ) level = float(ducking["speech" if s["overlaps_speech"] else "quiet"]) duck_windows.append((float(s["timeline_start"]), hold_end, level, restore_at)) # --- video track: each clip carries its original audio + ducking automation video_clip_objs = [] for c in video_clips: ts, te = float(c["timeline_start"]), float(c["timeline_end"]) audio = {"role": "original", "volume_keyframes": []} if ducking is not None: audio["volume_keyframes"] = release_ducking_keyframes( duck_windows, ducking["idle"], ducking["fade"], ts, te, bridge=ducking["bridge"]) audio["base_gain"] = round(float(ducking["idle"]), 4) else: audio["base_gain"] = 1.0 video_clip = { "source_path": c["source_path"], "source_start": round(float(c["source_start"]), 4), "source_end": round(float(c["source_end"]), 4), "timeline_start": round(ts, 4), "timeline_end": round(te, 4), "audio": audio, } for key in ( "chroma", "compound", "flip", "green_background", "lut", "mask", "opacity", "position", "reverse", "reverse_path", "rotation_degrees", "scale", "speed", "transition", ): if key in c: video_clip[key] = deepcopy(c[key]) video_clip_objs.append(video_clip) tracks = [{"kind": "video", "name": "video", "clips": video_clip_objs}] # --- narration track narr_segs = [] for s in placed: narration = { "source_path": s["source_path"], "timeline_start": _floor_time(s["timeline_start"], 4), "timeline_end": _ceil_time(s["timeline_end"], 4), "gain": round(float(s["gain"]), 4), "text": s["text"], "overlaps_speech": bool(s["overlaps_speech"]), } for key in ("source_duck_end", "source_restore_at", "source_handoff_status", "source_entry_status"): if key in s: narration[key] = deepcopy(s[key]) if "speed" in s: narration["speed"] = float(s["speed"]) narr_segs.append(narration) if narr_segs: tracks.append({"kind": "audio", "name": "narration", "role": "narration", "segments": narr_segs}) # --- bgm track (optional, looped, ducked under narration) if bgm: base = float(bgm["volume"]) duck = float(bgm["ducking_volume"]) fade = float(bgm["fade"]) kfs = ducking_keyframes(windows, base, duck, fade, 0.0, duration_s, bridge=ducking["bridge"] if ducking else None) tracks.append({ "kind": "audio", "name": "bgm", "role": "bgm", "loop": True, "segments": [{ "source_path": bgm["source_path"], "timeline_start": 0.0, "timeline_end": round(float(duration_s), 4), "gain": round(base, 4), "volume_keyframes": kfs, }], }) # --- subtitle (text) track: empty cues carry nothing to display text_source = subtitle_segments if subtitle_segments is not None else narration_segments text_segs = [] for s in text_source: if not s.get("text"): continue ts, te = float(s["timeline_start"]), float(s["timeline_end"]) if te <= ts: continue text_segment = { "text": s["text"], "timeline_start": round(ts, 4), "timeline_end": round(te, 4), } for key in ("flip", "opacity", "position", "rotation_degrees", "scale", "style", "style_id", "words"): if key in s: text_segment[key] = deepcopy(s[key]) text_segs.append(text_segment) if text_segs: tracks.append({"kind": "text", "name": "subtitle", "segments": text_segs}) # --- local image overlays (optional, timeline schema v2). Transform fields are # optional authoring extensions; the JianYing exporter validates and defaults them. images = [] for segment in image_segments: image = { "source_path": segment["source_path"], "timeline_start": round(float(segment["timeline_start"]), 4), "timeline_end": round(float(segment["timeline_end"]), 4), } for key in ("flip", "lut", "mask", "opacity", "position", "rotation_degrees", "scale", "speed", "transition"): if key in segment: image[key] = deepcopy(segment[key]) images.append(image) if images: tracks.append({"kind": "image", "name": "image", "segments": images}) tracks.extend(deepcopy(track) for track in extra_tracks) timeline = { "schema_version": SCHEMA_VERSION, "canvas": {"width": int(canvas["width"]), "height": int(canvas["height"]), "fps": float(canvas["fps"])}, "duration": round(float(duration_s), 4), "tracks": tracks, } if resource_packages: timeline["resource_packages"] = deepcopy(resource_packages) if style_presets: timeline["style_presets"] = deepcopy(style_presets) return timeline def save_timeline(timeline, path): with open(path, "w", encoding="utf-8") as f: json.dump(timeline, f, ensure_ascii=False, indent=2) return path def load_timeline(path): with open(path, encoding="utf-8") as f: return json.load(f) -
timeline_emit.py 7.7 KB
"""Backend-neutral timeline emission for the video-assemble skill.""" from pathlib import Path from audio_mix import _seg_place_window from lib import CONFIG, log from media import _build_video_clips from source_subtitles import _combined_subtitle_entries from timeline import build_timeline, save_timeline import packaging def _timeline_subtitle_segments(tts_segments, work_dir, duration_s): """Display-ready subtitle cues for timeline/export text tracks. The narration audio track keeps raw semantic text for editor reference; this payload mirrors SRT/ASS display policy, including terminal-punctuation cleanup and original-dialogue gap subtitles when configured. """ return [ { "text": entry["text"], "timeline_start": float(entry["start"]), "timeline_end": float(entry["end"]), } for entry in _combined_subtitle_entries(tts_segments, work_dir, duration_s) ] def _emit_timeline(input_video, tts_segments, work_dir, duration_s, canvas, has_bgm, *, audio_mode="narration", selected_audio_stream=0, explicit_audio_mix=None): """Build and persist the backend-neutral multi-track timeline.json.""" if audio_mode == "narration" and explicit_audio_mix is None: video_clips = _build_video_clips(input_video, work_dir, duration_s) else: # Non-narration sound comes from the actual current picture input as one # complete interval. Re-expanding an old cut plan would substitute different # source sound and make the optional editor project misrepresent the render. video_clips = [{ "source_path": str(Path(input_video)), "source_start": 0.0, "source_end": float(duration_s), "timeline_start": 0.0, "timeline_end": float(duration_s), }] # Mix segments are 1:1 with the UNFILTERED narration list, so they must be looked # up by their own index; a skipped beat would otherwise shift every later gain and # sample bound onto the wrong segment. mix_by_index = ( {item["index"]: item for item in explicit_audio_mix["segments"]} if explicit_audio_mix is not None else {} ) narration_segments = [] placed_indices = [] for seg in tts_segments: s, e = _seg_place_window(seg) if e <= s: continue narration_item = { # JianYing must consume the exact WAV written into narration.wav. In # particular, a tempo-adjusted beat cannot reference its longer pre-fit # source or the editor will trim its final words at timeline_end. "source_path": seg["placed_audio_path"], "timeline_start": s, "timeline_end": e, "text": seg["narration"], "overlaps_speech": seg["overlaps_speech"], "gain": ( mix_by_index[seg["index"]]["gain"] if explicit_audio_mix is not None else 1.0 ), } for key in ("source_duck_end", "source_restore_at", "source_handoff_status", "source_entry_status"): if key in seg: narration_item[key] = seg[key] narration_segments.append(narration_item) placed_indices.append(seg["index"]) fade = CONFIG["duck_fade_seconds"] bgm = None if has_bgm and explicit_audio_mix is None: bgm = {"source_path": CONFIG["bgm_path"], "volume": CONFIG["bgm_volume"], "ducking_volume": CONFIG["bgm_ducking_volume"], "fade": fade} # carry ducking automation whenever ducking is on at all; even under sidechain # mode the draft gets editable volume keyframes (ffmpeg stays the canonical mix) ducking = None if audio_mode == "narration" and explicit_audio_mix is None \ and CONFIG["ducking_mode"] != "none": ducking = {"idle": CONFIG["idle_orig_volume"], "speech": CONFIG["speech_ducking_volume"], "quiet": CONFIG["zone_ducking_volume"], "fade": fade, "bridge": CONFIG["duck_bridge_seconds"]} subtitle_segments = _timeline_subtitle_segments(tts_segments, work_dir, duration_s) timeline = build_timeline(canvas, duration_s, video_clips, narration_segments, bgm=bgm, ducking=ducking, subtitle_segments=subtitle_segments, image_segments=packaging.timeline_image_segments( packaging.load_packaging_layers(work_dir, canvas), canvas, duration_s)) if explicit_audio_mix is not None: for clip in timeline["tracks"][0]["clips"]: clip["audio"] = { "role": "picture_audio_not_consumed", "base_gain": 0.0, "volume_keyframes": [], } narration_track = next( (track for track in timeline["tracks"] if track.get("name") == "narration"), None ) if narration_track: for segment, index in zip(narration_track["segments"], placed_indices): adopted = mix_by_index[index] segment.update({ "gain": adopted["gain"], "output_start_sample": adopted["output_start_sample"], "output_end_sample": adopted["output_end_sample"], "sample_rate": 48_000, }) timeline["tracks"].append({ "kind": "audio", "name": "prepared_bed", "role": "prepared_bed", "segments": [{ "source_path": explicit_audio_mix["prepared"]["prepared_bed.wav"]["path"], "timeline_start": 0.0, "timeline_end": float(duration_s), "gain": 1.0, }], }) timeline["audio_delivery"] = { "mode": "explicit_adopted_full_sound", "sample_rate": 48_000, "total_samples": explicit_audio_mix["format"]["total_samples"], "master_gain_db": explicit_audio_mix["master_gain_db"], "canonical_renderer": "ffmpeg_explicit_mix", "reconstructable": False, "reconstructable_reason": ( "optional editor export does not implement the adopted 48 kHz " "prepared-bed, per-segment gain, and fixed-master chain" ), } if audio_mode != "narration": source_gain = 1.0 if audio_mode == "adopted-packet-copy" else CONFIG["idle_orig_volume"] for clip in timeline["tracks"][0]["clips"]: clip["audio"].update({ "base_gain": round(float(source_gain), 4), "selected_stream": selected_audio_stream, "mode": audio_mode, }) timeline["audio_delivery"] = { "mode": audio_mode, "selected_stream": selected_audio_stream, "packet_frozen": audio_mode == "adopted-packet-copy", "reconstructable": ( audio_mode == "adopted-packet-copy" and selected_audio_stream == 0 ), "canonical_renderer": "stream_copy" if audio_mode == "adopted-packet-copy" else "ffmpeg_mix", } degraded = [ {"source_path": clip["source_path"], "reason": clip["provenance_reason"]} for clip in video_clips if clip.get("provenance_degraded") ] if degraded: timeline["provenance"] = {"degraded": True, "degraded_clips": degraded} log(f" ⚠️ 时间线 provenance 降级: {degraded[0]['reason']} ({len(degraded)} clip)") else: timeline["provenance"] = {"degraded": False} out = Path(work_dir) / "timeline.json" save_timeline(timeline, out) log(f"时间线模型: {out} ({len(timeline['tracks'])} 轨)") return timeline -
visual_render.py 16.8 KB
"""Visual overlays, subtitle layout QC, masking, and video filter helpers.""" import json import re from pathlib import Path from assemble_constants import ( SUBTITLE_STYLE_REF_H, VISUAL_OVERLAYS, VISUAL_QC, _SUPPORTED_VISUAL_OVERLAY_TYPES, ) from audio_automation import coalesce_duck_windows from audio_mix import _seg_place_window from lib import CONFIG from source_subtitles import ( _combined_subtitle_entries, _original_gap_subtitle_entries, _source_subtitle_mask_policy, ) from subtitles.core import ( _measured_subtitle_band, _measured_subtitle_safe_area, _normalize_subtitle_text, _style_for_measured_subtitle_band, _subtitle_style_config, ) _OVERFLOW_KINDS = { "max_lines_exceeded": "line_count", "safe_width_exceeded": "line_width", "safe_height_exceeded": "safe_area", } def _visual_text_units(text): """Approximate visual text width in em units for deterministic geometry QC.""" units = 0.0 for ch in text: if ch.isspace(): units += 0.35 elif ord(ch) < 128: units += 0.56 else: units += 1.0 return units def _subtitle_layout_qc(entries, style, safe_area=None): """Machine-check subtitle safe-area/multiline/overflow facts for visual_qc.json.""" play_x = int(style["play_res_x"]) play_y = int(style["play_res_y"]) margin_l = int(style["margin_l"]) margin_r = int(style["margin_r"]) margin_v = int(style["margin_v"]) font_size = float(style["font_size"]) max_lines = CONFIG["subtitle_max_lines"] if safe_area is None: safe_area = { "x": margin_l, "y": margin_v, "width": max(1, play_x - margin_l - margin_r), "height": max(1, play_y - 2 * margin_v), "bottom_margin": margin_v, } usable_w = float(safe_area["width"]) line_h = font_size * 1.25 overflow_entries = [] violations = [] multi_line_entries = [] max_observed_lines = 0 entry_facts = [] for i, entry in enumerate(entries): raw_text = _normalize_subtitle_text(entry["text"]) lines = [ln for ln in re.split(r"(?:\\N|\n)+", raw_text) if ln != ""] or [""] line_count = len(lines) max_observed_lines = max(max_observed_lines, line_count) max_w = max(_visual_text_units(line) * font_size for line in lines) band_h = line_count * line_h + float(style["outline"]) * 2 + float(style["shadow"]) overflow_reasons = [] if line_count > max_lines: overflow_reasons.append("max_lines_exceeded") if max_w > usable_w + 1e-6: overflow_reasons.append("safe_width_exceeded") if band_h > safe_area["height"] + 1e-6: overflow_reasons.append("safe_height_exceeded") fact = { "index": i, "start": round(float(entry["start"]), 3), "end": round(float(entry["end"]), 3), "line_count": line_count, "max_line_width": round(max_w, 2), "safe_width": round(usable_w, 2), "band_height": round(band_h, 2), "overflow": bool(overflow_reasons), "overflow_reasons": overflow_reasons, } entry_facts.append(fact) if line_count > 1: multi_line_entries.append(i) if overflow_reasons: overflow_entries.append(fact) violations.extend( {"index": i, "kind": _OVERFLOW_KINDS[reason], "reason": reason} for reason in overflow_reasons ) return { "enabled": CONFIG["burn_subtitles"], "renderer": "ass" if CONFIG["burn_subtitles"] else "sidecar_srt", "style": { "font_size": int(font_size), "max_chars": int(style["max_chars"]), "max_lines": max_lines, "play_res_x": play_x, "play_res_y": play_y, "alignment": int(style["alignment"]), "margin_l": margin_l, "margin_r": margin_r, "margin_v": margin_v, }, "safe_area": safe_area, "entries": len(entry_facts), "max_lines": max_observed_lines, "max_observed_lines": max_observed_lines, "multi_line": bool(multi_line_entries), "multi_line_entries": multi_line_entries, "overflow": bool(overflow_entries), "overflow_entries": overflow_entries, "violations": violations, "entry_facts": entry_facts, } def _load_visual_overlays(work_dir): """Return (overlays, source) from the canonical visual_overlays.json handoff.""" path = Path(work_dir) / VISUAL_OVERLAYS if not path.exists(): return [], {"present": False, "path": str(path)} try: data = json.loads(path.read_text(encoding="utf-8")) except json.JSONDecodeError as exc: raise ValueError(f"{path}: JSON 无效") from exc if ( not isinstance(data, dict) or type(data.get("schema_version")) is not int or data["schema_version"] != 1 or not isinstance(data.get("overlays"), list) or not all(isinstance(item, dict) for item in data["overlays"]) ): raise ValueError(f"{path}: visual_overlays.json schema 无效") source = { "present": True, "path": str(path), "schema_version": 1, } return data["overlays"], source def _escape_drawtext_text(text): return ( text .replace("\\", "\\\\") .replace(":", "\\:") .replace("'", "\\'") .replace("%", "\\%") .replace("\n", "\\n") ) def _overlay_time_window(overlay, video_duration): start = float(overlay.get("start", 0.0)) end = float(overlay.get("end", video_duration)) return start, max(start, end) def _overlay_bbox(overlay, canvas, *, default_y): width = canvas["width"] height = canvas["height"] text = overlay["text"] font_size = int(overlay.get("font_size", max(18, round(height * 0.045)))) lines = [ln for ln in text.splitlines() if ln.strip()] or [text] max_w = max(_visual_text_units(ln) * font_size for ln in lines) text_h = len(lines) * font_size * 1.25 if overlay["type"] == "top_title": x = max(0.0, (width - max_w) / 2) y = float(overlay.get("y", default_y)) else: # Fractions of the canvas in [0, 1] are normalized coordinates; larger values are pixels. x = float(overlay.get("x", 0.08)) y = float(overlay.get("y", 0.25)) if 0.0 <= x <= 1.0: x *= width if 0.0 <= y <= 1.0: y *= height return { "x": round(x, 2), "y": round(y, 2), "width": round(max_w, 2), "height": round(text_h, 2), "font_size": font_size, "line_count": len(lines), "overflow": x < 0 or y < 0 or x + max_w > width or y + text_h > height, } def _visual_overlay_filters(work_dir, canvas, video_duration): """Render the first-release canonical visual_overlays.json contract. Only two semantic renderers are supported: top_title and inline_label_or_callout. Unsupported types are QC-blocking and deliberately do not silently render. """ overlays, source = _load_visual_overlays(work_dir) default_top_y = max(24, round(canvas["height"] * 0.05)) filters = [] facts = [] unsupported = [] overflow = [] for idx, overlay in enumerate(overlays): typ = overlay.get("type") text = overlay.get("text", "").strip() if typ not in _SUPPORTED_VISUAL_OVERLAY_TYPES: unsupported.append({"index": idx, "type": typ, "reason": "unsupported_overlay_type"}) continue if not text: unsupported.append({"index": idx, "type": typ, "reason": "missing_text"}) continue start, end = _overlay_time_window(overlay, video_duration) bbox = _overlay_bbox(overlay, canvas, default_y=default_top_y) if bbox["overflow"]: overflow.append({"index": idx, "type": typ, "bbox": bbox}) font_size = bbox["font_size"] safe_text = _escape_drawtext_text(text) enable = f"between(t\\,{start:.3f}\\,{end:.3f})" if typ == "top_title": filt = ( "drawtext=" f"{_drawtext_font_option()}text='{safe_text}':x=(w-text_w)/2:y={int(bbox['y'])}:" f"fontsize={font_size}:fontcolor=white:borderw=2:bordercolor=black@0.85:" f"box=1:boxcolor=black@0.35:boxborderw=12:enable='{enable}'" ) else: filt = ( "drawtext=" f"{_drawtext_font_option()}text='{safe_text}':x={int(bbox['x'])}:y={int(bbox['y'])}:" f"fontsize={font_size}:fontcolor=white:borderw=2:bordercolor=black@0.85:" f"box=1:boxcolor=black@0.45:boxborderw=8:enable='{enable}'" ) filters.append(filt) facts.append({ "index": idx, "type": typ, "text_chars": len(text), "start": round(start, 3), "end": round(end, 3), "bbox": bbox, }) qc = { "source": source, "supported_types": sorted(_SUPPORTED_VISUAL_OVERLAY_TYPES), "present": source["present"], "count": len(overlays), "rendered": len(facts), "facts": facts, "unsupported": unsupported, "overflow": overflow, } return filters, qc def _build_visual_qc(tts_segments, work_dir, video_duration, canvas, *, overlay_qc=None, mask_filter=None): entries = _combined_subtitle_entries(tts_segments, work_dir, video_duration) style = _style_for_measured_subtitle_band(_subtitle_style_config(canvas), canvas) subtitle_layout = _subtitle_layout_qc( entries, style, safe_area=_measured_subtitle_safe_area(style, canvas) ) mask = _source_subtitle_mask_policy(work_dir) mask.update({ "ratio": min(0.5, CONFIG["source_subtitle_mask_ratio"]) if mask["active"] else None, "filter": "drawbox" if mask_filter else None, "opacity": CONFIG["subtitle_mask_opacity"], "timing": CONFIG["source_subtitle_mask_timing"], "subtitle_y_top": CONFIG["subtitle_y_top"], "subtitle_y_bot": CONFIG["subtitle_y_bot"], }) if overlay_qc is None: overlay_qc = _visual_overlay_filters(work_dir, canvas, video_duration)[1] blocking_codes = [] if mask["blocking"]: blocking_codes.append("mask_policy_not_explicit") if subtitle_layout["overflow"]: blocking_codes.append("subtitle_overflow") if overlay_qc["unsupported"]: blocking_codes.append("unsupported_visual_overlay") if overlay_qc["overflow"]: blocking_codes.append("visual_overlay_overflow") return { "schema_version": 1, "artifact": VISUAL_QC, "verdict": "FAIL" if blocking_codes else "PASS", "blocking": bool(blocking_codes), "blocking_codes": blocking_codes, "geometry": { "canvas": { "width": canvas["width"], "height": canvas["height"], "fps": canvas["fps"], }, "storage": { "width": canvas["storage_width"], "height": canvas["storage_height"], }, "rotation": canvas["rotation"], "sample_aspect_ratio": canvas["sample_aspect_ratio"], "display_aspect_ratio": canvas["display_aspect_ratio"], }, "subtitles": subtitle_layout, "mask": mask, "overlays": overlay_qc, "summary": { "subtitle_entries": subtitle_layout["entries"], "subtitle_overflow": subtitle_layout["overflow"], "subtitle_multi_line": subtitle_layout["multi_line"], "mask_policy": mask["policy"], "mask_active": mask["active"], "overlay_rendered": overlay_qc["rendered"], "overlay_unsupported": len(overlay_qc["unsupported"]), }, } def _write_visual_qc(work_dir, qc): path = Path(work_dir) / VISUAL_QC path.write_text(json.dumps(qc, ensure_ascii=False, indent=2), encoding="utf-8") return path def _escape_subtitle_filter_path(path): """Escape a path for ffmpeg subtitle/ass video filter arguments.""" text = str(path).replace("\\", "/") for raw, escaped in ( ("\\", "\\\\"), (":", "\\:"), ("'", "\\'"), (",", "\\,"), ("[", "\\["), ("]", "\\]"), ): text = text.replace(raw, escaped) return text def _subtitle_burn_filter(subtitle_path): """Build the ffmpeg video filter used for hard-sub rendering. With SUBTITLE_FONT_FILE set, libass also loads the fonts in that file's directory so the ASS style's family name resolves to the declared file instead of a system font. """ filt = f"subtitles=filename='{_escape_subtitle_filter_path(subtitle_path)}'" font_file = CONFIG["subtitle_font_file"] if font_file: fonts_dir = Path(font_file).expanduser().resolve().parent filt += f":fontsdir='{_escape_subtitle_filter_path(fonts_dir)}'" return filt def _drawtext_font_option(): font_file = CONFIG["subtitle_font_file"] if not font_file: return "" return f"fontfile='{_escape_subtitle_filter_path(Path(font_file).expanduser().resolve())}':" def _output_downscale_filter(max_h): """Lanczos downscale that forces BOTH output dimensions even (libx264/yuv420p need it). -2 keeps the aspect ratio with an even width; 2*trunc(min(ih,H)/2) caps the height at H yet forces it even, so an odd OUTPUT_MAX_HEIGHT (e.g. 721) cannot produce an odd height that makes libx264 abort with an empty output file. 'min(ih,H)' only ever shrinks. """ return f"scale=-2:'2*trunc(min(ih,{max_h})/2)':flags=lanczos" def _source_subtitle_mask_filter(canvas, work_dir, tts_segments, video_duration): """Return source-subtitle drawbox filters, optionally scoped to narration windows. Many source videos (e.g. 庆余年) ship hardcoded subtitles; without this the recap shows the original subs AND our narration subs stacked. Once masking is explicitly enabled, the enhanced default is a measured, translucent narration-only band; opacity and timing remain configurable. """ policy = _source_subtitle_mask_policy(work_dir) if not policy["active"]: return None opacity = CONFIG["subtitle_mask_opacity"] timing = CONFIG["source_subtitle_mask_timing"] if timing not in {"all", "narration"}: raise ValueError(f"SOURCE_SUBTITLE_MASK_TIMING 必须是 all 或 narration,当前为 {timing!r}") band = _measured_subtitle_band(canvas) if band is not None: y_top, y_bot = band padding = CONFIG["subtitle_mask_padding"] mask_top = max(0, y_top - padding) mask_bot = min(canvas["height"], y_bot + padding) geometry = f"x=0:y={mask_top}:w=iw:h={mask_bot - mask_top}" else: # Our subtitle cues are one line. Keep the mask large enough for that line and its # margin, but never regress to the old two-line bar that hid ~23% of the image. style = _subtitle_style_config(canvas) play_res_y = float(style["play_res_y"]) line_h = float(style["font_size"]) * 1.25 pad = 10.0 * play_res_y / SUBTITLE_STYLE_REF_H sub_band = (float(style["margin_v"]) + line_h + pad) / play_res_y ratio = min(0.5, max(CONFIG["source_subtitle_mask_ratio"], sub_band)) geometry = f"x=0:y=ih-ih*{ratio:.3f}:w=iw:h=ih*{ratio:.3f}" base = f"drawbox={geometry}:color=black@{opacity:.2f}:t=fill" filters = [] if opacity > 0: if timing == "all": filters.append(base) else: windows = [ (start, end, 0.0) for start, end in map(_seg_place_window, tts_segments) if end > start ] # Avoid overlapping drawboxes: stacking two 60%-black masks would darken the # overlap to 84%. Coalescing also keeps long filter chains smaller. filters.extend( f"{base}:enable='between(t,{start:.3f},{end:.3f})'" for start, end, _ in coalesce_duck_windows(windows, bridge=0.001) ) # A translucent mask deliberately leaves the source glyphs visible. Whenever we burn a # replacement original-dialogue subtitle into a gap, cover that exact window opaquely first; # otherwise the source hard-sub and replacement text are stacked on top of each other. if not (timing == "all" and opacity >= 1.0 - 1e-9): replacement_windows = [ (entry["start"], entry["end"], 0.0) for entry in _original_gap_subtitle_entries(tts_segments, work_dir, video_duration) ] opaque = f"drawbox={geometry}:color=black@1.00:t=fill" filters.extend( f"{opaque}:enable='between(t,{start:.3f},{end:.3f})'" for start, end, _ in coalesce_duck_windows(replacement_windows, bridge=0.001) ) return ",".join(filters) if filters else None
-
-
SKILL.md 8.7 KB
--- name: video-assemble user-invocable: false description: > 合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录, 最后做响度标准化。作为最终合成阶段使用。输入源视频、tts_meta.json 与旁白位置; 输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。 --- ## 1. 定位 本技能负责最终合成: 1. 把各段旁白音频放到视频时间线上。 2. 在旁白窗口内压低原声,支持 fixed / sidechain / zone 模式。 3. 根据旁白位置生成 `subtitles.srt`;默认同时生成并烧录 `subtitles.ass`,`--no-burn-subtitles` 可关闭。 4. 可选把最终响度标准化到目标 LUFS。 ## 2. 声音收尾契约 合成阶段只实现创作决定,不凭空制造决定。Agent 在写旁白位置前,已在 `visual_audio_board.json` 为每个 beat 指定 `audio_owner`: - `original_dialogue` - `action_sound` - `ambience` / `music` - `silence` - `narration` 因此,旁白间隙是主动选择,不是必须填满的空白。不要为了“更满”而加入通用 BGM、压住必须听见的台词或消除有意义的沉默。 当前渲染器不解析 `visual_audio_board.json`;Agent 通过旁白时间、`overlaps_speech`、原声留白与现有混音参数落实这些决定。 ## 3. 输入契约 - `<video>`:源视频;cut 模式下为 `edited_source.mp4`。 - `work_dir/tts_meta.json`:默认 `narration` 模式必需;配音阶段写出的 `{segments: [...]}`。每段包含 `audio_path`、时间、`pause_after_ms`、`overlaps_speech` 和用于混音/字幕的位置。显式 `source-mix` / `adopted-packet-copy` 模式不读取它。 - 已采用的配音使用显式 `--tts-meta` 和 `--narration-adoption`:后者由调用方独立确认文字、请求的引擎/声线和速度策略,不能从待消费元数据自动“批准”出来。完整格式与记录边界见 `references/narration-adoption.md`。 - 已采用的完整声音底轨与逐段配音可再传 `--audio-mix-adoption`;严格格式、48 kHz 声道矩阵和双 binding 事务见 `references/explicit-audio-mix.md`。 下面的 `scripts/...` 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。 ## 4. 运行命令 ```bash python3 scripts/assemble.py <video> --work-dir <work_dir> \ [--audio-mode narration|source-mix|adopted-packet-copy] [--audio-stream-index <N>] \ [--tts-meta <tts_meta.json> --narration-adoption <narration_adoption.json>] \ [--audio-mix-adoption <audio_mix_adoption.json>] \ [--recap-stem <name>] [--output-dir <dir>] [--no-burn-subtitles] \ [--subtitle-y-top <inclusive-y> --subtitle-y-bot <exclusive-y>] \ [--source-video <orig.mp4>] [--export-jianying [--jianying-out <dir>]] ``` ## 5. 输出契约 - `recap_<stem>.mp4`:稳定的最终输出别名;每次运行覆盖更新。 - `work_dir/output.mp4`:工作目录内成片。 - `subtitles.srt`:旁白字幕;烧录时另有 `subtitles.ass`。 - `timeline.json`:后端无关的多轨模型,包含视频、原声、旁白、BGM、字幕和 ducking 自动化。 - `_placed_*.wav`:实际写入主混音的完整逐段旁白 PCM;时间线与剪映只引用这些文件。 - `narration_input_binding.json`:旁白输入、转换、实际放置、旁白总轨和最终音轨的消费记录(路径、PCM 参数、packet 计数)。区分未采用与已绑定采用决定两种状态;不等于声线鉴定或听审。 - `audio_mix_binding.json`:显式完整声音分支消费的画面时钟、底轨、48 kHz 配音放置、premaster、固定 master gain、最终 PCM/AAC 事实与 narration binding 路径的记录。 - `assembly_manifest.json`:输入来源、cut 来源标识(路径、大小、mtime)、渲染设置与最终输出路径。 - `assembly_qc.json`:旁白完整性、原声句末交接、时间线素材时长与交付质量的发布门禁。 - 剪映草稿目录:仅 `--export-jianying` 时生成,包含 `draft_content.json`、`draft_info.json` 与 `draft_meta_info.json`。 ## 6. 合成规则 - 音频模式的处理与冻结语义见 `references/audio-modes.md`。默认仍为 `narration`;另外两种模式必须显式选择。 - `--audio-mix-adoption` 只与显式 `--tts-meta`、`--narration-adoption` 同时使用;它保留 `narration` 模式名,但跳过旧速度/适配、原声 handoff、环境 BGM、duck、loudnorm 和 limiter。 - 音频按轨道混合:原声、可选 BGM 与旁白各自独立。 - 旁白不做任何容差裁尾;温和加速后仍放不下即 `no_safe_fit`。每段 `_placed_*.wav` 必须与序列化后的时间线区间等长或更短,否则 `timeline_audio_mismatch` 阻断。 - 已采用配音的 v1 合同只支持原速、禁止段内适速;不能让环境默认 1.15 倍速或旧缓存覆盖它。放不下就修订安排,不裁尾。严格运行使用新工作目录与新输出路径;输入/实际混音来源变动或 QC 失败时,不发布候选成片。没有采用文件的旧入口仍是兼容模式,不自动获得同等证据。 - 原声在旁白结束后保持压低到下一可靠句末的 `pause_start`,只在实测停顿内渐强, 于 `source_restore_at` 回满;无后续锚点时保持压低到时间线末端,而不是放出半句。 - `--export-jianying` / `EXPORT_JIANYING=1` 可把 `timeline.json` 导出为可编辑草稿。cut 模式应传 `--source-video <orig>`,让草稿引用真实原片区间。 - 剪映导出默认把视频、音频与图片复制到 `Resources/local/{video,audio,image}`,保持草稿可搬迁;`--jianying-no-bundle-media` 只适合原路径始终可访问的情况。 - 重叠覆盖物会拆到编号轨道;非空目标目录不会覆盖,而会创建编号兄弟目录。 - 常速、倒放、变换、富文本、转场、蒙版、LUT、绿幕复合草稿及显式特效轨道通过 timeline v2 扩展表达。需要素材包的功能只接受调用方合法提供的离线资源。 - 剪映草稿引用未烧录的源视频,因此原片硬字幕仍会保留,必要时在剪映内另行遮罩。 - 字幕外观可用 `SUBTITLE_FONT_SIZE`、`SUBTITLE_MARGIN_V`、`SUBTITLE_MAX_CHARS` 等控制。 - `SUBTITLE_Y_TOP/BOT` 把 ASS 基线放到测得的原片字幕区域,坐标为半开 `[top, bot)`;显式遮罩策略下默认 `SUBTITLE_MASK_OPACITY=0.6`,`SOURCE_SUBTITLE_MASK_TIMING=narration`。 - 原声在旁白间隙回到 `IDLE_ORIG_VOLUME`,旁白下压到 `SPEECH_DUCKING_VOLUME`;`DUCK_FADE_SECONDS` 控制过渡。还可配置 `DUCKING_MODE`、`ZONE_DUCKING_VOLUME`、`FINAL_LOUDNORM` 与 `TARGET_LUFS`。 - 可通过 `BGM_PATH` 指定 BGM;它会循环到成片长度,并按 `BGM_VOLUME` / `BGM_DUCKING_VOLUME` 混音。不要在没有创作依据时设置通用 BGM。 - 烧录字幕需要带 `subtitles` / libass 的 ffmpeg;合成阶段会预检并在缺失时明确失败。 - 原声留白中的对白字幕优先读取 Agent 校对的 `original_subtitles.json`;否则保守映射 ASR。只有遮罩覆盖留白或用户字幕明确要求替换时才烧录原声对白,并用 `「」` 与旁白区分。 ### 按原片区间准备声音,而不是整体压低旧成片 已有多段原声取舍和独立 BGM 决定时,先用 `references/source-score.md` 的独立 `source_score.py` 从原片声音流按精确帧区间重建原声轨、音乐轨及两者之和;它只输出 声音底轨和来源回执。要与逐段已采用配音合成,再由调用方提供 `references/explicit-audio-mix.md` 的严格 adoption;不要将底轨塞入旧入口自动 duck,也不要从含旧解说的成片取整条声音冒充干净原声。 ## 7. 字幕与可选包装 先锁定画面、剪点、旁白和混音,再投入字幕动画或边框包装;字幕样式不能掩盖叙事、剪点或声音问题。 普通交付优先使用现有 ASS 路径;只有用户需要逐 cue 排版、动画或透明图层时,才用项目级代码渲染器, 并按 `references/foreground-compose.md` 把它生成的 RGBA 序列叠到锁定母版。包装顺序与样帧抽检清单见 `references/packaging.md`。 ## 8. 能力边界 - 不生成旁白文字,不合成 TTS,不重新转写视频。 - 字幕烧录默认开启;关闭时不会重编码绘制字幕区域。 显式输出轴字幕轨的独立合同、完整替换语义和当前边界见 `references/subtitle-track.md`。 画面回原片重建后,若需保留另一文件中的已采用完整混音,先按 `references/pair-media.md` 显式配对独立画面与音轨。配对只复制流,不补字幕或片名卡; 后续字幕轨必须重新绑定配对后的容器与 `a:0`,不能继续沿用旧版本身份。
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.