ffmpeg-ops
Comprehensive ffmpeg/ffprobe media processing: transcode, cut/trim/concat, color grading, loudness normalization, subtitles, GIFs, HLS packaging, hardware encoding, and quality gates (VMAF). Triggers on: ffmpeg, ffprobe, transcode, compress/convert video, extract audio, color gra
Install
npx skills add https://github.com/0xDarkMatter/claude-mods/tree/main/skills/ffmpeg-ops
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install 0xdarkmatter-claude-mods@llmmart
git clone https://github.com/0xDarkMatter/claude-mods.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole 0xdarkmatter/claude-mods collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
ffmpeg Operations
Operational expertise for ffmpeg/ffprobe: the ~30 commands that cover most real work, the footguns that silently ruin output, EDL-driven editing (edit-as-code), and eight scripts that replace the logic an agent would otherwise re-derive every task.
Doctrine: probe first
Never transcode, cut, or filter blind. Every media task starts by probing the input — codec, duration, frame rate (constant or variable?), pixel format, rotation, stream layout. Half of all "ffmpeg did something weird" reports are a property of the input the command never checked.
python skills/ffmpeg-ops/scripts/probe-media.py input.mp4 # human summary
python skills/ffmpeg-ops/scripts/probe-media.py --doctor input.mp4 # TRIAGE: hazards + exact fixes
python skills/ffmpeg-ops/scripts/probe-media.py --json input.mp4 | jq '.data.streams'
python skills/ffmpeg-ops/scripts/probe-media.py --keyframes-near 92.5 input.mp4
--doctor makes the doctrine self-enforcing: VFR, HDR transfer, rotation
metadata, interlacing, non-yuv420p delivery, and moov-at-EOF each come back as a
finding with the exact fix command, and exit 10 means "fix before processing".
The --keyframes-near form answers "can I stream-copy a cut at 92.5s?" — it
reports the nearest keyframes so you know whether a copy cut will snap (see
Footguns). When a command fails with a cryptic message, decode it:
references/error-decoder.md.
Before recommending an encoder, verify the build has it. Installed ffmpeg builds vary wildly (especially hardware encoders — listed ≠ working):
bash skills/ffmpeg-ops/scripts/capability-scan.sh # full: proof-encodes each hw encoder
bash skills/ffmpeg-ops/scripts/capability-scan.sh --quick # list-only, no GPU touch
Cookbook
Commands are bash-form; they run unchanged in PowerShell except where the
Windows notes say otherwise. Replace -y/-n (overwrite/never)
consciously — never leave an agent-run command interactive.
Convert and compress
# Web-compatible H.264 — THE default delivery encode. yuv420p + faststart are not
# optional: without them Safari/QuickTime/old devices show black video, and the
# moov atom sits at EOF so browsers can't start playback until fully downloaded.
ffmpeg -i in.mov -c:v libx264 -crf 20 -preset slow -pix_fmt yuv420p \
-c:a aac -b:a 192k -movflags +faststart out.mp4
# H.265/HEVC — ~40% smaller at same quality, slower encode, less universal playback.
# -tag:v hvc1 is required for Apple players to recognize the stream.
ffmpeg -i in.mp4 -c:v libx265 -crf 24 -preset slow -tag:v hvc1 \
-c:a copy -movflags +faststart out.mp4
# AV1 via SVT-AV1 (libaom is 10-50x slower; only use it for research-grade encodes).
# preset 0-13: lower = slower/better; 6 is the quality/speed sweet spot.
ffmpeg -i in.mp4 -c:v libsvtav1 -crf 32 -preset 6 -c:a libopus -b:a 128k out.webm
# Remux only — change container, zero quality loss, near-instant. Try this FIRST
# when the ask is "make this .mkv play in X": often the codecs are fine.
ffmpeg -i in.mkv -c copy -movflags +faststart out.mp4
# Normalize a problem source (HEVC/VFR phone footage, Zoom/Loom exports) before ANY
# downstream editing. VFR breaks cut math, concat sync, and Remotion/player seeking.
ffmpeg -i in.mov -c:v libx264 -crf 18 -preset fast -pix_fmt yuv420p \
-fps_mode cfr -r 30 -c:a aac -b:a 192k normalized.mp4
# Archival master — FFV1 lossless in MKV (the preservation standard).
ffmpeg -i in.mp4 -c:v ffv1 -level 3 -g 1 -slicecrc 1 -c:a flac archive.mkv
# "Make it fit in 25MB" — computed two-pass bitrate, auto audio/downscale, VERIFIED:
python skills/ffmpeg-ops/scripts/smart-compress.py --target 25MB video.mp4
Codec choice, CRF/preset matrices, two-pass bitrate targeting, per-platform social targets: references/encoding.md + assets/encoding-presets.json.
Cut and join
# Fast lossless trim (stream copy). -ss/-to BEFORE -i = input seek, absolute times.
# CAVEAT: with -c copy the start snaps to the previous keyframe — can be seconds
# early, or give frozen/black lead-in. Check first with probe-media.py --keyframes-near.
ffmpeg -ss 00:01:30 -to 00:02:00 -i in.mp4 -c copy -avoid_negative_ts make_zero cut.mp4
# Frame-accurate trim (re-encode). Input-side -ss IS frame-accurate when re-encoding
# (ffmpeg decodes from the prior keyframe and discards) — fast AND exact. The old
# "put -ss after -i for accuracy" advice costs a full decode from 0:00 for nothing.
ffmpeg -ss 00:01:30 -to 00:02:00 -i in.mp4 -c:v libx264 -crf 18 -c:a aac cut.mp4
# Join files with IDENTICAL codec/params — concat demuxer, no re-encode.
printf "file '%s'\n" seg1.mp4 seg2.mp4 seg3.mp4 > concat.txt
ffmpeg -f concat -safe 0 -i concat.txt -c copy joined.mp4
# Join files with DIFFERENT codecs/sizes — concat filter, re-encodes.
ffmpeg -i a.mp4 -i b.mov -filter_complex \
"[0:v][0:a][1:v][1:a]concat=n=2:v=1:a=1[v][a]" \
-map "[v]" -map "[a]" -c:v libx264 -crf 20 -c:a aac joined.mp4
# Remove a middle segment (keep 0-60s and 120s-end): cut both keeps, then concat.
# For multi-cut edits, write an EDL and use cut-from-edl.py instead (see EDL workflow).
-ss semantics in full, keyframe theory, concat ×3 (demuxer/filter/protocol),
edit-decision-list editing: references/trim-concat.md
and references/edit-as-code.md.
Resize, transform, retime
# Resize to width, keep aspect. ALWAYS -2 (not -1): yuv420p needs even dimensions.
ffmpeg -i in.mp4 -vf "scale=1280:-2" -c:a copy out.mp4
# Crop (w:h:x:y from top-left); cropdetect finds black bars for you:
ffmpeg -i in.mp4 -vf cropdetect -frames:v 120 -f null - 2>&1 | rg crop=
ffmpeg -i in.mp4 -vf "crop=1920:800:0:140" -c:a copy out.mp4
# Vertical 9:16 from landscape — blurred-pad pattern (social standard):
ffmpeg -i in.mp4 -filter_complex \
"[0:v]scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,boxblur=20[bg];
[0:v]scale=1080:-2[fg];[bg][fg]overlay=(W-w)/2:(H-h)/2" -c:a copy vertical.mp4
# Rotate: fix metadata only (instant) vs bake pixels (re-encode).
ffmpeg -display_rotation 90 -i in.mp4 -c copy out.mp4 # metadata flip (ffmpeg 6+)
ffmpeg -i in.mp4 -vf "transpose=1" -c:a copy out.mp4 # transpose=1: 90° clockwise
# Frame-rate change (drops/dups frames; for smooth slow-mo see minterpolate below)
ffmpeg -i in.mp4 -vf "fps=30" -c:a copy out.mp4
# 2x speed-up: video PTS halved + audio atempo (atempo accepts 0.5-100; chain
# atempo=0.5,atempo=0.5 for 0.25x). -map ordering keeps streams paired.
ffmpeg -i in.mp4 -filter_complex \
"[0:v]setpts=0.5*PTS[v];[0:a]atempo=2.0[a]" -map "[v]" -map "[a]" fast.mp4
# Interpolated slow-mo (synthesizes in-between frames — slow but smooth):
ffmpeg -i in.mp4 -vf "minterpolate=fps=60:mi_mode=mci:mc_mode=aobmc,setpts=2*PTS" -an slow.mp4
# Timelapse from photos (and the reverse: video -> frames, under Images below)
ffmpeg -framerate 24 -pattern_type glob -i 'photos/*.jpg' \
-c:v libx264 -crf 20 -pix_fmt yuv420p timelapse.mp4
Filtergraph syntax (labels, chains, split), speed ramps, full filter cookbook: references/filtergraph.md.
Overlay, text, subtitles
# Watermark bottom-right with 24px margin (W/H = video, w/h = overlay dims):
ffmpeg -i in.mp4 -i logo.png -filter_complex \
"overlay=W-w-24:H-h-24:format=auto" -c:a copy out.mp4
# Burn a running timecode (note %{pts\:hms} — the colon must be escaped INSIDE
# the drawtext argument; see Windows notes for fontfile paths):
ffmpeg -i in.mp4 -vf \
"drawtext=text='%{pts\:hms}':fontsize=48:fontcolor=white:box=1:boxcolor=black@0.5:x=24:y=24" \
-c:a copy out.mp4
# Burn-in subtitles (hard subs; needs libass). Pragmatic path rule: cd to the
# subtitle's directory and use a bare relative filename — the filter's path
# escaping is the single worst quoting trap in ffmpeg, especially on Windows.
ffmpeg -i in.mp4 -vf "subtitles=subs.srt" -c:a copy burned.mp4
# Soft subtitles (toggleable, instant — no re-encode):
ffmpeg -i in.mp4 -i subs.srt -map 0 -map 1 -c copy -c:s mov_text soft.mp4 # mp4
ffmpeg -i in.mkv -i subs.srt -map 0 -map 1 -c copy -c:s srt soft.mkv # mkv
Styling (ASS force_style), extraction, format conversion, STT round-trip: references/subtitles.md.
Audio
# Extract audio without re-encoding (copy the stream as-is; pick the container
# matching the codec — probe first: aac->.m4a, opus->.opus/.ogg, mp3->.mp3):
ffmpeg -i in.mp4 -vn -c:a copy out.m4a
# Extract + transcode to Opus (best codec per bit: voice 24-32k mono, music 96-128k):
ffmpeg -i in.mp4 -vn -c:a libopus -b:a 128k out.opus
# Replace a video's audio track (keep video untouched):
ffmpeg -i video.mp4 -i music.m4a -map 0:v -map 1:a -c:v copy -c:a aac -shortest out.mp4
# Mix two audio inputs (normalize=0 stops amix halving the volume of each input):
ffmpeg -i voice.wav -i music.mp3 -filter_complex \
"[1:a]volume=0.25[m];[0:a][m]amix=inputs=2:duration=first:normalize=0[a]" \
-map "[a]" -c:a aac mixed.m4a
# Loudness-normalize, one-pass (quick; DYNAMIC mode — fine for drafts).
# Two-pass linear mode is measurably better: use loudnorm-scan.py (Scripts below).
# loudnorm internally upsamples to 192kHz — the -ar 48000 puts it back.
ffmpeg -i in.mp4 -af "loudnorm=I=-16:TP=-1.5:LRA=11" -ar 48000 -c:v copy out.mp4
# Trim leading/trailing silence:
ffmpeg -i in.wav -af \
"silenceremove=start_periods=1:start_threshold=-40dB:detection=peak,areverse,silenceremove=start_periods=1:start_threshold=-40dB:detection=peak,areverse" \
trimmed.wav
Targets: -14 LUFS streaming platforms, -16 podcasts, -23 EBU R128 broadcast. Channel mapping, multi-track, restoration filters: references/audio.md.
Speech-to-text prep (Whisper-family)
# THE canonical STT extraction — 16 kHz mono 16-bit PCM (what whisper.cpp /
# faster-whisper actually resample to; doing it here is faster and deterministic):
ffmpeg -i in.mp4 -vn -ac 1 -ar 16000 -c:a pcm_s16le stt.wav
# Pipe raw PCM straight to whisper.cpp — no temp file:
ffmpeg -v error -i in.mp4 -vn -ac 1 -ar 16000 -f s16le - | whisper-cli -m model.bin -f -
# Chunk long audio ON SILENCE BOUNDARIES (never mid-word) for parallel transcription:
python skills/ffmpeg-ops/scripts/detect-segments.py --silence --json in.mp4 \
| jq '.data.speech[]'
Pre-STT cleanup (when afftdn/highpass help vs hurt accuracy), WhisperX word-level
alignment (±50 ms), transcript JSON shape, the summarisation pipeline:
references/stt-whisper.md.
Images, GIFs, frames
# Thumbnail at a timestamp (input-side -ss: instant even at 2h offsets):
ffmpeg -ss 00:00:05 -i in.mp4 -frames:v 1 -q:v 2 thumb.jpg
# Contact sheet: 1 frame every 10s, tiled 4x3 (visual summary / scrub preview):
ffmpeg -i in.mp4 -vf "fps=1/10,scale=320:-2,tile=4x3" -frames:v 1 sheet.png
# High-quality GIF — palettegen/paletteuse is THE difference between a 256-color
# dithered mess and a clean GIF. Single pass via split:
ffmpeg -ss 5 -to 8 -i in.mp4 -filter_complex \
"fps=12,scale=480:-1:flags=lanczos,split[s0][s1];[s0]palettegen=max_colors=128[p];[s1][p]paletteuse=dither=bayer:bayer_scale=4" \
out.gif
# Embedded chapters from scene/silence detection (or YouTube description text):
python skills/ffmpeg-ops/scripts/make-chapters.py --from-scenes --media talk.mp4 \
--min-gap 30 --write chaptered.mp4
python skills/ffmpeg-ops/scripts/make-chapters.py --from-silence --media lecture.mp4 \
--format youtube
# Frames for ML datasets — fixed fps, model-square crop:
ffmpeg -i in.mp4 -vf "fps=1,scale=512:512:force_original_aspect_ratio=increase,crop=512:512" \
frames/%06d.png
# Image sequence -> video:
ffmpeg -framerate 24 -i frames/%06d.png -c:v libx264 -crf 18 -pix_fmt yuv420p out.mp4
# Player scrub-preview sprites + the WebVTT thumbnail track that maps them:
python skills/ffmpeg-ops/scripts/make-sprites.py --interval 5 video.mp4
Sprite sheets for web players, AVIF/WebP stills, dataset prep patterns: references/images-gif.md.
Diagnostics and validation
# Corruption / decode-error check (exit code is NOT the signal — the log is):
ffmpeg -v error -i in.mp4 -f null - 2> errors.log && [ ! -s errors.log ] && echo CLEAN
# Per-frame hashes — prove two pipelines produce identical frames:
ffmpeg -i in.mp4 -map 0:v -f framemd5 -
# Strip ALL metadata (GPS, device info — privacy before sharing phone video).
# -map_metadata -1 keeps rotation side-data; verify orientation after.
ffmpeg -i in.mp4 -map_metadata -1 -c copy clean.mp4
# Quick probes (machine-readable; prefer probe-media.py for the full picture):
ffprobe -v error -show_entries format=duration -of default=nw=1:nk=1 in.mp4
ffprobe -v error -select_streams v:0 -show_entries stream=codec_name,width,height,r_frame_rate -of csv=p=0 in.mp4
Safe re-encode of untrusted uploads, scene-change detection, integrity in CI: references/analysis-validation.md.
yt-dlp interop
yt-dlp embeds ffmpeg for merge/remux; these are the post-download patterns:
# Prefer h264+m4a at download time (avoids a transcode entirely):
yt-dlp -S "res:1080,vcodec:h264,acodec:m4a" --remux-video mp4 URL
# Clip a section AT download (server-side range requests; much faster than full DL):
yt-dlp --download-sections "*10:00-12:30" -S "res:1080,vcodec:h264" URL
# Audio-only for STT/summarisation:
yt-dlp -x --audio-format opus URL
# Already downloaded a VP9/AV1 .webm that needs to be H.264 .mp4: that is a normal
# transcode — use the web-compatible H.264 recipe above, NOT --recode-video.
Generative/test sources
# Synthetic video+audio — fixtures, pipeline tests, alignment checks (no real media
# needed; this is how tests/run.sh builds its fixtures):
ffmpeg -f lavfi -i testsrc2=duration=2:size=640x360:rate=30 \
-f lavfi -i "sine=frequency=440:duration=2" \
-c:v libx264 -pix_fmt yuv420p -c:a aac fixture.mp4
Audio-reactive visuals (showwaves/showspectrum), podcast audiograms: references/visualization.md.
Footguns
The table that pays this skill's rent. Each row is a class of silent failure.
| Footgun | The trap | The rule |
|---|---|---|
-ss + -c copy |
Cut starts seconds early or with frozen/black lead-in (snapped to prior keyframe) | Copy cuts snap. Check probe-media.py --keyframes-near; re-encode when exact |
Output-side -to after input-side -ss |
Timestamps reset at the seek point, so -to silently becomes a duration |
Keep -ss/-to on the same side of -i (both input-side is fast and absolute) |
Missing -pix_fmt yuv420p |
Encode "works" but Safari/QuickTime/TVs show black or refuse to play (defaulted to yuv444p/yuv422p from a high-quality source) | Always set it for delivery H.264/H.265 |
Missing -movflags +faststart |
Browser can't start playback until the whole file downloads (moov at EOF) | Always set it for web-served MP4 |
| Default stream selection | ffmpeg picks ONE stream per type (highest-res video, most-channels audio) — extra audio tracks and all subs are silently dropped | -map 0 to keep everything, explicit -map otherwise |
-vf + -c:v copy together |
Hard error — filters require decoding | Filtering implies re-encode; pick one |
| VFR source (phone/Zoom/Loom/screen-rec) | Cut math drifts, concat desyncs, players stutter | Normalize first: -fps_mode cfr -r 30 + re-encode (cookbook) |
-vsync (deprecated) |
Old flag, removed direction | Use -fps_mode (cfr/vfr/passthrough) |
scale=W:-1 |
Odd height → encoder error with yuv420p | Always -2 |
| concat demuxer on mismatched inputs | "Works" then glitches/desyncs at boundaries (codec/timebase mismatch) | Demuxer = identical params only; else concat filter with re-encode |
| amix default | Each input's volume halved (normalize defaults on) | amix=...:normalize=0 + explicit volume= |
| One-pass loudnorm | Dynamic mode pumps quiet passages; output silently 192 kHz | Two-pass linear via loudnorm-scan.py; add -ar 48000 |
-shortest absent on audio-replace |
Output runs as long as the LONGEST input (silence or frozen frame tail) | Add -shortest when muxing separate A/V |
| BT.601/709 colour shift | Slightly wrong colours after scaling SD↔HD (matrix guessed from resolution) | Tag explicitly when it matters: see references/color-hdr.md |
| drawtext/subtitles path escaping | Filter args re-parse : and \ — Windows paths like C:\x explode inside filter strings |
cd to the asset's dir and use bare relative names; or escape as C\:/path |
| Interactive overwrite prompt | Agent-run command hangs forever on "File exists. Overwrite? [y/N]" | Always pass -y or -n explicitly |
% in cmd.exe |
%06d patterns and %{pts} get mangled by cmd variable expansion |
Use PowerShell or bash; in .bat double to %% |
Windows notes
Platform-agnostic commands, but when running on Windows:
- PowerShell quoting is friendlier than bash here: single quotes are fully
literal, so
-vf 'scale=1280:-2,fps=30'needs no escaping. Double quotes only interpolate$and backtick — filtergraphs rarely contain either. NULnot/dev/nullfor two-pass logs:-passlogfiledefaults are fine, butffmpeg ... -f null NUL(PowerShell also accepts-f null -, which is portable — prefer it).- Font paths in drawtext:
fontfile='C\:/Windows/Fonts/arial.ttf'— forward slashes, escaped drive colon, inside the filter string. - Prefer
-f null -and relative paths to sidestep both quoting tables at once.
Decision trees
Codec — H.264 (libx264): default; universal playback, fast, good per-bit at
-crf 18..23. → H.265 (libx265): same quality ~40% smaller; slower; needs
-tag:v hvc1 for Apple; fine for storage/modern devices. → AV1 (libsvtav1): best
compression, royalty-free, web-first (YouTube/Netflix path); encode cost highest;
playback on older hardware is software-only. → VP9: only when a pipeline demands
webm and AV1 is unavailable. → FFV1: archival masters only.
Cut method — Need exact frames OR applying any filter → re-encode (input-side
-ss, -crf 18). Cut points happen to sit on keyframes (verify with
--keyframes-near) OR a ±2s slop is acceptable → stream copy with
-avoid_negative_ts make_zero. Many cuts from one source → EDL workflow below.
CPU vs hardware encode — Hardware (NVENC/QSV/AMF/VideoToolbox) is 5-20× faster
but worse quality per bit than libx264/x265 at slow presets. Use hardware for:
speed-critical batch work, live/streaming, drafts, "good enough" deliveries (bump
bitrate ~30% to compensate). Use CPU for: final masters, size-constrained targets,
quality comparisons. Always capability-scan.sh first — listed encoders fail at
runtime on driver mismatches. Details: references/hardware-accel.md.
EDL workflow (edit-as-code)
For any multi-cut edit, do not fire ad-hoc trim commands. Write an edit decision list — a JSON file naming every clip, time range, and why — then cut from it. The edit becomes reviewable (rationale is written down), rerunnable (regenerate the output any time), and diffable (versions of the edit are git history).
# 1. Find candidate cut points (silence = clean speech boundaries):
python skills/ffmpeg-ops/scripts/detect-segments.py --silence --json take3.mp4
# 2. Author the EDL (schema: assets/edl-schema.json) with per-scene rationale.
# 3. Dry-run prints every ffmpeg command it would run (default — nothing executes):
python skills/ffmpeg-ops/scripts/cut-from-edl.py edit.json
# 4. Execute: cuts + concat -> final. Re-encodes by default for frame accuracy;
# --copy for keyframe-aligned EDLs.
python skills/ffmpeg-ops/scripts/cut-from-edl.py edit.json --execute -o final.mp4
Rules that make this work (from the Fable launch-video pipeline): cuts must land in silence; the model reasons over transcripts, not frames; after cutting, re-transcribe the output to verify (no filler words survived, no words clipped). Full architecture, EDL schema, verification loop: references/edit-as-code.md.
Color grading
# Apply a .cube LUT (tetrahedral = highest quality interpolation):
ffmpeg -i in.mp4 -vf "lut3d=file=grade.cube:interp=tetrahedral" \
-c:v libx264 -crf 18 -c:a copy graded.mp4
# Generate a family of grade candidates + an HTML still-chooser:
python skills/ffmpeg-ops/scripts/gen-luts.py --variants all --out-dir work/luts \
--previews in.mp4
The human picks the grade. Generate variants, render preview stills, present a chooser — never auto-select a look. Grading is a taste call; the agent's job is the lattice math and the apply command. LUT format, log-footage normalization (S-Log3/V-Log → Rec.709), curves/eq safe ranges, checking work with ffmpeg's built-in scopes (waveform/vectorscope): references/color-grading.md. The 25-look recipe catalog — film stocks (Kodachrome, CineStill halation, Technicolor, Eterna), signature grades (Mad Max, Fincher, Matrix, BR2049, Amélie…), era/genre moods, Sin City selective color — every chain build-validated, plus the Hald-CLUT match-any-look workflow and scope-matching ladder: references/look-recipes.md. Pipeline correctness (pix_fmt, HDR→SDR tonemapping, range/matrix tagging): references/color-hdr.md.
Quality gates
# VMAF/SSIM/PSNR verdict on an encode (exit 10 = below threshold -> branch on it):
python skills/ffmpeg-ops/scripts/quality-compare.py reference.mp4 encoded.mp4 \
--metrics ssim,psnr
python skills/ffmpeg-ops/scripts/quality-compare.py reference.mp4 encoded.mp4 \
--metrics vmaf --min-vmaf 90 --json | jq '.data.vmaf'
VMAF ≥ 93 at 1080p ≈ visually transparent; 80-93 = noticeable on inspection.
Side-by-side visual A/B (hstack), metric interpretation, encode-ladder tuning:
references/quality-metrics.md.
Scripts
All eleven follow the Skill Resource Protocol:
--help with examples, stdout = data only, --json envelopes
(claude-mods.ffmpeg-ops.*/v1), semantic exit codes (0 ok, 2 usage, 3 input
missing, 4 invalid input, 5 missing dependency, 7 ffmpeg unavailable,
10 domain finding).
| Script | Job | Worked invocation |
|---|---|---|
probe-media.py |
Normalized inspection, keyframe proximity, --doctor triage (hazard → fix command, exit 10) |
probe-media.py --doctor in.mp4 |
capability-scan.sh |
What can THIS ffmpeg build do (proof-encodes hw encoders; --quick skips) |
capability-scan.sh --json \| jq '.data.encoders' — exit 10 = a listed encoder failed verification |
quality-compare.py |
VMAF/SSIM/PSNR gate | quality-compare.py ref.mp4 enc.mp4 --min-vmaf 90 — exit 10 = below threshold |
loudnorm-scan.py |
Two-pass loudnorm: measures pass 1, emits exact pass-2 filter | loudnorm-scan.py -I -16 in.mp4 --json \| jq -r '.data.pass2_filter' |
detect-segments.py |
Silence/scene boundaries as JSON segments (STT chunking, dead-air cuts, shot splits) | detect-segments.py --scenes --json in.mp4 \| jq '.data.segments' |
cut-from-edl.py |
EDL JSON → validated cuts + concat (dry-run by default) | cut-from-edl.py edit.json --execute -o final.mp4 |
make-chapters.py |
Scene/silence points (or explicit JSON) → embedded chapters / YouTube text / WebVTT | make-chapters.py --from-scenes --media talk.mp4 --write chaptered.mp4 |
smart-compress.py |
Fit a size cap: computed two-pass bitrate, auto audio/downscale, size-verified (exit 10 = still over) | smart-compress.py --target 25MB video.mp4 |
make-sprites.py |
Scrub-preview sprite sheets + WebVTT thumbnail track (#xywh) | make-sprites.py --interval 5 video.mp4 |
gen-luts.py |
Emit .cube grade variants (+ --previews still chooser) |
gen-luts.py --variants warm_filmic,punchy --out-dir luts/ |
verify-commands.sh |
Staleness verifier: --offline structural (CI), --live checks docs against the installed build |
verify-commands.sh --live — exit 10 = doc drift, 7 = no ffmpeg |
References
Load on demand — one concept per file:
| Reference | Load when |
|---|---|
| encoding.md | Choosing codec/CRF/preset, two-pass, social platform targets, archival |
| hardware-accel.md | NVENC/QSV/AMF/VideoToolbox/VAAPI flags, quality caveats, detection |
| filtergraph.md | Any -filter_complex, labels/chains/split, speed ramps, xstack |
| trim-concat.md | Cut accuracy, keyframes, concat selection, segment removal |
| edit-as-code.md | Multi-cut edits, EDL schema, transcript-driven editing, verify loop |
| audio.md | Loudness, mixing, channel layout, audio repair |
| stt-whisper.md | Whisper/WhisperX prep, chunking, transcript JSON, summarisation pipeline |
| subtitles.md | Burn vs soft, styling, extraction, format conversion |
| color-grading.md | LUTs, .cube format, log normalization, scopes, grade workflow |
| look-recipes.md | 25-look catalog (film stocks, signature movie grades, era/genre moods), Hald-CLUT extraction, scope-matching |
| color-hdr.md | pix_fmt, HDR→SDR tonemap, BT.601/709 tagging, 10-bit |
| quality-metrics.md | VMAF/SSIM interpretation, visual A/B, ladder tuning |
| streaming-hls.md | HLS/DASH packaging, ABR ladders, live restream |
| images-gif.md | GIF quality, sprite sheets, dataset frame extraction |
| restoration.md | Deinterlace, denoise, deband, stabilize, audio cleanup |
| analysis-validation.md | Corruption checks, hashing, metadata stripping, untrusted uploads |
| capture-devices.md | Screen/webcam capture per OS (gdigrab/dshow, avfoundation, x11grab) |
| error-decoder.md | An ffmpeg command failed with a cryptic message — symptom → cause → fix |
| visualization.md | Waveform/spectrogram videos, audiograms, comparison grids |
Assets: encoding-presets.json (recipe data incl. date-stamped social targets), hls-ladder.json (ABR ladder), edl-schema.json (the cut-from-edl.py contract).
Self-test
bash skills/ffmpeg-ops/tests/run.sh # offline suite; synthesizes fixtures via lavfi
Structural assertions always run; media round-trips run only when ffmpeg is on PATH (loud skip otherwise — never a silent false-clean).
Files (claude-mods)
-
assets
-
edl-schema.json 2.9 KB
{ "$schema": "http://json-schema.org/draft-07/schema#", "$id": "claude-mods.ffmpeg-ops.edl-schema/v1", "title": "Edit Decision List", "description": "The contract between shot selection and cut-from-edl.py. The edit is this file: reviewable (rationale written down), rerunnable (regenerate output any time), diffable (versions are git history). Times are seconds from the start of each source file.", "type": "object", "required": ["scenes"], "properties": { "title": { "type": "string", "description": "Human title for the edit" }, "output": { "type": "string", "description": "Default output path, relative to this file (cut-from-edl.py -o overrides)" }, "source_notes": { "type": "string", "description": "Anything the next reader needs about the footage (takes layout, which re-shoots exist, transcript locations)" }, "scenes": { "type": "array", "minItems": 1, "items": { "type": "object", "required": ["clips"], "properties": { "scene": { "type": ["integer", "string"], "description": "Scene number or id" }, "title": { "type": "string" }, "candidate_takes": { "type": "array", "items": { "type": "string" }, "description": "Takes that were considered — documentation of the search space" }, "selection_rationale": { "type": "string", "description": "WHY these clips won. Write it down; this is what makes the EDL reviewable. E.g. 'C003 is the cleanest complete take: zero ums, clean ending; C017 disqualified - 5.8s dead pause mid-sentence.'" }, "clips": { "type": "array", "minItems": 1, "items": { "type": "object", "required": ["file", "start", "end"], "properties": { "file": { "type": "string", "description": "Source path, relative to this EDL file (or absolute)" }, "start": { "type": "number", "minimum": 0, "description": "In-point, seconds. Should sit in silence — verify with detect-segments.py --silence" }, "end": { "type": "number", "exclusiveMinimum": 0, "description": "Out-point, seconds (absolute in the source, not a duration). Must be > start" }, "first_words": { "type": "string", "description": "First words spoken in this range — a human-checkable anchor against the transcript" }, "note": { "type": "string" } } } } } } } } } -
encoding-presets.json 3.7 KB
{ "schema": "claude-mods.ffmpeg-ops.encoding-presets/v1", "updated": "2026-06-12", "note": "Recipe data the agent queries instead of re-deriving. Args are ffmpeg argv fragments; {input}/{output} are placeholders. Social specs drift - check the 'updated' stamp and the platform's current docs before trusting a target older than ~6 months.", "delivery": { "web_h264": { "use": "Default web/share delivery - universal playback", "args": "-c:v libx264 -crf 20 -preset slow -pix_fmt yuv420p -c:a aac -b:a 192k -movflags +faststart", "container": "mp4" }, "web_h264_small": { "use": "Size-sensitive H.264 (email, chat upload caps)", "args": "-c:v libx264 -crf 24 -preset slow -pix_fmt yuv420p -vf scale=-2:720 -c:a aac -b:a 128k -movflags +faststart", "container": "mp4" }, "hevc_storage": { "use": "Personal library/storage - ~40% smaller, modern devices", "args": "-c:v libx265 -crf 24 -preset slow -tag:v hvc1 -pix_fmt yuv420p -c:a aac -b:a 160k -movflags +faststart", "container": "mp4" }, "av1_web": { "use": "Best compression for web-first targets; encode cost highest", "args": "-c:v libsvtav1 -crf 32 -preset 6 -pix_fmt yuv420p10le -c:a libopus -b:a 128k", "container": "webm" }, "archive_ffv1": { "use": "Lossless preservation master (archival standard)", "args": "-c:v ffv1 -level 3 -g 1 -slicecrc 1 -c:a flac", "container": "mkv" }, "normalize_source": { "use": "Fix problem sources (HEVC/VFR/Zoom/Loom) before editing", "args": "-c:v libx264 -crf 18 -preset fast -pix_fmt yuv420p -fps_mode cfr -r 30 -c:a aac -b:a 192k -ar 48000", "container": "mp4" } }, "audio": { "podcast_opus": { "use": "Voice distribution - best codec per bit", "args": "-vn -c:a libopus -b:a 32k -ac 1 -application voip" }, "music_opus": { "use": "Music/general audio", "args": "-vn -c:a libopus -b:a 128k" }, "stt_prep": { "use": "Whisper-family input - 16 kHz mono PCM", "args": "-vn -ac 1 -ar 16000 -c:a pcm_s16le", "container": "wav" }, "loudness_targets_lufs": { "streaming_platforms": -14, "podcast": -16, "ebu_r128_broadcast": -23 } }, "social": { "_note": "Specs as of the 'updated' stamp. aspect = canvas; fit landscape sources with the blurred-pad pattern in SKILL.md.", "youtube_standard": { "aspect": "16:9", "resolution": "1920x1080", "fps_max": 60, "args": "-c:v libx264 -crf 19 -preset slow -pix_fmt yuv420p -c:a aac -b:a 192k -ar 48000 -movflags +faststart" }, "youtube_shorts": { "aspect": "9:16", "resolution": "1080x1920", "duration_max_s": 180, "args": "-c:v libx264 -crf 20 -preset slow -pix_fmt yuv420p -c:a aac -b:a 192k -movflags +faststart" }, "instagram_reel": { "aspect": "9:16", "resolution": "1080x1920", "duration_max_s": 900, "fps_max": 60, "args": "-c:v libx264 -crf 21 -preset slow -pix_fmt yuv420p -c:a aac -b:a 128k -movflags +faststart" }, "tiktok": { "aspect": "9:16", "resolution": "1080x1920", "duration_max_s": 600, "args": "-c:v libx264 -crf 21 -preset slow -pix_fmt yuv420p -c:a aac -b:a 128k -movflags +faststart" }, "twitter_x": { "aspect": "16:9 or 1:1", "resolution": "1920x1080", "duration_max_s": 140, "size_max_mb": 512, "args": "-c:v libx264 -crf 22 -preset slow -pix_fmt yuv420p -c:a aac -b:a 128k -movflags +faststart" } }, "gif": { "standard": { "use": "README/PR demo GIF - palettegen quality at sane size", "filter": "fps=12,scale=480:-1:flags=lanczos,split[s0][s1];[s0]palettegen=max_colors=128[p];[s1][p]paletteuse=dither=bayer:bayer_scale=4" } } } -
hls-ladder.json 1.6 KB
{ "schema": "claude-mods.ffmpeg-ops.hls-ladder/v1", "updated": "2026-06-12", "note": "ABR ladder reference derived from Apple's HLS authoring guidelines (H.264, 16:9). Bitrates are video-only targets in kb/s; pair with 128k AAC audio. Trim rungs you don't need - a 3-rung ladder (e.g. 1080/720/360) covers most non-broadcast uses. Segment duration 6s; keyframe interval must equal segment duration x fps.", "ladder": [ { "name": "2160p", "resolution": "3840x2160", "video_kbps": 16800, "profile": "high", "level": "5.1", "fps": "source" }, { "name": "1440p", "resolution": "2560x1440", "video_kbps": 9600, "profile": "high", "level": "5.0", "fps": "source" }, { "name": "1080p", "resolution": "1920x1080", "video_kbps": 6000, "profile": "high", "level": "4.2", "fps": "source" }, { "name": "720p", "resolution": "1280x720", "video_kbps": 3000, "profile": "main", "level": "4.0", "fps": "source" }, { "name": "540p", "resolution": "960x540", "video_kbps": 2000, "profile": "main", "level": "3.1", "fps": "source" }, { "name": "360p", "resolution": "640x360", "video_kbps": 730, "profile": "main", "level": "3.0", "fps": "source" }, { "name": "270p", "resolution": "480x270", "video_kbps": 365, "profile": "baseline", "level": "3.0", "fps": "source" } ], "audio": { "codec": "aac", "kbps": 128, "sample_rate": 48000 }, "packaging": { "segment_duration_s": 6, "playlist_type": "vod", "keyframe_rule": "force keyframes at segment boundaries: -g <fps*6> -keyint_min <fps*6> -sc_threshold 0, or -force_key_frames expr:gte(t,n_forced*6)" } }
-
-
references
-
analysis-validation.md 3.4 KB
# Analysis & validation — integrity, hashing, metadata, untrusted media ## Corruption / decode check ```bash ffmpeg -v error -i in.mp4 -f null - 2> errors.log # exit code alone is NOT the verdict — partial corruption decodes "successfully". # empty errors.log = clean; lines name the damaged streams/timestamps. ``` Fast container-level check (no full decode): `ffprobe -v error in.mp4` — catches truncation and broken headers in milliseconds; use it as the cheap first gate in batch jobs, the full decode as the thorough second. ## Frame hashing — prove pipelines identical ```bash ffmpeg -i a.mp4 -map 0:v -f framemd5 a.md5 ffmpeg -i b.mp4 -map 0:v -f framemd5 b.md5 diff a.md5 b.md5 # identical = bit-identical decoded frames ``` Use cases: verify a remux didn't touch frames, prove an FFV1 archival round-trip is lossless, CI-assert a refactored pipeline produces identical output. `-f streamhash` (one hash per stream) is the cheap whole-file variant. ## Metadata: inspect and strip ```bash ffprobe -v error -show_format -show_entries format_tags in.mp4 # what's in there # strip everything (GPS, device model, creation time — privacy before sharing): ffmpeg -i in.mp4 -map_metadata -1 -map 0 -c copy clean.mp4 ``` Two traps: (1) **rotation** — stripping can drop the display matrix on phone video; probe the output (`probe-media.py`) and re-apply `-display_rotation` if needed. (2) **chapters** survive `-map_metadata -1`; add `-map_chapters -1` to drop those too. ```bash # add useful metadata (chapters from scene detection, title, language): -metadata title="..." -metadata:s:a:0 language=eng ``` ## Untrusted uploads (server-side discipline) A user-supplied "video" is attacker-controlled input to a large C codebase. The pattern: 1. **Validate cheaply first** — `ffprobe -v error` with a timeout; reject on any error, absurd stream counts, or absurd dimensions/duration vs your product limits. 2. **Never trust the extension** — probe reports the real container. 3. **Re-encode, don't copy** — a full decode→encode discards container exploits, weird private streams, and metadata payloads in one move (the normalize recipe in SKILL.md is the right shape). 4. **Cap resources** — wall-clock timeout per job, `-t <max>` duration cap; ffmpeg happily eats a 10-hour 8K input otherwise. 5. Strip metadata on output (`-map_metadata -1`) — it re-encodes *in*, otherwise. ## Scene/content probes ```bash # scene-change list (chapters, shot logs): python skills/ffmpeg-ops/scripts/detect-segments.py --scenes --json in.mp4 | jq '.data.cuts' # black-frame / freeze detection (broken renders, dead air): ffmpeg -i in.mp4 -vf "blackdetect=d=0.5:pix_th=0.10" -an -f null - 2>&1 | rg black_ ffmpeg -i in.mp4 -vf "freezedetect=n=-60dB:d=2" -an -f null - 2>&1 | rg freeze_ # bitrate-over-time (find the spike that breaks a streaming budget): ffprobe -v error -select_streams v:0 -show_entries packet=pts_time,size -of csv=p=0 in.mp4 \ | awk -F, '{b[int($1)]+=$2} END{for(s in b) printf "%d\t%.0f kb/s\n", s, b[s]*8/1000}' | sort -n ``` ## CI gates for media artifacts A render pipeline's test suite, in three asserts: container parses (`ffprobe -v error`), duration within tolerance (`format=duration` vs expected), decode clean (`-v error -f null -` with empty log). Add `quality-compare.py --min-ssim 0.97` against a golden reference when the pipeline is supposed to be visually stable — see [quality-metrics.md](quality-metrics.md). -
audio.md 3.4 KB
# Audio — loudness, mixing, channels, repair ## Loudness (EBU R128) Targets: **-14 LUFS** streaming platforms, **-16** podcasts, **-23** broadcast. True peak ceiling -1.5 dBTP (-2 for lossy delivery). One-pass `loudnorm` is *dynamic* mode (a compressor — pumps quiet passages). For anything that ships, use two-pass **linear** mode; the measurement dance is automated: ```bash python skills/ffmpeg-ops/scripts/loudnorm-scan.py -I -16 in.mp4 --json \ | jq -r '.data.pass2_command' # prints the exact pass-2 ffmpeg command with measured_* values filled in ``` Check where you stand without changing anything: ```bash ffmpeg -i in.mp4 -af ebur128 -f null - 2>&1 | tail -12 # integrated I, LRA, peaks ``` `loudnorm` outputs 192 kHz internally — always pair with `-ar 48000`. ## Mixing and ducking ```bash # voice over music, music ducked 12dB whenever voice is present (sidechain): ffmpeg -i voice.wav -i music.mp3 -filter_complex \ "[1:a][0:a]sidechaincompress=threshold=0.05:ratio=8:attack=20:release=400[duck]; [0:a][duck]amix=inputs=2:duration=first:normalize=0[a]" \ -map "[a]" -c:a aac mixed.m4a # plain mix at set levels (amix halves inputs unless normalize=0): -filter_complex "[1:a]volume=0.25[m];[0:a][m]amix=inputs=2:duration=first:normalize=0[a]" # concatenate audio files losslessly (same codec) / with re-encode: ffmpeg -f concat -safe 0 -i list.txt -c copy out.mp3 -filter_complex "[0:a][1:a]concat=n=2:v=0:a=1[a]" ``` ## Channels ```bash # stereo -> mono (downmix), mono -> "stereo" (duplicate): -ac 1 # downmix -ac 2 # duplicate mono to both # keep ONE channel of a stereo file (e.g. lav mic on left only): -af "pan=mono|c0=c0" # left; c0=c1 for right # swap channels / manual stereo from two mono files: -af "pan=stereo|c0=c1|c1=c0" ffmpeg -i L.wav -i R.wav -filter_complex "[0:a][1:a]join=inputs=2:channel_layout=stereo[a]" -map "[a]" out.wav # pick the 3rd audio track from a multi-track recording (OBS etc.): -map 0:a:2 ``` ## Sync repair ```bash # audio late by 300ms -> advance it (itsoffset on the AUDIO input): ffmpeg -i in.mp4 -itsoffset -0.3 -i in.mp4 -map 0:v -map 1:a -c copy fixed.mp4 # constant drift (audio runs long) -> resample-stretch: -af "atempo=1.001" # tune factor = video_duration / audio_duration ``` ## Repair & cleanup ```bash -af "highpass=f=100" # rumble/handling noise -af "afftdn=nf=-25" # broadband denoise (use ears; see stt-whisper.md caveat) -af "adeclick" # vinyl/mouth clicks -af "deesser" # sibilance -af "compand=attacks=0.05:decays=0.3:points=-80/-80|-45/-15|-27/-9|0/-7|20/-7" # leveler for speech -af "alimiter=limit=0.97" # brickwall before lossy encode ``` Order matters: **subtractive first** (highpass → denoise → declick), then dynamics (compand), then loudness (loudnorm), limiter last. ## Format notes - Sample rate: keep 48 kHz for video work (44.1 kHz is a music-CD convention; mixing the two invites resample drift in long files). - `aresample=async=1` repairs streams with small timestamp gaps (common in screen-recorder output) — add it when concat output crackles at boundaries. - Bit depth: `pcm_s16le` for interchange, `pcm_s24le` when the source is 24-bit; never "upgrade" 16→24 (it's free silence). -
capture-devices.md 2.8 KB
# Capture — screen and devices, per OS Capture is the one genuinely platform-specific corner of ffmpeg. Same downstream processing everywhere; only the input device differs. ## Windows (gdigrab / dshow / ddagrab) ```bash # full screen (gdigrab — works everywhere, CPU-based): ffmpeg -f gdigrab -framerate 30 -i desktop -c:v libx264 -preset ultrafast -crf 23 \ -pix_fmt yuv420p cap.mp4 # region / single window: ffmpeg -f gdigrab -framerate 30 -offset_x 100 -offset_y 100 -video_size 1280x720 -i desktop ... ffmpeg -f gdigrab -framerate 30 -i title="Exact Window Title" ... # modern GPU path (Win10+, much lower overhead, needs d3d11 build): ffmpeg -f ddagrab -framerate 60 -i 0 -c:v h264_nvenc -cq 23 cap.mp4 # webcam + mic (dshow): FIRST list devices, then use exact names: ffmpeg -list_devices true -f dshow -i dummy ffmpeg -f dshow -rtbufsize 256M -i video="HD Webcam":audio="Microphone (Realtek)" \ -c:v libx264 -preset veryfast -crf 22 -c:a aac cam.mp4 # system audio loopback: ffmpeg has no native WASAPI-loopback input — install # the VB-Cable/virtual-audio-capturer dshow device, or capture with OBS instead. ``` `-rtbufsize 256M` on dshow prevents the "real-time buffer too full" frame drops. ## macOS (avfoundation) ```bash ffmpeg -f avfoundation -list_devices true -i "" # indices change; always list # screen 1 + default mic ("1:0" = video-index:audio-index): ffmpeg -f avfoundation -framerate 30 -capture_cursor 1 -i "1:0" \ -c:v libx264 -preset veryfast -crf 22 -pix_fmt yuv420p cap.mp4 ``` Screen Recording permission (System Settings → Privacy) must be granted to the *terminal* running ffmpeg — the failure is a black recording, not an error. System-audio capture needs a loopback driver (BlackHole). ## Linux (x11grab / kmsgrab / v4l2 / pulse) ```bash # X11 screen + pulse audio: ffmpeg -f x11grab -framerate 30 -video_size 1920x1080 -i :0.0 \ -f pulse -i default -c:v libx264 -preset veryfast -crf 22 -pix_fmt yuv420p cap.mp4 # webcam: ffmpeg -f v4l2 -framerate 30 -video_size 1280x720 -i /dev/video0 cam.mp4 ``` Wayland blocks x11grab — capture via `pipewiregrab`/`kmsgrab` (build-dependent) or use OBS as the capture layer. ## Capture-encode discipline (all platforms) - **Capture cheap, compress later.** `-preset ultrafast -crf 18` (or hardware encode) during capture; transcode to delivery settings afterwards ([encoding.md](encoding.md)). Dropped frames during capture are unfixable; large intermediates are. - Screen content is **full-range RGB** — the range/matrix tagging trap in [color-hdr.md](color-hdr.md) applies to every screen recording. - Capture is inherently VFR-ish under load: run the normalize recipe before editing captures. - Long captures: `-f segment -segment_time 600 -reset_timestamps 1` so a crash loses ten minutes, not three hours. -
color-grading.md 4.2 KB
# Color grading — LUTs, log footage, scopes, the grade workflow Creative color. Pipeline *correctness* (pix_fmt, HDR, range/matrix tags) is [color-hdr.md](color-hdr.md) — read that first if colors look *wrong* rather than *unstyled*. Recipes for *named* looks (teal & orange, pastel, noir, VHS…) and the Hald-CLUT / scope-matching techniques: [look-recipes.md](look-recipes.md). ## The workflow 1. **Normalize log footage** to Rec.709 (below) — grade on display-referred video. 2. **Generate candidates**, render preview stills, build the chooser: ```bash python skills/ffmpeg-ops/scripts/gen-luts.py --variants all \ --out-dir work/luts --previews footage.mp4 --frame-at 12.5 # -> work/luts/*.cube + preview_*.png + index.html ``` 3. **The human picks.** Never auto-select a grade — taste is a human gate. 4. **Apply:** ```bash ffmpeg -i in.mp4 -vf "lut3d=file=work/luts/warm_filmic.cube:interp=tetrahedral" \ -c:v libx264 -crf 18 -c:a copy graded.mp4 ``` 5. **Check against scopes**, not eyeballs (below). Order of operations when combining with other work: denoise → normalize log → grade (LUT) → sharpen → encode. Grading before denoise amplifies chroma noise. ## Log footage ("why does my drone/mirrorless footage look washed out") Log profiles (S-Log3, V-Log, D-Log, C-Log) pack wide dynamic range into a flat image; they *require* a conversion to Rec.709. Options: - `gen-luts.py --input-space slog3` bakes the S-Log3→Rec.709 conversion into every generated look (one LUT, one filter pass). - Camera vendors ship official conversion LUTs (Sony/Panasonic/DJI download pages) — highest fidelity; apply the official .cube first, then grade: `-vf "lut3d=vendor_to709.cube,lut3d=grade.cube"`. If footage is HLG/PQ rather than log, that's tonemapping, not grading — [color-hdr.md](color-hdr.md). ## .cube format (hand-writable, generatable) Plain ASCII: `TITLE`, `LUT_3D_SIZE N` (17/33/65 — 33 is the sweet spot), optional `DOMAIN_MIN/MAX`, then N³ lines of `R G B` floats 0–1, **red varying fastest**. That's why an agent (or `gen-luts.py`) can write one directly. ffmpeg's `lut3d` reads .cube/.3dl/.dat/.m3d; `interp=tetrahedral` is the quality option. ## Direct-filter grading (no LUT) For one-off tweaks; safe ranges that don't destroy footage: ```bash -vf "eq=brightness=0.03:contrast=1.08:saturation=1.1" # brightness ±0.1, contrast 0.9-1.3, sat 0-1.5 -vf "colortemperature=temperature=5500" # WB fix: 4000 warm <-> 7000 cool -vf "colorbalance=rs=0.05:bs=-0.05" # shadows toward orange (rs+) / teal (bs-) -vf "curves=preset=increase_contrast" # also: lighter, darker, vintage -vf "curves=master='0/0.04 0.5/0.5 1/0.96'" # custom: gentle film fade -vf "vibrance=intensity=0.4" # saturation that protects skin tones -vf "unsharp=5:5:0.8" # output sharpen, AFTER grade ``` ## Scopes — grade against measurements ffmpeg ships the same scopes a colorist uses; preview with `ffplay` or render a scope strip beside the image: ```bash # waveform (exposure): legal video sits 0-100%; clipping = flat line at top ffplay -i graded.mp4 -vf "split[a][b];[b]waveform=mode=column:display=stack[w];[a][w]vstack" # vectorscope (color cast/saturation): cast = trace off-center; skin tones hug # the I-line (~33° toward red-yellow) ffplay -i graded.mp4 -vf "split[a][b];[b]vectorscope=mode=color3[v];[a][v]hstack" # histogram per channel ffplay -i graded.mp4 -vf "split[a][b];[b]histogram[h];[a][h]hstack" ``` Mechanical checks: blown highlights = waveform pinned at 100% across a region; crushed blacks = pinned at 0; white-balance error = vectorscope centroid displaced on the B–R axis. ## Batch consistency Same grade across a folder = same LUT applied in a loop — this is the point of LUT-based grading (one decision, n applications): ```bash for f in clips/*.mp4; do ffmpeg -y -i "$f" -vf "lut3d=file=work/luts/warm_filmic.cube:interp=tetrahedral" \ -c:v libx264 -crf 18 -c:a copy "graded/$(basename "$f")" done ``` Shot-to-shot exposure differences need a per-clip `eq` *before* the shared LUT — match waveforms first, then the look lands identically. -
color-hdr.md 3.1 KB
# Color correctness — pix_fmt, HDR→SDR, range and matrix tags The "colors look *wrong*" file (washed out / too dark / slightly shifted / black video on Apple devices). For creative grading see [color-grading.md](color-grading.md). ## Pixel formats | pix_fmt | Use | |---|---| | `yuv420p` | **Every delivery encode.** The only universally-played option | | `yuv420p10le` | 10-bit: HEVC/AV1 delivery, HDR (mandatory), banding-prone gradients | | `yuv422p/444p` | Intermediates only — players choke | | `rgb24 / rgba` | Image outputs, overlays with alpha | ffmpeg preserves the source format when it can: encode from a screen recording or PNG sequence without `-pix_fmt yuv420p` and you silently get yuv444p → black/unplayable on QuickTime/Safari/TVs. **This is the single most common "ffmpeg broke my video" cause.** ## Range: limited (TV) vs full (PC) Video is normally limited range (16–235); PC/screen content is full (0–255). Mis-tagged range = washed-out blacks or crushed shadows *only in some players*. ```bash # screen recordings (full) -> delivery (limited), tagged correctly: -vf "scale=in_range=full:out_range=limited" -color_range tv ``` If output looks fine in one player and washed out in another, suspect range tags before anything else. Probe: `probe-media.py --json | jq '.data.video'`. ## Matrix: BT.601 vs BT.709 (the subtle skin-tone shift) SD is 601, HD is 709. Scaling SD↔HD without saying so makes the *scaler guess*, and a wrong guess shifts greens/skin slightly. Force it when crossing the line: ```bash # SD source upscaled to HD, explicit matrix conversion + tag: -vf "scale=1920:1080:in_color_matrix=bt601:out_color_matrix=bt709" \ -colorspace bt709 -color_primaries bt709 -color_trc bt709 ``` The three `-color*` flags only *tag* (they don't convert); the scale options *convert*. You usually want both. ## HDR → SDR tonemapping ("phone HDR video looks grey/flat after processing") iPhone/modern-camera HDR is HLG or PQ (probe shows `color_transfer=arib-std-b67` or `smpte2084`). Re-encoding without tonemapping produces the classic grey washed-out look. Convert properly (needs libzimg — check `capability-scan.sh`): ```bash ffmpeg -i hdr.mov -vf \ "zscale=t=linear:npl=100,format=gbrpf32le,zscale=p=bt709,tonemap=tonemap=hable:desat=0,zscale=t=bt709:m=bt709:r=tv,format=yuv420p" \ -c:v libx264 -crf 20 -c:a copy sdr.mp4 ``` Tonemap operators: `hable` (filmic, safe default), `mobius` (preserves mids), `reinhard` (flat), `linear` (clips). Without libzimg, a rougher fallback: `-vf "tonemapx=..."` builds vary — prefer installing a full build. **Keeping HDR:** copy streams (`-c copy`) keeps HDR metadata intact; re-encoding HDR10 properly requires x265 with `hdr10=1` master-display params — niche; verify with a probe that `color_transfer=smpte2084` survived. ## Alpha (transparency) ```bash # video with alpha -> overlay-ready formats: -c:v prores_ks -profile:v 4444 -pix_fmt yuva444p10le out.mov # NLE-grade -c:v libvpx-vp9 -pix_fmt yuva420p out.webm # web ``` MP4/H.264 has **no alpha** — requests for "transparent mp4" need webm or ProRes 4444 (or a separate matte). -
edit-as-code.md 3.9 KB
# Edit-as-Code — EDL-driven editing The pattern behind Anthropic's Fable launch video (edited entirely through Claude Code — transcription → shot-selection JSON → ffmpeg → LUTs, no NLE): **the edit is files, not timeline state.** Every stage's output is a reviewable, rerunnable, diffable artifact. ## The pipeline ``` raw takes (+ script if scripted) │ ▼ 1 TRANSCRIBE word-level JSON per take → work/transcripts/*.json ▼ 2 SELECT reason over transcripts → work/final-edit.json (EDL) ▼ 3 CUT cut-from-edl.py --execute → work/edl-cuts/ + final.mp4 ▼ 4 VERIFY re-transcribe the output → no clipped words, no filler ▼ 5 GRADE gen-luts.py + HUMAN picks → graded.mp4 ▼ 6 PACKAGE loudnorm pass-2, faststart, gates → deliverable ``` Stages 5–6 are [color-grading.md](color-grading.md) and [audio.md](audio.md)/[quality-metrics.md](quality-metrics.md); transcription is [stt-whisper.md](stt-whisper.md). This file owns stages 2–4. ## The EDL is the deliverable artifact Schema: [../assets/edl-schema.json](../assets/edl-schema.json). The load-bearing field is `selection_rationale` — *why* each take won, written down: ```json { "scene": 1, "title": "Part 1: Intro", "candidate_takes": ["C001", "C002", "C003", "C017 (re-shoot)"], "selection_rationale": "C017 disqualified - 5.8s dead pause mid-sentence. C003 is the cleanest complete take: zero ums, clean ending.", "clips": [{ "file": "takes/A004C003.mp4", "start": 1.89, "end": 60.81, "first_words": "Hey everyone, it's..." }] } ``` A reviewer reads the rationale instead of scrubbing footage. Git diffs of the EDL *are* the edit history. Re-running `cut-from-edl.py` regenerates the output identically. ## Shot-selection heuristics (multi-take footage) The agent reasons over **transcripts, not frames** — it cannot watch video. Per scene, read every candidate take's transcript and apply: - **Fewest filler words** ("um", "uh", restarts) wins, all else equal. - **Prefer later takes** — speakers warm up; the last full take is usually best. - **Disqualify** takes with dead pauses > ~2 s mid-sentence or that never complete the scripted line. - **Trim warm-up openers** ("Hey [name]…" used to start a sentence warm): cut at the silent gap *after* the warm-up, never mid-word. - Record disqualifications in the rationale too — the search space is part of the review. For many scenes, fan out one agent per scene (each reads only its candidates) and have a verifier pass check the assembled EDL — this maps directly onto the Workflow tool's pipeline+verify pattern. ## The two verification rules **1. Every cut boundary must land in silence.** Words are clipped by cuts that "look right" numerically. Mechanically check each in/out against measured silence: ```bash python skills/ffmpeg-ops/scripts/detect-segments.py --silence --json take.mp4 \ | jq --argjson t 60.81 '.data.silences[] | select(.start <= $t and .end >= $t)' # empty result = the proposed cut at 60.81 is NOT in silence — move it ``` **2. Re-transcribe the output.** After `cut-from-edl.py --execute`, run the final video back through transcription and assert: every scene's `first_words` appears, no filler words survived, no sentence is truncated at a boundary. This catches off-by-keyframe and timestamp-unit errors that no amount of EDL review will. ## Cut mode choice `cut-from-edl.py` re-encodes by default (frame-accurate, normalizes mixed sources, concat always safe). Use `--copy` only when the EDL was authored against measured keyframes (`probe-media.py --keyframes-near` for every in-point) — e.g. when the source is an all-intra mezzanine ([encoding.md](encoding.md)). ## Human gates Taste calls stay human: the grade pick ([color-grading.md](color-grading.md)), final timing, sound design. The agent's job ends at presenting options with evidence — never auto-select past these gates. -
encoding.md 4 KB
# Encoding — codecs, CRF, presets, two-pass, targets Recipe data lives in [../assets/encoding-presets.json](../assets/encoding-presets.json) (query it; don't re-derive). This file is the *why* behind those numbers. ## CRF — constant quality, the default rate mode CRF encodes to a perceptual quality level; size falls where it falls. Use CRF for everything except a hard size/bandwidth budget (then two-pass, below). | Encoder | Range | Visually lossless | Good delivery | Small | Notes | |---|---|---|---|---|---| | libx264 | 0–51 | 17–18 | 20–23 | 26–28 | +6 ≈ half the size | | libx265 | 0–51 | 20–21 | 23–26 | 28–30 | x265 CRF ≈ x264 CRF + 3 for similar quality | | libsvtav1 | 0–63 | 25–28 | 30–35 | 38–45 | scale differs — do not map 1:1 from x264 | | libvpx-vp9 | 0–63 | 24–28 | 31–36 | 40+ | needs `-b:v 0` for pure CRF mode | **VP9 trap:** `-crf 32` alone is *constrained* quality; pure CRF needs `-c:v libvpx-vp9 -crf 32 -b:v 0`. ## Presets — speed vs compression efficiency Preset changes *size at the same quality*, not the quality itself (CRF pins that). - **libx264/libx265:** `ultrafast..placebo`. `slow` is the sweet spot for delivery; `fast`/`medium` for intermediates; never `placebo` (≈1% gain, 2× time over veryslow). - **libsvtav1:** numeric `0–13`, lower = slower. `6` balanced, `4` quality-leaning, `8–10` for drafts. - Rule of thumb: if encode time doesn't matter, drop one preset slower rather than lowering CRF — better size/quality trade. ## Tune (libx264) `-tune film` (live action grain), `-tune animation` (flat areas + lines), `-tune grain` (preserve heavy grain — also consider this for film scans), `-tune stillimage`, `-tune zerolatency` (streaming only — disables lookahead). Don't set tune at all when unsure. ## 10-bit `-pix_fmt yuv420p10le` reduces banding in gradients (skies, dark scenes) even for 8-bit sources, at ~5% size cost. x265 and SVT-AV1 handle it natively; for H.264 it breaks too many players — keep H.264 8-bit. HDR requires 10-bit (see [color-hdr.md](color-hdr.md)). ## Two-pass — when you have a size budget Target bitrate = (size_MB × 8192 ÷ seconds) − audio_kbps. ```bash # 700 MB target for a 1h video with 128k audio → (700*8192/3600)-128 ≈ 1465k ffmpeg -y -i in.mp4 -c:v libx264 -b:v 1465k -preset slow -pass 1 -an -f null - ffmpeg -i in.mp4 -c:v libx264 -b:v 1465k -preset slow -pass 2 \ -c:a aac -b:a 128k -movflags +faststart out.mp4 ``` Pass 1 writes `ffmpeg2pass-0.log` in the CWD — run both passes from the same directory. On Windows `-f null -` works in PowerShell; no need for `NUL`. ## Audio codec choice | Codec | Use | Bitrates | |---|---|---| | libopus | Best per-bit; anything not chained to MP4-only players | voice 24–32k mono, music 96–128k stereo | | aac (native) | MP4 delivery default; fine at ≥128k stereo | 128–192k | | libmp3lame | Legacy compat only | `-q:a 2` (~190k VBR) | | flac / pcm_s16le | Archival / editing intermediates | lossless | Opus-in-MP4 exists but player support is patchy — Opus belongs in webm/mka/opus. ## Intermediates for editing Long-GOP H.264/HEVC is miserable to scrub/cut repeatedly. For multi-step edit pipelines, transcode once to an all-intra mezzanine and work on that: ```bash ffmpeg -i in.mp4 -c:v libx264 -crf 14 -preset fast -g 1 -c:a pcm_s16le mezz.mov ``` (`-g 1` = every frame a keyframe: any cut point is copy-safe, scrubbing is instant. ProRes via `-c:v prores_ks -profile:v 3` if the destination is an NLE.) ## Archival FFV1 level 3 in MKV is the preservation standard (lossless, checksummed, seekable): ```bash ffmpeg -i in.mp4 -c:v ffv1 -level 3 -g 1 -slicecrc 1 -c:a flac archive.mkv ``` Verify the round trip with `-f framemd5` (see [analysis-validation.md](analysis-validation.md)). ## Hard size caps (upload limits) CRF first, then check, then two-pass only if over: ```bash ffmpeg -i in.mp4 -c:v libx264 -crf 23 -preset slow -pix_fmt yuv420p \ -c:a aac -b:a 128k -movflags +faststart try.mp4 # over budget? compute bitrate for the cap and two-pass (above), or step CRF +2 ``` -
error-decoder.md 7.1 KB
# Error decoder — cryptic ffmpeg message → cause → fix ffmpeg's errors describe the *symptom at the C layer*, not the cause. This table maps the messages agents actually hit to what went wrong and the move that fixes it. Match on the quoted fragment (messages vary slightly across versions). ## Container / file errors | Message fragment | Actual cause | Fix | |---|---|---| | `moov atom not found` | MP4 truncated mid-write (crashed recorder, interrupted download, still-recording file) — the index never got written | If the recorder is still running, wait. Else recover with an untruncated reference file from the same device (untrunc) — ffmpeg alone cannot rebuild a missing moov | | `Invalid data found when processing input` | File isn't what the extension claims, is corrupt, or is encrypted (DRM) | `ffprobe -v error file` to see what it really is; check size > 0; DRM content is out of scope, full stop | | `Error opening output files: Invalid argument` (output side) | ffmpeg couldn't infer the muxer — usually a non-standard output extension (`.tmp`, no extension) | Name the format explicitly: `-f mp4 out.tmp`, or use a real extension | | `Unable to choose an output format ... use a standard extension` | Same as above, said more politely | Same fix | | `Permission denied` on output | Output open in a player (Windows file lock), or writing into a read-only dir | Close the player; write elsewhere; never edit a file in place — write new + rename | | `No such file or directory` but the path looks right | Shell quoting ate part of the path (spaces, `&`, parentheses), or a filter arg consumed it | Quote the whole path; for paths *inside* filter args see the quoting row below | ## Codec / stream errors | Message fragment | Actual cause | Fix | |---|---|---| | `Filtering and streamcopy cannot be used together` | `-vf`/`-af`/`-filter_complex` combined with `-c copy` on the same stream | Filters require re-encoding — drop `-c copy` (or only copy the *other* stream: `-c:a copy` with a video filter is fine) | | `height not divisible by 2` (or width) | yuv420p needs even dimensions; a `scale=W:-1` produced an odd size | Use `scale=W:-2` (and `-2` for width too) | | `Unknown encoder 'libx265'` (libvmaf, libsvtav1, …) | This build doesn't include the library — common with distro/minimal builds | `capability-scan.sh` to see what you have; install a full build (gyan.dev "full" on Windows, BtbN on Linux) | | `Specified pixel format ... is invalid or not supported` | Hardware encoder fed a CPU pixel format (or 10-bit into an 8-bit-only encoder) | NVENC/QSV need `format=nv12`/hwupload chains — see hardware-accel.md; or drop to a software encoder | | `No capable devices found` / `Cannot load nvcuda.dll` | NVENC listed in the build but no working NVIDIA driver/GPU | `capability-scan.sh` confirms (listed-but-failed = exit 10); use libx264 or fix the driver | | `Conversion failed!` as the only error | The real error is 5–20 lines earlier in stderr | Read upward; with `-v error` the first printed line IS the cause | | `Too many packets buffered for output stream` | Muxer starved — usually one stream much shorter than another in a filter graph | Add `-shortest`, or fix the graph so both streams cover the same span | | `Non-monotonic DTS` / `non monotonically increasing dts` warnings | Timestamp disorder — VFR source, sloppy cut, or concat of mismatched segments | Usually survivable as a warning; if A/V drifts: re-encode with `-fps_mode cfr`, or remux with `-fflags +genpts` | ## Filter errors | Message fragment | Actual cause | Fix | |---|---|---| | `No such filter: 'xyz'` | Typo, or build-optional filter absent (drawtext needs libfreetype, subtitles needs libass, …) | `ffmpeg -filters \| rg xyz`; full build if missing | | `Unable to parse option value "..." ` inside a filter | The filter-arg parser ate a `:` or `,` — classically a **Windows drive colon** (`lut3d=file=C:/...`) or a timecode | Escape (`C\:/path`) or — better — `cd` to the asset's directory and use a bare relative filename | | `Error initializing filter 'subtitles'` / `Unable to open ...srt` | Path escaping (above), or the build lacks libass | Relative filename from the subs' directory; check `capability-scan.sh` | | `Cannot find a matching stream for unlabeled input pad` | A filtergraph input wasn't connected — wrong `[0:v]` index or a consumed-twice stream | Label every pad explicitly; `split` a stream before feeding two filters | | `Media type mismatch between the ... filter` | Audio stream wired into a video filter or vice versa (`[0:a]` into `scale`, …) | Check the `[n:v]`/`[n:a]` selectors at each filter boundary | | `Padded dimensions cannot be smaller than input dimensions` | `pad=` target smaller than the (already-scaled) frame | Scale down first in the same chain, or enlarge the pad target | ## Seek / cut errors | Message fragment | Actual cause | Fix | |---|---|---| | No error at all, but the output is empty/0 bytes | Input-side `-ss` seeked PAST the end of the file — ffmpeg exits 0 having written nothing | Probe duration first (`probe-media.py`); treat empty output as failure in scripts, never trust exit 0 alone for frame extraction | | `Non full-range YUV is non-standard` then encoder fails (writing .jpg) | The mjpeg encoder refuses full-range input under default strictness — common when grabbing stills from full-range/PC-range sources | Output `.png` instead, or add `-strict unofficial` for jpg | | Output starts with frozen/black video after a copy cut | Cut point wasn't a keyframe; player shows nothing until the next IDR | `probe-media.py --keyframes-near <t>`; re-encode the cut or move it to a keyframe | | `-to value smaller than -ss; aborting` | `-ss` input-side + `-to` output-side: timestamps reset at the seek, so your absolute `-to` is now "before" 0 | Keep `-ss`/`-to` on the same side of `-i` | | First frames of a concat glitch/flash | concat demuxer fed segments with mismatched codec params or timebases | Identical params only for the demuxer; otherwise concat *filter* + re-encode (trim-concat.md) | ## Audio errors | Message fragment | Actual cause | Fix | |---|---|---| | `Invalid audio stream. Exactly one MP3 audio stream is required` | Muxing video (e.g. cover art counts!) or 2+ streams into `.mp3` | `-vn -map 0:a:0` for mp3; or use a real container (m4a/mka) | | Output much quieter than inputs after `amix` | amix normalizes (divides) by input count by default | `amix=...:normalize=0` + explicit `volume=` per input | | `The encoder 'aac' is experimental` (very old builds) | Ancient ffmpeg | Upgrade; (historic workaround was `-strict -2`) — if you see this, the build is too old to trust for anything | ## Reading errors efficiently ```bash ffmpeg -v error -i in.mp4 ... 2>&1 | head -5 # first error line = the cause ffmpeg -v verbose ... # when error mode hides context ffmpeg -h filter=scale # option ranges when "Invalid argument" comes from a filter ``` The single most useful habit: when a long command fails, re-run with `-v error` and **read the first line, not the last** — ffmpeg prints the root cause first and generic wrappers ("Conversion failed!", "Error while processing") last. -
filtergraph.md 3.4 KB
# Filtergraphs — syntax, labels, chains, and the patterns that need them ## Grammar in 60 seconds ``` -vf "f1=a=1:b=2,f2" simple: one video chain, commas join filters -af "f1,f2" same for audio -filter_complex "[0:v]f1[x];[x][1:v]f2[out]" multiple inputs/outputs need labels ``` - `,` chains filters; `;` separates parallel chains. - `[0:v]` `[1:a]` = input file 0's video, file 1's audio. `[label]` = your wire. - Every labeled output must be consumed (or `-map`ped). Unconsumed = error. - One stream cannot feed two filters — `split`/`asplit` it first. - `-vf`/`-af` and `-filter_complex` are mutually exclusive per stream; filters and `-c copy` are mutually exclusive, full stop. **Escaping (three layers deep):** the filter arg parser eats `:` and `,`, the graph parser eats `;` and `[]`, then your shell takes a pass. Inside a filter argument, escape with `\` (e.g. `drawtext=text='1\:30'`). Avoid the whole topic where possible: relative paths for files, single-quoted graphs (bash *and* PowerShell), no spaces in asset names. ## split — the fan-out primitive ```bash # blurred-background vertical (one decode, two consumers): -filter_complex "[0:v]split[a][b];[a]scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,boxblur=20[bg];[b]scale=1080:-2[fg];[bg][fg]overlay=(W-w)/2:(H-h)/2" ``` ## Common graph patterns ```bash # picture-in-picture, top-right, 1/4 size -filter_complex "[1:v]scale=iw/4:-1[pip];[0:v][pip]overlay=W-w-24:24" # overlay visible only between 5s and 12s -filter_complex "[0:v][1:v]overlay=24:24:enable='between(t,5,12)'" # crossfade two clips (1s, starting at 4s into clip A) — video and audio -filter_complex "[0:v][1:v]xfade=transition=fade:duration=1:offset=4[v];[0:a][1:a]acrossfade=d=1[a]" # side-by-side A/B (heights must match; scale first if not) -filter_complex "[0:v][1:v]hstack" # 2x2 grid -filter_complex "[0:v][1:v][2:v][3:v]xstack=inputs=4:layout=0_0|w0_0|0_h0|w0_h0" ``` ## Time manipulation ```bash # constant speed: video PTS x factor, audio atempo (0.5-100; chain for <0.5) -filter_complex "[0:v]setpts=0.5*PTS[v];[0:a]atempo=2.0[a]" # 2x -filter_complex "[0:v]setpts=4*PTS[v];[0:a]atempo=0.5,atempo=0.5[a]" # 0.25x # speed RAMP (slow-mo a highlight 10-12s, normal speed around it): cut three # ranges with trim/atrim, retime the middle, concat — see trim-concat.md's # remove-middle pattern with setpts=2*PTS added to the middle chain. # interpolated 60fps slow-mo (synthesizes frames; slow, occasionally wobbly # around fast motion — check the output) -vf "minterpolate=fps=60:mi_mode=mci:mc_mode=aobmc:vsbmc=1,setpts=2*PTS" -an ``` ## Expressions Filter args accept expressions: `t` (seconds), `n` (frame), `w/h`/`iw/ih` (sizes), `main_w/overlay_w` in overlay. Useful forms: ```bash overlay=x='if(gte(t,3),24,-w)' # slide in at t=3 drawtext=...:x=(w-text_w)/2:y=h-th-40 # centered lower third select='not(mod(n,30))' # every 30th frame fade=t=in:st=0:d=1,fade=t=out:st=9:d=1 # fade in/out (10s clip) ``` ## Per-filter docs without leaving the terminal ```bash ffmpeg -h filter=xfade # all options + ranges for one filter ffmpeg -filters | rg blur # discover what this build has ``` Niche corners worth knowing exist: `v360` (360°/VR re-projection), `geq` (per-pixel expressions), `sendcmd` (timed parameter changes), `zmq` (live parameter control). -
hardware-accel.md 3.2 KB
# Hardware acceleration — NVENC, QSV, AMF, VideoToolbox, VAAPI ## The one paragraph that prevents most hw-encode mistakes Hardware encoders are **5–20× faster** and **worse per bit** than libx264/x265 at slow presets — a dedicated ASIC doing fewer optimization passes. Use them for batch/draft/realtime work and bump bitrate ~30% to compensate; use CPU for final masters and size-constrained encodes. And **always verify, never trust the list**: ```bash bash skills/ffmpeg-ops/scripts/capability-scan.sh # proof-encodes each hw encoder ``` An encoder appearing in `ffmpeg -encoders` only means it was compiled in; NVENC fails at runtime on driver/CUDA mismatches, QSV without the right GPU/driver, VAAPI without a render node. Exit 10 from capability-scan = listed-but-broken. ## NVENC (NVIDIA) ```bash # quality-targeted VBR (the CRF-like mode; -cq lower = better, ~19-28) ffmpeg -i in.mp4 -c:v h264_nvenc -preset p5 -tune hq -rc vbr -cq 23 -b:v 0 \ -pix_fmt yuv420p -c:a copy out.mp4 ffmpeg -i in.mp4 -c:v hevc_nvenc -preset p6 -tune hq -rc vbr -cq 26 -tag:v hvc1 ... ``` - Presets are `p1`(fast)–`p7`(quality); the old `slow/fast/ll*` names are legacy. - `-rc vbr -cq N -b:v 0` ≈ constant quality; omit `-b:v 0` and ffmpeg imposes a default bitrate cap (classic "why is NVENC output blurry" cause). - Full decode→encode on GPU: `-hwaccel cuda -hwaccel_output_format cuda` before `-i`, GPU-side `scale_cuda`/`scale_npp` for resizing. - Consumer GeForce caps concurrent NVENC sessions (driver-dependent, typically 5–8). ## QSV (Intel Quick Sync) ```bash ffmpeg -init_hw_device qsv=hw -i in.mp4 -vf "format=nv12,hwupload" \ -c:v h264_qsv -global_quality 23 -preset slower -c:a copy out.mp4 ``` `-global_quality` is the CRF-analog (ICQ mode). Common failure: iGPU disabled in BIOS or no Intel media driver — capability-scan catches both. ## AMF (AMD, Windows) ```bash ffmpeg -i in.mp4 -c:v h264_amf -quality quality -rc cqp -qp_i 22 -qp_p 24 -c:a copy out.mp4 ``` Weakest quality-per-bit of the four; prefer CPU unless speed is the whole point. ## VideoToolbox (macOS) ```bash ffmpeg -i in.mp4 -c:v hevc_videotoolbox -q:v 55 -tag:v hvc1 -c:a copy out.mp4 ``` `-q:v` 1–100 (higher = better, ~50–65 typical). Apple Silicon VT is fast and respectable; still below libx265 slow for size-critical work. ## VAAPI (Linux) ```bash ffmpeg -vaapi_device /dev/dri/renderD128 -i in.mp4 \ -vf "format=nv12,hwupload" -c:v h264_vaapi -qp 23 -c:a copy out.mp4 ``` Needs a render node and the right driver (iHD for modern Intel, Mesa for AMD). The `format=nv12,hwupload` dance is mandatory — software frames must be uploaded. ## Hardware DECODE (often the better win) Decode acceleration helps any pipeline bottlenecked on reading high-res sources (4K HEVC preview/thumbnail/analysis jobs), independent of encode choice: ```bash ffmpeg -hwaccel auto -i 4k_hevc.mp4 -vf scale=1280:-2 -c:v libx264 -crf 20 out.mp4 ``` `-hwaccel auto` falls back to software silently — safe to include by default. Caveat: filters run on CPU frames unless you keep the pipeline on-GPU (`-hwaccel_output_format cuda` + `*_cuda` filters); mixing GPU decode with CPU filters costs a download/upload round trip and can be *slower* than pure CPU for filter-heavy graphs. Measure before assuming. -
images-gif.md 3.8 KB
# Images, GIFs, frames — thumbnails, sprites, datasets ## Thumbnails ```bash # at a timestamp (input-side -ss = instant, even 2h into the file): ffmpeg -ss 00:12:05 -i in.mp4 -frames:v 1 -q:v 2 thumb.jpg # "a representative frame" — the thumbnail filter scans and picks: ffmpeg -i in.mp4 -vf "thumbnail=300" -frames:v 1 thumb.jpg # one per chapter/scene: feed timestamps from detect-segments.py --scenes: python skills/ffmpeg-ops/scripts/detect-segments.py --scenes --json in.mp4 \ | jq -r '.data.cuts[]' | while read -r t; do ffmpeg -y -v error -ss "$t" -i in.mp4 -frames:v 1 "thumbs/scene_${t}.jpg" done ``` `-q:v` for JPEG: 2 ≈ excellent … 31 ≈ awful. PNG/WebP/AVIF by extension (`-c:v libwebp -quality 85`, AVIF needs libaom/libsvtav1 still support). ## Contact sheets & sprite sheets ```bash # contact sheet: 1 frame / 10s, 4x3 grid (visual summary of a video): ffmpeg -i in.mp4 -vf "fps=1/10,scale=320:-2,tile=4x3" -frames:v 1 sheet.png # scrub-preview sprite sheet for a web player (1/s, 10x10 pages, numbered): ffmpeg -i in.mp4 -vf "fps=1,scale=160:-2,tile=10x10" sprites_%02d.jpg # (player WebVTT thumbnail tracks map time -> sheet offset: t seconds = tile t%100) # GOTCHA: tile fed from an IMAGE-SEQUENCE input (-i seq_%02d.png) partial-fills # the grid on some builds (observed: ffmpeg 8.0 Windows) even though every # frame decodes. Deterministic fallback = explicit stack graph: # ffmpeg -i a.png -i b.png ... -filter_complex \ # "[0:v][1:v][2:v]hstack=3[r0];[3:v][4:v][5:v]hstack=3[r1];[r0][r1]vstack=2" ``` ## GIF (the palettegen discipline) GIF is 256 colors with no partial transparency; quality is *entirely* about the palette and dithering: ```bash # single-pass via split (fps and scale BEFORE palettegen — palette should be # computed on the frames that will actually be in the GIF): ffmpeg -ss 5 -to 8 -i in.mp4 -filter_complex \ "fps=12,scale=480:-1:flags=lanczos,split[s0][s1];[s0]palettegen=max_colors=128[p];[s1][p]paletteuse=dither=bayer:bayer_scale=4" \ out.gif ``` Size levers, in order of impact: duration → dimensions → fps (8–15 is plenty) → max_colors → dither (`bayer` smallest, `floyd_steinberg`/`sierra2_4a` prettiest). `palettegen=stat_mode=diff` helps when only a small region moves. **Modern check first:** most "GIF" destinations (Slack, GitHub, web) accept MP4 or animated WebP at a tenth the size — `-c:v libwebp -loop 0 -quality 80 out.webp`. ## Frame extraction ```bash ffmpeg -i in.mp4 frames/%06d.png # every frame ffmpeg -i in.mp4 -vf "fps=2" frames/%06d.png # 2 per second ffmpeg -i in.mp4 -vf "select='eq(pict_type,I)'" -fps_mode vfr keyframes/%04d.png ffmpeg -ss 12.500 -i in.mp4 -frames:v 1 exact.png # the frame at 12.5s ``` ## ML dataset prep ```bash # fixed-rate, model-square (center crop), consistent naming: ffmpeg -i in.mp4 -vf "fps=1,scale=512:512:force_original_aspect_ratio=increase,crop=512:512" \ ds/vid01_%06d.png # letterbox instead of crop (keep full frame): -vf "fps=1,scale=512:512:force_original_aspect_ratio=decrease,pad=512:512:(ow-iw)/2:(oh-ih)/2" # dedupe near-identical frames (slideshows, talking heads) before extraction: -vf "mpdecimate,fps=1" -fps_mode vfr ``` PNG for training (lossless); JPEG `-q:v 2` only when storage forces it. Keep the source-time mapping recoverable: either fixed fps (frame n ÷ fps = seconds) or `-frame_pts 1` to name files by PTS. ## Sequences → video ```bash ffmpeg -framerate 24 -i frames/%06d.png -c:v libx264 -crf 18 -pix_fmt yuv420p out.mp4 ffmpeg -framerate 24 -pattern_type glob -i 'shots/*.png' ... # unnumbered names (not on Windows cmd) ``` `-framerate` (input, before `-i`) sets how fast stills are read — forgetting it gives the 25fps default regardless of intent. And `-pix_fmt yuv420p` again: PNG sources otherwise produce yuv444p output (see [color-hdr.md](color-hdr.md)). -
look-recipes.md 17.5 KB
# Look recipes — matching named aesthetics Known-good starting points for recognizable looks, plus the two techniques for matching a look you can *see* but can't name (Hald-CLUT, scope-matching). Workflow, scopes, and LUT machinery: [color-grading.md](color-grading.md) — normalize log footage to Rec.709 FIRST, grade second. Two delivery forms per look: **LUT** (consistent, reusable — `gen-luts.py --variants <name>` where a parametric variant exists) and **direct filter chain** (tweakable per shot). Values are starting points for normal-exposure Rec.709; expect ±30% adjustment. `colorbalance` options are per-band per-channel: `rs/gs/bs` shadows, `rm/gm/bm` midtones, `rh/gh/bh` highlights, each −1..1. **Skin-tone caveat** (verified on the Kodak test portraits): looks that deepen shadows — kodachrome, noir, bleach bypass, day-for-night, horror — crush facial detail in darker skin. Lift midtones first (`eq=gamma=1.05` to `1.15` before the look) and verify on the waveform that face luma stays in the ~25–65% band. Always test a grade on the *darkest-skinned* person in the footage, not the lightest. **Contents:** [Film stocks & processes](#film-stocks--processes) · [Signature movie grades](#signature-movie-grades) · [Era looks](#era-looks) · [Genre & mood](#genre--mood) · [Stylized effects](#stylized-effects) · [Matching techniques](#matching-a-look-you-can-see-but-cant-name) · [At scale](#applying-any-of-this-at-scale) ## Film stocks & processes ### Kodachrome High contrast that deepens in the shadows, warm reds, glowing skin — the mid-century slide look: ```bash -vf "curves=master='0/0 0.25/0.20 0.75/0.78 1/0.98',colorbalance=rh=.04:rm=.02,eq=saturation=1.15:contrast=1.08" ``` Scope: shadows genuinely DARK (waveform floor at 0), skin warm of the I-line. ### CineStill 800T (with halation) Tungsten-cool base + the signature red-orange glow bleeding around bright lights. The glow is a composite, not a color shift — extract highlights, blur, tint red, screen-blend back: ```bash # Three load-bearing details, all visually verified: (1) format=rgb24 pins the # split — float filters like colortemperature push the negotiated format past # 8-bit and the branch's auto-conversions then corrupt into a full-frame # magenta wash; (2) the threshold is maxval-relative so it survives any bit # depth; (3) u/v are forced neutral so only luminance carries into the halo. # Scale sigma with resolution (~14 at 1080p); raise rr toward 1.5 for a # stronger glow. -filter_complex "[0:v]colortemperature=temperature=7800,eq=saturation=0.95,format=rgb24,split[a][b];[b]lutyuv=y='if(gt(val,0.78*maxval),val,0)':u='(maxval+minval)/2':v='(maxval+minval)/2',gblur=sigma=14,colorchannelmixer=rr=1:gg=0.3:bb=0.15[halo];[a][halo]blend=all_mode=screen" ``` Only reads as 800T on footage WITH point lights/neon in frame. Scope: cool centroid, but red channel spikes hugging every highlight. ### Technicolor 2-strip (1920s) The whole world collapses onto a red↔cyan axis (no blue/yellow existed): ```bash -vf "colorchannelmixer=rr=1:gg=.6:gb=.4:bg=.4:bb=.6,eq=saturation=1.2:contrast=1.1" ``` LUT form: `gen-luts.py --variants technicolor2`. Skies go teal, lips go lipstick-red, yellows die — that's correct, that's the look. ### Technicolor 3-strip (glorious era) Lush saturated primaries, marble-glow skin, no cast: ```bash -vf "vibrance=intensity=0.35,eq=contrast=1.12,curves=master='0/0.01 1/0.99'" ``` `vibrance` (not `eq=saturation`) is the point — it boosts muted colors while protecting already-saturated skin. ### Fuji Eterna The modern "cinema flat" — low saturation, low contrast, long highlight roll: ```bash -vf "eq=saturation=0.82:contrast=0.92,curves=master='0/0.05 0.7/0.66 1/0.93'" ``` ### Cross-process (E6-in-C41) Green-yellow highlights, cyan-blue shadows, punchy contrast — the skate-video look: ```bash -vf "curves=r='0/0 0.5/0.42 1/0.95':g='0/0.03 0.5/0.52 1/1':b='0/0.12 0.5/0.50 1/0.85',eq=saturation=1.2:contrast=1.1" ``` ### Sepia (the real matrix, not a tint) ```bash -vf "colorchannelmixer=rr=.393:rg=.769:rb=.189:gr=.349:gg=.686:gb=.168:br=.272:bg=.534:bb=.131" ``` LUT form: `gen-luts.py --variants sepia`. For *toned* B&W instead (sepia highlights, neutral shadows): `hue=s=0,colorbalance=rh=.12:gh=.06`. ### Vintage film fade (Kodachrome-adjacent print fade) `gen-luts.py --variants film_fade`, or built-in `curves=preset=vintage`. Waveform floor ~5–8%, never 0. Sell it with `,noise=alls=6:allf=t+u`. ## Signature movie grades ### Blockbuster teal & orange (Transformers-era default) Skin warm, shadows teal — complementary separation. `gen-luts.py --variants teal_orange`, or: ```bash -vf "colorbalance=rs=-.06:bs=.08:rm=.04:bm=-.03:rh=.05:bh=-.06,eq=saturation=1.12" ``` Vectorscope: two lobes, skin ON the I-line. Fails on faceless footage and tungsten interiors. ### Mad Max: Fury Road (graphic-novel chrome) Not bleached apocalypse — the opposite: hyper-saturated teal/orange, crunchy contrast, sharpened grit: ```bash -vf "eq=saturation=1.4:contrast=1.25,colorbalance=rs=-.08:bs=.10:rm=.06:rh=.08:bh=-.08,unsharp=5:5:0.8" ``` (Night scenes in the film are graded BLUE day-for-night — combine with that recipe below.) ### The Matrix (digital green) Green pushed into midtones+shadows, slightly sick skin, crushed-but-readable: ```bash -vf "colorbalance=gs=.05:gm=.08:gh=.03,eq=saturation=0.85:contrast=1.10,curves=master='0/0.02 1/0.95'" ``` LUT form: `gen-luts.py --variants matrix_green`. ### Fincher (Se7en/Gone Girl murk) Cool, green-yellow undertone, HIGHLIGHTS PULLED DOWN (nothing ever blooms), shadow detail retained: ```bash -vf "colortemperature=temperature=6800,colorbalance=gs=.02:gm=.03:bh=-.04,eq=saturation=0.90:contrast=1.08,curves=master='0/0.01 0.8/0.72 1/0.90'" ``` Scope: waveform ceiling ~90%, never 100 — the pulled highlight IS the look. ### O Brother, Where Art Thou? (sepia wasteland) The first full-DI grade: desaturated, golden-burnt, green grass turned hay: ```bash -vf "eq=saturation=0.65,colorbalance=rm=.08:gm=.04:bm=-.08:rh=.06:bh=-.06,curves=master='0/0.03 1/0.95'" ``` ### Amélie (golden Paris) Warm gold + a deliberate green undertone, high saturation, cozy: ```bash -vf "colorbalance=rm=.06:gm=.05:bm=-.06:gs=.04,eq=saturation=1.25:contrast=1.08" ``` ### Blade Runner 2049 (orange smog) A monochromatic orange ENVELOPE — everything breathes the same dust: ```bash -vf "colorbalance=rm=.10:gm=.03:bm=-.12:rh=.08:bh=-.10,eq=saturation=0.90:contrast=1.05,curves=master='0/0.04 1/0.96'" ``` The interior-neon scenes are the [neon night](#neon-night--cyberpunk) recipe instead — the film alternates the two. ### Twilight (melodrama blue) The heavy blue wash: ```bash -vf "colortemperature=temperature=9500,colorbalance=bs=.08:bm=.06,eq=saturation=0.75:contrast=1.05" ``` ### In the Mood for Love (crimson & emerald) Reds and greens saturated past realism, everything else muted — color as character: ```bash -vf "vibrance=intensity=0.5:rbal=1.6:gbal=1.2:bbal=0.4,eq=contrast=1.10,curves=master='0/0.02 1/0.97'" ``` ### Fantastic Mr. Fox (autumn box) The whole frame inside yellows/browns/oranges, cool tones nearly banned: ```bash -vf "colorbalance=rm=.07:gm=.04:bm=-.10:rs=.03:bs=-.06:bh=-.08,eq=saturation=1.1:contrast=1.05" ``` ## Era looks ### Golden hour / filmic warm `gen-luts.py --variants golden_hour` (or `warm_filmic` subtler), or: ```bash -vf "colortemperature=temperature=4400,colorbalance=rh=.05:rm=.03:bh=-.03,eq=saturation=1.08,curves=master='0/0.02 1/0.97'" ``` Whites stay ≤ ~10% off-center on the vectorscope; highlights unclipped. ### Pastel (Wes Anderson) `gen-luts.py --variants pastel`, or: ```bash -vf "eq=saturation=0.72:contrast=0.88:brightness=0.04,curves=master='0/0.08 1/0.92'" ``` Half art-direction — only reads on composed frames. Waveform lives in 8–92%. ### 70s cinema (warm faded New Hollywood) Film-fade plus era warmth and soft contrast: ```bash -vf "curves=master='0/0.06 1/0.92',colorbalance=rm=.05:gm=.02:bh=-.04,eq=saturation=0.92:contrast=0.96,noise=alls=7:allf=t+u" ``` ### VHS / camcorder Color is a third of it — softness and chroma error carry it: ```bash -vf "eq=saturation=0.85:contrast=0.95,curves=master='0/0.06 1/0.94',gblur=sigma=0.6,chromashift=cbh=2:crh=-2,noise=alls=10:allf=t" ``` Full commitment: `scale=640:480,setsar=1` + `-ar 32000` audio. ## Genre & mood ### Film noir (B&W) ```bash -vf "hue=s=0,eq=contrast=1.25:brightness=-0.02,vignette=PI/5" ``` LUT: `gen-luts.py --variants noir_bw` (+ `vignette` at apply time — spatial ops don't fit in a LUT). Red-filter sky drama: `colorchannelmixer=.7:.2:.1` before `hue=s=0`. Waveform must use the FULL range — noir is contrast. ### Bleach bypass (war grit) `gen-luts.py --variants bleach_bypass`, or `eq=saturation=0.45:contrast=1.3,unsharp=5:5:0.4`. ### Horror sick-green Desaturated, green-poisoned shadows, everything slightly too dark: ```bash -vf "colorbalance=gs=.05:gm=.04:rs=-.03,eq=saturation=0.70:contrast=1.15:brightness=-0.05" ``` ### Grimdark battlefield (worked scope-extraction example) Extracted from a real graded reference (a 1080p fantasy-series trailer) with the [scope-matching ladder](#scope-matching-align-to-a-reference-clip-by-numbers) run in reverse — measure the reference with `signalstats`, then tune until your footage's numbers land in the same band. Measured (1,261 frames, cleaned per the caveats below): **SATAVG ≈ 7** (vivid footage runs 30–60), **UAVG 125.6 / VAVG 129.8** (a *warm-ash* cast — not blue), day exteriors **YAVG ≈ 110** with global ≈ 58 (night scenes), blacks ≈ 7: ```bash -vf "eq=saturation=0.33,colorbalance=rm=.02:gm=.012:bm=-.02,curves=master='0/0.03 0.5/0.42 1/0.95'" ``` LUT form: `gen-luts.py --variants grimdark` (calibrated to the day-exterior key; deepen `curves` mids toward `0.5/0.30` for the night cluster). vs Nordic noir: grimdark is warm-ash; Nordic is cool and flatter. **Measuring a trailer (or any edited reference) honestly:** 1. **Crop the letterbox first** (`cropdetect`, then `crop=`) — baked bars drag every luma stat down. 2. **Drop fades/title cards**: filter per-frame stats to `YAVG > 25` before averaging, else cut transitions poison the mean. 3. **Expect scene clusters**: shows grade per scene-type (this reference's banquet interiors are warm amber, nothing like its exteriors). The *chroma fingerprint* (SATAVG + U/V cast) is usually consistent — transfer that globally; match *key* (YAVG) per scene-type, never to the global mean. 4. Verify a transfer by re-measuring the graded result: `ffmpeg -i graded.mp4 -vf signalstats,metadata=print:file=- -f null -`. ### Nordic noir (Scandinavian bleak) Desaturated, cool, FLAT — the anti-blockbuster: ```bash -vf "colortemperature=temperature=7500,eq=saturation=0.65:contrast=0.95,curves=master='0/0.04 1/0.90'" ``` vs Twilight blue: this one is low-contrast and barely saturated; Twilight is a saturated blue *wash*. ### Romance soft glow Warm, lifted, gentle bloom on highlights: ```bash -filter_complex "[0:v]colorbalance=rh=.04:rm=.02,eq=saturation=1.05:contrast=0.94,curves=master='0/0.05 1/0.97'[base];[base]split[a][b];[b]gblur=sigma=8[soft];[a][soft]blend=all_mode=screen:all_opacity=0.18" ``` ### Neon night / cyberpunk ```bash -vf "eq=saturation=1.25:contrast=1.1,colorbalance=bs=.15:bm=.05:rs=-.05,curves=b='0/0.08 1/1':r='0/0 1/0.95'" ``` Needs practicals/neon in frame; on daylight it's just a bad cool cast. ### Day-for-night ```bash -vf "eq=brightness=-0.15:saturation=0.55,colorbalance=bs=0.25:bm=0.12,curves=master='0/0 0.7/0.45 1/0.8'" ``` No visible sky/sun, no blown highlights, or it never sells. ## Stylized effects ### Sin City selective color Everything monochrome EXCEPT one hue (`colorhold` keeps a color, greys the rest): ```bash -vf "colorhold=color=red:similarity=0.35:blend=0.1,eq=contrast=1.3" ``` Works for any anchor color (`color=0x00a0ff` etc.). High-contrast B&W base is what makes the held color violent. ### Tone maps: monotone / duotone / tritone One mechanism, three intensities: desaturate, then re-map the tonal axis onto 2 or 3 color stops with per-channel curves. **Chroma of the look = how far the stops sit from the neutral grey axis** — monotones barely leave it (darkroom chemical tones), muted duotones use tertiary/greyed pairs, poster duotones live far out. Every variant below is also a parametric LUT: `gen-luts.py --variants mono_selenium,tri_tobacco --previews footage.mp4`. The chain template (3 stops; drop the `0.5/` midpoints for a 2-stop duotone — stop values are the color's channels /255): ```bash -vf "hue=s=0,curves=r='0/<Rs> 0.5/<Rm> 1/<Rh>':g='0/<Gs> 0.5/<Gm> 1/<Gh>':b='0/<Bs> 0.5/<Bm> 1/<Bh>'" # worked example — selenium monotone: -vf "hue=s=0,curves=r='0/0.05 0.5/0.48 1/0.96':g='0/0.04 0.5/0.46 1/0.95':b='0/0.07 0.5/0.52 1/0.97'" ``` | Variant | Stops (shadow → [mid →] highlight) | Use | |---|---|---| | **Monotones** (single chemical tone, near-grey chroma) | | | | `mono_selenium` | (.05,.04,.07) → (.48,.46,.52) → (.96,.95,.97) | Fine-print B&W with the cool violet selenium whisper | | `mono_platinum` | (.07,.07,.06) → (.52,.51,.49) → (.97,.96,.94) | Warm-neutral platinum print; the most archival-looking B&W | | `mono_coffee` | (.08,.05,.03) → (.55,.47,.40) → (.96,.92,.87) | Warm brown tone, gentler than sepia | | `mono_steel` | (.04,.06,.09) → (.46,.50,.55) → (.94,.96,.98) | Cool documentary B&W | | **Muted duotones** (tertiary pairs) | | | | `duo_ash_rose` | (.23,.20,.22) → (.85,.78,.76) | Fashion/editorial soft; flattering on skin | | `duo_olive_bone` | (.18,.20,.14) → (.90,.88,.81) | Field/military/heritage | | `duo_petrol_paper` | (.12,.23,.24) → (.93,.91,.86) | Calm tech/industrial editorial | | `duo_indigo_parchment` | (.16,.23,.33) → (.91,.89,.82) | Faded-cyanotype archival — the muted cousin of `duo_cyanotype` | | `duo_slate_ice` | (.11,.15,.20) → (.95,.97,.98) | Corporate/tech-keynote neutral | | **Poster duotones** (high chroma, deliberate) | | | | `duo_navy` | (.05,.08,.25) → (.98,.93,.80) | Editorial/magazine classic | | `duo_cyanotype` | (.04,.16,.29) → (.92,.96,1.0) | Blueprint/architectural | | `duo_sunset` | (.23,.06,.36) → (1.0,.78,.34) | Festival poster | | `duo_forest` | (.06,.24,.18) → (.91,.85,.63) | Organic/outdoor brand | | `duo_crimson` | (.10,.02,.03) → (1.0,.88,.86) | Sports/thriller key art | | `duo_synthwave` | (.35,.06,.42) → (.42,.91,1.0) | Retro-tech/vaporwave | | **Tritones** (distinct shadow/mid/highlight hues) | | | | `tri_split_classic` | (.06,.07,.12) → (.50,.49,.48) → (.98,.94,.86) | THE darkroom split: cool shadows, neutral mids, warm highlights | | `tri_tobacco` | (.05,.04,.02) → (.45,.40,.28) → (.95,.88,.70) | Western/whiskey-ad warmth with real blacks | | `tri_arctic` | (.03,.05,.09) → (.42,.50,.58) → (.93,.97,1.0) | Expedition/documentary cold | Tuning rules: contrast BEFORE the map widens the spread (`eq=contrast=1.1,hue=s=0,...`); to mute any variant, pull its stops toward the grey diagonal (average each stop with its own luma); the mid stop is where skin lives — keep it near-neutral unless the face *is* the poster. ## Matching a look you can see but can't name ### Hald-CLUT: grade one frame anywhere, get a video LUT for free A Hald image is a LUT unrolled into a PNG — any **global** color edit applied to it becomes applicable to video: ```bash # 1. identity Hald (level 8 = 64^3 lattice) ffmpeg -f lavfi -i haldclutsrc=8 -frames:v 1 hald.png # 2. open hald.png in ANY photo editor with a still from your footage; design # the look on the still; apply the IDENTICAL adjustments to hald.png # 3. the edited Hald IS your LUT: ffmpeg -i in.mp4 -i hald_graded.png -filter_complex "[0:v][1:v]haldclut" \ -c:v libx264 -crf 18 -c:a copy graded.mp4 ``` **Stealing a look**: any editor preset / Lightroom recipe / .acv curve applied to the Hald identity is thereby extracted as a LUT. Photoshop curves apply directly too: `curves=psfile=their_grade.acv`. **The one rule**: only GLOBAL color ops survive — curves, levels, WB, HSL, balance, saturation. Spatial ops (vignette, sharpen, local contrast, dehaze, grain, healing) corrupt the lattice; do those in the filter chain. Fidelity note: visually identical, not bit-identical — the 8-bit lattice quantizes (measured SSIM ≈ 0.95 vs the same chain applied directly). For very steep curves prefer the direct chain or a 16-bit TIFF Hald. ### Scope-matching: align to a reference clip by numbers ffmpeg has no automatic shot-matcher. **The governing rule: transfer the chroma fingerprint (SATAVG + U/V cast) globally — it's what stays constant across a graded work; match key (YAVG) per scene-type, never to the global mean** (night scenes drag any edited reference's average far below what a day scene should hit — see the grimdark example's measurement checklist). The manual ladder (scope views from [color-grading.md](color-grading.md), reference and target side-by-side via `hstack`): 1. **Black/white points** (waveform): `curves=master='0/<floor> 1/<ceil>'`. 2. **Midtone brightness** (waveform mass): `eq=gamma=`. 3. **Cast** (vectorscope centroid): `colortemperature` + `colorbalance`. 4. **Saturation** (vectorscope spread): `eq=saturation=`. 5. **Verify on skin**: both clips' faces hug the I-line equally. Order matters — each step changes the reading of the ones after it; never start with saturation. ## Applying any of this at scale One look across a project = bake the chain into a LUT once (`gen-luts.py` variant, or render the chain through a Hald identity and use `haldclut` everywhere). Match per-clip exposure FIRST with `eq`, apply the shared look second — the batch-consistency section of [color-grading.md](color-grading.md). Composite looks (halation, bloom) keep their spatial half in the filter chain; only their color half bakes into the LUT. -
quality-metrics.md 3.1 KB
# Quality metrics — VMAF, SSIM, PSNR, visual A/B ## The tool ```bash python skills/ffmpeg-ops/scripts/quality-compare.py original.mp4 encoded.mp4 \ --metrics vmaf --min-vmaf 90 # exit 10 below threshold -> branch on it ``` Handles resolution mismatch (auto-scales distorted to reference), parses the filters' log output, returns one envelope. Use it instead of hand-running the metric filters. ## Reading the numbers | Metric | Transparent | Good | Visible degradation | Notes | |---|---|---|---|---| | **VMAF** | ≥ 93 | 85–93 | < 80 | Perceptual model (Netflix); the one to trust. Trained at 1080p living-room viewing | | **SSIM** | ≥ 0.99 | 0.97–0.99 | < 0.95 | Structural; cheap, no libvmaf needed | | **PSNR** | ≥ 45 dB | 38–45 | < 35 | Naive signal ratio; only comparable between encodes of the *same* source | - Check VMAF **min** (worst moment), not just mean — a 95-mean encode with a 62-min scene has a visible glitch. `quality-compare.py --json | jq '.data.vmaf'` reports mean/min/harmonic_mean. - VMAF on 4K-viewed-at-4K: use the 4K model variant if available in your build; otherwise treat scores as optimistic. - Comparing two *different sources* with PSNR/SSIM is meaningless; metrics judge an encode against *its own* reference. ## Workflow: tune CRF mechanically Find the highest CRF (smallest file) that stays above your VMAF floor: ```bash for crf in 20 23 26 29; do ffmpeg -y -v error -i ref.mp4 -c:v libx264 -crf $crf -preset slow -an "t$crf.mp4" python skills/ffmpeg-ops/scripts/quality-compare.py ref.mp4 "t$crf.mp4" \ --metrics vmaf --json | jq -r --arg c $crf '"crf=\($c) vmaf=\(.data.vmaf.mean) min=\(.data.vmaf.min)"' done ``` Encode a representative 60–90 s slice, not the whole file (`-ss <busy-section> -t 60` on both reference cut and encodes — cut the reference first so they align). ## Visual A/B (the human half) ```bash # side-by-side (label which is which!) ffmpeg -i ref.mp4 -i enc.mp4 -filter_complex \ "[0:v]drawtext=text='REF':fontsize=36:fontcolor=white:box=1:boxcolor=black@0.5:x=24:y=24[a]; [1:v]drawtext=text='ENC':fontsize=36:fontcolor=white:box=1:boxcolor=black@0.5:x=24:y=24[b]; [a][b]hstack" -c:v libx264 -crf 16 -an ab.mp4 # difference view — what the encoder actually changed (grey = identical): ffmpeg -i ref.mp4 -i enc.mp4 -filter_complex "blend=all_mode=difference,eq=brightness=0.3" -an diff.mp4 # wipe split-screen (left=ref, right=enc, hard seam at 50%): ffmpeg -i ref.mp4 -i enc.mp4 -filter_complex "[1:v][0:v]overlay=x='-W/2'" -an wipe.mp4 ``` Where codecs fail first (look here in the A/B): dark gradients (banding), fast motion (blocking), fine texture like grass/water (smearing), red saturated areas (chroma 4:2:0). ## When the numbers and your eyes disagree Believe your eyes, then find out why: wrong reference alignment (an offset frame ruins every metric — verify identical frame counts), range/matrix mismatch ([color-hdr.md](color-hdr.md)) penalizing colors uniformly, or grain (encoders denoise; metrics partially forgive it, viewers notice). Banding specifically is under-penalized by all three metrics — check dark scenes by eye at viewing brightness. -
restoration.md 2.7 KB
# Restoration — deinterlace, denoise, deband, stabilize, repair Order of operations: **deinterlace → denoise → deband → stabilize → grade → sharpen → encode.** (Stabilize after denoise: noise defeats motion estimation.) ## Deinterlace (combing artifacts on motion = interlaced source) ```bash # probe says field_order=tt/bb (or you see combing): ffmpeg -i dvd.vob -vf "bwdif=mode=send_field" -c:v libx264 -crf 19 out.mp4 ``` `bwdif` beats the older `yadif`; `mode=send_field` doubles frame rate (50i→50p, correct for sports/motion), `mode=send_frame` keeps it (fine for films). **Telecined film** (24fps in 30i — duplicate-ish frames in a 3:2 pattern) wants inverse telecine instead: `-vf "fieldmatch,decimate"`. ## Denoise ```bash -vf "hqdn3d=4:3:6:4.5" # fast, general (luma-spatial:chroma-spatial:luma-temporal:chroma-temporal) -vf "nlmeans=s=4" # much slower, much better on heavy noise -vf "atadenoise" # temporal-only; preserves detail on static shots ``` Start gentle (hqdn3d defaults), inspect at 100% zoom, increase until noise is acceptable — over-denoising produces the plastic-skin look that's worse than grain. For *intentional* film grain, don't denoise; encode with `-tune grain` ([encoding.md](encoding.md)). ## Deband (visible steps in skies/gradients) ```bash -vf "deband=1thr=0.015:2thr=0.015:3thr=0.015" # prevention on re-encode: 10-bit output kills most banding at the source -pix_fmt yuv420p10le # (HEVC/AV1 — see color-hdr.md) ``` ## Stabilize (vidstab two-pass; needs libvidstab — check capability-scan) ```bash # pass 1: analyze motion -> transforms.trf ffmpeg -i shaky.mp4 -vf "vidstabdetect=shakiness=6:accuracy=15:result=transforms.trf" -f null - # pass 2: apply + crop the wobble margin + mild sharpen ffmpeg -i shaky.mp4 -vf \ "vidstabtransform=input=transforms.trf:zoom=2:smoothing=24,unsharp=5:5:0.6" \ -c:v libx264 -crf 19 -c:a copy stable.mp4 ``` `smoothing` ≈ frames of camera-path averaging (higher = floatier); `zoom` crops the edges that stabilization exposes. The single-pass `deshake` filter is a quick-and-dirty fallback when libvidstab is absent. ## Old/odd footage misc ```bash # wrong speed (PAL 25fps of a 23.976 film, pitch off): retime v+a together -filter_complex "[0:v]setpts=PTS*25/23.976[v];[0:a]atempo=0.95904[a]" # VHS-style chroma bleed: mild chroma denoise + slight desat -vf "hqdn3d=0:6:0:6,eq=saturation=0.92" # duplicate-frame removal (bad pulldown, stuttery web rips): -vf "mpdecimate" -fps_mode vfr ``` ## Audio repair Lives in [audio.md](audio.md) (highpass → afftdn → declick → compand chain). For damaged *files* (truncated/corrupt) see [analysis-validation.md](analysis-validation.md) — remux first (`-c copy -fflags +genpts`), repair second. -
streaming-hls.md 2.9 KB
# Streaming — HLS/DASH packaging, ABR ladders, live restream ## Single-rendition HLS VOD (the 80% case) ```bash ffmpeg -i in.mp4 -c:v libx264 -crf 21 -preset slow -pix_fmt yuv420p \ -g 180 -keyint_min 180 -sc_threshold 0 \ -c:a aac -b:a 128k -ar 48000 \ -f hls -hls_time 6 -hls_playlist_type vod \ -hls_segment_filename 'out/seg_%04d.ts' out/index.m3u8 ``` The keyframe rule is the part everyone misses: **segment boundaries must be keyframes**, so `-g`/`-keyint_min` = fps × hls_time (here 30×6=180) and `-sc_threshold 0` stops scene-detection from inserting extras. Without this, segment durations drift and players stall on seeks. fMP4 segments instead of TS (required for HEVC-in-HLS, nicer for CMAF): `-hls_segment_type fmp4`. ## ABR ladder (multi-rendition) Ladder data: [../assets/hls-ladder.json](../assets/hls-ladder.json) — trim to 3 rungs for non-broadcast use. One-command master playlist via `-var_stream_map`: ```bash ffmpeg -i in.mp4 \ -filter_complex "[0:v]split=3[v1][v2][v3];[v1]scale=-2:1080[v1o];[v2]scale=-2:720[v2o];[v3]scale=-2:360[v3o]" \ -map "[v1o]" -c:v:0 libx264 -b:v:0 6000k -maxrate:v:0 6600k -bufsize:v:0 12000k \ -map "[v2o]" -c:v:1 libx264 -b:v:1 3000k -maxrate:v:1 3300k -bufsize:v:1 6000k \ -map "[v3o]" -c:v:2 libx264 -b:v:2 730k -maxrate:v:2 800k -bufsize:v:2 1460k \ -map a:0 -map a:0 -map a:0 -c:a aac -b:a 128k -ar 48000 \ -preset slow -pix_fmt yuv420p -g 180 -keyint_min 180 -sc_threshold 0 \ -f hls -hls_time 6 -hls_playlist_type vod \ -master_pl_name master.m3u8 \ -var_stream_map "v:0,a:0 v:1,a:1 v:2,a:2" \ -hls_segment_filename 'out/%v/seg_%04d.ts' 'out/%v/index.m3u8' ``` ABR uses **capped bitrate** (`-b:v` + `-maxrate` + `-bufsize` ≈ 2× maxrate), not CRF — the ladder's promise to the player is a bandwidth, not a quality. ## DASH Same encode discipline; `-f dash`: ```bash ffmpeg -i in.mp4 ... -f dash -seg_duration 6 -use_template 1 -use_timeline 1 out/manifest.mpd ``` For both-HLS-and-DASH from one encode, encode renditions to fMP4 once and package with a dedicated packager (shaka-packager) rather than encoding twice. ## Live restream (RTMP push) ```bash # screen/webcam/file -> YouTube/Twitch ingest. zerolatency + CBR-ish + 2s GOP: ffmpeg -re -i source.mp4 -c:v libx264 -preset veryfast -tune zerolatency \ -b:v 4500k -maxrate 4500k -bufsize 9000k -pix_fmt yuv420p -g 60 \ -c:a aac -b:a 128k -ar 44100 \ -f flv rtmp://a.rtmp.youtube.com/live2/STREAM_KEY ``` `-re` paces a *file* to realtime (never use it for live capture inputs). NVENC (`h264_nvenc -preset p4 -tune ll`) is the right call here — encode speed matters more than per-bit quality. Stream keys are secrets: env var, not command line, on shared machines. ## Serving HLS locally (testing) Any static server works (`python -m http.server`) — HLS is just files + correct MIME (`.m3u8` = application/vnd.apple.mpegurl, `.ts` = video/mp2t). Browsers other than Safari need hls.js; quick check without a page: `ffplay out/index.m3u8`. -
stt-whisper.md 3.9 KB
# STT / Whisper prep — audio in, transcripts out ffmpeg is the universal front-end for Whisper-family transcription; the prep step is where transcription quality is silently won or lost. ## The canonical extraction Whisper models consume 16 kHz mono. Resampling in ffmpeg (not in the STT tool) is faster and deterministic: ```bash ffmpeg -i in.mp4 -vn -ac 1 -ar 16000 -c:a pcm_s16le stt.wav # no temp file — pipe raw PCM straight in (whisper.cpp shown): ffmpeg -v error -i in.mp4 -vn -ac 1 -ar 16000 -f s16le - | whisper-cli -m model.bin -f - ``` ## Pre-cleanup: what helps and what hurts | Filter | Effect on STT accuracy | |---|---| | `loudnorm` / `dynaudnorm` | Helps quiet/uneven recordings — Whisper mis-segments very quiet audio | | `highpass=f=100` | Helps rumble/handling noise; harmless otherwise | | `afftdn` (denoise) | Helps **only** on genuinely noisy audio; on clean audio it smears consonants and *hurts* | | Aggressive `silenceremove` | **Hurts** — Whisper uses silence for sentence segmentation; removing it merges sentences and breaks timestamps relative to the original media | Rule: normalize loudness, high-pass at 100 Hz, denoise only when you can hear the noise. Never strip silence from audio you'll want timestamps against. ## Chunking long audio Chunk **on silence boundaries, never mid-word** — and overlap is unnecessary when the boundaries are real silences: ```bash python skills/ffmpeg-ops/scripts/detect-segments.py --silence --min-silence 0.6 \ --json long.mp4 | jq -r '.data.speech[] | "\(.start) \(.end)"' | while read -r s e; do ffmpeg -v error -ss "$s" -to "$e" -i long.mp4 -vn -ac 1 -ar 16000 \ -c:a pcm_s16le "chunks/chunk_${s}.wav" done ``` Chunk filenames carry the source offset, so chunk-local timestamps convert back to source time by adding `s`. Group speech segments into ~5–10 min batches for parallel transcription (one agent/process per batch). ## The transcript-JSON contract Normalize every engine's output into this shape (the WhisperX word form) — it is the contract that shot selection ([edit-as-code.md](edit-as-code.md)), cut verification, caption timing, and overlay placement all read: ```json { "words": [ { "word": "Hey", "start": 1.02, "end": 1.50 }, { "word": "it's", "start": 1.90, "end": 2.04 } ], "segments": [ { "text": "Hey it's ...", "start": 1.02, "end": 3.38 } ] } ``` **ASR spelling is unreliable; timings are the product.** A name misheard as "Sark" still carries correct timestamps — match phrases fuzzily, trust the times. ## Engine notes - **whisper.cpp** — local, fast, no Python; word timestamps approximate (segment cross-fade heuristics). - **faster-whisper** — local Python (CTranslate2), `word_timestamps=True`; good default. - **WhisperX** — adds wav2vec2 forced alignment on top of Whisper: word timestamps to ±50 ms (vanilla Whisper ≈ ±500 ms), plus VAD pre-filter and optional diarization. **Use when timestamps drive cuts or captions.** - Managed APIs (ElevenLabs, Deepgram, AssemblyAI) — fine; normalize their response into the contract above. ## Round trip to subtitles STT output → SRT/VTT → mux or burn ([subtitles.md](subtitles.md)): ```bash # most engines emit SRT directly; converting is one command anyway: ffmpeg -i transcript.vtt captions.srt ffmpeg -i in.mp4 -i captions.srt -map 0 -map 1 -c copy -c:s mov_text captioned.mp4 ``` ## The summarisation pipeline (daily-driver workflow) ``` source (file or ytdlp audio-only) → ffmpeg 16k mono extraction (above) → transcribe (engine of choice, word JSON) → THE AGENT summarises the transcript ← ffmpeg's job ended one step ago → optional visual pass: detect-segments.py --scenes + a contact sheet (images-gif.md) → chapter list with thumbnails ``` For downloaded sources, prefer `yt-dlp -x --audio-format opus URL` (or `--download-sections` for a time range) so you never pull video bytes you only needed audio from. -
subtitles.md 2.6 KB
# Subtitles — burn vs soft, styling, extraction ## Decision | Need | Method | |---|---| | Toggleable, instant, preserves quality | **Soft** (mux as a stream, `-c copy`) | | Always visible (social video, players without sub support) | **Burn-in** (re-encode, `subtitles` filter) | | Styled karaoke/positioned text | ASS soft in MKV, or burn | ## Soft subtitles (mux) ```bash ffmpeg -i in.mp4 -i subs.srt -map 0 -map 1 -c copy -c:s mov_text out.mp4 # mp4 ffmpeg -i in.mkv -i subs.srt -map 0 -map 1 -c copy -c:s srt out.mkv # mkv (srt/ass) # language tag + default flag (players auto-select): ... -metadata:s:s:0 language=eng -disposition:s:0 default ``` MP4 only carries `mov_text` (and loses ASS styling); MKV carries srt/ass/pgs natively. WebVTT for web (`-c:s webvtt`, .vtt). ## Burn-in ```bash # cd to the subtitle's directory first — the filter's path escaping is the worst # quoting trap in ffmpeg (a Windows drive colon needs C\\:/ escaping INSIDE the arg) ffmpeg -i in.mp4 -vf "subtitles=subs.srt" -c:v libx264 -crf 20 -c:a copy out.mp4 # styled burn (libass force_style; fontconfig resolves the family name): -vf "subtitles=subs.srt:force_style='FontName=Arial,FontSize=28,PrimaryColour=&H00FFFFFF,OutlineColour=&H00000000,Outline=2,MarginV=40'" # burn the EMBEDDED subtitle track of an mkv (note: input path, stream index si): -vf "subtitles=in.mkv:si=0" # bitmap subs (PGS/dvd_subtitle) cannot go through `subtitles` — overlay them: -filter_complex "[0:v][0:s:0]overlay" ``` `force_style` colors are ASS `&HAABBGGRR` (blue-green-red, not RGB; alpha 00 = opaque). ## Extract / convert ```bash ffprobe -v error -show_entries stream=index,codec_name:stream_tags=language \ -select_streams s -of csv=p=0 in.mkv # what sub tracks exist ffmpeg -i in.mkv -map 0:s:0 subs.srt # extract first text track ffmpeg -i subs.srt subs.vtt # srt <-> vtt <-> ass conversion ``` Bitmap tracks (pgs, dvd_subtitle) can't convert to text via ffmpeg — that's an OCR job (external tooling). ## Timing repair ```bash # subs 2.5s late -> shift earlier: ffmpeg -itsoffset -2.5 -i subs.srt -c copy shifted.srt ``` For *rate* drift (23.976 vs 25 fps subs), retiming = multiply timestamps — external tools or regenerate from STT ([stt-whisper.md](stt-whisper.md)). ## From STT Whisper-family engines emit SRT/VTT directly; the round trip (extract audio → transcribe → mux back) is in [stt-whisper.md](stt-whisper.md). Caption-quality rule for generated subs: ≤ 2 lines, ≤ ~42 chars/line, segments split on the word-level timestamps at phrase boundaries — not the engine's raw 30-word blobs. -
trim-concat.md 4.1 KB
# Trim & Concat — seek semantics, keyframes, joining ## `-ss` semantics (the most misunderstood flag in ffmpeg) | Placement | With `-c copy` | With re-encode | |---|---|---| | **Before `-i`** (input seek) | Fast; **snaps to the previous keyframe** — start can be seconds early, or players show frozen/black until the first keyframe | Fast **and frame-accurate** (decodes from the prior keyframe, discards up to the target) | | **After `-i`** (output seek) | Decodes everything from 0:00 then discards — slow, accurate | Slow, accurate — *no advantage* over input seek on modern ffmpeg | **Modern rule: put `-ss` before `-i` always.** The "put it after for accuracy" advice predates ffmpeg 2.1 and now only costs time. `-to` vs `-t`: `-to` = absolute end position, `-t` = duration. **Keep `-ss` and `-to` on the same side of `-i`.** With input-side `-ss` and *output-side* `-to`, timestamps have already been reset at the seek point, so `-to 60` means "60s after the cut start", not "at 60s in the source" — a silent off-by-`ss` error. ## Keyframes and copy cuts A stream-copied cut can only begin at a keyframe (IDR). Typical delivery files have keyframes every 2–10 s, so a copy cut at an arbitrary point either: 1. snaps the start earlier (most players), or 2. keeps audio from the requested point but shows frozen video until the next keyframe (some players). Decide mechanically: ```bash python skills/ffmpeg-ops/scripts/probe-media.py --keyframes-near 92.5 in.mp4 # copy_cut_drift_s tells you how far the copy cut would land from your target ``` Drift acceptable → copy. Not → re-encode just that cut (`-crf 18` keeps it visually identical). For many cuts from one source, an all-intra mezzanine (see [encoding.md](encoding.md)) makes *every* point copy-safe. Always add `-avoid_negative_ts make_zero` to copy cuts — some muxers otherwise write leading negative timestamps that desync players. ## The three concats | Method | When | Cost | |---|---|---| | **concat demuxer** | Same codec, resolution, fps, timebase (e.g. segments you cut from one source) | zero — stream copy | | **concat filter** | Different codecs/sizes/fps | full re-encode | | **concat protocol** | MPEG-TS only (`concat:a.ts\|b.ts`) — rarely what you want | zero | ```bash # demuxer (the workhorse) printf "file '%s'\n" a.mp4 b.mp4 c.mp4 > concat.txt ffmpeg -f concat -safe 0 -i concat.txt -c copy -movflags +faststart out.mp4 # filter (mixed sources) — normalize geometry inline ffmpeg -i a.mp4 -i b.mov -filter_complex \ "[0:v]scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,fps=30[v0]; [1:v]scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,fps=30[v1]; [v0][0:a][v1][1:a]concat=n=2:v=1:a=1[v][a]" \ -map "[v]" -map "[a]" -c:v libx264 -crf 20 -c:a aac out.mp4 ``` concat.txt paths are relative **to the concat.txt file**, not the CWD. `-safe 0` is required for absolute paths. Windows paths work with forward slashes: `file 'X:/clips/a.mp4'`. **Audio gotcha:** mismatched sample rates/channel layouts break the demuxer too — not just video params. When in doubt: probe both, or re-encode via the filter. ## Removing a middle section Two keeps + concat (simple, recommended), or one command with trim filters: ```bash # keep 0-60 and 120-end in one pass (re-encode) ffmpeg -i in.mp4 -filter_complex \ "[0:v]trim=0:60,setpts=PTS-STARTPTS[v0];[0:a]atrim=0:60,asetpts=PTS-STARTPTS[a0]; [0:v]trim=start=120,setpts=PTS-STARTPTS[v1];[0:a]atrim=start=120,asetpts=PTS-STARTPTS[a1]; [v0][a0][v1][a1]concat=n=2:v=1:a=1[v][a]" \ -map "[v]" -map "[a]" -c:v libx264 -crf 18 -c:a aac out.mp4 ``` `setpts=PTS-STARTPTS` after every trim is mandatory — trim keeps original timestamps and the concat misbehaves without the reset. For 3+ cuts, stop hand-writing graphs: author an EDL and use `cut-from-edl.py` ([edit-as-code.md](edit-as-code.md)). ## Segmenting (the reverse of concat) ```bash # split into ~5-minute pieces at keyframes, no re-encode ffmpeg -i in.mp4 -f segment -segment_time 300 -reset_timestamps 1 -c copy part%03d.mp4 ``` Segment boundaries snap to keyframes in copy mode — pieces won't be exactly 300s. -
visualization.md 3.2 KB
# Visualization — audio-reactive video, audiograms, spectrograms Turning sound into pixels: podcast clips for social, waveform "audiograms", debugging audio by looking at it. ## Waveform video (the podcast audiogram) ```bash # scrolling waveform over a brand background + episode title: ffmpeg -i episode.mp3 -loop 1 -i bg_1080x1920.png -filter_complex \ "[0:a]showwaves=s=1080x300:mode=cline:colors=white:rate=30[w]; [1:v][w]overlay=0:1200:shortest=1,drawtext=text='EP 42 — Title':fontsize=56:fontcolor=white:x=(w-text_w)/2:y=320" \ -c:v libx264 -crf 21 -preset fast -pix_fmt yuv420p -c:a aac -b:a 128k -shortest audiogram.mp4 ``` `shortest=1` on the overlay + `-shortest` at the end stop the looped image from running forever. `showwaves` modes: `cline` (filled, the podcast look), `line`, `p2p`, `point`. ## Spectrum styles ```bash # frequency bars (the "visualizer" look): "[0:a]showfreqs=s=1280x420:mode=bar:fscale=log[v]" # scrolling spectrogram (also the debugging view — see below): "[0:a]showspectrum=s=1280x720:mode=combined:color=intensity:scale=log:slide=scroll[v]" # musical/CQT spectrum (notes align to rows — lovely for music): "[0:a]showcqt=s=1280x720[v]" # minimal volume meter / phase scope: "[0:a]avectorscope=s=720x720:zoom=1.5[v]" ``` All consume `[0:a]` and produce a video stream — overlay/hstack them like any other video ([filtergraph.md](filtergraph.md)). ## Static waveform / spectrogram images ```bash # waveform PNG (one image of the whole file — episode art, quick inspection): ffmpeg -i in.mp3 -filter_complex "showwavespic=s=1920x480:colors=#3aa3ff" -frames:v 1 wave.png # spectrogram PNG — the audio-debugging x-ray: ffmpeg -i in.wav -lavfi "showspectrumpic=s=1920x1080:scale=log" -frames:v 1 spec.png ``` Reading the spectrogram: a hard ceiling at ~16 kHz = the file was once a lossy 128k MP3 regardless of its current extension; mains hum = a solid line at 50/60 Hz (kill with `highpass`); clicks = vertical needles. Faster than ears for "is this 'lossless' file actually lossless". ## Audio-reactive overlays (beyond fixed shapes) ffmpeg-only reactivity is limited to the built-in scopes. For brand-grade audio-reactive motion (pulsing logos, beat-synced glow), render with a composition tool (hyperframes' audio-reactive bindings or Remotion's `useAudioData`) and use ffmpeg for the I/O around it: extract the audio (`-vn`), supply stems, encode/package the rendered result ([encoding.md](encoding.md)). ## Comparison grids (encode A/B, model-output review) ```bash # 2x2 labelled grid of four variants: ffmpeg -i a.mp4 -i b.mp4 -i c.mp4 -i d.mp4 -filter_complex \ "[0:v]drawtext=text='crf20':fontsize=36:fontcolor=white:box=1:boxcolor=black@0.5:x=12:y=12[a]; [1:v]drawtext=text='crf26':fontsize=36:fontcolor=white:box=1:boxcolor=black@0.5:x=12:y=12[b]; [2:v]drawtext=text='nvenc':fontsize=36:fontcolor=white:box=1:boxcolor=black@0.5:x=12:y=12[c]; [3:v]drawtext=text='av1':fontsize=36:fontcolor=white:box=1:boxcolor=black@0.5:x=12:y=12[d]; [a][b][c][d]xstack=inputs=4:layout=0_0|w0_0|0_h0|w0_h0" -an grid.mp4 ``` Inputs must share dimensions (scale first if not). The 2-input case (`hstack` + difference blend) lives in [quality-metrics.md](quality-metrics.md).
-
-
scripts
-
capability-scan.sh 5.9 KB
#!/usr/bin/env bash # What can THIS ffmpeg build actually do — encoders, hwaccels, key filters. # # Listing an encoder is not the same as it working: hardware encoders (NVENC/QSV/ # AMF/VideoToolbox/VAAPI) routinely appear in `-encoders` yet fail at runtime on # driver/device mismatches. Default mode therefore PROOF-ENCODES 10 frames of # lavfi testsrc2 through every present hw encoder; --quick skips that (list-only). # # Usage: capability-scan.sh [--quick] [--json] [-q] # Input: none (inspects the ffmpeg on PATH) # Output: stdout = TSV records (kind, name, listed, verified), or --json envelope # (schema claude-mods.ffmpeg-ops.capability/v1) # Stderr: headers, progress, errors # Exit: 0 ok, 2 usage, 5 ffmpeg missing (jq missing for --json), # 10 at least one LISTED hw encoder FAILED its proof-encode # # Examples: # capability-scan.sh # capability-scan.sh --quick # capability-scan.sh --json | jq '.data.encoders[] | select(.hw and .listed)' # capability-scan.sh --json | jq -r '.data.recommended_hw // "none"' set -uo pipefail EXIT_OK=0; EXIT_USAGE=2; EXIT_MISSING_DEP=5; EXIT_FAILED_VERIFY=10 SCHEMA="claude-mods.ffmpeg-ops.capability/v1" QUICK=0; JSON=0; QUIET=0 while [[ $# -gt 0 ]]; do case "$1" in --quick) QUICK=1 ;; --json) JSON=1 ;; -q|--quiet) QUIET=1 ;; -h|--help) sed -n '2,23p' "$0" | sed 's/^# \{0,1\}//'; exit "$EXIT_OK" ;; *) echo "ERROR: unknown argument: $1 (try --help)" >&2; exit "$EXIT_USAGE" ;; esac shift done command -v ffmpeg >/dev/null 2>&1 || { [[ "$JSON" -eq 1 ]] && echo '{"error":{"code":"MISSING_DEPENDENCY","message":"ffmpeg not on PATH"}}' echo "ERROR: ffmpeg not found on PATH" >&2; exit "$EXIT_MISSING_DEP"; } HAS_JQ=0; command -v jq >/dev/null 2>&1 && HAS_JQ=1 [[ "$JSON" -eq 1 && "$HAS_JQ" -eq 0 ]] && { echo '{"error":{"code":"MISSING_DEPENDENCY","message":"jq required for --json"}}' echo "ERROR: jq required for --json" >&2; exit "$EXIT_MISSING_DEP"; } emit() { [[ "$QUIET" -eq 1 ]] && return 0; printf '%s\n' "$1" >&2; } VERSION="$(ffmpeg -hide_banner -version 2>/dev/null | head -1)" ENCODERS_RAW="$(ffmpeg -hide_banner -encoders 2>/dev/null)" HWACCELS="$(ffmpeg -hide_banner -hwaccels 2>/dev/null | tail -n +2 | tr -d ' ' | grep -v '^$' || true)" FILTERS_RAW="$(ffmpeg -hide_banner -filters 2>/dev/null)" emit "== capability-scan: $VERSION" # Hardware encoders worth knowing about, in rough preference order per vendor. HW_ENCODERS=(h264_nvenc hevc_nvenc av1_nvenc h264_qsv hevc_qsv av1_qsv h264_amf hevc_amf av1_amf h264_videotoolbox hevc_videotoolbox h264_vaapi hevc_vaapi av1_vaapi) # Software encoders + filters the cookbook leans on. SW_ENCODERS=(libx264 libx265 libsvtav1 libaom-av1 libvpx-vp9 aac libopus libmp3lame ffv1) KEY_FILTERS=(scale crop pad overlay drawtext subtitles loudnorm silencedetect silenceremove lut3d curves eq zscale tonemap minterpolate vidstabdetect vidstabtransform bwdif hqdn3d nlmeans palettegen paletteuse libvmaf ssim psnr xstack showwaves showspectrum) # Flags-column width varies across ffmpeg majors (3 chars <=7.x, 2 in 8.x). listed_encoder() { grep -qE "^ [A-Z.]{6} +$1 " <<<"$ENCODERS_RAW"; } listed_filter() { grep -qE "^ +[A-Z.|]+ +$1 +" <<<"$FILTERS_RAW"; } proof_encode() { # $1 = encoder name; returns 0 verified, 1 failed local enc="$1" extra=() case "$enc" in *_vaapi) extra=(-vaapi_device /dev/dri/renderD128 -vf format=nv12,hwupload) ;; *_qsv) extra=(-vf format=nv12) ;; esac ffmpeg -v error -y -f lavfi -i testsrc2=duration=1:size=640x360:rate=30 \ "${extra[@]+"${extra[@]}"}" -frames:v 10 -c:v "$enc" -f null - >/dev/null 2>&1 } failed_verify=0 ROWS=() # tsv rows for stdout JSON_ENC=() # jq-built objects RECOMMENDED="" for enc in "${HW_ENCODERS[@]}"; do listed=false verified=null if listed_encoder "$enc"; then listed=true if [[ "$QUICK" -eq 1 ]]; then verified=null emit " hw $enc listed (proof-encode skipped: --quick)" elif proof_encode "$enc"; then verified=true [[ -z "$RECOMMENDED" ]] && RECOMMENDED="$enc" emit " hw $enc VERIFIED" else verified=false; failed_verify=1 emit " hw $enc LISTED BUT FAILED proof-encode (driver/device mismatch?)" fi fi ROWS+=("$(printf 'encoder\t%s\thw\t%s\t%s' "$enc" "$listed" "$verified")") [[ "$HAS_JQ" -eq 1 ]] && JSON_ENC+=("$(jq -cn --arg n "$enc" --argjson l "$listed" \ --argjson v "$verified" '{name:$n, hw:true, listed:$l, verified:$v}')") done for enc in "${SW_ENCODERS[@]}"; do listed=false; listed_encoder "$enc" && listed=true ROWS+=("$(printf 'encoder\t%s\tsw\t%s\tnull' "$enc" "$listed")") [[ "$HAS_JQ" -eq 1 ]] && JSON_ENC+=("$(jq -cn --arg n "$enc" --argjson l "$listed" \ '{name:$n, hw:false, listed:$l, verified:null}')") done JSON_FILT=() missing_filters=() for f in "${KEY_FILTERS[@]}"; do present=false; listed_filter "$f" && present=true [[ "$present" == false ]] && missing_filters+=("$f") ROWS+=("$(printf 'filter\t%s\t-\t%s\tnull' "$f" "$present")") [[ "$HAS_JQ" -eq 1 ]] && JSON_FILT+=("$(jq -cn --arg n "$f" --argjson p "$present" \ '{name:$n, present:$p}')") done [[ ${#missing_filters[@]} -gt 0 ]] && \ emit " note: filters not in this build: ${missing_filters[*]}" if [[ "$JSON" -eq 1 ]]; then printf '%s\n' "${JSON_ENC[@]}" | jq -s \ --arg version "$VERSION" --arg schema "$SCHEMA" \ --arg rec "$RECOMMENDED" --argjson quick "$([[ $QUICK -eq 1 ]] && echo true || echo false)" \ --argjson hwaccels "$(printf '%s\n' $HWACCELS | jq -Rn '[inputs | select(length>0)]')" \ --argjson filters "$(printf '%s\n' "${JSON_FILT[@]}" | jq -s '.')" \ '{data:{version:$version, quick:$quick, hwaccels:$hwaccels, encoders:., filters:$filters, recommended_hw:(if $rec=="" then null else $rec end)}, meta:{schema:$schema}}' else printf '%s\n' "${ROWS[@]}" fi [[ "$failed_verify" -eq 1 ]] && exit "$EXIT_FAILED_VERIFY" exit "$EXIT_OK" -
cut-from-edl.py 11.3 KB
#!/usr/bin/env python3 """EDL JSON -> validated cuts + concat: the deterministic core of edit-as-code. Reads an edit decision list (schema: assets/edl-schema.json — scenes, clips, time ranges, written rationale), validates it, and produces the final video via per-clip cuts + the concat demuxer. DRY-RUN BY DEFAULT: prints every command it would run and touches nothing until --execute. Re-encode mode (default) is frame-accurate and normalizes codec/resolution/fps across clips so the concat is always safe; --copy is faster but requires keyframe-aligned cut points and identical source parameters. Usage: cut-from-edl.py [--execute] [--copy] [-o OUT] [--workdir DIR] [--json] <edl.json> Input: EDL JSON as positional; clip paths resolve relative to the EDL's directory Output: stdout = planned/executed command list (or --json envelope, schema claude-mods.ffmpeg-ops.edl/v1) Stderr: progress, warnings, errors Exit: 0 ok, 2 usage, 3 EDL or source file missing, 4 EDL invalid, 5 ffmpeg missing (--execute only) Examples: cut-from-edl.py edit.json # dry-run: show the plan cut-from-edl.py edit.json --execute -o final.mp4 cut-from-edl.py edit.json --execute --copy # keyframe-aligned EDLs only cut-from-edl.py edit.json --json | jq '.data.commands' """ import argparse import json import shutil import subprocess import sys from pathlib import Path from typing import NoReturn SCHEMA = "claude-mods.ffmpeg-ops.edl/v1" EXIT_OK, EXIT_USAGE, EXIT_NOT_FOUND, EXIT_VALIDATION, EXIT_MISSING_DEP = 0, 2, 3, 4, 5 def err(json_mode: bool, code: str, message: str, exit_code: int) -> NoReturn: if json_mode: print(json.dumps({"error": {"code": code, "message": message, "details": {}}})) print(f"ERROR: {message}", file=sys.stderr) sys.exit(exit_code) def validate_edl(edl: dict) -> list: """Stdlib structural validation mirroring assets/edl-schema.json.""" problems = [] scenes = edl.get("scenes") if not isinstance(scenes, list) or not scenes: return ["'scenes' must be a non-empty array"] for i, scene in enumerate(scenes): where = f"scenes[{i}]" if not isinstance(scene, dict): problems.append(f"{where} must be an object") continue clips = scene.get("clips") if not isinstance(clips, list) or not clips: problems.append(f"{where}.clips must be a non-empty array") continue for j, clip in enumerate(clips): cw = f"{where}.clips[{j}]" if not isinstance(clip, dict): problems.append(f"{cw} must be an object") continue if not isinstance(clip.get("file"), str) or not clip.get("file"): problems.append(f"{cw}.file must be a non-empty string") start, end = clip.get("start"), clip.get("end") if not isinstance(start, (int, float)) or start < 0: problems.append(f"{cw}.start must be a number >= 0") if not isinstance(end, (int, float)): problems.append(f"{cw}.end must be a number") elif isinstance(start, (int, float)) and end <= start: problems.append(f"{cw}: end ({end}) must be > start ({start})") return problems def video_props(ffprobe: str, path: Path) -> dict: proc = subprocess.run( [ffprobe, "-v", "error", "-select_streams", "v:0", "-show_entries", "stream=width,height,r_frame_rate", "-of", "csv=p=0", str(path)], capture_output=True, text=True) parts = proc.stdout.strip().split(",") if len(parts) == 3: try: num, den = parts[2].split("/") fps = round(int(num) / int(den), 3) if int(den) else 0 return {"width": int(parts[0]), "height": int(parts[1]), "fps": fps} except (ValueError, ZeroDivisionError): pass return {} def main() -> int: ap = argparse.ArgumentParser( description="Cut + concat a final video from an EDL JSON (dry-run by default).", epilog="Examples:\n" " cut-from-edl.py edit.json\n" " cut-from-edl.py edit.json --execute -o final.mp4\n", formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("edl", help="EDL JSON file (see assets/edl-schema.json)") ap.add_argument("--execute", action="store_true", help="actually run the cuts (default: dry-run print only)") ap.add_argument("--copy", action="store_true", help="stream-copy cuts (fast; needs keyframe-aligned points " "and identical source params)") ap.add_argument("-o", "--output", default=None, help="final output path, resolved against the CWD (default: the " "EDL 'output' field resolved against the EDL file, else final.mp4)") ap.add_argument("--workdir", default=None, help="directory for cut segments (default: <edl-dir>/edl-cuts)") ap.add_argument("--json", action="store_true", help="emit JSON envelope on stdout") args = ap.parse_args() edl_path = Path(args.edl) if not edl_path.is_file(): err(args.json, "NOT_FOUND", f"EDL not found: {edl_path}", EXIT_NOT_FOUND) try: edl = json.loads(edl_path.read_text(encoding="utf-8")) except json.JSONDecodeError as e: err(args.json, "VALIDATION", f"EDL is not valid JSON: {e}", EXIT_VALIDATION) problems = validate_edl(edl) if problems: err(args.json, "VALIDATION", "EDL failed validation: " + "; ".join(problems[:5]) + (f" (+{len(problems) - 5} more)" if len(problems) > 5 else ""), EXIT_VALIDATION) base = edl_path.resolve().parent workdir = Path(args.workdir) if args.workdir else base / "edl-cuts" # CLI -o resolves against the CWD (normal CLI convention); the EDL's own # 'output' field resolves against the EDL file (schema contract). if args.output: output = Path(args.output).resolve() else: output = Path(edl.get("output") or "final.mp4") if not output.is_absolute(): output = base / output # Resolve and existence-check sources (fatal in execute, warning in dry-run). clips, missing = [], [] for scene in edl["scenes"]: for clip in scene["clips"]: src = Path(clip["file"]) if not src.is_absolute(): src = base / src if not src.is_file(): missing.append(str(src)) clips.append({"scene": scene.get("scene"), "src": src, "start": float(clip["start"]), "end": float(clip["end"])}) if missing: for m in missing: print(f"warning: source missing: {m}", file=sys.stderr) if args.execute: err(args.json, "NOT_FOUND", f"{len(missing)} source file(s) missing (first: {missing[0]})", EXIT_NOT_FOUND) ffmpeg = shutil.which("ffmpeg") ffprobe = shutil.which("ffprobe") if args.execute and not ffmpeg: err(args.json, "MISSING_DEPENDENCY", "ffmpeg not found on PATH", EXIT_MISSING_DEP) # Re-encode mode normalizes every segment to the first clip's geometry/fps, # which is what makes the concat demuxer unconditionally safe. norm_filter = "" if not args.copy and ffprobe and not missing: props = [video_props(ffprobe, c["src"]) for c in clips] props = [p for p in props if p] if props: w, h, fps = props[0]["width"], props[0]["height"], props[0]["fps"] or 30 if any((p["width"], p["height"]) != (w, h) or p["fps"] != props[0]["fps"] for p in props): print(f"note: mixed source params — normalizing all segments to " f"{w}x{h} @ {fps}fps", file=sys.stderr) norm_filter = (f"scale={w}:{h}:force_original_aspect_ratio=decrease," f"pad={w}:{h}:(ow-iw)/2:(oh-ih)/2,fps={fps}") commands, concat_lines = [], [] for n, clip in enumerate(clips, 1): seg = workdir / f"seg{n:03d}.mp4" cmd = ["ffmpeg", "-y", "-ss", f"{clip['start']}", "-to", f"{clip['end']}", "-i", str(clip["src"])] if args.copy: cmd += ["-c", "copy", "-avoid_negative_ts", "make_zero"] else: if norm_filter: cmd += ["-vf", norm_filter] cmd += ["-c:v", "libx264", "-crf", "18", "-preset", "fast", "-pix_fmt", "yuv420p", "-c:a", "aac", "-b:a", "192k", "-ar", "48000"] cmd.append(str(seg)) commands.append(cmd) concat_lines.append(f"file '{seg.as_posix()}'") concat_txt = workdir / "concat.txt" final_cmd = ["ffmpeg", "-y", "-f", "concat", "-safe", "0", "-i", str(concat_txt), "-c", "copy", "-movflags", "+faststart", str(output)] data = { "edl": str(edl_path), "mode": "copy" if args.copy else "reencode", "executed": bool(args.execute), "workdir": str(workdir), "output": str(output), "segments": len(clips), "missing_sources": missing, "commands": [" ".join(c) for c in commands] + [" ".join(final_cmd)], } if not args.execute: if args.json: print(json.dumps({"data": data, "meta": {"schema": SCHEMA}}, indent=2)) else: print(f"# DRY-RUN — {len(clips)} segment(s) -> {output}") for c in data["commands"][:-1]: print(c) print(f"# concat.txt:\n" + "\n".join(f"# {l}" for l in concat_lines)) print(data["commands"][-1]) print("dry-run only; pass --execute to run", file=sys.stderr) return EXIT_OK workdir.mkdir(parents=True, exist_ok=True) for n, cmd in enumerate(commands, 1): print(f"cutting segment {n}/{len(commands)}...", file=sys.stderr) proc = subprocess.run(cmd, capture_output=True, text=True) if proc.returncode != 0: err(args.json, "VALIDATION", f"segment {n} failed: {(proc.stderr.strip().splitlines() or ['?'])[-1]}", EXIT_VALIDATION) concat_txt.write_text("\n".join(concat_lines) + "\n", encoding="utf-8") # Atomic final write: concat to a temp name, then rename over the # destination. The temp KEEPS the real extension — ffmpeg infers the muxer # from it, and "final.mp4.tmp" would fail with "Invalid argument". tmp_out = output.with_name(output.stem + ".tmp" + output.suffix) final_cmd[-1] = str(tmp_out) # the destination dir must exist BEFORE ffmpeg opens the temp output - # otherwise concat dies with a cryptic "Error opening output files" output.parent.mkdir(parents=True, exist_ok=True) print("concatenating...", file=sys.stderr) proc = subprocess.run(final_cmd, capture_output=True, text=True) if proc.returncode != 0: err(args.json, "VALIDATION", f"concat failed: {(proc.stderr.strip().splitlines() or ['?'])[-1]}", EXIT_VALIDATION) tmp_out.replace(output) if args.json: print(json.dumps({"data": data, "meta": {"schema": SCHEMA}}, indent=2)) else: print(str(output)) print(f"done: {output} ({len(clips)} segments)", file=sys.stderr) print("next: re-transcribe the output and verify no words were clipped " "(see references/edit-as-code.md)", file=sys.stderr) return EXIT_OK if __name__ == "__main__": sys.exit(main()) -
detect-segments.py 7.7 KB
#!/usr/bin/env python3 """Silence/scene boundaries as JSON segments — for STT chunking, dead-air cuts, shot splits. ffmpeg's silencedetect and scene-score output is human-oriented log text on stderr; this script runs the right filter and parses it into clean segments. --silence also derives the inverse (speech segments), which is what STT chunking and the cuts-land-in-silence EDL verification actually consume. Usage: detect-segments.py [--silence | --scenes] [options] [--json] <file> Input: one media file as positional Output: stdout = TSV segments (kind, start, end, duration), or --json envelope (schema claude-mods.ffmpeg-ops.segments/v1) Stderr: progress, errors Exit: 0 ok, 2 usage, 3 file not found, 4 stream missing for mode / parse failure, 5 ffmpeg missing Examples: detect-segments.py --silence interview.mp4 detect-segments.py --silence --noise -35dB --min-silence 0.8 --json in.mp4 | jq '.data.speech' detect-segments.py --scenes --scene-threshold 0.3 --json in.mp4 | jq '.data.cuts' """ import argparse import json import re import shutil import subprocess import sys from pathlib import Path from typing import NoReturn SCHEMA = "claude-mods.ffmpeg-ops.segments/v1" EXIT_OK, EXIT_USAGE, EXIT_NOT_FOUND, EXIT_VALIDATION, EXIT_MISSING_DEP = 0, 2, 3, 4, 5 def err(json_mode: bool, code: str, message: str, exit_code: int) -> NoReturn: if json_mode: print(json.dumps({"error": {"code": code, "message": message, "details": {}}})) print(f"ERROR: {message}", file=sys.stderr) sys.exit(exit_code) def media_duration(ffprobe: str, path: Path) -> float: proc = subprocess.run( [ffprobe, "-v", "error", "-show_entries", "format=duration", "-of", "default=nw=1:nk=1", str(path)], capture_output=True, text=True) try: return float(proc.stdout.strip()) except ValueError: return 0.0 def detect_silence(ffmpeg: str, path: Path, noise: str, min_silence: float, duration: float) -> dict: proc = subprocess.run( [ffmpeg, "-hide_banner", "-nostats", "-i", str(path), "-af", f"silencedetect=noise={noise}:d={min_silence}", "-vn", "-f", "null", "-"], capture_output=True, text=True) if proc.returncode != 0: return {"_error": (proc.stderr.strip().splitlines() or ["unknown"])[-1]} starts = [float(m) for m in re.findall(r"silence_start:\s*(-?[\d.]+)", proc.stderr)] ends = [float(m) for m in re.findall(r"silence_end:\s*(-?[\d.]+)", proc.stderr)] # A silence running to EOF has a start but no end line. if len(starts) == len(ends) + 1: ends.append(duration) silences = [{"start": round(max(0.0, s), 3), "end": round(e, 3), "duration": round(e - s, 3)} for s, e in zip(starts, ends)] speech, cursor = [], 0.0 for sil in silences: if sil["start"] > cursor + 0.01: speech.append({"start": round(cursor, 3), "end": sil["start"], "duration": round(sil["start"] - cursor, 3)}) cursor = sil["end"] if duration > cursor + 0.01: speech.append({"start": round(cursor, 3), "end": round(duration, 3), "duration": round(duration - cursor, 3)}) return {"silences": silences, "speech": speech} def detect_scenes(ffmpeg: str, path: Path, threshold: float, duration: float) -> dict: # metadata=print:file=- routes the per-frame report to STDOUT — a clean parse, # unlike silencedetect which only logs to stderr. proc = subprocess.run( [ffmpeg, "-hide_banner", "-nostats", "-i", str(path), "-vf", f"select='gt(scene,{threshold})',metadata=print:file=-", "-an", "-f", "null", "-"], capture_output=True, text=True) if proc.returncode != 0: return {"_error": (proc.stderr.strip().splitlines() or ["unknown"])[-1]} cuts, scores = [], [] pts_re = re.compile(r"pts_time:(-?[\d.]+)") score_re = re.compile(r"lavfi\.scene_score=([\d.]+)") pending_pts = None for line in proc.stdout.splitlines(): m = pts_re.search(line) if m: pending_pts = float(m.group(1)) continue m = score_re.search(line) if m and pending_pts is not None: cuts.append(round(pending_pts, 3)) scores.append(float(m.group(1))) pending_pts = None segments, cursor = [], 0.0 for c in cuts: if c > cursor + 0.01: segments.append({"start": round(cursor, 3), "end": c, "duration": round(c - cursor, 3)}) cursor = c if duration > cursor + 0.01: segments.append({"start": round(cursor, 3), "end": round(duration, 3), "duration": round(duration - cursor, 3)}) return {"cuts": cuts, "scores": scores, "segments": segments} def main() -> int: ap = argparse.ArgumentParser( description="Detect silence or scene-change boundaries as JSON segments.", epilog="Examples:\n" " detect-segments.py --silence interview.mp4\n" " detect-segments.py --scenes --json in.mp4 | jq '.data.cuts'\n", formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("file", help="media file to analyze") mode = ap.add_mutually_exclusive_group() mode.add_argument("--silence", action="store_true", help="detect audio silence + derive speech segments (default)") mode.add_argument("--scenes", action="store_true", help="detect video scene changes") ap.add_argument("--noise", default="-30dB", help="silence threshold, e.g. -30dB (default) or -35dB") ap.add_argument("--min-silence", type=float, default=0.5, help="minimum silence duration in seconds (default 0.5)") ap.add_argument("--scene-threshold", type=float, default=0.4, help="scene-change score threshold 0..1 (default 0.4)") ap.add_argument("--json", action="store_true", help="emit JSON envelope on stdout") args = ap.parse_args() ffmpeg, ffprobe = shutil.which("ffmpeg"), shutil.which("ffprobe") if not ffmpeg or not ffprobe: err(args.json, "MISSING_DEPENDENCY", "ffmpeg/ffprobe not found on PATH", EXIT_MISSING_DEP) path = Path(args.file) if not path.is_file(): err(args.json, "NOT_FOUND", f"file not found: {path}", EXIT_NOT_FOUND) duration = media_duration(ffprobe, path) mode_name = "scenes" if args.scenes else "silence" print(f"detecting {mode_name} in {path.name}...", file=sys.stderr) if args.scenes: result = detect_scenes(ffmpeg, path, args.scene_threshold, duration) params = {"scene_threshold": args.scene_threshold} else: result = detect_silence(ffmpeg, path, args.noise, args.min_silence, duration) params = {"noise": args.noise, "min_silence_s": args.min_silence} if "_error" in result: err(args.json, "VALIDATION", f"{mode_name} analysis failed (missing stream for mode?): {result['_error']}", EXIT_VALIDATION) data = {"file": str(path), "mode": mode_name, "duration_s": round(duration, 3), "params": params, **result} if args.json: print(json.dumps({"data": data, "meta": {"schema": SCHEMA}}, indent=2)) return EXIT_OK if args.scenes: for seg in data["segments"]: print(f"scene\t{seg['start']}\t{seg['end']}\t{seg['duration']}") else: for seg in data["silences"]: print(f"silence\t{seg['start']}\t{seg['end']}\t{seg['duration']}") for seg in data["speech"]: print(f"speech\t{seg['start']}\t{seg['end']}\t{seg['duration']}") return EXIT_OK if __name__ == "__main__": sys.exit(main()) -
gen-luts.py 15.5 KB
#!/usr/bin/env python3 """Generate .cube 3D LUT grade variants (+ optional preview stills with HTML chooser). A .cube LUT is plain ASCII (an N^3 lattice of RGB triples), so grade candidates can be computed rather than hand-tuned in an NLE. This emits a family of looks — optionally on top of an S-Log3 -> Rec.709 conversion for log footage — and, with --previews, renders one still per look plus an index.html so a HUMAN can choose. THE AGENT NEVER PICKS THE GRADE. Generate, render previews, present the chooser, wait. Grading is a taste call (see SKILL.md / references/color-grading.md). Usage: gen-luts.py [--variants LIST|all] [--size N] [--input-space slog3|rec709] [--out-dir DIR] [--previews MEDIA [--frame-at S]] [--json] Input: no positional; --previews takes a video/image to grade stills from Output: stdout = one line per written file (or --json manifest envelope, schema claude-mods.ffmpeg-ops.luts/v1) Stderr: progress, the human-picks-the-grade reminder, errors Exit: 0 ok, 2 usage, 3 preview source missing, 5 ffmpeg missing (--previews only) Examples: gen-luts.py --variants all --out-dir work/luts gen-luts.py --variants warm_filmic,punchy,teal_orange --input-space slog3 gen-luts.py --variants all --out-dir work/luts --previews footage.mp4 --frame-at 12.5 gen-luts.py --variants all --json | jq -r '.data.files[]' """ import argparse import json import shutil import subprocess import sys from pathlib import Path from typing import NoReturn SCHEMA = "claude-mods.ffmpeg-ops.luts/v1" EXIT_OK, EXIT_USAGE, EXIT_NOT_FOUND, EXIT_MISSING_DEP = 0, 2, 3, 5 # Each look: white-balance temp (+warm/-cool), lift/gamma/gain (master), # per-channel gain tweaks, contrast (pivot 0.5), saturation, fade (black lift). # Optional "mix": a 3x3 channel-mix matrix applied first (rows = output R,G,B # as weights of input r,g,b) — what makes sepia/Technicolor expressible. LOOKS = { "neutral709": dict(temp=0.00, lift=0.000, gamma=1.00, gain=1.00, rgb_gain=(1.00, 1.00, 1.00), contrast=1.00, sat=1.00, fade=0.00), "warm_filmic": dict(temp=0.06, lift=0.005, gamma=0.98, gain=1.00, rgb_gain=(1.02, 1.00, 0.97), contrast=1.08, sat=1.05, fade=0.03), "punchy": dict(temp=0.01, lift=-0.010, gamma=1.00, gain=1.02, rgb_gain=(1.00, 1.00, 1.00), contrast=1.22, sat=1.25, fade=0.00), "teal_orange": dict(temp=0.02, lift=0.000, gamma=1.00, gain=1.00, rgb_gain=(1.05, 1.00, 0.94), contrast=1.10, sat=1.10, fade=0.01, shadow_teal=0.04), "cool_desat": dict(temp=-0.05, lift=0.005, gamma=1.00, gain=0.99, rgb_gain=(0.97, 1.00, 1.03), contrast=1.04, sat=0.80, fade=0.02), "bleach_bypass": dict(temp=0.00, lift=-0.005, gamma=1.00, gain=0.98, rgb_gain=(1.00, 1.00, 1.00), contrast=1.30, sat=0.45, fade=0.00), "film_fade": dict(temp=0.02, lift=0.010, gamma=1.02, gain=0.99, rgb_gain=(1.01, 1.00, 0.99), contrast=0.96, sat=0.90, fade=0.06), "golden_hour": dict(temp=0.10, lift=0.005, gamma=1.01, gain=1.00, rgb_gain=(1.04, 1.01, 0.95), contrast=1.05, sat=1.08, fade=0.02), "pastel": dict(temp=0.01, lift=0.015, gamma=1.05, gain=0.99, rgb_gain=(1.00, 1.00, 1.00), contrast=0.88, sat=0.72, fade=0.08), "noir_bw": dict(temp=0.00, lift=-0.005, gamma=1.00, gain=1.00, rgb_gain=(1.00, 1.00, 1.00), contrast=1.25, sat=0.00, fade=0.00), "sepia": dict(temp=0.00, lift=0.005, gamma=1.00, gain=1.00, rgb_gain=(1.00, 1.00, 1.00), contrast=1.02, sat=1.00, fade=0.02, mix=((.393, .769, .189), (.349, .686, .168), (.272, .534, .131))), "technicolor2": dict(temp=0.00, lift=0.000, gamma=1.00, gain=1.00, rgb_gain=(1.00, 1.00, 1.00), contrast=1.10, sat=1.20, fade=0.00, mix=((1.0, 0.0, 0.0), (0.0, 0.6, 0.4), (0.0, 0.4, 0.6))), "matrix_green": dict(temp=0.00, lift=0.005, gamma=1.00, gain=1.00, rgb_gain=(0.97, 1.06, 0.98), contrast=1.10, sat=0.85, fade=0.02), # Scope-extracted from reference footage (see look-recipes.md grimdark): # warm-ash desat, pulled mids, true-ish blacks, controlled ceiling. "grimdark": dict(temp=0.015, lift=0.000, gamma=0.93, gain=0.97, rgb_gain=(1.02, 1.01, 0.98), contrast=1.04, sat=0.33, fade=0.03), } # Tone-map variants: gradient-map luma onto 2 stops (duotone) or 3 stops # (tritone/monotone: shadow, mid, highlight), all 0..1 RGB. Chroma of the look # = how far the stops sit from the neutral grey axis - monotones barely leave # it, poster duotones live far out. "contrast" applies pre-map (widens spread). _TONE_BASE = dict(temp=0.0, lift=0.0, gamma=1.0, gain=1.0, rgb_gain=(1.0, 1.0, 1.0), contrast=1.05, sat=1.0, fade=0.0) LOOKS.update({ # poster-strength duotones "duo_navy": {**_TONE_BASE, "tones": ((.05, .08, .25), (.98, .93, .80))}, "duo_cyanotype": {**_TONE_BASE, "tones": ((.04, .16, .29), (.92, .96, 1.0))}, "duo_sunset": {**_TONE_BASE, "tones": ((.23, .06, .36), (1.0, .78, .34))}, "duo_forest": {**_TONE_BASE, "tones": ((.06, .24, .18), (.91, .85, .63))}, "duo_crimson": {**_TONE_BASE, "tones": ((.10, .02, .03), (1.0, .88, .86))}, "duo_synthwave": {**_TONE_BASE, "tones": ((.35, .06, .42), (.42, .91, 1.0))}, # muted / tertiary duotones "duo_ash_rose": {**_TONE_BASE, "tones": ((.23, .20, .22), (.85, .78, .76))}, "duo_olive_bone": {**_TONE_BASE, "tones": ((.18, .20, .14), (.90, .88, .81))}, "duo_petrol_paper": {**_TONE_BASE, "tones": ((.12, .23, .24), (.93, .91, .86))}, "duo_indigo_parchment": {**_TONE_BASE, "tones": ((.16, .23, .33), (.91, .89, .82))}, "duo_slate_ice": {**_TONE_BASE, "tones": ((.11, .15, .20), (.95, .97, .98))}, # monotones (darkroom chemical tones - chroma barely off the grey axis) "mono_selenium": {**_TONE_BASE, "tones": ((.05, .04, .07), (.48, .46, .52), (.96, .95, .97))}, "mono_platinum": {**_TONE_BASE, "tones": ((.07, .07, .06), (.52, .51, .49), (.97, .96, .94))}, "mono_coffee": {**_TONE_BASE, "tones": ((.08, .05, .03), (.55, .47, .40), (.96, .92, .87))}, "mono_steel": {**_TONE_BASE, "tones": ((.04, .06, .09), (.46, .50, .55), (.94, .96, .98))}, # tritones (distinct shadow / mid / highlight hues) "tri_split_classic": {**_TONE_BASE, "tones": ((.06, .07, .12), (.50, .49, .48), (.98, .94, .86))}, "tri_tobacco": {**_TONE_BASE, "tones": ((.05, .04, .02), (.45, .40, .28), (.95, .88, .70))}, "tri_arctic": {**_TONE_BASE, "tones": ((.03, .05, .09), (.42, .50, .58), (.93, .97, 1.0))}, }) def err(json_mode: bool, code: str, message: str, exit_code: int) -> NoReturn: if json_mode: print(json.dumps({"error": {"code": code, "message": message, "details": {}}})) print(f"ERROR: {message}", file=sys.stderr) sys.exit(exit_code) def clamp(x: float) -> float: return 0.0 if x < 0.0 else 1.0 if x > 1.0 else x def slog3_to_linear(x: float) -> float: """Sony S-Log3 EOTF (input 0..1 code value -> scene linear).""" if x >= 171.2102946929 / 1023.0: return (10.0 ** ((x * 1023.0 - 420.0) / 261.5)) * 0.19 - 0.01 return (x * 1023.0 - 95.0) * 0.01125 / (171.2102946929 - 95.0) def linear_to_rec709(x: float) -> float: """BT.709 OETF with a Reinhard-style shoulder for >1.0 scene values.""" x = max(0.0, x) x = x / (1.0 + 0.35 * x) # soft highlight roll-off if x < 0.018: return 4.5 * x return 1.099 * (x ** 0.45) - 0.099 def apply_look(r: float, g: float, b: float, p: dict) -> tuple: # Channel mix first (sepia/Technicolor-class looks), then white balance. mix = p.get("mix") if mix: r, g, b = (mix[0][0] * r + mix[0][1] * g + mix[0][2] * b, mix[1][0] * r + mix[1][1] * g + mix[1][2] * b, mix[2][0] * r + mix[2][1] * g + mix[2][2] * b) t = p["temp"] r, b = r * (1.0 + t), b * (1.0 - t) # Lift / gamma / gain (master), then per-channel gain. out = [] for c, cg in zip((r, g, b), p["rgb_gain"]): c = c * p["gain"] * cg + p["lift"] * (1.0 - c) c = clamp(c) ** (1.0 / p["gamma"]) out.append(c) r, g, b = out # Teal/orange split-tone: push shadows toward teal (complement of the warm gain). st = p.get("shadow_teal", 0.0) if st: luma = 0.2126 * r + 0.7152 * g + 0.0722 * b w = (1.0 - luma) ** 2 # weight shadows only r, b = r - st * w, b + st * w # Contrast around mid pivot. k = p["contrast"] r, g, b = (0.5 + (c - 0.5) * k for c in (r, g, b)) # Tone gradient map (replaces saturation): 2 stops = duotone lerp, # 3 stops = piecewise shadow->mid (luma 0..0.5) -> highlight (0.5..1). tones = p.get("tones") luma = 0.2126 * r + 0.7152 * g + 0.0722 * b if tones: luma = clamp(luma) if len(tones) == 3: lo, hi = (tones[0], tones[1]) if luma < 0.5 else (tones[1], tones[2]) f2 = luma * 2 if luma < 0.5 else (luma - 0.5) * 2 else: lo, hi, f2 = tones[0], tones[1], luma r, g, b = (lo[i] + f2 * (hi[i] - lo[i]) for i in range(3)) else: s = p["sat"] r, g, b = (luma + s * (c - luma) for c in (r, g, b)) # Fade (lifted blacks). f = p["fade"] r, g, b = (f + c * (1.0 - f) for c in (r, g, b)) return clamp(r), clamp(g), clamp(b) def write_cube(path: Path, name: str, size: int, input_space: str, params: dict) -> None: lines = [f'# generated by claude-mods ffmpeg-ops gen-luts.py', f'# look={name} input_space={input_space}', f'TITLE "{name}"', f'LUT_3D_SIZE {size}', 'DOMAIN_MIN 0.0 0.0 0.0', 'DOMAIN_MAX 1.0 1.0 1.0'] n = size - 1 for bi in range(size): # .cube order: red varies fastest for gi in range(size): for ri in range(size): r, g, b = ri / n, gi / n, bi / n if input_space == "slog3": r, g, b = (linear_to_rec709(slog3_to_linear(c)) for c in (r, g, b)) r, g, b = apply_look(r, g, b, params) lines.append(f"{r:.6f} {g:.6f} {b:.6f}") tmp = path.with_suffix(".cube.tmp") tmp.write_text("\n".join(lines) + "\n", encoding="ascii") tmp.replace(path) def render_previews(ffmpeg: str, media: Path, luts: list, out_dir: Path, frame_at: float) -> list: stills = [] base_png = out_dir / "preview_original.png" runs = [(None, base_png)] + [(p, out_dir / f"preview_{p.stem}.png") for p in luts] media_abs = str(media.resolve()) for lut, png in runs: cmd = [ffmpeg, "-y", "-v", "error", "-ss", str(frame_at), "-i", media_abs] if lut: # Run from out_dir and reference the LUT by bare filename — a full # path inside the filter arg hits the drive-colon escaping trap # ("lut3d=file=C:/..." parses ':' as an option separator). cmd += ["-vf", f"lut3d=file={lut.name}:interp=tetrahedral"] cmd += ["-frames:v", "1", png.name] proc = subprocess.run(cmd, capture_output=True, text=True, cwd=str(out_dir)) if proc.returncode == 0: stills.append(png) else: print(f"warning: preview failed for {lut.name if lut else 'original'}: " f"{(proc.stderr.strip().splitlines() or ['?'])[-1]}", file=sys.stderr) cells = "\n".join( f'<figure><img src="{p.name}" loading="lazy">' f"<figcaption>{p.stem.replace('preview_', '')}</figcaption></figure>" for p in stills) (out_dir / "index.html").write_text( "<!doctype html><meta charset='utf-8'><title>Pick a grade</title>" "<style>body{background:#111;color:#eee;font:14px system-ui;margin:24px}" "main{display:grid;grid-template-columns:repeat(auto-fill,minmax(420px,1fr));gap:16px}" "img{width:100%;border-radius:6px}figcaption{margin-top:4px;text-align:center}" "</style><h1>Pick a grade</h1><main>" + cells + "</main>\n", encoding="utf-8") return stills def main() -> int: ap = argparse.ArgumentParser( description="Generate .cube grade variants; optionally render a preview chooser.", epilog="Examples:\n" " gen-luts.py --variants all --out-dir work/luts\n" " gen-luts.py --variants all --previews footage.mp4 --frame-at 12.5\n", formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("--variants", default="all", help=f"comma list or 'all' of: {', '.join(LOOKS)} (default all)") ap.add_argument("--size", type=int, default=33, choices=(17, 33, 65), help="lattice points per axis (default 33)") ap.add_argument("--input-space", default="rec709", choices=("rec709", "slog3"), help="source space; slog3 bakes an S-Log3->Rec.709 conversion in") ap.add_argument("--out-dir", default="luts", help="output directory (default ./luts)") ap.add_argument("--previews", default=None, metavar="MEDIA", help="render a graded still per LUT from this video/image + index.html") ap.add_argument("--frame-at", type=float, default=5.0, help="timestamp for the preview frame (default 5.0s)") ap.add_argument("--json", action="store_true", help="emit JSON manifest on stdout") args = ap.parse_args() if args.variants.strip().lower() == "all": names = list(LOOKS) else: names = [v.strip() for v in args.variants.split(",") if v.strip()] unknown = [n for n in names if n not in LOOKS] if unknown or not names: err(args.json, "USAGE", f"unknown look(s): {', '.join(unknown) or '(none given)'} " f"(available: {', '.join(LOOKS)})", EXIT_USAGE) ffmpeg = None media = None if args.previews: ffmpeg = shutil.which("ffmpeg") if not ffmpeg: err(args.json, "MISSING_DEPENDENCY", "ffmpeg not found on PATH (required for --previews)", EXIT_MISSING_DEP) media = Path(args.previews) if not media.is_file(): err(args.json, "NOT_FOUND", f"preview source not found: {media}", EXIT_NOT_FOUND) out_dir = Path(args.out_dir) out_dir.mkdir(parents=True, exist_ok=True) written = [] for name in names: path = out_dir / f"{name}.cube" print(f"writing {path.name} ({args.size}^3, {args.input_space})...", file=sys.stderr) write_cube(path, name, args.size, args.input_space, LOOKS[name]) written.append(path) stills = [] if args.previews and ffmpeg and media: print("rendering preview stills...", file=sys.stderr) stills = render_previews(ffmpeg, media, written, out_dir, args.frame_at) data = {"out_dir": str(out_dir), "size": args.size, "input_space": args.input_space, "files": [str(p) for p in written], "previews": [str(p) for p in stills], "chooser": str(out_dir / "index.html") if stills else None} if args.json: print(json.dumps({"data": data, "meta": {"schema": SCHEMA}}, indent=2)) else: for p in written + stills: print(p) if stills: print(out_dir / "index.html") print("REMINDER: present the chooser to the human — never auto-pick a grade.", file=sys.stderr) return EXIT_OK if __name__ == "__main__": sys.exit(main()) -
loudnorm-scan.py 5.2 KB
#!/usr/bin/env python3 """Two-pass EBU R128 loudness: run the measurement pass, emit the exact pass-2 filter. One-pass loudnorm runs in dynamic mode (pumps quiet passages). Proper linear normalization needs the measured values fed back in — this script runs pass 1, parses loudnorm's JSON report off stderr, and prints the ready-to-paste pass-2 filter string (and full command), so the agent never re-derives the dance. Usage: loudnorm-scan.py [-I LUFS] [--tp dBTP] [--lra LU] [--json] <file> Input: one media file with an audio stream Output: stdout = measured values + pass-2 filter (or --json envelope, schema claude-mods.ffmpeg-ops.loudnorm/v1) Stderr: progress, errors Exit: 0 ok, 2 usage, 3 file not found, 4 no audio / parse failure, 5 ffmpeg missing Targets: -14 streaming platforms, -16 podcasts (default), -23 EBU R128 broadcast. Examples: loudnorm-scan.py podcast.wav loudnorm-scan.py -I -14 --json music.mp4 | jq -r '.data.pass2_filter' loudnorm-scan.py -I -23 --tp -2 --lra 7 broadcast.mov """ import argparse import json import shutil import subprocess import sys from pathlib import Path from typing import NoReturn SCHEMA = "claude-mods.ffmpeg-ops.loudnorm/v1" EXIT_OK, EXIT_USAGE, EXIT_NOT_FOUND, EXIT_VALIDATION, EXIT_MISSING_DEP = 0, 2, 3, 4, 5 def err(json_mode: bool, code: str, message: str, exit_code: int) -> NoReturn: if json_mode: print(json.dumps({"error": {"code": code, "message": message, "details": {}}})) print(f"ERROR: {message}", file=sys.stderr) sys.exit(exit_code) def main() -> int: ap = argparse.ArgumentParser( description="Measure loudness (pass 1) and emit the exact pass-2 loudnorm filter.", epilog="Examples:\n" " loudnorm-scan.py podcast.wav\n" " loudnorm-scan.py -I -14 --json music.mp4 | jq -r '.data.pass2_filter'\n", formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("file", help="media file with an audio stream") ap.add_argument("-I", "--target-i", type=float, default=-16.0, help="integrated loudness target, LUFS (default -16)") ap.add_argument("--tp", type=float, default=-1.5, help="true-peak ceiling, dBTP (default -1.5)") ap.add_argument("--lra", type=float, default=11.0, help="loudness range target, LU (default 11)") ap.add_argument("--json", action="store_true", help="emit JSON envelope on stdout") args = ap.parse_args() ffmpeg = shutil.which("ffmpeg") if not ffmpeg: err(args.json, "MISSING_DEPENDENCY", "ffmpeg not found on PATH", EXIT_MISSING_DEP) path = Path(args.file) if not path.is_file(): err(args.json, "NOT_FOUND", f"file not found: {path}", EXIT_NOT_FOUND) base = f"I={args.target_i:g}:TP={args.tp:g}:LRA={args.lra:g}" print(f"measuring loudness of {path.name} (pass 1)...", file=sys.stderr) proc = subprocess.run( [ffmpeg, "-hide_banner", "-nostats", "-i", str(path), "-af", f"loudnorm={base}:print_format=json", "-f", "null", "-"], capture_output=True, text=True) # loudnorm prints its JSON report as the last {...} block on stderr. stderr = proc.stderr or "" start, end = stderr.rfind("{"), stderr.rfind("}") if proc.returncode != 0 or start == -1 or end <= start: detail = stderr.strip().splitlines()[-1] if stderr.strip() else "no detail" err(args.json, "VALIDATION", f"loudnorm measurement failed (no audio stream?): {detail}", EXIT_VALIDATION) try: m = json.loads(stderr[start:end + 1]) except json.JSONDecodeError: err(args.json, "VALIDATION", "could not parse loudnorm JSON report", EXIT_VALIDATION) pass2_filter = ( f"loudnorm={base}" f":measured_I={m['input_i']}:measured_TP={m['input_tp']}" f":measured_LRA={m['input_lra']}:measured_thresh={m['input_thresh']}" f":offset={m['target_offset']}:linear=true" ) # loudnorm internally resamples to 192 kHz — the -ar 48000 puts it back. pass2_command = (f'ffmpeg -y -i "{path}" -af "{pass2_filter}" -ar 48000 ' f'-c:v copy "{path.stem}.normalized{path.suffix}"') data = { "file": str(path), "target": {"I": args.target_i, "TP": args.tp, "LRA": args.lra}, "measured": { "input_i": float(m["input_i"]), "input_tp": float(m["input_tp"]), "input_lra": float(m["input_lra"]), "input_thresh": float(m["input_thresh"]), "target_offset": float(m["target_offset"]), }, "normalization_mode": m.get("normalization_type", ""), "pass2_filter": pass2_filter, "pass2_command": pass2_command, } if args.json: print(json.dumps({"data": data, "meta": {"schema": SCHEMA}}, indent=2)) else: print(f"measured I={m['input_i']} LUFS TP={m['input_tp']} dBTP " f"LRA={m['input_lra']} LU thresh={m['input_thresh']}") print(f"target I={args.target_i:g} TP={args.tp:g} LRA={args.lra:g}") print(f"pass2 {pass2_filter}") print(f"command {pass2_command}") return EXIT_OK if __name__ == "__main__": sys.exit(main()) -
make-chapters.py 11.4 KB
#!/usr/bin/env python3 """Chapter authoring: scene/silence boundaries or explicit JSON -> embedded chapters. Derives chapter points (scene detection, speech-after-silence starts, or an explicit chapters JSON), merges points closer than --min-gap, and emits any of: ffmetadata (the format ffmpeg muxes), YouTube description text, WebVTT chapters, or JSON. --write muxes the chapters INTO a stream-copy of the media (atomic, original untouched). Usage: make-chapters.py (--from-scenes | --from-silence | --chapters FILE) [--media FILE] [--min-gap S] [--duration S] [--format ffmetadata|youtube|vtt|json] [--write OUT] [--json] Input: --media for detection modes and --write; --chapters JSON is [{"start": 0, "title": "Intro"}, ...] (or {"chapters": [...]}) Output: stdout = the chosen format (default ffmetadata); --json = envelope (schema claude-mods.ffmpeg-ops.chapters/v1) Stderr: progress, YouTube-rule warnings, errors Exit: 0 ok, 2 usage, 3 media/chapters file missing, 4 invalid chapters JSON, 5 ffmpeg/ffprobe missing when required Examples: make-chapters.py --from-scenes --media talk.mp4 --min-gap 30 make-chapters.py --from-silence --media lecture.mp4 --write chaptered.mp4 make-chapters.py --chapters chapters.json --duration 3600 --format youtube make-chapters.py --from-scenes --media in.mp4 --format json | jq '.data.chapters' """ import argparse import json import shutil import subprocess import sys from pathlib import Path from typing import NoReturn SCHEMA = "claude-mods.ffmpeg-ops.chapters/v1" EXIT_OK, EXIT_USAGE, EXIT_NOT_FOUND, EXIT_VALIDATION, EXIT_MISSING_DEP = 0, 2, 3, 4, 5 def err(json_mode: bool, code: str, message: str, exit_code: int) -> NoReturn: if json_mode: print(json.dumps({"error": {"code": code, "message": message, "details": {}}})) print(f"ERROR: {message}", file=sys.stderr) sys.exit(exit_code) def media_duration(path: Path, json_mode: bool) -> float: ffprobe = shutil.which("ffprobe") if not ffprobe: err(json_mode, "MISSING_DEPENDENCY", "ffprobe not found on PATH", EXIT_MISSING_DEP) proc = subprocess.run( [ffprobe, "-v", "error", "-show_entries", "format=duration", "-of", "default=nw=1:nk=1", str(path)], capture_output=True, text=True) try: return float(proc.stdout.strip()) except ValueError: err(json_mode, "VALIDATION", f"could not read duration of {path.name}", EXIT_VALIDATION) def detect_points(mode: str, media: Path, json_mode: bool) -> list: """Shell out to the sibling detect-segments.py — one detection implementation.""" sibling = Path(__file__).resolve().parent / "detect-segments.py" flag = "--scenes" if mode == "scenes" else "--silence" proc = subprocess.run( [sys.executable, str(sibling), flag, "--json", str(media)], capture_output=True, text=True) if proc.returncode != 0: err(json_mode, "VALIDATION", f"detect-segments {flag} failed (exit {proc.returncode}): " f"{(proc.stderr.strip().splitlines() or ['?'])[-1]}", proc.returncode) data = json.loads(proc.stdout)["data"] if mode == "scenes": return [float(c) for c in data.get("cuts", [])] # silence mode: a chapter candidate is where speech RESUMES return [float(seg["start"]) for seg in data.get("speech", [])] def load_chapters_file(path: Path, json_mode: bool) -> list: if not path.is_file(): err(json_mode, "NOT_FOUND", f"chapters file not found: {path}", EXIT_NOT_FOUND) try: raw = json.loads(path.read_text(encoding="utf-8")) except json.JSONDecodeError as e: err(json_mode, "VALIDATION", f"chapters file is not valid JSON: {e}", EXIT_VALIDATION) items = raw.get("chapters") if isinstance(raw, dict) else raw if not isinstance(items, list) or not items: err(json_mode, "VALIDATION", 'chapters JSON must be a non-empty array of {"start": s, "title": "..."}', EXIT_VALIDATION) chapters = [] for i, c in enumerate(items): if not isinstance(c, dict) or not isinstance(c.get("start"), (int, float)): err(json_mode, "VALIDATION", f"chapters[{i}] needs a numeric 'start'", EXIT_VALIDATION) chapters.append({"start": float(c["start"]), "title": str(c.get("title") or f"Chapter {i + 1}")}) return sorted(chapters, key=lambda c: c["start"]) def build_chapters(points: list, min_gap: float, duration: float) -> list: """Merge close points, force a chapter at 0, attach END times.""" merged = [0.0] for p in sorted(p for p in points if p > 0): if p - merged[-1] >= min_gap and (duration <= 0 or duration - p >= min_gap): merged.append(round(p, 3)) return [{"start": s, "title": f"Chapter {i + 1}"} for i, s in enumerate(merged)] def attach_ends(chapters: list, duration: float) -> list: out = [] for i, c in enumerate(chapters): end = chapters[i + 1]["start"] if i + 1 < len(chapters) else duration out.append({**c, "end": round(max(end, c["start"]), 3)}) return out def esc_ffmeta(s: str) -> str: for ch in ("\\", "=", ";", "#"): s = s.replace(ch, "\\" + ch) return s.replace("\n", " ") def fmt_ffmetadata(chapters: list) -> str: lines = [";FFMETADATA1"] for c in chapters: lines += ["[CHAPTER]", "TIMEBASE=1/1000", f"START={int(c['start'] * 1000)}", f"END={int(c['end'] * 1000)}", f"title={esc_ffmeta(c['title'])}"] return "\n".join(lines) + "\n" def ts_youtube(s: float) -> str: h, rem = divmod(int(s), 3600) m, sec = divmod(rem, 60) return f"{h}:{m:02d}:{sec:02d}" if h else f"{m}:{sec:02d}" def ts_vtt(s: float) -> str: h, rem = divmod(int(s), 3600) m, sec = divmod(rem, 60) return f"{h:02d}:{m:02d}:{sec:02d}.{int(round((s % 1) * 1000)):03d}" def fmt_youtube(chapters: list) -> str: # YouTube parses chapters only if: first at 0:00, >= 3 chapters, each >= 10 s. if chapters and chapters[0]["start"] != 0: print("warning: YouTube requires the first chapter at 0:00", file=sys.stderr) if len(chapters) < 3: print("warning: YouTube needs >= 3 chapters to render them", file=sys.stderr) if any(c["end"] - c["start"] < 10 for c in chapters): print("warning: YouTube ignores chapter lists with any chapter < 10 s", file=sys.stderr) return "\n".join(f"{ts_youtube(c['start'])} {c['title']}" for c in chapters) + "\n" def fmt_vtt(chapters: list) -> str: blocks = [f"{ts_vtt(c['start'])} --> {ts_vtt(c['end'])}\n{c['title']}" for c in chapters] return "WEBVTT\n\n" + "\n\n".join(blocks) + "\n" def mux_chapters(media: Path, meta: str, out: Path, json_mode: bool) -> None: ffmpeg = shutil.which("ffmpeg") if not ffmpeg: err(json_mode, "MISSING_DEPENDENCY", "ffmpeg not found on PATH (--write)", EXIT_MISSING_DEP) meta_file = out.parent / (out.stem + ".ffmeta.tmp") tmp_out = out.with_name(out.stem + ".tmp" + out.suffix) out.parent.mkdir(parents=True, exist_ok=True) meta_file.write_text(meta, encoding="utf-8") try: proc = subprocess.run( [ffmpeg, "-y", "-v", "error", "-i", str(media), "-f", "ffmetadata", "-i", str(meta_file), "-map", "0", "-map_metadata", "0", "-map_chapters", "1", "-c", "copy", str(tmp_out)], capture_output=True, text=True) if proc.returncode != 0: err(json_mode, "VALIDATION", f"chapter mux failed: {(proc.stderr.strip().splitlines() or ['?'])[-1]}", EXIT_VALIDATION) tmp_out.replace(out) finally: meta_file.unlink(missing_ok=True) tmp_out.unlink(missing_ok=True) def main() -> int: ap = argparse.ArgumentParser( description="Derive chapters and emit ffmetadata/YouTube/VTT or mux them in.", epilog="Examples:\n" " make-chapters.py --from-scenes --media talk.mp4 --min-gap 30\n" " make-chapters.py --chapters ch.json --duration 3600 --format youtube\n", formatter_class=argparse.RawDescriptionHelpFormatter) src = ap.add_mutually_exclusive_group(required=True) src.add_argument("--from-scenes", action="store_true", help="chapter points from video scene changes") src.add_argument("--from-silence", action="store_true", help="chapter points where speech resumes after silence") src.add_argument("--chapters", metavar="FILE", help='explicit JSON: [{"start": s, "title": "..."}]') ap.add_argument("--media", metavar="FILE", help="media file (required for detection modes and --write)") ap.add_argument("--min-gap", type=float, default=15.0, help="merge detected points closer than this, seconds (default 15)") ap.add_argument("--duration", type=float, default=None, help="total duration override (skips the ffprobe lookup)") ap.add_argument("--format", default="ffmetadata", choices=("ffmetadata", "youtube", "vtt", "json"), help="stdout format (default ffmetadata)") ap.add_argument("--write", metavar="OUT", default=None, help="mux chapters into a stream-copy of --media at this path") ap.add_argument("--json", action="store_true", help="emit JSON envelope on stdout (same as --format json)") args = ap.parse_args() json_mode = args.json or args.format == "json" detection = args.from_scenes or args.from_silence if (detection or args.write) and not args.media: err(json_mode, "USAGE", "--media is required for --from-scenes/--from-silence/--write", EXIT_USAGE) media = Path(args.media) if args.media else None if media and not media.is_file(): err(json_mode, "NOT_FOUND", f"media not found: {media}", EXIT_NOT_FOUND) if args.duration is not None: duration = args.duration elif media: duration = media_duration(media, json_mode) else: err(json_mode, "USAGE", "--duration is required when no --media is given", EXIT_USAGE) if args.chapters: chapters = load_chapters_file(Path(args.chapters), json_mode) chapters = [{**c} for c in chapters] else: mode = "scenes" if args.from_scenes else "silence" print(f"deriving chapter points from {mode}...", file=sys.stderr) points = detect_points(mode, media, json_mode) # type: ignore[arg-type] chapters = build_chapters(points, args.min_gap, duration) chapters = attach_ends(chapters, duration) meta = fmt_ffmetadata(chapters) written = None if args.write: mux_chapters(media, meta, Path(args.write), json_mode) # type: ignore[arg-type] written = str(Path(args.write)) print(f"chapters muxed -> {written}", file=sys.stderr) if json_mode: data = {"media": str(media) if media else None, "duration_s": duration, "count": len(chapters), "chapters": chapters, "written": written} print(json.dumps({"data": data, "meta": {"schema": SCHEMA}}, indent=2)) elif args.format == "youtube": sys.stdout.write(fmt_youtube(chapters)) elif args.format == "vtt": sys.stdout.write(fmt_vtt(chapters)) else: sys.stdout.write(meta) return EXIT_OK if __name__ == "__main__": sys.exit(main()) -
make-sprites.py 6.5 KB
#!/usr/bin/env python3 """Scrub-preview sprites + WebVTT thumbnail track for web players. Renders tiled sprite sheets at a fixed interval and writes the thumbs.vtt that maps each time range to its sprite region (#xywh media fragments) — the format Video.js / JW Player / Plyr / hls.js preview plugins consume. The geometry math (page, row, column per thumb) is exactly the part worth never re-deriving. Usage: make-sprites.py [--interval S] [--width PX] [--cols N] [--rows N] [--out-dir DIR] [--json] <media> Input: one video file as positional Output: stdout = written file list (or --json envelope, schema claude-mods.ffmpeg-ops.sprites/v1) Stderr: progress, errors Exit: 0 ok, 2 usage, 3 file not found, 4 probe/render failure, 5 ffmpeg missing Examples: make-sprites.py --interval 5 video.mp4 make-sprites.py --interval 10 --width 240 --out-dir previews/ lecture.mp4 make-sprites.py --json video.mp4 | jq -r '.data.vtt' """ import argparse import json import math import shutil import subprocess import sys from pathlib import Path from typing import NoReturn SCHEMA = "claude-mods.ffmpeg-ops.sprites/v1" EXIT_OK, EXIT_USAGE, EXIT_NOT_FOUND, EXIT_VALIDATION, EXIT_MISSING_DEP = 0, 2, 3, 4, 5 def err(json_mode: bool, code: str, message: str, exit_code: int) -> NoReturn: if json_mode: print(json.dumps({"error": {"code": code, "message": message, "details": {}}})) print(f"ERROR: {message}", file=sys.stderr) sys.exit(exit_code) def probe(ffprobe: str, path: Path) -> dict: # Full -show_streams, not selective -show_entries: the rotation side data # (side_data_list) is silently omitted by entry-filtered queries on some # ffprobe versions, which made rotated sources produce squashed thumbs. proc = subprocess.run( [ffprobe, "-v", "error", "-select_streams", "v:0", "-print_format", "json", "-show_streams", "-show_format", str(path)], capture_output=True, text=True) if proc.returncode != 0: return {} raw = json.loads(proc.stdout) streams = raw.get("streams", []) if not streams: return {} s = streams[0] rotation = 0 for sd in s.get("side_data_list", []) or []: try: rotation = int(sd.get("rotation", 0)) % 360 except (TypeError, ValueError): pass w, h = s.get("width", 0), s.get("height", 0) if rotation in (90, 270): # ffmpeg autorotates on decode; sprites show display dims w, h = h, w return {"width": w, "height": h, "duration": float(raw.get("format", {}).get("duration", 0) or 0)} def ts(seconds: float) -> str: h, rem = divmod(int(seconds), 3600) m, s = divmod(rem, 60) return f"{h:02d}:{m:02d}:{s:02d}.{int(round((seconds % 1) * 1000)):03d}" def main() -> int: ap = argparse.ArgumentParser( description="Sprite sheets + WebVTT thumbnail track for player scrub previews.", epilog="Examples:\n" " make-sprites.py --interval 5 video.mp4\n" " make-sprites.py --interval 10 --width 240 --out-dir previews/ in.mp4\n", formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("file", help="video file") ap.add_argument("--interval", type=float, default=5.0, help="seconds per thumbnail (default 5)") ap.add_argument("--width", type=int, default=160, help="thumbnail width in px (default 160)") ap.add_argument("--cols", type=int, default=10, help="grid columns (default 10)") ap.add_argument("--rows", type=int, default=10, help="grid rows (default 10)") ap.add_argument("--out-dir", default="sprites", help="output dir (default ./sprites)") ap.add_argument("--json", action="store_true", help="emit JSON envelope on stdout") args = ap.parse_args() if args.interval <= 0 or args.width < 16 or args.cols < 1 or args.rows < 1: err(args.json, "USAGE", "interval/width/cols/rows out of range", EXIT_USAGE) ffmpeg, ffprobe = shutil.which("ffmpeg"), shutil.which("ffprobe") if not ffmpeg or not ffprobe: err(args.json, "MISSING_DEPENDENCY", "ffmpeg/ffprobe not found on PATH", EXIT_MISSING_DEP) path = Path(args.file) if not path.is_file(): err(args.json, "NOT_FOUND", f"file not found: {path}", EXIT_NOT_FOUND) info = probe(ffprobe, path) if not info or not info["width"] or info["duration"] <= 0: err(args.json, "VALIDATION", "no probeable video stream/duration", EXIT_VALIDATION) # Explicit even thumb height so our geometry and ffmpeg's agree exactly. tw = args.width // 2 * 2 th = max(2, round(tw * info["height"] / info["width"] / 2) * 2) per_page = args.cols * args.rows n_thumbs = max(1, math.ceil(info["duration"] / args.interval)) n_pages = math.ceil(n_thumbs / per_page) out_dir = Path(args.out_dir) out_dir.mkdir(parents=True, exist_ok=True) print(f"{n_thumbs} thumbs ({tw}x{th}) on {n_pages} sheet(s)...", file=sys.stderr) proc = subprocess.run( [ffmpeg, "-y", "-v", "error", "-i", str(path.resolve()), "-vf", f"fps=1/{args.interval},scale={tw}:{th},tile={args.cols}x{args.rows}", "-q:v", "3", "sprite_%02d.jpg"], capture_output=True, text=True, cwd=str(out_dir)) if proc.returncode != 0: err(args.json, "VALIDATION", f"sprite render failed: {(proc.stderr.strip().splitlines() or ['?'])[-1]}", EXIT_VALIDATION) sheets = sorted(out_dir.glob("sprite_*.jpg")) lines = ["WEBVTT", ""] for i in range(n_thumbs): t0 = i * args.interval t1 = min((i + 1) * args.interval, info["duration"]) page = i // per_page + 1 idx = i % per_page x, y = (idx % args.cols) * tw, (idx // args.cols) * th lines += [f"{ts(t0)} --> {ts(t1)}", f"sprite_{page:02d}.jpg#xywh={x},{y},{tw},{th}", ""] vtt = out_dir / "thumbs.vtt" vtt.write_text("\n".join(lines), encoding="utf-8") data = {"media": str(path), "thumbs": n_thumbs, "thumb_size": [tw, th], "grid": [args.cols, args.rows], "interval_s": args.interval, "sheets": [str(p) for p in sheets], "vtt": str(vtt)} if args.json: print(json.dumps({"data": data, "meta": {"schema": SCHEMA}}, indent=2)) else: for p in [*sheets, vtt]: print(p) print(f"done: point the player's thumbnail track at {vtt.name} " f"(URLs resolve relative to the VTT)", file=sys.stderr) return EXIT_OK if __name__ == "__main__": sys.exit(main()) -
probe-media.py 14.3 KB
#!/usr/bin/env python3 """Normalized media inspection via ffprobe — the probe-first doctrine's tool. Wraps ffprobe's verbose, build-varying JSON into one stable, compact envelope: container, duration, per-stream codec/dimensions/fps/pix_fmt/color/rotation, and (on request) the keyframes nearest a timestamp so the agent can decide whether a stream-copy cut is safe. --doctor turns the probe into triage: each detected processing hazard (VFR, HDR transfer, rotation metadata, interlacing, non-yuv420p delivery, moov at EOF) is reported WITH the exact fix command, and the exit code becomes a branchable signal. Usage: probe-media.py [--json] [--keyframes-near SECONDS] [--doctor] <file> Input: one media file path as positional Output: stdout = human summary, or envelope {"data":...,"meta":...} with --json (schema claude-mods.ffmpeg-ops.probe/v1) Stderr: warnings, errors Exit: 0 ok, 2 usage, 3 file not found, 4 not parseable media, 5 ffprobe missing, 10 --doctor found at least one issue Examples: probe-media.py input.mp4 probe-media.py --json input.mp4 | jq '.data.video.fps' probe-media.py --keyframes-near 92.5 input.mp4 probe-media.py --doctor input.mp4 || echo "fix before processing" probe-media.py --doctor --json input.mp4 | jq -r '.data.doctor.findings[].fix' """ import argparse import json import shutil import subprocess import sys from fractions import Fraction from pathlib import Path from typing import NoReturn SCHEMA = "claude-mods.ffmpeg-ops.probe/v1" EXIT_OK, EXIT_USAGE, EXIT_NOT_FOUND, EXIT_VALIDATION, EXIT_MISSING_DEP = 0, 2, 3, 4, 5 EXIT_FINDINGS = 10 def err(args_json: bool, code: str, message: str, exit_code: int) -> NoReturn: if args_json: print(json.dumps({"error": {"code": code, "message": message, "details": {}}})) print(f"ERROR: {message}", file=sys.stderr) sys.exit(exit_code) def parse_rate(rate: str) -> float: """ffprobe rates arrive as '30000/1001' or '25/1'; '0/0' means unknown.""" try: f = Fraction(rate) return round(float(f), 3) if f else 0.0 except (ValueError, ZeroDivisionError): return 0.0 def stream_rotation(stream: dict) -> int: # Modern ffprobe: displaymatrix side data; legacy: tags.rotate. for sd in stream.get("side_data_list", []) or []: if "rotation" in sd: try: return int(sd["rotation"]) % 360 except (TypeError, ValueError): pass try: return int(stream.get("tags", {}).get("rotate", 0)) % 360 except (TypeError, ValueError): return 0 def normalize(raw: dict, path: Path) -> dict: fmt = raw.get("format", {}) out = { "file": str(path), "container": fmt.get("format_name", ""), "duration_s": round(float(fmt.get("duration", 0) or 0), 3), "size_bytes": int(fmt.get("size", 0) or 0), "bitrate_bps": int(fmt.get("bit_rate", 0) or 0), "stream_count": int(fmt.get("nb_streams", 0) or 0), "video": None, "audio": [], "subtitles": [], "streams": [], } for s in raw.get("streams", []): kind = s.get("codec_type", "unknown") entry = { "index": s.get("index"), "type": kind, "codec": s.get("codec_name", ""), "profile": s.get("profile", ""), "language": (s.get("tags", {}) or {}).get("language", ""), "default": bool((s.get("disposition", {}) or {}).get("default", 0)), } if kind == "video": avg = parse_rate(s.get("avg_frame_rate", "0/0")) real = parse_rate(s.get("r_frame_rate", "0/0")) entry.update({ "width": s.get("width", 0), "height": s.get("height", 0), "fps": avg or real, # avg != r is the cheap variable-frame-rate tell. "vfr_suspect": bool(avg and real and abs(avg - real) > 0.01), "pix_fmt": s.get("pix_fmt", ""), "field_order": s.get("field_order", ""), "color_space": s.get("color_space", ""), "color_transfer": s.get("color_transfer", ""), "color_primaries": s.get("color_primaries", ""), "rotation_deg": stream_rotation(s), "bitrate_bps": int(s.get("bit_rate", 0) or 0), }) if out["video"] is None and not s.get("disposition", {}).get("attached_pic"): out["video"] = entry elif kind == "audio": entry.update({ "sample_rate": int(s.get("sample_rate", 0) or 0), "channels": s.get("channels", 0), "channel_layout": s.get("channel_layout", ""), "bitrate_bps": int(s.get("bit_rate", 0) or 0), }) out["audio"].append(entry) elif kind == "subtitle": out["subtitles"].append(entry) out["streams"].append(entry) return out def moov_after_mdat(path: Path) -> bool: """Walk top-level MP4/MOV atoms: True if moov sits after mdat (no faststart).""" try: with path.open("rb") as f: pos, size = 0, path.stat().st_size seen_mdat = False while pos + 8 <= size: f.seek(pos) header = f.read(16) if len(header) < 8: break box_len = int.from_bytes(header[0:4], "big") box_type = header[4:8] if box_len == 1 and len(header) >= 16: # 64-bit largesize box_len = int.from_bytes(header[8:16], "big") elif box_len == 0: # box runs to EOF box_len = size - pos if box_len < 8: break if box_type == b"mdat": seen_mdat = True elif box_type == b"moov": return seen_mdat pos += box_len except OSError: pass return False def doctor(data: dict, path: Path) -> list: """Triage: each finding pairs the hazard with the exact fix command.""" findings = [] q = f'"{path}"' v = data["video"] def add(severity: str, issue: str, why: str, fix: str) -> None: findings.append({"severity": severity, "issue": issue, "why": why, "fix": fix}) if v: if v["vfr_suspect"]: add("warn", "variable frame rate (VFR) suspected", "cut math drifts, concat desyncs, players/editors stutter", f"ffmpeg -i {q} -c:v libx264 -crf 18 -preset fast -pix_fmt yuv420p " f"-fps_mode cfr -r {round(v['fps']) or 30} -c:a aac -b:a 192k normalized.mp4") if v["color_transfer"] in ("smpte2084", "arib-std-b67"): kind = "PQ/HDR10" if v["color_transfer"] == "smpte2084" else "HLG" add("warn", f"HDR transfer ({kind})", "re-encoding without tonemapping produces grey, washed-out SDR", f"ffmpeg -i {q} -vf \"zscale=t=linear:npl=100,format=gbrpf32le," f"zscale=p=bt709,tonemap=tonemap=hable:desat=0," f"zscale=t=bt709:m=bt709:r=tv,format=yuv420p\" " f"-c:v libx264 -crf 20 -c:a copy sdr.mp4") if v["rotation_deg"]: add("warn", f"rotation metadata ({v['rotation_deg']} deg)", "filters/thumbnails operate on unrotated pixels; some pipelines drop the flag", f"ffmpeg -display_rotation 0 -i {q} -c copy upright.mp4 " f"# or bake: -vf transpose + re-encode") if v["field_order"] not in ("", "progressive", "unknown"): add("warn", f"interlaced (field_order={v['field_order']})", "combing artifacts on motion after any scale/re-encode", f"ffmpeg -i {q} -vf bwdif=mode=send_field -c:v libx264 -crf 19 " f"-c:a copy deinterlaced.mp4") # H.264 delivery must be 8-bit 4:2:0; HEVC Main10 (yuv420p10le) is a # legitimate delivery profile (and mandatory for HDR10) — don't flag it. ok_pix = ("", "yuv420p") if v["codec"] == "h264" else \ ("", "yuv420p", "yuv420p10le") if v["codec"] in ("h264", "hevc") and v["pix_fmt"] not in ok_pix: add("warn", f"pix_fmt {v['pix_fmt']} on a delivery codec", "Safari/QuickTime/TVs show black or refuse playback on >4:2:0", f"ffmpeg -i {q} -c:v libx264 -crf 18 -pix_fmt yuv420p -c:a copy " f"-movflags +faststart compatible.mp4") elif data["audio"]: add("info", "no video stream (audio-only)", "video operations will fail; audio/STT workflows are fine", "") if "mp4" in data["container"] or "mov" in data["container"]: if moov_after_mdat(path): add("warn", "moov atom after mdat (no faststart)", "browsers must download the whole file before playback starts", f"ffmpeg -i {q} -c copy -movflags +faststart faststart.mp4") if data["duration_s"] <= 0: add("warn", "container reports no duration", "truncated/still-recording file, or a stream needing -fflags +genpts", f"ffmpeg -v error -i {q} -f null - # decode check; then remux -c copy") return findings def keyframes_near(ffprobe: str, path: Path, ts: float, window: float = 30.0) -> dict: start = max(0.0, ts - window) proc = subprocess.run( [ffprobe, "-v", "error", "-select_streams", "v:0", "-show_entries", "packet=pts_time,flags", "-of", "csv=p=0", "-read_intervals", f"{start}%{ts + window}", str(path)], capture_output=True, text=True) keys = [] for line in proc.stdout.splitlines(): parts = line.strip().split(",") if len(parts) >= 2 and "K" in parts[1]: try: keys.append(float(parts[0])) except ValueError: continue keys.sort() prev = max((k for k in keys if k <= ts), default=None) nxt = min((k for k in keys if k > ts), default=None) return { "target_s": ts, "prev_keyframe_s": prev, "next_keyframe_s": nxt, "copy_cut_drift_s": round(ts - prev, 3) if prev is not None else None, "window_scanned_s": [round(start, 3), round(ts + window, 3)], } def main() -> int: ap = argparse.ArgumentParser( description="Normalized media inspection via ffprobe.", epilog="Examples:\n" " probe-media.py input.mp4\n" " probe-media.py --json input.mp4 | jq '.data.video.fps'\n" " probe-media.py --keyframes-near 92.5 input.mp4\n", formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("file", help="media file to probe") ap.add_argument("--json", action="store_true", help="emit JSON envelope on stdout") ap.add_argument("--keyframes-near", type=float, metavar="SECONDS", default=None, help="also report nearest keyframes to this timestamp") ap.add_argument("--doctor", action="store_true", help="triage mode: report processing hazards with exact fix " "commands; exit 10 if any found") args = ap.parse_args() ffprobe = shutil.which("ffprobe") if not ffprobe: err(args.json, "MISSING_DEPENDENCY", "ffprobe not found on PATH (install ffmpeg)", EXIT_MISSING_DEP) path = Path(args.file) if not path.is_file(): err(args.json, "NOT_FOUND", f"file not found: {path}", EXIT_NOT_FOUND) proc = subprocess.run( [ffprobe, "-v", "error", "-print_format", "json", "-show_format", "-show_streams", str(path)], capture_output=True, text=True) if proc.returncode != 0 or not proc.stdout.strip(): err(args.json, "VALIDATION", f"ffprobe could not parse '{path.name}' as media: " f"{proc.stderr.strip().splitlines()[-1] if proc.stderr.strip() else 'no detail'}", EXIT_VALIDATION) data = normalize(json.loads(proc.stdout), path) if args.keyframes_near is not None: if data["video"] is None: err(args.json, "VALIDATION", "no video stream; --keyframes-near needs one", EXIT_VALIDATION) data["keyframes"] = keyframes_near(ffprobe, path, args.keyframes_near) findings = [] if args.doctor: findings = doctor(data, path) has_warn = any(f["severity"] != "info" for f in findings) data["doctor"] = {"findings": findings, "clean": not has_warn} if args.json: print(json.dumps({"data": data, "meta": {"schema": SCHEMA}}, indent=2)) if args.doctor and not data["doctor"]["clean"]: return EXIT_FINDINGS return EXIT_OK # Human summary (stdout is still the data product — keep it grep-friendly). v = data["video"] print(f"file {data['file']}") print(f"container {data['container']} " f"{data['duration_s']}s {data['size_bytes']} bytes " f"{data['bitrate_bps'] // 1000} kb/s {data['stream_count']} streams") if v: vfr = " VFR-SUSPECT" if v["vfr_suspect"] else "" rot = f" rotation={v['rotation_deg']}" if v["rotation_deg"] else "" print(f"video {v['codec']} {v['width']}x{v['height']} " f"{v['fps']}fps {v['pix_fmt']}{rot}{vfr}") if v["color_space"] or v["color_transfer"]: print(f"color space={v['color_space'] or '?'} " f"transfer={v['color_transfer'] or '?'} " f"primaries={v['color_primaries'] or '?'}") for a in data["audio"]: print(f"audio #{a['index']} {a['codec']} {a['sample_rate']}Hz " f"{a['channels']}ch {a['channel_layout']} lang={a['language'] or '-'}") for s in data["subtitles"]: print(f"subs #{s['index']} {s['codec']} lang={s['language'] or '-'}") if "keyframes" in data: k = data["keyframes"] print(f"keyframes target={k['target_s']}s " f"prev={k['prev_keyframe_s']}s next={k['next_keyframe_s']}s " f"copy-cut-drift={k['copy_cut_drift_s']}s") if args.doctor: if not findings: print("doctor clean — no processing hazards detected") for f in findings: print(f"doctor [{f['severity']}] {f['issue']} — {f['why']}") if f["fix"]: print(f" fix: {f['fix']}") if not data["doctor"]["clean"]: return EXIT_FINDINGS return EXIT_OK if __name__ == "__main__": sys.exit(main()) -
quality-compare.py 8.7 KB
#!/usr/bin/env python3 """Objective quality verdict on an encode — VMAF/SSIM/PSNR vs the reference. Closes the encode loop: "did my compression actually look ok" becomes a number and an exit code the caller can branch on. Handles the resolution mismatch case (distorted is auto-scaled to reference dimensions before comparison) and parses the metric filters' log-text output so the agent never has to. Usage: quality-compare.py [--metrics LIST] [--min-vmaf N] [--min-ssim N] [--json] <reference> <distorted> Input: reference (original) and distorted (encoded) files as positionals Output: stdout = metric lines (or --json envelope, schema claude-mods.ffmpeg-ops.quality/v1) Stderr: progress, errors Exit: 0 ok / at-or-above thresholds, 2 usage, 3 input missing, 4 metric parse failure, 5 ffmpeg missing (or libvmaf absent when vmaf requested), 10 BELOW a requested threshold Guide: VMAF >= 93 at 1080p ~ visually transparent; 80-93 noticeable on inspection; < 80 visibly degraded. SSIM >= 0.98 ~ excellent. Examples: quality-compare.py original.mp4 encoded.mp4 quality-compare.py original.mp4 encoded.mp4 --metrics vmaf --min-vmaf 90 quality-compare.py original.mp4 encoded.mp4 --metrics ssim,psnr --min-ssim 0.97 quality-compare.py original.mp4 encoded.mp4 --metrics vmaf --json | jq '.data.vmaf' """ import argparse import json import re import shutil import subprocess import sys import tempfile from pathlib import Path from typing import NoReturn, Optional SCHEMA = "claude-mods.ffmpeg-ops.quality/v1" EXIT_OK, EXIT_USAGE, EXIT_NOT_FOUND, EXIT_VALIDATION = 0, 2, 3, 4 EXIT_MISSING_DEP, EXIT_BELOW_THRESHOLD = 5, 10 def err(json_mode: bool, code: str, message: str, exit_code: int) -> NoReturn: if json_mode: print(json.dumps({"error": {"code": code, "message": message, "details": {}}})) print(f"ERROR: {message}", file=sys.stderr) sys.exit(exit_code) def video_dims(ffprobe: str, path: Path) -> Optional[tuple]: proc = subprocess.run( [ffprobe, "-v", "error", "-select_streams", "v:0", "-show_entries", "stream=width,height", "-of", "csv=p=0", str(path)], capture_output=True, text=True) parts = proc.stdout.strip().split(",") if len(parts) == 2 and all(p.isdigit() for p in parts): return int(parts[0]), int(parts[1]) return None def has_filter(ffmpeg: str, name: str) -> bool: proc = subprocess.run([ffmpeg, "-hide_banner", "-filters"], capture_output=True, text=True) return bool(re.search(rf"^\s+[A-Z.|]+\s+{re.escape(name)}\s+", proc.stdout, re.MULTILINE)) def run_metric(ffmpeg: str, ref: Path, dist: Path, scale: str, metric_filter: str, cwd: Optional[str] = None) -> subprocess.CompletedProcess: # libvmaf/ssim/psnr convention: first input = distorted, second = reference. # cwd is set for vmaf so log_path can be a bare filename — a full Windows # path inside the filter arg hits the drive-colon escaping trap. graph = f"[0:v]{scale}[d];[d][1:v]{metric_filter}" if scale \ else f"[0:v][1:v]{metric_filter}" return subprocess.run( [ffmpeg, "-hide_banner", "-nostats", "-i", str(dist.resolve()), "-i", str(ref.resolve()), "-filter_complex", graph, "-f", "null", "-"], capture_output=True, text=True, cwd=cwd) def main() -> int: ap = argparse.ArgumentParser( description="VMAF/SSIM/PSNR quality verdict: encoded vs reference.", epilog="Examples:\n" " quality-compare.py original.mp4 encoded.mp4\n" " quality-compare.py original.mp4 encoded.mp4 --metrics vmaf --min-vmaf 90\n", formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("reference", help="original/reference file") ap.add_argument("distorted", help="encoded/processed file to judge") ap.add_argument("--metrics", default="ssim,psnr", help="comma list of ssim,psnr,vmaf (default ssim,psnr)") ap.add_argument("--min-vmaf", type=float, default=None, help="exit 10 if VMAF score is below this") ap.add_argument("--min-ssim", type=float, default=None, help="exit 10 if SSIM (All) is below this") ap.add_argument("--json", action="store_true", help="emit JSON envelope on stdout") args = ap.parse_args() metrics = [m.strip().lower() for m in args.metrics.split(",") if m.strip()] bad = [m for m in metrics if m not in ("ssim", "psnr", "vmaf")] if bad or not metrics: err(args.json, "USAGE", f"unknown metric(s): {', '.join(bad) or '(none)'}", EXIT_USAGE) if args.min_vmaf is not None and "vmaf" not in metrics: metrics.append("vmaf") ffmpeg, ffprobe = shutil.which("ffmpeg"), shutil.which("ffprobe") if not ffmpeg or not ffprobe: err(args.json, "MISSING_DEPENDENCY", "ffmpeg/ffprobe not found on PATH", EXIT_MISSING_DEP) ref, dist = Path(args.reference), Path(args.distorted) for p in (ref, dist): if not p.is_file(): err(args.json, "NOT_FOUND", f"file not found: {p}", EXIT_NOT_FOUND) if "vmaf" in metrics and not has_filter(ffmpeg, "libvmaf"): err(args.json, "MISSING_DEPENDENCY", "this ffmpeg build lacks libvmaf (install a full build, e.g. " "gyan.dev 'full' on Windows, or use --metrics ssim,psnr)", EXIT_MISSING_DEP) ref_dims, dist_dims = video_dims(ffprobe, ref), video_dims(ffprobe, dist) if not ref_dims or not dist_dims: err(args.json, "VALIDATION", "could not read video dimensions from inputs", EXIT_VALIDATION) scale = "" if ref_dims != dist_dims: scale = f"scale={ref_dims[0]}:{ref_dims[1]}:flags=bicubic" print(f"note: scaling distorted {dist_dims[0]}x{dist_dims[1]} -> " f"{ref_dims[0]}x{ref_dims[1]} for comparison", file=sys.stderr) results: dict = {} for metric in metrics: print(f"running {metric}...", file=sys.stderr) if metric == "vmaf": with tempfile.TemporaryDirectory() as td: log = Path(td) / "vmaf.json" proc = run_metric(ffmpeg, ref, dist, scale, "libvmaf=log_fmt=json:log_path=vmaf.json", cwd=td) if proc.returncode != 0 or not log.is_file(): err(args.json, "VALIDATION", f"vmaf run failed: {(proc.stderr.strip().splitlines() or ['?'])[-1]}", EXIT_VALIDATION) vmaf_data = json.loads(log.read_text()) pooled = vmaf_data.get("pooled_metrics", {}).get("vmaf", {}) results["vmaf"] = {"mean": round(pooled.get("mean", 0.0), 2), "min": round(pooled.get("min", 0.0), 2), "harmonic_mean": round(pooled.get("harmonic_mean", 0.0), 2)} elif metric == "ssim": proc = run_metric(ffmpeg, ref, dist, scale, "ssim") m = re.search(r"SSIM.*All:([\d.]+)", proc.stderr) if not m: err(args.json, "VALIDATION", "could not parse SSIM output", EXIT_VALIDATION) results["ssim"] = {"all": float(m.group(1))} elif metric == "psnr": proc = run_metric(ffmpeg, ref, dist, scale, "psnr") m = re.search(r"PSNR.*average:([\d.]+|inf)", proc.stderr) if not m: err(args.json, "VALIDATION", "could not parse PSNR output", EXIT_VALIDATION) val = m.group(1) results["psnr"] = {"average_db": float("inf") if val == "inf" else float(val)} below = [] if args.min_vmaf is not None and results.get("vmaf", {}).get("mean", 1e9) < args.min_vmaf: below.append(f"vmaf {results['vmaf']['mean']} < {args.min_vmaf}") if args.min_ssim is not None and results.get("ssim", {}).get("all", 1e9) < args.min_ssim: below.append(f"ssim {results['ssim']['all']} < {args.min_ssim}") data = {"reference": str(ref), "distorted": str(dist), "scaled_for_comparison": bool(scale), "thresholds": {"min_vmaf": args.min_vmaf, "min_ssim": args.min_ssim}, "below_threshold": below, **results} if args.json: print(json.dumps({"data": data, "meta": {"schema": SCHEMA}}, indent=2)) else: for name, vals in results.items(): flat = " ".join(f"{k}={v}" for k, v in vals.items()) print(f"{name}\t{flat}") for b in below: print(f"below-threshold\t{b}") if below: print(f"VERDICT: below threshold ({'; '.join(below)})", file=sys.stderr) return EXIT_BELOW_THRESHOLD return EXIT_OK if __name__ == "__main__": sys.exit(main()) -
smart-compress.py 10.1 KB
#!/usr/bin/env python3 """Target-size compression: 'make this fit in 25MB' as one verified command. Computes the video bitrate from the size budget (duration-aware, audio and mux overhead subtracted), auto-selects an audio bitrate and a downscale rung when the bits-per-pixel would be hopeless at source resolution, runs a two-pass encode (predictable size, unlike CRF), and VERIFIES the result actually landed under the cap — retrying once at -8% if not. Usage: smart-compress.py --target SIZE [-o OUT] [--codec x264|x265] [--preset P] [--no-downscale] [--json] <file> Input: one media file as positional; SIZE like 25MB, 8M, 512KB, 1.5GB Output: stdout = result line (or --json envelope, schema claude-mods.ffmpeg-ops.compress/v1) Stderr: progress, plan explanation, errors Exit: 0 ok and under target, 2 usage, 3 input missing, 4 encode failure, 5 ffmpeg missing, 10 best effort still OVER target (kept, caller decides) Examples: smart-compress.py --target 25MB video.mp4 # Discord/email cap smart-compress.py --target 8MB -o clip_small.mp4 clip.mov smart-compress.py --target 50MB --codec x265 lecture.mp4 smart-compress.py --target 10MB --json in.mp4 | jq '.data.final_bytes' """ import argparse import json import re import shutil import subprocess import sys import tempfile from pathlib import Path from typing import NoReturn, Optional SCHEMA = "claude-mods.ffmpeg-ops.compress/v1" EXIT_OK, EXIT_USAGE, EXIT_NOT_FOUND, EXIT_VALIDATION = 0, 2, 3, 4 EXIT_MISSING_DEP, EXIT_OVER_TARGET = 5, 10 MUX_OVERHEAD = 0.98 # reserve 2% of the budget for container overhead DOWNSCALE_LADDER = [1080, 720, 540, 360, 270] MIN_BPP = 0.045 # below this bits-per-pixel, downscale instead def err(json_mode: bool, code: str, message: str, exit_code: int) -> NoReturn: if json_mode: print(json.dumps({"error": {"code": code, "message": message, "details": {}}})) print(f"ERROR: {message}", file=sys.stderr) sys.exit(exit_code) def parse_size(s: str) -> Optional[int]: m = re.fullmatch(r"([\d.]+)\s*([KMG]i?B?|B)?", s.strip(), re.IGNORECASE) if not m: return None mult = {"": 1, "B": 1, "K": 1000, "M": 1000**2, "G": 1000**3, "KI": 1024, "MI": 1024**2, "GI": 1024**3} unit = (m.group(2) or "").upper().rstrip("B") try: return int(float(m.group(1)) * mult[unit]) except (KeyError, ValueError): return None def probe(ffprobe: str, path: Path) -> dict: proc = subprocess.run( [ffprobe, "-v", "error", "-print_format", "json", "-show_format", "-show_streams", str(path)], capture_output=True, text=True) if proc.returncode != 0: return {} raw = json.loads(proc.stdout) out = {"duration": float(raw.get("format", {}).get("duration", 0) or 0), "size": int(raw.get("format", {}).get("size", 0) or 0), "width": 0, "height": 0, "fps": 30.0, "has_audio": False} for s in raw.get("streams", []): if s.get("codec_type") == "video" and not out["width"]: out["width"], out["height"] = s.get("width", 0), s.get("height", 0) try: num, den = s.get("avg_frame_rate", "30/1").split("/") out["fps"] = (int(num) / int(den)) if int(den) else 30.0 except (ValueError, ZeroDivisionError): pass elif s.get("codec_type") == "audio": out["has_audio"] = True return out def plan_encode(info: dict, target_bytes: int, allow_downscale: bool) -> dict: budget_kbps = (target_bytes * 8 / 1000) / info["duration"] * MUX_OVERHEAD # Audio gets ~12% of the budget, clamped to sane speech/music rates. audio_kbps = int(min(160, max(48, budget_kbps * 0.12))) if info["has_audio"] else 0 video_kbps = budget_kbps - audio_kbps w, h, fps = info["width"], info["height"], info["fps"] or 30.0 scaled_h = None if allow_downscale and w and h: bpp = video_kbps * 1000 / (w * h * fps) if bpp < MIN_BPP: for rung in DOWNSCALE_LADDER: if rung >= h: continue rw = w * rung / h if video_kbps * 1000 / (rw * rung * fps) >= MIN_BPP: scaled_h = rung break else: scaled_h = DOWNSCALE_LADDER[-1] if h > DOWNSCALE_LADDER[-1] else None return {"video_kbps": int(video_kbps), "audio_kbps": audio_kbps, "scale_height": scaled_h} def two_pass(ffmpeg: str, path: Path, out: Path, plan: dict, codec: str, preset: str, json_mode: bool) -> None: enc = {"x264": "libx264", "x265": "libx265"}[codec] vf = ["-vf", f"scale=-2:{plan['scale_height']}"] if plan["scale_height"] else [] audio = (["-c:a", "aac", "-b:a", f"{plan['audio_kbps']}k", "-ar", "48000"] if plan["audio_kbps"] else ["-an"]) tag = ["-tag:v", "hvc1"] if codec == "x265" else [] with tempfile.TemporaryDirectory() as td: passlog = str(Path(td) / "ffpass") base = [ffmpeg, "-y", "-v", "error", "-i", str(path), "-c:v", enc, "-b:v", f"{plan['video_kbps']}k", "-preset", preset, "-pix_fmt", "yuv420p", *tag, *vf, "-passlogfile", passlog] p1 = subprocess.run([*base, "-pass", "1", "-an", "-f", "null", "NUL" if sys.platform == "win32" else "/dev/null"], capture_output=True, text=True) if p1.returncode != 0: err(json_mode, "VALIDATION", f"pass 1 failed: {(p1.stderr.strip().splitlines() or ['?'])[-1]}", EXIT_VALIDATION) p2 = subprocess.run([*base, "-pass", "2", *audio, "-movflags", "+faststart", str(out)], capture_output=True, text=True) if p2.returncode != 0: err(json_mode, "VALIDATION", f"pass 2 failed: {(p2.stderr.strip().splitlines() or ['?'])[-1]}", EXIT_VALIDATION) def main() -> int: ap = argparse.ArgumentParser( description="Compress a video to fit a size target (two-pass, verified).", epilog="Examples:\n" " smart-compress.py --target 25MB video.mp4\n" " smart-compress.py --target 8MB -o small.mp4 clip.mov\n", formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("file", help="input media file") ap.add_argument("--target", required=True, metavar="SIZE", help="size cap, e.g. 25MB, 8M, 1.5GB (MiB/GiB also accepted)") ap.add_argument("-o", "--output", default=None, help="output path (default <stem>.compressed.mp4)") ap.add_argument("--codec", default="x264", choices=("x264", "x265"), help="x264 = universal (default); x265 = ~40%% smaller, modern players") ap.add_argument("--preset", default="slow", help="encoder preset (default slow; use medium/fast for speed)") ap.add_argument("--no-downscale", action="store_true", help="never lower resolution, even at hopeless bits-per-pixel") ap.add_argument("--json", action="store_true", help="emit JSON envelope on stdout") args = ap.parse_args() target = parse_size(args.target) if not target or target <= 0: err(args.json, "USAGE", f"could not parse --target size: {args.target!r}", EXIT_USAGE) ffmpeg, ffprobe = shutil.which("ffmpeg"), shutil.which("ffprobe") if not ffmpeg or not ffprobe: err(args.json, "MISSING_DEPENDENCY", "ffmpeg/ffprobe not found on PATH", EXIT_MISSING_DEP) path = Path(args.file) if not path.is_file(): err(args.json, "NOT_FOUND", f"file not found: {path}", EXIT_NOT_FOUND) info = probe(ffprobe, path) if not info or info["duration"] <= 0: err(args.json, "VALIDATION", "could not probe input (no duration)", EXIT_VALIDATION) if info["size"] and info["size"] <= target: print(f"input is already {info['size']} bytes <= target {target} — " f"no encode needed (copy it as-is)", file=sys.stderr) out = Path(args.output) if args.output else path.with_name( path.stem + ".compressed.mp4") plan = plan_encode(info, target, not args.no_downscale) if plan["video_kbps"] < 50: err(args.json, "VALIDATION", f"budget gives only {plan['video_kbps']} kb/s video for " f"{info['duration']:.0f}s — target too small; trim the video or raise it", EXIT_VALIDATION) scale_note = f", downscale to {plan['scale_height']}p" if plan["scale_height"] else "" print(f"plan: video {plan['video_kbps']}k + audio {plan['audio_kbps']}k " f"({args.codec}, two-pass, preset {args.preset}{scale_note})", file=sys.stderr) attempts = [] current = dict(plan) for attempt in (1, 2): print(f"encoding (attempt {attempt})...", file=sys.stderr) two_pass(ffmpeg, path, out, current, args.codec, args.preset, args.json) size = out.stat().st_size attempts.append({"video_kbps": current["video_kbps"], "bytes": size}) if size <= target: break # Two-pass overshoot is rare but real on short/complex content: -8%. print(f"over target ({size} > {target}); retrying at -8% bitrate", file=sys.stderr) current["video_kbps"] = int(current["video_kbps"] * 0.92) final = out.stat().st_size data = {"input": str(path), "output": str(out), "target_bytes": target, "final_bytes": final, "under_target": final <= target, "plan": plan, "attempts": attempts} if args.json: print(json.dumps({"data": data, "meta": {"schema": SCHEMA}}, indent=2)) else: print(f"{out}\t{final}\t{'OK' if final <= target else 'OVER'}\t{target}") if final > target: print(f"best effort is still over target — kept at {final} bytes; " f"trim duration or accept a lower resolution", file=sys.stderr) return EXIT_OVER_TARGET print(f"done: {final} bytes ({100 * final / target:.0f}% of budget)", file=sys.stderr) return EXIT_OK if __name__ == "__main__": sys.exit(main()) -
verify-commands.sh 6.9 KB
#!/usr/bin/env bash # Staleness verifier for ffmpeg-ops docs — offline structural + live build-drift. # # --offline (default): structural integrity, NO ffmpeg needed. Assets parse as # JSON, every reference/script/asset on disk is cited from SKILL.md, and every # relative link in SKILL.md resolves. Runs in PR CI; may block. # --live: does the documentation still match an actual ffmpeg? Extracts the # encoders/filters the docs rely on and checks them against the INSTALLED # build (`-encoders`/`-filters`/`-h full`). Core items missing = drift # (exit 10); build-optional items (libx265, libvmaf, ...) only warn. # Runs in the scheduled freshness workflow; never blocks a PR. # # Usage: verify-commands.sh [--offline | --live] [-q] # Input: none (inspects the skill's own files; --live also the ffmpeg on PATH) # Output: stdout = findings (one per line, "DRIFT:" / "STRUCT:" prefixed) # Stderr: progress, warnings # Exit: 0 clean, 2 usage, 7 ffmpeg unavailable (--live only; advisory), # 10 drift/structural finding # # Examples: # verify-commands.sh --offline # verify-commands.sh --live # verify-commands.sh --live -q; echo "exit=$?" set -uo pipefail EXIT_OK=0; EXIT_USAGE=2; EXIT_UNAVAILABLE=7; EXIT_DRIFT=10 MODE="offline"; QUIET=0 while [[ $# -gt 0 ]]; do case "$1" in --offline) MODE="offline" ;; --live) MODE="live" ;; -q|--quiet) QUIET=1 ;; -h|--help) sed -n '2,26p' "$0" | sed 's/^# \{0,1\}//'; exit "$EXIT_OK" ;; *) echo "ERROR: unknown argument: $1 (try --help)" >&2; exit "$EXIT_USAGE" ;; esac shift done SKILL_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" SKILL_MD="$SKILL_DIR/SKILL.md" findings=0 emit() { [[ "$QUIET" -eq 1 ]] || printf '%s\n' "$1" >&2; } finding() { printf '%s\n' "$1"; findings=$((findings + 1)); } # Pick a working python for JSON validation (Windows Store stub exits non-zero). PYTHON="" for c in python3 python py; do if command -v "$c" >/dev/null 2>&1 && "$c" -c "" >/dev/null 2>&1; then PYTHON="$c"; break; fi done # ── offline: structural ────────────────────────────────────────────────────── offline_checks() { emit "== verify-commands --offline (structural)" [[ -f "$SKILL_MD" ]] || { finding "STRUCT: SKILL.md missing"; return; } # 1. assets parse as JSON for a in "$SKILL_DIR"/assets/*.json; do [[ -e "$a" ]] || continue if [[ -n "$PYTHON" ]]; then "$PYTHON" -c "import json,sys; json.load(open(sys.argv[1], encoding='utf-8'))" "$a" \ >/dev/null 2>&1 || finding "STRUCT: asset not valid JSON: $(basename "$a")" elif command -v jq >/dev/null 2>&1; then jq empty "$a" >/dev/null 2>&1 || finding "STRUCT: asset not valid JSON: $(basename "$a")" fi done # 2. every shipped resource is cited from SKILL.md (dead weight check) for d in references scripts assets; do for f in "$SKILL_DIR/$d"/*; do [[ -f "$f" ]] || continue base="$(basename "$f")" [[ "$base" == ".gitkeep" ]] && continue grep -q "$base" "$SKILL_MD" \ || finding "STRUCT: $d/$base exists on disk but is never cited from SKILL.md" done done # 3. every relative resource link in SKILL.md resolves while IFS= read -r path; do [[ -e "$SKILL_DIR/$path" ]] \ || finding "STRUCT: SKILL.md links to missing file: $path" done < <(grep -oE '\]\((references|assets|scripts|tests)/[^)#]+\)' "$SKILL_MD" \ | sed -E 's/^\]\(//; s/\)$//' | sort -u) } # ── live: docs vs the installed build ──────────────────────────────────────── live_checks() { emit "== verify-commands --live (installed-build drift)" if ! command -v ffmpeg >/dev/null 2>&1; then echo "ffmpeg not on PATH — live check unavailable (advisory, not a failure)" >&2 exit "$EXIT_UNAVAILABLE" fi local encoders filters hfull docs encoders="$(ffmpeg -hide_banner -encoders 2>/dev/null)" filters="$(ffmpeg -hide_banner -filters 2>/dev/null)" hfull="$(ffmpeg -hide_banner -h full 2>/dev/null)" docs="$(cat "$SKILL_MD" "$SKILL_DIR"/references/*.md 2>/dev/null)" # Filters that exist in EVERY ffmpeg build — absence means the filter was # renamed/removed upstream, i.e. our docs drifted. local core_filters=(scale crop pad fps overlay concat setpts atempo amix silencedetect silenceremove loudnorm palettegen paletteuse select tile transpose trim atrim split format) # Build-optional (external libs / hw): warn only. local optional_tokens=(libx264 libx265 libsvtav1 libaom-av1 libvpx-vp9 libopus libmp3lame drawtext subtitles lut3d zscale tonemap libvmaf minterpolate vidstabdetect vidstabtransform bwdif hqdn3d nlmeans xstack showwaves showspectrum colorbalance colortemperature colorchannelmixer colorhold vibrance haldclut chromashift) # CLI options the cookbook depends on; renamed/removed = drift (-vsync class). local core_options=(fps_mode movflags avoid_negative_ts map_metadata filter_complex frames pix_fmt) # NOTE: the flags column width varies across ffmpeg majors (3 chars <=7.x, # 2 chars in 8.x) — match any flag run, then the exact filter name token. for f in "${core_filters[@]}"; do grep -qE "^ +[A-Z.|]+ +$f +" <<<"$filters" \ || finding "DRIFT: core filter '$f' not in installed ffmpeg (renamed/removed upstream?)" done for opt in "${core_options[@]}"; do grep -q -- "-$opt" <<<"$hfull" \ || finding "DRIFT: documented option '-$opt' unknown to installed ffmpeg" done # Every software encoder the docs name must at least be a known encoder name # in this build — missing here is a warning (build config), not drift, EXCEPT # the universal natives (aac, ffv1) which every build ships. for enc in aac ffv1; do grep -qE "^ [A-Z.]{6} +$enc " <<<"$encoders" \ || finding "DRIFT: native encoder '$enc' not in installed ffmpeg" done for tok in "${optional_tokens[@]}"; do if grep -qF "$tok" <<<"$docs"; then grep -qE "(^ [A-Z.]{6} +$tok )|(^ +[A-Z.|]+ +$tok +)" <<<"$encoders"$'\n'"$filters" \ || emit " warn: '$tok' documented but absent from this build (build-optional — not drift)" fi done # Deprecated-flag tripwire: docs must not RECOMMEND -vsync (mentioning it as # deprecated in footgun tables is fine; a code fence using it is not). if grep -E '^\s*ffmpeg .*-vsync ' <<<"$docs" | grep -vq 'fps_mode'; then finding "DRIFT: a documented command still uses deprecated -vsync (use -fps_mode)" fi } case "$MODE" in offline) offline_checks ;; live) live_checks ;; esac if [[ "$findings" -eq 0 ]]; then emit "verify-commands ($MODE): clean" exit "$EXIT_OK" fi emit "verify-commands ($MODE): $findings finding(s)" exit "$EXIT_DRIFT"
-
-
tests
-
run.sh 15.7 KB
#!/usr/bin/env bash # Self-test for ffmpeg-ops scripts. # # Structural assertions always run (no ffmpeg needed): --help contracts, # py_compile/bash -n, documented exit codes on bad input, pure-python LUT # generation, EDL validation + dry-run, offline staleness verifier, asset JSON. # Media round-trips run ONLY when ffmpeg is on PATH — fixtures are synthesized # with lavfi (testsrc2/sine), so no binary fixtures live in the repo. Without # ffmpeg the media section is a LOUD skip, never a silent false-clean. # # Usage: bash tests/run.sh # Exit: 0 all pass, 1 one or more failures set -uo pipefail HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" SKILL="$(dirname "$HERE")" S="$SKILL/scripts" # Pick a python that actually executes (Windows Store python3 stub exits non-zero). PYTHON="" for c in python python3 py; do if command -v "$c" >/dev/null 2>&1 && "$c" -c "" >/dev/null 2>&1; then PYTHON="$c"; break; fi done [[ -z "$PYTHON" ]] && { echo "no working python found" >&2; exit 1; } SB="$(mktemp -d)"; trap 'rm -rf "$SB"' EXIT PASS=0; FAIL=0 ok() { PASS=$((PASS+1)); printf ' PASS %s\n' "$1"; } no() { FAIL=$((FAIL+1)); printf ' FAIL %s\n' "$1"; } expect_exit() { [[ "$2" == "$3" ]] && ok "$1 (exit $3)" || no "$1 (want $2 got $3)"; } expect_has() { case "$3" in *"$2"*) ok "$1";; *) no "$1 (missing '$2')";; esac; } echo "=== ffmpeg-ops self-test ===" # ── structural: every script honors the contract ──────────────────────────── echo "-- contracts --" for py in probe-media.py loudnorm-scan.py detect-segments.py quality-compare.py \ cut-from-edl.py gen-luts.py make-chapters.py smart-compress.py \ make-sprites.py; do "$PYTHON" -m py_compile "$S/$py" 2>/dev/null && ok "py_compile $py" || no "py_compile $py" "$PYTHON" "$S/$py" --help >/dev/null 2>&1; expect_exit "$py --help" 0 $? out="$("$PYTHON" "$S/$py" --help 2>/dev/null)"; expect_has "$py --help has Examples" "xamples" "$out" done for sh in capability-scan.sh verify-commands.sh; do bash -n "$S/$sh" 2>/dev/null && ok "bash -n $sh" || no "bash -n $sh" bash "$S/$sh" --help >/dev/null 2>&1; expect_exit "$sh --help" 0 $? done bash "$S/capability-scan.sh" --bogus-flag >/dev/null 2>&1; expect_exit "capability-scan unknown flag -> 2" 2 $? bash "$S/verify-commands.sh" --bogus >/dev/null 2>&1; expect_exit "verify-commands unknown flag -> 2" 2 $? # ── structural: documented failure exits ──────────────────────────────────── echo "-- exit codes --" "$PYTHON" "$S/probe-media.py" "$SB/nope.mp4" >/dev/null 2>&1 rc=$?; [[ "$rc" == 3 || "$rc" == 5 ]] && ok "probe missing file -> 3 (or 5 sans ffprobe; got $rc)" \ || no "probe missing file (want 3/5 got $rc)" "$PYTHON" "$S/cut-from-edl.py" "$SB/nope.json" >/dev/null 2>&1; expect_exit "edl missing -> 3" 3 $? printf 'not json' > "$SB/bad.json" "$PYTHON" "$S/cut-from-edl.py" "$SB/bad.json" >/dev/null 2>&1; expect_exit "edl not json -> 4" 4 $? printf '{"scenes":[]}' > "$SB/empty.json" "$PYTHON" "$S/cut-from-edl.py" "$SB/empty.json" >/dev/null 2>&1; expect_exit "edl empty scenes -> 4" 4 $? printf '{"scenes":[{"clips":[{"file":"a.mp4","start":5,"end":2}]}]}' > "$SB/inv.json" "$PYTHON" "$S/cut-from-edl.py" "$SB/inv.json" >/dev/null 2>&1; expect_exit "edl end<=start -> 4" 4 $? "$PYTHON" "$S/gen-luts.py" --variants not_a_look >/dev/null 2>&1; expect_exit "gen-luts unknown look -> 2" 2 $? "$PYTHON" "$S/quality-compare.py" --metrics bogus a b >/dev/null 2>&1; expect_exit "quality bad metric -> 2" 2 $? "$PYTHON" "$S/make-chapters.py" --from-scenes >/dev/null 2>&1; expect_exit "chapters detection w/o --media -> 2" 2 $? "$PYTHON" "$S/smart-compress.py" --target not_a_size x.mp4 >/dev/null 2>&1; expect_exit "smart-compress bad size -> 2" 2 $? "$PYTHON" "$S/make-sprites.py" --interval 0 x.mp4 >/dev/null 2>&1; expect_exit "make-sprites bad interval -> 2" 2 $? "$PYTHON" "$S/make-chapters.py" --chapters "$SB/nope.json" --duration 60 >/dev/null 2>&1; expect_exit "chapters file missing -> 3" 3 $? printf 'not json' > "$SB/badch.json" "$PYTHON" "$S/make-chapters.py" --chapters "$SB/badch.json" --duration 60 >/dev/null 2>&1; expect_exit "chapters bad json -> 4" 4 $? # ── structural: chapter formatting (no ffmpeg required via --duration) ─────── echo "-- make-chapters formats --" printf '[{"start":0,"title":"Intro"},{"start":65,"title":"Topic = One"},{"start":130,"title":"Wrap"}]' > "$SB/ch.json" out="$("$PYTHON" "$S/make-chapters.py" --chapters "$SB/ch.json" --duration 200 2>/dev/null)"; rc=$? expect_exit "ffmetadata from explicit JSON -> 0" 0 "$rc" expect_has "ffmetadata header" ";FFMETADATA1" "$out" expect_has "ffmetadata escapes '=' in title" 'Topic \= One' "$out" out="$("$PYTHON" "$S/make-chapters.py" --chapters "$SB/ch.json" --duration 200 --format youtube 2>/dev/null)" expect_has "youtube format starts at 0:00" "0:00 Intro" "$out" out="$("$PYTHON" "$S/make-chapters.py" --chapters "$SB/ch.json" --duration 200 --format vtt 2>/dev/null)" expect_has "vtt format header" "WEBVTT" "$out" # ── structural: EDL dry-run (no ffmpeg required) ───────────────────────────── echo "-- cut-from-edl dry-run --" printf '{"scenes":[{"scene":1,"selection_rationale":"test","clips":[{"file":"takes/a.mp4","start":1.5,"end":4.0}]}]}' > "$SB/edit.json" out="$("$PYTHON" "$S/cut-from-edl.py" "$SB/edit.json" 2>/dev/null)"; rc=$? expect_exit "dry-run with absent sources -> 0" 0 "$rc" expect_has "dry-run prints ffmpeg commands" "ffmpeg" "$out" expect_has "dry-run includes concat step" "concat" "$out" # ── structural: pure-python LUT generation ─────────────────────────────────── echo "-- gen-luts --" out="$("$PYTHON" "$S/gen-luts.py" --variants warm_filmic --size 17 --out-dir "$SB/luts" 2>/dev/null)"; rc=$? expect_exit "gen-luts size 17 -> 0" 0 "$rc" [[ -f "$SB/luts/warm_filmic.cube" ]] && ok "cube file written" || no "cube file written" # Captured, not `head | grep -q`: under `set -o pipefail` grep -q's early exit # SIGPIPEs the producer (141) and flakes the assert even on a match. cube_hdr="$(head -5 "$SB/luts/warm_filmic.cube")" expect_has "cube header size" "LUT_3D_SIZE 17" "$cube_hdr" rows="$(grep -cE '^[0-9]' "$SB/luts/warm_filmic.cube")" [[ "$rows" == "4913" ]] && ok "cube row count 17^3" || no "cube row count (want 4913 got $rows)" out="$("$PYTHON" "$S/gen-luts.py" --variants neutral709 --size 17 --out-dir "$SB/luts" --json 2>/dev/null)" expect_has "gen-luts --json envelope" '"schema": "claude-mods.ffmpeg-ops.luts/v1"' "$out" "$PYTHON" "$S/gen-luts.py" --variants noir_bw,pastel,golden_hour,sepia,technicolor2,matrix_green --size 17 --out-dir "$SB/luts" >/dev/null 2>&1 expect_exit "gen-luts look-recipe variants -> 0" 0 $? # noir_bw is sat=0: every lattice row must be greyscale (R==G==B) nongrey="$(grep -E '^[0-9]' "$SB/luts/noir_bw.cube" | awk '{if ($1!=$2 || $2!=$3) c++} END{print c+0}')" [[ "$nongrey" == "0" ]] && ok "noir_bw LUT is true greyscale" || no "noir_bw LUT greyscale ($nongrey colored rows)" # sepia channel-mix: mid-grey input (grid 8,8,8 of 17^3 = data row 2457) maps warm (R>G>B) grep -E '^[0-9]' "$SB/luts/sepia.cube" | awk 'NR==2457{ok=($1>$2 && $2>$3)} END{exit !ok}' \ && ok "sepia LUT maps mid-grey warm (R>G>B)" || no "sepia LUT mid-grey warmth" # duotone gradient map: cyanotype black input -> shadow color (B dominant), white -> highlight (near-white) "$PYTHON" "$S/gen-luts.py" --variants duo_cyanotype,duo_synthwave --size 17 --out-dir "$SB/luts" >/dev/null 2>&1 expect_exit "gen-luts duotone variants -> 0" 0 $? grep -E '^[0-9]' "$SB/luts/duo_cyanotype.cube" | awk 'NR==1{ok=($3>$1)} END{exit !ok}' \ && ok "cyanotype LUT black -> blue shadow" || no "cyanotype LUT black -> blue shadow" grep -E '^[0-9]' "$SB/luts/duo_cyanotype.cube" | awk 'NR==4913{ok=($1>0.85 && $2>0.9 && $3>0.95)} END{exit !ok}' \ && ok "cyanotype LUT white -> paper highlight" || no "cyanotype LUT white -> paper highlight" # tritone split: cool shadows (B>R at black) AND warm highlights (R>B at white) "$PYTHON" "$S/gen-luts.py" --variants tri_split_classic --size 17 --out-dir "$SB/luts" >/dev/null 2>&1 expect_exit "gen-luts tritone variant -> 0" 0 $? grep -E '^[0-9]' "$SB/luts/tri_split_classic.cube" | awk 'NR==1{ok=($3>$1)} END{exit !ok}' \ && ok "tritone split: cool shadows" || no "tritone split: cool shadows" grep -E '^[0-9]' "$SB/luts/tri_split_classic.cube" | awk 'NR==4913{ok=($1>$3)} END{exit !ok}' \ && ok "tritone split: warm highlights" || no "tritone split: warm highlights" # ── structural: offline staleness verifier + assets ───────────────────────── echo "-- verify-commands --offline / assets --" bash "$S/verify-commands.sh" --offline >/dev/null 2>&1; expect_exit "verifier --offline clean" 0 $? for a in "$SKILL"/assets/*.json; do "$PYTHON" -c "import json,sys; json.load(open(sys.argv[1], encoding='utf-8'))" "$a" \ >/dev/null 2>&1 && ok "asset parses: $(basename "$a")" || no "asset parses: $(basename "$a")" done # ── media round-trips (only with ffmpeg on PATH) ───────────────────────────── if ! command -v ffmpeg >/dev/null 2>&1 || ! command -v ffprobe >/dev/null 2>&1; then echo "" echo " SKIP ffmpeg/ffprobe not on PATH — media round-trip tests NOT run." echo " (structural suite above still gates; install ffmpeg for full coverage)" else echo "-- media round-trips (lavfi fixtures) --" FIX="$SB/fixture.mp4" ffmpeg -v error -y -f lavfi -i testsrc2=duration=2:size=320x180:rate=30 \ -f lavfi -i "sine=frequency=440:duration=2" \ -c:v libx264 -pix_fmt yuv420p -c:a aac -shortest "$FIX" 2>/dev/null [[ -f "$FIX" ]] && ok "fixture synthesized" || no "fixture synthesized" out="$("$PYTHON" "$S/probe-media.py" "$FIX" 2>/dev/null)"; rc=$? expect_exit "probe fixture -> 0" 0 "$rc" expect_has "probe reports video" "h264 320x180" "$out" out="$("$PYTHON" "$S/probe-media.py" --json "$FIX" 2>/dev/null)" expect_has "probe --json envelope" '"schema": "claude-mods.ffmpeg-ops.probe/v1"' "$out" "$PYTHON" "$S/probe-media.py" --keyframes-near 1.0 "$FIX" >/dev/null 2>&1 expect_exit "probe --keyframes-near -> 0" 0 $? printf 'plain text' > "$SB/not-media.mp4" "$PYTHON" "$S/probe-media.py" "$SB/not-media.mp4" >/dev/null 2>&1 expect_exit "probe non-media -> 4" 4 $? # tone (1s) + silence (1s): silence and speech segments both detectable WAV="$SB/tone-silence.wav" ffmpeg -v error -y -f lavfi -i "sine=frequency=440:duration=1" \ -af "apad=pad_dur=1" -c:a pcm_s16le "$WAV" 2>/dev/null out="$("$PYTHON" "$S/detect-segments.py" --silence --min-silence 0.4 "$WAV" 2>/dev/null)"; rc=$? expect_exit "detect-segments --silence -> 0" 0 "$rc" expect_has "finds the silence" "silence" "$out" expect_has "derives speech segment" "speech" "$out" "$PYTHON" "$S/detect-segments.py" --scenes "$FIX" >/dev/null 2>&1 expect_exit "detect-segments --scenes -> 0" 0 $? out="$("$PYTHON" "$S/loudnorm-scan.py" "$FIX" --json 2>/dev/null)"; rc=$? expect_exit "loudnorm-scan -> 0" 0 "$rc" expect_has "emits pass-2 filter" "measured_I" "$out" "$PYTHON" "$S/quality-compare.py" "$FIX" "$FIX" --metrics ssim >/dev/null 2>&1 expect_exit "quality self-compare -> 0" 0 $? out="$("$PYTHON" "$S/quality-compare.py" "$FIX" "$FIX" --metrics ssim --json 2>/dev/null)" expect_has "ssim of identical ~1" '"all": 1' "$out" # Captured, not `-filters | grep -q`: under pipefail a SIGPIPE'd ffmpeg (141) # would silently take the no-vmaf branch even when libvmaf is present. filter_list="$(ffmpeg -hide_banner -filters 2>/dev/null)" if grep -q libvmaf <<<"$filter_list"; then "$PYTHON" "$S/quality-compare.py" "$FIX" "$FIX" --metrics vmaf --min-vmaf 95 >/dev/null 2>&1 expect_exit "vmaf self-compare above threshold -> 0" 0 $? else echo " SKIP vmaf (libvmaf not in this build)" fi printf '{"scenes":[{"scene":1,"clips":[{"file":"%s","start":0.2,"end":1.0},{"file":"%s","start":1.2,"end":1.8}]}]}' \ "$(basename "$FIX")" "$(basename "$FIX")" > "$SB/cutme.json" "$PYTHON" "$S/cut-from-edl.py" "$SB/cutme.json" --execute -o "$SB/final.mp4" >/dev/null 2>&1 rc=$? expect_exit "cut-from-edl --execute -> 0" 0 "$rc" [[ -f "$SB/final.mp4" ]] && ok "EDL final output exists" || no "EDL final output exists" dur="$(ffprobe -v error -show_entries format=duration -of default=nw=1:nk=1 "$SB/final.mp4" 2>/dev/null)" "$PYTHON" -c "import sys; d=float(sys.argv[1]); sys.exit(0 if 1.0 < d < 1.9 else 1)" "${dur:-0}" \ && ok "EDL output duration ~1.4s (got ${dur}s)" || no "EDL output duration (got ${dur}s)" # regression (live E2E find): -o resolves against the CWD and the output dir # is created BEFORE ffmpeg opens the temp file (was: mkdir after concat -> # cryptic "Error opening output files" for any -o into a new directory) ( cd "$SB" && "$PYTHON" "$S/cut-from-edl.py" cutme.json --execute -o "newdir/out2.mp4" >/dev/null 2>&1 ) expect_exit "cut-from-edl -o cwd-relative into new dir -> 0" 0 $? [[ -f "$SB/newdir/out2.mp4" ]] && ok "-o resolved vs CWD, dest dir auto-created" \ || no "-o resolved vs CWD, dest dir auto-created" "$PYTHON" "$S/make-chapters.py" --from-silence --media "$FIX" --min-gap 0.2 \ --write "$SB/chaptered.mp4" >/dev/null 2>&1 expect_exit "make-chapters --write -> 0" 0 $? nch="$(ffprobe -v error -show_entries chapter=start_time -of csv=p=0 "$SB/chaptered.mp4" 2>/dev/null | grep -c .)" [[ "${nch:-0}" -ge 1 ]] && ok "muxed file has chapters ($nch)" || no "muxed file has chapters (got ${nch:-0})" out="$("$PYTHON" "$S/gen-luts.py" --variants neutral709 --size 17 --out-dir "$SB/lutsprev" \ --previews "$FIX" --frame-at 0.5 2>/dev/null)"; rc=$? expect_exit "gen-luts --previews -> 0" 0 "$rc" [[ -f "$SB/lutsprev/preview_neutral709.png" ]] && ok "preview still rendered" || no "preview still rendered" [[ -f "$SB/lutsprev/index.html" ]] && ok "chooser index.html written" || no "chooser index.html written" # --doctor: synthesized fixture has moov AFTER mdat (no faststart) -> finding out="$("$PYTHON" "$S/probe-media.py" --doctor "$FIX" 2>/dev/null)"; rc=$? expect_exit "doctor flags non-faststart fixture -> 10" 10 "$rc" expect_has "doctor names the moov issue" "faststart" "$out" ffmpeg -v error -y -i "$FIX" -c copy -movflags +faststart "$SB/fast.mp4" 2>/dev/null "$PYTHON" "$S/probe-media.py" --doctor "$SB/fast.mp4" >/dev/null 2>&1 expect_exit "doctor clean after faststart remux -> 0" 0 $? "$PYTHON" "$S/probe-media.py" --doctor "$WAV" >/dev/null 2>&1 expect_exit "doctor on audio-only -> 0 (info only)" 0 $? "$PYTHON" "$S/smart-compress.py" --target 150KB --preset fast \ -o "$SB/small.mp4" "$SB/fast.mp4" >/dev/null 2>&1 expect_exit "smart-compress -> 0" 0 $? sz="$(wc -c < "$SB/small.mp4" 2>/dev/null | tr -d ' ')" [[ "${sz:-999999}" -le 150000 ]] && ok "compressed under target ($sz <= 150000)" \ || no "compressed under target (got ${sz:-missing})" "$PYTHON" "$S/make-sprites.py" --interval 0.5 --width 64 --cols 2 --rows 2 \ --out-dir "$SB/sprites" "$FIX" >/dev/null 2>&1 expect_exit "make-sprites -> 0" 0 $? [[ -f "$SB/sprites/sprite_01.jpg" ]] && ok "sprite sheet written" || no "sprite sheet written" grep -q "xywh=64,0,64" "$SB/sprites/thumbs.vtt" 2>/dev/null \ && ok "vtt has correct xywh geometry" || no "vtt has correct xywh geometry" bash "$S/capability-scan.sh" --quick >/dev/null 2>&1 expect_exit "capability-scan --quick -> 0" 0 $? bash "$S/verify-commands.sh" --live >/dev/null 2>&1 rc=$?; [[ "$rc" == 0 ]] && ok "verifier --live clean against installed build" \ || no "verifier --live (got $rc — investigate drift findings)" fi echo "" echo "=== $PASS passed, $FAIL failed ===" [[ "$FAIL" -eq 0 ]] || exit 1 exit 0
-
-
SKILL.md 27.3 KB
--- name: ffmpeg-ops description: "Comprehensive ffmpeg/ffprobe media processing: transcode, cut/trim/concat, color grading, loudness normalization, subtitles, GIFs, HLS packaging, hardware encoding, and quality gates (VMAF). Triggers on: ffmpeg, ffprobe, transcode, compress/convert video, extract audio, color grade, hls." license: MIT compatibility: "ffmpeg 5.0+ (6.0+ recommended). Scripts: bash + python3.10+. Optional per task: libvmaf, libass, libzimg, libvidstab." allowed-tools: "Read Write Edit Bash Glob Grep" metadata: author: claude-mods related-skills: color-ops, debug-ops --- # ffmpeg Operations Operational expertise for ffmpeg/ffprobe: the ~30 commands that cover most real work, the footguns that silently ruin output, EDL-driven editing (edit-as-code), and eight scripts that replace the logic an agent would otherwise re-derive every task. ## Doctrine: probe first **Never transcode, cut, or filter blind.** Every media task starts by probing the input — codec, duration, frame rate (constant or variable?), pixel format, rotation, stream layout. Half of all "ffmpeg did something weird" reports are a property of the *input* the command never checked. ```bash python skills/ffmpeg-ops/scripts/probe-media.py input.mp4 # human summary python skills/ffmpeg-ops/scripts/probe-media.py --doctor input.mp4 # TRIAGE: hazards + exact fixes python skills/ffmpeg-ops/scripts/probe-media.py --json input.mp4 | jq '.data.streams' python skills/ffmpeg-ops/scripts/probe-media.py --keyframes-near 92.5 input.mp4 ``` `--doctor` makes the doctrine self-enforcing: VFR, HDR transfer, rotation metadata, interlacing, non-yuv420p delivery, and moov-at-EOF each come back as a finding **with the exact fix command**, and exit 10 means "fix before processing". The `--keyframes-near` form answers "can I stream-copy a cut at 92.5s?" — it reports the nearest keyframes so you know whether a copy cut will snap (see Footguns). When a command fails with a cryptic message, decode it: [references/error-decoder.md](references/error-decoder.md). **Before recommending an encoder, verify the build has it.** Installed ffmpeg builds vary wildly (especially hardware encoders — *listed* ≠ *working*): ```bash bash skills/ffmpeg-ops/scripts/capability-scan.sh # full: proof-encodes each hw encoder bash skills/ffmpeg-ops/scripts/capability-scan.sh --quick # list-only, no GPU touch ``` ## Cookbook Commands are bash-form; they run unchanged in PowerShell except where the [Windows notes](#windows-notes) say otherwise. Replace `-y`/`-n` (overwrite/never) consciously — never leave an agent-run command interactive. ### Convert and compress ```bash # Web-compatible H.264 — THE default delivery encode. yuv420p + faststart are not # optional: without them Safari/QuickTime/old devices show black video, and the # moov atom sits at EOF so browsers can't start playback until fully downloaded. ffmpeg -i in.mov -c:v libx264 -crf 20 -preset slow -pix_fmt yuv420p \ -c:a aac -b:a 192k -movflags +faststart out.mp4 # H.265/HEVC — ~40% smaller at same quality, slower encode, less universal playback. # -tag:v hvc1 is required for Apple players to recognize the stream. ffmpeg -i in.mp4 -c:v libx265 -crf 24 -preset slow -tag:v hvc1 \ -c:a copy -movflags +faststart out.mp4 # AV1 via SVT-AV1 (libaom is 10-50x slower; only use it for research-grade encodes). # preset 0-13: lower = slower/better; 6 is the quality/speed sweet spot. ffmpeg -i in.mp4 -c:v libsvtav1 -crf 32 -preset 6 -c:a libopus -b:a 128k out.webm # Remux only — change container, zero quality loss, near-instant. Try this FIRST # when the ask is "make this .mkv play in X": often the codecs are fine. ffmpeg -i in.mkv -c copy -movflags +faststart out.mp4 # Normalize a problem source (HEVC/VFR phone footage, Zoom/Loom exports) before ANY # downstream editing. VFR breaks cut math, concat sync, and Remotion/player seeking. ffmpeg -i in.mov -c:v libx264 -crf 18 -preset fast -pix_fmt yuv420p \ -fps_mode cfr -r 30 -c:a aac -b:a 192k normalized.mp4 # Archival master — FFV1 lossless in MKV (the preservation standard). ffmpeg -i in.mp4 -c:v ffv1 -level 3 -g 1 -slicecrc 1 -c:a flac archive.mkv # "Make it fit in 25MB" — computed two-pass bitrate, auto audio/downscale, VERIFIED: python skills/ffmpeg-ops/scripts/smart-compress.py --target 25MB video.mp4 ``` Codec choice, CRF/preset matrices, two-pass bitrate targeting, per-platform social targets: [references/encoding.md](references/encoding.md) + [assets/encoding-presets.json](assets/encoding-presets.json). ### Cut and join ```bash # Fast lossless trim (stream copy). -ss/-to BEFORE -i = input seek, absolute times. # CAVEAT: with -c copy the start snaps to the previous keyframe — can be seconds # early, or give frozen/black lead-in. Check first with probe-media.py --keyframes-near. ffmpeg -ss 00:01:30 -to 00:02:00 -i in.mp4 -c copy -avoid_negative_ts make_zero cut.mp4 # Frame-accurate trim (re-encode). Input-side -ss IS frame-accurate when re-encoding # (ffmpeg decodes from the prior keyframe and discards) — fast AND exact. The old # "put -ss after -i for accuracy" advice costs a full decode from 0:00 for nothing. ffmpeg -ss 00:01:30 -to 00:02:00 -i in.mp4 -c:v libx264 -crf 18 -c:a aac cut.mp4 # Join files with IDENTICAL codec/params — concat demuxer, no re-encode. printf "file '%s'\n" seg1.mp4 seg2.mp4 seg3.mp4 > concat.txt ffmpeg -f concat -safe 0 -i concat.txt -c copy joined.mp4 # Join files with DIFFERENT codecs/sizes — concat filter, re-encodes. ffmpeg -i a.mp4 -i b.mov -filter_complex \ "[0:v][0:a][1:v][1:a]concat=n=2:v=1:a=1[v][a]" \ -map "[v]" -map "[a]" -c:v libx264 -crf 20 -c:a aac joined.mp4 # Remove a middle segment (keep 0-60s and 120s-end): cut both keeps, then concat. # For multi-cut edits, write an EDL and use cut-from-edl.py instead (see EDL workflow). ``` `-ss` semantics in full, keyframe theory, concat ×3 (demuxer/filter/protocol), edit-decision-list editing: [references/trim-concat.md](references/trim-concat.md) and [references/edit-as-code.md](references/edit-as-code.md). ### Resize, transform, retime ```bash # Resize to width, keep aspect. ALWAYS -2 (not -1): yuv420p needs even dimensions. ffmpeg -i in.mp4 -vf "scale=1280:-2" -c:a copy out.mp4 # Crop (w:h:x:y from top-left); cropdetect finds black bars for you: ffmpeg -i in.mp4 -vf cropdetect -frames:v 120 -f null - 2>&1 | rg crop= ffmpeg -i in.mp4 -vf "crop=1920:800:0:140" -c:a copy out.mp4 # Vertical 9:16 from landscape — blurred-pad pattern (social standard): ffmpeg -i in.mp4 -filter_complex \ "[0:v]scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,boxblur=20[bg]; [0:v]scale=1080:-2[fg];[bg][fg]overlay=(W-w)/2:(H-h)/2" -c:a copy vertical.mp4 # Rotate: fix metadata only (instant) vs bake pixels (re-encode). ffmpeg -display_rotation 90 -i in.mp4 -c copy out.mp4 # metadata flip (ffmpeg 6+) ffmpeg -i in.mp4 -vf "transpose=1" -c:a copy out.mp4 # transpose=1: 90° clockwise # Frame-rate change (drops/dups frames; for smooth slow-mo see minterpolate below) ffmpeg -i in.mp4 -vf "fps=30" -c:a copy out.mp4 # 2x speed-up: video PTS halved + audio atempo (atempo accepts 0.5-100; chain # atempo=0.5,atempo=0.5 for 0.25x). -map ordering keeps streams paired. ffmpeg -i in.mp4 -filter_complex \ "[0:v]setpts=0.5*PTS[v];[0:a]atempo=2.0[a]" -map "[v]" -map "[a]" fast.mp4 # Interpolated slow-mo (synthesizes in-between frames — slow but smooth): ffmpeg -i in.mp4 -vf "minterpolate=fps=60:mi_mode=mci:mc_mode=aobmc,setpts=2*PTS" -an slow.mp4 # Timelapse from photos (and the reverse: video -> frames, under Images below) ffmpeg -framerate 24 -pattern_type glob -i 'photos/*.jpg' \ -c:v libx264 -crf 20 -pix_fmt yuv420p timelapse.mp4 ``` Filtergraph syntax (labels, chains, split), speed ramps, full filter cookbook: [references/filtergraph.md](references/filtergraph.md). ### Overlay, text, subtitles ```bash # Watermark bottom-right with 24px margin (W/H = video, w/h = overlay dims): ffmpeg -i in.mp4 -i logo.png -filter_complex \ "overlay=W-w-24:H-h-24:format=auto" -c:a copy out.mp4 # Burn a running timecode (note %{pts\:hms} — the colon must be escaped INSIDE # the drawtext argument; see Windows notes for fontfile paths): ffmpeg -i in.mp4 -vf \ "drawtext=text='%{pts\:hms}':fontsize=48:fontcolor=white:box=1:boxcolor=black@0.5:x=24:y=24" \ -c:a copy out.mp4 # Burn-in subtitles (hard subs; needs libass). Pragmatic path rule: cd to the # subtitle's directory and use a bare relative filename — the filter's path # escaping is the single worst quoting trap in ffmpeg, especially on Windows. ffmpeg -i in.mp4 -vf "subtitles=subs.srt" -c:a copy burned.mp4 # Soft subtitles (toggleable, instant — no re-encode): ffmpeg -i in.mp4 -i subs.srt -map 0 -map 1 -c copy -c:s mov_text soft.mp4 # mp4 ffmpeg -i in.mkv -i subs.srt -map 0 -map 1 -c copy -c:s srt soft.mkv # mkv ``` Styling (ASS force_style), extraction, format conversion, STT round-trip: [references/subtitles.md](references/subtitles.md). ### Audio ```bash # Extract audio without re-encoding (copy the stream as-is; pick the container # matching the codec — probe first: aac->.m4a, opus->.opus/.ogg, mp3->.mp3): ffmpeg -i in.mp4 -vn -c:a copy out.m4a # Extract + transcode to Opus (best codec per bit: voice 24-32k mono, music 96-128k): ffmpeg -i in.mp4 -vn -c:a libopus -b:a 128k out.opus # Replace a video's audio track (keep video untouched): ffmpeg -i video.mp4 -i music.m4a -map 0:v -map 1:a -c:v copy -c:a aac -shortest out.mp4 # Mix two audio inputs (normalize=0 stops amix halving the volume of each input): ffmpeg -i voice.wav -i music.mp3 -filter_complex \ "[1:a]volume=0.25[m];[0:a][m]amix=inputs=2:duration=first:normalize=0[a]" \ -map "[a]" -c:a aac mixed.m4a # Loudness-normalize, one-pass (quick; DYNAMIC mode — fine for drafts). # Two-pass linear mode is measurably better: use loudnorm-scan.py (Scripts below). # loudnorm internally upsamples to 192kHz — the -ar 48000 puts it back. ffmpeg -i in.mp4 -af "loudnorm=I=-16:TP=-1.5:LRA=11" -ar 48000 -c:v copy out.mp4 # Trim leading/trailing silence: ffmpeg -i in.wav -af \ "silenceremove=start_periods=1:start_threshold=-40dB:detection=peak,areverse,silenceremove=start_periods=1:start_threshold=-40dB:detection=peak,areverse" \ trimmed.wav ``` Targets: -14 LUFS streaming platforms, -16 podcasts, -23 EBU R128 broadcast. Channel mapping, multi-track, restoration filters: [references/audio.md](references/audio.md). ### Speech-to-text prep (Whisper-family) ```bash # THE canonical STT extraction — 16 kHz mono 16-bit PCM (what whisper.cpp / # faster-whisper actually resample to; doing it here is faster and deterministic): ffmpeg -i in.mp4 -vn -ac 1 -ar 16000 -c:a pcm_s16le stt.wav # Pipe raw PCM straight to whisper.cpp — no temp file: ffmpeg -v error -i in.mp4 -vn -ac 1 -ar 16000 -f s16le - | whisper-cli -m model.bin -f - # Chunk long audio ON SILENCE BOUNDARIES (never mid-word) for parallel transcription: python skills/ffmpeg-ops/scripts/detect-segments.py --silence --json in.mp4 \ | jq '.data.speech[]' ``` Pre-STT cleanup (when `afftdn`/`highpass` help vs hurt accuracy), WhisperX word-level alignment (±50 ms), transcript JSON shape, the summarisation pipeline: [references/stt-whisper.md](references/stt-whisper.md). ### Images, GIFs, frames ```bash # Thumbnail at a timestamp (input-side -ss: instant even at 2h offsets): ffmpeg -ss 00:00:05 -i in.mp4 -frames:v 1 -q:v 2 thumb.jpg # Contact sheet: 1 frame every 10s, tiled 4x3 (visual summary / scrub preview): ffmpeg -i in.mp4 -vf "fps=1/10,scale=320:-2,tile=4x3" -frames:v 1 sheet.png # High-quality GIF — palettegen/paletteuse is THE difference between a 256-color # dithered mess and a clean GIF. Single pass via split: ffmpeg -ss 5 -to 8 -i in.mp4 -filter_complex \ "fps=12,scale=480:-1:flags=lanczos,split[s0][s1];[s0]palettegen=max_colors=128[p];[s1][p]paletteuse=dither=bayer:bayer_scale=4" \ out.gif # Embedded chapters from scene/silence detection (or YouTube description text): python skills/ffmpeg-ops/scripts/make-chapters.py --from-scenes --media talk.mp4 \ --min-gap 30 --write chaptered.mp4 python skills/ffmpeg-ops/scripts/make-chapters.py --from-silence --media lecture.mp4 \ --format youtube # Frames for ML datasets — fixed fps, model-square crop: ffmpeg -i in.mp4 -vf "fps=1,scale=512:512:force_original_aspect_ratio=increase,crop=512:512" \ frames/%06d.png # Image sequence -> video: ffmpeg -framerate 24 -i frames/%06d.png -c:v libx264 -crf 18 -pix_fmt yuv420p out.mp4 # Player scrub-preview sprites + the WebVTT thumbnail track that maps them: python skills/ffmpeg-ops/scripts/make-sprites.py --interval 5 video.mp4 ``` Sprite sheets for web players, AVIF/WebP stills, dataset prep patterns: [references/images-gif.md](references/images-gif.md). ### Diagnostics and validation ```bash # Corruption / decode-error check (exit code is NOT the signal — the log is): ffmpeg -v error -i in.mp4 -f null - 2> errors.log && [ ! -s errors.log ] && echo CLEAN # Per-frame hashes — prove two pipelines produce identical frames: ffmpeg -i in.mp4 -map 0:v -f framemd5 - # Strip ALL metadata (GPS, device info — privacy before sharing phone video). # -map_metadata -1 keeps rotation side-data; verify orientation after. ffmpeg -i in.mp4 -map_metadata -1 -c copy clean.mp4 # Quick probes (machine-readable; prefer probe-media.py for the full picture): ffprobe -v error -show_entries format=duration -of default=nw=1:nk=1 in.mp4 ffprobe -v error -select_streams v:0 -show_entries stream=codec_name,width,height,r_frame_rate -of csv=p=0 in.mp4 ``` Safe re-encode of untrusted uploads, scene-change detection, integrity in CI: [references/analysis-validation.md](references/analysis-validation.md). ### yt-dlp interop yt-dlp embeds ffmpeg for merge/remux; these are the post-download patterns: ```bash # Prefer h264+m4a at download time (avoids a transcode entirely): yt-dlp -S "res:1080,vcodec:h264,acodec:m4a" --remux-video mp4 URL # Clip a section AT download (server-side range requests; much faster than full DL): yt-dlp --download-sections "*10:00-12:30" -S "res:1080,vcodec:h264" URL # Audio-only for STT/summarisation: yt-dlp -x --audio-format opus URL # Already downloaded a VP9/AV1 .webm that needs to be H.264 .mp4: that is a normal # transcode — use the web-compatible H.264 recipe above, NOT --recode-video. ``` ### Generative/test sources ```bash # Synthetic video+audio — fixtures, pipeline tests, alignment checks (no real media # needed; this is how tests/run.sh builds its fixtures): ffmpeg -f lavfi -i testsrc2=duration=2:size=640x360:rate=30 \ -f lavfi -i "sine=frequency=440:duration=2" \ -c:v libx264 -pix_fmt yuv420p -c:a aac fixture.mp4 ``` Audio-reactive visuals (showwaves/showspectrum), podcast audiograms: [references/visualization.md](references/visualization.md). ## Footguns The table that pays this skill's rent. Each row is a class of silent failure. | Footgun | The trap | The rule | |---|---|---| | `-ss` + `-c copy` | Cut starts seconds early or with frozen/black lead-in (snapped to prior keyframe) | Copy cuts snap. Check `probe-media.py --keyframes-near`; re-encode when exact | | Output-side `-to` after input-side `-ss` | Timestamps reset at the seek point, so `-to` silently becomes a *duration* | Keep `-ss`/`-to` on the same side of `-i` (both input-side is fast and absolute) | | Missing `-pix_fmt yuv420p` | Encode "works" but Safari/QuickTime/TVs show black or refuse to play (defaulted to yuv444p/yuv422p from a high-quality source) | Always set it for delivery H.264/H.265 | | Missing `-movflags +faststart` | Browser can't start playback until the whole file downloads (moov at EOF) | Always set it for web-served MP4 | | Default stream selection | ffmpeg picks ONE stream per type (highest-res video, most-channels audio) — extra audio tracks and all subs are silently dropped | `-map 0` to keep everything, explicit `-map` otherwise | | `-vf` + `-c:v copy` together | Hard error — filters require decoding | Filtering implies re-encode; pick one | | VFR source (phone/Zoom/Loom/screen-rec) | Cut math drifts, concat desyncs, players stutter | Normalize first: `-fps_mode cfr -r 30` + re-encode (cookbook) | | `-vsync` (deprecated) | Old flag, removed direction | Use `-fps_mode` (cfr/vfr/passthrough) | | `scale=W:-1` | Odd height → encoder error with yuv420p | Always `-2` | | concat demuxer on mismatched inputs | "Works" then glitches/desyncs at boundaries (codec/timebase mismatch) | Demuxer = identical params only; else concat *filter* with re-encode | | amix default | Each input's volume halved (normalize defaults on) | `amix=...:normalize=0` + explicit `volume=` | | One-pass loudnorm | Dynamic mode pumps quiet passages; output silently 192 kHz | Two-pass linear via `loudnorm-scan.py`; add `-ar 48000` | | `-shortest` absent on audio-replace | Output runs as long as the LONGEST input (silence or frozen frame tail) | Add `-shortest` when muxing separate A/V | | BT.601/709 colour shift | Slightly wrong colours after scaling SD↔HD (matrix guessed from resolution) | Tag explicitly when it matters: see [references/color-hdr.md](references/color-hdr.md) | | drawtext/subtitles path escaping | Filter args re-parse `:` and `\` — Windows paths like `C:\x` explode inside filter strings | cd to the asset's dir and use bare relative names; or escape as `C\:/path` | | Interactive overwrite prompt | Agent-run command hangs forever on "File exists. Overwrite? [y/N]" | Always pass `-y` or `-n` explicitly | | `%` in cmd.exe | `%06d` patterns and `%{pts}` get mangled by cmd variable expansion | Use PowerShell or bash; in .bat double to `%%` | ### Windows notes Platform-agnostic commands, but when running on Windows: - **PowerShell quoting is friendlier than bash here**: single quotes are fully literal, so `-vf 'scale=1280:-2,fps=30'` needs no escaping. Double quotes only interpolate `$` and backtick — filtergraphs rarely contain either. - **`NUL` not `/dev/null`** for two-pass logs: `-passlogfile` defaults are fine, but `ffmpeg ... -f null NUL` (PowerShell also accepts `-f null -`, which is portable — prefer it). - **Font paths in drawtext**: `fontfile='C\:/Windows/Fonts/arial.ttf'` — forward slashes, escaped drive colon, inside the filter string. - **Prefer `-f null -` and relative paths** to sidestep both quoting tables at once. ## Decision trees **Codec** — `H.264 (libx264)`: default; universal playback, fast, good per-bit at `-crf 18..23`. → `H.265 (libx265)`: same quality ~40% smaller; slower; needs `-tag:v hvc1` for Apple; fine for storage/modern devices. → `AV1 (libsvtav1)`: best compression, royalty-free, web-first (YouTube/Netflix path); encode cost highest; playback on older hardware is software-only. → `VP9`: only when a pipeline demands webm and AV1 is unavailable. → `FFV1`: archival masters only. **Cut method** — Need exact frames OR applying any filter → re-encode (input-side `-ss`, `-crf 18`). Cut points happen to sit on keyframes (verify with `--keyframes-near`) OR a ±2s slop is acceptable → stream copy with `-avoid_negative_ts make_zero`. Many cuts from one source → EDL workflow below. **CPU vs hardware encode** — Hardware (NVENC/QSV/AMF/VideoToolbox) is 5-20× faster but **worse quality per bit** than libx264/x265 at slow presets. Use hardware for: speed-critical batch work, live/streaming, drafts, "good enough" deliveries (bump bitrate ~30% to compensate). Use CPU for: final masters, size-constrained targets, quality comparisons. Always `capability-scan.sh` first — listed encoders fail at runtime on driver mismatches. Details: [references/hardware-accel.md](references/hardware-accel.md). ## EDL workflow (edit-as-code) For any multi-cut edit, do not fire ad-hoc trim commands. Write an **edit decision list** — a JSON file naming every clip, time range, and *why* — then cut from it. The edit becomes reviewable (rationale is written down), rerunnable (regenerate the output any time), and diffable (versions of the edit are git history). ```bash # 1. Find candidate cut points (silence = clean speech boundaries): python skills/ffmpeg-ops/scripts/detect-segments.py --silence --json take3.mp4 # 2. Author the EDL (schema: assets/edl-schema.json) with per-scene rationale. # 3. Dry-run prints every ffmpeg command it would run (default — nothing executes): python skills/ffmpeg-ops/scripts/cut-from-edl.py edit.json # 4. Execute: cuts + concat -> final. Re-encodes by default for frame accuracy; # --copy for keyframe-aligned EDLs. python skills/ffmpeg-ops/scripts/cut-from-edl.py edit.json --execute -o final.mp4 ``` Rules that make this work (from the Fable launch-video pipeline): cuts must land in **silence**; the model reasons over **transcripts, not frames**; after cutting, **re-transcribe the output to verify** (no filler words survived, no words clipped). Full architecture, EDL schema, verification loop: [references/edit-as-code.md](references/edit-as-code.md). ## Color grading ```bash # Apply a .cube LUT (tetrahedral = highest quality interpolation): ffmpeg -i in.mp4 -vf "lut3d=file=grade.cube:interp=tetrahedral" \ -c:v libx264 -crf 18 -c:a copy graded.mp4 # Generate a family of grade candidates + an HTML still-chooser: python skills/ffmpeg-ops/scripts/gen-luts.py --variants all --out-dir work/luts \ --previews in.mp4 ``` **The human picks the grade.** Generate variants, render preview stills, present a chooser — never auto-select a look. Grading is a taste call; the agent's job is the lattice math and the apply command. LUT format, log-footage normalization (S-Log3/V-Log → Rec.709), curves/eq safe ranges, checking work with ffmpeg's built-in scopes (waveform/vectorscope): [references/color-grading.md](references/color-grading.md). The 25-look recipe catalog — film stocks (Kodachrome, CineStill halation, Technicolor, Eterna), signature grades (Mad Max, Fincher, Matrix, BR2049, Amélie…), era/genre moods, Sin City selective color — every chain build-validated, plus the Hald-CLUT match-any-look workflow and scope-matching ladder: [references/look-recipes.md](references/look-recipes.md). Pipeline correctness (pix_fmt, HDR→SDR tonemapping, range/matrix tagging): [references/color-hdr.md](references/color-hdr.md). ## Quality gates ```bash # VMAF/SSIM/PSNR verdict on an encode (exit 10 = below threshold -> branch on it): python skills/ffmpeg-ops/scripts/quality-compare.py reference.mp4 encoded.mp4 \ --metrics ssim,psnr python skills/ffmpeg-ops/scripts/quality-compare.py reference.mp4 encoded.mp4 \ --metrics vmaf --min-vmaf 90 --json | jq '.data.vmaf' ``` VMAF ≥ 93 at 1080p ≈ visually transparent; 80-93 = noticeable on inspection. Side-by-side visual A/B (`hstack`), metric interpretation, encode-ladder tuning: [references/quality-metrics.md](references/quality-metrics.md). ## Scripts All eleven follow the [Skill Resource Protocol](../../docs/SKILL-RESOURCE-PROTOCOL.md): `--help` with examples, stdout = data only, `--json` envelopes (`claude-mods.ffmpeg-ops.*/v1`), semantic exit codes (`0` ok, `2` usage, `3` input missing, `4` invalid input, `5` missing dependency, `7` ffmpeg unavailable, `10` domain finding). | Script | Job | Worked invocation | |---|---|---| | `probe-media.py` | Normalized inspection, keyframe proximity, `--doctor` triage (hazard → fix command, exit 10) | `probe-media.py --doctor in.mp4` | | `capability-scan.sh` | What can THIS ffmpeg build do (proof-encodes hw encoders; `--quick` skips) | `capability-scan.sh --json \| jq '.data.encoders'` — exit 10 = a listed encoder failed verification | | `quality-compare.py` | VMAF/SSIM/PSNR gate | `quality-compare.py ref.mp4 enc.mp4 --min-vmaf 90` — exit 10 = below threshold | | `loudnorm-scan.py` | Two-pass loudnorm: measures pass 1, emits exact pass-2 filter | `loudnorm-scan.py -I -16 in.mp4 --json \| jq -r '.data.pass2_filter'` | | `detect-segments.py` | Silence/scene boundaries as JSON segments (STT chunking, dead-air cuts, shot splits) | `detect-segments.py --scenes --json in.mp4 \| jq '.data.segments'` | | `cut-from-edl.py` | EDL JSON → validated cuts + concat (dry-run by default) | `cut-from-edl.py edit.json --execute -o final.mp4` | | `make-chapters.py` | Scene/silence points (or explicit JSON) → embedded chapters / YouTube text / WebVTT | `make-chapters.py --from-scenes --media talk.mp4 --write chaptered.mp4` | | `smart-compress.py` | Fit a size cap: computed two-pass bitrate, auto audio/downscale, size-verified (exit 10 = still over) | `smart-compress.py --target 25MB video.mp4` | | `make-sprites.py` | Scrub-preview sprite sheets + WebVTT thumbnail track (#xywh) | `make-sprites.py --interval 5 video.mp4` | | `gen-luts.py` | Emit .cube grade variants (+ `--previews` still chooser) | `gen-luts.py --variants warm_filmic,punchy --out-dir luts/` | | `verify-commands.sh` | Staleness verifier: `--offline` structural (CI), `--live` checks docs against the installed build | `verify-commands.sh --live` — exit 10 = doc drift, 7 = no ffmpeg | ## References Load on demand — one concept per file: | Reference | Load when | |---|---| | [encoding.md](references/encoding.md) | Choosing codec/CRF/preset, two-pass, social platform targets, archival | | [hardware-accel.md](references/hardware-accel.md) | NVENC/QSV/AMF/VideoToolbox/VAAPI flags, quality caveats, detection | | [filtergraph.md](references/filtergraph.md) | Any `-filter_complex`, labels/chains/split, speed ramps, xstack | | [trim-concat.md](references/trim-concat.md) | Cut accuracy, keyframes, concat selection, segment removal | | [edit-as-code.md](references/edit-as-code.md) | Multi-cut edits, EDL schema, transcript-driven editing, verify loop | | [audio.md](references/audio.md) | Loudness, mixing, channel layout, audio repair | | [stt-whisper.md](references/stt-whisper.md) | Whisper/WhisperX prep, chunking, transcript JSON, summarisation pipeline | | [subtitles.md](references/subtitles.md) | Burn vs soft, styling, extraction, format conversion | | [color-grading.md](references/color-grading.md) | LUTs, .cube format, log normalization, scopes, grade workflow | | [look-recipes.md](references/look-recipes.md) | 25-look catalog (film stocks, signature movie grades, era/genre moods), Hald-CLUT extraction, scope-matching | | [color-hdr.md](references/color-hdr.md) | pix_fmt, HDR→SDR tonemap, BT.601/709 tagging, 10-bit | | [quality-metrics.md](references/quality-metrics.md) | VMAF/SSIM interpretation, visual A/B, ladder tuning | | [streaming-hls.md](references/streaming-hls.md) | HLS/DASH packaging, ABR ladders, live restream | | [images-gif.md](references/images-gif.md) | GIF quality, sprite sheets, dataset frame extraction | | [restoration.md](references/restoration.md) | Deinterlace, denoise, deband, stabilize, audio cleanup | | [analysis-validation.md](references/analysis-validation.md) | Corruption checks, hashing, metadata stripping, untrusted uploads | | [capture-devices.md](references/capture-devices.md) | Screen/webcam capture per OS (gdigrab/dshow, avfoundation, x11grab) | | [error-decoder.md](references/error-decoder.md) | An ffmpeg command failed with a cryptic message — symptom → cause → fix | | [visualization.md](references/visualization.md) | Waveform/spectrogram videos, audiograms, comparison grids | Assets: [encoding-presets.json](assets/encoding-presets.json) (recipe data incl. date-stamped social targets), [hls-ladder.json](assets/hls-ladder.json) (ABR ladder), [edl-schema.json](assets/edl-schema.json) (the cut-from-edl.py contract). ## Self-test ```bash bash skills/ffmpeg-ops/tests/run.sh # offline suite; synthesizes fixtures via lavfi ``` Structural assertions always run; media round-trips run only when ffmpeg is on PATH (loud skip otherwise — never a silent false-clean).
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.