ai-media
Use when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with ffmpeg (mux, duck, loudnorm, concat). NOT still-image generation/editing (that is `replicate-images`);
Install
npx skills add https://github.com/ericrisco/rsc-harness/tree/main/skills/ai-media
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install ericrisco-rsc-harness@llmmart
git clone https://github.com/ericrisco/rsc-harness.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole ericrisco/rsc-harness collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
ai-media
You are the cross-modal director. You decide which generative-media model to call per modality, in what order, with what params, then assemble the pieces with ffmpeg into one finished file. You do not own a single provider's API surface and you do not prompt still images — you orchestrate and glue.
Pipeline shape — decide what the goal needs
Map the goal to modalities and an ordered step list, and lock that plan before you generate a single asset — media generation is slow and metered, so a re-roll of a 10 s Veo clip or a 90 s music track costs real money and minutes. Fixing the scene list, aspect ratio, target loudness and model per modality first is cheaper than discovering at mux time that your clips are 9:16 and your VO is the wrong sample rate. The "delegate to" column is where the actual call mechanics live — you pick the model and params, those skills run the call.
| Goal | Needs | Ordered steps | Delegate calls to |
|---|---|---|---|
| Narrated explainer | stills + img→video + VO + music | script → per-scene stills → clip per scene → VO → music → conform → concat → mix+duck → loudnorm → MP4 | replicate-images, fal/replicate |
| Product teaser (1 hero) | 1 still + img→video + music | still → clip → music → mix → loudnorm → MP4 | replicate-images, fal/replicate |
| Faceless short | stills + img→video + VO + music + captions | (explainer pipeline) + burn captions | replicate-images; ../video-shorts/SKILL.md for the script |
| Just a voiceover | VO only | script → TTS → loudnorm | — |
| Just a clip from a still | img→video only | still (input) → clip | fal/replicate |
| Code-rendered explainer | none of the above | render from React/TS | stop — route to remotion-video |
If the video is rendered from data/code (charts, timelines, JSON-driven scenes), this is not your job → ../remotion-video/SKILL.md. You handle model-generated + ffmpeg-glued.
Modality 1 — Voice (TTS)
ElevenLabs Python SDK. The call is convert(text, voice_id, model_id, output_format); auth via ELEVENLABS_API_KEY.
from elevenlabs.client import ElevenLabs
client = ElevenLabs() # reads ELEVENLABS_API_KEY
audio = client.text_to_speech.convert(
text="Your narration script here.",
voice_id="JBFqnCBsd6RMkjVDRZzb",
model_id="eleven_multilingual_v2", # final-quality VO
output_format="mp3_44100_128", # codec_samplerate_bitrate
)
with open("vo.mp3", "wb") as f:
for chunk in audio:
f.write(chunk)
Pick the model tier by what the job needs:
| Model | When | Latency | Cost lever |
|---|---|---|---|
eleven_v3 |
most expressive, hero final VO — verify availability first (see caveat) | higher | most credits/char |
eleven_multilingual_v2 |
high-quality multilingual VO (default for finals) | medium | medium |
eleven_flash_v2_5 |
real-time / batch / scale / drafts | ~75 ms | cheapest |
Do not hardcode eleven_v3 blind. It shipped to the API in alpha (Aug 2025) and the docs model list now carries it, but the text_to_speech.convert API reference still documents the default as eleven_multilingual_v2 and does not enumerate eleven_v3 as a guaranteed value. Before you build a final pass on it, confirm it returns from GET /v1/models for your key (or just call once and check) — otherwise default to eleven_multilingual_v2, which is the safe, always-available hero tier.
output_format is codec_samplerate_bitrate — e.g. mp3_44100_128, mp3_22050_32. Match the VO sample rate to your assembly target, do not master the VO loud and hope.
Bad → Good:
- Bad: generate VO at
mp3_22050_32, then mux onto a 48 kHz video — ffmpeg silently resamples, you get artifacts and a level mismatch. - Good: generate VO at the rate you will master at (e.g.
mp3_44100_128), and set final loudness withloudnormin assembly, not by cranking the TTS.
TTS is billed per character/token (~0.5–1 credit/char on the Flash/Turbo lines). ElevenLabs cut TTS API pricing up to 55% on 2026-05-07 (e.g. Flash on Creator $0.11→$0.05 / 1k tokens) — that figure is TTS-specific, not the Music cut. Pricing staling fast: these are point-in-time numbers from elevenlabs.io/pricing/api as of 2026-06-02 — re-check the page before quoting a budget. Shorter scripts and Flash on drafts are the cost levers.
Modality 2 — Image-to-video
The still is an input, not your output. Generate or edit the source image in ../replicate-images/SKILL.md, then animate it here. Reality check: every serious 2026 model does 1080p or native 4K — resolution is no longer the differentiating axis. The hard limit is per-generation duration (~5–15 s, model-dependent). Long pieces are one clip per scene, then concat — never one long take.
Durations below are from each vendor's own pages (as of 2026-06; see references/models-and-params.md for the citations) — they move with releases, so verify on the catalog before a final run:
| Model | Duration | Aspect / max res | Control surface | Native audio | Open-source |
|---|---|---|---|---|---|
| Google Veo 3.1 | 8 s / generation | 16:9 / 9:16, up to 4K | high | yes — synced 48 kHz dialogue/SFX | no |
| Kling 3.0 | up to 15 s | flexible, 4K | strong identity/temporal, lip-sync | no | no |
| Runway Gen-4.5 | 2–10 s | flexible | best — motion brushes, camera control, reference image | no | no |
| MiniMax Hailuo 02 | 6 s or 10 s (1080p caps at 6 s) | up to 1080p | medium | no | no |
| Wan 2.6 | up to 15 s | up to 1080p | first/last-frame control, A/V sync | no (sync) | yes (Apache) |
Choose by the binding constraint: need synced dialogue → Veo 3.1; need precise camera/motion control → Runway Gen-4.5; need identity consistency across scenes or the longest single take → Kling 3.0 / Wan 2.6; need open-source/self-host → Wan 2.6; cost-sensitive 1080p → Hailuo 02. Endpoint ids and per-call mechanics live in ../fal/SKILL.md / ../replicate/SKILL.md (both rails carry these models). See references/models-and-params.md for endpoint ids and current limits.
Modality 3 — Music / score
Costs are per-minute and plan-dependent — treat them as approximate and verify on the vendor pricing page (figures as of 2026-06; sources in references/models-and-params.md):
| Model | Cost (approx, verify) | Licensing story | Control |
|---|---|---|---|
| ElevenLabs Music v2 | per-minute, ~$0.15–0.50/min depending on plan (Music API pricing cut up to 50% at v2 launch — separate from the 55% TTS cut) | cleanest — vendor states trained only on licensed data, cleared for commercial use (Believe collaboration named at launch) | genre-switch mid-track |
| Suno v5 | plan-based | usage rights on paid plans post Nov-2025 label settlements (rights, not ownership) | vendor blind-test benchmark ELO ~1293 |
| Udio | $30/mo Pro plan (commercial rights); no official public API — third-party gateways only | UMG-licensed platform announced for 2026 | — |
Confirm commercial rights before you ship. Licensing differs per model and per plan; "I generated it" is not "I may sell the ad with it." For a clean commercial story with an official API, ElevenLabs Music v2 is the safe default — Udio has no first-party API, so do not plan a programmatic pipeline around it. The rest is the same fal/replicate call mechanics.
Assembly with ffmpeg
Four operations. Each is a copy-paste recipe; full filter graphs, caption burning and pitfalls are in references/ffmpeg-assembly.md.
(a) Mux VO onto video — map both streams, copy video, take the shorter duration:
ffmpeg -i scene.mp4 -i vo.mp3 \
-map 0:v -map 1:a -c:v copy -shortest out.mp4
(b) Duck music under the VO — sidechaincompress keys the music off the voice so it drops when narration plays (pro DAW ducking, no manual keyframes):
ffmpeg -i vo.mp3 -i music.mp3 -filter_complex \
"[1:a][0:a]sidechaincompress=threshold=0.03:ratio=8:attack=20:release=300[duck]; \
[0:a][duck]amix=inputs=2:duration=longest[aout]" \
-map "[aout]" -c:a aac mix.m4a
Cheaper static fallback when sidechain is overkill — fix the music low under a full VO:
ffmpeg -i vo.mp3 -i music.mp3 -filter_complex \
"[1:a]volume=0.3[m];[0:a][m]amix=inputs=2:duration=longest[aout]" \
-map "[aout]" mix.m4a
(c) Loudnorm to a target LUFS (two-pass) — measure, then apply. Target -14 LUFS for social/streaming, -16 for podcast-style VO. Normalize per track before mixing.
# pass 1: measure (read the JSON it prints)
ffmpeg -i mix.m4a -af loudnorm=I=-14:TP=-1.5:LRA=11:print_format=json -f null -
# pass 2: apply with the measured values
ffmpeg -i mix.m4a -af \
loudnorm=I=-14:TP=-1.5:LRA=11:measured_I=-20.1:measured_TP=-4.2:measured_LRA=6.0:measured_thresh=-30.8:offset=0.5:linear=true \
master.m4a
(d) Concat scenes — conform first. Same-codec/res/fps clips → fast demuxer with -c copy. Mismatched clips → re-encode and scale first, then concat. Never -c copy-concat mismatched clips — you get desync or a corrupt stream.
# all clips identical codec/res/fps:
printf "file '%s'\n" scene1.mp4 scene2.mp4 scene3.mp4 > list.txt
ffmpeg -f concat -safe 0 -i list.txt -c copy joined.mp4
# mismatched: conform each, then concat filter
ffmpeg -i s1.mp4 -i s2.mp4 -filter_complex \
"[0:v]scale=1920:1080,fps=30,setsar=1[v0];[1:v]scale=1920:1080,fps=30,setsar=1[v1]; \
[v0][1:a?][v1][1:a?]concat=n=2:v=1:a=0[v]" -map "[v]" joined.mp4
End-to-end worked pipeline (narrated explainer)
Ordered command list — generate once, assemble deterministically:
- Lock the plan — scene list, aspect (e.g. 16:9 1080p 30fps), target -14 LUFS, models chosen.
- Stills per scene →
../replicate-images/SKILL.md(one prompt per scene). - Clip per scene (img→video, ~5–15 s each, model cap) via your chosen model on
fal/replicate. - VO → ElevenLabs
convert(...)at the master sample rate. - Music → ElevenLabs Music v2, length = total runtime, confirm rights.
- Conform every clip to 1920x1080/30fps/SAR 1.
- Concat the conformed clips →
body.mp4. - Loudnorm VO and music tracks (two-pass) to consistent levels.
- Mix + duck music under VO →
mix.m4a. - Final loudnorm the mix to -14 LUFS →
master.m4a. - Mux
master.m4aontobody.mp4with-shortest→final.mp4.
Emit this as a runnable script. scripts/verify.sh lints it (loudnorm present, conform-before-concat, final MP4 target).
Cost & regen discipline
- Draft small, then final. Generate clips short/low-res and VO on Flash to lock timing and the cut; only the final pass spends on hero quality. Re-rolling locked scenes is the biggest waste.
- Per-modality levers: shorter scripts (TTS per-char), fewer scene re-rolls (img→video), fewer music minutes (music is billed per minute — verify the current rate on the vendor pricing page).
- Spend tracking as a discipline →
../fal/SKILL.md/../replicate/SKILL.mdfor per-call cost; treat budget as a constraint you set before generating.
Anti-patterns
| Anti-pattern | Why it bites | Do instead |
|---|---|---|
| Generating assets before locking the pipeline | aspect/sample-rate/duration mismatches surface at mux time, forcing paid re-rolls | lock scene list, aspect, LUFS, models first |
| One long video-gen call for the whole piece | models cap at ~5–15 s; you fight the limit and waste rolls | one clip per scene, then concat |
Mixing tracks without per-track loudnorm |
VO buried or blasting over music; inconsistent levels | two-pass loudnorm each track before mix |
-c copy-concat of mismatched clips |
desync, corrupt stream, wrong frame timing | conform res/fps/SAR, then concat |
| Music low set by ear / static only when VO needs space | narration gets masked under the bed | sidechaincompress keyed off the VO |
| Shipping generated music without checking rights | "generated" ≠ "licensed to sell" — legal exposure | confirm commercial rights per model/plan |
| Prompting/editing the still inside this skill | duplicates replicate-images' job, worse prompts |
delegate the still, consume it here |
| Mastering loudness by cranking the TTS | clipping, no true-peak control | set level with loudnorm, not the generator |
Files (rsc-harness)
-
evals
-
cases.yaml 3.4 KB
skill: ai-media should_trigger: - prompt: "Turn this 60-second script into a narrated video with AI voice, some B-roll clips, and background music." why: Full cross-modal pipeline — TTS + image-to-video + music + ffmpeg assembly, the core job. - prompt: "Generate an AI voiceover for my explainer and mix it over the footage, ducking the music when the voice plays." why: Non-obvious — the assembly core (sidechaincompress ducking + mux) is what makes this ai-media, not a single-model call. - prompt: "I have a product photo — make it into a 9:16 video clip with subtle motion and add a music bed." why: Image-to-video + music where the still is an input (consumed), not generated here. - prompt: "Genérame una voz en off en español y un vídeo a partir de este guion, con música de fondo." why: Spanish phrasing for the full narrated-video pipeline. - prompt: "I generated 5 scene clips with Kling — stitch them into one MP4, normalize the loudness, and lay my voiceover on top." why: Non-obvious — assembly-only (conform + concat + loudnorm + mux), zero generation, still squarely this skill. - prompt: "Posa música de fons al meu vídeo i abaixa-la quan parla la veu, després deixa-ho tot a -14 LUFS." why: Catalan — ducking + loudnorm mastering, the assembly discipline. should_not_trigger: - prompt: "Generate a photorealistic product hero image and inpaint out the background." route_to: replicate-images why: Still-image creation/editing — the named boundary; ai-media consumes stills, does not prompt them. - prompt: "How do I verify the fal webhook signature and poll the queue for my video job?" route_to: fal why: Provider call mechanics (queue/webhook), not cross-modal media strategy. - prompt: "Build a data-driven animated explainer in React where text and charts render from JSON." route_to: remotion-video why: Code-rendered, frame-exact compositing — not model-generated + ffmpeg-glued. - prompt: "Write me a punchy 30-second TikTok script with a 3-second hook and on-screen text." route_to: video-shorts why: Script/cut creative craft — ai-media renders/voices a script, it does not write the beat sheet. - prompt: "Set up the full podcast episode: run-of-show, multitrack recording, RSS item and show notes." route_to: podcast why: Audio-episode lifecycle — ai-media makes a VO/score, it does not run a podcast. capability: - scenario: "Given a 4-scene narrated-ad script, produce the end-to-end pipeline and assemble the final 1080p 16:9 MP4." must_include: - Plans the whole pipeline (scene list, aspect, target LUFS, models) before generating any asset - Delegates the still images to replicate-images rather than prompting them here - Picks a TTS model with rationale (Flash for drafts/scale vs multilingual_v2 as the safe hero default; treats eleven_v3 as verify-availability-first, not a blind hardcode) - Generates one clip per scene (~5-15 s, model cap), not one long take, because models cap clip duration - Applies per-track loudnorm before mixing the audio - Ducks the music under the VO (sidechaincompress, or static volume fallback) - Conforms clips to one resolution/fps/SAR before concat (no -c copy on mismatched clips) - Final mux yields one MP4 carrying both a video and an audio stream - Confirms the music's commercial rights before shipping -
README.md 781 B
# Evals — ai-media `cases.yaml` holds the trigger boundary and one capability rubric for this skill. There is no automated runner here: feed each `should_trigger` / `should_not_trigger` prompt to your skill-routing setup and confirm ai-media is selected (or that the prompt routes to the named sibling). For the `capability` case, run the scenario end to end and check the produced pipeline + plan against every line in `must_include` by hand — each item is a pass/fail bullet. The point is that the agent plans the cross-modal pipeline before generating, delegates stills to `replicate-images`, generates per-scene (not one long take), and masters with per-track `loudnorm` + ducking before a final mux. Pair this with `scripts/verify.sh` to lint an emitted assembly script.
-
-
references
-
ffmpeg-assembly.md 5 KB
# ffmpeg assembly cookbook Branch-specific to the assembly step. Every operation here is deterministic — same inputs, same output. Lock generation first; this is the glue. ## Mux: audio file onto a video Map the video from input 0 and the audio from input 1, copy the video stream (no re-encode), trim to the shorter of the two: ```bash ffmpeg -i scene.mp4 -i vo.mp3 \ -map 0:v -map 1:a -c:v copy -c:a aac -shortest out.mp4 ``` - `-map 0:v -map 1:a` — pick exactly the streams you want; without it ffmpeg guesses. - `-c:v copy` — never re-encode video you do not need to touch. - `-shortest` — stops at the shorter input. Watch this: if the VO is longer than the clip, the tail is cut. Pad the video or trim the VO deliberately. ## Two-pass loudnorm (ITU-R BS.1770) Pass 1 measures, pass 2 applies with the measured values. Targets: **-14 LUFS** social/streaming, **-16 LUFS** podcast-style VO. Always set a true-peak ceiling (`TP=-1.5`). ```bash # pass 1 — read the JSON block it prints to stderr ffmpeg -i input.wav -af loudnorm=I=-14:TP=-1.5:LRA=11:print_format=json -f null - ``` Copy `input_i`, `input_tp`, `input_lra`, `input_thresh`, `target_offset` into pass 2: ```bash ffmpeg -i input.wav -af \ loudnorm=I=-14:TP=-1.5:LRA=11:measured_I=-20.1:measured_TP=-4.2:measured_LRA=6.0:measured_thresh=-30.8:offset=0.5:linear=true \ -ar 48000 output.wav ``` Normalize **each track before mixing**, then loudnorm the final mix once more. One-pass loudnorm is acceptable for quick drafts; two-pass is the accurate path. `slhck/ffmpeg-normalize` wraps this if you want it scripted. ## Duck music under the VO — sidechaincompress The music (carrier) is keyed by the VO (sidechain): when the voice plays, the compressor pulls the music down; when it stops, the music recovers. This emulates a DAW ducking automation without manual keyframes. ```bash ffmpeg -i vo.wav -i music.wav -filter_complex " [1:a][0:a]sidechaincompress=threshold=0.03:ratio=8:attack=20:release=300[ducked]; [0:a][ducked]amix=inputs=2:duration=longest:dropout_transition=0[aout] " -map "[aout]" -c:a aac mix.m4a ``` Tuning: - `threshold` — how loud the VO must be to trigger ducking (lower = more sensitive). - `ratio` — how hard the music drops (8 = strong). - `attack` (ms) — how fast it ducks; keep small so the music dips immediately. - `release` (ms) — how slowly it recovers; 300–500 avoids pumping. ### Static fallback (no sidechain) When you do not need dynamic ducking, just fix the music low under a full-length VO: ```bash ffmpeg -i vo.wav -i music.wav -filter_complex \ "[1:a]volume=0.3[m];[0:a][m]amix=inputs=2:duration=longest[aout]" \ -map "[aout]" mix.m4a ``` `amix` averages levels and can lower perceived loudness — re-run loudnorm on the result. ## Concat scenes ### Demuxer (fast, no re-encode) — only for identical codec/res/fps/SAR ```bash printf "file '%s'\n" s1.mp4 s2.mp4 s3.mp4 > list.txt ffmpeg -f concat -safe 0 -i list.txt -c copy joined.mp4 ``` If clips differ in any of codec/resolution/fps/SAR, this produces desync or a broken stream. Conform first. ### Concat filter (re-encode) — for mismatched clips Scale + set fps + normalize SAR per input, then concat: ```bash ffmpeg -i s1.mp4 -i s2.mp4 -i s3.mp4 -filter_complex " [0:v]scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,fps=30,setsar=1[v0]; [1:v]scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,fps=30,setsar=1[v1]; [2:v]scale=1920:1080:force_original_aspect_ratio=decrease,pad=1920:1080:(ow-iw)/2:(oh-ih)/2,fps=30,setsar=1[v2]; [v0][v1][v2]concat=n=3:v=1:a=0[v] " -map "[v]" -r 30 joined.mp4 ``` ## Conforming — the gotchas - **SAR/DAR mismatch** — clips with different sample aspect ratios concat to skewed frames. `setsar=1` on every input. - **fps mismatch** — `fps=30` (or your target) on every input; otherwise concat timing drifts. - **Resolution** — `scale=...:force_original_aspect_ratio=decrease` + `pad` letterboxes without distortion. Plain `scale=W:H` stretches. ## Burn captions / subtitles ```bash ffmpeg -i joined.mp4 -vf "subtitles=captions.srt:force_style='FontSize=24,PrimaryColour=&H00FFFFFF'" \ -c:a copy captioned.mp4 ``` Burned-in (hardsub) for social where soft subs are ignored. Use `-c:s mov_text` to mux a soft subtitle track instead when the player supports it. ## Pitfalls checklist - `-shortest` silently truncates — confirm which input is shorter and whether that is intended. - Audio sample-rate mismatch → ffmpeg resamples and may shift level; set `-ar 48000` consistently. - `-c copy`-concat across mismatched clips → desync/corruption; conform first. - `amix` lowers perceived loudness → loudnorm the result. - Forgetting `setsar=1` → skewed frames after concat. ## Sources - mux.com "combine audio and video with FFmpeg" / FFmpeg mixing guide; cloudinary FFmpeg add-audio guide. (accessed 2026-06-02) - legacistudios.com FFmpeg mixing/ducking guide. (accessed 2026-06-02) - slhck/ffmpeg-normalize; ffmpeg loudnorm & concat docs. (accessed 2026-06-02) -
models-and-params.md 7.2 KB
# Models & params — current map (2026-06) Fast-staling reference. Verify endpoint ids and limits on the provider catalog before a final run. Call mechanics live in `../fal/SKILL.md` and `../replicate/SKILL.md`. ## TTS — ElevenLabs SDK: `elevenlabs` (Python). Auth: `ELEVENLABS_API_KEY`. Call: ```python client.text_to_speech.convert( text=..., voice_id=..., model_id=..., output_format=..., ) ``` ### Model tiers | model_id | Profile | Latency | Use | |----------|---------|---------|-----| | `eleven_v3` | most expressive, highest quality | higher | hero final VO, emotional delivery — **verify availability first** | | `eleven_multilingual_v2` | high quality, nuanced, multilingual | medium | default for final VO (always-available) | | `eleven_flash_v2_5` | ultra-low ~75 ms latency | ~75 ms | real-time, batch, scale, drafts | **`eleven_v3` availability caveat.** v3 shipped to the API in *alpha* (elevenlabs.io/blog/eleven-v3-alpha-now-available-in-the-api, 2025-08-20) and now appears in `/docs/overview/models`, but the `text_to_speech.convert` API reference still documents the default as `eleven_multilingual_v2` and does not enumerate `eleven_v3` as a guaranteed value on that endpoint. Do not hardcode `model_id="eleven_v3"` for a production pass without first confirming it returns from `GET /v1/models` for your key (or test-calling once). Default to `eleven_multilingual_v2` when in doubt — it is the safe always-available hero tier. ### output_format — `codec_samplerate_bitrate` | Code | Meaning | |------|---------| | `mp3_44100_128` | MP3, 44.1 kHz, 128 kbps — good master default | | `mp3_22050_32` | MP3, 22.05 kHz, 32 kbps — small/drafts only | | (others) | PCM / µ-law variants per docs | Match the sample rate to your assembly master rate. Do not resample at mux time. ### Cost Billed per character/token. ~0.5–1 credit/char on Flash/Turbo lines. **TTS** API pricing dropped up to 55% on 2026-05-07 (updated 2026-05-27) — the primary blog quotes Flash on Creator going $0.11→$0.05 per 1,000 tokens; Multilingual v2/v3 is roughly double Flash per character. This is the TTS figure; the Music API cut was a separate up-to-50% (see Music section). **Point-in-time as of 2026-06-02 — re-verify on elevenlabs.io/pricing/api before quoting a budget.** Levers: shorter scripts, Flash for drafts/scale. ## Image-to-video Resolution is no longer the differentiator — every serious model hits 1080p or native 4K. The binding constraint is **per-generation duration (~5–15 s, model-dependent)**: long pieces = clip per scene + concat. Durations below are each vendor's own published figure as of 2026-06 (citations in Sources) — they move with releases, so confirm on the catalog before a final run. | Model | Duration / generation | Aspect / max res | Control | Native audio | Open-source | Pick when | |-------|-----------------------|------------------|---------|--------------|-------------|-----------| | Google **Veo 3.1** | **8 s** | 16:9 & 9:16, 720p/1080p/4K | high | **yes** (synced 48 kHz dialogue/SFX) | no | need synced spoken dialogue/SFX | | **Kling 3.0** | **up to 15 s** | 4K | strong identity/temporal, lip-sync | no | no | identity consistency across scenes; longest single take | | **Runway Gen-4.5** | **2–10 s** | flexible | **best** — motion brushes, camera control, reference image | no | no | precise motion/camera control | | **MiniMax Hailuo 02** | **6 s or 10 s** (1080p caps at 6 s) | 768p / 1080p | medium | no | no | cost-sensitive 1080p | | **Wan 2.6** | **up to 15 s** | up to 1080p | first/last-frame control, A/V sync | no (sync) | **yes** (Apache 2.0) | self-host / open-source | Endpoint ids: look up the current model slug on the fal model catalog or the Replicate model catalog (both rails carry Veo/Kling/Wan/Hailuo; TTS and music endpoints also on fal). The slugs change — do not hardcode from memory. ## Music / score Costs are per-minute and plan-dependent — approximate and fast-staling, verify on the vendor pricing page. | Model | Cost (approx as of 2026-06, verify) | Licensing | Duration / control | |-------|-------------------------------------|-----------|--------------------| | **ElevenLabs Music v2** (announced 2026-05-26, upd 2026-05-31) | per-minute, ~$0.15–0.50/min depending on plan/source; Music v1/v2 API pricing cut up to 50% at launch (Creative self-serve up to ~40%) — distinct from the 55% *TTS* cut | vendor states trained **only on licensed data, cleared for commercial use** (Believe collaboration named in the launch post; Merlin/Kobalt not specifically cited there) — cleanest commercial story | genre-switch mid-track | | **Suno v5** | plan-based | usage rights on paid plans post Nov-2025 label settlements (rights, not copyright ownership) | vendor blind-test benchmark, ELO ~1293 (v5, 2025-09) | | **Udio** | $30/mo Pro plan (commercial rights); **no official public API** as of this window — third-party gateways only | UMG-licensed platform announced for 2026 | — | **Confirm commercial rights before shipping.** Terms differ per model and per plan. ElevenLabs Music v2 is the safe default for ads/commercial output and is the only one of the three with a first-party API — do not plan a programmatic pipeline around Udio. The ElevenLabs Music per-minute dollar rate spread (a primary `/pricing/api` render showed ~$0.15/min; aggregators report ~$0.50/min) is exactly why this number is a verify-on-catalog range, not a hardcoded fact. ## Sources Primary vendor docs first; aggregators only where a primary page is JS-gated and the figure is corroborated across several. - **TTS (ElevenLabs):** github.com/elevenlabs/elevenlabs-python README; elevenlabs.io/docs/api-reference/text-to-speech/convert (default `eleven_multilingual_v2`, no enumerated `eleven_v3`); elevenlabs.io/docs/overview/models (lists `eleven_v3`); elevenlabs.io/blog/eleven-v3-alpha-now-available-in-the-api (v3 shipped *alpha*, 2025-08-20). (accessed 2026-06-02) - **TTS pricing:** elevenlabs.io/blog/weve-lowered-api-agents-pricing-and-introduced-pay-as-you-go (up-to-55% TTS cut, Flash Creator $0.11→$0.05/1k tokens, 2026-05-07 upd 2026-05-27); elevenlabs.io/pricing/api. (accessed 2026-06-02) - **Image-to-video (per-vendor durations):** Veo 3.1 — deepmind.google/models/veo (8 s, up to 4K, native 48 kHz audio; released 2025-10-14). Kling 3.0 — ir.kuaishou.com Kling 3.0 launch release (up to 15 s; 10 s was the 2.6 ceiling). Runway Gen-4.5 — help.runwayml.com "Creating with Gen-4.5" (2–10 s). MiniMax Hailuo 02 — replicate.com/minimax/hailuo-02 + minimax.io news (6 s/10 s, 768p/1080p, 1080p caps at 6 s). Wan 2.6 — alibabacloud.com Wan2.6 announcement (up to 15 s, A/V sync, Apache 2.0). (accessed 2026-06-02) - **Music:** elevenlabs.io/blog/introducing-music-v2 (2026-05-26 upd 2026-05-31; trained only on licensed data, Believe collaboration; Music API cut up to 50% / Creative ~40%); elevenlabs.io/pricing/api + help.elevenlabs.io "How much does Eleven Music cost" (per-minute billing; dollar rate plan-dependent, ~$0.15–0.50/min spread across primary vs aggregator). Suno v5 ELO ~1293 — Suno-published blind-test benchmark (v5, 2025-09). Udio — udio.com/pricing ($30/mo Pro, commercial rights; no first-party API). (accessed 2026-06-02) - **Delivery rails:** fal.ai model catalog; replicate.com model catalog. (accessed 2026-06-02)
-
-
scripts
-
verify.sh 3.2 KB
#!/usr/bin/env bash # verify.sh — lint an emitted ai-media assembly pipeline (read-only). # # Checks, over *.sh shell scripts in a target dir: # 1. ffmpeg assembly applies `loudnorm` somewhere (mastering present). # 2. multi-scene jobs conform (scale/fps/re-encode) before `concat`, rather # than `-c copy`-concatenating mismatched clips. # 3. a final muxed .mp4 output target is named. # 4. if a final .mp4 exists AND ffprobe is available, it carries both a # video and an audio stream. # # Lint + optional ffprobe presence check — never renders, never writes. # Exits 0 on an empty/clean target (no false failure). # # Usage: scripts/verify.sh [TARGET_DIR] (default: .) set -uo pipefail TARGET="${1:-.}" fail=0 warn=0 note() { printf ' %s\n' "$*"; } bad() { printf 'FAIL: %s\n' "$*"; fail=1; } soft() { printf 'WARN: %s\n' "$*"; warn=1; } if [ ! -d "$TARGET" ]; then echo "verify.sh: target dir not found: $TARGET" exit 1 fi # Collect candidate pipeline scripts that actually invoke ffmpeg. scripts=() while IFS= read -r f; do if grep -lq 'ffmpeg' "$f" 2>/dev/null; then scripts+=("$f") fi done < <(find "$TARGET" -type f -name '*.sh' 2>/dev/null) if [ "${#scripts[@]}" -eq 0 ]; then echo "verify.sh: no ffmpeg pipeline scripts found under $TARGET — nothing to lint." exit 0 fi echo "Linting ${#scripts[@]} pipeline script(s) under $TARGET" for s in "${scripts[@]}"; do echo "- $s" # 1. loudnorm mastering present. if grep -Eq 'loudnorm' "$s"; then note "loudnorm present (mastering ok)" else bad "no 'loudnorm' — audio is not loudness-normalized before final mux ($s)" fi # 2. conform-before-concat for multi-scene jobs. if grep -Eq 'concat' "$s"; then if grep -Eq '(scale=|fps=|setsar=|concat=n=)' "$s"; then note "concat conforms clips (scale/fps/setsar or concat filter) — ok" elif grep -Eq 'concat[^=]*-c[: ]*v?[: ]*copy|-c copy' "$s"; then bad "uses concat with '-c copy' but no scale/fps/setsar conform — mismatched clips will desync ($s)" else soft "concat present but conform step not detected — confirm clips share codec/res/fps ($s)" fi fi # 3. a final .mp4 output target is named. if grep -Eq '[A-Za-z0-9_./-]+\.mp4' "$s"; then note "final .mp4 output target named — ok" else bad "no '.mp4' output target named in the pipeline ($s)" fi done # 4. optional ffprobe stream check on any rendered .mp4. if command -v ffprobe >/dev/null 2>&1; then while IFS= read -r mp4; do streams="$(ffprobe -v error -show_entries stream=codec_type \ -of default=nw=1:nk=1 "$mp4" 2>/dev/null)" has_v=0; has_a=0 echo "$streams" | grep -q '^video$' && has_v=1 echo "$streams" | grep -q '^audio$' && has_a=1 if [ "$has_v" -eq 1 ] && [ "$has_a" -eq 1 ]; then note "ffprobe: $mp4 has both video + audio — ok" else bad "ffprobe: $mp4 missing $( [ $has_v -eq 0 ] && echo 'video' ) $( [ $has_a -eq 0 ] && echo 'audio' ) stream" fi done < <(find "$TARGET" -type f -name '*.mp4' 2>/dev/null) else echo "(ffprobe not installed — skipping rendered-MP4 stream check)" fi if [ "$fail" -ne 0 ]; then echo "verify.sh: FAILED" exit 1 fi if [ "$warn" -ne 0 ]; then echo "verify.sh: passed with warnings" exit 0 fi echo "verify.sh: OK" exit 0
-
-
SKILL.md 13 KB
--- name: ai-media description: "Use when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with ffmpeg (mux, duck, loudnorm, concat). NOT still-image generation/editing (that is `replicate-images`); NOT code-rendered React compositing (that is `remotion-video`)." tags: [ai-media, text-to-speech, image-to-video, voiceover, music-generation, ffmpeg, media-pipeline, elevenlabs] recommends: [replicate-images, fal, replicate, remotion-video, video-shorts] origin: risco --- # ai-media You are the cross-modal director. You decide **which** generative-media model to call per modality, in **what order**, with **what params**, then **assemble** the pieces with ffmpeg into one finished file. You do not own a single provider's API surface and you do not prompt still images — you orchestrate and glue. ## Pipeline shape — decide what the goal needs Map the goal to modalities and an ordered step list, and **lock that plan before you generate a single asset** — media generation is slow and metered, so a re-roll of a 10 s Veo clip or a 90 s music track costs real money and minutes. Fixing the scene list, aspect ratio, target loudness and model per modality *first* is cheaper than discovering at mux time that your clips are 9:16 and your VO is the wrong sample rate. The "delegate to" column is where the actual call mechanics live — you pick the model and params, those skills run the call. | Goal | Needs | Ordered steps | Delegate calls to | |------|-------|---------------|-------------------| | Narrated explainer | stills + img→video + VO + music | script → per-scene stills → clip per scene → VO → music → conform → concat → mix+duck → loudnorm → MP4 | `replicate-images`, `fal`/`replicate` | | Product teaser (1 hero) | 1 still + img→video + music | still → clip → music → mix → loudnorm → MP4 | `replicate-images`, `fal`/`replicate` | | Faceless short | stills + img→video + VO + music + captions | (explainer pipeline) + burn captions | `replicate-images`; `../video-shorts/SKILL.md` for the script | | Just a voiceover | VO only | script → TTS → loudnorm | — | | Just a clip from a still | img→video only | still (input) → clip | `fal`/`replicate` | | Code-rendered explainer | none of the above | render from React/TS | **stop — route to `remotion-video`** | If the video is *rendered from data/code* (charts, timelines, JSON-driven scenes), this is not your job → `../remotion-video/SKILL.md`. You handle *model-generated + ffmpeg-glued*. ## Modality 1 — Voice (TTS) ElevenLabs Python SDK. The call is `convert(text, voice_id, model_id, output_format)`; auth via `ELEVENLABS_API_KEY`. ```python from elevenlabs.client import ElevenLabs client = ElevenLabs() # reads ELEVENLABS_API_KEY audio = client.text_to_speech.convert( text="Your narration script here.", voice_id="JBFqnCBsd6RMkjVDRZzb", model_id="eleven_multilingual_v2", # final-quality VO output_format="mp3_44100_128", # codec_samplerate_bitrate ) with open("vo.mp3", "wb") as f: for chunk in audio: f.write(chunk) ``` Pick the model tier by what the job needs: | Model | When | Latency | Cost lever | |-------|------|---------|-----------| | `eleven_v3` | most expressive, hero final VO — **verify availability first (see caveat)** | higher | most credits/char | | `eleven_multilingual_v2` | high-quality multilingual VO (default for finals) | medium | medium | | `eleven_flash_v2_5` | real-time / batch / scale / drafts | ~75 ms | cheapest | **Do not hardcode `eleven_v3` blind.** It shipped to the API in *alpha* (Aug 2025) and the docs model list now carries it, but the `text_to_speech.convert` API reference still documents the default as `eleven_multilingual_v2` and does not enumerate `eleven_v3` as a guaranteed value. Before you build a final pass on it, confirm it returns from `GET /v1/models` for your key (or just call once and check) — otherwise default to `eleven_multilingual_v2`, which is the safe, always-available hero tier. `output_format` is `codec_samplerate_bitrate` — e.g. `mp3_44100_128`, `mp3_22050_32`. **Match the VO sample rate to your assembly target**, do not master the VO loud and hope. Bad → Good: - **Bad:** generate VO at `mp3_22050_32`, then mux onto a 48 kHz video — ffmpeg silently resamples, you get artifacts and a level mismatch. - **Good:** generate VO at the rate you will master at (e.g. `mp3_44100_128`), and set final loudness with `loudnorm` in assembly, not by cranking the TTS. TTS is billed per character/token (~0.5–1 credit/char on the Flash/Turbo lines). ElevenLabs cut *TTS* API pricing up to 55% on 2026-05-07 (e.g. Flash on Creator $0.11→$0.05 / 1k tokens) — that figure is TTS-specific, not the Music cut. **Pricing staling fast: these are point-in-time numbers from elevenlabs.io/pricing/api as of 2026-06-02 — re-check the page before quoting a budget.** Shorter scripts and Flash on drafts are the cost levers. ## Modality 2 — Image-to-video **The still is an input, not your output.** Generate or edit the source image in `../replicate-images/SKILL.md`, then animate it here. Reality check: every serious 2026 model does 1080p or native 4K — **resolution is no longer the differentiating axis**. The hard limit is **per-generation duration (~5–15 s, model-dependent)**. Long pieces are **one clip per scene, then concat** — never one long take. Durations below are from each vendor's own pages (as of 2026-06; see `references/models-and-params.md` for the citations) — they move with releases, so verify on the catalog before a final run: | Model | Duration | Aspect / max res | Control surface | Native audio | Open-source | |-------|----------|------------------|-----------------|--------------|-------------| | Google **Veo 3.1** | **8 s** / generation | 16:9 / 9:16, up to 4K | high | **yes** — synced 48 kHz dialogue/SFX | no | | **Kling 3.0** | **up to 15 s** | flexible, 4K | strong identity/temporal, lip-sync | no | no | | **Runway Gen-4.5** | **2–10 s** | flexible | **best** — motion brushes, camera control, reference image | no | no | | **MiniMax Hailuo 02** | **6 s or 10 s** (1080p caps at 6 s) | up to 1080p | medium | no | no | | **Wan 2.6** | **up to 15 s** | up to 1080p | first/last-frame control, A/V sync | no (sync) | **yes** (Apache) | Choose by the binding constraint: need synced dialogue → Veo 3.1; need precise camera/motion control → Runway Gen-4.5; need identity consistency across scenes or the longest single take → Kling 3.0 / Wan 2.6; need open-source/self-host → Wan 2.6; cost-sensitive 1080p → Hailuo 02. Endpoint ids and per-call mechanics live in `../fal/SKILL.md` / `../replicate/SKILL.md` (both rails carry these models). See `references/models-and-params.md` for endpoint ids and current limits. ## Modality 3 — Music / score Costs are per-minute and plan-dependent — treat them as approximate and **verify on the vendor pricing page** (figures as of 2026-06; sources in `references/models-and-params.md`): | Model | Cost (approx, verify) | Licensing story | Control | |-------|-----------------------|-----------------|---------| | **ElevenLabs Music v2** | per-minute, ~$0.15–0.50/min depending on plan (Music API pricing cut up to 50% at v2 launch — separate from the 55% *TTS* cut) | **cleanest** — vendor states trained *only on licensed data, cleared for commercial use* (Believe collaboration named at launch) | genre-switch mid-track | | **Suno v5** | plan-based | usage rights on paid plans post Nov-2025 label settlements (rights, not ownership) | vendor blind-test benchmark ELO ~1293 | | **Udio** | $30/mo Pro plan (commercial rights); **no official public API** — third-party gateways only | UMG-licensed platform announced for 2026 | — | **Confirm commercial rights before you ship.** Licensing differs per model and per plan; "I generated it" is not "I may sell the ad with it." For a clean commercial story with an official API, ElevenLabs Music v2 is the safe default — Udio has no first-party API, so do not plan a programmatic pipeline around it. The rest is the same fal/replicate call mechanics. ## Assembly with ffmpeg Four operations. Each is a copy-paste recipe; full filter graphs, caption burning and pitfalls are in `references/ffmpeg-assembly.md`. **(a) Mux VO onto video** — map both streams, copy video, take the shorter duration: ```bash ffmpeg -i scene.mp4 -i vo.mp3 \ -map 0:v -map 1:a -c:v copy -shortest out.mp4 ``` **(b) Duck music under the VO** — `sidechaincompress` keys the music off the voice so it drops when narration plays (pro DAW ducking, no manual keyframes): ```bash ffmpeg -i vo.mp3 -i music.mp3 -filter_complex \ "[1:a][0:a]sidechaincompress=threshold=0.03:ratio=8:attack=20:release=300[duck]; \ [0:a][duck]amix=inputs=2:duration=longest[aout]" \ -map "[aout]" -c:a aac mix.m4a ``` Cheaper static fallback when sidechain is overkill — fix the music low under a full VO: ```bash ffmpeg -i vo.mp3 -i music.mp3 -filter_complex \ "[1:a]volume=0.3[m];[0:a][m]amix=inputs=2:duration=longest[aout]" \ -map "[aout]" mix.m4a ``` **(c) Loudnorm to a target LUFS (two-pass)** — measure, then apply. Target -14 LUFS for social/streaming, -16 for podcast-style VO. Normalize per track *before* mixing. ```bash # pass 1: measure (read the JSON it prints) ffmpeg -i mix.m4a -af loudnorm=I=-14:TP=-1.5:LRA=11:print_format=json -f null - # pass 2: apply with the measured values ffmpeg -i mix.m4a -af \ loudnorm=I=-14:TP=-1.5:LRA=11:measured_I=-20.1:measured_TP=-4.2:measured_LRA=6.0:measured_thresh=-30.8:offset=0.5:linear=true \ master.m4a ``` **(d) Concat scenes — conform first.** Same-codec/res/fps clips → fast demuxer with `-c copy`. Mismatched clips → re-encode and scale first, then concat. **Never `-c copy`-concat mismatched clips** — you get desync or a corrupt stream. ```bash # all clips identical codec/res/fps: printf "file '%s'\n" scene1.mp4 scene2.mp4 scene3.mp4 > list.txt ffmpeg -f concat -safe 0 -i list.txt -c copy joined.mp4 # mismatched: conform each, then concat filter ffmpeg -i s1.mp4 -i s2.mp4 -filter_complex \ "[0:v]scale=1920:1080,fps=30,setsar=1[v0];[1:v]scale=1920:1080,fps=30,setsar=1[v1]; \ [v0][1:a?][v1][1:a?]concat=n=2:v=1:a=0[v]" -map "[v]" joined.mp4 ``` ## End-to-end worked pipeline (narrated explainer) Ordered command list — generate once, assemble deterministically: 1. **Lock the plan** — scene list, aspect (e.g. 16:9 1080p 30fps), target -14 LUFS, models chosen. 2. **Stills per scene** → `../replicate-images/SKILL.md` (one prompt per scene). 3. **Clip per scene** (img→video, ~5–15 s each, model cap) via your chosen model on `fal`/`replicate`. 4. **VO** → ElevenLabs `convert(...)` at the master sample rate. 5. **Music** → ElevenLabs Music v2, length = total runtime, confirm rights. 6. **Conform** every clip to 1920x1080/30fps/SAR 1. 7. **Concat** the conformed clips → `body.mp4`. 8. **Loudnorm** VO and music tracks (two-pass) to consistent levels. 9. **Mix + duck** music under VO → `mix.m4a`. 10. **Final loudnorm** the mix to -14 LUFS → `master.m4a`. 11. **Mux** `master.m4a` onto `body.mp4` with `-shortest` → `final.mp4`. Emit this as a runnable script. `scripts/verify.sh` lints it (loudnorm present, conform-before-concat, final MP4 target). ## Cost & regen discipline - **Draft small, then final.** Generate clips short/low-res and VO on Flash to lock timing and the cut; only the final pass spends on hero quality. Re-rolling locked scenes is the biggest waste. - **Per-modality levers:** shorter scripts (TTS per-char), fewer scene re-rolls (img→video), fewer music minutes (music is billed per minute — verify the current rate on the vendor pricing page). - Spend tracking *as a discipline* → `../fal/SKILL.md` / `../replicate/SKILL.md` for per-call cost; treat budget as a constraint you set before generating. ## Anti-patterns | Anti-pattern | Why it bites | Do instead | |--------------|--------------|------------| | Generating assets before locking the pipeline | aspect/sample-rate/duration mismatches surface at mux time, forcing paid re-rolls | lock scene list, aspect, LUFS, models first | | One long video-gen call for the whole piece | models cap at ~5–15 s; you fight the limit and waste rolls | one clip per scene, then concat | | Mixing tracks without per-track `loudnorm` | VO buried or blasting over music; inconsistent levels | two-pass loudnorm each track before mix | | `-c copy`-concat of mismatched clips | desync, corrupt stream, wrong frame timing | conform res/fps/SAR, then concat | | Music low set by ear / static only when VO needs space | narration gets masked under the bed | `sidechaincompress` keyed off the VO | | Shipping generated music without checking rights | "generated" ≠ "licensed to sell" — legal exposure | confirm commercial rights per model/plan | | Prompting/editing the still inside this skill | duplicates `replicate-images`' job, worse prompts | delegate the still, consume it here | | Mastering loudness by cranking the TTS | clipping, no true-peak control | set level with `loudnorm`, not the generator |
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.