render-flat-vector-explainer
Assemble the FREE steps of the flat-vector-explainer video format — a flat-illustration creator-character walks a countable N-step product routine, one step per beat, and Remotion composites every chip/numeral/tagline/slate/CTA as an animated DOM overlay ON TOP of the Kling i2v c
Install
npx skills add https://github.com/gooseworks-ai/goose-skills/tree/main/skills/ads/capabilities/render-flat-vector-explainer
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install gooseworks-ai-goose-skills@llmmart
git clone https://github.com/gooseworks-ai/goose-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole gooseworks-ai/goose-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
render-flat-vector-explainer
Assembles a flat-vector product-routine explainer: one illustrated creator-character walks through a countable N-step routine (e.g. collagen -> serum -> eye cream -> hair), one step per beat, each beat carrying a large corner numeral, a labelled chip + one-line tagline, and the step's real product photo, closing on an "N products" grid + brand CTA. It reads as a premium DTC explainer (Spotify/Anchor flat-vector lineage), not UGC.
This capability is documentation-grade. The content-goose molecule is a documented recipe, not a runnable end-to-end app, so this capability ships the config schema (scripts/config.example.json), the field-to-script map (scripts/PIPELINE.md), and a README (scripts/README.md) describing the FREE assembly steps the agent runs by hand with ffmpeg + Remotion + PIL. The paid generative steps are separate capabilities the recipe orchestrates and gates.
The two non-negotiable separations
- Motion layer != text layer. Animate a text-stripped clean plate with Kling i2v (subtle motion, style-preserving negative, cfg 0.5), then composite every chip / numeral / tagline / slate / CTA as an animated Remotion DOM overlay on top. Baking text into the keyframe before i2v warps the type and forfeits the ability to retime/restyle it — this separation is the format's whole credibility.
- Real assets != AI assets. The per-step product photo and the closing "N products" grid are real product webps composited with PIL (AI duplicates SKUs in a grid). Only the character vignettes and stylized backgrounds are generative.
Free assembly steps (this capability)
The agent runs these deterministic, $0 steps by hand — see scripts/README.md for the ffmpeg/Remotion/PIL detail:
- Remotion overlay — import each Kling clip as the moving base; composite chips / numerals / taglines / slate / grid / CTA as animated DOM on top -> the animated silent master. Slate/grid/CTA beats are Remotion text with no i2v.
- PIL product grid — composite the N real product webps on the brand ground for the closing lockup; preserve each aspect (never stretch, never AI-dupe).
- Captions — word-by-word burned from the eleven_v3 with-timestamps char timings (libass); suppress on slate/grid/CTA scenes so two text layers don't collide.
- Audio mix + master — place each VO line at its scene start, duck the music under VO (sidechaincompress),
loudnorm I=-15VO-forward, mux, burn captions LAST ->finals/master-final.mp4(~50s). - 30s cut — slice each beat's region OUT of the animated silent master (never a static intermediate); trim short beats, gently slow long beats (setpts <=1.6x), re-burn scaled captions ->
finals/master-final-30s-v1.mp4.
Paid gen steps (separate capabilities)
The recipe orchestrates and gates these; they are not part of this capability:
- Flat-vector character anchor + per-scene keyframes + clean plates ->
create-image-fal(nano-banana; re-render a FRESH flat-vector anchor, never chain a photoreal ref). - Kling i2v on the character scenes ->
create-video-fal(Kling 2.5-turbo/pro, cfg 0.5, style-preserving negative, low motion; TEST one scene before batching). - Full-sentence VO ->
create-vo-elevenlabs(eleven_v3, with-timestamps). - Lo-fi music bed ->
create-music-elevenlabs.
Contract
- Documentation-grade + FREE assembly (Remotion + PIL + FFmpeg); no paid calls in this capability, no AI-rendered text.
- Text is an overlay, never baked. Strip to a clean plate -> i2v -> composite text as Remotion DOM.
- Any multi-SKU grid is PIL of the real product webps; preserve each aspect ratio.
- Kling holds the 2D flat-vector look only at LOW motion (cfg 0.5 + style-preserving negative). Aggressive motion drifts to photoreal.
- Cut down from the ANIMATED master, never a static intermediate; frame-diff to prove localized motion.
- The paid steps — keyframes/clean plates, Kling i2v, VO, music — are separate capabilities (create-image-fal, create-video-fal, create-vo-elevenlabs, create-music-elevenlabs); the recipe orchestrates them and gates the spend.
Files (goose-skills)
-
scripts
-
config.example.json 9.9 KB
{ "_comment": "Spoiled Child 'The Perfect Morning Routine = 4 Products' \u2014 the worked example (v3 shipped recipe). Copy to config.json and edit. A flat-vector creator-character walks a countable 4-step routine, one step per beat: large corner numeral + labelled chip + tagline + the step's REAL product photo. Character scenes are Kling i2v off text-stripped clean plates (subtle motion, style-preserving negative); ALL text is an animated Remotion DOM overlay (NEVER baked); slate/grid/CTA scenes get no i2v; the closing grid is a PIL composite of real product webps. eleven_v3 full-sentence VO drives word-by-word burned captions over a VO-forward music bed. Built as a ~50s animated master, then re-cut to 30s FROM the animated master. Values lifted from the source design-brief scene table + meta.json. The runnable scripts live in the source project's working/ (see PIPELINE.md).", "brand_name": "Spoiled Child", "concept": { "single_point": "The perfect morning routine is just 4 products (collagen drink, face serum, eye cream, hair treatment) \u2014 not twenty.", "motif_phrase": "THE PERFECT ROUTINE = 4 PRODUCTS", "closing_phrase": "4 is enough.", "hook_line": "Your morning routine has twenty different products. Why?", "concept_label": "morning routine (NOT 'skincare' \u2014 step 4 is a hair treatment)" }, "width": 1080, "height": 1920, "fps": 30, "master_duration_sec": 50.0, "deliverable_duration_sec": 30.0, "brand_palette": [ "#FAFAFA", "#C96E2F", "#F2C9C2", "#0A0B0D", "#88B6A4", "#B3A7C4" ], "display_font": "Inter", "character": { "mode": "anchor-ref (fresh flat-vector \u2014 NOT a chained photoreal ref)", "anchor_prompt": "Flat-vector illustration of a woman, late 20s, soft brown wavy shoulder-length hair, warm-medium skin tone, cream slip-dress. Clean 2D flat-vector style (Spotify/Anchor lineage), no gradients, no photoreal shading. Consistent character across every scene.", "anchor_path": "assets/character-lock/creator-anchor.png", "style_reference_note": "The brand's existing character (clients/spoiled-child/shared/characters/anchor/character-anchor-05-home-aesthetic.png) is photoreal \u2014 use it ONLY as a written descriptor source, NEVER as a `medias` chained ref (it flattens gens to photoreal).", "keyframe_variants": [ "counter-overwhelmed", "smile-shrug", "lifting-spoon", "serum-pump", "under-eye-dab", "hands-through-hair", "mirror-satisfied" ] }, "kling": { "model": "fal-ai/kling-video/v2.5-turbo/pro/image-to-video", "duration_sec": 5, "cfg_scale": 0.5, "negative_prompt": "photorealistic, 3D render, realistic skin, style change, gradient shading, text warp, camera shake", "_note": "Holds flat-vector style at LOW motion only. TEST one scene before batching." }, "scenes": [ { "n": 1, "kind": "character", "duration_sec": 4.0, "motion": "gentle overwhelmed breathing + one blink", "keyframe_prompt": "Creator at a bathroom counter overstuffed with 20+ jars/bottles/tubes from many brands; faintly comedic, overwhelmed.", "overlay": null, "vo": "Your morning routine has twenty different products." }, { "n": 2, "kind": "character", "duration_sec": 2.5, "motion": "small natural smile + shrug; clutter settles", "keyframe_prompt": "Same creator, small smile + shrug at the cluttered counter.", "overlay": null, "vo": "Why?" }, { "n": 3, "kind": "slate", "duration_sec": 3.5, "motion": "Remotion slate slam-in", "keyframe_prompt": null, "overlay": { "headline": "THE PERFECT ROUTINE = 4 PRODUCTS", "ground": "#C96E2F" }, "vo": "The perfect routine is just four products." }, { "n": 4, "kind": "character", "duration_sec": 6.5, "motion": "lifts a tablespoon of amber liquid toward camera; minimal drift", "keyframe_prompt": "Creator lifts a tablespoon of amber liquid E27 collagen; teal accent panel.", "overlay": { "numeral": "1", "chip": "01 \u00b7 COLLAGEN DRINK", "tagline": "Liquid collagen for glowing skin from within.", "product_photo": "assets/products/e27-main-bottle.webp", "accent": "#88B6A4" }, "vo": "One \u2014 liquid collagen. A spoonful at breakfast for skin from within." }, { "n": 5, "kind": "character", "duration_sec": 6.5, "motion": "presses serum pump near cheek", "keyframe_prompt": "Creator presses S33 serum pump near cheek; peach accent panel.", "overlay": { "numeral": "2", "chip": "02 \u00b7 FACE SERUM", "tagline": "Vitamin-C serum for that morning glow that makes you feel radiant.", "product_photo": "assets/products/s33-bottle-close.webp", "accent": "#F2C9C2" }, "vo": "Two \u2014 a vitamin-C serum for that morning glow that makes you feel radiant." }, { "n": 6, "kind": "character", "duration_sec": 5.5, "motion": "dabs eye cream under-eye", "keyframe_prompt": "Creator dabs eye cream under-eye; lavender accent panel.", "overlay": { "numeral": "3", "chip": "03 \u00b7 EYE CREAM", "tagline": "An eye cream \u2014 your eyes will de-puff in minutes.", "product_photo": "assets/products/t31-main-product.webp", "accent": "#B3A7C4" }, "vo": "Three \u2014 an eye cream, and your eyes will de-puff in minutes." }, { "n": 7, "kind": "character", "duration_sec": 6.0, "motion": "runs hands through hair, soft steam", "keyframe_prompt": "Creator runs hands through hair, soft steam; warm rust accent panel.", "overlay": { "numeral": "4", "chip": "04 \u00b7 HAIR TREATMENT", "tagline": "A rinse-in treatment that leaves your hair shining all day.", "product_photo": "assets/products/h30-main-product.webp", "accent": "#C96E2F" }, "vo": "Four \u2014 a rinse-in hair treatment that leaves your hair shining all day." }, { "n": 8, "kind": "grid", "duration_sec": 3.0, "motion": "PIL grid + Remotion slate", "keyframe_prompt": null, "overlay": { "headline": "4 is enough.", "ground": "#C96E2F" }, "vo": "Four products. That's it." }, { "n": 9, "kind": "character", "duration_sec": 5.0, "motion": "satisfied smile, holding 1 product; clean bathroom (callback to scene 1)", "keyframe_prompt": "Creator at mirror, satisfied smile, holding one product; clean bathroom.", "overlay": null, "vo": "Skip the twenty-step routine. Build the four-product one." }, { "n": 10, "kind": "cta", "duration_sec": 8.0, "motion": "Remotion CTA", "keyframe_prompt": null, "overlay": { "wordmark": "Spoiled Child", "cta_pill": "Build your morning routine \u2192", "url": "spoiledchild.com", "ground": "#C96E2F" }, "vo": "Spoiled Child. Your morning routine, simplified. Spoiledchild dot com." } ], "product_grid": { "method": "PIL composite of the 4 REAL product webps (never AI \u2014 AI duplicates SKUs)", "ground": "#C96E2F", "layout": "2x2", "preserve_aspect": true, "images": [ "assets/products/e27-main-bottle.webp", "assets/products/s33-bottle-close.webp", "assets/products/t31-main-product.webp", "assets/products/h30-main-product.webp" ] }, "voice": { "engine": "ElevenLabs", "model": "eleven_v3", "endpoint": "text-to-speech/with-timestamps", "voice_chosen": "Eryn", "voice_id": "dMyQqiVXTU80dDl2eNK8", "casting_ab": [ "Eryn (dMyQqiVXTU80dDl2eNK8)", "Angela (FUfBrNit0NNZAwb58KWH)" ], "settings": { "stability": 0.45, "similarity_boost": 0.8, "style": 0.1, "use_speaker_boost": true, "speed": 1.1 }, "_note": "Write VO as FULL SENTENCES (not keyword fragments). with-timestamps char-level timings drive the word-by-word captions." }, "music": { "engine": "ElevenLabs", "prompt": "Lo-fi pop bed, 95-105 BPM, builds across sections (intro pad -> groove -> warm -> lo-fi build -> resolved). NOT cinematic, NOT moody. No melody hooks competing with VO. Instrumental only.", "length_ms": 52000, "force_instrumental": true, "mix_profile": "VO-forward: vo loudnorm I=-15, music loudnorm I=-25 + volume 0.5, sidechaincompress duck, final alimiter limit=0.89, target ~-15 LUFS / -1 dBTP" }, "captions": { "method": "word-by-word burned from eleven_v3 with-timestamps (libass; Klap is the hosted equivalent)", "style": "Inter/Arial 56 white + 4px outline, bottom-third (MarginV 360)", "burned_last": true, "suppress_on_scenes": [ 3, 8, 10 ], "_suppress_note": "slate/grid/CTA scenes carry their own on-screen text \u2014 two text layers collide." }, "cutdowns": { "source": "working/silent-master.mp4 (the ANIMATED master \u2014 NEVER a static intermediate)", "rules": "trim short beats (never freeze); gently slow long beats (setpts, clamp <=1.6x); re-burn scaled captions (timestamps / vo_speedup, offset to new scene starts); suppress on 3/8/10", "variants": { "v1": "full 10-beat story, VO @1.25x (atempo), silences trimmed, ~30s \u2014 THE SHIPPED DELIVERABLE", "v2": "hook-led montage: hook -> slate -> grid pulled forward -> rapid 4-product montage (number + product only) -> CTA, ~30s (>=4 cuts/10s)" } }, "post_production": { "music": { "default": "on", "note": "lo-fi pop bed under the VO, default on" }, "captions": { "default": "on", "note": "word-by-word burned captions (libass/Klap), default on" }, "end_card": { "default": "on", "note": "CTA + brand wordmark scene, default on" } } } -
PIPELINE.md 4 KB
# PIPELINE — flat-vector-explainer engine map This molecule is **documentation-grade**. Rather than re-implement a 14-scene Remotion app here, it maps every `config.json` field to the **real, runnable script** that produced the worked example. The executable reference is the source project: ``` clients/spoiled-child/video-11-routine-broken/ ``` Its `HOW_TO.md` + `LEARNINGS.md` are the authoritative v3 recipe (Kling i2v + Remotion overlay). Run the pipeline there (or port these scripts into a new brand's project folder), then bring the config here as the recipe of record. ## Config field → source script | `config.json` field | Source script | Phase | Paid? | |---|---|---|---| | `character.anchor_prompt`, `character.keyframe_variants`, `scenes[].keyframe_prompt` | `working/gen_keyframes.py` | 1 — character lock + per-scene keyframes (nano-banana, flat-vector; re-render a FRESH anchor — never chain a photoreal ref) | **PAID** | | (clean plates — strip baked text) | `working/clean_plate.py` | 2 — nano-banana edit removes chips/numerals/taglines/badges from character keyframes → clean plates for i2v | **PAID** | | `scenes[].motion`, `kling.*` | `working/kling_i2v.py` | 3 — Kling 2.5-turbo/pro i2v on character scenes only; cfg 0.5 + style-preserving negative + subtle motion; TEST one scene first | **PAID** | | `scenes[].overlay` (chips, numerals, taglines, slate, CTA) | `working/remotion/` (Remotion project) | 4 — imports each Kling clip as the moving base and composites ALL text as animated DOM ON TOP → animated silent master. Text is NEVER baked into a keyframe | free | | `product_grid.*` | `working/scripts/build_scene08.py` | 4 — PIL composite of the N real product webps on the brand ground (preserve each aspect; AI duplicates SKUs so this is PIL, not AI) | free | | `voice.*`, `scenes[].vo` | `working/scripts/render_vo.py` | 5 — ElevenLabs eleven_v3, `text-to-speech/with-timestamps`; full-sentence per-scene VO + char-level timestamps | **PAID** | | `music.*` | `working/gen_music.py` | 5 — ElevenLabs music bed, lo-fi pop, VO-forward loudnorm | **PAID** | | (VO+music mix) | `working/scripts/mix_audio.sh` | 5 — place each VO line at its scene start, duck music under VO (sidechaincompress), `loudnorm I=-15`, VO-forward | free | | `captions.*` | `working/scripts/build_captions.py` | 5 — word-by-word burned captions from the VO char-timestamps (libass); suppress on slate/grid/CTA scenes | free | | (silent master assembly) | `working/scripts/build_silent.sh` | 4 — stitch the Remotion-composited scenes into `working/silent-master.mp4` (the ANIMATED master) | free | | (50s master) | `working/build_master.py` | 5 — mux silent master + mixed audio, burn captions LAST → `finals/master-final.mp4` | free | | `cutdowns.*` | `working/build_30s.py` | 6 — slice each beat's region OUT of `silent-master.mp4` (the ANIMATED master, NEVER a static intermediate); trim short / slow long (≤1.6×); re-burn scaled captions → `finals/master-final-30s-v1.mp4` | free | | (QC) | `working/motion_probe.py` | 7 — frame-diff proof of localized motion (+ `/watch` the final) | free | ## The two non-negotiable separations 1. **Motion layer ≠ text layer.** `clean_plate.py` strips text → `kling_i2v.py` animates the clean plate → `remotion/` composites text as DOM on top. Never `kling_i2v.py` a keyframe with baked text (i2v warps it, and you can't retime/restyle it). 2. **Real assets ≠ AI assets.** `build_scene08.py` PIL-composites the real product webps. AI gen duplicates SKUs in a grid — it is only for the character vignettes + backgrounds. ## The cut-down trap (LEARNINGS L6) `build_30s.py` MUST slice from `working/silent-master.mp4` (the Kling-animated master), NOT from `working/segs/` (an earlier static Ken-Burns round). The mtime is the tell. A finished-looking audio mix can hide frozen characters — always run `motion_probe.py` frame-diff on a character scene and confirm **localized** face/hand glow (whole-outline glow = pan-only = wrong source). -
README.md 5.4 KB
# render-flat-vector-explainer — free assembly how-to This capability is **documentation-grade**: the flat-vector-explainer molecule is a documented recipe, not a runnable end-to-end app. There are no standalone scripts here to `--config` and fire. Instead this folder ships: - **`config.example.json`** — the full config schema (the Spoiled Child "Perfect Morning Routine = 4 Products" worked example). Copy to `config.json` and edit. - **`PIPELINE.md`** — the field-by-field map from every `config.json` key to the real, runnable source script that produced the worked example (in the source project's `working/`), and which phase / paid-or-free it is. - **this README** — the FREE assembly steps the agent runs by hand with ffmpeg + Remotion + PIL after the paid gen steps have produced the character clips, VO, and music. The paid generative steps (keyframes, clean plates, Kling i2v, VO, music) are **separate capabilities** the recipe orchestrates and gates — see the recipe. Everything below is free, deterministic, $0. ## The two separations that make this format work 1. **Motion layer != text layer.** i2v only ever animates a **text-stripped clean plate**. ALL words/numerals/chips/taglines/slate/CTA are composited on top as an **animated Remotion DOM overlay** — never baked into the keyframe. Baked text warps under i2v and you lose the ability to retime/restyle it. 2. **Real assets != AI assets.** The per-step product photo and the closing "N products" grid are the **real product webps composited with PIL**. AI duplicates SKUs in a grid. Only the character vignettes + backgrounds are generative. ## Free assembly steps Run these after the paid steps have delivered the per-scene Kling clips (character beats), the VO (with char-level timestamps), and the music bed. ### 1. Remotion overlay -> animated silent master (character-locked) Build a Remotion composition that imports each Kling clip as the moving base and composites the scene's `overlay` (numeral, chip, tagline, slate headline, CTA pill, wordmark, url) as **animated DOM on top**. Keep the **motion layer and text layer strictly separate** — the Kling footage is the only moving base; text is DOM. - `kind: character` scenes -> Kling clip base + DOM overlay. - `kind: slate` / `grid` / `cta` scenes -> Remotion text only, **no i2v** (a solid brand `ground` color + the headline / grid / CTA). - Use the brand palette + display font from the config for crisp glyphs at the exact brand color. Real DOM type = crisp; never let i2v render the type. Render the composition to `working/silent-master.mp4` — the **animated** master. Everything downstream slices from this, never from an earlier static intermediate. ### 2. PIL real-product grid Composite the closing "N products" lockup from `product_grid.images` (the real product webps) onto `product_grid.ground` in the configured `layout` (e.g. 2x2). **Preserve each product's aspect ratio — never stretch, never AI-generate the grid** (AI dupes SKUs). The same real webps are also the per-step callout photos in the character beats. Product webps are git-LFS in the brand folder — fetch + checkout first (pointers are ~131-byte stubs). ### 3. Caption burn (word-synced) Build word-by-word burned captions from the eleven_v3 with-timestamps **char-level timings** (libass; Klap is the hosted equivalent). Style per `captions.style` (e.g. Inter 56 white + 4px outline, bottom-third MarginV 360). **Suppress captions on the slate/grid/CTA scenes** (`captions.suppress_on_scenes`) — those carry their own on-screen text and two text layers collide. Burn captions **LAST**, after the audio mux. ### 4. Audio mix + 50s master Place each scene's VO line at its scene start; duck the music under the VO with `sidechaincompress`; `loudnorm I=-15` VO-forward (music `loudnorm I=-25 + volume 0.5`, final `alimiter limit=0.89`, ~-15 LUFS / -1 dBTP). Mux the silent master + mixed audio, then burn captions last -> `finals/master-final.mp4` (~50s full-story master). ### 5. 50s -> 30s cut (from the ANIMATED master) Slice each beat's region OUT of `working/silent-master.mp4` — the **animated** master, **never** an earlier static Ken-Burns / segs round (the mtime is the tell; a finished audio mix can hide frozen characters). Trim short beats (never freeze a frame); gently slow long beats (`setpts`, clamp <=1.6x) so motion stretches instead of holding. Re-burn scaled captions offset to the new scene starts; keep suppression on the slate/grid/CTA scenes. Output `finals/master-final-30s-v1.mp4` — the shipped 30s deliverable. Optional v2 cutdown: hook-led montage (hook -> slate -> grid pulled forward -> rapid 4-product montage of number + product only -> CTA), >=4 cuts / 10s. ## QC before ship (mandatory) - **Frame-diff a character scene** (blend=difference heatmap): **localized** face/hand glow = real motion; whole-outline glow = a static pan = you sliced the wrong (static) source — re-point the cut at `silent-master.mp4`. - `/watch` the final: real localized motion, correct duration (~30s ±0.3s), caption sync, correct numerals/labels, no AI text leak in the i2v footage, N distinct real products in the grid (no AI dupes), VO-forward mix (~-15 LUFS). - Canvas 1080x1920, 30fps, h264, aac present. ## Requires Node + Remotion, `Pillow` (PIL), `ffmpeg` with libass. All free — the paid keys (`FAL_API_KEY`, `ELEVENLABS_API_KEY`) belong to the separate paid capabilities, which route through the GooseWorks proxies so the calls bill the Ads agent.
-
-
tests
-
smoke-test.md 1.2 KB
# Smoke Test This capability is documentation-grade — it ships the config schema + assembly recipe, not a runnable end-to-end script. The check is that the docs + config are present and coherent. Pass when: - `scripts/config.example.json` parses as valid JSON and carries the flat-vector-explainer schema (concept + single countable point, character anchor, ordered `scenes[]` with `kind` in character/slate/grid/cta, `product_grid` of real webps, voice, music, captions with `suppress_on_scenes`, cutdowns, post_production toggles). - `scripts/PIPELINE.md` maps every config field to its source script + phase + paid/free. - `scripts/README.md` documents the FREE assembly steps: Remotion text/DOM overlay kept SEPARATE from the Kling i2v motion layer (character-locked), PIL real-product grid, word-synced caption burn, VO-forward audio mix, and the 50s -> 30s cut taken FROM the animated silent master (never a static intermediate). - The SKILL.md description is a single line with no ": " and states the paid gen steps (keyframes, Kling i2v, VO, music) are separate capabilities the recipe orchestrates. - No paid call is made by this capability. Paid caps route through the GooseWorks proxies (bill the agent, no direct provider host).
-
-
SKILL.md 5 KB
--- name: render-flat-vector-explainer description: Assemble the FREE steps of the flat-vector-explainer video format — a flat-illustration creator-character walks a countable N-step product routine, one step per beat, and Remotion composites every chip/numeral/tagline/slate/CTA as an animated DOM overlay ON TOP of the Kling i2v character clips (text is NEVER baked into a keyframe — i2v warps type), the closing 'N products' grid is a PIL composite of the REAL product photos (not AI), full-sentence VO drives word-by-word burned captions over a VO-forward music bed, and the ~50s animated silent master is re-cut to a 30s deliverable FROM the animated master (never a static intermediate). Documentation-grade — ships config.example.json + PIPELINE.md + a README of the free assembly; the paid gen steps (keyframes, Kling i2v, VO, music) are separate capabilities the recipe orchestrates. Use for the flat-vector-explainer format. status: active --- # render-flat-vector-explainer Assembles a flat-vector product-routine explainer: one illustrated creator-character walks through a countable N-step routine (e.g. collagen -> serum -> eye cream -> hair), one step per beat, each beat carrying a large corner numeral, a labelled chip + one-line tagline, and the step's real product photo, closing on an "N products" grid + brand CTA. It reads as a premium DTC explainer (Spotify/Anchor flat-vector lineage), not UGC. This capability is **documentation-grade**. The content-goose molecule is a documented recipe, not a runnable end-to-end app, so this capability ships the **config schema** (`scripts/config.example.json`), the **field-to-script map** (`scripts/PIPELINE.md`), and a **README** (`scripts/README.md`) describing the FREE assembly steps the agent runs by hand with ffmpeg + Remotion + PIL. The paid generative steps are separate capabilities the recipe orchestrates and gates. ## The two non-negotiable separations 1. **Motion layer != text layer.** Animate a **text-stripped clean plate** with Kling i2v (subtle motion, style-preserving negative, cfg 0.5), then composite every chip / numeral / tagline / slate / CTA as an **animated Remotion DOM overlay** on top. Baking text into the keyframe before i2v warps the type and forfeits the ability to retime/restyle it — this separation is the format's whole credibility. 2. **Real assets != AI assets.** The per-step product photo and the closing "N products" grid are **real product webps composited with PIL** (AI duplicates SKUs in a grid). Only the character vignettes and stylized backgrounds are generative. ## Free assembly steps (this capability) The agent runs these deterministic, $0 steps by hand — see `scripts/README.md` for the ffmpeg/Remotion/PIL detail: - **Remotion overlay** — import each Kling clip as the moving base; composite chips / numerals / taglines / slate / grid / CTA as animated DOM on top -> the animated silent master. Slate/grid/CTA beats are Remotion text with no i2v. - **PIL product grid** — composite the N real product webps on the brand ground for the closing lockup; preserve each aspect (never stretch, never AI-dupe). - **Captions** — word-by-word burned from the eleven_v3 with-timestamps char timings (libass); suppress on slate/grid/CTA scenes so two text layers don't collide. - **Audio mix + master** — place each VO line at its scene start, duck the music under VO (sidechaincompress), `loudnorm I=-15` VO-forward, mux, burn captions LAST -> `finals/master-final.mp4` (~50s). - **30s cut** — slice each beat's region OUT of the **animated silent master** (never a static intermediate); trim short beats, gently slow long beats (setpts <=1.6x), re-burn scaled captions -> `finals/master-final-30s-v1.mp4`. ## Paid gen steps (separate capabilities) The recipe orchestrates and gates these; they are not part of this capability: - Flat-vector character anchor + per-scene keyframes + clean plates -> `create-image-fal` (nano-banana; re-render a FRESH flat-vector anchor, never chain a photoreal ref). - Kling i2v on the character scenes -> `create-video-fal` (Kling 2.5-turbo/pro, cfg 0.5, style-preserving negative, low motion; TEST one scene before batching). - Full-sentence VO -> `create-vo-elevenlabs` (eleven_v3, with-timestamps). - Lo-fi music bed -> `create-music-elevenlabs`. ## Contract - Documentation-grade + FREE assembly (Remotion + PIL + FFmpeg); no paid calls in this capability, no AI-rendered text. - Text is an overlay, never baked. Strip to a clean plate -> i2v -> composite text as Remotion DOM. - Any multi-SKU grid is PIL of the real product webps; preserve each aspect ratio. - Kling holds the 2D flat-vector look only at LOW motion (cfg 0.5 + style-preserving negative). Aggressive motion drifts to photoreal. - Cut down from the ANIMATED master, never a static intermediate; frame-diff to prove localized motion. - The paid steps — keyframes/clean plates, Kling i2v, VO, music — are separate capabilities (create-image-fal, create-video-fal, create-vo-elevenlabs, create-music-elevenlabs); the recipe orchestrates them and gates the spend. -
skill.meta.json 333 B
{ "slug": "render-flat-vector-explainer", "category": "capabilities", "domain": "ads", "tags": [ "ads" ], "installation": { "base_command": "npx goose-skills install render-flat-vector-explainer", "supports": [ "claude", "cursor", "codex" ] }, "requires_skills": [ "watch" ] }
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.