Claude Skill

render-flat-vector-explainer

Assemble the FREE steps of the flat-vector-explainer video format — a flat-illustration creator-character walks a countable N-step product routine, one step per beat, and Remotion composites every chip/numeral/tagline/slate/CTA as an animated DOM overlay ON TOP of the Kling i2v c

LLM Mart · 0 points · 5 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download gooseworks-ai-goose-skills-skills_ads_capabilities_render-flat-vector-explainer-e1592ee.zip · 11 KB
Part of gooseworks-ai/goose-skills — 44 skills

Install

skills CLI npx skills add https://github.com/gooseworks-ai/goose-skills/tree/main/skills/ads/capabilities/render-flat-vector-explainer
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install gooseworks-ai-goose-skills@llmmart
Git git clone https://github.com/gooseworks-ai/goose-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole gooseworks-ai/goose-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

render-flat-vector-explainer

Assembles a flat-vector product-routine explainer: one illustrated creator-character walks through a countable N-step routine (e.g. collagen -> serum -> eye cream -> hair), one step per beat, each beat carrying a large corner numeral, a labelled chip + one-line tagline, and the step's real product photo, closing on an "N products" grid + brand CTA. It reads as a premium DTC explainer (Spotify/Anchor flat-vector lineage), not UGC.

This capability is documentation-grade. The content-goose molecule is a documented recipe, not a runnable end-to-end app, so this capability ships the config schema (scripts/config.example.json), the field-to-script map (scripts/PIPELINE.md), and a README (scripts/README.md) describing the FREE assembly steps the agent runs by hand with ffmpeg + Remotion + PIL. The paid generative steps are separate capabilities the recipe orchestrates and gates.

The two non-negotiable separations

  1. Motion layer != text layer. Animate a text-stripped clean plate with Kling i2v (subtle motion, style-preserving negative, cfg 0.5), then composite every chip / numeral / tagline / slate / CTA as an animated Remotion DOM overlay on top. Baking text into the keyframe before i2v warps the type and forfeits the ability to retime/restyle it — this separation is the format's whole credibility.
  2. Real assets != AI assets. The per-step product photo and the closing "N products" grid are real product webps composited with PIL (AI duplicates SKUs in a grid). Only the character vignettes and stylized backgrounds are generative.

Free assembly steps (this capability)

The agent runs these deterministic, $0 steps by hand — see scripts/README.md for the ffmpeg/Remotion/PIL detail:

  • Remotion overlay — import each Kling clip as the moving base; composite chips / numerals / taglines / slate / grid / CTA as animated DOM on top -> the animated silent master. Slate/grid/CTA beats are Remotion text with no i2v.
  • PIL product grid — composite the N real product webps on the brand ground for the closing lockup; preserve each aspect (never stretch, never AI-dupe).
  • Captions — word-by-word burned from the eleven_v3 with-timestamps char timings (libass); suppress on slate/grid/CTA scenes so two text layers don't collide.
  • Audio mix + master — place each VO line at its scene start, duck the music under VO (sidechaincompress), loudnorm I=-15 VO-forward, mux, burn captions LAST -> finals/master-final.mp4 (~50s).
  • 30s cut — slice each beat's region OUT of the animated silent master (never a static intermediate); trim short beats, gently slow long beats (setpts <=1.6x), re-burn scaled captions -> finals/master-final-30s-v1.mp4.

Paid gen steps (separate capabilities)

The recipe orchestrates and gates these; they are not part of this capability:

  • Flat-vector character anchor + per-scene keyframes + clean plates -> create-image-fal (nano-banana; re-render a FRESH flat-vector anchor, never chain a photoreal ref).
  • Kling i2v on the character scenes -> create-video-fal (Kling 2.5-turbo/pro, cfg 0.5, style-preserving negative, low motion; TEST one scene before batching).
  • Full-sentence VO -> create-vo-elevenlabs (eleven_v3, with-timestamps).
  • Lo-fi music bed -> create-music-elevenlabs.

Contract

  • Documentation-grade + FREE assembly (Remotion + PIL + FFmpeg); no paid calls in this capability, no AI-rendered text.
  • Text is an overlay, never baked. Strip to a clean plate -> i2v -> composite text as Remotion DOM.
  • Any multi-SKU grid is PIL of the real product webps; preserve each aspect ratio.
  • Kling holds the 2D flat-vector look only at LOW motion (cfg 0.5 + style-preserving negative). Aggressive motion drifts to photoreal.
  • Cut down from the ANIMATED master, never a static intermediate; frame-diff to prove localized motion.
  • The paid steps — keyframes/clean plates, Kling i2v, VO, music — are separate capabilities (create-image-fal, create-video-fal, create-vo-elevenlabs, create-music-elevenlabs); the recipe orchestrates them and gates the spend.
Files (goose-skills)
  • scripts
    • config.example.json 9.9 KB
      {
        "_comment": "Spoiled Child 'The Perfect Morning Routine = 4 Products' \u2014 the worked example (v3 shipped recipe). Copy to config.json and edit. A flat-vector creator-character walks a countable 4-step routine, one step per beat: large corner numeral + labelled chip + tagline + the step's REAL product photo. Character scenes are Kling i2v off text-stripped clean plates (subtle motion, style-preserving negative); ALL text is an animated Remotion DOM overlay (NEVER baked); slate/grid/CTA scenes get no i2v; the closing grid is a PIL composite of real product webps. eleven_v3 full-sentence VO drives word-by-word burned captions over a VO-forward music bed. Built as a ~50s animated master, then re-cut to 30s FROM the animated master. Values lifted from the source design-brief scene table + meta.json. The runnable scripts live in the source project's working/ (see PIPELINE.md).",
        "brand_name": "Spoiled Child",
        "concept": {
          "single_point": "The perfect morning routine is just 4 products (collagen drink, face serum, eye cream, hair treatment) \u2014 not twenty.",
          "motif_phrase": "THE PERFECT ROUTINE = 4 PRODUCTS",
          "closing_phrase": "4 is enough.",
          "hook_line": "Your morning routine has twenty different products. Why?",
          "concept_label": "morning routine (NOT 'skincare' \u2014 step 4 is a hair treatment)"
        },
        "width": 1080,
        "height": 1920,
        "fps": 30,
        "master_duration_sec": 50.0,
        "deliverable_duration_sec": 30.0,
        "brand_palette": [
          "#FAFAFA",
          "#C96E2F",
          "#F2C9C2",
          "#0A0B0D",
          "#88B6A4",
          "#B3A7C4"
        ],
        "display_font": "Inter",
        "character": {
          "mode": "anchor-ref (fresh flat-vector \u2014 NOT a chained photoreal ref)",
          "anchor_prompt": "Flat-vector illustration of a woman, late 20s, soft brown wavy shoulder-length hair, warm-medium skin tone, cream slip-dress. Clean 2D flat-vector style (Spotify/Anchor lineage), no gradients, no photoreal shading. Consistent character across every scene.",
          "anchor_path": "assets/character-lock/creator-anchor.png",
          "style_reference_note": "The brand's existing character (clients/spoiled-child/shared/characters/anchor/character-anchor-05-home-aesthetic.png) is photoreal \u2014 use it ONLY as a written descriptor source, NEVER as a `medias` chained ref (it flattens gens to photoreal).",
          "keyframe_variants": [
            "counter-overwhelmed",
            "smile-shrug",
            "lifting-spoon",
            "serum-pump",
            "under-eye-dab",
            "hands-through-hair",
            "mirror-satisfied"
          ]
        },
        "kling": {
          "model": "fal-ai/kling-video/v2.5-turbo/pro/image-to-video",
          "duration_sec": 5,
          "cfg_scale": 0.5,
          "negative_prompt": "photorealistic, 3D render, realistic skin, style change, gradient shading, text warp, camera shake",
          "_note": "Holds flat-vector style at LOW motion only. TEST one scene before batching."
        },
        "scenes": [
          {
            "n": 1,
            "kind": "character",
            "duration_sec": 4.0,
            "motion": "gentle overwhelmed breathing + one blink",
            "keyframe_prompt": "Creator at a bathroom counter overstuffed with 20+ jars/bottles/tubes from many brands; faintly comedic, overwhelmed.",
            "overlay": null,
            "vo": "Your morning routine has twenty different products."
          },
          {
            "n": 2,
            "kind": "character",
            "duration_sec": 2.5,
            "motion": "small natural smile + shrug; clutter settles",
            "keyframe_prompt": "Same creator, small smile + shrug at the cluttered counter.",
            "overlay": null,
            "vo": "Why?"
          },
          {
            "n": 3,
            "kind": "slate",
            "duration_sec": 3.5,
            "motion": "Remotion slate slam-in",
            "keyframe_prompt": null,
            "overlay": {
              "headline": "THE PERFECT ROUTINE = 4 PRODUCTS",
              "ground": "#C96E2F"
            },
            "vo": "The perfect routine is just four products."
          },
          {
            "n": 4,
            "kind": "character",
            "duration_sec": 6.5,
            "motion": "lifts a tablespoon of amber liquid toward camera; minimal drift",
            "keyframe_prompt": "Creator lifts a tablespoon of amber liquid E27 collagen; teal accent panel.",
            "overlay": {
              "numeral": "1",
              "chip": "01 \u00b7 COLLAGEN DRINK",
              "tagline": "Liquid collagen for glowing skin from within.",
              "product_photo": "assets/products/e27-main-bottle.webp",
              "accent": "#88B6A4"
            },
            "vo": "One \u2014 liquid collagen. A spoonful at breakfast for skin from within."
          },
          {
            "n": 5,
            "kind": "character",
            "duration_sec": 6.5,
            "motion": "presses serum pump near cheek",
            "keyframe_prompt": "Creator presses S33 serum pump near cheek; peach accent panel.",
            "overlay": {
              "numeral": "2",
              "chip": "02 \u00b7 FACE SERUM",
              "tagline": "Vitamin-C serum for that morning glow that makes you feel radiant.",
              "product_photo": "assets/products/s33-bottle-close.webp",
              "accent": "#F2C9C2"
            },
            "vo": "Two \u2014 a vitamin-C serum for that morning glow that makes you feel radiant."
          },
          {
            "n": 6,
            "kind": "character",
            "duration_sec": 5.5,
            "motion": "dabs eye cream under-eye",
            "keyframe_prompt": "Creator dabs eye cream under-eye; lavender accent panel.",
            "overlay": {
              "numeral": "3",
              "chip": "03 \u00b7 EYE CREAM",
              "tagline": "An eye cream \u2014 your eyes will de-puff in minutes.",
              "product_photo": "assets/products/t31-main-product.webp",
              "accent": "#B3A7C4"
            },
            "vo": "Three \u2014 an eye cream, and your eyes will de-puff in minutes."
          },
          {
            "n": 7,
            "kind": "character",
            "duration_sec": 6.0,
            "motion": "runs hands through hair, soft steam",
            "keyframe_prompt": "Creator runs hands through hair, soft steam; warm rust accent panel.",
            "overlay": {
              "numeral": "4",
              "chip": "04 \u00b7 HAIR TREATMENT",
              "tagline": "A rinse-in treatment that leaves your hair shining all day.",
              "product_photo": "assets/products/h30-main-product.webp",
              "accent": "#C96E2F"
            },
            "vo": "Four \u2014 a rinse-in hair treatment that leaves your hair shining all day."
          },
          {
            "n": 8,
            "kind": "grid",
            "duration_sec": 3.0,
            "motion": "PIL grid + Remotion slate",
            "keyframe_prompt": null,
            "overlay": {
              "headline": "4 is enough.",
              "ground": "#C96E2F"
            },
            "vo": "Four products. That's it."
          },
          {
            "n": 9,
            "kind": "character",
            "duration_sec": 5.0,
            "motion": "satisfied smile, holding 1 product; clean bathroom (callback to scene 1)",
            "keyframe_prompt": "Creator at mirror, satisfied smile, holding one product; clean bathroom.",
            "overlay": null,
            "vo": "Skip the twenty-step routine. Build the four-product one."
          },
          {
            "n": 10,
            "kind": "cta",
            "duration_sec": 8.0,
            "motion": "Remotion CTA",
            "keyframe_prompt": null,
            "overlay": {
              "wordmark": "Spoiled Child",
              "cta_pill": "Build your morning routine \u2192",
              "url": "spoiledchild.com",
              "ground": "#C96E2F"
            },
            "vo": "Spoiled Child. Your morning routine, simplified. Spoiledchild dot com."
          }
        ],
        "product_grid": {
          "method": "PIL composite of the 4 REAL product webps (never AI \u2014 AI duplicates SKUs)",
          "ground": "#C96E2F",
          "layout": "2x2",
          "preserve_aspect": true,
          "images": [
            "assets/products/e27-main-bottle.webp",
            "assets/products/s33-bottle-close.webp",
            "assets/products/t31-main-product.webp",
            "assets/products/h30-main-product.webp"
          ]
        },
        "voice": {
          "engine": "ElevenLabs",
          "model": "eleven_v3",
          "endpoint": "text-to-speech/with-timestamps",
          "voice_chosen": "Eryn",
          "voice_id": "dMyQqiVXTU80dDl2eNK8",
          "casting_ab": [
            "Eryn (dMyQqiVXTU80dDl2eNK8)",
            "Angela (FUfBrNit0NNZAwb58KWH)"
          ],
          "settings": {
            "stability": 0.45,
            "similarity_boost": 0.8,
            "style": 0.1,
            "use_speaker_boost": true,
            "speed": 1.1
          },
          "_note": "Write VO as FULL SENTENCES (not keyword fragments). with-timestamps char-level timings drive the word-by-word captions."
        },
        "music": {
          "engine": "ElevenLabs",
          "prompt": "Lo-fi pop bed, 95-105 BPM, builds across sections (intro pad -> groove -> warm -> lo-fi build -> resolved). NOT cinematic, NOT moody. No melody hooks competing with VO. Instrumental only.",
          "length_ms": 52000,
          "force_instrumental": true,
          "mix_profile": "VO-forward: vo loudnorm I=-15, music loudnorm I=-25 + volume 0.5, sidechaincompress duck, final alimiter limit=0.89, target ~-15 LUFS / -1 dBTP"
        },
        "captions": {
          "method": "word-by-word burned from eleven_v3 with-timestamps (libass; Klap is the hosted equivalent)",
          "style": "Inter/Arial 56 white + 4px outline, bottom-third (MarginV 360)",
          "burned_last": true,
          "suppress_on_scenes": [
            3,
            8,
            10
          ],
          "_suppress_note": "slate/grid/CTA scenes carry their own on-screen text \u2014 two text layers collide."
        },
        "cutdowns": {
          "source": "working/silent-master.mp4 (the ANIMATED master \u2014 NEVER a static intermediate)",
          "rules": "trim short beats (never freeze); gently slow long beats (setpts, clamp <=1.6x); re-burn scaled captions (timestamps / vo_speedup, offset to new scene starts); suppress on 3/8/10",
          "variants": {
            "v1": "full 10-beat story, VO @1.25x (atempo), silences trimmed, ~30s \u2014 THE SHIPPED DELIVERABLE",
            "v2": "hook-led montage: hook -> slate -> grid pulled forward -> rapid 4-product montage (number + product only) -> CTA, ~30s (>=4 cuts/10s)"
          }
        },
        "post_production": {
          "music": {
            "default": "on",
            "note": "lo-fi pop bed under the VO, default on"
          },
          "captions": {
            "default": "on",
            "note": "word-by-word burned captions (libass/Klap), default on"
          },
          "end_card": {
            "default": "on",
            "note": "CTA + brand wordmark scene, default on"
          }
        }
      }
      
    • PIPELINE.md 4 KB
      # PIPELINE — flat-vector-explainer engine map
      
      This molecule is **documentation-grade**. Rather than re-implement a 14-scene Remotion
      app here, it maps every `config.json` field to the **real, runnable script** that
      produced the worked example. The executable reference is the source project:
      
      ```
      clients/spoiled-child/video-11-routine-broken/
      ```
      
      Its `HOW_TO.md` + `LEARNINGS.md` are the authoritative v3 recipe (Kling i2v + Remotion
      overlay). Run the pipeline there (or port these scripts into a new brand's project
      folder), then bring the config here as the recipe of record.
      
      ## Config field → source script
      
      | `config.json` field | Source script | Phase | Paid? |
      |---|---|---|---|
      | `character.anchor_prompt`, `character.keyframe_variants`, `scenes[].keyframe_prompt` | `working/gen_keyframes.py` | 1 — character lock + per-scene keyframes (nano-banana, flat-vector; re-render a FRESH anchor — never chain a photoreal ref) | **PAID** |
      | (clean plates — strip baked text) | `working/clean_plate.py` | 2 — nano-banana edit removes chips/numerals/taglines/badges from character keyframes → clean plates for i2v | **PAID** |
      | `scenes[].motion`, `kling.*` | `working/kling_i2v.py` | 3 — Kling 2.5-turbo/pro i2v on character scenes only; cfg 0.5 + style-preserving negative + subtle motion; TEST one scene first | **PAID** |
      | `scenes[].overlay` (chips, numerals, taglines, slate, CTA) | `working/remotion/` (Remotion project) | 4 — imports each Kling clip as the moving base and composites ALL text as animated DOM ON TOP → animated silent master. Text is NEVER baked into a keyframe | free |
      | `product_grid.*` | `working/scripts/build_scene08.py` | 4 — PIL composite of the N real product webps on the brand ground (preserve each aspect; AI duplicates SKUs so this is PIL, not AI) | free |
      | `voice.*`, `scenes[].vo` | `working/scripts/render_vo.py` | 5 — ElevenLabs eleven_v3, `text-to-speech/with-timestamps`; full-sentence per-scene VO + char-level timestamps | **PAID** |
      | `music.*` | `working/gen_music.py` | 5 — ElevenLabs music bed, lo-fi pop, VO-forward loudnorm | **PAID** |
      | (VO+music mix) | `working/scripts/mix_audio.sh` | 5 — place each VO line at its scene start, duck music under VO (sidechaincompress), `loudnorm I=-15`, VO-forward | free |
      | `captions.*` | `working/scripts/build_captions.py` | 5 — word-by-word burned captions from the VO char-timestamps (libass); suppress on slate/grid/CTA scenes | free |
      | (silent master assembly) | `working/scripts/build_silent.sh` | 4 — stitch the Remotion-composited scenes into `working/silent-master.mp4` (the ANIMATED master) | free |
      | (50s master) | `working/build_master.py` | 5 — mux silent master + mixed audio, burn captions LAST → `finals/master-final.mp4` | free |
      | `cutdowns.*` | `working/build_30s.py` | 6 — slice each beat's region OUT of `silent-master.mp4` (the ANIMATED master, NEVER a static intermediate); trim short / slow long (≤1.6×); re-burn scaled captions → `finals/master-final-30s-v1.mp4` | free |
      | (QC) | `working/motion_probe.py` | 7 — frame-diff proof of localized motion (+ `/watch` the final) | free |
      
      ## The two non-negotiable separations
      
      1. **Motion layer ≠ text layer.** `clean_plate.py` strips text → `kling_i2v.py` animates
         the clean plate → `remotion/` composites text as DOM on top. Never `kling_i2v.py` a
         keyframe with baked text (i2v warps it, and you can't retime/restyle it).
      2. **Real assets ≠ AI assets.** `build_scene08.py` PIL-composites the real product webps.
         AI gen duplicates SKUs in a grid — it is only for the character vignettes + backgrounds.
      
      ## The cut-down trap (LEARNINGS L6)
      
      `build_30s.py` MUST slice from `working/silent-master.mp4` (the Kling-animated master),
      NOT from `working/segs/` (an earlier static Ken-Burns round). The mtime is the tell.
      A finished-looking audio mix can hide frozen characters — always run `motion_probe.py`
      frame-diff on a character scene and confirm **localized** face/hand glow (whole-outline
      glow = pan-only = wrong source).
      
    • README.md 5.4 KB
      # render-flat-vector-explainer — free assembly how-to
      
      This capability is **documentation-grade**: the flat-vector-explainer molecule is a
      documented recipe, not a runnable end-to-end app. There are no standalone scripts here to
      `--config` and fire. Instead this folder ships:
      
      - **`config.example.json`** — the full config schema (the Spoiled Child "Perfect Morning
        Routine = 4 Products" worked example). Copy to `config.json` and edit.
      - **`PIPELINE.md`** — the field-by-field map from every `config.json` key to the real,
        runnable source script that produced the worked example (in the source project's
        `working/`), and which phase / paid-or-free it is.
      - **this README** — the FREE assembly steps the agent runs by hand with ffmpeg + Remotion +
        PIL after the paid gen steps have produced the character clips, VO, and music.
      
      The paid generative steps (keyframes, clean plates, Kling i2v, VO, music) are **separate
      capabilities** the recipe orchestrates and gates — see the recipe. Everything below is
      free, deterministic, $0.
      
      ## The two separations that make this format work
      
      1. **Motion layer != text layer.** i2v only ever animates a **text-stripped clean plate**.
         ALL words/numerals/chips/taglines/slate/CTA are composited on top as an **animated
         Remotion DOM overlay** — never baked into the keyframe. Baked text warps under i2v and
         you lose the ability to retime/restyle it.
      2. **Real assets != AI assets.** The per-step product photo and the closing "N products"
         grid are the **real product webps composited with PIL**. AI duplicates SKUs in a grid.
         Only the character vignettes + backgrounds are generative.
      
      ## Free assembly steps
      
      Run these after the paid steps have delivered the per-scene Kling clips (character beats),
      the VO (with char-level timestamps), and the music bed.
      
      ### 1. Remotion overlay -> animated silent master (character-locked)
      
      Build a Remotion composition that imports each Kling clip as the moving base and
      composites the scene's `overlay` (numeral, chip, tagline, slate headline, CTA pill,
      wordmark, url) as **animated DOM on top**. Keep the **motion layer and text layer
      strictly separate** — the Kling footage is the only moving base; text is DOM.
      
      - `kind: character` scenes -> Kling clip base + DOM overlay.
      - `kind: slate` / `grid` / `cta` scenes -> Remotion text only, **no i2v** (a solid brand
        `ground` color + the headline / grid / CTA).
      - Use the brand palette + display font from the config for crisp glyphs at the exact brand
        color. Real DOM type = crisp; never let i2v render the type.
      
      Render the composition to `working/silent-master.mp4` — the **animated** master.
      Everything downstream slices from this, never from an earlier static intermediate.
      
      ### 2. PIL real-product grid
      
      Composite the closing "N products" lockup from `product_grid.images` (the real product
      webps) onto `product_grid.ground` in the configured `layout` (e.g. 2x2). **Preserve each
      product's aspect ratio — never stretch, never AI-generate the grid** (AI dupes SKUs). The
      same real webps are also the per-step callout photos in the character beats. Product webps
      are git-LFS in the brand folder — fetch + checkout first (pointers are ~131-byte stubs).
      
      ### 3. Caption burn (word-synced)
      
      Build word-by-word burned captions from the eleven_v3 with-timestamps **char-level
      timings** (libass; Klap is the hosted equivalent). Style per `captions.style` (e.g. Inter
      56 white + 4px outline, bottom-third MarginV 360). **Suppress captions on the
      slate/grid/CTA scenes** (`captions.suppress_on_scenes`) — those carry their own on-screen
      text and two text layers collide. Burn captions **LAST**, after the audio mux.
      
      ### 4. Audio mix + 50s master
      
      Place each scene's VO line at its scene start; duck the music under the VO with
      `sidechaincompress`; `loudnorm I=-15` VO-forward (music `loudnorm I=-25 + volume 0.5`,
      final `alimiter limit=0.89`, ~-15 LUFS / -1 dBTP). Mux the silent master + mixed audio,
      then burn captions last -> `finals/master-final.mp4` (~50s full-story master).
      
      ### 5. 50s -> 30s cut (from the ANIMATED master)
      
      Slice each beat's region OUT of `working/silent-master.mp4` — the **animated** master,
      **never** an earlier static Ken-Burns / segs round (the mtime is the tell; a finished audio
      mix can hide frozen characters). Trim short beats (never freeze a frame); gently slow long
      beats (`setpts`, clamp <=1.6x) so motion stretches instead of holding. Re-burn scaled
      captions offset to the new scene starts; keep suppression on the slate/grid/CTA scenes.
      Output `finals/master-final-30s-v1.mp4` — the shipped 30s deliverable.
      
      Optional v2 cutdown: hook-led montage (hook -> slate -> grid pulled forward -> rapid
      4-product montage of number + product only -> CTA), >=4 cuts / 10s.
      
      ## QC before ship (mandatory)
      
      - **Frame-diff a character scene** (blend=difference heatmap): **localized** face/hand
        glow = real motion; whole-outline glow = a static pan = you sliced the wrong (static)
        source — re-point the cut at `silent-master.mp4`.
      - `/watch` the final: real localized motion, correct duration (~30s ±0.3s), caption sync,
        correct numerals/labels, no AI text leak in the i2v footage, N distinct real products in
        the grid (no AI dupes), VO-forward mix (~-15 LUFS).
      - Canvas 1080x1920, 30fps, h264, aac present.
      
      ## Requires
      
      Node + Remotion, `Pillow` (PIL), `ffmpeg` with libass. All free — the paid keys
      (`FAL_API_KEY`, `ELEVENLABS_API_KEY`) belong to the separate paid capabilities, which
      route through the GooseWorks proxies so the calls bill the Ads agent.
      
  • tests
    • smoke-test.md 1.2 KB
      # Smoke Test
      
      This capability is documentation-grade — it ships the config schema + assembly recipe, not
      a runnable end-to-end script. The check is that the docs + config are present and coherent.
      
      Pass when:
      - `scripts/config.example.json` parses as valid JSON and carries the flat-vector-explainer
        schema (concept + single countable point, character anchor, ordered `scenes[]` with
        `kind` in character/slate/grid/cta, `product_grid` of real webps, voice, music, captions
        with `suppress_on_scenes`, cutdowns, post_production toggles).
      - `scripts/PIPELINE.md` maps every config field to its source script + phase + paid/free.
      - `scripts/README.md` documents the FREE assembly steps: Remotion text/DOM overlay kept
        SEPARATE from the Kling i2v motion layer (character-locked), PIL real-product grid,
        word-synced caption burn, VO-forward audio mix, and the 50s -> 30s cut taken FROM the
        animated silent master (never a static intermediate).
      - The SKILL.md description is a single line with no ": " and states the paid gen steps
        (keyframes, Kling i2v, VO, music) are separate capabilities the recipe orchestrates.
      - No paid call is made by this capability. Paid caps route through the GooseWorks proxies
        (bill the agent, no direct provider host).
      
  • SKILL.md 5 KB
    ---
    name: render-flat-vector-explainer
    description: Assemble the FREE steps of the flat-vector-explainer video format — a flat-illustration creator-character walks a countable N-step product routine, one step per beat, and Remotion composites every chip/numeral/tagline/slate/CTA as an animated DOM overlay ON TOP of the Kling i2v character clips (text is NEVER baked into a keyframe — i2v warps type), the closing 'N products' grid is a PIL composite of the REAL product photos (not AI), full-sentence VO drives word-by-word burned captions over a VO-forward music bed, and the ~50s animated silent master is re-cut to a 30s deliverable FROM the animated master (never a static intermediate). Documentation-grade — ships config.example.json + PIPELINE.md + a README of the free assembly; the paid gen steps (keyframes, Kling i2v, VO, music) are separate capabilities the recipe orchestrates. Use for the flat-vector-explainer format.
    status: active
    ---
    
    # render-flat-vector-explainer
    
    Assembles a flat-vector product-routine explainer: one illustrated creator-character walks through a countable N-step routine (e.g. collagen -> serum -> eye cream -> hair), one step per beat, each beat carrying a large corner numeral, a labelled chip + one-line tagline, and the step's real product photo, closing on an "N products" grid + brand CTA. It reads as a premium DTC explainer (Spotify/Anchor flat-vector lineage), not UGC.
    
    This capability is **documentation-grade**. The content-goose molecule is a documented recipe, not a runnable end-to-end app, so this capability ships the **config schema** (`scripts/config.example.json`), the **field-to-script map** (`scripts/PIPELINE.md`), and a **README** (`scripts/README.md`) describing the FREE assembly steps the agent runs by hand with ffmpeg + Remotion + PIL. The paid generative steps are separate capabilities the recipe orchestrates and gates.
    
    ## The two non-negotiable separations
    
    1. **Motion layer != text layer.** Animate a **text-stripped clean plate** with Kling i2v (subtle motion, style-preserving negative, cfg 0.5), then composite every chip / numeral / tagline / slate / CTA as an **animated Remotion DOM overlay** on top. Baking text into the keyframe before i2v warps the type and forfeits the ability to retime/restyle it — this separation is the format's whole credibility.
    2. **Real assets != AI assets.** The per-step product photo and the closing "N products" grid are **real product webps composited with PIL** (AI duplicates SKUs in a grid). Only the character vignettes and stylized backgrounds are generative.
    
    ## Free assembly steps (this capability)
    
    The agent runs these deterministic, $0 steps by hand — see `scripts/README.md` for the ffmpeg/Remotion/PIL detail:
    
    - **Remotion overlay** — import each Kling clip as the moving base; composite chips / numerals / taglines / slate / grid / CTA as animated DOM on top -> the animated silent master. Slate/grid/CTA beats are Remotion text with no i2v.
    - **PIL product grid** — composite the N real product webps on the brand ground for the closing lockup; preserve each aspect (never stretch, never AI-dupe).
    - **Captions** — word-by-word burned from the eleven_v3 with-timestamps char timings (libass); suppress on slate/grid/CTA scenes so two text layers don't collide.
    - **Audio mix + master** — place each VO line at its scene start, duck the music under VO (sidechaincompress), `loudnorm I=-15` VO-forward, mux, burn captions LAST -> `finals/master-final.mp4` (~50s).
    - **30s cut** — slice each beat's region OUT of the **animated silent master** (never a static intermediate); trim short beats, gently slow long beats (setpts <=1.6x), re-burn scaled captions -> `finals/master-final-30s-v1.mp4`.
    
    ## Paid gen steps (separate capabilities)
    
    The recipe orchestrates and gates these; they are not part of this capability:
    
    - Flat-vector character anchor + per-scene keyframes + clean plates -> `create-image-fal` (nano-banana; re-render a FRESH flat-vector anchor, never chain a photoreal ref).
    - Kling i2v on the character scenes -> `create-video-fal` (Kling 2.5-turbo/pro, cfg 0.5, style-preserving negative, low motion; TEST one scene before batching).
    - Full-sentence VO -> `create-vo-elevenlabs` (eleven_v3, with-timestamps).
    - Lo-fi music bed -> `create-music-elevenlabs`.
    
    ## Contract
    
    - Documentation-grade + FREE assembly (Remotion + PIL + FFmpeg); no paid calls in this capability, no AI-rendered text.
    - Text is an overlay, never baked. Strip to a clean plate -> i2v -> composite text as Remotion DOM.
    - Any multi-SKU grid is PIL of the real product webps; preserve each aspect ratio.
    - Kling holds the 2D flat-vector look only at LOW motion (cfg 0.5 + style-preserving negative). Aggressive motion drifts to photoreal.
    - Cut down from the ANIMATED master, never a static intermediate; frame-diff to prove localized motion.
    - The paid steps — keyframes/clean plates, Kling i2v, VO, music — are separate capabilities (create-image-fal, create-video-fal, create-vo-elevenlabs, create-music-elevenlabs); the recipe orchestrates them and gates the spend.
    
  • skill.meta.json 333 B
    {
      "slug": "render-flat-vector-explainer",
      "category": "capabilities",
      "domain": "ads",
      "tags": [
        "ads"
      ],
      "installation": {
        "base_command": "npx goose-skills install render-flat-vector-explainer",
        "supports": [
          "claude",
          "cursor",
          "codex"
        ]
      },
      "requires_skills": [
        "watch"
      ]
    }
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related