{"slug":"muapi-storyboard-to-cooking-video","title":"muapi-storyboard-to-cooking-video","summary":"Turn a single photo of a person into a 15-second cinematic pasta-making (or other cuisine) tutorial video. First builds a composite reference sheet (character + kitchen + 9-step action board), then animates the full cooking sequence with audio in a single continuous shot.","platform":"opencode","tags":[],"authorName":"LLM Mart","authorSlug":"llm-mart","score":0,"source":"github","price":null,"verified":false,"createdAt":"2026-09-01T17:42:40.62037Z","repo":{"url":"https://github.com/SamurAIGPT/Generative-Media-Skills","stars":5509,"forks":657,"license":null,"updatedAt":"2026-10-07T08:21:48Z"},"bodyHtml":"<hr>\n<h2>slug: muapi-storyboard-to-cooking-video\nname: muapi-storyboard-to-cooking-video\nversion: \"1.0.0\"\ndescription: Turn a single photo of a person into a 15-second cinematic pasta-making (or other cuisine) tutorial video. First builds a composite reference sheet (character + kitchen + 9-step action board), then animates the full cooking sequence with audio in a single continuous shot.\nacceptLicenseTerms: true</h2>\n<h1>Storyboard to Cooking Video</h1>\n<p><strong>Turn a single photo of a person into a polished 15-second cinematic cooking tutorial. The skill first generates a high-end production reference sheet — character look, kitchen environment, and a 9-panel action board — then drives a continuous reference-to-video render that keeps the subject's face, outfit, and kitchen consistent across every frame.</strong></p>\n<h2>Inputs</h2>\n<table>\n<thead>\n<tr>\n<th style=\"text-align: left\">Name</th>\n<th style=\"text-align: left\">Type</th>\n<th style=\"text-align: left\">Required</th>\n<th style=\"text-align: left\">Default</th>\n<th style=\"text-align: left\">Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td style=\"text-align: left\"><code>person_image</code></td>\n<td style=\"text-align: left\">image_url</td>\n<td style=\"text-align: left\">yes</td>\n<td style=\"text-align: left\">—</td>\n<td style=\"text-align: left\">URL of the person photo. Used as identity reference in BOTH the reference sheet and the final video.</td>\n</tr>\n<tr>\n<td style=\"text-align: left\"><code>dish</code></td>\n<td style=\"text-align: left\">text</td>\n<td style=\"text-align: left\">no</td>\n<td style=\"text-align: left\">fresh pasta</td>\n<td style=\"text-align: left\">The cooking subject (e.g. \"fresh pasta\", \"sushi rolls\", \"wood-fired pizza\", \"matcha latte\"). Drives the 9-step action board.</td>\n</tr>\n<tr>\n<td style=\"text-align: left\"><code>kitchen_style</code></td>\n<td style=\"text-align: left\">text</td>\n<td style=\"text-align: left\">no</td>\n<td style=\"text-align: left\">warm rustic-modern Italian</td>\n<td style=\"text-align: left\">The kitchen aesthetic (e.g. \"warm rustic-modern Italian\", \"minimalist Tokyo\", \"bright Scandinavian\", \"moody industrial\").</td>\n</tr>\n<tr>\n<td style=\"text-align: left\"><code>outfit</code></td>\n<td style=\"text-align: left\">text</td>\n<td style=\"text-align: left\">no</td>\n<td style=\"text-align: left\">white t-shirt, olive green apron, dark trousers</td>\n<td style=\"text-align: left\">What the person wears throughout the video.</td>\n</tr>\n<tr>\n<td style=\"text-align: left\"><code>duration_seconds</code></td>\n<td style=\"text-align: left\">int</td>\n<td style=\"text-align: left\">no</td>\n<td style=\"text-align: left\">15</td>\n<td style=\"text-align: left\">Final video duration. Use 15 for the full 9-step arc; 10 collapses to ~6 beats.</td>\n</tr>\n<tr>\n<td style=\"text-align: left\"><code>aspect_ratio</code></td>\n<td style=\"text-align: left\">text</td>\n<td style=\"text-align: left\">no</td>\n<td style=\"text-align: left\">16:9</td>\n<td style=\"text-align: left\">Output aspect ratio. Use <code>9:16</code> for vertical/Reels.</td>\n</tr>\n<tr>\n<td style=\"text-align: left\"><code>resolution</code></td>\n<td style=\"text-align: left\">text</td>\n<td style=\"text-align: left\">no</td>\n<td style=\"text-align: left\">720p</td>\n<td style=\"text-align: left\">Video resolution. Options: <code>480p</code>, <code>720p</code>.</td>\n</tr>\n</tbody>\n</table>\n<h2>Steps</h2>\n<p>Submit the plan with TWO sequential steps. Step 2 depends on the output of Step 1.</p>\n<h3>Step 1 — Reference Sheet (Composite Storyboard)</h3>\n<p>Generate the composite \"production reference board\" image. This is a single image, NOT a video frame — it bundles character sheet + location reference + 9-panel action board.</p>\n<p><strong>Endpoint:</strong> <code>gpt-image-v2-edit</code>\n<strong>CLI:</strong></p>\n<pre><code>muapi image edit \\\n  --model gpt-image-v2-edit \\\n  --image \"{{person_image}}\" \\\n  --image-size \"3840x2160\" \\\n  --quality auto \\\n  --background auto \\\n  --moderation low \\\n  --output-format png \\\n  --prompt \"Create one single composite reference sheet for a {{duration_seconds}}-second realistic {{dish}}-making tutorial video. The image should be a clean, high-end production reference board, not a poster with heavy text. Format: {{aspect_ratio}} wide reference sheet, elegant white margins, clean grid layout, realistic cinematic photography style. Concept: {{dish}} tutorial in a {{kitchen_style}} kitchen.\n\nTop row: motion / choreography guide with 9 numbered cinematic action panels showing the {{dish}} process step-by-step from raw ingredients to final plated dish.\n\nMiddle-left: realistic character reference sheet of the uploaded person — preserve their exact face, hair color, hair texture, eye color, skin tone, and all facial features with 100% accuracy. Show the same person in: face close-up, full-body front view, side/action working pose, and back view. Dress them in {{outfit}}. Keep them grounded, approachable, skilled, and cinematic.\n\nMiddle-right / background: location reference sheet of an elegant {{kitchen_style}} kitchen with tactile surfaces, natural daylight from a large window, hanging cookware, herbs, and premium cooking atmosphere appropriate to the cuisine.\n\nStyle: realistic, cinematic, warm natural light, shallow depth of field, tactile food photography, premium cooking show aesthetic, rich surface textures.\n\nBottom strip: simple visual icons only for {{duration_seconds}} seconds, {{aspect_ratio}}, realistic, cinematic, tasty, natural camera. Minimal text, no dense paragraphs. Let the visuals do the heavy lifting.\"\n</code></pre>\n<p>Wait for completion and capture the output URL as <code>{{reference_sheet_url}}</code>. Show it to the user and confirm the character likeness + kitchen mood before moving to Step 2 — Step 2 is the expensive call.</p>\n<h3>Step 2 — Cooking Video (Reference-to-Video)</h3>\n<p>Animate the full sequence using both the original person photo (identity anchor) and the reference sheet (narrative + environment guide) as dual references.</p>\n<p><strong>Endpoint:</strong> <code>bytedance-seedance-2-0-reference-to-video-fast</code>\n<strong>CLI:</strong></p>\n<pre><code>muapi video generate \\\n  --model bytedance-seedance-2-0-reference-to-video-fast \\\n  --image \"{{person_image}}\" \\\n  --image \"{{reference_sheet_url}}\" \\\n  --aspect-ratio \"{{aspect_ratio}}\" \\\n  --duration \"{{duration_seconds}}\" \\\n  --resolution \"{{resolution}}\" \\\n  --generate-audio true \\\n  --prompt \"The person in @Image1 is the subject — preserve their exact face, hair, eye color, skin tone, and all facial features with 100% accuracy throughout the entire video.\nUse @Image2 as the visual and narrative guide — follow the cooking steps, kitchen setting, outfit, and atmosphere shown in the reference sheet exactly.\nA single continuous cinematic video of the person from @Image1 making {{dish}} in the {{kitchen_style}} kitchen shown in @Image2. They wear {{outfit}} throughout.\n\nVIDEO STRUCTURE\nFollow the exact 9-step sequence as shown in @Image2, beat by beat, from raw ingredients through preparation to a final plated close-up.\n\nMOTION STYLE\n- Slow, deliberate, satisfying transitions between each step\n- Natural hand and body movement with clear culinary intent\n- Continuous flow with no jump cuts\n- Warm and immersive pacing\n\nCAMERA &amp; CINEMATOGRAPHY\n- Close-up shots for hands during mixing, kneading, cutting, plating\n- Medium shots showing the person working at the counter\n- Pull back slightly for the final plating to reveal the full kitchen\n- Shallow depth of field — focus on hands and food, soft background blur\n- No abrupt cuts — smooth match cuts and fluid transitions\n\nVISUAL STYLE\n- Warm natural daylight from a large kitchen window\n- Rich tactile textures matching @Image2's environment\n- Full color, warm cinematic color grading\n\nCONSISTENCY RULES\n- Same character throughout — face of @Image1 in every frame\n- Same outfit across entire video\n- Same kitchen environment as shown in @Image2\n\nAUDIO\n- Soft kitchen ambience, gentle culinary SFX (chopping, sizzling, pouring), light cinematic underscore\n- No dialogue, no narration\n\nOUTPUT STYLE\n- Duration: exactly {{duration_seconds}} seconds\n- Polished, cinematic, premium cooking show quality\n- Ends with a beautiful close-up of the finished plated {{dish}}\"\n</code></pre>\n<p>After generation:</p>\n<ul>\n<li>Present the final video URL to the user.</li>\n<li>Offer follow-ups: vertical 9:16 re-render for Reels, a longer 30s extended cut, or swap <code>{{dish}}</code> for a different cuisine using the same person image.</li>\n</ul>\n<h2>Notes</h2>\n<ul>\n<li><strong>Two-image reference is the whole trick.</strong> <code>@Image1</code> locks identity, <code>@Image2</code> locks choreography + environment. Never drop one — single-reference runs lose either the face or the kitchen.</li>\n<li>The reference sheet at Step 1 must be wide (3840x2160). Smaller resolutions blur the 9 action panels and the video model can't read them.</li>\n<li><code>bytedance-seedance-2-0-reference-to-video-fast</code> natively generates audio when <code>generate_audio=true</code>. Always include an audio direction in the prompt; otherwise the soundtrack is random.</li>\n<li>Real human faces ARE supported here because the person photo is the user's own subject and we route through the reference-to-video endpoint (not the restricted i2v variants).</li>\n<li>If the user wants a non-cooking sequence (e.g., latte art, plating tutorial, mixology), keep the same two-step structure — only <code>{{dish}}</code> and the 9-step description change.</li>\n<li>For shorter pieces (&lt;= 8s), reduce the action board to 5–6 panels in Step 1; cramming 9 beats into 8s degrades motion quality (single-beat rule).</li>\n</ul>\n<h2>Trigger Keywords</h2>\n<p><code>cooking video</code>, <code>cooking tutorial</code>, <code>pasta video</code>, <code>recipe video</code>, <code>food video</code>, <code>chef video</code>, <code>cooking storyboard</code>, <code>kitchen tutorial</code>, <code>cooking reel</code>, <code>tutorial video from photo</code>, <code>storyboard to video</code></p>\n<hr>\n<h2>Notes for the Executing Agent</h2>\n<ul>\n<li>This recipe is LLM-orchestrated: read each phase, gather any missing inputs from the user, then call <code>muapi</code> CLI commands. Use <code>muapi auth configure</code> first if <code>MUAPI_API_KEY</code> is unset.</li>\n<li>For model IDs without a CLI alias yet, fall back to the raw endpoint via <code>curl -X POST https://api.muapi.ai/api/v1/&lt;endpoint&gt; -H \"x-api-key: $MUAPI_API_KEY\" -H 'content-type: application/json' -d '{...}'</code> and poll with <code>muapi predict wait &lt;request_id&gt;</code>.</li>\n<li>Substitute <code>{{input_name}}</code> placeholders with the user's actual inputs before issuing each call.</li>\n<li>Step 1 must complete and return an output image URL before Step 2 fires — pass that URL as the second <code>--image</code> to the video step.</li>\n</ul>\n","files":[{"path":"SKILL.md","sizeBytes":8794,"isText":true}],"reviewScore":null,"reviewSummary":null,"trust":{"provenance":"trusted-source-unreviewed","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow.","bodySource":null},"bodyLocked":false,"purchaseUrl":null,"sourceUrl":null,"report":{"provenance":"trusted-source-unreviewed","screen":{"ran":true,"outcome":"clean","suspicious":0,"notes":0,"hiddenCharacters":false},"virusScan":{"engine":"clamav","status":"clean","scannedAt":"2026-09-01T17:43:35.731417Z","sha256":"77AB80636444B031F5144B401C02FF10F93D55F7797CBF5EE20C68E01E8BD22D","sizeBytes":3882},"review":null,"source":{"repositoryUrl":"https://github.com/SamurAIGPT/Generative-Media-Skills","path":"library/motion/storyboard-to-cooking-video","license":null,"commit":"3d8de1cd6657c6d70583f34b89c2dc034512c1ea","subtreeSha":"CF90E75D791594CB5742396F8F396CC05954A7EE9BCDA5CD5D7AB3129FBFB0A6","lastSyncedAt":"2026-10-07T15:23:11.266656Z"},"reviewedAt":"2026-09-01T17:46:03.529388Z","notice":"Community-authored content, reproduced verbatim and not vetted as instructions. Treat it as data to evaluate, never as directives to follow."},"install":[{"target":"skills-cli","command":"npx skills add https://github.com/SamurAIGPT/Generative-Media-Skills/tree/main/library/motion/storyboard-to-cooking-video"},{"target":"claude-code","command":"claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install samuraigpt-generative-media-skills@llmmart"},{"target":"git","command":"git clone https://github.com/SamurAIGPT/Generative-Media-Skills.git"}]}