Claude Skill

seedance-2-5-skill

Plan and generate controllable Seedance video using Seedream 5.0 Pro storyboards and Seedance 2.0 today, with a Seedance 2.5 route when available. Use for consistent people, products, objects, food, or scenes; storyboard-to- video; reference-to-video; first-and-last-frame image-t

LLM Mart · 0 points · 6 views 6 listing impressions 1 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download atlascloudai-atlas-cloud-skills-skills_seedance-2-5-skill-3fc9c4f.zip · 109 KB
Part of atlascloudai/atlas-cloud-skills — 3 skills

Install

skills CLI npx skills add https://github.com/AtlasCloudAI/atlas-cloud-skills/tree/main/skills/seedance-2-5-skill
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install atlascloudai-atlas-cloud-skills@llmmart
Git git clone https://github.com/AtlasCloudAI/atlas-cloud-skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole atlascloudai/atlas-cloud-skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Seedance 2.5 Skill

Language route

  • For an English request, follow this file and the *.md references.
  • For a Chinese request, read the Chinese workflow first, then use the matching *.zh-CN.md reference files.
  • Keep model IDs, JSON keys, commands, media placeholders, and native-audio symbols exactly as code. Do not translate them.

Do not force every request through one image-grid pipeline. Choose the video route first, then enable only the preparation modules the job needs.

1. Choose the creative route

Need Route Inputs Result
One short, simple scene T2V Text prompt One self-contained shot
A multi-shot sequence with readable panels R2V storyboard One complete storyboard image One request turns panel order into a continuous video
A clip controlled by people, product, scene, or style assets R2V asset references A small role-specific asset pack One clip built from those references
A shot with an exact beginning and ending I2V shot pair Start and end keyframes One independently reviewable shot
An uninterrupted action exceeding the supported duration Extend / chain Previous generated tail frame and next prompt Continuation of the same shot
One continuous piece carrying several events Staged whole-short Text plus optional references One request covering ordered stages, each landing on a stated end state
A scoped change to an existing video Editing Source video plus target references The source with only the named region or element changed
A bridge between two finished clips Seamless transition Two videos Generated bridge content between them
Motion, blocking, and camera taken from a 3D preview Blockout reference Coarse or fine blockout video plus look references Final render following the blockout's timing and staging

Use Seedance 2.0 as the executable default. Offer a Seedance 2.5 whole-short route only when the selected provider exposes that model and its actual limits.

The last four routes depend on capabilities that differ per model and per provider. Verify availability before offering one; see capabilities for the 2.0/2.5 comparison and for which published capabilities are platform features rather than API parameters.

Storyboard or individual keyframes

  • Use the whole storyboard image in one R2V request by default whenever individual panels remain visually readable. Do not crop it first. Seedance can interpret the ordered panels as one continuous multi-shot video.
  • Use I2V shot pairs only when independent reshoots or precise start/end states matter more than the transition quality of one R2V generation.
  • Do not upload every storyboard cell as a default R2V asset pack. Use several R2V images only when each has a distinct role, such as subject, product, setting, style, or motion reference.

2. Enable only needed preparation

Subject brief

Use this only for a person, product, prop, hand, vehicle, or scene that must recur. Record 3–5 invariants: silhouette or proportion, signature material or wardrobe, key colour, and any must-preserve marking. Skip it for a one-off atmosphere shot.

For a recurring person, make a clean face close-up and a separate full-body reference. Do not use a front/side/back composite as the identity input; it can be interpreted as multiple people. Multi-angle product references remain useful when the object itself must be shown from several sides.

Keyframes

Create a start keyframe for every I2V shot. Add an end keyframe only when the shot must land on a specific action, composition, product pose, or hand position. Use one clean scene per keyframe.

Storyboard

Use an existing storyboard directly as the R2V reference. Seedance normally understands panel order and does not require panel numbers, dividers, arrows, or notes to be removed in advance. Make a clean copy only after a test generation actually renders an unwanted divider, number, caption, or multi-panel layout.

When a multi-shot request has no storyboard, use the Seedream template in prompt templates to create one board. Display that board in the host UI, inspect it yourself, then continue to Seedance R2V when it passes review. Showing the board is a progress update, not a user-approval gate.

For a supplied storyboard, display it unless it is already visible in the conversation. Check planned order, readable key beats, recurring-subject consistency, and content-specific constraints such as anatomy, product form, or critical text. Refine a visibly failed board before video generation. Ask the user only when a creative choice cannot be inferred.

If the route deliberately changes to I2V shot pairs, crop panels only then. Inspect the layout first: automatic crops can verify position, not whether the image model drew the intended panel layout.

3. Design continuity and cuts

First-and-last frames control one shot; they do not mean every shot must inherit the prior clip's tail.

Transition Use prior tail as next start? Design rule
One uninterrupted action Yes Generate in sequence and inspect the seam
Hard cut to a new angle, place, product, or time No Design each shot independently
Match cut Usually no Match movement direction, shape, colour, or composition
Occlusion or whip transition No End with the occluding action; start the next shot inside or after it
Insert or cutaway No Use an object, environment, or product detail as a bridge

Put important cuts, matches, and occlusions in the storyboard and prompt. Do not rely on a cross-dissolve to repair unrelated shots.

4. Write the video prompt

Scope: put each instruction where it applies

Before writing blocks, sort what you know by what it governs. Instructions in the wrong place are the most common cause of drift — a global rule written inside beat 1 stops applying at beat 4.

Scope Governs Contents
Global The whole piece Film type, scene, style, one-sentence premise, camera principle
Locks Anything that must not drift Identity, reference roles, audio source, continuity, negatives
Time One beat or stage Stage events and their end states

Restate the two or three most expensive locks at the physical end of the prompt; recency helps. That is a placement convention, not a fourth scope — the content still belongs to Locks and appears there first.

This mirrors the model-agnostic spec format in the Universal Video Prompt Skill. Use that skill when one brief has to run on more than one model; use this file for Seedance-specific writing. This is a required companion for a complete Seedance 2.5 Skill setup. If the link does not resolve, help the user install universal-video-prompt-skill before continuing; do not invent the missing shared specification.

Blocks

Use only the blocks that affect the shot:

[subject/reference binding]
+ [one observable action]
+ [space and important object relationships]
+ [one primary camera move, coherent composite move, or a cut]
+ [light/style when it matters]
+ [audio or dialogue when enabled]
+ [end state, for any stage that must land somewhere specific]
+ [must-preserve constraints]

End states carry multi-event work

For anything with more than one event, state what is visibly true when each stage ends. This is the highest-leverage single addition to a multi-stage prompt: it converts "keep it consistent" into something the model can target and you can check.

weak:   the two of them keep working on the bouquet
strong: end state: the florist holds the bouquet in the left hand;
        the scissors are back on the right side of the bench

An end state must be visible. "She feels relieved" is not one; "her shoulders drop and the frown clears" is. Read long video for the staged structure in full.

Time granularity: decide before writing beats

Granularity is a prior decision. Writing beats at second precision and then downgrading means rewriting them.

Granularity Use when
None — event order only One continuous action, mood pieces, single shots. Timestamps here fragment the shot
Stages + end states Most narrative work. Default
Second-level Only under an external hard constraint: music, lip sync, reference handoff, a beat that must land at a fixed time

Infer it when the input settles it — a supplied music or voiceover track means second-level, a stated mood piece means none, an explicit fixed beat means second-level. When the request is a multi-event narrative with no external constraint, ask, and recommend with a reason rather than presenting a bare menu.

Timestamps allocate a time budget; they are not frame-accurate edit points, and actions may land slightly before or after a boundary. Do not demand impossible density such as three distinct actions inside one second.

  • Name reference roles explicitly, for example Image 1: person, Image 2: product, Image 3: kitchen setting.
  • For multi-shot R2V, list Shot 1, Shot 2, and Shot 3 in event order. Set duration in provider controls rather than forcing exact seconds in text.
  • Prefer one primary movement per shot. A composite movement is valid when its direction, relation to the subject, and speed express one synchronized intent.
  • When native audio is enabled, use () for music, <> for sound effects, {} for dialogue, and 【】 for on-screen captions.
  • State only constraints that are costly to redo.

Read the reference matching the job:

File Read it for
prompt templates Route-specific templates
prompt blocks Reusable camera, audio, constraint patterns
long video Staged structure, end states, timestamp rules
multi reference Binding many assets without confusing them
real person Believable human subjects, and when to omit the detail
transitions Which transitions to generate and which to edit
editing and extension Changing or continuing existing video
capabilities 2.0 vs 2.5 limits; platform features vs API parameters
model profile Measured per-model behaviour and compile notes
cinematography Detailed visual decisions
troubleshooting Fault-specific fixes
execution adapters Runner and adapter configuration

5. Generate, review, and finish

  1. For a generated storyboard, create only the still first, display it in the conversation, and inspect it before any video request.
  2. For a supplied storyboard, display the input unless it is already visible, then verify it fits the chosen route.
  3. If the board passes review, generate one representative video pass without waiting for approval. Otherwise refine or regenerate the board first.
  4. Review identity, locks, stage end states, composition, motion, seam, and audio in that order, and stop at the first failure — later checks are wasted effort on a wrong identity. Regenerate only the failed shot or segment.
  5. Generate chains in order because the next segment needs the real prior tail. Generate cut-based clips independently and edit the planned transition.

The bundled script is a draft assembler, not a colour-grading or music-mixing system.

Atlas execution layer

Keep creative route selection independent from how a job is submitted. The default models are Seedream 5.0 Pro for stills and Seedance 2.0 for video; users may override them only after verifying route support.

In an agent conversation, use the Atlas Cloud Skill as the default direct generation route. It can discover a model, upload local media, submit an image or video request, poll, and retrieve outputs. Report Execution: atlas-skill only when it actually submitted the generation.

Use atlas-mcp only when the user explicitly selects MCP and its generation tools are exposed. Use atlas-cli only when the user explicitly selects a terminal, script, CI, or batch run. If the Atlas Cloud Skill is missing, help install AtlasCloudAI/atlas-cloud-skills before selecting a fallback.

Before reporting that an Atlas Cloud API key is missing, check the credentials in the selected execution process. For the REST runner, check ATLASCLOUD_API_KEY first and ATLAS_CLOUD_API_KEY as a compatibility alias. Do not infer credential availability from a different provider, plugin, or process; each execution channel can have an independent credential scope.

If neither key exists, direct the user to https://www.atlascloud.ai/console/api-keys?utm_source=github&utm_campaign=awesome-seedance-2.5-prompts-skills. Never ask them to paste the key in chat. Tell them to set ATLASCLOUD_API_KEY in the submitting process or the host's secure environment settings, then refresh or restart the execution session if needed. If the key exists in a parent or host configuration but is absent from the submitting process, report an environment-scope mismatch instead of saying the user has no key.

Billable task state machine

Apply these rules to every image and video generation:

  1. After submission, record the prediction ID and logical stage immediately.
  2. Treat starting, queued, pending, and processing as active. Poll the same ID every 2 seconds; never submit another task for that stage.
  3. Treat completed and succeeded as successful terminal states. Download and inspect the output before starting a dependent stage.
  4. Treat failed, timeout, and canceled as terminal failures. A new task requires an explicit retry decision; report the old ID and possible extra cost first.
  5. A zero or missing processing-time field, delayed output, a local polling timeout, a stopped turn, or a temporary status-query error is not proof of failure. Preserve the ID and resume polling.
  6. Interpret continue as “resume the existing task,” never as permission to retry. Do not submit video while its required storyboard is still active.

The 2-second interval applies to every Atlas execution route in this workflow. For atlas-skill, repeat its prediction-result step with the same ID. For atlas-mcp, call atlas_get_prediction with the same ID every 2 seconds. The MCP server performs a single status lookup per tool call; the agent owns the loop. The bundled REST and CLI adapters enforce the interval in code. A status lookup is read-only and must never be replaced with another generation call.

When the runner must resume, set execution.resumePredictionIds.<stage> to the existing ID. Supported stage keys include grid, ref1, ref2, seg1, shot1, and clip1. Never create a replacement merely because a prior polling process ended.

scripts/generate.mjs cannot invoke an agent Skill or MCP server. It is a separate batch runner: it defaults to atlas-rest, and can use atlas-cli only when explicitly set in execution.adapter. It never selects CLI automatically. See execution adapters.

# Generate a storyboard only. The runner prints [storyboard-preview] with an
# absolute path; display and inspect that image before starting video work.
GRID_ONLY=1 node scripts/generate.mjs scripts/myjob.json

# Run only the first independent clip or segment as a quality gate.
CLIPS_MAX=1 node scripts/generate.mjs scripts/myjob.json
SEGS_MAX=1 node scripts/generate.mjs scripts/myjob.json
{ "execution": { "adapter": "atlas-cli" } }

Available runner modes:

  • grid: crop a model storyboard and make one independent I2V clip per panel; suitable for a hard-cut montage.
  • shot-pairs: make independent I2V shots from segments[], with a first and optional last keyframe index.
  • reference: send a small role-specific set of R2V references. A reference may be generated from a prompt or read from a local file.
  • chain: continue one action by feeding the generated tail frame to the next segment. It is an alignment aid, not a guarantee of an invisible seam.

For fault-specific fixes, read troubleshooting.

Files (atlas-cloud-skills)
  • references
    • capabilities.md 5.1 KB
      # Capabilities: 2.0 versus 2.5, and what to promise
      
      Two separate questions get conflated constantly:
      
      1. What did the **model launch** announce?
      2. What can **this provider's API** actually accept today?
      
      Treat every number below as the first until verified as the second. Announced
      capability is not an API contract, and several widely-quoted 2.5 headline
      capabilities are platform features rather than API parameters.
      
      ## Reference material limits
      
      | Type | 2.0 | 2.5 |
      |---|---|---|
      | Images per request | 0–9 | up to 30 |
      | Video references | up to 3, single clip 2–15s, total ≤15s | up to 10, single clip 2–30s, total ≤30s |
      | Audio references | up to 3, single clip 2–15s, total ≤15s | up to 10, single clip 2–30s, total ≤30s |
      | Audio-only reference | not supported — at least one image or video required | supported |
      | Combined ceiling | — | around 50 materials |
      
      Effective clip bounds run marginally wider than the nominal ones (roughly 1.8s at
      the low end, and slightly over the stated ceiling at the top) because the output
      runs a little longer than the requested duration.
      
      ## Duration
      
      | | 2.0 | 2.5 |
      |---|---|---|
      | Range | 4–15s | 4–30s |
      
      ## Resolution
      
      2.0 exposes 480p and 720p. **For 2.5, read the provider's model page** — launch
      material describes higher output than 2.0, but the enumerated values differ by
      provider and are the only thing worth quoting. Do not restate launch-material
      resolution as an API parameter.
      
      ## Stability ranges
      
      Recommended ranges are about stability, not hard limits. Exceeding them is
      allowed and gets less predictable — expect to regenerate more.
      
      | Input | Stable | Possible, less reliable |
      |---|---|---|
      | Distinct subjects in subject images | 1–8 | 9–12 |
      | Distinct subjects in subject audio/video | 1–5 | 6–10 |
      | Reference clip duration per subject | 5–10s | longer |
      | Source video for editing | under ~20s | longer |
      | Reference images for video editing | 1–5 | 6–8 |
      
      **Views:** up to about five subjects, single-view and multi-view both work. Past
      that, prefer single-view. When several views are needed, separate images per view
      are more stable than one collage.
      
      ## Automatically locked parameters
      
      Some task types derive parameters from the input and will not let you set them:
      
      | Task | Aspect ratio | Duration |
      |---|---|---|
      | Video editing | Inherits the source; cannot be set | Approximately the source's; cannot be set. Frame handling can shift it by up to ~0.3s |
      | First frame, or first-and-last frame | Inherits the **first** image | Can be set |
      | Video extension | Inherits the source | Extension length can be set |
      
      For first-and-last-frame work, give both images the **same aspect ratio** —
      mismatched ratios stretch the last frame.
      
      ## Platform features versus API parameters
      
      The following appear in 2.5 launch and product documentation. They are features of
      the **first-party creation platform**. Whether any of them is reachable through a
      given API depends entirely on that provider, and several are UI-driven by nature:
      
      | Feature | Why it may not be an API parameter |
      |---|---|
      | Long-video mode well beyond 30s in one pass | A distinct product mode, not a duration argument |
      | Nested extension stacking past a single request's ceiling | An iterative UI flow |
      | Mark-based editing (box select, brush, anchor points) | Requires on-frame annotation input |
      | DCC blockout plugins (Maya, Blender) | A separate integration, not a model parameter |
      | One-click assembly from a set of images | A product workflow above the model |
      | Seamless bridging between two finished clips | May or may not be exposed |
      
      **Do not quote these as model specifications.** The common error is repeating a
      maximum duration from launch material as though any API call can request it. State
      what the provider's model page enumerates; describe the rest as announced platform
      capability if it needs mentioning at all.
      
      ## Before offering a route
      
      1. Confirm the provider exposes the model. A 200 response is not confirmation —
         fallback pages return 200 too. Check the page's title and body, not the status
         code, and never infer an endpoint from a guessed org/model slug.
      2. Read the enumerated parameters from that page rather than from doc examples.
         Parameter names sometimes differ between an external reference and a live page;
         when they disagree, the live page wins.
      3. Confirm which task types are exposed. Text-to-video availability says nothing
         about whether editing or extension is available.
      4. Record what you found in [model profile](model-profile.md) so the next run does
         not re-probe.
      
      ## What no capability tier fixes
      
      - Text that must read exactly — subtitles, formulas, signage, product specs.
        Prepare it as an asset or add it in post.
      - Frame-accurate timing. Timestamps allocate a budget, not an edit point.
      - Pixel-identical preservation across an edit or a boundary. Editing preserves
        content and event order substantially, not exactly.
      
      ## Related
      
      - [model profile](model-profile.md) · [long video](long-video.md) ·
        [multi reference](multi-reference.md) ·
        [editing and extension](editing-and-extension.md)
      
    • capabilities.zh-CN.md 4.9 KB
      # 能力:2.0 与 2.5,以及该承诺什么
      
      两个完全不同的问题经常被混在一起:
      
      1. **模型发布**宣布了什么?
      2. **这家服务的 API** 今天真正能接受什么?
      
      下面所有数字在被核实为第二类之前,都按第一类对待。宣布的能力不是 API 契约,而且几条被广泛引用的 2.5 头部能力属于**平台功能**而非 API 参数。
      
      ## 参考素材上限
      
      | 类型 | 2.0 | 2.5 |
      |---|---|---|
      | 每次请求图片数 | 0–9 | 最多 30 |
      | 参考视频 | 最多 3 条,单条 2–15s,合计 ≤15s | 最多 10 条,单条 2–30s,合计 ≤30s |
      | 参考音频 | 最多 3 条,单条 2–15s,合计 ≤15s | 最多 10 条,单条 2–30s,合计 ≤30s |
      | 仅传音频 | 不支持——至少需要一张图或一条视频 | 支持 |
      | 合计上限 | — | 约 50 份素材 |
      
      实际的单条边界比标称值略宽(低端约 1.8s,高端略超标称上限),因为输出会比请求时长略长一点。
      
      ## 时长
      
      | | 2.0 | 2.5 |
      |---|---|---|
      | 范围 | 4–15s | 4–30s |
      
      ## 分辨率
      
      2.0 提供 480p 与 720p。**2.5 请读所在服务的模型页**——发布材料描述了比 2.0 更高的输出,但枚举值各家不同,而且只有枚举值值得引用。不要把发布材料里的分辨率当成 API 参数复述。
      
      ## 稳定范围
      
      推荐范围关乎稳定性,不是硬上限。超出是允许的,只是可预测性下降——预期要多重跑几次。
      
      | 输入 | 稳定 | 可行但不稳 |
      |---|---|---|
      | 主体图里的不同主体 | 1–8 个 | 9–12 个 |
      | 主体音/视频里的不同主体 | 1–5 个 | 6–10 个 |
      | 每个主体的参考片时长 | 5–10s | 更长 |
      | 编辑用的源视频 | 20 秒以内 | 更长 |
      | 视频编辑的参考图 | 1–5 张 | 6–8 张 |
      
      **视角**:约五个主体以内,单视角与多视角都行。超过之后优先单视角。需要多个视角时,一个视角一张图比拼成一张更稳。
      
      ## 会被自动锁定的参数
      
      有些任务类型从输入推导参数,不允许你设定:
      
      | 任务 | 画幅 | 时长 |
      |---|---|---|
      | 视频编辑 | 继承源片,不能单独设 | 约等于源片,不能单独设。帧处理可能带来最多约 0.3 秒差异 |
      | 首帧、或首尾帧 | 继承**第一张**图 | 可以设 |
      | 视频延长 | 继承源片 | 延长长度可以设 |
      
      首尾帧的两张图要用**相同画幅**——画幅不一致会拉伸尾帧。
      
      ## 平台功能与 API 参数
      
      以下内容出现在 2.5 的发布与产品文档里。它们是**第一方创作平台的功能**。任何一项能否通过某个 API 触达,完全取决于那家服务,而且其中几项本质上是 UI 驱动的:
      
      | 功能 | 为什么可能不是 API 参数 |
      |---|---|
      | 远超 30s 的长视频模式一次直出 | 那是一个独立的产品模式,不是时长参数 |
      | 嵌套式延长叠加超过单次请求上限 | 一个迭代式的 UI 流程 |
      | 标注式编辑(框选、笔刷、锚点) | 需要在画面上做标注输入 |
      | DCC 白模插件(Maya、Blender) | 一个独立集成,不是模型参数 |
      | 从一组图片一键成片 | 模型之上的产品工作流 |
      | 两条成片之间的无缝桥接 | 可能暴露、也可能不暴露 |
      
      **不要把这些当成模型规格引用。** 常见错误是把发布材料里的最大时长当作任何 API 调用都能请求的值。只陈述该服务模型页枚举出来的东西;其余若必须提及,描述为「已宣布的平台能力」。
      
      ## 提供某条路线之前
      
      1. 确认该服务真的暴露了这个模型。**200 响应不算确认**——兜底页也返回 200。查页面标题与正文,不是状态码;绝不要从猜测的 org/model slug 推断端点
      2. 从那个页面读枚举参数,而不是从文档样例。参数名有时在外部文档与实际页面之间不一致;两者冲突时**以实际页面为准**
      
         ⚠️ **但页面自己的参数表也可能是错的。** 本平台已确认三次:文档写的范围**比实际能提交的更窄**(H3 t2v 文档写 `5-10`、实测 15s 可用;Kling 两个档位都写 `5 or 10`、实测 15s 均可用)。**直接提交你真正想要的值,让 API 来回答**——被拒不创建任务、不花钱
      3. 确认哪些任务类型被暴露。文生视频可用,完全不代表编辑或延长可用
      4. 把查到的结果记进 [model-profile.zh-CN.md](model-profile.zh-CN.md),下次运行不必重新探测
      
      ## 任何能力档次都解决不了的事
      
      - 必须准确无误的文字——字幕、公式、招牌、产品规格。做成素材或后期加上
      - 帧精确的时序。时间戳分配预算,不是剪辑点
      - 编辑或边界处的逐像素保留。编辑大体上保留内容与事件顺序,不是精确保留
      
      ## 关联
      
      - [model-profile.zh-CN.md](model-profile.zh-CN.md) · [long-video.zh-CN.md](long-video.zh-CN.md) ·
        [multi-reference.zh-CN.md](multi-reference.zh-CN.md) ·
        [editing-and-extension.zh-CN.md](editing-and-extension.zh-CN.md)
      
    • cinematography.md 6.9 KB
      # Cinematography vocabulary: camera, light, and composition
      
      Use this vocabulary to describe how a shot is filmed. It is a creative library,
      not an API parameter list or an outcome guarantee.
      
      ## Working rules
      
      1. Pick one primary movement per shot. Composite movement is valid when
         direction, subject relation, and speed form one intention.
      2. Use `Shot 1`, `Shot 2`, and `Shot 3` for storyboard order; set duration in
         provider controls instead of writing a timestamp schedule in the prompt.
      3. Tight close-ups combined with large rotations are less stable than wider
         framing or a planned cut.
      4. Choose the frame ratio and safe areas for the user's actual delivery context.
      
      ## Uncommon terms: keep the term, add the description
      
      Terms whose recognition varies — niche vocabulary, terms with inconsistent
      industry usage, anything named after a film, director, or platform trend — get
      written twice:
      
      ```text
      <term> + <target subject> + <visible change> + <foreground/background relation>
            + <direction or speed>
      ```
      
      ```text
      rack focus: shift focus from the foreground leaves to the person behind them.
        The leaves go soft while the face resolves from soft to sharp.
      
      bullet time: freeze the moment bat meets ball; the camera orbits clockwise
        around the contact point while debris hangs in the air.
      ```
      
      A model that knows the term takes the shortcut; one that does not follows the
      description. The same prompt then works on both, which is why this is preferable
      to maintaining a per-model vocabulary list.
      
      **Usually safe unqualified:** shot scales, the basic moves in the table below,
      basic positions (low angle, overhead, first-person).
      
      **Usually needs the description:** dolly zoom, bullet time, speed ramp, bounce
      ramp, rack focus, whip-pan transition, match cut.
      
      Aperture, focal length, and shutter values are allowed, but the intended visible
      result controls more reliably than a number alone.
      
      ## Camera
      
      ### Shot scale
      
      | Scale | What it shows | Typical use |
      |---|---|---|
      | Extreme close-up | Eye, mouth, droplet, texture | Tension or a key material detail |
      | Close-up | Face or one object | Emotion, decision, product texture |
      | Medium close-up | Head and shoulders | Dialogue and expression |
      | Medium shot | Waist-up subject | Everyday action and expression |
      | Wide shot | Full person and setting | Space and full-body action |
      | Extreme wide / establishing shot | Whole environment | Opening tone, scale, or release |
      
      Move wide to close to concentrate attention; move close to wide to release or
      finish. Open with a scale that reads immediately, such as a strong detail or
      an establishing shot with clear scale.
      
      ### Camera angle
      
      - Eye level: objective, equal, natural.
      - Low angle: power, heroism, pressure.
      - High angle: vulnerability, overview, observation.
      - Top-down: pattern, layout, preparation, spatial explanation.
      - First-person view: participation and immersion.
      
      ### Movement
      
      | Movement | Effect | Best use | Caution |
      |---|---|---|---|
      | Slow push-in | Focus and rising emotion | Reveal, decision, key detail | Avoid combining with a large orbit |
      | Slow pull-out | Release and closure | Ending, spatial reveal | Keep the subject readable |
      | Smooth lateral move | Smooth display or accompaniment | Product sweep, side-following walk | Avoid unexplained speed changes |
      | Tracking shot | Motion and immersion | Run, walk, drive | Excessive handheld shake can distract |
      | Pan or tilt | Guide attention and reveal | Move from A to B | Fast moves can blur |
      | Crane move | Scale and emotional rise/fall | Opening lift, ending descent | Preserve a clear subject anchor |
      | Smooth orbit | Three-dimensional display | Product or medium/wide character | Avoid large close-up face rotations |
      | Handheld drift | Documentary tension or realism | Street, suspense, lived-in scene | Avoid for polished product work |
      | Static camera | Stability and observation | Dialogue, still life, atmosphere | Give a person a subtle natural action |
      
      Examples of coherent composite movement: `track the person with a steady
      rightward lateral move`; `slowly push in while rising slightly to reveal the
      subject`. Do not stack unrelated push-ins, orbits, and pans without a shared
      purpose.
      
      ### Lens, depth, and transitions
      
      - Shallow depth of field: isolates the subject and gives texture or intimacy.
      - Deep focus: keeps environment and subject clear for documentary or landscape.
      - Wide lens feeling: spatial energy, perspective, slight distortion.
      - Long-lens feeling: compressed space and background separation.
      - Hard cut: rhythm, energy, process montage.
      - Dissolve: genuine change of time or place.
      - One take: one continuous movement chain with compatible segment boundaries.
      
      ## Lighting
      
      ### Lighting setups
      
      - Three-point lighting: key, fill, and rim; a reliable dimensional baseline.
      - Rembrandt light: a 45-degree side key with a small cheek triangle; dramatic portraiture.
      - Butterfly light: high frontal key; beauty and fashion.
      - Backlit silhouette: light behind the subject; atmosphere, emotion, epic scale.
      - Low key: mostly dark with controlled highlights; suspense, force, luxury.
      - High key: bright, low-shadow image; clean, fresh, commerce.
      - Practical light: a visible source such as lamp, stove, neon, or window; natural atmosphere.
      
      Use one or two visual anchors in a video prompt: for example, `warm practical
      light from the stove, soft rim light` or `cool neon backlight with a magenta
      accent`. Keep key direction and warm/cool relationships consistent across shots
      that share one place and time.
      
      ### Palette and mood
      
      | Palette | Visual anchor | Use |
      |---|---|---|
      | Golden hour | Warm key, deep-blue shadows | Warmth, romance, final release |
      | Blue hour | Cool environment, small warm points | Quiet, solitude, night opening |
      | Noon hard light | Bright highlights, hard shadows | Reality, sport, documentary |
      | Overcast soft light | Diffuse grey-white light, low saturation | Poetic, subdued, gentle |
      | Neon night | Magenta and cyan contrast | Cyber city, fashion, nightlife |
      | Warm kitchen | Yellow practical light, wood brown | Food, comfort, home |
      | Cool technology | Cool white key, restrained brand colour | Clean product or finance |
      
      ## Composition
      
      - Rule of thirds: stable general-purpose placement.
      - Central symmetry: ritual, product hero, confrontation.
      - Leading lines: direct attention with architecture, props, or light.
      - Frame within a frame: use doorways, windows, or foreground objects for depth.
      - Diagonal or triangular arrangement: movement or stable multi-subject layout.
      - Foreground, midground, and background: plan three layers to avoid a flat image.
      - One visual centre: one brightest, sharpest, or most active focal point.
      - Negative space: leave room when simplicity and focus are the goal.
      
      For a key product moment, place the decisive detail clearly in the composition,
      give it a compatible light cue, and use a restrained push-in. Never impose a
      single frame ratio or universal safe-area percentage; determine both from the
      user's intended delivery context.
      
    • cinematography.zh-CN.md 6 KB
      # 电影语言词库:运镜、光影与构图
      
      本词库用于描述镜头“怎么拍”。它是创作词库,不是 API 参数,也不保证模型一定实现。
      
      ## 使用原则
      
      1. 每个镜头先确定一个主导运镜。复合运镜可以使用,但方向、主体关系和速度必须共同服务于一个意图。
      2. Storyboard 用 `镜头1`、`镜头2`、`镜头3` 表达顺序;时长交给服务参数,不在提示词中硬写逐秒计划。
      3. 近距离脸部特写配大幅旋转更不稳定,优先改用较宽景别或计划好的剪辑。
      4. 画幅和安全区由用户实际投放场景决定,不预设固定比例。
      
      ## 冷僻术语:保留术语,追加描述
      
      识别度不确定的术语——冷僻词、行业用法不统一的词、以电影/导演/平台风潮命名的词——都写两遍:
      
      ```text
      <术语> + <目标主体> + <可见变化> + <前后景关系> + <方向或速度>
      ```
      
      ```text
      rack focus:焦点从前景的叶片移到它身后的人。叶片转虚,人脸由虚转实。
      bullet time:冻结球棒击中球的瞬间;镜头顺时针绕接触点环绕,碎屑悬在空中。
      ```
      
      认识术语的模型走捷径,不认识的走描述。同一份提示词在两类模型上都成立,这也是它优于维护逐模型词表的原因。
      
      **通常可以不加限定的**:景别、下表里的基础运镜、基础机位(低角度、俯拍、第一人称)。
      
      **通常需要加描述的**:dolly zoom、bullet time、speed ramp、bounce ramp、rack focus、whip-pan 转场、match cut。
      
      光圈、焦段、快门值可以写,但**说明想要的可见结果通常比单给数值更可控**。
      
      ## 镜头
      
      ### 景别
      
      | 景别 | 画面范围 | 常见用途 |
      |---|---|---|
      | 极致特写 | 眼睛、嘴、水珠、纹理 | 张力或关键材质细节 |
      | 特写 | 脸或单个物体 | 情绪、决心、产品质感 |
      | 中近景 | 头肩范围 | 对话和表情 |
      | 中景 | 腰部以上 | 日常动作与表情 |
      | 全景 | 完整人物与环境 | 空间和全身动作 |
      | 大远景或建立镜头 | 完整环境 | 开场定调、气势、收尾 |
      
      远到近会收拢注意力;近到远会释放或收尾。开场优先选择一眼能读懂的景别,例如强细节或具有尺度感的建立镜头。
      
      ### 机位
      
      - 平视:客观、平等、自然。
      - 低角度仰拍:力量、英雄感、压迫感。
      - 高角度俯拍:脆弱、总览、观察。
      - 顶视:图案感、布局、备料、空间说明。
      - 第一人称:参与感与沉浸感。
      
      ### 运镜
      
      | 运镜 | 传达 | 适合 | 注意 |
      |---|---|---|---|
      | 缓慢推近 | 聚焦、情绪升温 | 揭示、决心、关键展示 | 不要随意叠加大幅环绕 |
      | 缓慢拉远 | 释放、收尾 | 结尾、空间揭示 | 保持主体可读 |
      | 平稳横移 | 平滑展示、陪伴 | 产品扫过、侧向跟随 | 避免无理由变速 |
      | 跟拍 | 动感、代入 | 奔跑、行走、驾驶 | 过强手持会分散注意 |
      | 摇镜 | 引导视线、揭示 | 从 A 到 B | 快速摇镜容易模糊 |
      | 升降 | 气势、情绪起落 | 开场抬升、结尾下降 | 保持清晰主体锚点 |
      | 平滑环绕 | 立体展示 | 产品或中远景人物 | 避免脸部近景大幅旋转 |
      | 手持漂移 | 纪实、紧张、真实 | 街头、悬疑、生活感 | 不适合精致产品展示 |
      | 固定机位 | 稳定、观察、留白 | 对话、静物、氛围 | 人物需要自然微动作 |
      
      可用的复合运镜,例如 `镜头向右平稳横移,保持跟随人物`,或者 `缓慢推近并轻微升起以揭示主体`。不要把推近、环绕和摇镜等无关动作随意叠加。
      
      ### 焦段、景深与转场
      
      - 浅景深:突出主体,适合人物、产品和材质细节。
      - 大景深:主体和环境都清晰,适合纪实和大场景。
      - 广角感:空间张力、透视、轻微畸变。
      - 长焦感:压缩空间、背景分离。
      - 硬切:节奏、能量、工序蒙太奇。
      - 叠化:真正的时间或地点变化。
      - 一镜到底:同一条连续运镜链,首尾状态必须兼容。
      
      ## 光影
      
      ### 布光
      
      - 三点布光:主光、补光、轮廓光,稳定且有层次。
      - 伦勃朗光:侧前方主光,适合戏剧人物质感。
      - 蝴蝶光:上方正面主光,适合美妆和时尚。
      - 逆光剪影:主体背后有光,适合情绪和史诗感。
      - 低调光:大面积暗、局部亮,适合悬疑、力量和奢华。
      - 高调光:整体明亮、阴影少,适合干净、清新和电商。
      - 实用光:画面内可见光源,例如台灯、灶火、霓虹或窗光,最自然。
      
      视频提示词通常只取一到两个视觉锚点,例如 `灶火提供暖色实用光,柔和轮廓光`,或 `冷色霓虹逆光,少量洋红点缀`。同一时空连续镜头保持主光方向和冷暖关系一致。
      
      ### 调色与情绪
      
      | 调色板 | 视觉锚点 | 用途 |
      |---|---|---|
      | 黄金时刻 | 暖金主光、深蓝阴影 | 温暖、浪漫、收尾 |
      | 蓝调时刻 | 冷蓝环境、少量暖点光 | 静谧、孤独、夜晚开场 |
      | 正午硬光 | 明亮高光、清晰硬阴影 | 真实、运动、纪实 |
      | 阴天柔光 | 灰白漫反射、低饱和 | 文艺、克制、温柔 |
      | 霓虹夜 | 品红与青蓝对比 | 都市、时尚、赛博 |
      | 暖厨房 | 暖黄实用光、木棕环境 | 美食、治愈、居家 |
      | 冷科技 | 冷白主光、克制品牌色 | 干净产品或金融场景 |
      
      ## 构图
      
      - 三分法:稳定的通用构图。
      - 中心对称:仪式感、产品主视觉、对峙。
      - 引导线:用建筑、道具或光线引导视线。
      - 框中框:用门窗或前景物增加纵深和聚焦。
      - 对角线或三角构图:制造动感或稳定多主体排布。
      - 前中后景:安排三层空间,避免画面扁平。
      - 单一视觉重心:只保留一个最亮、最清晰或最活跃的焦点。
      - 留白:在需要极简和聚焦时留出呼吸空间。
      
      关键产品瞬间应让卖点在构图中清楚可见,配合兼容的光影提示和克制推近。不要预设固定画幅或通用安全区百分比,必须根据用户实际投放场景决定。
      
    • editing-and-extension.md 5.6 KB
      # Editing and extension
      
      Three jobs that all start from existing video. What they share: the source video
      is the authority, and the prompt's job is to say precisely **what changes and what
      must not.**
      
      Editing and extension also lock some parameters automatically based on the input.
      Check [capabilities](capabilities.md) before promising an aspect ratio or duration.
      
      ## Editing an existing video
      
      ```text
      [Edit goal]
      Edit @video1. Within <the whole video | a time range>, <add | remove | replace |
      adjust> <object, region, or audio category>.
      
      [Source role]
      @video1 is the sole editing master. It defines <characters, scene, actions,
      composition, camera movement, occlusion relationships, audio, event order>.
      
      [Target material role]
      @image1 defines <the specified attributes of the target>.
        Do not use <its background, composition, or unrelated objects>.
      
      [Edit scope]
      Modify only <object, region, time range, or audio category>.
      
      [Preserve]
      Keep <everything that must not change> from @video1.
      ```
      
      Four things make this work:
      
      **Sole editing master.** Say it. Without it, supplied reference images compete
      with the source for control of composition and staging.
      
      **Scope stated twice** — once as what to change, once as what to keep. The second
      half is not redundant; it is what stops collateral changes.
      
      **Target count.** When replacing an object, state how many exist:
      `exactly one white lamp appears throughout; replace only the original yellow one`.
      
      **Timeline inheritance.** A replaced object inherits everything the original did:
      
      ```text
      The white lamp inherits every appearance, rotation, hand occlusion, and exit of
      the original yellow lamp, including timing, path, and speed changes.
      ```
      
      Without that line the replacement often appears correctly but moves differently.
      
      ### Natural responses are allowed
      
      Scope should not be so tight that the result looks pasted. When changing light
      colour, allow skin tone to respond: `change only the light colour on the right
      wall and the area it illuminates; allow the character's skin tone to respond
      naturally to the environmental light.`
      
      ### Audio can be edited separately
      
      Dialogue, language, voice, music, and effects are independently addressable. Name
      the speaker or sound category, the change, and what must stay:
      
      ```text
      Edit @video1. Remove only the original background music. Keep the character
      dialogue, lip sync, ambience, and action sound effects; preserve the visuals,
      camera treatment, and editing rhythm from @video1.
      ```
      
      ## Forward extension — after the source
      
      Describe the boundary state first, then what happens next.
      
      ```text
      Extend @video1 forward. The first frame of the extended segment continues directly
      from the last frame of @video1. Maintain continuity in <subject pose and
      orientation>, <prop position>, <background and spatial relationships>, <camera
      position and composition>, <lighting>, and <motion direction>.
      
      Then, <the new action, event, camera treatment, or audio>.
      
      Throughout, maintain <character identity and clothing>, <key props>, <background
      layout>, and <axis of action>. Keep each subject one continuous instance: do not
      duplicate or split it.
      ```
      
      This is the context where seam defaults are correct: **require natural
      continuation, smooth motion connection, no rigid cut, no objects appearing from
      nowhere.** Those constraints belong here, not globally — see
      [transitions](transitions.md).
      
      ## Backward extension — before the source
      
      The direction that goes wrong, and the reason is specific.
      
      Describe the new content first, then state the source's **first frame as the
      explicit end state** of what you are generating:
      
      ```text
      Extend @video1 backward. Before the source begins, <the preceding action or event>.
      
      The last frame of the extended segment connects naturally to the first frame of
      @video1: <subject pose and orientation>, <prop position and state>, <other
      subjects' positions>. Match the <camera position and composition>, <lighting>, and
      <motion direction> of @video1's first frame.
      
      <Materials that should only appear after the source begins> must not appear early
      in the backward extension.
      ```
      
      **Writing only "then it connects to the source video" is the failure mode.** Two
      things go wrong: characters or effects that belong *after* the source show up
      early, and the image keeps changing after it has already reached the target state.
      Both are fixed by treating the source's first frame as a stated end state — the
      same device that carries staged long video.
      
      The last line matters when references are supplied. Say which materials are for the
      preceding segment and which must wait.
      
      ## Chaining
      
      For an action longer than one request supports:
      
      - Split where a stage genuinely ends, not mid-action.
      - Feed the **real generated tail** forward — not the tail you intended.
      - Generate chained segments in order; they depend on each other.
      - Generate cut-separated shots independently and assemble in the edit.
      
      A chain is an alignment aid, not a guarantee of an invisible seam.
      
      ## Review
      
      Boundary frames connect **visually**. They are not pixel-identical splices, and
      expecting that leads to endless regeneration of acceptable output.
      
      Check, in order:
      
      1. Both sides of the boundary — pose, props, light, motion direction
      2. The complete extended segment, not only the seam
      3. Whether anything appeared early (backward extension) or persisted wrongly
      4. Audio continuity across the boundary; the extended segment's level may differ
         slightly from the source
      
      Fix the failed segment only.
      
      ## Related
      
      - [long video](long-video.md) — the end-state device this reuses
      - [transitions](transitions.md) — where seam constraints belong
      - [capabilities](capabilities.md) — what these tasks lock automatically
      
    • editing-and-extension.zh-CN.md 5 KB
      # 视频编辑与延长
      
      三类都从已有视频出发。共同点:**源视频是权威,提示词的任务是精确说明什么变、什么绝不能变。**
      
      编辑与延长还会根据输入**自动锁定**部分参数。承诺画幅或时长之前先看 [capabilities.zh-CN.md](capabilities.zh-CN.md)。
      
      ## 编辑已有视频
      
      ```text
      [编辑目标]
      编辑 @video1。在<整段视频 | 某个时间范围>内,<添加 | 移除 | 替换 | 调整>
      <物体、区域或音频类别>。
      
      [源视频角色]
      @video1 是唯一编辑母版。它定义<人物、场景、动作、构图、运镜、遮挡关系、音频、
      事件顺序>。
      
      [目标素材角色]
      @image1 定义<目标物体的指定属性>。
        不要用<它的背景、构图或无关物体>。
      
      [编辑范围]
      只修改<物体、区域、时间范围或音频类别>。
      
      [需要保留]
      保留 @video1 的<绝不能变的全部内容>。
      ```
      
      四件事让它成立:
      
      **唯一编辑母版。** 要写出来。不写这句,提供的参考图会跟源视频争夺构图与调度的控制权。
      
      **范围写两遍** —— 一遍是要改什么,一遍是要保留什么。第二半不是冗余,它是防连带改动的东西。
      
      **目标数量。** 替换物体时写明存在几个:`全片只有一盏白色台灯;只替换原来那盏黄色的`。
      
      **时间轴继承。** 被替换的物体继承原物体做过的一切:
      
      ```text
      白色台灯继承原来那盏黄色台灯的全部出现、灯臂旋转、手部遮挡与离场,
      包括时机、路径与速度变化。
      ```
      
      不写这行,替换物经常外形正确但运动方式不一样。
      
      ### 允许自然反应
      
      范围不该收得太紧、以致结果看起来像贴上去的。改灯光颜色时,允许皮肤跟着环境光变:`只改右墙的灯光颜色及其照亮的区域;允许人物肤色对环境光作出自然反应。`
      
      ### 音频可以单独编辑
      
      台词、语言、音色、音乐、音效各自可寻址。写明说话人或声音类别、要改什么、什么必须保留:
      
      ```text
      编辑 @video1。只移除原有背景音乐。保留人物台词、唇形同步、环境声与动作音效;
      保留 @video1 的画面、运镜与剪辑节奏。
      ```
      
      ## 向前延长 —— 接在源片之后
      
      先描述边界状态,再写接下来发生什么。
      
      ```text
      向前延长 @video1。延长段的第一帧直接承接 @video1 的最后一帧。保持
      <主体姿态与朝向>、<道具位置>、<背景与空间关系>、<机位与构图>、<光照>、
      <运动方向>的连续性。
      
      然后,<新的动作、事件、运镜或音频>。
      
      全程保持<人物身份与服装>、<关键道具>、<背景布局>、<动作轴线>的连续性。
      每个主体保持为同一个连续实体:不许复制或分裂。
      ```
      
      这正是接缝默认值成立的上下文:**要求自然延续、动作衔接顺畅、禁止生硬切断、禁止物体凭空出现。** 这些约束属于这里,不属于全局——见 [transitions.zh-CN.md](transitions.zh-CN.md)。
      
      ## 向后延长 —— 接在源片之前
      
      会出错的那个方向,原因很具体。
      
      先写新内容,再把源片的**首帧写成你正在生成这一段的明确末态**:
      
      ```text
      向后延长 @video1。在源视频开始之前,<在前面发生的动作或事件>。
      
      延长段的最后一帧自然衔接到 @video1 的第一帧:<主体姿态与朝向>、
      <道具位置与状态>、<其他主体的位置>。匹配 @video1 首帧的<机位与构图>、
      <光照>、<运动方向>。
      
      <只应在源视频开始之后出现的素材>不得在向后延长段中提前出现。
      ```
      
      **只写「然后接上源视频」就是那个翻车写法。** 两件事会出错:属于源片**之后**的角色或效果提前露面;以及画面到达目标状态之后继续变。两者都靠把源片首帧当成写明的末态来解决——正是承载分段长视频的同一个装置。
      
      最后那行在有参考素材时尤其重要。说清哪些素材属于前面这一段、哪些必须等。
      
      ## 接龙
      
      一个动作长于单次请求能支持的范围时:
      
      - 在某一段**真正结束**的地方切,不在动作中途切
      - 把**真实生成的尾帧**往前喂——不是你设想的那个尾帧
      - 接龙段落**按顺序**生成,它们互相依赖
      - 被切点分开的镜头独立生成,在剪辑台上拼
      
      接龙是对齐辅助,**不保证接缝隐形**。
      
      ## 复查
      
      边界帧在**视觉上**衔接,它们不是逐像素的剪辑拼接——期待那个会导致对可用输出的无限重做。
      
      按顺序检查:
      
      1. 边界两侧——姿态、道具、光、运动方向
      2. 完整的延长段,不只是接缝
      3. 有没有东西提前出现(向后延长),或错误地留存
      4. 跨边界的音频连续性;延长段的音量可能与源片略有差异
      
      只重做失败的那一段。
      
      ## 关联
      
      - [long-video.zh-CN.md](long-video.zh-CN.md) — 这里复用的末态装置
      - [transitions.zh-CN.md](transitions.zh-CN.md) — 接缝约束的正确作用域
      - [capabilities.zh-CN.md](capabilities.zh-CN.md) — 这些任务会自动锁定什么
      
    • execution-adapters.md 6.4 KB
      # Atlas execution adapters
      
      Creative planning and the model profile are independent from the execution
      channel. Default profile: Seedream 5.0 Pro for stills and Seedance 2.0 for
      video. Verify that a chosen model supports the required route before a billable
      generation.
      
      ## Default agent route
      
      In an agent conversation, the **Atlas Cloud Skill** is the default direct
      generation route. Use it to discover models, inspect parameters, upload local
      assets, submit image or video generation, poll, and retrieve results. Report
      `Execution: atlas-skill` only when the Skill actually submitted the job.
      After submission, repeat the Skill's prediction-result step with the same ID
      every 2 seconds until a terminal state is returned.
      
      If it is unavailable, help install it before choosing another route:
      
      ```bash
      npx skills add AtlasCloudAI/atlas-cloud-skills
      ```
      
      Then refresh the agent's Skill registry if the host requires it.
      
      ## Explicit alternatives
      
      | Choice | Use it when | Availability check |
      |---|---|---|
      | `atlas-skill` | Default for an interactive agent generation | The Atlas Cloud Skill can submit media in this client |
      | `atlas-mcp` | The user explicitly selects MCP | Atlas MCP generation tools are exposed in the current client |
      | `atlas-cli` | The user explicitly selects terminal, script, CI, or batch execution | `atlas auth status`, then `atlas models get` |
      | `atlas-rest` | The bundled Node runner or a controlled compatibility run | `ATLASCLOUD_API_KEY` and live model verification |
      | `manual` | No authenticated executor is available | Return prompts, model IDs, and asset list only |
      
      The MCP server does not own a background polling loop. On the `atlas-mcp`
      route, the agent must call `atlas_get_prediction` with the same prediction ID
      every 2 seconds until a terminal state is returned.
      
      Do not say that a Node runner invoked MCP or an agent Skill. They are
      agent-level routes. Do not let a terminal runner silently choose the CLI merely
      because it is installed.
      
      ## API key discovery and setup
      
      Check credentials in the process that will actually submit the task. For the
      REST runner, use this order:
      
      1. `process.env.ATLASCLOUD_API_KEY`
      2. `process.env.ATLAS_CLOUD_API_KEY` as a compatibility alias
      
      Do not infer Atlas credentials from a different provider, plugin, or process.
      Each execution channel can have an independent credential scope.
      
      If no key is visible, direct the user to
      `https://www.atlascloud.ai/console/api-keys?utm_source=github&utm_campaign=awesome-seedance-2.5-prompts-skills`.
      Do not request or print the key in
      chat. Offer either a current-shell setup:
      
      ```bash
      export ATLASCLOUD_API_KEY="<your-key>"
      ```
      
      or the host application's secure environment or secret settings. Refresh or
      restart the execution session after changing persistent configuration when
      required. If the key exists in a parent or host configuration but the submitting
      process cannot see it, report an environment-scope mismatch and use a process
      that receives the configured environment. Do not claim that the user has no key.
      
      ## Bundled runner contract
      
      `scripts/generate.mjs` implements only `atlas-rest` and explicit `atlas-cli`.
      It defaults to `atlas-rest` when `execution.adapter` is absent. The legacy value
      `"auto"` remains accepted but resolves to `atlas-rest`; it no longer probes or
      selects the CLI. This prevents a local installation from changing a user's
      execution path or billing destination.
      
      ```json
      {
        "execution": {
          "adapter": "atlas-rest",
          "apiKeyEnv": "ATLASCLOUD_API_KEY",
          "verifyModels": true
        },
        "modelProfile": "seedance-default"
      }
      ```
      
      Set `adapter` to `atlas-cli` only when the CLI route was explicitly selected.
      `atlas-skill` and `atlas-mcp` are intentionally rejected by the Node runner so
      the agent can execute them directly rather than misreporting the channel.
      
      For an R2V job, keep the submitted storyboard image within the live provider's
      accepted reference-image constraints. If a board exceeds a provider constraint,
      re-layout the same ordered panels; do not change the intended output frame
      ratio as a workaround.
      
      If a submitted runner task loses its local polling process, set a segment's
      prediction ID under `execution.resumePredictionIds` using its stage key, for
      example:
      
      ```json
      {
        "execution": {
          "resumePredictionIds": {
            "grid": "existing-image-prediction-id",
            "seg1": "existing-video-prediction-id"
          }
        }
      }
      ```
      
      The runner resumes polling and download rather than submitting a second
      billable task. Active states are `starting`, `queued`, `pending`, and
      `processing`; successful states are `completed` and `succeeded`; terminal
      failure states are `failed`, `timeout`, and `canceled`. Missing inference time,
      an expired local polling window, an interrupted turn, or a temporary query
      failure never authorizes a replacement task. Only an explicit retry decision
      after a terminal failure may create a new prediction.
      
      All four Atlas execution routes use a 2-second status-query interval in this
      workflow. The agent owns the loop for `atlas-skill` and `atlas-mcp`; the MCP
      operation is `atlas_get_prediction`. The bundled REST and CLI adapters enforce
      the interval in code. Their progress logs are still emitted every 32 seconds,
      and the default local polling window remains 900 seconds. This workflow-level
      rule takes precedence over a generic lower-level example with a different
      cadence.
      
      ## CLI route
      
      When the user explicitly selects the CLI and it is not installed, use the
      official installer and then authenticate:
      
      ```bash
      curl -fsSL https://raw.githubusercontent.com/AtlasCloudAI/cli/main/install.sh | sh
      atlas auth login
      ```
      
      Run `atlas auth login` interactively. The adapter will not log in for you by
      passing an API key on the command line, because command-line arguments are
      readable by other processes on the host. If you must automate it on a
      single-user machine, set `ATLAS_CLI_ALLOW_ARGV_TOKEN=1` and accept that
      exposure; otherwise prefer `execution.adapter: "atlas-rest"`, which sends the
      key only in an HTTPS request header.
      
      Verify the selected model before generation:
      
      ```bash
      atlas models search seedance --type video --json
      atlas models get <model-id> --json
      atlas generate cost video <model-id> -p "<prompt>" --json
      ```
      
      The CLI adapter uses `atlas generate image|video <model> -p <prompt>` with
      non-blocking submission and polling. Local media is passed as `@/absolute/path`;
      data URLs and remote URLs are passed directly. It maps images, references,
      start/end frames, duration, resolution, ratio, audio, and additional model
      parameters.
      
    • execution-adapters.zh-CN.md 5.7 KB
      # Atlas 执行适配层
      
      创作规划与模型配置相互独立。默认使用 Seedream 5.0 Pro 生图、Seedance 2.0 生视频。任何计费生成前,都先核实所选模型支持需要的路线。
      
      ## 智能体默认路线
      
      在智能体对话中,**Atlas Cloud Skill** 是默认的直接生成路线。用它查询模型、检查参数、上传本地素材、提交图片或视频任务、轮询并获取结果。只有 Skill 实际提交了任务,才报告 `Execution: atlas-skill`。
      
      提交后,必须每 2 秒用同一 ID 重复执行 Skill 的预测结果查询步骤,直到任务进入终态。
      
      若它不可用,先协助安装:
      
      ```bash
      npx skills add AtlasCloudAI/atlas-cloud-skills
      ```
      
      如果宿主需要,再刷新 Skill 注册表。
      
      ## 明确选择的替代路线
      
      | 选择 | 使用时机 | 可用性检查 |
      |---|---|---|
      | `atlas-skill` | 智能体交互生成的默认选择 | Atlas Cloud Skill 能在当前客户端提交媒体任务 |
      | `atlas-mcp` | 用户明确选择 MCP | 当前客户端暴露 Atlas MCP 生成工具 |
      | `atlas-cli` | 用户明确选择终端、脚本、CI 或批量执行 | `atlas auth status`,再执行 `atlas models get` |
      | `atlas-rest` | 使用内置 Node 脚本或受控兼容执行 | `ATLASCLOUD_API_KEY` 和在线模型核验 |
      | `manual` | 没有已认证执行通道 | 只输出提示词、模型 ID 和素材清单 |
      
      MCP 服务端本身不维护后台轮询循环。走 `atlas-mcp` 时,智能体必须每 2 秒用同一预测 ID 调用一次 `atlas_get_prediction`,直到任务进入终态。
      
      不要声称 Node 脚本调用了 MCP 或智能体 Skill;两者都属于智能体层路线。也不要因为 CLI 已安装就让脚本静默切换到 CLI。
      
      ## API Key 检查与配置
      
      必须在真正提交任务的进程中检查凭据。REST 脚本按以下顺序读取:
      
      1. `process.env.ATLASCLOUD_API_KEY`
      2. 兼容变量 `process.env.ATLAS_CLOUD_API_KEY`
      
      不要根据另一个服务商、插件或进程的配置状态推断 Atlas 凭据。每个执行通道可能拥有独立的凭据作用域。
      
      如果执行进程读不到 Key,引导用户前往 `https://www.atlascloud.ai/console/api-keys?utm_source=github&utm_campaign=awesome-seedance-2.5-prompts-skills` 获取。不要要求用户在对话中粘贴或展示 Key。可以指导用户临时设置当前终端:
      
      ```bash
      export ATLASCLOUD_API_KEY="<your-key>"
      ```
      
      也可以写入宿主应用提供的安全环境变量或密钥设置。修改持久配置后,必要时刷新或重启执行会话。如果 Key 已存在于宿主或父进程配置,但实际提交任务的进程不可见,应报告“环境作用域不一致”,并改用能继承该环境的执行进程,不能说用户没有 Key。
      
      ## 内置脚本约定
      
      `scripts/generate.mjs` 只实现 `atlas-rest` 和显式 `atlas-cli`。未设置 `execution.adapter` 时默认 `atlas-rest`。旧值 `"auto"` 仍可使用,但固定解析为 `atlas-rest`,不再探测或选择 CLI。这样本机安装状态不会悄悄改变用户的执行路径或计费出口。
      
      ```json
      {
        "execution": {
          "adapter": "atlas-rest",
          "apiKeyEnv": "ATLASCLOUD_API_KEY",
          "verifyModels": true
        },
        "modelProfile": "seedance-default"
      }
      ```
      
      只有明确选择 CLI 时才设置 `adapter` 为 `atlas-cli`。`atlas-skill` 和 `atlas-mcp` 会被 Node 脚本明确拒绝,让智能体直接执行,避免错误报告执行通道。
      
      R2V 任务的 Storyboard 参考图必须符合在线服务当前接受的图片限制。若超出限制,只重新排布同一套有序分镜,不要为此改变用户希望的视频画幅。
      
      如果脚本提交后本地轮询进程中断,把已有 Atlas 预测 ID 按阶段键写入 `execution.resumePredictionIds`,例如:
      
      ```json
      {
        "execution": {
          "resumePredictionIds": {
            "grid": "existing-image-prediction-id",
            "seg1": "existing-video-prediction-id"
          }
        }
      }
      ```
      
      脚本会恢复轮询和下载,不会再次提交计费任务。`starting`、`queued`、`pending`、`processing` 是进行中状态;`completed`、`succeeded` 是成功终态;`failed`、`timeout`、`canceled` 是失败终态。推理时间缺失、本地轮询窗口结束、对话中断或临时查询失败都不能授权创建替代任务。只有任务明确进入失败终态并作出明确重试决定后,才可以创建新预测任务。
      
      本工作流的四种 Atlas 执行路线都使用 2 秒状态查询周期。`atlas-skill` 和 `atlas-mcp` 的循环由智能体负责,其中 MCP 调用 `atlas_get_prediction`;内置 REST 和 CLI 适配器则在代码中强制执行。REST/CLI 的进度日志仍每 32 秒输出一次,默认本地轮询窗口仍为 900 秒。若底层通用示例使用其他周期,以本工作流的 2 秒规则为准。
      
      ## CLI 路线
      
      当用户明确选择 CLI 且本机未安装时,使用官方安装方式,然后登录:
      
      ```bash
      curl -fsSL https://raw.githubusercontent.com/AtlasCloudAI/cli/main/install.sh | sh
      atlas auth login
      ```
      
      请交互式执行 `atlas auth login`。适配器不会替你用命令行参数传 API key 去登录,
      因为命令行参数可被本机其他进程读取。若必须在单用户机器上自动化,可设置
      `ATLAS_CLI_ALLOW_ARGV_TOKEN=1` 并接受该暴露;否则请优先使用
      `execution.adapter: "atlas-rest"`,它只在 HTTPS 请求头中发送 key。
      
      生成前核实模型:
      
      ```bash
      atlas models search seedance --type video --json
      atlas models get <model-id> --json
      atlas generate cost video <model-id> -p "<prompt>" --json
      ```
      
      CLI 适配器通过 `atlas generate image|video <model> -p <prompt>` 非阻塞提交并轮询。当地文件使用 `@/absolute/path`,数据 URL 和远程 URL 直接传递。它会映射图片、参考图、首尾帧、时长、分辨率、画幅、音频和额外模型参数。
      
    • long-video.md 5.1 KB
      # Long video: stages and end states
      
      A longer duration is not a longer prompt. It is a **sequence of stages, each with
      one primary change and a stated end state.** Get that structure right and the
      piece holds; get it wrong and extra length only produces more drift.
      
      ## Structure
      
      ```text
      [Goal]
        One sentence: what kind of piece, the central subject, the primary event.
      
      [Global]
        Environment and texture · visual style · camera language · subject styling ·
        performance core · prohibitions
      
      [Stage 1]
        Initial state: characters, props, scene as the piece opens
        Primary event: ONE action or change
        End state: what is visibly true when this stage ends
      
      [Stage 2]
        Continue from: what must remain unchanged from the previous stage
        Primary event: ONE action or change
        End state: ...
      
      [Stage n]
        Primary event: the closing event
        End state: the final visible state
      
      [Consistency]
        What holds across every stage: identity, subject count, wardrobe, prop
        ownership, spatial direction, audio relationships
      ```
      
      Restate the two or three most expensive consistency items at the physical end.
      
      ## The rules that matter
      
      **One primary change per stage.** Two changes in one stage means the model
      chooses which to honour. Split them.
      
      **Every stage ends on something visible.** Not a mood, not "continues" — a
      checkable state:
      
      ```text
      End state: the florist holds the bouquet in the left hand; the scissors are
      back on the right side of the bench.
      ```
      
      Positions, who holds what, what is open or closed, what has left frame. If you
      cannot check it on a still frame, it is not an end state.
      
      **Each stage names what carries over.** `Continue from the previous stage: both
      characters keep the same identities and clothing; the florist still holds the
      bouquet.` This is where identity drift gets caught before it starts.
      
      **The closing stage needs an end state too.** The most common omission, and the
      reason pieces trail off instead of landing.
      
      ## Timestamps: only when something external demands them
      
      Stages are the default. Reach for one-second precision only for a critical
      handoff, an entrance or exit, a transition, or a beat that must land at a fixed
      time.
      
      | Pattern | Form | Use for |
      |---|---|---|
      | Time range | `0–5s ... 5–10s ... 10–15s` | Allocating pacing across stages |
      | Exact point | `at 5s the camera whip-pans left and completes the transition` | One critical beat |
      | Relative | `three seconds after the button is pressed, the lights fade` | A delay between events |
      
      Rules:
      
      - Ranges must be **consecutive and non-overlapping**.
      - A range is a **time budget, not an edit point.** Actions may land slightly
        before or after a boundary.
      - **Too little content in a range** gives the model more freedom; **too much**
        causes over-cutting or dropped events.
      - Never demand a frequency such as "complete three actions in one second".
      - Combine ranges with end states. A range says how long; the end state says what
        must be true when it is over. The end state is doing the real work.
      
      If a model's timing adherence is weak, drop to stages — a timestamp the model
      does not honour is worse than no timestamp, because it looks like control.
      See [model profile](model-profile.md).
      
      ## Worked shape
      
      ```text
      [Goal]
      An instructional video showing a flower shop's order-packing process. A florist
      and a store assistant arrange, wrap, and hand off a bouquet together.
      
      [Stage 1]
      Initial state: the florist stands behind the workbench; loose stems, scissors,
      and wrapping paper lie on the tabletop.
      Primary event: the florist arranges the stems and trims them to length.
      End state: the florist holds the bouquet in the left hand; the scissors are back
      on the right side of the workbench.
      
      [Stage 2]
      Continue from: both characters keep the same identities and clothing; the florist
      still holds the bouquet.
      Primary event: the assistant unfolds the wrapping paper; the florist places the
      bouquet inside and ties it with a green ribbon.
      End state: the wrapped bouquet lies flat in the centre of the workbench, ribbon
      bow facing the camera.
      
      [Stage 3]
      Primary event: the assistant lifts the bouquet onto the pickup shelf.
      End state: the bouquet sits centred on the pickup shelf; both characters stand
      behind the workbench inspecting the finished order.
      
      [Consistency]
      Keep both identities, both outfits, the workbench orientation, the scissors
      position, and bouquet ownership consistent throughout.
      ```
      
      Note the absence of adjective stacking and per-stage camera vocabulary. The
      Global block carries the look; the end states carry the structure.
      
      ## Beyond a single request
      
      When the piece exceeds what one request supports, chain instead of cramming:
      
      - Split at a point where a stage genuinely ends, not mid-action.
      - Feed the real generated tail forward — see
        [editing and extension](editing-and-extension.md).
      - Each segment still needs its own end states.
      
      A chain is an alignment aid, not a guarantee of an invisible seam. Inspect both
      sides of every boundary.
      
      ## Related
      
      - [prompt blocks](prompt-blocks.md) · [multi reference](multi-reference.md) ·
        [capabilities](capabilities.md) · [editing and extension](editing-and-extension.md)
      
    • long-video.zh-CN.md 4.5 KB
      # 长视频:分段与末态
      
      更长的时长不等于更长的提示词,而是**一串段落,每段一个主要变化、一个写明的末态**。结构对了片子就立得住;结构错了,多出来的长度只会产出更多漂移。
      
      ## 结构
      
      ```text
      [生成目标]
        一句话:什么类型的片子、中心主体是谁、主要事件是什么。
      
      [全局设定]
        基础环境与质感 · 视觉风格 · 镜头语言 · 主体造型 · 表演核心 · 禁止项
      
      [第 1 段]
        初始状态:开场时的人物、道具、场景
        主要事件:一个动作或变化
        末态:这一段结束时什么是可见的
      
      [第 2 段]
        承接上段:必须保持不变的是什么
        主要事件:一个动作或变化
        末态:……
      
      [第 n 段]
        主要事件:收尾事件
        末态:最终的可见状态
      
      [保持一致]
        跨全部段落必须一致的:身份、主体数量、服装、道具归属、空间方向、音频关系
      ```
      
      把两三条最贵的一致性项在物理结尾复述一次。
      
      ## 真正管用的几条规则
      
      **每段只有一个主要变化。** 一段里塞两个变化,等于让模型自己挑一个执行。拆开。
      
      **每段都收在看得见的东西上。** 不是情绪、不是「继续」,而是可检查的状态:
      
      ```text
      末态:花艺师左手持花束;剪刀回到工作台右侧。
      ```
      
      位置、谁拿着什么、什么是打开或关闭的、什么已经离开画面。**在一张静帧上检查不了的,就不是末态。**
      
      **每段点明什么承接下来。** `承接上段:两人身份与服装不变,花艺师仍持花束。` 身份漂移就是在这里被提前拦住的。
      
      **收尾那一段也需要末态。** 最常被漏掉的一项,也是片子会「散掉」而不是「收住」的原因。
      
      ## 时间戳:只在外部有硬约束时才用
      
      分段是默认。只在关键交接、进出场、转场,或必须落在固定时间的拍点上,才动用一秒精度。
      
      | 写法 | 形式 | 用于 |
      |---|---|---|
      | 时间范围 | `0–5s … 5–10s … 10–15s` | 在段落之间分配节奏 |
      | 精确时点 | `第 5 秒摄影机快速左摇完成转场` | 一个关键拍点 |
      | 相对时间 | `按下按钮三秒后灯光渐暗` | 事件之间的延迟 |
      
      规则:
      
      - 范围必须**连续且不重叠**
      - 范围是**时间预算,不是剪辑点**。动作可能落在边界前后
      - **范围里内容太少**给模型更多自由;**太多**会导致过度切分或漏事件
      - 绝不要求「一秒内完成三个动作」这类频率
      - 范围要与末态配合。范围说多长,末态说结束时什么必须为真——**真正干活的是末态**
      
      如果某个模型时序遵循度弱,直接降到分段——一个模型不遵守的时间戳比没有时间戳更糟,因为它看起来像控制。见 [model-profile.zh-CN.md](model-profile.zh-CN.md)。
      
      ## 示例骨架
      
      ```text
      [生成目标]
      一条展示花店打包订单流程的教学视频。花艺师与店员一起整理、包装并交付一束花。
      
      [第 1 段]
      初始状态:花艺师站在工作台后;台面上散放着花枝、剪刀和包装纸。
      主要事件:花艺师整理花枝并修剪到合适长度。
      末态:花艺师左手持花束;剪刀回到工作台右侧。
      
      [第 2 段]
      承接上段:两人身份与服装不变,花艺师仍持花束。
      主要事件:店员展开包装纸;花艺师把花束放进去,用绿色缎带系好。
      末态:包好的花束平放在工作台中央,蝴蝶结朝向镜头。
      
      [第 3 段]
      主要事件:店员拿起花束放到取货架上。
      末态:花束居中放在取货架上;两人站在工作台后查看这份完成的订单。
      
      [保持一致]
      两人身份、两套服装、工作台朝向、剪刀位置、花束归属全程保持一致。
      ```
      
      注意里面没有形容词堆、也没有逐段的运镜词汇。全局设定承载观感,末态承载结构。
      
      ## 超出单次请求时
      
      片子长度超过一次请求能支持的范围时,接龙而不是硬塞:
      
      - 在某一段**真正结束**的地方切,不在动作中途切
      - 把**真实生成的尾帧**往前喂——不是你设想的尾帧,见 [editing-and-extension.zh-CN.md](editing-and-extension.zh-CN.md)
      - 每一段仍然需要自己的末态
      
      接龙是对齐辅助,**不保证接缝隐形**。边界两侧都要检查。
      
      ## 关联
      
      - [prompt-blocks.zh-CN.md](prompt-blocks.zh-CN.md) · [multi-reference.zh-CN.md](multi-reference.zh-CN.md) ·
        [capabilities.zh-CN.md](capabilities.zh-CN.md) · [editing-and-extension.zh-CN.md](editing-and-extension.zh-CN.md)
      
    • model-profile.md 7.2 KB
      # Model profiles: Seedance family
      
      Measured behaviour, not documented behaviour. The compiler reads these fields to
      decide what to emit and what to degrade.
      
      Field definitions and the blank template live in
      [the Universal Video Prompt Skill](../../universal-video-prompt-skill/references/model-profile-schema.md).
      If that link does not resolve, the Seedance 2.5 Skill installation is incomplete.
      Help the user install `universal-video-prompt-skill` before continuing and **do not
      invent a schema**.
      
      **Unknown is a valid value.** An empty field prompts a probe; a guessed field
      silently corrupts every run built on it.
      
      ---
      
      ## `bytedance/seedance-2.0`
      
      Last verified: 2026-08-03
      
      ### Capability layer
      
      | Field | Value | How verified |
      |---|---|---|
      | Reference addressing | `@image1`, `@image2`, … | Generation |
      | Max references | 9 images | Provider docs |
      | Multi-shot in one generation | yes — ordered segments with cuts | 15s multi-segment run |
      | Hard cut support | yes | Same |
      | Duration range | 4–15s | Provider docs |
      | Resolution | 480p, 720p, 1080p depending on provider | Provider model page |
      | Native audio | yes | Generation with audio enabled |
      | Timing adherence | Requested beats land **~2s late** across a 15s piece; segment **order holds** | Controlled 15s run against a timestamped prompt |
      | Recommended granularity | **stages** | Derived from the row above |
      | Extension / chaining | via tail-frame chaining | — |
      | Audio-only reference | not supported | Provider docs |
      
      ### Bias layer
      
      | Field | Value | How verified |
      |---|---|---|
      | Default aesthetic bias | Strong cinematic priors; interprets sparse prompts well and improvises sensibly | A/B against an over-specified variant of the same segment |
      | Effective anti-default phrasing | **Name the drawing tool, not the abstract property.** `crayon / coloured pencil / coarse brush, visible stroke direction, uneven fill, ragged edges` produced genuinely hand-drawn marks | Controlled A/B on one hand-drawn VFX spec |
      | Ineffective or overshooting | **`graphically flat, never photoreal` — satisfied exactly and uselessly.** The model honoured it with smooth neon-tube vector outlines and even fill: flat, but not hand-drawn at all. Swapping this one lock for tool names reversed the result on the same spec | Same A/B as above |
      | Ineffective or overshooting | **Storyboard-grid over-specification scores worse than text-only staging** on heavy-VFX segments — it suppresses the camera priors that make those shots work | A/B on the same segment |
      | Negative-lock behaviour | Respected. Front-load them | Iteration |
      | Reference-versus-text priority | **The image wins.** Composition references override written composition | Iteration |
      | Handheld / POV | Executes handheld POV convincingly, including motion blur on fast follows and a hand entering frame from below | Hand-drawn VFX spec, v2 |
      | Live-action texture retention | Strong. Keeps real bone, glass, stone, and ceiling fixtures at their own colour while a drawn layer sits on top | Same |
      | **Prompt-length tolerance** | **High.** A longer revision kept the hand, the opening transformation, the drawn texture, and gained the new beats. A comparison model on the same two prompts lost all three | v3 of the same spec |
      | Contact shadows unprompted | Adds a cast shadow under a drawn object sitting on a real surface, without being asked. This is the strongest available evidence of the "both media share one physical space" contract | v3 |
      | Spatial containment | Reads `open plinth` correctly — renders plinth surface, support rod, and shadow, not a vitrine | v3 |
      
      ### Known failure modes
      
      | Symptom | Detail | Handling |
      |---|---|---|
      | Small text errors | UI labels and signage render with character-level errors and garbled small type | Post-production for anything that must read exactly |
      | Dropped list items | Enumerated menu entries partially render — some items simply absent | Reduce the count, or add the text in post |
      | Camera move ignored | A stated move is skipped while the rest of the segment lands correctly | Re-state the move as the segment's primary intent, or accept and reframe |
      | Timeline drift | Whole timeline shifts later, ~2s over 15s | Use stages; if second-level is required, write beats early and verify |
      | **Dropped stages under load** | A 15s piece with **four** staged events rendered only stages 1 and 4 — the two middle stages were skipped entirely, not merely delayed. Raising event density per stage while keeping five shorter stages did **not** reproduce the drop | Treat ~4 distinct state-changes in 15s as the ceiling. If the spec needs more, either shorten each stage and raise density, or split across requests |
      
      ### Compile notes
      
      - Emit `@imageN`. Never bracketed or spelled-out reference labels.
      - Default to `stages`.
      - Prefer text-driven staging over feeding a storyboard grid for heavy-VFX and
        large-scale segments.
      - When a composition reference and written composition conflict, drop the written
        one — it will lose anyway and only adds noise.
      
      ---
      
      ## `bytedance/seedance-2.5`
      
      Last verified: — · Status: **not yet measured**
      
      Availability differs by provider. Verify against the provider's model page before
      offering 2.5-specific routes; see [capabilities](capabilities.md).
      
      ### Capability layer
      
      | Field | Value | How verified |
      |---|---|---|
      | Reference addressing | `@Image 1` / `@image1` — **unverified**, confirm against the provider | — |
      | Max references | up to 30 images, 10 video, 10 audio, ~50 combined | Launch material |
      | Multi-shot in one generation | expected yes | Not measured |
      | Duration range | 4–30s | Launch material |
      | Resolution | read the provider's model page | — |
      | Native audio | yes | Launch material |
      | Audio-only reference | supported | Launch material |
      | Timing adherence | **not measured** — the single most valuable probe | — |
      | Recommended granularity | **unknown** — do not assume it differs from 2.0 | — |
      | Editing / extension / bridging | announced; API exposure varies | — |
      
      ### Bias layer
      
      Not measured. Do **not** copy 2.0's bias-layer entries here — anti-default
      phrasing is per-model by definition, and launch material specifically claims
      changed default behaviour around unrequested subtitles and music, which is exactly
      the kind of thing a 2.0-tuned negative list would now be redundantly fighting.
      
      ### First probes to run, in order
      
      1. **Reference addressing** — one generation with two bound references. Everything
         else is blocked on getting this right.
      2. **Timing adherence** — the same spec at `stages` and at `second-level`; measure
         the drift. Fills timing adherence, recommended granularity, and usually a
         failure mode in one run.
      3. **Unrequested subtitles and music** — generate with no negatives at all. If
         they no longer appear, 2.0's negative lines are dead weight here.
      4. **Reference ceiling in practice** — where stability actually degrades, not where
         the documented limit sits.
      
      Record results here as they land, including negative results.
      
      ---
      
      ## Related
      
      - [model profile schema](../../universal-video-prompt-skill/references/model-profile-schema.md) — field definitions
      - [capabilities](capabilities.md) — documented limits and platform-versus-API
      - [troubleshooting](troubleshooting.md) — symptom-driven fixes
      
    • model-profile.zh-CN.md 6.6 KB
      # 模型档案:Seedance 家族
      
      记实测行为,不记文档行为。编译器读这些字段来决定输出什么、降级什么。
      
      字段定义与空白模板在[通用视频提示词 Skill](../../universal-video-prompt-skill/references/model-profile-schema.zh-CN.md)。
      若此链接不可达,说明 Seedance 2.5 Skill 安装不完整。继续前先协助用户安装 `universal-video-prompt-skill`,并且**不要自行编造字段定义**。
      
      **「未知」是合法值。** 空字段会触发一次探测;猜出来的字段会静默污染每一次基于它的运行。
      
      ---
      
      ## `bytedance/seedance-2.0`
      
      最后核实:2026-08-03 · 服务:Atlas Cloud
      
      ### 能力位层
      
      | 字段 | 值 | 如何核实 |
      |---|---|---|
      | 素材寻址 | `@image1`、`@image2`… | 生成实测 |
      | 素材上限 | 9 张图 | 厂商文档 |
      | 一次生成能否多镜头 | 能——带切点的有序段落 | 15 秒多段生成 |
      | 是否支持硬切 | 能 | 同上 |
      | 时长范围 | 4–15s | 厂商文档 |
      | 分辨率 | 480p、720p、1080p,视服务而定 | 服务模型页 |
      | 原生音频 | 有 | 开启音频的生成 |
      | 时序遵循度 | 15 秒的片子里请求拍点**整体晚约 2 秒**;段落**顺序稳** | 对着带时间戳的提示词做的受控 15 秒实测 |
      | 建议粒度 | **分段** | 由上一行推出 |
      | 延长/接龙 | 通过尾帧接龙 | — |
      | 仅传音频 | 不支持 | 厂商文档 |
      | **prompt 长度耐受性** | **高。** 加长版本保住了手、开场变形、手绘质感,还吃下了新增拍点;对照模型在同样两份提示词上把这三样全丢了 | 同一份 spec 的 v3 |
      
      ### 偏置层
      
      | 字段 | 值 | 如何核实 |
      |---|---|---|
      | 默认审美偏置 | 镜头先验很强;稀疏提示词理解得好,会合理即兴 | 与过度指定版本做 A/B |
      | 有效的反默认写法 | **按工具命名,不要命名抽象属性。** `蜡笔/彩铅/粗笔刷,看得见排线方向、涂色不均、毛边` 给出了真正的手绘笔触 | 同一份手绘 VFX spec 的受控 A/B |
      | 无效或过冲的写法 | **`保持平面图形、绝不写实` —— 被精确满足、且毫无用处。** 模型用光滑的霓虹管矢量描边加均匀填充满足了它:平面,但完全不手绘。在同一份 spec 上只把这一条锁换成工具名,结果直接反转 | 同上 |
      | 无效或过冲的写法 | **喂故事板网格过度指定,比纯文字 staging 更差** —— 会压掉让重特效镜头成立的运镜先验 | 同一段落的 A/B |
      | 负向锁行为 | 会被尊重。前置 | 迭代 |
      | 参考图与文字的优先级 | **图赢。** 构图参考会覆盖写出来的构图 | 迭代 |
      | 手持 / POV | 能可信地执行手持 POV,包括快速跟随时的运动模糊、以及手从画面下方入画 | 手绘 VFX spec,v2 |
      | 实拍质感保持 | 强。手绘层叠在上面时,真实的骨、玻璃、石材、天花灯具都保持自己的颜色 | 同上 |
      | 未经要求的接触阴影 | 手绘物落在真实表面上时会自己加投影,没有被要求过。这是「两种媒介共享一个物理空间」能拿到的最强证据 | v3 |
      | 空间容纳关系 | 能正确读懂 `开放式展台` ——会渲染台面、支撑杆和投影,不是展柜 | v3 |
      
      ### 已知失败模式
      
      | 症状 | 细节 | 处理 |
      |---|---|---|
      | 小字出错 | UI 标签与招牌出现字符级错误,小字乱码 | 需要完全准确的文字走后期 |
      | 列表项漏画 | 枚举的菜单项只画出一部分——有些条目直接缺失 | 减少条目数,或文字走后期 |
      | 运镜被忽略 | 写明的运镜被跳过,段落其余部分正常落地 | 把运镜改写成该段的主要意图,或接受并重新构图 |
      | 时间轴漂移 | 整条时间轴后移,15 秒里约 2 秒 | 用分段;若必须秒级,把拍点写早并核实 |
      | **超载时丢段** | 15 秒装**四段**时只渲染了第 1 段和第 4 段——中间两段**整个跳过**,不是延迟。改成五段更短、每段事件更密之后**没有**复现 | 把 15 秒里约 4 个不同状态变化当作上限。需要更多就缩短每段、提高密度,或拆成多次请求 |
      
      ### 编译注记
      
      - 输出 `@imageN`。绝不用方括号或拼写形式的素材标签
      - 默认 `分段`
      - 重特效与大场面段落优先文字驱动 staging,不要喂故事板网格
      - 构图参考与写出来的构图冲突时,删掉写的那份——它反正会输,只是增加噪音
      
      ---
      
      ## `bytedance/seedance-2.5`
      
      最后核实:— · 状态:**尚未实测**
      
      可用性因服务而异。提供 2.5 专属路线之前,先核实该服务的模型页;见 [capabilities.zh-CN.md](capabilities.zh-CN.md)。
      
      ### 能力位层
      
      | 字段 | 值 | 如何核实 |
      |---|---|---|
      | 素材寻址 | `@Image 1` / `@image1` —— **未核实**,向服务确认 | — |
      | 素材上限 | 最多 30 图、10 视频、10 音频,合计约 50 | 发布材料 |
      | 一次生成能否多镜头 | 预期能 | 未实测 |
      | 时长范围 | 4–30s | 发布材料 |
      | 分辨率 | 读服务的模型页 | — |
      | 原生音频 | 有 | 发布材料 |
      | 仅传音频 | 支持 | 发布材料 |
      | 时序遵循度 | **未实测** —— 最值得跑的那次探测 | — |
      | 建议粒度 | **未知** —— 不要假设它与 2.0 不同 | — |
      | 编辑/延长/桥接 | 已宣布;API 暴露情况因服务而异 | — |
      
      ### 偏置层
      
      未实测。**不要**把 2.0 的偏置层条目抄过来——反默认写法按定义是逐模型的,而且发布材料明确声称默认行为在「多余字幕与音乐」方面有变化,那恰恰是一份为 2.0 调过的负向清单现在可能在多余对抗的东西。
      
      ### 该先跑的探测,按顺序
      
      1. **素材寻址** —— 一次生成,绑两份素材。这条不对,后面全部被卡住
      2. **时序遵循度** —— 同一份 spec 分别用 `分段` 与 `秒级` 各跑一遍,量偏差。一次填满时序遵循度、建议粒度,通常还附带一个失败模式
      3. **多余字幕与音乐** —— 完全不写负向跑一次。如果它们不再出现,2.0 那套负向行在这里是多余负担
      4. **素材上限的实际拐点** —— 稳定性真正开始下降的位置,不是文档上限的位置
      
      结果落地就记回本文件,**包括负面结果**。
      
      ⚠️ **别信参数表。** 本平台已确认三次文档比实际更窄(H3 t2v、Kling 两档)。**直接提交你真正想要的值**——被拒不创建任务、不花钱。
      
      ---
      
      ## 关联
      
      - [模型档案字段定义](../../universal-video-prompt-skill/references/model-profile-schema.zh-CN.md)
      - [capabilities.zh-CN.md](capabilities.zh-CN.md) — 文档上限、以及平台功能与 API 的分界
      - [troubleshooting.zh-CN.md](troubleshooting.zh-CN.md) — 按症状查修法
      
    • multi-reference.md 5.6 KB
      # Multi-reference binding
      
      The goal is **not** to make every reference appear at once. It is to let the model
      select the right materials for the moment it is generating. More references
      without clearer relationships produces worse output, not richer output.
      
      Work in this order. Each step exists because skipping it produces a specific
      failure.
      
      ## Step 1 — Bind each subject individually
      
      One line per subject. State what the reference controls **and** what not to take
      from it.
      
      ```text
      Character A corresponds to @image1. Use only the appearance, hairstyle, and clothing.
      Character B corresponds to @image2. Use only the appearance, hairstyle, and clothing.
      Prop A corresponds to @image3. Use only the structure, material, and colour.
      Scene A references @image4. Use only the spatial layout and lighting.
        Do not use the people in the image.
      ```
      
      **Never write** `@image1 through @image4 define four characters respectively`. That
      states no mapping at all — it tells the model there are four characters and four
      images, not which is which.
      
      When several images show one object from different angles, say so and state the
      output count:
      
      ```text
      @image1 defines the front of the folding desk lamp.
      @image2 defines its left-side structure.
      @image3 defines its right-side structure.
      All three images define one lamp. Exactly one lamp appears throughout.
      ```
      
      Without the count line, multi-view references are a common cause of duplicated
      objects.
      
      ## Step 2 — Group by type
      
      Once there are more than a handful, group them. Grouping is what makes the
      relationships readable rather than a flat list.
      
      ```text
      [Characters]
      Conservator → @image1. Appearance, hairstyle, clothing only.
      Registrar   → @image2. Appearance, hairstyle, clothing only.
      Do not interchange their appearances, clothing, actions, positions, or dialogue.
      
      [Props]
      Sample Case  → @image5, belongs only to the Conservator.
      Record Board → @image6, belongs only to the Registrar.
      
      [Scenes]
      Conservation Lab → @image7. Space, materials, lighting only.
      Gallery          → @image8. Space, materials, lighting only.
      
      [Motion and audio]
      @video1 defines the motion of the Conservator opening the Sample Case.
        Do not use the person or scene from the video.
      @audio1 defines the Guide's voice and dialogue.
      ```
      
      **Prop ownership** is the line people omit and then need. `belongs only to X` is
      what prevents props migrating between characters mid-piece.
      
      ## Step 3 — Profile the subjects that recur
      
      When one subject uses several references across several scenes, collect them:
      
      ```text
      [Subject profile: Conservator]
      Appearance and clothing: @image1
      Fixed prop: Sample Case from @image5
      Locations: Conservation Lab, Gallery
      Motion references: case-opening from @video1, sample-placement from @video2
      Do not use: other characters' clothing. Do not give this character the
      Record Board or guide equipment.
      ```
      
      A profile is worth writing when a subject appears in more than two scenes.
      Below that it is overhead.
      
      ## Step 4 — Select references per scene
      
      This is the step that makes many references work. Each scene names only what it
      uses, plus its end state:
      
      ```text
      Scene 1 | Inspection in the Conservation Lab
      Use: Conservator, Sample Case, Conservation Lab, case-opening motion from @video1.
      Event: the Conservator opens the Sample Case at the workbench and inspects it.
      End state: the Conservator remains on the inner side of the workbench; the Sample
      Case stays beside their right hand, on the left of frame.
      
      Scene 2 | Registration in the Gallery
      Use: Registrar, Record Board, Gallery.
      Event: the Registrar checks the number on the Record Board beside the display case.
      End state: the Registrar still holds the Record Board with both hands; no other
      character enters the display-case area.
      ```
      
      Without per-scene selection, the model tries to satisfy every reference in every
      shot, which is the actual mechanism behind crowded, incoherent output.
      
      ## Choosing what to supply
      
      Recommended ranges improve stability; they are not capability limits. Exceeding
      them is allowed and gets less predictable.
      
      | Material | Prefer | Why |
      |---|---|---|
      | Subject images | 1–8 distinct subjects | Beyond this, identity separation degrades |
      | Subject video/audio | 1–5 subjects, 5–10s each | Longer references dilute what is inherited |
      | Video editing | source under ~20s, 1–5 reference images | Scope stays controllable |
      
      **Views:** with up to about five subjects, single-view and multi-view both work.
      Past that, prefer single-view. When multiple views are needed, **separate images
      per view beat one collage** — a composite is frequently read as several subjects.
      
      **Faces:** use a clean face close-up plus a separate full-body image. Do not use a
      front/side/back turnaround sheet as an identity input; it reads as multiple people.
      
      **Motion references:** when a reference video already defines motion, camera, and
      sequence accurately, state only which attributes to inherit. Re-describing the
      same motion in text competes with the reference. A coarse blockout is the
      exception — it supplies motion and space only, so the prompt must still define
      subjects, scene, action, and style.
      
      ## Check before submitting
      
      - [ ] Every subject bound individually, one line each
      - [ ] Every reference states both what it controls and what not to use
      - [ ] Multi-view sets state the output count
      - [ ] Prop ownership stated where props could migrate
      - [ ] References selected per scene, not required all at once
      - [ ] Reference count within what the target model accepts —
            see [capabilities](capabilities.md)
      
      ## Related
      
      - [long video](long-video.md) · [real person](real-person.md) ·
        [capabilities](capabilities.md) · [troubleshooting](troubleshooting.md)
      
    • multi-reference.zh-CN.md 5.1 KB
      # 多素材绑定
      
      目标**不是**让每份素材都出现,而是让模型为它正在生成的那个时刻选对素材。素材更多但关系不更清楚,产出的是更差的结果,不是更丰富的结果。
      
      按顺序做。每一步的存在都对应一个跳过它就会出现的具体失败。
      
      ## 第 1 步 —— 逐个绑定主体
      
      一个主体一行。写明这份素材**控什么**,以及**不许从它拿什么**。
      
      ```text
      角色 A 对应 @image1。只使用外形、发型和服装。
      角色 B 对应 @image2。只使用外形、发型和服装。
      道具 A 对应 @image3。只使用结构、材质和颜色。
      场景 A 参考 @image4。只使用空间布局与光照。不要用图中的人。
      ```
      
      **绝不要写** `@image1 到 @image4 分别定义四个角色`。那等于什么映射都没说——它只告诉模型「有四个角色和四张图」,没说哪个对哪个。
      
      几张图拍同一个物体的不同角度时,要说明,并写明输出数量:
      
      ```text
      @image1 定义折叠台灯的正面。
      @image2 定义它的左侧结构。
      @image3 定义它的右侧结构。
      三张图定义的是同一盏灯。全片只出现一盏灯。
      ```
      
      不写数量那一行,多视角素材是「物体被复制」最常见的成因。
      
      ## 第 2 步 —— 按类型分组
      
      素材超过几份之后就分组。分组才让关系可读,而不是一张扁平的清单。
      
      ```text
      [角色]
      修复师 → @image1。仅外形、发型、服装。
      登记员 → @image2。仅外形、发型、服装。
      四个角色的外形、服装、动作、位置、台词不得互换。
      
      [道具]
      样品箱   → @image5,只属于修复师。
      记录板   → @image6,只属于登记员。
      
      [场景]
      修复实验室 → @image7。仅空间、材质、光照。
      展厅       → @image8。仅空间、材质、光照。
      
      [动作与音频]
      @video1 定义修复师打开样品箱的动作。不要用视频里的人或场景。
      @audio1 定义导览员的音色与指定台词。
      ```
      
      **道具归属**是最常被省略、然后又最需要的那一行。`只属于 X` 正是防止道具在片中途换主人的东西。
      
      ## 第 3 步 —— 给反复出现的主体建档
      
      同一个主体在多个场景里用到多份素材时,把它们收在一起:
      
      ```text
      [主体档案:修复师]
      外形与服装:@image1
      固定道具:@image5 的样品箱
      所在地点:修复实验室、展厅
      动作参考:@video1 的开箱动作、@video2 的放样动作
      不要使用:其他角色的服装。不要把记录板或导览设备给这个角色。
      ```
      
      一个主体出现在两个以上场景时才值得建档,低于这个数量属于额外负担。
      
      ## 第 4 步 —— 按场景挑素材
      
      这一步才是让「素材多」真的能用的那一步。每个场景只点它要用的,再加上末态:
      
      ```text
      场景 1 | 修复实验室内的检查
      使用:修复师、样品箱、修复实验室、@video1 的开箱动作。
      事件:修复师在工作台上打开样品箱,检查里面的样品。
      末态:修复师仍在工作台内侧;样品箱留在其右手旁、位于画面左侧。
      
      场景 2 | 展厅内的登记
      使用:登记员、记录板、展厅。
      事件:登记员在展柜旁核对记录板上的编号。
      末态:登记员仍双手持记录板;没有其他角色进入展柜区域。
      ```
      
      不做按场景挑选,模型会试图在每个镜头里满足所有素材——这正是画面拥挤、逻辑混乱的实际成因。
      
      ## 该给多少素材
      
      推荐范围是为了稳定性,不是能力上限。超出是允许的,只是可预测性下降。
      
      | 素材 | 稳定 | 可行但不稳 |
      |---|---|---|
      | 主体图里的不同主体 | 1–8 个 | 9–12 个 |
      | 主体音/视频里的不同主体 | 1–5 个 | 6–10 个 |
      | 每个主体的参考片时长 | 5–10s | 更长 |
      | 编辑用的源视频 | 20 秒以内 | 更长 |
      | 视频编辑的参考图 | 1–5 张 | 6–8 张 |
      
      **视角**:约五个主体以内,单视角与多视角都行。超过之后优先单视角。需要多个视角时,**一个视角一张图胜过拼成一张**——拼图经常被读成多个主体。
      
      **人脸**:用一张干净的脸部特写加一张单独的全身图。**不要**把正/侧/背三视图当作身份输入——它会被读成多个人。
      
      **动作参考**:参考视频已经准确定义了运动、运镜和顺序时,**只写要继承哪些属性**。在文字里把同一段运动再描述一遍会跟参考素材打架。粗白模是例外——它只提供运动和空间,所以提示词仍必须定义主体、场景、动作和风格。
      
      ## 提交前检查
      
      - [ ] 每个主体单独绑定,一个一行
      - [ ] 每份素材都写了控什么、以及不许拿什么
      - [ ] 多视角组写明了输出数量
      - [ ] 道具可能换主人的地方写了归属
      - [ ] 素材按场景挑选,不是要求全部同时出现
      - [ ] 素材数量在目标模型接受范围内——见 [capabilities.zh-CN.md](capabilities.zh-CN.md)
      
      ## 关联
      
      - [long-video.zh-CN.md](long-video.zh-CN.md) · [real-person.zh-CN.md](real-person.zh-CN.md) ·
        [capabilities.zh-CN.md](capabilities.zh-CN.md) · [troubleshooting.zh-CN.md](troubleshooting.zh-CN.md)
      
    • prompt-blocks.md 4.4 KB
      # Prompt blocks and conventions
      
      Use only the blocks that affect the shot.
      
      | Block | Purpose | Example |
      |---|---|---|
      | Reference binding | State what an input controls | `Use the insulated cup in Image 1 as the subject` |
      | Action | Give one observable event | `A hand naturally lifts and slightly turns the cup` |
      | Space | Lock an important relation | `The cup sits at the centre of a wood table; window light enters from the left` |
      | Camera | State framing and camera intent | `Medium close-up; smooth rightward move while tracking the cup` |
      | Light/style | Unify the look when needed | `Warm practical light; hand-drawn food-animation texture` |
      | Audio | Add only when native audio is enabled | `(upbeat jazz)<oil sizzling>` |
      | End state | Land a stage on a checkable state | `End state: the cup sits centred on the shelf; both hands have left frame` |
      | Constraint | Preserve expensive-to-redo details | `Keep the label readable; no subtitles` |
      
      For whole-storyboard R2V, order events as `Shot 1 / Shot 2 / Shot 3`. For I2V,
      write one shot and its start-to-end change.
      
      ## End states
      
      Any stage that must land somewhere specific gets one. This is the block that
      carries multi-event work — it converts "keep it consistent" into something the
      model can target and you can check on a still frame.
      
      ```text
      weak:   they keep working on the bouquet
      strong: End state: the florist holds the bouquet in the left hand; the scissors
              are back on the right side of the bench
      ```
      
      Rules:
      
      - It must be **visible**. `she feels relieved` is not an end state;
        `her shoulders drop and the frown clears` is.
      - Positions, who holds what, what is open or closed, what has left frame.
      - The **final** stage needs one too — the most common omission, and the reason
        pieces trail off instead of landing.
      - One primary change per stage. Two changes means the model picks one.
      
      The same device has an image form (a keyframe sequence, where each image *is* one
      stage's end state) and a video form (the boundary frame of an extension). See
      [long video](long-video.md) and
      [editing and extension](editing-and-extension.md).
      
      ## Camera
      
      Choose one primary movement per shot:
      
      - `Static camera`: observation, product stillness, dialogue.
      - `Slow push-in`: reveal, emotion, key detail.
      - `Slow pull-out`: release, ending, space reveal.
      - `Smooth lateral move`: product display or a walking subject.
      - `Tracking shot`: travel or action with a clear direction.
      - `Pan or tilt`: reveal a second subject or location.
      - `Crane move`: opening scale or final lift.
      - `Smooth orbit`: product or medium/wide character shot; avoid large rotations in tight face close-ups.
      
      A composite move is valid when every part supports one synchronized intent. For
      example, `move smoothly to the right while tracking the person` is a side-tracking
      shot. State direction, subject relation, and speed. Split at a planned cut when
      movements are independent or compete for attention.
      
      ## Audio symbols
      
      | Type | Symbol | Example |
      |---|---|---|
      | Music | `()` | `(fast rock music plays in the background)` |
      | Sound effect | `<>` | `<a dog barks in the distance>` |
      | Dialogue | `{}` | `She says in Japanese {こんにちは}` |
      | Caption | `【】` | `【Chapter One: Departure】` |
      
      The runner enables native audio by default. Set `generate_audio:false` only
      when the user explicitly wants a fully replaced post-production soundtrack. Keep
      native dialogue in one language and use the symbols above.
      
      ## Constraints
      
      Use constraints for specific risks rather than a long negative-prompt list:
      
      - `Preserve the bottle proportion and readable label in Image 1.`
      - `No subtitles.`
      - `No unrelated text, platform UI, or corner badge.`
      - `Keep only one corresponding person in frame; do not duplicate similar-looking people.`
      
      Put the few most important constraints last. Do not require an exact
      second-by-second schedule in text; use provider duration controls.
      
      ## Cut vocabulary
      
      - `Hard cut to ...`: deliberate energy or angle change.
      - `Match cut on <shape, colour, or movement> to ...`: match cut.
      - `After steam, a person, or an object fully occludes the lens, cut to ...`: occlusion cut.
      - `Insert <product or environmental detail>, then cut back ...`: cutaway.
      
      Use `continue seamlessly from the final frame of the prior segment` only for a
      true extension of the same shot. For whole-storyboard R2V, state a transition
      only when it matters to the story; otherwise let ordered panels establish cuts.
      
    • prompt-blocks.zh-CN.md 4 KB
      # 提示词组件库与约定
      
      只使用会影响当前镜头的组件,不要默认填满。
      
      | 组件 | 作用 | 示例 |
      |---|---|---|
      | 参考绑定 | 指定输入图片控制什么 | `图片1中的保温杯作为主体` |
      | 动作 | 给镜头一个可观察事件 | `手自然拿起杯身并轻微旋转` |
      | 空间 | 锁定重要关系 | `杯子位于木桌中央,窗光从左侧照入` |
      | 镜头 | 指定景别与镜头意图 | `中近景,镜头向右平稳横移,保持跟随杯子` |
      | 光影与风格 | 必要时统一视觉 | `暖黄实用光,手绘食物番质感` |
      | 音频 | 仅在启用原生音频时加入 | `(轻快爵士乐)<油锅滋滋声>` |
      | 末态 | 让某一段落在可检查的状态上 | `末态:杯子居中放在架子上;双手已离开画面` |
      | 约束 | 保住重做代价高的内容 | `保持标签清晰,无字幕` |
      
      整张 Storyboard R2V 按 `镜头1 / 镜头2 / 镜头3` 写事件顺序。I2V 只写单个镜头从开始到结束的变化。
      
      ## 末态
      
      任何必须落在特定位置的段落都要写一个。这是承载多事件片子的那个组件——它把「保持一致」变成模型能瞄准、你也能在一张静帧上检查的东西。
      
      ```text
      弱:  两个人继续弄那束花
      强:  末态:花艺师左手持花束;剪刀回到工作台右侧
      ```
      
      规则:
      
      - 必须**看得见**。`她感到释然` 不是末态;`她的肩膀落下、皱起的眉松开` 是
      - 位置、谁拿着什么、什么是打开或关闭的、什么已经离开画面
      - **最后一段也需要**——最常被漏掉的一项,也是片子会散掉而不是收住的原因
      - 每段一个主要变化。两个变化意味着模型自己挑一个
      
      同一个装置还有图像形态(关键帧序列——每张图**就是**某一段的末态)和视频形态(延长的边界帧)。见 [long-video.zh-CN.md](long-video.zh-CN.md) 与 [editing-and-extension.zh-CN.md](editing-and-extension.zh-CN.md)。
      
      ## 运镜
      
      每镜头先选一个主导运镜:
      
      - `固定机位`:观察、产品静态、对话。
      - `缓慢推近`:揭示、情绪、关键展示。
      - `缓慢拉远`:释放、收尾、展示空间。
      - `平稳横移`:展示产品或陪伴人物行走。
      - `跟拍`:有明确方向的移动或动作。
      - `摇镜`:揭示第二主体或地点。
      - `升降`:开场尺度或收尾上扬。
      - `平滑环绕`:产品或中远景人物;避免近距离脸部大幅环绕。
      
      当多个运动共同服务于一个同步意图时,可以使用复合运镜。例如 `镜头向右平稳横移,保持跟随人物` 是侧向跟拍。写清方向、与主体关系和速度。若多个运动彼此独立或抢注意力,应在计划好的剪辑处分开。
      
      ## 音频符号
      
      | 类型 | 符号 | 示例 |
      |---|---|---|
      | 音乐 | `()` | `(背景中播放着快节奏的摇滚乐)` |
      | 音效 | `<>` | `<远处传来狗叫声>` |
      | 台词 | `{}` | `用日语说道{こんにちは}` |
      | 画面文字 | `【】` | `【第一章:启程】` |
      
      脚本默认开启原生音频。只有用户明确希望完全使用后期替换音轨时,才设置 `generate_audio:false`。需要原生台词时,保持同一种语言并使用以上符号。
      
      ## 约束
      
      约束用于具体风险,不要写成长篇负面词:
      
      - `保持图片1中产品的瓶身比例和标签清晰`。
      - `保持无字幕`。
      - `不要出现无关文字、平台 UI 或角标`。
      - `同框仅保留一个对应人物,避免出现相同外形的重复人物`。
      
      把少数最重要的约束放在末尾。不要在文本里要求精确逐秒时间表;使用服务的时长参数。
      
      ## 剪辑词汇
      
      - `硬切到……`:节奏、能量或机位切换。
      - `以<形状、颜色或动作>匹配切到……`:匹配剪辑。
      - `蒸汽、人物或物体掠过镜头遮满画面后切到……`:遮挡剪辑。
      - `插入<产品或环境细节>,再切回……`:插入镜头。
      
      只有真正延长同一个镜头时才写 `无缝承接上一段最后画面`。整张 Storyboard R2V 只有在转场对故事真正重要时才写明,否则让有序分镜自行表达剪辑。
      
    • prompt-templates.md 5.3 KB
      # Prompt templates
      
      Replace angle-bracket placeholders. Choose one route; do not combine every
      template into one request.
      
      ## Subject reference: image model, optional
      
      Use this only when a person, product, object, or scene must recur.
      
      ```text
      Create a clean reference image for <subject>.
      Keep only these invariants: <silhouette or proportion>, <signature material or
      wardrobe>, <key colour>, <must-preserve marking>.
      <For a person: neutral face close-up, simple background, no text.>
      <For a product: clear three-quarter view, clean background, no text.>
      ```
      
      For a recurring person, generate a face close-up and a separate full-body image.
      Use a multi-view turnaround only as an image-design aid; do not submit it as a
      person's video identity reference.
      
      ## Seedream storyboard reference: default Seedance R2V route
      
      Use this to generate the one storyboard image uploaded directly as `Image 1` to
      Seedance. Panel numbers, arrows, and concise notes are allowed when they help
      explain sequence.
      
      ```text
      Generate one <rows>x<cols> visual storyboard grid for this sequence, read left
      to right and top to bottom:
      <sequence summary>.
      
      Panel 1: <shot purpose, composition, action>.
      Panel 2: <shot purpose, composition, action>.
      ...
      
      Keep <subject / product / scene> consistent across panels: <invariants>.
      Each panel must visibly show its planned composition and action. For a continuing
      action, show the next meaningful visible moment; for a cut, show the new shot
      clearly. Add panel numbers, concise notes, arrows, and clear cut or transition
      intentions when they help explain order or motion.
      ```
      
      Inspect panel order, planned beats, essential identity or scene details, and
      task-critical visual elements before R2V. Regenerate the whole board when a
      critical panel is wrong. Submit the accepted board as one image; create a clean
      copy only if a test generation renders labels or dividers into the video.
      
      ## Clean storyboard: optional retry or I2V crop source
      
      ```text
      Generate one clean <rows>x<cols> visual storyboard grid for this sequence:
      <sequence summary>.
      
      Each panel contains only the planned shot composition and action. Keep
      <subject / product / scene> consistent: <invariants>.
      Use thin neutral panel dividers only. Do not include numbers, captions, speech
      bubbles, labels, subtitles, logos, unrelated overlay marks, UI, or contact-sheet
      text.
      ```
      
      Use this only after a previous generation leaked labels or dividers, or when a
      deliberate I2V crop route needs it. Inspect layout and aspect ratios before
      cropping.
      
      ## Whole-storyboard R2V: preferred while panels remain readable
      
      Upload the complete storyboard as `Image 1`.
      
      ```text
      Follow the storyboard in Image 1 in left-to-right, top-to-bottom order. Generate
      one complete shot at a time.
      
      Shot 1: <panel 1 action and camera>.
      Shot 2: <panel 2 action and its cut or transition from the prior shot>.
      Shot 3: <panel 3 action and camera>.
      
      Keep <subject / product / scene invariants> throughout; <global lighting or
      style rule>. <Audio rule, if enabled>.
      Use Image 1 panel dividers, numbers, arrows, and notes only to understand the
      shot order and intention; do not render them into the final video.
      ```
      
      Do not split the board first. Panel count is not decisive; readable visual detail
      is. Switch to I2V shot pairs only when independent reshoots or precise start/end
      states matter more than one-pass transition quality.
      
      ## R2V asset references: one clip from a small asset pack
      
      Upload only assets with different jobs.
      
      ```text
      Use the <person / product> in Image 1 as the subject. Use Image 2 as the
      <scene or wardrobe> reference and Image 3 as the <visual style or composition>
      reference. <State the role of Video 1 or Audio 1 when supplied.>
      
      <One main action>. <Camera and light>.
      Keep <the two or three essential invariants>.
      ```
      
      Do not treat a large set of sequential storyboard cells as an asset pack. Use a
      whole storyboard or separate I2V shots for a sequence.
      
      ## I2V shot pair: start and end keyframes
      
      Upload the start keyframe as `Image 1` and the target end keyframe as `Image 2`.
      
      ```text
      Start from Image 1. The subject is <subject>.
      <One continuous, observable action>.
      Camera: <one move>. <Space or important object relationship>.
      Naturally resolve into the composition and subject state in Image 2.
      <Light or style>. Preserve <identity, product label, or scene invariants>.
      ```
      
      For an intentional hard cut, do not use `Image 2` as the following shot's start
      image. The next shot uses its own pair.
      
      ## Continuous extension or chain
      
      Use the prior generated tail only for the same uninterrupted shot.
      
      ```text
      Continue seamlessly from the final frame of the prior segment: <continue the
      same action, subject, space, light, and camera intention>. <One next beat>.
      ```
      
      Inspect the boundary before later segments. If the scene, angle, time, or story
      emphasis changes, end the chain and design a cut.
      
      ## Native-audio markers
      
      ```text
      (<music description>)
      <sound-effect description>
      <character> says {dialogue}.
      On-screen caption: 【caption text】.
      ```
      
      Keep dialogue in one language. Enable native audio only when the chosen provider
      supports it and the shot needs it.
      
      ## Video-edit variant
      
      ```text
      Strictly edit Video 1: change <one specified property> to <new value>.
      Keep the subject, action, setting, lighting, and camera movement unchanged.
      ```
      
      For editing, refer to `Video 1` directly; do not call it a reference video.
      
    • prompt-templates.zh-CN.md 5.1 KB
      # 提示词模板
      
      把尖括号中的占位内容替换成具体信息。每次选择一条路线,不要把所有模板塞进同一请求。
      
      ## 主体参考图:生图模型,可选
      
      只有人物、产品、物体或场景需要反复出现时才使用。
      
      ```text
      为<主体>生成一张干净的参考图。
      只固定这些不变量:<轮廓或比例>、<标志性材质或服装>、<主色>、<不可改变的标识>。
      <人物:自然表情的脸部近景,简单背景,无文字。>
      <产品:清晰的三分之四视角,干净背景,无文字。>
      ```
      
      重复人物分别生成脸部近景和全身图。多视角转身图只可用于生图阶段的设计辅助,不要作为人物视频身份参考上传。
      
      ## Seedream Storyboard 参考图:Seedance R2V 默认路线
      
      用它生成将直接作为 `图片1` 上传到 Seedance 的一张 Storyboard。编号、箭头和简短注释可以保留,只要它们有助于理解顺序。
      
      ```text
      为以下连续内容生成一张<行数>x<列数>的视觉分镜图,阅读顺序为从左到右、从上到下:
      <整体剧情或动作概述>。
      
      第1格:<镜头目的、构图、动作>。
      第2格:<镜头目的、构图、动作>。
      ……
      
      所有格子中的<人物、产品或场景>保持一致:<不变量>。
      每格必须清楚呈现计划中的构图和动作。连续动作展示下一个有意义的可见瞬间;发生剪辑时清楚展示新镜头。需要帮助表达顺序或运动时,可以加入编号、简短注释、箭头,以及明确的剪辑或转场意图。
      ```
      
      生成后先检查格子顺序、关键节奏、主体或场景一致性,以及任务关键的视觉元素是否可读。关键格出错时重做这张整图。通过后整张上传给 Seedance R2V;只有测试视频真的把备注或编号带进画面时才做清理版。
      
      ## 清理版 Storyboard:可选重试或 I2V 切格来源
      
      ```text
      为以下连续内容生成一张干净的<行数>x<列数>视觉分镜图:
      <整体剧情或动作概述>。
      
      每格只包含计划中的镜头构图和动作。<人物、产品或场景>保持一致:<不变量>。
      只使用细的中性格线。不要加入编号、说明文字、气泡、标签、字幕、标志、无关叠加标记、界面元素或拼贴说明文字。
      ```
      
      只有原分镜实测把文字或格线带入视频,或明确需要切格用于 I2V 时才使用。切格前检查布局、位置和比例。
      
      ## 整张 Storyboard R2V:格子仍清晰可读时的首选
      
      把整张 Storyboard 作为 `图片1` 上传。
      
      ```text
      严格按图片1中的分镜图从左到右、从上到下的顺序生成短片,一次只出现一个完整镜头。
      
      镜头1:<第1格的动作和镜头语言>。
      镜头2:<第2格的动作,以及与上一镜头的剪辑或转场关系>。
      镜头3:<第3格的动作和镜头语言>。
      
      全片保持<主体、产品或场景不变量>;<全局光影或风格规则>。
      <启用时的声音规则>。
      图片1中的格线、编号、箭头和备注只用于理解分镜顺序与镜头意图,不要将它们渲染为最终视频画面。
      ```
      
      不要先切格。决定因素不是格子数量,而是每格视觉细节是否仍清晰可读。只有单镜头重拍或精确起止状态比一次生成的自然转场更重要时,才改用 I2V 首尾帧。
      
      ## R2V 参考素材:少量素材生成一条片段
      
      只上传职责不同的素材。
      
      ```text
      以图片1中的<人物或产品>为主体;图片2作为<场景或服装>参考;
      图片3作为<画面风格或构图>参考;<视频1或音频1的职责,如有>。
      
      <一个主要动作>。
      <运镜和光影>。
      保持<2至3条关键不变量>。
      ```
      
      不要把大量连续分镜格伪装成素材包。连续剧情使用整张 Storyboard 路线或独立 I2V 镜头。
      
      ## I2V 首尾帧:一个镜头的起止控制
      
      首帧作为 `图片1`,目标尾帧作为 `图片2` 上传。
      
      ```text
      从图片1的画面开始,主体为<主体>。
      <一个连续、可观察的动作>。
      镜头<一个运镜>,<空间或重要物体关系>。
      画面自然收束到图片2的构图与主体状态。
      <光影或风格>。
      保持<身份、产品标签或场景不变量>。
      ```
      
      如果是有意硬切,不要把 `图片2` 作为下一镜头的首帧;下一镜头使用自己的首尾帧。
      
      ## 连续延长或接龙
      
      只有同一镜头、同一不中断动作才用上一段真实尾帧作为下一段首帧。
      
      ```text
      无缝承接上一段的最后画面:<延续同一动作、主体、空间、光影和镜头意图>。
      <下一个动作节点>。
      ```
      
      继续下一段前检查边界。发生场景、机位、时间或叙事重点变化时,结束接龙并设计剪辑。
      
      ## 原生音频标记
      
      ```text
      (背景中播放着<音乐描述>)
      <音效描述>
      <角色>说道{台词}。
      画面出现【字幕内容】。
      ```
      
      台词只使用一种语言。只有所选服务支持且镜头确实需要时才启用原生音频。
      
      ## 视频编辑变体
      
      ```text
      严格编辑视频1,将<一个明确属性>修改为<新值>,
      其余主体、动作、场景、光影和运镜保持不变。
      ```
      
      编辑时直接引用 `视频1`,不要写 `参考视频1`。
      
    • real-person.md 5.4 KB
      # Believable human subjects
      
      Two failures recur with generated people: an unmistakable synthetic look, and
      multiple characters converging on the same face. Both are addressed by describing
      along fixed dimensions rather than piling on adjectives.
      
      **Read the omission rules first.** This detail is written for text-driven
      subjects. With an identity reference supplied, most of it becomes harmful.
      
      ## The dimensions
      
      ```text
      [age / ethnicity]      specific age + specific origin + style adjective + face-shape noun
      [skin tone / texture]  warm-cool + specific tone + texture adjective + realism suffix
      [facial features]      eye shape + brow structure + nose + lips + jaw  (3–4 is enough)
      [eyes / interiority]   quality of gaze + what it conveys + the emotion underneath
      [hair]                 colour + condition/texture + specific style + how it interacts
                             with the environment
      [wardrobe / fabric]    cut + colour + garment noun + material and wear + how it is worn
      [build / bearing]      frame and shoulders + framing + action or gaze + overall aura
      ```
      
      Examples of each, kept short deliberately:
      
      ```text
      [age / ethnicity]      22-year-old East Asian woman with a classical, gentle face
      [skin]                 cool-toned fair skin with a delicate sheen, retaining real
                             fine pores and natural skin texture
      [facial features]      slender eyes with slightly moist rims, relaxed brows, a
                             straight delicate nose bridge, a soft jawline
      [eyes]                 a gentle, focused gaze carrying deep attachment and a trace
                             of reluctance
      [hair]                 jet-black hair in a low classical bun held by a plain jade
                             pin, a few loose strands at the cheeks moving in the breeze
      [wardrobe]             minimalist cross-collar white robe in soft silk with a faint
                             lustre, collar slightly loosened
      [build / bearing]      slender frame, narrow shoulders; chest-up framing with direct
                             eye contact; gentle classical bearing
      ```
      
      ## Where the anti-synthetic effect actually comes from
      
      Two dimensions carry most of it:
      
      **Skin texture.** The realism suffix — retaining real fine pores and natural skin
      texture, optionally freckles or blemishes — counteracts a smoothing default.
      
      **Interiority.** A gaze that conveys something specific is what separates a
      photograph of a person from a rendering of a face.
      
      The rest of the dimensions mainly do a different job: **keeping characters
      distinct from one another.** Two characters described along the same seven
      dimensions with different values will not converge; two characters described as
      "a handsome man" and "a beautiful woman" will.
      
      ### The realism suffix is model-dependent
      
      It exists to counteract a *particular* model's default. It does not transfer
      blindly:
      
      - On a model that already renders coarse skin, it **overshoots** — faces come out
        dirty or aged.
      - The correct test is cheap: same prompt with and without the suffix. If the
        version without it already lands, drop it.
      - Record both outcomes in [model profile](model-profile.md), including the
        overshoot. A phrase that overshoots is worse than one that does nothing.
      
      ## When to omit — per dimension, not all-or-nothing
      
      With an identity reference supplied, written detail **competes** with the image,
      and the image usually wins. But switching the whole block off throws away the two
      dimensions a reference cannot carry.
      
      | Situation | Dimensions to write |
      |---|---|
      | Text-driven, no identity reference | All of them |
      | Reference locks style or scene only | All of them |
      | **Reference locks the face or identity** | **Drop** facial features, skin, hair, wardrobe. **Keep** build/bearing and interiority |
      
      Build, bearing, and interiority stay because a still reference cannot express
      them — it has no gaze *behaviour* and no posture *over time*. Those are exactly
      what text is for.
      
      When a reference is supplied, describing the face again is not redundancy; it is
      a competing instruction. Expect drift.
      
      ## Performance, not appearance
      
      Appearance is static; a video needs behaviour. Add a performance line and keep it
      observable — two to four cues per emotional turn:
      
      ```text
      [performance core] restrained and nuanced: a shifting gaze, the rise and fall of
      breathing, a slight tremble at the lips, one tear tracking down without being
      wiped away
      ```
      
      Do not write `she looks sad`. Draw from: gaze direction and shift, brow tension,
      mouth movement, breathing, throat, hands, posture. Listing every facial detail
      does not increase control — the cues start competing.
      
      For several emotional turns, trigger each on an event rather than a timestamp:
      `when she hears the applause, her fingers stop on the programme`. Event-triggered
      turns do not depend on timing adherence, so they survive a model change.
      
      ## Negatives worth stating
      
      Only when the piece is actually at risk of them:
      
      ```text
      no exaggerated crying · no rapid cuts · no large body movements ·
      no extra dialogue or background music · no distorted features or extra fingers
      ```
      
      Keep this short and specific to the piece. An inherited blocklist makes the prompt
      longer without making it better, and dilutes the negatives that matter.
      
      ## Related
      
      - [multi reference](multi-reference.md) — binding an identity reference properly
      - [model profile](model-profile.md) — where the realism-suffix findings live
      - [troubleshooting](troubleshooting.md) — duplicate people, face drift
      
    • real-person.zh-CN.md 4.8 KB
      # 可信的人物
      
      生成人物反复出现两类失败:一眼能看出的合成感,以及多个角色长得越来越像。两者都靠**沿固定维度描述**来解决,而不是靠堆形容词。
      
      **先读省略规则。** 这套细节是给纯文字驱动的主体写的。一旦有身份参考图,其中大部分反而有害。
      
      ## 各维度
      
      ```text
      [年龄/人种]     具体年龄 + 具体出身 + 风格形容词 + 脸型名词
      [肤色/肤质]     冷暖调 + 具体色调名词 + 质感形容词 + 真实度后缀
      [面部细节]       眼型 + 眉骨 + 鼻梁 + 唇形 + 下颌  (3–4 项就够)
      [眼神/内在]     眼神的质地 + 它传达什么 + 底下压着的情绪
      [发型发色]       颜色 + 状态/质感 + 具体发型 + 与环境的互动
      [服装/面料]     版型 + 颜色 + 服装名词 + 材质与新旧 + 穿法细节
      [体型/气质]     骨架与肩线 + 构图 + 动作或视线 + 整体气场
      ```
      
      各维度示例,刻意写短:
      
      ```text
      [年龄/人种]  22 岁东亚女性,古典温婉的电影感面孔
      [肤质]        冷调白皙肤色,带细腻润泽感,保留真实的微毛孔与皮肤纹理
      [面部细节]    细长的眼型(眼眶微润)、眉形舒展、鼻梁小而直、下颌线柔和
      [眼神]        眼神温柔专注,含着深深的依恋与一丝不舍
      [发型]        黑发盘成随意而雅致的古典低髻,只用一支素玉簪固定,
                    几缕碎发落在脸颊边随风轻动
      [服装]        极简纯白交领古装,柔软有光泽的丝或轻纱,领口略松
      [体型/气质]  身形纤细、肩线窄;胸以上半身特写,直视镜头;
                    整体是温婉深情的古典气质
      ```
      
      ## 去合成感真正来自哪两维
      
      **肤质。** 那个真实度后缀——保留真实的微毛孔与皮肤纹理,必要时加雀斑或瑕疵——反的是磨皮默认值。
      
      **内在。** 一个传达了具体内容的眼神,才是「一张人的照片」与「一张脸的渲染」之间的区别。
      
      其余维度主要干的是另一件事:**让角色之间彼此区分**。沿同样七个维度、填不同值描述的两个角色不会趋同;描述成「一个帅气男人」和「一个漂亮女人」的两个角色一定趋同。
      
      ### 真实度后缀是模型相关的
      
      它的存在是为了反掉**某个特定模型的**默认值,不能盲目移植:
      
      - 在本来就渲染粗糙皮肤的模型上会**过冲**——脸会脏、会显老
      - 测试很便宜:同一份提示词,带这句跑一次、不带跑一次。不带的那版已经达标,就在这个模型上删掉它
      - **两个结果都记进** [model-profile.zh-CN.md](model-profile.zh-CN.md),包括过冲。过冲的写法比毫无作用的写法更糟
      
      ## 什么时候省略——逐维度,不是一刀切
      
      有身份参考图时,写出来的细节会跟图**竞争**,而且图通常会赢。但整块关掉,等于把参考图给不了的那两维一起丢掉。
      
      | 情况 | 该写哪些维度 |
      |---|---|
      | 纯文字驱动,无身份参考 | 全写 |
      | 参考图只锁风格或场景 | 全写 |
      | **参考图锁脸/锁身份** | **删掉**面部细节、肤质、发型、服装。**保留**体型气质与内在 |
      
      体型、气场、内在要保留,因为一张静态参考图表达不了它们——它没有眼神的**行为**,也没有随时间展开的**体态**。那恰恰是文字该干的活。
      
      有参考图时把脸再描述一遍,不是冗余,是一条**竞争性指令**,会引起漂移。
      
      ## 表演,不是外貌
      
      外貌是静态的,视频需要行为。加一条表演行,保持可观察——一次情绪转折 2–4 个线索:
      
      ```text
      [表演核心] 内敛克制:视线的游移、呼吸的起伏、唇线的轻颤,
                 一滴眼泪滑落而不去擦
      ```
      
      不要写「她看起来很难过」。可选:视线方向与移动、眉的张力、嘴部动作、呼吸、喉部、手、体态。把每个面部细节都列出来不会增加控制力——线索之间会开始互相打架。
      
      多次情绪转折时,用**事件**触发而不是时间点:`当她听到掌声,手指停在节目单上`。事件触发不依赖时序遵循度,所以能跨模型存活。
      
      ## 值得写的负向
      
      只在这条片子真的有这个风险时才写:
      
      ```text
      不许夸张哭泣 · 不许快速切镜 · 不许大幅肢体动作 ·
      不许多余台词或背景音乐 · 不许五官变形或多手指
      ```
      
      保持简短、针对本片。抄来的黑名单只会让提示词变长而不变好,而且会稀释真正重要的那几条负向。
      
      ## 关联
      
      - [multi-reference.zh-CN.md](multi-reference.zh-CN.md) — 身份参考图的正确绑定方式
      - [model-profile.zh-CN.md](model-profile.zh-CN.md) — 真实度后缀的实测结论存放处
      - [troubleshooting.zh-CN.md](troubleshooting.zh-CN.md) — 人物重复、脸部漂移
      
    • transitions.md 4.5 KB
      # Transitions
      
      Two decisions, in this order:
      
      1. **Should the model generate this transition at all,** or should the edit own it?
      2. If the model, **name the type at the cut point.**
      
      ## First: does this belong in the edit?
      
      A generation costs money and gives you less control than a timeline. Several
      common transitions are two seconds of work in any editor:
      
      | Transition | Where it belongs | Why |
      |---|---|---|
      | Straight cut | Edit | Nothing to generate |
      | Fade to/from black or white | Edit | Exact, adjustable, free |
      | Cross dissolve | Edit | Same, and easier to time |
      | Flash frame (white/black) | Edit | Frame-accurate there, approximate here |
      | Wipe / slide | Edit | A geometric effect, not a photographed event |
      
      Generate the transitions that are **photographed events** — things an editor
      cannot fabricate from two finished clips:
      
      | Transition | What it is | Suits |
      |---|---|---|
      | **Occlusion** | Camera moves until an object fills frame; the next shot pulls out of it | Space and time jumps |
      | **Match object** | The outgoing frame's shape, outline, or colour becomes the incoming one | Montage, conceptual links |
      | **Motion / whip** | Rapid camera movement blurs out and resolves in the next scene | Speed, action, travel |
      | **Action relay** | A subject's movement carries across: they leave frame one way and enter the next continuing it | Outfit changes, location hops |
      | **Push / pull through** | A detail is magnified until it becomes the next scene, or the reverse | Scale shifts, worlds-within-worlds |
      | **Material spread** | A medium — ink, smoke, light — spreads across frame and the next scene emerges from it | Stylised and culturally specific work |
      
      These need generation because the camera and the subject have to *do* something
      during the transition. That is also why they read as craft rather than as effects.
      
      ## Second: name it at the cut point
      
      The skeleton is one line:
      
      ```text
      Use a <transition type> at the cut.
      ```
      
      Then, if the transition needs staging, describe the mechanics — what triggers it,
      which direction things move, and what the next shot opens on:
      
      ```text
      Use an occlusion transition at the cut. The camera pushes forward until a person's
      back fills the frame and the image goes dark; the next shot pulls out of the
      darkness onto the gallery interior.
      ```
      
      For the trickier types, state: **trigger action → camera movement → the visual
      transformation → the arrival state.** Arrival state is the one people skip, and it
      is what stops the incoming shot from drifting after the transition completes.
      
      ## Do not attach "no hard cut" by default
      
      `no hard cut` and `nothing appears from nowhere` are **extension and continuation
      defaults**. In that context they are correct: a broken seam and objects
      materialising are the two characteristic failures of continuing existing footage.
      
      Everywhere else, a hard cut or a sudden appearance is a **technique** — teleports,
      jump scares, magic reveals, comedic timing. Attaching those constraints globally
      removes tools you may want.
      
      Enable them as a scoped preset:
      
      - Extending or chaining a continuous action → **on**
      - Deliberate discontinuity → **off**
      - Everything else → decide per piece
      
      The general rule this instantiates: a default that is correct in one scope is not
      a global rule. The same mistake shows up as "every beat must end on a camera move"
      — true in a piece where camera movement *is* the premise, wrong as a universal.
      
      ## Letting the model choose
      
      When the piece has a consistent style but no specific transition requirement, a
      constrained menu works better than either silence or a hard pick:
      
      ```text
      At each cut, choose whichever suits this piece from: occlusion, match object,
      motion, or material spread.
      ```
      
      Silence gets you an arbitrary transition; a hard pick may fight the material. A
      menu keeps the choice inside a set you have already approved.
      
      ## Multi-shot support is model-dependent
      
      All of this assumes the model can produce ordered shots with cuts in one
      generation. Some models always bridge shots with movement instead. Verify before
      building a spec around cuts — see [model profile](model-profile.md) and
      [capabilities](capabilities.md).
      
      When multi-shot is unavailable, generate one shot per request and assemble the
      cuts in the edit. That also moves the five edit-owned transitions above back where
      they belong.
      
      ## Related
      
      - [editing and extension](editing-and-extension.md) — where the seam defaults apply
      - [cinematography](cinematography.md) — camera vocabulary
      - [troubleshooting](troubleshooting.md) — seams that jump or feel arbitrary
      
    • transitions.zh-CN.md 4.1 KB
      # 转场
      
      两个决定,按这个顺序:
      
      1. **这个转场该不该由模型生成**,还是该由剪辑来做?
      2. 如果由模型,**在切点写明类型**。
      
      ## 先问:这个该由剪辑做吗
      
      一次生成要花钱,而且比时间线给你的控制少。几种常见转场在任何剪辑软件里都是两秒钟的事:
      
      | 转场 | 该在哪做 | 为什么 |
      |---|---|---|
      | 硬切 | 剪辑 | 没有东西需要生成 |
      | 淡入淡出黑/白 | 剪辑 | 精确、可调、免费 |
      | 叠化 | 剪辑 | 同上,而且时间点更好控 |
      | 闪白/闪黑 | 剪辑 | 那里是帧精确的,这里只是近似 |
      | 擦除/推移 | 剪辑 | 这是几何效果,不是被拍下来的事件 |
      
      值得生成的是那些**被拍摄下来的事件**——剪辑无法从两条成片里凭空造出来的:
      
      | 转场 | 是什么 | 适合 |
      |---|---|---|
      | **遮罩** | 镜头前进直到某个物体填满画面,下一镜从中拉出 | 空间与时间的跳跃 |
      | **相似物** | 前一镜末帧的形状、轮廓或颜色变成下一镜的首帧 | 蒙太奇、概念性连接 |
      | **动作/甩镜** | 快速运镜糊掉,在下一场景中解析出来 | 速度、动作、旅行 |
      | **动态接力** | 主体的动作跨镜延续:从一侧出画,在下一镜延续着入画 | 换装、地点跳跃 |
      | **推拉穿越** | 一个细节被放大到成为下一场景,或反过来 | 尺度切换、世界中的世界 |
      | **材质蔓延** | 一种媒介——墨、烟、光——铺满画面,下一场景从中浮现 | 风格化与文化特定的作品 |
      
      这些需要生成,因为转场期间**镜头和主体必须真的做点什么**。这也是它们读起来像手艺而不是效果的原因。
      
      ## 然后:在切点写明类型
      
      骨架只有一行:
      
      ```text
      在切点使用<转场类型>。
      ```
      
      如果这个转场需要调度,再写清机制——什么触发它、往哪个方向动、下一镜开在什么状态:
      
      ```text
      在切点使用遮罩转场。摄影机向前推进,直到一个人的背影填满画面、画面转暗;
      下一镜从黑暗中拉出,露出展厅内部。
      ```
      
      对较难的几类,要写明:**触发动作 → 运镜 → 视觉变形 → 到达状态**。到达状态是最常被省略的一项,也正是它防止转场完成之后新镜头继续漂。
      
      ## 不要默认挂「禁止硬切」
      
      `禁止硬切` 和 `禁止物体凭空出现` 是**延长与续写的默认值**。在那个上下文里它们是对的:接缝断裂和物体凭空冒出来是续写已有素材的两个典型翻车。
      
      在别处,硬切和凭空出现是**手法**——闪现、跳惊、魔法显形、喜剧节奏。全局挂上等于把你可能想用的工具删掉。
      
      按场景 opt-in:
      
      - 延长或接龙一个连续动作 → **开**
      - 刻意制造不连续 → **关**
      - 其余情况 → 每片一决定
      
      这条实例化的通则是:**在一个作用域里正确的默认值,不是全局规则。** 同一个错误的另一副面孔是「每一拍都必须以运镜收尾」——在某条以运镜为命题的片子里成立,当普适规则就错了。
      
      ## 让模型自己挑
      
      片子风格一致、但没有具体转场要求时,**给一个受限菜单**比沉默或硬指定都好:
      
      ```text
      每个切点从遮罩、相似物、动作、材质蔓延中选择最适合本片的一种。
      ```
      
      沉默会得到一个随意的转场;硬指定可能跟素材打架。菜单把选择限制在你已经认可的集合里。
      
      ## 多镜头支持是模型相关的
      
      以上全部假设模型能在一次生成里产出带切点的有序镜头。有些模型总是用运动把镜头连起来。围绕切点搭 spec 之前先核实——见 [model-profile.zh-CN.md](model-profile.zh-CN.md) 与 [capabilities.zh-CN.md](capabilities.zh-CN.md)。
      
      多镜头不可用时,一镜一次请求,切点在剪辑台上完成。那也顺带把上面那五种「剪辑该做的」转场放回了它们该在的地方。
      
      ## 关联
      
      - [editing-and-extension.zh-CN.md](editing-and-extension.zh-CN.md) — 接缝约束该用在哪
      - [cinematography.zh-CN.md](cinematography.zh-CN.md) — 运镜词汇
      - [troubleshooting.zh-CN.md](troubleshooting.zh-CN.md) — 接缝跳动或转场生硬
      
    • troubleshooting.md 4.8 KB
      # Troubleshooting
      
      Fix the failed shot or seam rather than restarting a complete sequence.
      
      ## Identity and reference problems
      
      | Symptom | Likely cause | Targeted fix |
      |---|---|---|
      | Person changes face or outfit | Weak or mixed identity reference | Use a clean face close-up and separate full-body image; name their roles clearly |
      | Duplicate people | A composite human reference reads as several subjects | Replace the turnaround or contact sheet with one single-person image; state the one-person constraint |
      | Product form or label changes | Product reference has insufficient priority | Use one clean product image, bind it as `Image 1`, and explicitly preserve proportion and label |
      | Style changes across cuts | Conflicting light or art direction | Lock the same dominant light direction, palette, and material language in every relevant shot |
      
      ## Motion and shot problems
      
      | Symptom | Likely cause | Targeted fix |
      |---|---|---|
      | Wobble or morphing | Too much action or competing camera instructions | Reduce to one continuous action and one primary move, or describe one synchronized composite move |
      | Face warps in close-up | Tight framing plus aggressive rotation | Use wider framing, gentler motion, or a planned cut |
      | Endpoint is missed | I2V lacks a usable end keyframe | Regenerate a clean end keyframe and use it as `last_image` |
      | R2V ignores storyboard order | Sequence is ambiguous or panels are too small | Write `Shot 1 / Shot 2 / Shot 3`; simplify only the ambiguous board; switch to I2V only when individual control is required |
      
      ## Seams and edit problems
      
      | Symptom | Likely cause | Targeted fix |
      |---|---|---|
      | A continuous chain jumps or rewinds | Generated tail does not align sufficiently with the next segment | Rework the ending action, regenerate the boundary segment, and trim only after visual inspection |
      | Hard cut feels arbitrary | No planned visual or audio relationship | Add a match cut, occlusion, cutaway, or strong beat before the cut |
      | Dissolve looks muddy | Two unrelated compositions are blended | Use a hard, match, or occlusion cut; reserve dissolves for true time or place changes |
      | Input audio fails | A supplied reference audio file or URL was rejected | Fix or replace `audio.references`; do not expect silent removal of the input audio |
      | Output audio fails after retries | Seedance returned no audio stream or reported audio-generation failure | Inspect the request and rerun that segment with `generate_audio:true` rather than silently falling back to silent video |
      | Native audio ends with a click | Generated clip audio is truncated | Regenerate the segment with native audio; add a short fade only after the new generation is acceptable |
      
      ## Multi-reference problems
      
      | Symptom | Likely cause | Targeted fix |
      |---|---|---|
      | Wrong reference applied to a subject | Bound as a group (`images 1–4 define four characters`) | One binding line per subject, each naming its own reference |
      | One object appears twice | Multi-view references read as several objects | State the output count: `all three images define one lamp; exactly one appears throughout` |
      | Props migrate between characters | Ownership never stated | Add `belongs only to <subject>` |
      | Every reference crowds into every shot | No per-scene selection | List which references each scene uses |
      | Reference background leaks in | Role stated one-sided | Add the `do not use` half |
      | Re-described motion conflicts with a motion reference | Text competes with the video | State only which attributes to inherit |
      
      ## Stage and timing problems
      
      | Symptom | Likely cause | Targeted fix |
      |---|---|---|
      | Stage ends in the wrong state | No end state written | State what is visibly true when the stage ends |
      | Piece trails off | Final stage has no end state | Add one — the most common omission |
      | Model invents pauses or extra cuts | Second-level granularity on a continuous action | Drop to stages, or to event order only |
      | Beats land consistently late | Model timing drift | Check the profile; write beats earlier, or drop to stages |
      | Events dropped inside a range | Too much content in one time range | Split into more stages |
      | Style holds early then drifts | A global rule was written inside stage 1 | Move it to the global block |
      | Later characters appear too early in a backward extension | Source's first frame not stated as an end state | State it explicitly, and name what must not appear early |
      
      ## Review order
      
      Review identity, locks, stage end states, composition, motion, seam, then sound.
      **Stop at the first failure** — later polish cannot repair wrong identity or a
      missed endpoint, so checking past the first break wastes the pass.
      
      When the same lock breaks repeatedly on one model, that is a profile finding, not
      a prompt problem. Record it in [model profile](model-profile.md) instead of
      rewriting the prompt again.
      
    • troubleshooting.zh-CN.md 4.5 KB
      # 问题排查
      
      默认只修复失败的镜头或接缝,不要因为一个问题重启整条序列。
      
      ## 身份与参考图问题
      
      | 症状 | 常见原因 | 定向修复 |
      |---|---|---|
      | 人物脸或服装变化 | 身份参考弱或职责混杂 | 使用干净脸部近景和独立全身图,并清楚说明各自职责 |
      | 出现重复人物 | 人物拼图被理解为多个主体 | 用单人图替换转身图或拼贴图,并明确同框只保留一个人物 |
      | 产品形态或标签变化 | 产品参考视觉优先级不足 | 使用一张干净产品图,将其绑定为 `图片1`,明确保持比例和标签 |
      | 剪辑后画风变化 | 光影或美术方向互相冲突 | 所有相关镜头锁定同一主光方向、色板和材质语言 |
      
      ## 运动与镜头问题
      
      | 症状 | 常见原因 | 定向修复 |
      |---|---|---|
      | 画面抖动或变形 | 动作过多或运镜指令冲突 | 减为一个连续动作和一个主导运镜,或写成一个同步的复合运镜 |
      | 脸部近景变形 | 景别太近加大幅旋转 | 换更宽景别、更轻动作,或使用计划好的剪辑 |
      | 没有达到目标尾帧 | I2V 没有可用尾帧 | 重做干净尾帧并将其设为 `last_image` |
      | R2V 忽略分镜顺序 | 顺序含糊或格子细节太小 | 写明 `镜头1 / 镜头2 / 镜头3`;只简化有歧义的分镜;只有需要逐镜头控制时才切 I2V |
      
      ## 接缝与剪辑问题
      
      | 症状 | 常见原因 | 定向修复 |
      |---|---|---|
      | 连续接龙跳帧或倒退 | 真实尾帧与下一段不够对齐 | 重做收尾动作,重生成边界片段,只在目视检查后裁切 |
      | 硬切显得突兀 | 没有视觉或声音关系 | 在剪辑前加入匹配、遮挡、插入镜头或强节奏点 |
      | 叠化显得浑浊 | 两个无关构图被强行混合 | 改为硬切、匹配切或遮挡切;叠化只用于真正的时空变化 |
      | 输入音频报错 | 提供的参考音频文件或 URL 被拒绝 | 修复或替换 `audio.references`;不要期待脚本静默移除输入音频 |
      | 重试后输出音频仍失败 | Seedance 未返回音频流或报告音频生成失败 | 检查请求并让该片段以 `generate_audio:true` 重生成,不要静默降级为无声视频 |
      | 原生音频末尾有爆音 | 生成音频被截断 | 先重生成该片段;新结果可接受后再考虑加短淡出 |
      
      ## 多素材问题
      
      | 症状 | 常见原因 | 定向修复 |
      |---|---|---|
      | 素材用错在了某个主体上 | 按组绑定(`图片1到4分别定义四个角色`) | 一个主体一行绑定,各自点明自己的素材 |
      | 同一个物体出现两次 | 多视角素材被读成多个物体 | 写明输出数量:`三张图定义同一盏灯;全片只出现一盏` |
      | 道具在角色之间转移 | 从未写明归属 | 加 `只属于<主体>` |
      | 每份素材都挤进每个镜头 | 没有按场景挑选 | 列出每个场景要用哪些素材 |
      | 参考图的背景漏进来 | 素材职责只写了一半 | 补上 `不要用……` 那半 |
      | 重复描述的运动与动作参考打架 | 文字在跟参考视频争夺控制 | 只写要继承哪些属性 |
      
      ## 分段与时序问题
      
      | 症状 | 常见原因 | 定向修复 |
      |---|---|---|
      | 段落结束时状态不对 | 没写末态 | 写明这一段结束时什么是可见的 |
      | 片子散掉、没有收住 | 最后一段没写末态 | 补上——最常见的遗漏 |
      | 模型造出多余的停顿或切换 | 对连续动作用了秒级粒度 | 降到分段,或只写事件顺序 |
      | 拍点稳定偏晚 | 该模型的时序漂移 | 查档案;把拍点写早,或降到分段 |
      | 某个时间范围里的事件被漏掉 | 一个范围里塞了太多内容 | 拆成更多段 |
      | 风格前面守住、后面漂了 | 全局规则被写进了第 1 段里 | 移到全局设定 |
      | 向后延长时后段角色提前出现 | 源片首帧没写成末态 | 明确写成末态,并点明什么不许提前出现 |
      
      ## 检查顺序
      
      按主体一致性、锁、段落末态、构图、运动、接缝、声音的顺序检查。**第一个失败就停**——后续润色无法弥补错误的身份或落点,继续往下查是浪费这一轮。
      
      同一条锁在某个模型上反复破,那是**档案发现**,不是提示词问题。记进 [model-profile.zh-CN.md](model-profile.zh-CN.md),不要反复改提示词。
      
      **⚠️ 只看抽帧有盲区。** 静帧能定质感、构图、身份、末态,但说不了动态质量、转场流畅度、节奏与音画同步。绝不要只凭静帧下整体结论——要么看回放,要么明说结论只覆盖哪一半。
      
    • workflow.zh-CN.md 13.6 KB
      # Seedance 2.5 Skill 中文工作流
      
      本文件是中文请求的主流程。模型 ID、JSON 字段、命令、媒体占位符和音频符号属于代码,不翻译。
      
      ## 1. 先选视频路线
      
      | 需求 | 路线 | 输入 | 结果 |
      |---|---|---|---|
      | 一个简短、单一的场景 | T2V | 文字提示词 | 一条自洽镜头 |
      | 多镜头且每格仍清晰可读 | 整张 Storyboard R2V | 一张完整分镜图 | 一次请求生成连续短片 |
      | 需要人物、产品、场景或风格素材控制 | R2V 参考素材 | 少量、分工明确的参考图 | 一条受素材控制的片段 |
      | 必须精确控制起止状态 | I2V 首尾帧 | 首帧和尾帧 | 一条可单独重做的镜头 |
      | 同一个不中断动作超过平台时长 | 延长或接龙 | 上一段真实尾帧和下一段提示词 | 同一镜头的延续 |
      | 一条片子里承载多个事件 | 分段直出 | 文字加可选参考素材 | 一次请求覆盖有序的多个段落,每段落在写明的末态上 |
      | 对已有视频做限定范围的改动 | 编辑 | 源视频加目标参考素材 | 只改指定区域或元素的源片 |
      | 两条成片之间的桥接 | 无缝转场 | 两条视频 | 生成中间的衔接内容 |
      | 从 3D 预览取运动、调度与运镜 | 白模参考 | 粗或精白模视频加外观参考 | 按白模的时序与调度渲出的成片 |
      
      当前默认可执行模型是 Seedance 2.0。只有所选服务已真实提供 Seedance 2.5 及其限制时,才使用 2.5 路线。
      
      后四条路线依赖的能力逐模型、逐服务不同。**提供之前先核实可用性**——2.0 与 2.5 的对照、以及哪些已发布能力属于平台功能而非 API 参数,见 [能力](capabilities.zh-CN.md)。
      
      ### 整张 Storyboard 与逐镜头的选择
      
      - 默认把**整张 Storyboard 作为一张参考图进行一次 R2V 请求**,前提是每格仍能看清。不要先切格。
      - 只有当需要单独重做某镜头,或必须精确指定开场和收束状态时,才选择 I2V 首尾帧。
      - 不要把每个分镜格都当作 R2V 素材图上传。多张 R2V 图只适用于它们分工不同的情况,例如人物、产品、场景、风格、动作参考。
      
      ## 2. 按需启用准备模块
      
      ### 主体设定
      
      只有会反复出现的人物、产品、道具、手、车辆或场景需要主体设定。保留 3–5 条不变量:轮廓或比例、标志性材质或服装、主色,以及不可改变的标识。一次性氛围镜头可以跳过。
      
      反复出现的人物使用一张干净脸部近景和一张独立全身参考图。不要把正侧背拼图当作人物身份参考,它可能被理解为多人。需要展示物体多个角度时,产品多角度图仍然有用。
      
      ### 关键帧
      
      每个 I2V 镜头都需要首帧。只有镜头必须落在特定动作、构图、产品姿态或手部位置时,才加尾帧。每张关键帧只包含一个干净场景。
      
      ### Storyboard
      
      已有 Storyboard 默认直接作为 R2V 参考。Seedance 通常能理解格子顺序;不需要预先删掉编号、格线、箭头或备注。只有实测视频把这些元素渲染进最终画面时,才生成清理版。
      
      没有 Storyboard 时,使用 [中文提示词模板](prompt-templates.zh-CN.md) 中的 Seedream 模板生成一张。**必须在宿主界面展示图片,自己检查后再继续 R2V。** 这只是中间状态展示,不需要等待用户确认。
      
      检查格子顺序、关键节奏是否可读、重复主体是否一致,以及该任务真正重要的结构条件,例如人物解剖、产品形态或关键文字。发现明显错误先自行重做;只有创意取舍无法推断时才询问用户。
      
      只有明确改为 I2V 首尾帧路线时才切格。切格前先检查布局;自动切图只能确认位置,不能判断模型是否真的画对了每格内容。
      
      ## 3. 连续性与转场
      
      首尾帧只控制**一个镜头**,不代表每个镜头都要继承上一段尾帧。
      
      | 转场类型 | 下一段是否用上一段尾帧作首帧 | 设计原则 |
      |---|---:|---|
      | 同一动作、同一镜头不中断 | 是 | 按顺序生成并检查接缝 |
      | 切换机位、地点、产品或时间 | 否 | 两段独立设计 |
      | 匹配剪辑 | 通常否 | 匹配动作方向、形状、颜色或构图 |
      | 遮挡或甩镜转场 | 否 | 上一镜头以遮挡结束,下一镜头从遮挡内或之后开始 |
      | 插入镜头或空镜 | 否 | 用物体、环境或产品细节做桥接 |
      
      把重要的剪辑、匹配和遮挡写进 Storyboard 和视频提示词。不要用叠化硬修两个无关镜头。
      
      ## 4. 写视频提示词
      
      ### 作用域:每条指令放在它管得住的地方
      
      写模块之前,先按**这行字管什么范围**把已知信息分类。放错位置是漂移最常见的原因——写在第 1 拍里的全局规则,到第 4 拍就不管用了。
      
      | 作用域 | 管什么 | 内容 |
      |---|---|---|
      | 全局 | 整条片子 | 片型、场景、风格、一句话导演命题、运镜原则 |
      | 锁 | 任何不许漂的东西 | 身份、参考素材职责、音频源、连续性、负向 |
      | 时序 | 某一拍或某一段 | 段落事件与各自末态 |
      
      把两三条最贵的锁在提示词的**物理结尾**复述一次(近因效应有用)。这是放置约定,不是第四个作用域——内容仍属「锁」,并且先在那里出现过。
      
      这与[通用视频提示词 Skill](../../universal-video-prompt-skill/references/workflow.zh-CN.md) 的模型无关 spec 格式一致。同一份需求要跑多个模型时用那个 Skill;本文件用于 Seedance 专属的写法。若此链接不可达,说明安装不完整;继续前先协助用户安装 `universal-video-prompt-skill`,不要自行编造缺失的共享规则。
      
      ### 模块
      
      只写会影响镜头的模块:
      
      ```text
      [主体或参考图绑定]
      + [一个可观察动作]
      + [空间与重要物体关系]
      + [一个主导运镜、同一意图的复合运镜,或剪辑]
      + [必要时的光影和风格]
      + [启用时的声音或台词]
      + [末态,任何必须落在特定位置的段落都要写]
      + [必须保持的约束]
      ```
      
      ### 末态承载多事件的片子
      
      只要片子有一个以上事件,就写明每段结束时**什么是可见的**。这是给多段提示词加的杠杆最大的一项:它把「保持一致」变成模型能瞄准、你也能检查的东西。
      
      ```text
      弱:  两个人继续弄那束花
      强:  末态:花艺师左手持花束;剪刀回到工作台右侧
      ```
      
      末态必须看得见。`她感到释然` 不是;`她的肩膀落下、皱起的眉松开` 是。完整分段结构见 [long-video.zh-CN.md](long-video.zh-CN.md)。
      
      ### 时序粒度:写拍点之前先定
      
      粒度是前置决策。按秒级写完拍点再想降级,等于重写。
      
      | 粒度 | 什么时候用 |
      |---|---|
      | 不写时序——只给事件顺序 | 单一连续动作、氛围片、单镜头。写了秒数反而把镜头切碎 |
      | **分段 + 末态** | 绝大多数叙事片。**默认** |
      | 秒级 | 只在有外部硬约束时:音乐、口型、素材交接、必须落在固定时间的拍点 |
      
      能从输入推出来的就推——给了音乐或口播音轨=秒级,说了氛围片=不写,有明确固定拍点=秒级。当需求是**多事件叙事、无外部约束**时,**要问,而且要给建议并说明理由**,不要甩一个空的选择题。
      
      时间戳分配的是时间预算,**不是帧精确的剪辑点**,动作可能落在边界前后。不要要求「一秒内完成三个动作」这类不可能的密度。
      
      - 明确每张参考图的职责,例如 `图片1中的人物`、`图片2中的产品`、`图片3中的厨房场景`。
      - 整张 Storyboard R2V 按 `镜头1`、`镜头2`、`镜头3` 写事件顺序;时长交给服务参数,不在文字里强写逐秒时间表。
      - 每个镜头优先一个主导运镜。允许复合运镜,但方向、与主体的关系、速度必须共同服务于一个意图。
      - 启用原生音频时,`()` 表示音乐,`<>` 表示音效,`{}` 表示台词,`【】` 表示画面文字。
      - 只写代价高的约束,避免冗长负面词。
      
      按任务读对应的参考文件:
      
      | 文件 | 什么时候读 |
      |---|---|
      | [提示词模板](prompt-templates.zh-CN.md) | 各路线的模板 |
      | [提示词组件库](prompt-blocks.zh-CN.md) | 可复用的镜头、声音、约束词块 |
      | [长视频](long-video.zh-CN.md) | 分段结构、末态、时间戳规则 |
      | [多素材绑定](multi-reference.zh-CN.md) | 素材多而不混淆 |
      | [人物](real-person.zh-CN.md) | 可信的人,以及什么时候该省略这些细节 |
      | [转场](transitions.zh-CN.md) | 哪些该生成、哪些该交给剪辑 |
      | [编辑与延长](editing-and-extension.zh-CN.md) | 改动或续写已有视频 |
      | [能力](capabilities.zh-CN.md) | 2.0 与 2.5 的上限;平台功能与 API 参数的分界 |
      | [模型档案](model-profile.zh-CN.md) | 实测的逐模型行为与编译注记 |
      | [电影语言词库](cinematography.zh-CN.md) | 细化镜头语言 |
      | [问题排查](troubleshooting.zh-CN.md) | 按症状查修法 |
      | [执行适配器](execution-adapters.zh-CN.md) | 脚本与适配器配置 |
      
      ## 5. 生成、检查、收尾
      
      1. 需要生成 Storyboard 时,先只生成静帧,在对话中展示,检查后才发起视频请求。
      2. 用户提供 Storyboard 时,若对话中还不可见则展示它,然后检查是否适合所选路线。
      3. 分镜通过后直接生成一条代表性视频,不等待用户确认;不通过则先优化分镜。
      4. 按主体一致性、锁、段落末态、构图、运动、接缝、音频的顺序检查,**第一个失败就停**——身份或落点错了,后面全白查。只重做出错镜头或片段。
      5. 接龙必须按顺序生成,因为下一段要用真实尾帧。剪辑型视频则独立生成各镜头,再在时间线中处理预先设计的转场。
      
      **⚠️ 只看抽帧有盲区。** 静帧能定质感、构图、身份、末态,但说不了动态质量、转场流畅度、节奏与音画同步。绝不要只凭静帧下整体结论——要么看回放,要么明说结论只覆盖哪一半。
      
      脚本只负责生成草稿和拼接,不负责专业调色或配乐混音。
      
      ## 6. Atlas 执行路线
      
      创作路线与提交方式相互独立。默认使用 Seedream 5.0 Pro 生图、Seedance 2.0 生视频;改模型前先核实其支持所需路线。
      
      在智能体对话中,默认由 **Atlas Cloud Skill 直接执行生成**。它可以查模型、上传本地素材、提交图片或视频任务、轮询并取回结果。只有它实际提交了任务,才报告 `Execution: atlas-skill`。
      
      只有用户明确选择 MCP 且当前客户端暴露了生成工具时,才用 `atlas-mcp`。只有用户明确选择终端、脚本、CI 或批量执行时,才用 `atlas-cli`。若 Atlas Cloud Skill 缺失,先协助安装 `AtlasCloudAI/atlas-cloud-skills`,再选择替代路径。
      
      报告“缺少 Atlas Cloud API Key”之前,必须检查**实际选中的执行进程**。REST 脚本先检查 `ATLASCLOUD_API_KEY`,再兼容检查 `ATLAS_CLOUD_API_KEY`。不要根据另一个服务商、插件或进程的状态推断凭据是否存在;不同执行通道可能拥有相互独立的凭据作用域。
      
      如果两个变量都不存在,引导用户前往 `https://www.atlascloud.ai/console/api-keys?utm_source=github&utm_campaign=awesome-seedance-2.5-prompts-skills` 获取 Key。不要让用户把 Key 粘贴到对话中;应指导其在真正提交任务的进程或宿主安全环境设置中配置 `ATLASCLOUD_API_KEY`,必要时刷新或重启执行会话。如果 Key 已存在于宿主或父进程配置,但提交进程读不到,应报告“环境作用域不一致”,不能说用户没有 Key。
      
      ### 计费任务状态机
      
      所有图片和视频生成都必须遵守:
      
      1. 提交后立即记录预测 ID 和对应创作阶段。
      2. `starting`、`queued`、`pending`、`processing` 都是进行中状态;必须每 2 秒查询同一 ID,禁止为同一阶段再提交任务。
      3. `completed`、`succeeded` 是成功终态;必须下载并检查输出,才能开始依赖它的下一阶段。
      4. `failed`、`timeout`、`canceled` 是失败终态;新任务必须经过明确重试决定,并先报告原任务 ID 和可能增加的费用。
      5. 推理时间为 0 或缺失、输出延迟、本地轮询超时、对话中断、临时查询失败,都不能证明任务失败;必须保留 ID 并恢复轮询。
      6. 用户说“继续”只表示恢复原任务,不代表允许重试。Storyboard 仍在生成时,禁止提交依赖它的视频任务。
      
      本工作流的 2 秒查询周期适用于所有 Atlas 执行路线。`atlas-skill` 每 2 秒用同一 ID 重复执行查询预测结果的步骤;`atlas-mcp` 每 2 秒用同一 ID 调用一次 `atlas_get_prediction`。MCP 服务端每次工具调用只查询一次状态,循环由智能体负责。内置 REST 和 CLI 适配器则在代码中实现 2 秒轮询。查询状态是只读操作,任何情况下都不能用新的生成调用代替。
      
      脚本恢复任务时,把原 ID 写入 `execution.resumePredictionIds.<阶段>`。阶段键包括 `grid`、`ref1`、`ref2`、`seg1`、`shot1`、`clip1` 等。不能因为上一轮轮询进程结束就创建替代任务。
      
      `scripts/generate.mjs` 不能调用智能体 Skill 或 MCP;它是单独的批处理工具。未设置时默认 `atlas-rest`,只有显式设置 `execution.adapter: "atlas-cli"` 才会使用 CLI,绝不因为本机安装了 CLI 就自动切换。
      
      ```bash
      GRID_ONLY=1 node scripts/generate.mjs scripts/myjob.json
      CLIPS_MAX=1 node scripts/generate.mjs scripts/myjob.json
      SEGS_MAX=1 node scripts/generate.mjs scripts/myjob.json
      ```
      
      脚本输出 `[storyboard-preview]` 绝对路径时,必须把该图展示给用户并自行检查,再继续视频任务。详细执行规则见 [中文执行适配层](execution-adapters.zh-CN.md)。
      
  • scripts
    • providers
      • atlas-cli.mjs 8.3 KB · in bundle
      • atlas-rest.mjs 6.6 KB · in bundle
      • index.mjs 1.6 KB · in bundle
    • config.chain.en.example.json 1.7 KB
      {
        "_comment": "Two 15-second continuous-action segments. The second segment starts from the real generated tail frame. Inspect the seam manually.",
        "mode": "chain",
        "outDir": "build",
        "modelProfile": "seedance-default",
        "execution": { "adapter": "atlas-rest", "apiKeyEnv": "ATLASCLOUD_API_KEY", "verifyModels": true },
        "grid": {
          "model": "bytedance/seedream-v5.0-pro/text-to-image",
          "size": "2048*1152",
          "rows": 1,
          "cols": 3,
          "prompt": "Create a 1-row, 3-panel visual storyboard for tomato-and-egg stir-fry. Warm hand-drawn food-animation style, cozy yellow kitchen, black iron wok, one consistent pair of clean hands only, no face. Panel 1: beaten egg beside diced tomato. Panel 2: tomato releases red juice in the wok. Panel 3: glossy finished tomato and egg on a white plate with scallion. Thin dividers, no unrelated text."
        },
        "video": { "model": "bytedance/seedance-2.0/image-to-video", "resolution": "1080p", "ratio": "9:16", "duration": 15, "generate_audio": true, "watermark": false },
        "styleSuffix": "Warm hand-drawn food-animation style; one consistent pair of hands, black iron wok, and cozy kitchen; soft warm light; natural continuous motion.",
        "segments": [
          { "first": 1, "last": 2, "duration": 15, "prompt": "A hand pours beaten egg into hot oil; it puffs and is gently separated into soft yellow curds, then removed. Tomato enters the same wok and releases bubbling red juice. Slow push-in." },
          { "chain": true, "last": 3, "duration": 15, "prompt": "Continue the same wok and action from the prior final frame. Return the egg to the tomatoes, toss until glossy, plate the dish, scatter scallion, and let steam rise. Slow push-in to the final composition." }
        ],
        "seamTrimFrames": 1
      }
      
    • config.chain.zh-CN.example.json 2.3 KB
      {
        "_comment": "chain(两段15秒连续动作)示例 · 番茄炒蛋。前段用 last_image 设计终点;后段从上一段真实尾帧续起。生成后必须人工检查衔接。rows*cols=关键帧数;segments 引用关键帧序号(1基)。",
        "mode": "chain",
        "outDir": "build",
        "modelProfile": "seedance-default",
        "execution": { "adapter": "atlas-rest", "apiKeyEnv": "ATLASCLOUD_API_KEY", "verifyModels": true },
      
        "grid": {
          "model": "bytedance/seedream-v5.0-pro/text-to-image",
          "size": "2048*1152",
          "rows": 1,
          "cols": 3,
          "prompt": "生成一张1行3列共3格的漫画分镜网格图,横向并排,吉卜力食物番手绘动画风格,暖黄色调的治愈厨房、黑铁锅,手绘颗粒质感、柔和光影。每格只有一双干净的手和食物、没有人脸、没有多余的手,格间细黑边。【任何格子都不要出现数字编号或文字】。从左到右是番茄炒蛋的三个关键节点:1一只手刚打好一碗金黄蛋液,旁边案板上是切好的红番茄块(起手备料);2番茄块在黑铁锅里煸炒、渐渐煸出红色汤汁咕嘟冒泡(中段);3成品番茄炒蛋盛入白瓷盘、红黄相间油亮、撒上翠绿葱花、热气升腾(收尾)。三格里这双手、黑铁锅与厨房完全一致;不要无关文字或角标。"
        },
      
        "video": {
          "model": "bytedance/seedance-2.0/image-to-video",
          "resolution": "1080p",
          "ratio": "9:16",
          "duration": 15,
          "generate_audio": true,
          "watermark": false
        },
      
        "styleSuffix": "吉卜力食物番手绘动画风格,竖屏,暖黄治愈厨房,只有一双干净的手和食物、没有人脸与多余的手,手绘颗粒质感、柔和暖光,动作低缓连贯。",
      
        "segments": [
          { "first": 1, "last": 2, "duration": 15, "prompt": "一只手把金黄蛋液缓缓倒入热油锅,蛋液迅速膨起、边缘冒泡,用锅铲轻轻划散成嫩黄蛋块并盛出;随后红番茄块下入同一口黑铁锅翻炒,渐渐煸出红色汤汁、开始咕嘟冒泡。镜头缓慢推近。" },
          { "chain": true, "last": 3, "duration": 15, "prompt": "嫩黄的炒蛋倒回番茄锅中,一只手颠勺让蛋块与番茄红汁快速翻炒裹匀、汤汁收浓油亮;最后盛入白瓷盘、撒上翠绿葱花,热气袅袅升起。镜头缓慢推近定格。" }
        ],
      
        "seamTrimFrames": 1
      }
      
    • config.grid.en.example.json 2 KB
      {
        "_comment": "Independent I2V grid example. Each panel becomes one clip for a hard-cut montage. rows*cols must equal shots.length.",
        "outDir": "build",
        "modelProfile": "seedance-default",
        "execution": { "adapter": "atlas-rest", "apiKeyEnv": "ATLASCLOUD_API_KEY", "verifyModels": true },
        "ffmpeg": "ffmpeg",
        "ffprobe": "ffprobe",
        "grid": {
          "model": "bytedance/seedream-v5.0-pro/text-to-image",
          "size": "1728*2304",
          "rows": 3,
          "cols": 2,
          "prompt": "Create a 3-row by 2-column storyboard grid for six independent cooking shots. Warm hand-drawn food-animation style, late-night kitchen, wood counter, black iron wok, soft warm practical light. One consistent pair of clean hands only; no face and no extra hands. Thin dark dividers. Show, in reading order: ingredients on board; pork searing; aromatics blooming; green peppers entering; stir-fry tossing; finished dish plated. Keep the hands, wok, kitchen, lighting, and style consistent. No unrelated text or corner marks."
        },
        "video": { "model": "bytedance/seedance-2.0/image-to-video", "resolution": "1080p", "ratio": "9:16", "duration": 5, "generate_audio": true, "watermark": false },
        "styleSuffix": "Warm hand-drawn food-animation style, vertical frame, one consistent pair of hands and black iron wok, no face or extra hands, natural motion, one clear camera intention per shot.",
        "shots": [
          "Slowly push in on sliced pork, green pepper, garlic, and fermented beans arranged on a board; one hand places the final pork slice.",
          "Pork sizzles in a hot black iron wok while one hand stirs; edges curl and turn golden.",
          "Garlic, ginger, and fermented beans bloom in the rendered pork fat; one hand gently tosses the wok.",
          "Green pepper enters the wok and is stir-fried over high heat until its skin begins to blister.",
          "Pork and pepper toss together with soy sauce; steam rises and the food becomes glossy.",
          "The finished dish is plated; one hand scatters scallion and the camera slowly pushes in."
        ]
      }
      
    • config.grid.zh-CN.example.json 2.8 KB
      {
        "_comment": "辣椒炒肉示例 · 独立镜头网格路线。每格生成一个独立 I2V 片段,适合硬切蒙太奇。rows*cols 必须等于 shots 数量。",
        "outDir": "build",
        "modelProfile": "seedance-default",
        "execution": {
          "adapter": "atlas-rest",
          "apiKeyEnv": "ATLASCLOUD_API_KEY",
          "verifyModels": true
        },
        "ffmpeg": "ffmpeg",
        "ffprobe": "ffprobe",
      
        "grid": {
          "model": "bytedance/seedream-v5.0-pro/text-to-image",
          "size": "1728*2304",
          "rows": 3,
          "cols": 2,
          "prompt": "生成一张3行2列共6格的漫画分镜网格图,吉卜力食物番手绘动画风格,暖黄色调的深夜治愈厨房,木质台面与黑铁锅,手绘颗粒质感、柔和光影。每格只有一双干净整洁的手和食物、没有人脸、没有多余的手。格与格之间是细黑边。【重要:任何格子里都不要出现数字编号或文字】。按从左到右、从上到下依次表现辣椒炒肉的真实炒制步骤:1木案板上摆好切片五花肉、斜切青辣椒段、蒜片姜片和黑豆豉(备料俯拍);2五花肉片在烧热的黑铁锅里煸炒出油、边缘金黄微卷;3下蒜片姜片和黑豆豉在肉油里爆香冒烟;4倒入翠绿青辣椒段大火煸炒、辣椒起虎皮;5淋酱油把肉和青椒大火翻炒混合、锅气腾起;6成品辣椒炒肉盛入白瓷盘、油亮红绿、热气升腾。六格里这双手、黑铁锅与厨房完全一致;不要无关文字或角标。"
        },
      
        "video": {
          "model": "bytedance/seedance-2.0/image-to-video",
          "resolution": "1080p",
          "ratio": "9:16",
          "duration": 5,
          "generate_audio": true,
          "watermark": false
        },
      
        "styleSuffix": "吉卜力食物番手绘动画风格,竖屏,暖黄治愈厨房,只有一双干净的手和食物、没有人脸与多余的手,手绘颗粒质感、柔和暖光,镜头运动平缓,以一种主导运镜为主;需要复合时保持一个同步意图。",
      
        "shots": [
          "镜头缓慢推近木案板上码好的五花肉片、青辣椒段、蒜片与黑豆豉,一只手轻轻把最后一片肉摆正,热气微微浮动。",
          "五花肉片在烧热的黑铁锅里滋滋翻动,一只手持锅铲煸炒,肥肉渗出油脂、边缘变得金黄微卷,油光闪动、白色蒸汽升腾。",
          "锅中下入蒜片、姜片和黑豆豉,在金黄的肉油里翻炒爆香,香气与蒸汽升腾,一只手轻轻颠动铁锅。",
          "翠绿的青辣椒段倒入锅中大火煸炒,一只手颠勺翻动,辣椒表皮渐渐起虎皮皱纹,锅气腾起。",
          "淋入酱油,五花肉与青椒在大火中快速翻炒混合,酱汁裹满食材油亮,锅边腾起镬气与一丝火光。",
          "一盘热气腾腾的辣椒炒肉盛入白瓷盘,油亮的红绿相间,一只手撒上翠绿葱花,蒸汽袅袅升起,镜头缓慢推近定格。"
        ]
      }
      
    • config.reference.en.example.json 2.4 KB
      {
        "_comment": "Role-specific R2V reference example. Each image has one job: hands, kitchen, style, finished dish, or opening state.",
        "mode": "reference",
        "outDir": "build",
        "modelProfile": "seedance-default",
        "execution": { "adapter": "atlas-rest", "apiKeyEnv": "ATLASCLOUD_API_KEY", "verifyModels": true },
        "refs": [
          { "prompt": "Hand-drawn food-animation style: one clean pair of hands with rolled shirt sleeves on a warm plain background, no face. Identity reference for the hands.", "size": "1152*2048" },
          { "prompt": "Hand-drawn food-animation style: empty late-night kitchen, wood counter, black iron wok, warm hanging lamp, night beyond the window, no people. Setting reference.", "size": "1152*2048" },
          { "prompt": "Hand-drawn food-animation style palette: warm yellow, soft lighting, subtle paper grain. Style reference.", "size": "1152*2048" },
          { "prompt": "Hand-drawn food-animation style: finished tomato-and-egg stir-fry on a white plate, glossy red and yellow, scallion and rising steam. Product reference.", "size": "1152*2048" },
          { "prompt": "Vertical opening frame in a warm kitchen: one hand holds beaten egg beside diced tomatoes, no face. Start-frame reference.", "size": "1152*2048" }
        ],
        "audio": { "_comment": "Optional reference audio. A rejected input audio reference stops the run and is not silently removed.", "outputRetries": 2, "references": [] },
        "video": { "refModel": "bytedance/seedance-2.0/reference-to-video", "resolution": "1080p", "ratio": "9:16", "duration": 15, "generate_audio": true, "watermark": false },
        "styleSuffix": "Warm hand-drawn food-animation style, vertical frame, one pair of hands, one black iron wok, no face or extra hands, soft warm light, one clear camera intention per shot.",
        "segments": [
          { "refs": [1,2,3,4,5], "duration": 15, "prompt": "Use Image 1 for the hands, Image 2 for the kitchen, and Image 3 for style. Start from Image 5: pour egg into the hot wok, form soft curds, remove them, then stir-fry tomato until red juice bubbles. Slow push-in; end on the tomato sauce state." },
          { "refs": [1,2,3,4,5], "chainPrevAsRef": true, "duration": 15, "prompt": "Continue from Image 6, the final frame of the prior segment, with the same hands, kitchen, and style: return egg to the tomato sauce, toss until glossy, plate it, add scallion, and let steam rise. Use Image 4 for the finished-dish reference. Slow pull-out to the final composition." }
        ]
      }
      
    • config.reference.zh-CN.example.json 2.8 KB
      {
        "_comment": "reference 模式示例 · 番茄炒蛋。refs[] 中每张图只承担一种角色(手、场景、风格、成品、开场)。提示词里的图片N按每段 reference_images 顺序对应。第二段仅在同一连续动作时才用上一段尾帧作附加参考。",
        "mode": "reference",
        "outDir": "build",
        "modelProfile": "seedance-default",
        "execution": { "adapter": "atlas-rest", "apiKeyEnv": "ATLASCLOUD_API_KEY", "verifyModels": true },
      
        "refs": [
          { "prompt": "吉卜力食物番手绘风,一双干净的手、卷起衬衫袖口,暖色纯背景,无人脸。角色(手)一致性参考。", "size": "1152*2048" },
          { "prompt": "吉卜力食物番手绘风,暖黄深夜治愈厨房空镜:木台面、黑铁锅、暖黄吊灯、窗外夜色,无人物。场景参考。", "size": "1152*2048" },
          { "prompt": "吉卜力食物番手绘风调性色板,暖黄主调、柔和光影、纸纹颗粒质感。整体风格参考。", "size": "1152*2048" },
          { "prompt": "吉卜力食物番手绘风,白瓷盘装好的番茄炒蛋,红黄相间油亮撒葱花热气升腾。成品参考。", "size": "1152*2048" },
          { "prompt": "吉卜力食物番手绘风竖屏开场,一只手端金黄蛋液碗、案板上切好的红番茄块,暖黄厨房无人脸。首帧参考。", "size": "1152*2048" }
        ],
      
        "audio": {
          "_comment": "可选参考音频。file 会先上传到 Atlas Cloud;若上传或模型校验失败会停止并提示,不会自动移除。空数组表示只由 Seedance 原生生成声音。audioRefs 使用 1 基编号选择某段使用的音频。",
          "outputRetries": 2,
          "references": []
        },
      
        "video": {
          "refModel": "bytedance/seedance-2.0/reference-to-video",
          "resolution": "1080p",
          "ratio": "9:16",
          "duration": 15,
          "generate_audio": true,
          "watermark": false
        },
      
        "styleSuffix": "吉卜力食物番手绘动画风格,竖屏,暖黄治愈厨房,只有一双干净的手和食物、没有人脸与多余的手,手绘颗粒质感、柔和暖光,每个镜头以一种主导运镜为主;需要复合时保持一个同步意图。",
      
        "segments": [
          { "refs": [1,2,3,4,5], "duration": 15, "prompt": "以图片1的手、图片2的厨房、图片3的风格为基准,从图片5开场:手把金黄蛋液倒入热油锅膨起冒泡、划散成嫩黄蛋块盛出;红番茄块下锅翻炒煸出红汁咕嘟冒泡。镜头缓慢推近,收尾定格在番茄出汁。" },
          { "refs": [1,2,3,4,5], "chainPrevAsRef": true, "duration": 15, "prompt": "以图片1手、图片2厨房、图片3风格为基准,无缝承接图片6(上一段最后画面)番茄出汁状态:炒蛋倒回锅中颠勺翻炒裹匀收汁油亮;盛入白瓷盘撒葱花热气升腾,镜头缓慢拉远定格。成品参考图片4。" }
        ]
      }
      
    • config.shot-pairs.en.example.json 1.8 KB
      {
        "_comment": "Two independent I2V shots. Panels 1/2 are the first/last frame of Shot 1; panels 3/4 are the first/last frame of Shot 2. The shots use a hard cut and do not inherit a tail frame.",
        "mode": "shot-pairs",
        "outDir": "build",
        "modelProfile": "seedance-default",
        "execution": { "adapter": "atlas-rest", "apiKeyEnv": "ATLASCLOUD_API_KEY", "verifyModels": true },
        "grid": { "model": "bytedance/seedream-v5.0-pro/text-to-image", "size": "2048*1152", "rows": 1, "cols": 4, "prompt": "Create a clean 1-row, 4-panel visual storyboard. Warm hand-drawn food-animation style, late-night kitchen, black iron wok, wood counter, one consistent pair of clean hands, no face. Panel 1: ingredients on board. Panel 2: pork searing in the wok. Panel 3: green peppers entering the wok. Panel 4: finished pork and peppers on a white plate. Keep all visual invariants consistent. Thin dividers only; no numbers, captions, bubbles, logos, or text." },
        "video": { "model": "bytedance/seedance-2.0/image-to-video", "resolution": "1080p", "ratio": "9:16", "duration": 5, "generate_audio": true, "watermark": false },
        "styleSuffix": "Warm hand-drawn food-animation style, vertical frame, one pair of hands and one black iron wok, no face or extra hands, natural motion, one clear camera intention per shot.",
        "segments": [
          { "first": 1, "last": 2, "duration": 5, "prompt": "Start from the ingredients in Image 1. A hand slides pork into the hot wok; it sizzles, renders fat, and turns golden at the edges. Slowly move from the top view into Image 2's medium-close composition." },
          { "first": 3, "last": 4, "duration": 5, "prompt": "Start from Image 3 as green pepper enters the wok. A hand gently stirs until the pepper blisters and the dish is plated. Smooth lateral move into Image 4's final composition." }
        ]
      }
      
    • config.shot-pairs.zh-CN.example.json 2.1 KB
      {
        "_comment": "shot-pairs 示例 · 两个独立 I2V 镜头。4 格分别是镜头1首/尾、镜头2首/尾;两镜头硬切,不延续上一段尾帧。",
        "mode": "shot-pairs",
        "outDir": "build",
        "modelProfile": "seedance-default",
        "execution": { "adapter": "atlas-rest", "apiKeyEnv": "ATLASCLOUD_API_KEY", "verifyModels": true },
      
        "grid": {
          "model": "bytedance/seedream-v5.0-pro/text-to-image",
          "size": "2048*1152",
          "rows": 1,
          "cols": 4,
          "prompt": "生成一张1行4列的无文字视觉分镜网格图。吉卜力食物番手绘动画风格,暖黄深夜厨房,黑铁锅、木台面、同一双干净的手,无人脸。第1格:俯拍,案板上摆好五花肉片和青辣椒;第2格:中近景,五花肉在黑铁锅中煸炒至金黄卷边;第3格:俯拍,翠绿青椒段刚倒入冒烟的锅中;第4格:中近景,辣椒炒肉盛入白瓷盘、热气升腾。四格保持同一双手、黑铁锅、厨房、暖黄柔光和手绘颗粒。只有细黑格线;不要数字、标题、字幕、气泡、logo或任何文字。"
        },
      
        "video": {
          "model": "bytedance/seedance-2.0/image-to-video",
          "resolution": "1080p",
          "ratio": "9:16",
          "duration": 5,
          "generate_audio": true,
          "watermark": false
        },
      
        "styleSuffix": "吉卜力食物番手绘动画风格,竖屏,暖黄治愈厨房,同一双干净的手和同一口黑铁锅,无人脸、无多余的手,动作低缓自然,每个镜头以一种主导运镜为主;需要复合时保持一个同步意图。",
      
        "segments": [
          {
            "first": 1,
            "last": 2,
            "duration": 5,
            "prompt": "从图片1的备料画面开始,一只手把五花肉片依次滑入烧热的黑铁锅,肉片滋滋作响、渗出油脂、边缘逐渐金黄微卷。镜头从俯拍自然收束到图片2的中近景构图,缓慢推近。"
          },
          {
            "first": 3,
            "last": 4,
            "duration": 5,
            "prompt": "从图片3的青椒下锅画面开始,一只手用锅铲轻轻翻炒,青椒表皮渐渐起虎皮,肉片与青椒油亮混合;随后盛入白瓷盘。镜头平稳横移,收束到图片4的成品构图。"
          }
        ]
      }
      
    • config.storyboard-reference.en.example.json 1.1 KB
      {
        "_comment": "Whole-storyboard R2V example. Place a readable board at assets/storyboard.png and submit it as Image 1 in one request. Do not split it. Make a clean copy only if a test video renders a number, divider, or note.",
        "mode": "reference",
        "outDir": "build",
        "modelProfile": "seedance-default",
        "execution": { "adapter": "atlas-rest", "apiKeyEnv": "ATLASCLOUD_API_KEY", "verifyModels": true },
        "refs": [{ "file": "../assets/storyboard.png" }],
        "video": { "refModel": "bytedance/seedance-2.0/reference-to-video", "resolution": "1080p", "ratio": "9:16", "duration": 15, "generate_audio": true, "watermark": false },
        "segments": [{ "refs": [1], "duration": 15, "prompt": "Follow the storyboard in Image 1 in left-to-right, top-to-bottom order. Generate one complete shot at a time. Naturally connect the subject, composition, action, and camera language in each panel. Keep the character, wardrobe, setting, lighting, and art direction consistent. Use Image 1 dividers, numbers, arrows, and notes only to understand the sequence and shot intention; do not render them into the final video." }]
      }
      
    • config.storyboard-reference.zh-CN.example.json 1.1 KB
      {
        "_comment": "整张分镜图 R2V 示例。只要每格画面仍清晰可读,就将分镜图放到 skill 根目录的 assets/storyboard.png,整张作为图片1一次提交;不切格。只有实测生成把编号、格线或注释带进视频时,才另做清理版。",
        "mode": "reference",
        "outDir": "build",
        "modelProfile": "seedance-default",
        "execution": { "adapter": "atlas-rest", "apiKeyEnv": "ATLASCLOUD_API_KEY", "verifyModels": true },
      
        "refs": [
          {
            "file": "../assets/storyboard.png"
          }
        ],
      
        "video": {
          "refModel": "bytedance/seedance-2.0/reference-to-video",
          "resolution": "1080p",
          "ratio": "9:16",
          "duration": 15,
          "generate_audio": true,
          "watermark": false
        },
      
        "segments": [
          {
            "refs": [1],
            "duration": 15,
            "prompt": "严格按图片1中的分镜图从左到右、从上到下的顺序生成短片,一次只出现一个完整镜头。每个镜头依照分镜中的主体、构图、动作与镜头语言自然衔接;保持人物、服装、场景、光影和画风一致。"
          }
        ]
      }
      
    • config.t2v-chain.en.example.json 1.5 KB
      {
        "_comment": "Text-to-video chain without reference images. The second segment continues from the real tail frame of the first. Inspect the seam manually; it is not guaranteed invisible.",
        "mode": "chain",
        "outDir": "build",
        "modelProfile": "seedance-default",
        "execution": { "adapter": "atlas-rest", "apiKeyEnv": "ATLASCLOUD_API_KEY", "verifyModels": true },
        "video": { "t2vModel": "bytedance/seedance-2.0/text-to-video", "model": "bytedance/seedance-2.0/image-to-video", "resolution": "1080p", "ratio": "9:16", "duration": 15, "generate_audio": true, "watermark": false },
        "styleSuffix": "Warm hand-drawn food-animation style, vertical frame, a late-night kitchen, black iron wok and stove flame, one pair of clean hands and food only, no face or extra hands, soft warm light, one clear camera intention per shot.",
        "segments": [
          { "t2v": true, "duration": 15, "prompt": "First half of a beef stir-fry. Open on an extreme close-up of thin beef hitting smoking oil and curling at the edges. One hand stirs over high heat; slowly push in. Move the beef aside, bloom garlic, ginger, fermented beans, and sliced chili. End with the aromatics smoking in the wok." },
          { "chain": true, "duration": 15, "prompt": "Continue from the prior final frame with the same wok and hands. Add green pepper and stir-fry until blistered; return the beef and toss over high heat; add sauce and scallion; plate the glossy dish as steam rises. Slowly pull out to the final composition." }
        ],
        "seamTrimFrames": 1
      }
      
    • config.t2v-chain.zh-CN.example.json 1.6 KB
      {
        "_comment": "t2v 接龙(无参考图)示例 · 小炒黄牛肉。首段文生视频取尾帧,后续段从该尾帧继续同一动作;衔接需人工审片,不是保证无缝。",
        "mode": "chain",
        "outDir": "build",
        "modelProfile": "seedance-default",
        "execution": { "adapter": "atlas-rest", "apiKeyEnv": "ATLASCLOUD_API_KEY", "verifyModels": true },
      
        "video": {
          "t2vModel": "bytedance/seedance-2.0/text-to-video",
          "model": "bytedance/seedance-2.0/image-to-video",
          "resolution": "1080p",
          "ratio": "9:16",
          "duration": 15,
          "generate_audio": true,
          "watermark": false
        },
      
        "styleSuffix": "吉卜力食物番手绘动画风格,竖屏,暖黄色调的深夜治愈厨房、黑铁锅与灶火,只有一双干净的手和食物、没有人脸与多余的手,手绘颗粒质感、柔和暖光,每个镜头以一种主导运镜为主;需要复合时保持一个同步意图。",
      
        "segments": [
          { "t2v": true, "duration": 15, "prompt": "小炒黄牛肉前半段。开场极近特写:黄牛肉薄片滑入冒烟热油刺啦变色卷边,形成强钩子;一只手锅铲大火划炒牛肉,镜头缓慢推近,灶火映亮锅沿;牛肉拨到锅边,下蒜片姜片黑豆豉小米椒圈爆香,蒸汽升腾。收尾定格在锅中爆香冒烟。" },
          { "chain": true, "duration": 15, "prompt": "无缝承接上一段:同一口黑铁锅同一双手,从爆香冒烟继续。下青椒段大火煸炒起虎皮;牛肉回锅合炒,颠勺锅气窜火光(money shot);淋生抽蚝油翻炒油亮,撒蒜苗;盛入白瓷盘红绿油亮热气升腾,镜头缓慢拉远定格。" }
        ],
      
        "seamTrimFrames": 1
      }
      
    • generate.mjs 21.2 KB · in bundle
    • split_grid.py 2.5 KB
      #!/usr/bin/env python3
      """Split an N x M storyboard grid image into individual numbered frames.
      
      The image models in this workflow can return the whole storyboard as one grid
      image (e.g. a 4x4 of 16 shots). This script is only for the intentional
      individual-I2V / shot-pair route. The default R2V storyboard route uploads the
      whole board directly and does not need cropping. When cropping is wanted, this
      does it deterministically in reading order (left-to-right, top-to-bottom):
      frame_01, frame_02, ...
      
      Usage:
          python split_grid.py grid.png --rows 4 --cols 4 --out frames/
          python split_grid.py grid.png --rows 4 --cols 4 --border 6   # trim gutters
      
      Requires Pillow:  pip install pillow
      """
      import argparse
      import os
      import sys
      
      
      def split_grid(image_path, rows, cols, out_dir, border=0):
          try:
              from PIL import Image
          except ImportError:
              sys.exit("Pillow is required. Install with:  pip install pillow")
      
          img = Image.open(image_path)
          w, h = img.size
          cell_w = w // cols
          cell_h = h // rows
      
          os.makedirs(out_dir, exist_ok=True)
          n = 0
          for r in range(rows):
              for c in range(cols):
                  left = c * cell_w + border
                  upper = r * cell_h + border
                  right = (c + 1) * cell_w - border
                  lower = (r + 1) * cell_h - border
                  if right <= left or lower <= upper:
                      sys.exit("border too large for this cell size; reduce --border")
                  frame = img.crop((left, upper, right, lower))
                  n += 1
                  path = os.path.join(out_dir, f"frame_{n:02d}.png")
                  frame.save(path)
                  print(f"wrote {path}  ({right - left}x{lower - upper})")
          print(f"\nDone: {n} frames from a {rows}x{cols} grid -> {out_dir}/")
          print("Next: inspect each frame as a single-shot still, then animate one at a "
                "time in Seedance only when you intentionally chose the I2V route.")
      
      
      def main():
          p = argparse.ArgumentParser(description="Split a storyboard grid into frames.")
          p.add_argument("image", help="path to the grid image")
          p.add_argument("--rows", type=int, required=True, help="number of rows")
          p.add_argument("--cols", type=int, required=True, help="number of columns")
          p.add_argument("--out", default="frames", help="output directory (default: frames/)")
          p.add_argument("--border", type=int, default=0,
                         help="pixels to trim off each cell edge to drop gutters/borders")
          args = p.parse_args()
          split_grid(args.image, args.rows, args.cols, args.out, args.border)
      
      
      if __name__ == "__main__":
          main()
      
  • SKILL.md 16.7 KB
    ---
    name: seedance-2-5-skill
    description: >-
      Plan and generate controllable Seedance video using Seedream 5.0 Pro
      storyboards and Seedance 2.0 today, with a Seedance 2.5 route when available.
      Use for consistent people, products, objects, food, or scenes; storyboard-to-
      video; reference-to-video; first-and-last-frame image-to-video; extensions;
      and Atlas Cloud media generation.
    ---
    
    # Seedance 2.5 Skill
    
    ## Language route
    
    - For an English request, follow this file and the `*.md` references.
    - For a Chinese request, read [the Chinese workflow](references/workflow.zh-CN.md)
      first, then use the matching `*.zh-CN.md` reference files.
    - Keep model IDs, JSON keys, commands, media placeholders, and native-audio
      symbols exactly as code. Do not translate them.
    
    Do not force every request through one image-grid pipeline. Choose the video
    route first, then enable only the preparation modules the job needs.
    
    ## 1. Choose the creative route
    
    | Need | Route | Inputs | Result |
    |---|---|---|---|
    | One short, simple scene | T2V | Text prompt | One self-contained shot |
    | A multi-shot sequence with readable panels | R2V storyboard | One complete storyboard image | One request turns panel order into a continuous video |
    | A clip controlled by people, product, scene, or style assets | R2V asset references | A small role-specific asset pack | One clip built from those references |
    | A shot with an exact beginning and ending | I2V shot pair | Start and end keyframes | One independently reviewable shot |
    | An uninterrupted action exceeding the supported duration | Extend / chain | Previous generated tail frame and next prompt | Continuation of the same shot |
    | One continuous piece carrying several events | Staged whole-short | Text plus optional references | One request covering ordered stages, each landing on a stated end state |
    | A scoped change to an existing video | Editing | Source video plus target references | The source with only the named region or element changed |
    | A bridge between two finished clips | Seamless transition | Two videos | Generated bridge content between them |
    | Motion, blocking, and camera taken from a 3D preview | Blockout reference | Coarse or fine blockout video plus look references | Final render following the blockout's timing and staging |
    
    Use Seedance 2.0 as the executable default. Offer a Seedance 2.5 whole-short
    route only when the selected provider exposes that model and its actual limits.
    
    The last four routes depend on capabilities that differ per model and per
    provider. Verify availability before offering one; see
    [capabilities](references/capabilities.md) for the 2.0/2.5 comparison and for
    which published capabilities are platform features rather than API parameters.
    
    ### Storyboard or individual keyframes
    
    - Use the **whole storyboard image in one R2V request** by default whenever
      individual panels remain visually readable. Do not crop it first. Seedance
      can interpret the ordered panels as one continuous multi-shot video.
    - Use **I2V shot pairs** only when independent reshoots or precise start/end
      states matter more than the transition quality of one R2V generation.
    - Do not upload every storyboard cell as a default R2V asset pack. Use several
      R2V images only when each has a distinct role, such as subject, product,
      setting, style, or motion reference.
    
    ## 2. Enable only needed preparation
    
    ### Subject brief
    
    Use this only for a person, product, prop, hand, vehicle, or scene that must
    recur. Record 3–5 invariants: silhouette or proportion, signature material or
    wardrobe, key colour, and any must-preserve marking. Skip it for a one-off
    atmosphere shot.
    
    For a recurring person, make a clean face close-up and a separate full-body
    reference. Do not use a front/side/back composite as the identity input; it can
    be interpreted as multiple people. Multi-angle product references remain useful
    when the object itself must be shown from several sides.
    
    ### Keyframes
    
    Create a start keyframe for every I2V shot. Add an end keyframe only when the
    shot must land on a specific action, composition, product pose, or hand
    position. Use one clean scene per keyframe.
    
    ### Storyboard
    
    Use an existing storyboard directly as the R2V reference. Seedance normally
    understands panel order and does not require panel numbers, dividers, arrows, or
    notes to be removed in advance. Make a clean copy only after a test generation
    actually renders an unwanted divider, number, caption, or multi-panel layout.
    
    When a multi-shot request has no storyboard, use the Seedream template in
    [prompt templates](references/prompt-templates.md) to create one board.
    **Display that board in the host UI, inspect it yourself, then continue to
    Seedance R2V when it passes review.** Showing the board is a progress update,
    not a user-approval gate.
    
    For a supplied storyboard, display it unless it is already visible in the
    conversation. Check planned order, readable key beats, recurring-subject
    consistency, and content-specific constraints such as anatomy, product form, or
    critical text. Refine a visibly failed board before video generation. Ask the
    user only when a creative choice cannot be inferred.
    
    If the route deliberately changes to I2V shot pairs, crop panels only then.
    Inspect the layout first: automatic crops can verify position, not whether the
    image model drew the intended panel layout.
    
    ## 3. Design continuity and cuts
    
    First-and-last frames control **one shot**; they do not mean every shot must
    inherit the prior clip's tail.
    
    | Transition | Use prior tail as next start? | Design rule |
    |---|---:|---|
    | One uninterrupted action | Yes | Generate in sequence and inspect the seam |
    | Hard cut to a new angle, place, product, or time | No | Design each shot independently |
    | Match cut | Usually no | Match movement direction, shape, colour, or composition |
    | Occlusion or whip transition | No | End with the occluding action; start the next shot inside or after it |
    | Insert or cutaway | No | Use an object, environment, or product detail as a bridge |
    
    Put important cuts, matches, and occlusions in the storyboard and prompt. Do
    not rely on a cross-dissolve to repair unrelated shots.
    
    ## 4. Write the video prompt
    
    ### Scope: put each instruction where it applies
    
    Before writing blocks, sort what you know by **what it governs**. Instructions in
    the wrong place are the most common cause of drift — a global rule written inside
    beat 1 stops applying at beat 4.
    
    | Scope | Governs | Contents |
    |---|---|---|
    | Global | The whole piece | Film type, scene, style, one-sentence premise, camera principle |
    | Locks | Anything that must not drift | Identity, reference roles, audio source, continuity, negatives |
    | Time | One beat or stage | Stage events and their end states |
    
    Restate the two or three most expensive locks at the **physical end** of the
    prompt; recency helps. That is a placement convention, not a fourth scope — the
    content still belongs to Locks and appears there first.
    
    This mirrors the model-agnostic spec format in
    [the Universal Video Prompt Skill](../universal-video-prompt-skill/SKILL.md). Use
    that skill when one brief has to run on more than one model; use this file for
    Seedance-specific writing. This is a required companion for a complete Seedance 2.5
    Skill setup. If the link does not resolve, help the user install
    `universal-video-prompt-skill` before continuing; do not invent the missing shared
    specification.
    
    ### Blocks
    
    Use only the blocks that affect the shot:
    
    ```text
    [subject/reference binding]
    + [one observable action]
    + [space and important object relationships]
    + [one primary camera move, coherent composite move, or a cut]
    + [light/style when it matters]
    + [audio or dialogue when enabled]
    + [end state, for any stage that must land somewhere specific]
    + [must-preserve constraints]
    ```
    
    ### End states carry multi-event work
    
    For anything with more than one event, state what is **visibly true** when each
    stage ends. This is the highest-leverage single addition to a multi-stage prompt:
    it converts "keep it consistent" into something the model can target and you can
    check.
    
    ```text
    weak:   the two of them keep working on the bouquet
    strong: end state: the florist holds the bouquet in the left hand;
            the scissors are back on the right side of the bench
    ```
    
    An end state must be visible. "She feels relieved" is not one; "her shoulders
    drop and the frown clears" is. Read [long video](references/long-video.md) for the
    staged structure in full.
    
    ### Time granularity: decide before writing beats
    
    Granularity is a prior decision. Writing beats at second precision and then
    downgrading means rewriting them.
    
    | Granularity | Use when |
    |---|---|
    | None — event order only | One continuous action, mood pieces, single shots. Timestamps here fragment the shot |
    | **Stages + end states** | Most narrative work. **Default** |
    | Second-level | Only under an external hard constraint: music, lip sync, reference handoff, a beat that must land at a fixed time |
    
    Infer it when the input settles it — a supplied music or voiceover track means
    second-level, a stated mood piece means none, an explicit fixed beat means
    second-level. When the request is a multi-event narrative with no external
    constraint, **ask, and recommend with a reason** rather than presenting a bare
    menu.
    
    Timestamps allocate a time budget; they are not frame-accurate edit points, and
    actions may land slightly before or after a boundary. Do not demand impossible
    density such as three distinct actions inside one second.
    
    - Name reference roles explicitly, for example `Image 1: person`, `Image 2:
      product`, `Image 3: kitchen setting`.
    - For multi-shot R2V, list `Shot 1`, `Shot 2`, and `Shot 3` in event order.
      Set duration in provider controls rather than forcing exact seconds in text.
    - Prefer one primary movement per shot. A composite movement is valid when its
      direction, relation to the subject, and speed express one synchronized intent.
    - When native audio is enabled, use `()` for music, `<>` for sound effects,
      `{}` for dialogue, and `【】` for on-screen captions.
    - State only constraints that are costly to redo.
    
    Read the reference matching the job:
    
    | File | Read it for |
    |---|---|
    | [prompt templates](references/prompt-templates.md) | Route-specific templates |
    | [prompt blocks](references/prompt-blocks.md) | Reusable camera, audio, constraint patterns |
    | [long video](references/long-video.md) | Staged structure, end states, timestamp rules |
    | [multi reference](references/multi-reference.md) | Binding many assets without confusing them |
    | [real person](references/real-person.md) | Believable human subjects, and when to omit the detail |
    | [transitions](references/transitions.md) | Which transitions to generate and which to edit |
    | [editing and extension](references/editing-and-extension.md) | Changing or continuing existing video |
    | [capabilities](references/capabilities.md) | 2.0 vs 2.5 limits; platform features vs API parameters |
    | [model profile](references/model-profile.md) | Measured per-model behaviour and compile notes |
    | [cinematography](references/cinematography.md) | Detailed visual decisions |
    | [troubleshooting](references/troubleshooting.md) | Fault-specific fixes |
    | [execution adapters](references/execution-adapters.md) | Runner and adapter configuration |
    
    ## 5. Generate, review, and finish
    
    1. For a generated storyboard, create only the still first, display it in the
       conversation, and inspect it before any video request.
    2. For a supplied storyboard, display the input unless it is already visible,
       then verify it fits the chosen route.
    3. If the board passes review, generate one representative video pass without
       waiting for approval. Otherwise refine or regenerate the board first.
    4. Review identity, locks, stage end states, composition, motion, seam, and audio
       in that order, and stop at the first failure — later checks are wasted effort
       on a wrong identity. Regenerate only the failed shot or segment.
    5. Generate chains in order because the next segment needs the real prior tail.
       Generate cut-based clips independently and edit the planned transition.
    
    The bundled script is a draft assembler, not a colour-grading or music-mixing
    system.
    
    ## Atlas execution layer
    
    Keep creative route selection independent from how a job is submitted. The
    default models are Seedream 5.0 Pro for stills and Seedance 2.0 for video;
    users may override them only after verifying route support.
    
    In an agent conversation, use the **Atlas Cloud Skill** as the default direct
    generation route. It can discover a model, upload local media, submit an image
    or video request, poll, and retrieve outputs. Report `Execution: atlas-skill`
    only when it actually submitted the generation.
    
    Use `atlas-mcp` only when the user explicitly selects MCP and its generation
    tools are exposed. Use `atlas-cli` only when the user explicitly selects a
    terminal, script, CI, or batch run. If the Atlas Cloud Skill is missing, help
    install `AtlasCloudAI/atlas-cloud-skills` before selecting a fallback.
    
    Before reporting that an Atlas Cloud API key is missing, check the credentials
    in the **selected execution process**. For the REST runner, check
    `ATLASCLOUD_API_KEY` first and `ATLAS_CLOUD_API_KEY` as a compatibility alias.
    Do not infer credential availability from a different provider, plugin, or
    process; each execution channel can have an independent credential scope.
    
    If neither key exists, direct the user to
    `https://www.atlascloud.ai/console/api-keys?utm_source=github&utm_campaign=awesome-seedance-2.5-prompts-skills`.
    Never ask them to paste the key in
    chat. Tell them to set `ATLASCLOUD_API_KEY` in the submitting process or the
    host's secure environment settings, then refresh or restart the execution
    session if needed. If the key exists in a parent or host configuration but is
    absent from the submitting process, report an environment-scope mismatch
    instead of saying the user has no key.
    
    ### Billable task state machine
    
    Apply these rules to every image and video generation:
    
    1. After submission, record the prediction ID and logical stage immediately.
    2. Treat `starting`, `queued`, `pending`, and `processing` as active. Poll the
       same ID every 2 seconds; never submit another task for that stage.
    3. Treat `completed` and `succeeded` as successful terminal states. Download and
       inspect the output before starting a dependent stage.
    4. Treat `failed`, `timeout`, and `canceled` as terminal failures. A new task
       requires an explicit retry decision; report the old ID and possible extra
       cost first.
    5. A zero or missing processing-time field, delayed output, a local polling
       timeout, a stopped turn, or a temporary status-query error is **not** proof
       of failure. Preserve the ID and resume polling.
    6. Interpret `continue` as “resume the existing task,” never as permission to
       retry. Do not submit video while its required storyboard is still active.
    
    The 2-second interval applies to every Atlas execution route in this workflow.
    For `atlas-skill`, repeat its prediction-result step with the same ID. For
    `atlas-mcp`, call `atlas_get_prediction` with the same ID every 2 seconds. The
    MCP server performs a single status lookup per tool call; the agent owns the
    loop. The bundled REST and CLI adapters enforce the interval in code. A status
    lookup is read-only and must never be replaced with another generation call.
    
    When the runner must resume, set `execution.resumePredictionIds.<stage>` to the
    existing ID. Supported stage keys include `grid`, `ref1`, `ref2`, `seg1`,
    `shot1`, and `clip1`. Never create a replacement merely because a prior polling
    process ended.
    
    `scripts/generate.mjs` cannot invoke an agent Skill or MCP server. It is a
    separate batch runner: it defaults to `atlas-rest`, and can use `atlas-cli`
    only when explicitly set in `execution.adapter`. It never selects CLI
    automatically. See [execution adapters](references/execution-adapters.md).
    
    ```bash
    # Generate a storyboard only. The runner prints [storyboard-preview] with an
    # absolute path; display and inspect that image before starting video work.
    GRID_ONLY=1 node scripts/generate.mjs scripts/myjob.json
    
    # Run only the first independent clip or segment as a quality gate.
    CLIPS_MAX=1 node scripts/generate.mjs scripts/myjob.json
    SEGS_MAX=1 node scripts/generate.mjs scripts/myjob.json
    ```
    
    ```json
    { "execution": { "adapter": "atlas-cli" } }
    ```
    
    Available runner modes:
    
    - `grid`: crop a model storyboard and make one independent I2V clip per panel;
      suitable for a hard-cut montage.
    - `shot-pairs`: make independent I2V shots from `segments[]`, with a `first`
      and optional `last` keyframe index.
    - `reference`: send a small role-specific set of R2V references. A reference
      may be generated from a prompt or read from a local file.
    - `chain`: continue one action by feeding the generated tail frame to the next
      segment. It is an alignment aid, not a guarantee of an invisible seam.
    
    For fault-specific fixes, read [troubleshooting](references/troubleshooting.md).
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related