Those currently challenging themselves with Seedance 2.5 are likely searching for the "perfect prompt" template and testing every idea they can think of.
However...
The scope of responsibility for a prompt is narrower than you think
Seedance 2.5 accepts 30 images, 10 videos, and 10 audio files as references in a single generation (announced by ByteDance official, July 31, 2026).
This specification implies that prompts are not intended to explain everything.

What to decide | Prompt suitability | Where to hold the "correct" answer |
|---|---|---|
Video purpose, worldview, atmosphere | High | Prompt |
Type of location, what happens | High | Prompt |
What to keep, what to change | Very High | Prompt |
Priorities | Very High | Prompt |
Lighting direction, lens feel | Medium | Prompt / Image |
Person's face, identity | Low | Image |
Accurate product shape | Low | Image |
Complex body movements | Low | Video |
Human contact, positioning, gaze | Low | Video |
Accurate camera trajectory | Low | Video / White model |
Frame-by-frame timing | Low | Video / Editing |
There is only one criterion for judgment: Who has the most accurate information?
The most accurate source for movement is the video of the movement itself, not a text description of it.
5 Things You Should Not Write in a Prompt
1. Atmospheric descriptions piled with adjectives

The official guide states: "Try not to pile on adjectives." It continues, "Being specific about movement and lens choice usually works better."
2. "Names" of movements
Conceptual names like playful cheeky mini-dance result in an output range as wide as the interpretation range.
In reality, this instruction returned a video of someone dancing in front of others: "approaching the partner -> putting hands forward -> looking at the partner." The goal was "walking alone while fooling around, slightly messy and fast movement," which isn't even a dance.
3. Decomposing body part movements into text

Even when rewriting by decomposing how arms swing, shoulders shake, and how the camera reacts in sync for each part, the output did not change. The official guide says motion references are "more useful than a long vague paragraph."
Even if you decompose and lengthen it, it won't reach the information density of a single reference.
4. Information the reference material already has

Official prompt examples specify what information to take from each reference.
" From @Clay Render 1, reference camera movement, pace, shot size transitions, subject trajectory, and positioning.
From @Image 2, reference character design, scene, materials, lighting, color, and atmosphere."
If you write information that the reference already possesses in the text, the reference and the text will start saying different things, leading to a breakdown.
5. Mutually contradictory instructions
Dense paragraphs cause instructions to clash internally. The moment an impossible instruction is mixed in, the AI decides for itself which one to discard.
5 Things You Should Write in a Prompt
1. The role of each reference file

The official guide instruction is to "Mention reference roles clearly." Listed roles include product identity, character consistency, movement via white models, green screen acting, audio, style, and partial corrections. Roles are not determined just by attaching files.
2. What NOT to use from that reference
While giving roles, clearly state information you don't want it to pick up. This phrasing actually worked:
Do not copy the location, clothing, text, people or story from @Video1. Copy only the body rhythm, arm gestures, walking rhythm and body-linked camera motion.
Separate it: "Mimic only the movement; do not mimic the location, clothes, or people."
3. Enumeration of non-negotiable elements

The official motion reference guide instructs prompt construction like this: "Start by naming the uploaded references, then list non-negotiable details."
- identity
- shape
- scale
- floor plan
- product design
- pose
- timing
- voice
- shot order
4. What to keep and what to change
Separately write what to keep (acting, timing, gaze, positioning, camera movement) and what to change (person, environment, costume). Without this distinction, the AI has the freedom to remake everything.
5. Priorities
Specify what to protect first in situations where everything cannot be followed simultaneously. In an actual prompt, it was written like this:
PRIORITY: 1. Match the body rhythm and arm movement of @Video1. 2. Match the body-linked camera motion of @Video1. 3. Preserve the POV perspective and environment from @Image1. 4. Keep the movement playful and casual rather than choreographed.
Decide the "Correct Answer" First
Break the desired video into elements and decide which "material holds the correct answer" for each element.
Element | Where to hold the correct answer | Reason |
|---|---|---|
Face / Person identity | Image | Text cannot maintain identity |
Body movement, positioning, gaze, timing | Video | Officially stated as control targets for video reference |
Spatial structure, camera trajectory | White model (3D without texture) | Official: "Effective when space, camera path, and relative position are more important than final texture" |
Human actions, product demonstration | Green screen material | Officially stated as a use case |
Worldview, location, meaning, priority | Prompt | The area where text is strongest |
Furniture shape, small items, background details | Leave to AI | No need to fix |


Furthermore, divide the strength of enforcement for each element into three levels:
- Fixed: Do not allow arbitrary remaking (person, product shape, original movement material).
- Guided: Give direction but allow interpretation (luxury apartment, night, warm lighting, fun atmosphere).
- Auto: Leave it to the AI (shape of the sofa, items on the bookshelf, details on the wall).
If you fix everything, the video becomes unnatural; if you leave everything to auto, it deviates from the goal. Designing where to fix and where to loosen is the job of the prompt.
3 Ways to Create. Shoot Difficult Movements First
Method | Suitable Situations |
|---|---|
Let AI create everything | Non-existent worlds, simple movements, strong worldview |
Provide samples to create | Face/product shape is fixed, want to continue composition |
Shoot first, then change | Complex movements, human contact, multiple people interacting |
The reason the third method is important is that ByteDance itself lists "stability of scenes where multiple subjects interact" as an area for improvement in Seedance 2.5. It's a losing battle to try to force through text an area the developer admits is difficult.

The "shoot first" method works because it replaces the AI's job from "inventing acting" to "changing the appearance of acting that is already established." Positioning, timing, gaze, and contact are held by the human-shot material, while the AI handles only identity and environment.
On X, a case was shared where a single piece of material featuring three different people was converted into the same person for all three. (Using Seedance 2.0, the poster explained no compositing or masks were used. As it is a third-party post, the details of the production process have not been verified.)
There are also clear cases where it is not suitable.
Generating everything without source material, contact is too complex, faces are completely hidden, who is where is ambiguous, or positioning is already broken at the source material stage.
If these five apply, shooting first won't solve it.
The Prompt Template Shown by Official
The basic order is:
Subject (Who/What) -> Action (Does what) -> Camera (How to shoot) -> Style (What texture)
For a 30-second single shot, the official even shows the time allocation:
- 00–06s: Show the situation with a wide shot
- 06–14s: Enter the main movement with a close-up
- 14–24s: Develop by moving the camera or adding inserts
- 24–30s: Converge
Template to fill in and use when using references👇
1REFERENCES:2@Image1 = [Role]. Use only [Information to use] from here. Do not use [Information not to use].3@Video1 = [Role]. Use only [Information to use] from here. Do not use [Information not to use].45ACTION:6[Who] [Where] [Does what].7[Quality of movement. Write whether it is fast or slow, messy or polite, and who it is directed at, rather than "a movement named X"]89CAMERA:10[Type of viewpoint / Lens angle of view].11[What the camera moves in sync with].12[Camera behaviors that must not be done].1314PERFORMANCE:15[Temperature of acting. Is it being shown to someone or done alone?]1617ENVIRONMENT:18[Location, time, character of light. Do not specify details]1920PRESERVE / CHANGE:21Keep = [Acting / Timing / Gaze / Positioning / Camera movement]22Change = [Person / Environment / Costume]2324PRIORITY:251. [Thing to protect with highest priority]262. [Thing to protect next]273. [Next]
This template is effective not because it reduces the amount of writing.
It's because the job of the text changes from "describing the video" to "directing the division of roles for the materials."
The video holds the correct answer for movement, the image holds the correct answer for the face, and the white model holds the correct answer for the space. The prompt decides their arrangement.
Thank you for reading to the end.
On X ([@casthirotaka](https://x.com/@casthirotaka)), I share stories that are effective for such practical work every day.
I would be happy if you followed me.





