How to Create Videos Exactly as Imagined: The Professional Way to Use Seedance 2.5

@casthirotaka
JAPANISCH15. Aug. 2026
159K
165
24
2
600

TL;DR

This guide explains how to master Seedance 2.5 by treating prompts as directorial instructions rather than descriptions, using video references to handle complex motions and character consistency.

Those currently challenging themselves with Seedance 2.5 are likely searching for the "perfect prompt" template and testing every idea they can think of.

However...

The scope of responsibility for a prompt is narrower than you think

Seedance 2.5 accepts 30 images, 10 videos, and 10 audio files as references in a single generation (announced by ByteDance official, July 31, 2026).

This specification implies that prompts are not intended to explain everything.

吉田洋孝|Cast Japan代表 - inline image

What to decide

Prompt suitability

Where to hold the "correct" answer

Video purpose, worldview, atmosphere

High

Prompt

Type of location, what happens

High

Prompt

What to keep, what to change

Very High

Prompt

Priorities

Very High

Prompt

Lighting direction, lens feel

Medium

Prompt / Image

Person's face, identity

Low

Image

Accurate product shape

Low

Image

Complex body movements

Low

Video

Human contact, positioning, gaze

Low

Video

Accurate camera trajectory

Low

Video / White model

Frame-by-frame timing

Low

Video / Editing

There is only one criterion for judgment: Who has the most accurate information?

The most accurate source for movement is the video of the movement itself, not a text description of it.

5 Things You Should Not Write in a Prompt

1. Atmospheric descriptions piled with adjectives

吉田洋孝|Cast Japan代表 - inline image

The official guide states: "Try not to pile on adjectives." It continues, "Being specific about movement and lens choice usually works better."

2. "Names" of movements

Conceptual names like playful cheeky mini-dance result in an output range as wide as the interpretation range.

In reality, this instruction returned a video of someone dancing in front of others: "approaching the partner -> putting hands forward -> looking at the partner." The goal was "walking alone while fooling around, slightly messy and fast movement," which isn't even a dance.

3. Decomposing body part movements into text

吉田洋孝|Cast Japan代表 - inline image

Even when rewriting by decomposing how arms swing, shoulders shake, and how the camera reacts in sync for each part, the output did not change. The official guide says motion references are "more useful than a long vague paragraph."

Even if you decompose and lengthen it, it won't reach the information density of a single reference.

4. Information the reference material already has

吉田洋孝|Cast Japan代表 - inline image

Official prompt examples specify what information to take from each reference.

" From @Clay Render 1, reference camera movement, pace, shot size transitions, subject trajectory, and positioning.

From @Image 2, reference character design, scene, materials, lighting, color, and atmosphere."

If you write information that the reference already possesses in the text, the reference and the text will start saying different things, leading to a breakdown.

5. Mutually contradictory instructions

Dense paragraphs cause instructions to clash internally. The moment an impossible instruction is mixed in, the AI decides for itself which one to discard.

5 Things You Should Write in a Prompt

1. The role of each reference file

吉田洋孝|Cast Japan代表 - inline image

The official guide instruction is to "Mention reference roles clearly." Listed roles include product identity, character consistency, movement via white models, green screen acting, audio, style, and partial corrections. Roles are not determined just by attaching files.

2. What NOT to use from that reference

While giving roles, clearly state information you don't want it to pick up. This phrasing actually worked:

Do not copy the location, clothing, text, people or story from @Video1. Copy only the body rhythm, arm gestures, walking rhythm and body-linked camera motion.

Separate it: "Mimic only the movement; do not mimic the location, clothes, or people."

3. Enumeration of non-negotiable elements

吉田洋孝|Cast Japan代表 - inline image

The official motion reference guide instructs prompt construction like this: "Start by naming the uploaded references, then list non-negotiable details."

  • identity
  • shape
  • scale
  • floor plan
  • product design
  • pose
  • timing
  • voice
  • shot order

4. What to keep and what to change

Separately write what to keep (acting, timing, gaze, positioning, camera movement) and what to change (person, environment, costume). Without this distinction, the AI has the freedom to remake everything.

5. Priorities

Specify what to protect first in situations where everything cannot be followed simultaneously. In an actual prompt, it was written like this:

PRIORITY: 1. Match the body rhythm and arm movement of @Video1. 2. Match the body-linked camera motion of @Video1. 3. Preserve the POV perspective and environment from @Image1. 4. Keep the movement playful and casual rather than choreographed.

Decide the "Correct Answer" First

Break the desired video into elements and decide which "material holds the correct answer" for each element.

Element

Where to hold the correct answer

Reason

Face / Person identity

Image

Text cannot maintain identity

Body movement, positioning, gaze, timing

Video

Officially stated as control targets for video reference

Spatial structure, camera trajectory

White model (3D without texture)

Official: "Effective when space, camera path, and relative position are more important than final texture"

Human actions, product demonstration

Green screen material

Officially stated as a use case

Worldview, location, meaning, priority

Prompt

The area where text is strongest

Furniture shape, small items, background details

Leave to AI

No need to fix

吉田洋孝|Cast Japan代表 - inline image
吉田洋孝|Cast Japan代表 - inline image

Furthermore, divide the strength of enforcement for each element into three levels:

  • Fixed: Do not allow arbitrary remaking (person, product shape, original movement material).
  • Guided: Give direction but allow interpretation (luxury apartment, night, warm lighting, fun atmosphere).
  • Auto: Leave it to the AI (shape of the sofa, items on the bookshelf, details on the wall).

If you fix everything, the video becomes unnatural; if you leave everything to auto, it deviates from the goal. Designing where to fix and where to loosen is the job of the prompt.

3 Ways to Create. Shoot Difficult Movements First

Method

Suitable Situations

Let AI create everything

Non-existent worlds, simple movements, strong worldview

Provide samples to create

Face/product shape is fixed, want to continue composition

Shoot first, then change

Complex movements, human contact, multiple people interacting

The reason the third method is important is that ByteDance itself lists "stability of scenes where multiple subjects interact" as an area for improvement in Seedance 2.5. It's a losing battle to try to force through text an area the developer admits is difficult.

吉田洋孝|Cast Japan代表 - inline image

The "shoot first" method works because it replaces the AI's job from "inventing acting" to "changing the appearance of acting that is already established." Positioning, timing, gaze, and contact are held by the human-shot material, while the AI handles only identity and environment.

On X, a case was shared where a single piece of material featuring three different people was converted into the same person for all three. (Using Seedance 2.0, the poster explained no compositing or masks were used. As it is a third-party post, the details of the production process have not been verified.)

There are also clear cases where it is not suitable.

Generating everything without source material, contact is too complex, faces are completely hidden, who is where is ambiguous, or positioning is already broken at the source material stage.

If these five apply, shooting first won't solve it.

The Prompt Template Shown by Official

The basic order is:

Subject (Who/What) -> Action (Does what) -> Camera (How to shoot) -> Style (What texture)

For a 30-second single shot, the official even shows the time allocation:

  • 00–06s: Show the situation with a wide shot
  • 06–14s: Enter the main movement with a close-up
  • 14–24s: Develop by moving the camera or adding inserts
  • 24–30s: Converge

Template to fill in and use when using references👇

text
1REFERENCES:
2@Image1 = [Role]. Use only [Information to use] from here. Do not use [Information not to use].
3@Video1 = [Role]. Use only [Information to use] from here. Do not use [Information not to use].
4
5ACTION:
6[Who] [Where] [Does what].
7[Quality of movement. Write whether it is fast or slow, messy or polite, and who it is directed at, rather than "a movement named X"]
8
9CAMERA:
10[Type of viewpoint / Lens angle of view].
11[What the camera moves in sync with].
12[Camera behaviors that must not be done].
13
14PERFORMANCE:
15[Temperature of acting. Is it being shown to someone or done alone?]
16
17ENVIRONMENT:
18[Location, time, character of light. Do not specify details]
19
20PRESERVE / CHANGE:
21Keep = [Acting / Timing / Gaze / Positioning / Camera movement]
22Change = [Person / Environment / Costume]
23
24PRIORITY:
251. [Thing to protect with highest priority]
262. [Thing to protect next]
273. [Next]

This template is effective not because it reduces the amount of writing.

It's because the job of the text changes from "describing the video" to "directing the division of roles for the materials."

The video holds the correct answer for movement, the image holds the correct answer for the face, and the white model holds the correct answer for the space. The prompt decides their arrangement.

Thank you for reading to the end.

On X ([@casthirotaka](https://x.com/@casthirotaka)), I share stories that are effective for such practical work every day.

I would be happy if you followed me.

In YouMind remixen

Turn one viral article into a full content workflow

Collect the source, decode the pattern, create assets, draft the story, and distribute from one AI workspace.

Explore YouMind
Für Creator

Verwandle dein Markdown in einen sauberen 𝕏-Artikel

Wenn du eigene Langtexte veröffentlichst, wird die 𝕏-Formatierung von Bildern, Tabellen und Codeblöcken mühsam. YouMind macht aus einem ganzen Markdown-Entwurf einen sauberen, sofort postbaren 𝕏-Artikel.

Markdown zu 𝕏 testen

Mehr Muster zum Entschlüsseln

Aktuelle virale Artikel

Mehr virale Artikel entdecken