I spent a few days with Seedance 2.5. Here's what I learned.

@minchoi
ENGLISCH06. Aug. 2026
189K
66
5
8
28

TL;DR

Min Choi shares a deep dive into Seedance 2.5 on CapCut, providing five detailed prompt templates and a professional workflow for consistent, high-quality AI video generation.

Seedance 2.5 is on CapCut now. 30-second scenes in a single generation. Up to 50 references. Real dialogue with lip sync. Motion control that sticks.

Here are 5 AI videos I made with the prompts so you can make your own.

The behind-the-scenes flood that's been going viral

You've seen this one on your timeline - a miniature city in a water tank on a soundstage, sluice gate opens, whole thing gets wiped out.

Here's how to actually make it.

No references. Just prompt (but you can also add input references if you want).

The entire thing rides on scale - the model has to feel enormous, which means the crew have to read as tiny:

the financial district towers stand between four and nine metres tall, and the crew walking the catwalks alongside them are dwarfed

Then you defend it, because the model will quietly shrink it into a tabletop diorama if you let it:

The tank is set into the stage floor - never on a table, trestles, benches or any raised surface. The model buildings are taller than the crew, always.

Full prompt:

markdown
128 seconds. 16:9. Amateur behind-the-scenes phone video shot inside an enormous film
2soundstage. Handheld, kinetic, one continuous take, phone-camera image quality — not
3cinematic.
4
5THE MODEL: A colossal practical miniature of Lower Manhattan built into a long concrete
6water tank set flush into the stage floor. This is a large-scale build, not a tabletop —
7the financial district towers stand between four and nine metres tall, and the crew
8walking the catwalks alongside them are dwarfed, reaching only partway up the lower
9buildings. Geography is accurate: the island tapers to a point at Battery Park at the
10left-hand end and widens to the right, the dense financial district cluster rising just
11behind the tip, avenues running true in a clear grid, and the Brooklyn and Manhattan
12bridges springing east across the tank water behind the island to a low Brooklyn shore.
13
14The build is convincingly weathered — streaked concrete, water-stained stone, rust on fire
15escapes, grime around window frames. Rooftops carry full clutter: HVAC units, water
16towers, vents, aerials, cables. Model traffic sits in the streets. Nothing is new or clean.
17
18CAMERA POSITION: at the tank rail on the model's west side, held by a crew member, not
19rigged. The island runs broadside across the frame — seen in profile, left to right, not
20receding into depth. Crew on the near catwalk pass through the foreground in silhouette.
21
22THE GATE: at the far left corner of the tank, diagonally opposite the camera, a colossal
23steel sluice gate is set into the tank wall — twelve metres wide, scarred and
24rust-streaked, holding back a full reservoir. Hydraulic rams on either side. Amber warning
25beacons above.
26
27THE STAGE: blue cyclorama with orange tracking crosses beyond the model. A technocrane
28overhead. Crew in plain grey shirts. Work lights on stands, cables taped down.
29
30SEQUENCE — 28 SECONDS
310-6s ESTABLISH SCALE. The stage is quiet, the model dry. The camera tracks slowly
32 left to right along the length of the island in profile, the whole skyline
33 reading across the frame. A crew member walks the near catwalk through the
34 foreground and is plainly, obviously smaller than the towers behind them. The
35 camera tilts up the tallest tower and the top is far above head height.
366-9s THE GATE. The operator pans left to the far corner. The steel gate fills that
37 end of the tank. Amber beacons start turning. A klaxon sounds twice. Crew step
38 back. Someone off-camera calls "clear."
399-11s ACTION. A voice off-camera calls "Action." The hydraulic rams haul the gate
40 upward, steel screaming against steel.
4111-14s RELEASE. Water blasts out beneath the rising gate under enormous pressure — a
42 solid white wall, not a build-up. It enters at the top left of frame and
43 immediately runs diagonally across the tank, the front cutting toward the
44 near-right and gaining height as it comes.
4514-22s IMPACT, IN PROFILE. The wave strikes the island along its western flank. Seen
46 side-on, the front stands roughly two thirds the height of the tallest towers —
47 the comparison is direct and unmistakable. Battery Park goes under first at
48 frame-left, then the impact runs rightward along the island like a fuse. Towers
49 shear at the base and topple toward camera with real weight and real slowness,
50 four and nine metre structures going over one after another. Water surges up the
51 avenues between them. Both bridge spans lift off their piers and break. Debris
52 the size of furniture tumbles in the flow. The leading edge reaches the near tank
53 wall and comes over the rail at the camera — the operator shouts, stumbles
54 backward, and the frame swings wildly before recovering.
5522-28s AFTERMATH. The surge subsides. Broken tower sections circle slowly in the tank.
56 Water pours off the cyclorama edge and drains across the stage floor. Someone
57 laughs off-camera. Two crew walk into frame in waders. The operator keeps
58 filming several seconds too long.
59
60IMAGE: deep depth of field throughout — everything from the near rail to the far cyclorama
61stays sharp. No shallow focus, no macro, no tilt-shift. Faint atmospheric haze at the far
62end of the stage. Phone-camera characteristics: aggressive auto-exposure, highlight
63clipping on the work lights, digital noise in the shadows.
64
65AUDIO: diegetic only — stage hum, the klaxon, the off-camera "clear" and "Action", steel
66grinding as the gate lifts, the enormous release of water, tower sections cracking and
67going over in sequence, water pouring over concrete, crew shouting and laughing, the phone
68mic distorting badly on the loudest moment. No music, no narration.
69
70CONSTRAINTS: the island is seen broadside in profile, never head-on and never receding into
71depth. The tank is set into the stage floor — never on a table, trestles, benches or any
72raised surface. The model buildings are taller than the crew, always. No people in the
73model, no model figures, no casualties, no injury. No readable signage or brand marks on
74any building. No brand marks on crew clothing. One continuous take, no cuts. No slow
75motion, no tilt-shift.

A selfie travel vlog that looks insanely realistic

0:00 / 0:28

Mykonos alleys, into a packed bar, out onto a terrace where the camera blows completely white before the sea comes back. A wave hits the wall, soaks the lens, she's laughing.

No references needed - but you can also drop in your own images, or AI avatar images if you're building content for an AI social account.

The trick is asking for everything you'd normally try to avoid:

Unstable grip, focus hunting, imperfect framing, over-zooming and correcting. Auto-exposure fighting the light constantly and losing.

Vague won't do it. "Amateur footage" gets you nothing. "The frame goes completely white, holds about a second, then slowly claws back" gets you the shot.

Last line of the prompt does the heavy lifting:

This should look recovered, not made.

Full prompt:

markdown
128 seconds. 9:16. Handheld consumer camcorder footage, filmed by the subject herself,
2early-2000s tape era.
3
4CAMERA / LOOK: She holds the camera herself throughout. Unstable grip, focus hunting,
5imperfect framing, over-zooming and correcting. Auto-exposure fighting the light
6constantly and losing. Soft tape-like image, compression artefacts, visible grain,
7clipped highlights, colour drifting as the white balance chases. She talks to the camera
8the way you talk to a friend holding it, not to an audience.
9
10SUBJECT: Woman, early twenties. Long blonde hair, salt-dried and pushed back off her
11face. Sunburn across the nose, no makeup. A thin white cotton shirt open over a swimsuit
12top, denim shorts, sandals. Same person, same clothes, throughout.
13
14SETTING: Mykonos Town, late afternoon going into sunset. Narrow whitewashed alleys with
15bougainvillaea over the walls and painted blue doors, opening onto Little Venice, where
16the bars sit directly on the water and waves break against the terrace walls below the
17tables. Busy — this is peak hour.
18
19STORYBOARD — 28 SECONDS
200-5s Alley. She films while walking, the frame swinging as she squeezes past people
21 coming the other way. Whitewash blown out, shadows crushed. She says something
22 short and laughs at herself.
235-10s She stops to film a cat asleep on a step beside a blue door, zooms in too far,
24 pulls back. Her hand comes into frame to point. Someone says something to her
25 off-camera and she answers without turning the camera.
2610-15s She goes into the bar. The image drops to near-black and hunts for several
27 seconds before it finds the level — bodies, a crowded bar top, bottles backlit,
28 low warm light. She threads through the crowd holding the camera up over her
29 head, filming the ceiling and the tops of heads as much as anything.
3015-19s She reaches the far door and steps through onto the terrace. The frame goes
31 completely white. It holds there for about a second, then slowly claws back, and
32 as it recovers the sea comes up out of the white — low sun on the water, the
33 horizon, tables right at the edge, the windmills on the headland in silhouette.
34 She says something quiet.
3519-24s She pans slowly along the terrace. People at tables inches from the drop, a
36 waiter edging past with a tray, the water directly below the wall. Lens flare
37 streaking across the frame every time she crosses the sun.
3824-28s A wave hits the wall below and throws spray straight up over the terrace. People
39 at the nearest table flinch and shout and laugh. She gets caught by it — the
40 frame jolts hard, droplets land on the lens and stay there, and everything after
41 is shot through them. She is laughing. She keeps filming the water for a few
42 seconds too long, then reaches for the record button.
43
44AUDIO: on-camera mic only. Alley voices and footsteps, then the bar — crowd noise,
45glassware, and low indistinct music with no discernible melody, buried under the room.
46Outside, wind buffeting the mic hard, sea against stone, the crowd reacting to the wave.
47Her voice unprojected, casual, half-swallowed at the ends of sentences. No soundtrack,
48no voiceover.
49
50CONSTRAINTS: no grading, no stabilisation, no crane or drone moves, no modern phones or
51cars, no burned-in text or timestamps, no polish. No bar name, signage, menu text, logo
52or brand mark anywhere in frame. No recognisable song. This should look recovered, not
53made.

A 30-second product ad for your own brand

0:00 / 0:24

Min Choi - inline image

A drink can. 110-degree orbit, macro on the wordmark, condensation running, a pour, back to the opening frame.

Generate a product sheet for your own brand, reference it in this prompt, and tweak from there. That's the whole workflow.

The reason AI product shots usually fail is the label - letters drop, the logo bends around the curve, you salvage four seconds out of twenty. One line fixes it:

THE PRODUCT IS NOT GENERATED. Treat the uploaded reference as a photograph of a real object composited into every shot. Where a shot would require the can to change, the shot changes instead.

That last clause is everything. You're giving it an escape hatch - if the angle gets hard, move the camera, don't redraw the product.

Then free what should move:

Condensation is NOT fixed. Droplets form, run, merge and fall freely throughout.

Full prompt:

markdown
124 seconds. 9:16. Luxury beverage commercial. 8K photorealistic, cinematic macro
2photography, physically accurate materials, shallow depth of field, smooth motion.
3
4THE PRODUCT IS NOT GENERATED. Treat the uploaded reference as a photograph of a real
5object that has been composited into every shot. Its proportions, wordmark, band position,
6lettering and finish arrive already fixed and are never redrawn. Where a shot would
7require the can to change, the shot changes instead.
8
9The can is slim, matte, deep forest green, with one brushed copper band at a third of its
10height. Above the band, large and clear across roughly sixty percent of the can's width,
11the wordmark HOLM — four letters, H, O, L, M, always all four, always fully formed. Below
12the band, GREEN MANDARIN in the same typeface at a third the size. Nothing else appears
13on the can.
14
15Condensation is NOT fixed. Droplets form, run, merge and fall freely throughout.
16
17SEQUENCE — 24 SECONDS
180-4s Wide still life, slow push-in. The can stands on wet black volcanic stone, label
19 square to camera. Cold mist rolls low. Dark background, soft green ambient light,
20 deep bokeh.
214-9s A slow 110-degree orbit — and no further. The label stays readable through the
22 entire move and is never seen edge-on. Copper rim-light rides the band.
23 Condensation tracks downward at real speed.
249-12s The camera returns square to the label and pushes to a tight macro. HOLM fills
25 the frame, all four letters sharp, edges clean, no distortion.
2612-16s Extreme macro on condensation. Droplets merge and release, running down the matte
27 surface with accurate surface tension.
2816-19s A hand enters, lifts the can cleanly out of frame. The camera stays. Mist
29 collapses into the space where it stood.
3019-22s Macro insert: the pour. Liquid enters a chilled glass with real viscosity, foam
31 forming and settling, light refracting through the column.
3222-24s The can is back on the stone beside the full glass, label square to camera, rim
33 light resolving. Final frame held.
34
35AUDIO: diegetic only — the crack of the seal, carbonation, liquid against glass, low room
36tone. No music, no voiceover.
37
38CONSTRAINTS: the wordmark is never partial, never cropped, never fewer than four letters;
39no rotation beyond 110 degrees; no text overlays, subtitles or watermarks; no second can;
40no label warping; no hand in frame except 16-19s; real-time throughout.

Hollywood is not ready for this

0:01 / 0:30

Min Choi - inline image
Min Choi - inline image

Two people in a hospital car park. Six shots, four lines of dialogue, full lip sync.

The prompt is built as shot blocks - camera, action, audio, environment, separated per beat:

markdown
1[CAM] MCU, LOCKOFF, eye level on Mara →
2[ACT] Mara's jaw sets. She looks at the ground, then up, decided →
3[AUDIO] Mara: "You already knew when you called me."

Two character reference sheets lock identity across all six shots. The dialogue lives in the [AUDIO] tags. Camera language stays in [CAM].

Swap in your own characters and lines and you've got a scene for your next AI film or short.

Full prompt:

markdown
128 seconds. 16:9. Photorealistic contemporary drama. Natural performance, no
2over-acting. Available light.
3@Image1 — identity reference for MARA. @Image2 — identity reference for DEV.
4Both identities locked across every shot. Wardrobe unchanged throughout.
5SETTING: A hospital car park at dusk. Sodium lamps just coming on, wet asphalt, a
6low concrete wall, distant traffic. Cold blue sky going to orange at the horizon.
7[CAM] MS, HANDHELD, two-shot, Mara frame-left →
8[ACT] They stand a metre apart, not looking at each other. Mara's arms are folded.
9Dev turns a set of car keys over in his hand, once, then again →
10[AUDIO] Distant traffic, a door closing somewhere across the lot →
11[ENV] Sodium lamps flickering to full.
12[CAM] MCU, LOCKOFF, eye level on Mara →
13[ACT] Mara's jaw sets. She looks at the ground, then up, decided →
14[AUDIO] Mara: "You already knew when you called me."
15[CAM] MCU, HANDHELD, OTS over Mara's right shoulder onto Dev →
16[ACT] Dev's mouth opens, closes. He looks away toward the building, then back. The
17keys stop moving →
18[AUDIO] Dev: "I didn't want to say it on the phone."
19[CAM] CU, slow push-in on Mara →
20[ACT] Her eyes fill but do not spill. A single blink held slightly too long →
21[AUDIO] Mara: "So you drove all this way to not say it here either."
22[CAM] MS, HANDHELD, profile two-shot, the gap between them centred →
23[ACT] Dev takes a half step forward. Mara does not move away, but does not close the
24gap. His hand lifts and settles on the wall instead of on her →
25[AUDIO] Dev: "Third floor. Room eleven. She's awake." → A long beat.
26[CAM] WS, STEADICAM, slow dolly out, both of them small against the building →
27[ACT] Mara unfolds her arms. She starts walking toward the entrance. Dev stays a
28second longer, then follows, half a pace behind →
29[AUDIO] Footsteps on wet asphalt, automatic doors, traffic. No music.
30CONSTRAINTS: exactly 28 seconds; cuts only at the shot boundaries above; dialogue
31delivered at natural conversational pace with accurate lip sync; no score, no
32subtitles; no additional dialogue invented; performances restrained throughout.

Motion control up to 30 seconds

Rooftop parkour. Kong vault, roll, wall run, precision jump, cat leap, front flip.

markdown
1
2Take any reference video with movement you want, and put a character of your choice into it - up to 30 seconds, holding consistent.
3Three sources, three jobs. Choreography from the video reference, body and wardrobe from one image, face and hair from another. They never met.
4The instruction that matters is about what doesn't carry over:
5Nothing else crosses over from @Video1 - including sound. Its subject's face, hair, build, clothing, surroundings and AUDIO are all irrelevant.
6I had to say that twice, in two places. Motion references leak - hand it a video and it wants everything in the video, including a man's breathing on a clip where the athlete is a woman.
7Full prompt:
820 seconds. 9:16. Photorealistic.
9SOURCES
10@Video1 — movement only. Supplies choreography, timing, trajectory, weight shift,
11rotation, and the exact moments of release and landing.
12@Image1 — the body. Supplies build, proportions and wardrobe from the collarbone down.
13@Image2 — the head. Supplies face, hair and identity. Skin tone runs continuous across
14the neck.
15Nothing else crosses over from @Video1 — including sound. Its subject's face, hair, build,
16clothing, surroundings and AUDIO are all irrelevant. @Video1 contributes no sound of any
17kind: its subject's voice, breathing, grunts and exertion are discarded entirely and are
18never heard. Read it for motion and discard everything else.
19MOVEMENT QUALITY: She is load-bearing in every frame. Muscle fires before the limb moves.
20When she meets a surface, the surface pushes back — ankles and knees compress on landing,
21fingers take weight on an edge, the torso counter-rotates to hold balance. Nothing about
22her drifts, hangs or arrives softly. This is a trained body under load.
23SETTING: A working rooftop at blue hour — vent housings, gravel, a water tank on legs, a
24low parapet, a taller stair-head wall, a gap between two roof levels, unremarkable
25mid-rise buildings beyond. No landmarks.
26SEQUENCE — 20 SECONDS
270-5s Wide, camera tracking laterally. She enters at a hard run along the roof edge and
28kong vaults the first vent housing — hands planting, legs shooting through
29between her arms. Gravel kicks on landing. Timing matched to @Video1 exactly.
305-11s The camera swings low and three-quarter. She rolls out of the landing over one
31shoulder and comes straight back up into the run, then runs three steps up the
32stair-head wall and pushes off it sideways, rotating 180 degrees in the air and
33landing already moving. Wardrobe from @Image1 moves with real fabric weight.
3411-16s Tight tracking from behind and to the side — she is never hidden by the wall.
35A precision jump across the gap between roof levels, stuck on the far edge, then
36a cat leap to the water tank platform: hands catching the top edge, feet planting
37on the face, and a hard pull over.
3816-20s She drops from the platform into a single front flip, rotating once, and lands.
39Impact absorbed deep through both knees, one hand down for balance, head coming
40up last. The camera settles to a held medium. Her face reads clearly and is
41exactly @Image2.
42AUDIO: diegetic only, and every human sound is FEMALE. The subject is a woman in her late
43twenties; all breathing, exhalation and effort sounds are hers — lighter, higher pitched,
44unmistakably a woman's voice throughout. Sharp exhalations through the vaults, controlled
45breathing between them, and a hard grunt on the final landing, all female. Plus footfall on
46gravel and metal, hands slapping concrete, fabric, wind at height, distant traffic. No male
47voice anywhere in the mix. No music.
48CONSTRAINTS: no identity drift; nothing from @Video1's wardrobe; no audio from @Video1; no
49male vocal sound of any kind; no wire-work weightlessness; no rotation beyond what @Video1
50contains; the body is never cropped or occluded during a rotation; no logos.

What these all have in common

The prompts aren't descriptions. They're mostly restrictions.

Where you want freedom, describe it. Where you don't, forbid it by name. Roughly half of every prompt above is telling the model what not to do.

And when you stack references, they fight each other. So you rank them:

@Image1 controls identity and is never overridden. @Image2 controls all clothing; if @Image1 shows different clothing, @Image2 wins.

That's the difference between a face that holds across six shots and one that slowly becomes someone else.

How I made these in CapCut

The loop is simple: come up with your scene → generate your characters, environments, and motion references → feed them back in with @ tags → generate.

That's it. The rest is knowing which settings to use.

Min Choi - inline image

In Video Studio you pick Standard or Director Mode. They're two different workflows and the choice matters more than it sounds.

Standard Mode is prompt → generate. Fast, single-pass, minimal steps. If you want one clean 15–30 second clip - a product shot, a social post, a one-off scene - this is fine and it's quicker.

Director Mode is the filmmaker workflow, and it's what I used for almost everything above. You work on an infinite canvas with scene cards. The agent helps develop the script, visualize characters and props, and keep them consistent across shots. The important part: you can rework one scene, one asset, or one frame without regenerating the whole project.

That last bit is the difference. When a face drifts on shot four, you fix shot four. You don't burn credits redoing all six.

Min Choi - inline image

The actual sequence:

  1. Work out the scene - what happens, how long, how many shots
  2. Generate the pieces that need to stay consistent - character sheets, wardrobe, product turnarounds, environment plates, motion reference video
  3. Bring those in and tag them in the prompt: @Image1, @Image2, @Video1
  4. Set model to Seedance 2.5, pick your mode, set aspect ratio and duration
  5. Write the prompt properly - this is where the work is
  6. Refine individual frames or scenes rather than regenerating
  7. Assemble and polish

Things that actually matter:

Prep beats prompt-and-pray. This model follows instructions faithfully, which means whatever you left unsaid, it decides for you. Every prompt above is long on purpose.

Timestamp everything. "0-6s establish scale, 6-9s the gate, 9-11s action, 11-14s release" is why 30 seconds holds together instead of drifting.

One good reference beats five mediocre ones. A single clean character or product sheet holds consistency across a full 30 seconds. Add more when you need more control, not because more feels safer.

Fix the frame, not the film. Biggest time-saver in Director Mode. Change an expression, an object, or one shot without touching the rest.

Finishing: trim the soft first frame, put your text hook in the first ~0.8 seconds instead of relying on generated text, layer real audio where it matters, grade for consistency, export high bitrate.

100% AI.

None of this needed a crew, a camera, or a location. Just knowing what to ask for.

In YouMind remixen

Turn one viral article into a full content workflow

Collect the source, decode the pattern, create assets, draft the story, and distribute from one AI workspace.

Explore YouMind
Für Creator

Verwandle dein Markdown in einen sauberen 𝕏-Artikel

Wenn du eigene Langtexte veröffentlichst, wird die 𝕏-Formatierung von Bildern, Tabellen und Codeblöcken mühsam. YouMind macht aus einem ganzen Markdown-Entwurf einen sauberen, sofort postbaren 𝕏-Artikel.

Markdown zu 𝕏 testen

Mehr Muster zum Entschlüsseln

Aktuelle virale Artikel

Mehr virale Artikel entdecken