Honestly, I did a double-take at the screen.
Nowadays, when you open TikTok or YouTube, vertical short dramas are constantly streaming.
Now, those can be created with a single prompt.
Dialogue, sound effects, BGM—everything included.
No filming. No actors. No editing.
The cost was approximately 1,400 yen per video.
Many people probably pass this by thinking, "I heard it's impressive." I understand that feeling all too well.
At first, I thought, "It's probably just a few seconds of CG video."
But what actually came out was a 30-second piece where the actor's lip movements matched, and it even included cuts I hadn't specified.
It wasn't just that "video was created"; it was that "a work of art emerged."
That's why I wrote this article.
I've included the full prompt that you can copy and paste to use immediately.
Also, whether you read it before or after this article, I run a LINE Open Chat where I share stories about AI utilization that I can't fit into articles. There, I'm giving away "32 Luxury AI Strategy Bonuses" for free. Participation is free, so feel free to take a look.
▼ Open Chat "Fuji AI [Latest Information]"
You can join via the link in my profile.
Let's get started.
◾️ What is Seedance 2.5?
Seedance 2.5 is a next-generation AI video generation model developed by ByteDance. From inputs like text, images, audio, and video, it can create a story-driven video up to 30 seconds long with audio in a single generation.
It was created by ByteDance, the parent company of TikTok. It was released on July 31, 2026.
You only need to remember three main features:
① 30 seconds in one generation, with sound
The previous generation, 2.0, was 15 seconds. This has exactly doubled. Moreover, the model handles the cuts itself. Even if you don't specify them, it automatically switches from wide shots to close-ups.
② Up to 50 reference materials can be provided
The breakdown is 30 images, 10 videos, and 10 audio files. You can fix the face, costume, location, and even the voice. This "voice fixing" becomes very effective later.
③ Simultaneous creation of video and sound
It doesn't create the video first and add sound later; it generates them together from the start. Therefore, the lip movements and voice do not go out of sync.
That covers the basics of Seedance 2.5. However, let me give you a warning first.
This is not a model that simply "increased image quality."
The official blog doesn't mention resolution once. In the actual service, you can only choose between 480P and 720P. The 1080P option doesn't even exist. Some internet articles claim a "great 4K version" was released, but that contradicts the official price list.
It is a model designed for length and consistency. That is the correct understanding.
◾️ The one reason it's the only choice for vertical dramas
The reason I chose this was simple: Length.
No matter how high the performance of Google's video AI is, it can only create 10 seconds at a time. This is fundamentally incompatible with vertical dramas. The structure of a vertical drama is "Hook → Build-up → Reversal → Punchline." This takes exactly 30 seconds. If you make this in three 10-second segments and join them later, the face changes at the cuts. The atmosphere changes. The protagonist in the third cut becomes a different person from the first. Seedance is currently the only one that can handle 30 seconds in one go. That's why for vertical dramas, you choose this over Google.
◾️ Let's try it (Just copy and paste)
It's not difficult. If you follow the steps, you can definitely do it.
[Step 1: Open Dreamina]

The gateway to using Seedance 2.5 from Japan is a service called Dreamina. It's under the same ByteDance umbrella as CapCut, and the interface is fully localized.
[Step 2: Log in]

You can choose from five login methods: Google, TikTok, Facebook, Email, or CapCut Mobile. Google is the fastest.
[Step 3: Switch the model to Seedance 2.5]

Click "AI Video" on the bottom left to see the model list. Select "Dreamina Seedance 2.5." Only 2.5 can output 30 seconds continuously.
[Step 4: Adjust three settings]
Set the buttons to the right of the model as follows:
- Generation Mode: Omni-reference (Default is fine)
- Aspect Ratio and Resolution: 9:16 / 720P
- Duration: 30 seconds
Crucially, the credit consumption is displayed to the left of the generate button. Each generation costs money, so get in the habit of checking this number.
[Step 5: Copy, paste, and send the following]
You don't need to change anything.
▼ Copy and paste everything below ▼
1Please create a 30-second vertical short drama (9:16). Use a tempo similar to Chinese vertical short dramas, with strong emotions and a structure where the power dynamic completely flips at the end.2[Character Definitions]3Subject 1 = Japanese man in his 40s. Short hair, wearing a black apron, the store manager.4Subject 2 = Japanese man in his late 20s. Black hair, grey hoodie, carrying a backpack.5Subject 3 = Japanese woman in her 30s. Navy suit, holding an envelope.6Always use these names hereafter.7[Shot Composition]8Shot 1 (0-3s)9- Camera: Fixed bust shot of Subject 110- Action: Subject 1 crosses arms, smirks, and looks down11- Sound: (No music, just store noise) / Says in Japanese {We don't need customers who don't even know how to queue.}12Shot 2 (3-9s)13- Camera: Slow zoom into Subject 2's face14- Action: Subject 2 looks down and slowly adjusts his backpack strap15- Sound: Says in Japanese {...I'm sorry} / <Small sound of dishes clinking>16Shot 3 (9-16s)17- Camera: Fixed medium shot of the store18- Action: Several surrounding customers awkwardly look away19- Sound: (A single low note) / <Sound of the sliding door opening quietly>20Shot 4 (16-22s)21- Camera: Slow movement from the entrance into the store22- Action: Subject 3 enters and presents an envelope and key to Subject 2 with both hands23- Sound: (Strings rise powerfully) / Subject 3 says in Japanese {I've come from headquarters. These are the documents for the manager change starting tomorrow.}24Shot 5 (22-27s)25- Camera: Fixed close-up of Subject 2's hands26- Action: Subject 2 quietly receives the key. Subject 1 stops with his mouth half-open27- Sound: Subject 2 says in Japanese {I wanted to decide after seeing the site myself} / Subject 1 says in Japanese {Eh?}28Shot 6 (27-30s)29- Camera: Fixed medium shot of all three30- Action: Subject 1 takes a half-step back. Subject 2 retorts after a beat. Subject 3 turns her face slightly away31- Sound: (Music suddenly lightens) / Subject 2 says in Japanese {Can you teach me how to queue starting tomorrow?} / Subject 3 says quietly in Japanese {Manager, that is the new manager.}32[Direction]33- Live-action / Ultra-realistic / Natural Japanese actors34- Small ramen shop at night. Warm lighting, low saturation35- Strong emotional acting. Only Subject 1 raises his voice36- 9:16, high quality, stable screen, no blur, no shake37[Consistency Specifications]38Same person, consistent clothing, unchanged hairstyle. Maintain appearance and clothing across all shots. Do not show two or more people with identical appearance/clothing (no twin effect).39[Prohibitions]40- No subtitles, captions, logos, watermarks, or text on screen41- No real people or real brand logos42- No readable text on signs, menus, or documents43- No sprinting, big jumps, or violent rotations44[Closing]45Convey the satisfaction of the moment the one looking down loses their position due to a single document. Only Subject 1 should be loud; the protagonist remains quiet until the end. Compose with expressions, acting, and camera work so the content is clear even without subtitles.
◾️ 6 things the official docs say "Don't do"

1. Don't provide character turnaround sheets
Unlike anime, providing front/side/back views makes the model think there are three different people, causing the face to break. Use one neutral headshot.
2. Don't use reference materials to the limit
You can add 50, but the recommendation is 4-5. Too many causes the model to lose track of priority.
3. Don't specify multiple camera movements in one shot
"Slowly zooming while panning and circling" causes the image to collapse. One shot = one camera movement.
4. Don't paste the script directly into the prompt
Redundant text confuses the model. Use a structured format like the prompt above.
5. Don't try to fix the seed
Seed fixing doesn't work for Seedance 2.0 series. Consistency is created via reference images and subject definitions.
6. Don't write violent actions
Avoid sprinting or jumping. Prioritize slow, continuous small movements.
◾️ Japanese dialogue will have an accent unless you write this

For Japanese dialogue, always specify the language name.
× "Sorry, I can't do it after all"
○ Says in Japanese {Sorry, I can't do it after all}
Without this, it might have a Chinese accent or start speaking English. Also, use the brackets correctly: () for music, <> for sound effects, {} for dialogue, and 【】 for subtitles.
◾️ If you rewrite it for your own story, use this order

Accurate Subject + Action Details + Location/Environment + Light/Tone + Camera Work + Video Style + Quality + Constraints
Think of it as an engineering specification, not poetry.
You can even specify the voice
You can specify emotion, tone, and speech rate in text:
Says in Japanese {Dialogue} with [Emotion] emotion, in a [Tone] tone, at a [Speed] speech rate.
◾️ How to fix it when the first try fails

The fix is not "adding words" but "closing escape routes."
- Delete all ambiguous words like "if possible" or "as much as possible."
- Don't write two instructions in one line.
- Write "how you want it to be seen" in the closing.
◾️ There is a template for successful 30-second videos

- 0-2s: Hook (Discomfort or Surprise)
- 2-18s: Conflict (Being looked down upon)
- 18-26s: Reversal (Show it with an object like a key or document)
- 26-30s: Twist/Resolution
◾️ For sequels, just provide the first video
To make a series, put the first video into the reference materials and specify it as @Video1 in the prompt. This maintains the face, clothes, and shop.
◾️ Two things I don't leave to the AI
- Audio: I don't rely on it entirely. For high quality, generate silent video and dub it later.
- Vertical generation: The official FAQ notes that generating in 9:16 has a higher chance of unwanted subtitles appearing. Generating in 16:9 and cropping to vertical is safer for mass production.
◾️ The honest truth about "1,400 yen per video"
This is the price if you succeed in one shot. In reality, with retries, it's more like 8,000 to 10,000 yen per episode. Still, compared to the 160,000 yen market price for human-filmed shorts, it's a different league.
◾️ Can you make money with this?
In China, 95% of vertical dramas are already AI-made. Revenue per 10,000 views has dropped by 90%. The barrier to entry has lowered, meaning competition has exploded. AI doesn't guarantee profit; it guarantees the ability to test many ideas cheaply. Japan is still at the entrance of this phase.
◾️ Summary
- Seedance 2.5 is for length (30s) and consistency, not 4K resolution.
- Use specific brackets and "In Japanese" for dialogue.
- Don't add words to fix errors; be absolute in prohibitions.
- AI makes the "success rate" no higher, but the "cost of trial" much lower.
◾️ If you do one thing today
Open Dreamina and paste the prompt from this article once. Don't just read and be convinced; move your hands and experience that moment of surprise.
Finally
To ensure you don't just "finish reading," I'm giving away 32 luxury bonuses for free, including guides for ChatGPT, Claude, and AI side hustles.






