Hello! I'm GENEL.
This time, I created a live-action x animation video where hand-drawn animations pop out in a real room and play around!
https://x.com/genel_ai/status/2081307056838263052
In this post, I will introduce:
- How to make this video
- The actual prompts used
Writing Prompts Randomly Leads to Failure
With this kind of video, if you start writing prompts randomly, it almost never turns out the way you want.
Here is what was generated with a prompt like "Hand-drawn animation merging with a live-action night amusement park & the camera follows it":

It has some good points, but some parts don't look like hand-drawn animation, and I felt it was a bit off.
So this time, I thought through the stage setting in detail, second by second, and turned it into a prompt.
First, Narrow Down the Structure to 4 Points
Before deciding on the exact seconds, I roughly decided what would happen within the 15 seconds. If you get too greedy here, it will definitely fall apart, so I narrowed the core of the story down to just four points.
- Animation is born from the hand
- It enters the PC screen
- It jumps out of the PC and moves around the room
- It transforms into a monster
It's a single flow where "doodles encroach upon a real room."

There were many other things I wanted it to do, but they wouldn't fit in 15 seconds. This is where you have to be patient.
Assigning to Time Segments
Once the structure is decided, map it out over time.
0-3 seconds: A white line is born from the hand, becomes a yellow star, and flies to the PC.
3-6 seconds: Inside the PC screen, it changes from a red line to a heart to a blue vortex, then jumps out.
6-10 seconds: Moves through the room while transforming: butterfly -> pink snake -> octopus monster.
10-13 seconds: Opening the refrigerator reveals a hamburger character -> transforms into a monster.
13-15 seconds: A giant eye appears on the wall, blinks, and turns into a smile.
As I wrote in my previous VFX article, dividing by time and writing in order makes it stable.
[Click here for the previous article]
https://note.com/genel/n/nb2c4b89fa1d8?sub_rt=share_sb
Instructions seem to pass through more easily when divided into segments rather than giving a long, rambling description of the situation!
Tools Used: ChatGPT and Seedance 2.0
Once this is decided, all that's left is to turn it into a video. I only used these two AI tools.
1. ChatGPT | Converts the blueprint into a prompt 2. Seedance 2.0 | Turns that prompt into a video
The key is not to just dump everything on ChatGPT and ask it to "think of a cool video."
If you give it the blueprint you've created yourself, it will take shape almost exactly as planned. Conversely, if you just leave it to the AI, you'll get something generic like the amusement park example.
However, if you just throw what you've decided at ChatGPT, it will return a disjointed, hard-to-read text. If you pass that directly to Seedance, it still won't be stable.
So, I even specified the order of the output.
1. The opening line (duration, aspect ratio, what is being merged with what) 2. The atmosphere of the live-action space and the texture of a smartphone recording 3. Movement by second (0-3s / 3-6s / 6-10s / 10-13s / 13-15s) 4. A final summary of the hand-drawn texture, camera tracking method, negative prompts, and ambient sound
Here is the instruction text I actually gave to ChatGPT:
1[Instructions for ChatGPT]2I am going to create a prompt for a 15-second video generation.3Please convert the following blueprint into a prompt to be passed to a video generation AI.45Blueprint6Setting: A dark, lived-in room lit only by warm indirect lighting and small decorative lights.7Video Type: Live-action x Hand-drawn glowing animation8Events:90-3s: A white line is born from a hand, becomes a yellow star, and flies to a PC.103-6s: Inside the PC screen, red line -> heart -> blue vortex, then jumps out.116-10s: Transforms into butterfly -> pink snake -> octopus monster while moving through the room.1210-13s: Opening the fridge reveals a hamburger character -> transforms into a monster.1313-15s: A giant eye appears on the wall, blinks, and turns into a smile.1415Writing Rules:16- When shapes change, write "transforms while collapsing into the next shape" instead of "disappears and the next appears."17- At the end of each segment, always describe the camera movement. The animation moves first, and the camera follows with a delay.18- Maintain continuity of walking and following within the same space without cutting.19- For the live-action side, simulate one-handed smartphone filming. Include camera shake, slight focus hunting, low-light noise, and auto-exposure fluctuations.20- For the hand-drawn animation, use textures like crayon, chalk, colored pencil, or rough brushes. Lines should jitter every frame with uneven coloring and fuzziness.2122Output Order:231. One line: "15 seconds, 16:9 horizontal video. A video merging live-action [location] with hand-drawn glowing animation."242. Atmosphere of the live-action space and smartphone-like texture.253. Each segment: 0-3s / 3-6s / 6-10s / 10-13s / 13-15s.264. Final summary of hand-drawn animation texture, camera tracking, negative prompts, and ambient sounds.2728Negative Prompts to include:29Too precise 3DCG, too smooth vector lines, neon tube-like expressions, perfectly stable gimbal shots, overly cinematic lighting, perfectly composed commercial photography, subtitles, logos, BGM.3031No headings or explanations needed. Output only the prompt text.

Actual ChatGPT chat screen
Here is the prompt that was created:
115 seconds, 16:9 horizontal video. A video merging a dark live-action interior with hand-drawn glowing animation.2A slightly messy, lived-in room lit only by warm indirect lighting and small decorative string lights. Strong camera shake as if filmed one-handed with a smartphone, slight focus hunting, low-light noise, and auto-exposure fluctuations. Real texture, not polished like a commercial, but as if captured accidentally by an individual in their room.30-3s:4A live-action hand held up on the left of the screen. A white hand-drawn line is born from the fingertips, clings to the live-action hand, and becomes the shape of another white hand. The lines are fuzzy as if drawn with crayons or chalk and jitter finely every frame. The white hand collapses while transforming into a small yellow star, which flies toward a laptop at the bottom right with a trail. The camera follows the star with a slight delay, panning sharply to the bottom right.53-6s:6The yellow star enters the laptop screen. Red hand-drawn lines scribble violently on the PC screen, instantly transforming into a red heart. The heart trembles and changes into a blue vortex, then vigorously jumps out of the PC screen. The blue line stretches like something between liquid and flame, and the camera quickly tilts upward to follow it. Natural motion blur from the sudden camera movement.76-10s:8The blue vortex turns into a light blue butterfly on the wall and flies toward the back of the room. The camera follows the butterfly while walking. The butterfly transforms into a pink snake-like line, crawling along the wall, door, and floor. In the center of the door, the pink line swirls and transforms into a small octopus-like hand-drawn monster. The monster looks at the camera for a moment and then flees toward the refrigerator. The camera hurriedly follows.910-13s:10Yellow hand-drawn light rays and arrows appear on the refrigerator door, and the door swings open. Inside the bluish refrigerator is a cute hamburger character with eyes and thin limbs. The next moment, the hamburger's mouth rips wide open, transforming into a comical monster with a red interior and triangular fangs. The camera shakes violently as if the filmer is startled, and the refrigerator door is quickly closed.1113-15s:12The camera moves backward to show the whole room. The outline of a giant white eye is drawn on the upper wall, and a vivid blue pupil appears inside. Immediately after the eye blinks slowly once, it transforms into a giant pink and white smile with many triangular teeth. The camera tilts in surprise, and the screen goes dark to end.13The hand-drawn animation is rough frame-by-frame animation, not smooth CG. Line thickness, outlines, and coloring constantly jitter slightly, retaining the texture of crayons, chalk, colored pencils, and rough brushes. Vivid glowing colors centered on white, yellow, light blue, pink, and orange. There is a weak glow around the animation, but it does not illuminate the live-action space too strongly.14The camera always follows slightly behind the animation. The animation moves out of frame first, and the handheld camera pans, tilts, or moves forward in that direction with a delay. Perfectly stable gimbal shots are prohibited. Do not use too precise 3DCG, too smooth vector lines, cinematic lighting, excessive depth of field, perfectly composed commercial photography, subtitles, logos, or BGM. Only include indoor sounds, rustling clothes, footsteps, small PC operation sounds, the opening and closing of the fridge, and light electronic sounds and comical cries from the hand-drawn creatures.

Actual ChatGPT chat screen
Now, just take this prompt and use the "Text to Video" feature in the video generation AI "Seedance 2.0"!
For more about Seedance 2.0, please refer to this Note article:
[Click here for the Seedance 2.0 article]
https://note.com/genel/n/n3aac3300f211?sub_rt=share_sb
After copying and pasting the prompt, setting the generation time to 15 seconds, and generating the video, it looks like this!

It turned out great! I added BGM generated with SUNO afterward.
How to Arrange Different Versions
This is the part I want to recommend the most this time 💡
You might be thinking:
"I can now make the exact same video as GENEL, but what should I do if I want to make a different video?"
📢 So, I've prepared instructions that allow you to arrange and mass-produce versions!!
Just by copying and pasting these instructions, you can generate videos like these one after another!


I'm introducing them in this Note article, so please check it out if you're interested ✨





