How to Replicate a Viral 700k-Follower Douyin Account Using AI: A Complete Tutorial

@leo_xiaolei
УПРОЩЁННЫЙ КИТАЙСКИЙ21 авг. 2026 г.
406K
1.1K
111
19
1.6K

Суть

This technical guide breaks down the workflow for replicating viral AI-generated short videos, focusing on character consistency cards, frame-by-frame action deconstruction, and realistic clothing physics prompts.

This story stems from a recently viral Douyin account: "Bu Chi Xia Fan Cai" (Doesn't Eat Side Dishes).

As of August 21, 2026, the homepage has only 36 works, yet followers have reached nearly 700,000.

The core is a small action that has contrast, is slightly curious, and keeps people watching. Each video only changes one creative point, while keeping everything else as static as possible.

What I want to verify is: can we use an AI character to reverse-engineer and replicate this content production method?

Character Cards

If a video is only made once, random character generation might work.

However, if our goal is to build an IP, the character must be recognizable at a glance.

Short legs today, long legs tomorrow; cold white skin in this one, yellowish arms in the next; chest shape concentrated one moment, outward-expanding the next—this makes every video feel like a random draw.

Character cards are used to end this randomness.

Regarding character cards, I already have a complete tutorial, see: https://x.com/leo_xiaolei/status/2077410771311468618

A single character card is sufficient for general image and video generation, but for the requirements of building an IP, a single character card has significant limitations—namely, insufficient detail. For AI, it's not comprehensive information, but rather information overload.

Conversely, synthesizing a character image for the current scene through multiple detail images is much more deterministic.

Therefore, character consistency is not one step, but two.

"Duo Rou's" character card went through quite a bit of rework. Initially, the upper body was too long and the legs too short; after lengthening the legs, it was too much, and users asked us to find a middle ground. Later, we decided to keep them longer, finally fixing it at 172 cm, 61 kg, and a 7.75 head-to-body ratio. Shoulders are narrow and naturally sloped, the abdomen is relaxed, and the upper arms, waist, hips, and thighs retain normal soft tissue—it shouldn't look like a hard, dry fitness model.

If a video is only made once, random character generation isn't unusable. But our goal is to change only one new creative point per video; the character must be recognizable at a glance. If body proportions, skin tone, chest shape, and hairstyles jump around, the account has no fixed persona, and every video is like the first draw.

Character cards end this randomness. They aren't responsible for making the character "pretty"; they are responsible for answering things every future video will encounter: who is this person, what are the body proportions, is the skin the same texture in different parts, where do the chest shape and volume go, what is the default hairstyle, and what can or cannot change during movement.

Chest shape descriptions like "large breasts" are useless; the model might make them outward-expanding, side-stacked, hollow in the middle, or just two hard spheres. Duo Rou's design goal is 98/73, 75F, concentrated, forward-facing, natural teardrop shape, retaining central volume without shrinking overall volume just to fix outward expansion.

刘多肉 - inline image

Benchmark Video Deconstruction

Taking the account's classic video—lifting the shirt—as an example:

I view the video in three layers:

The complete upward roll reference is split into ten segments:

Production Knowledge

At this point, you'll realize AI video isn't just about writing prompts. You must simultaneously handle several types of knowledge.

The first is camera and composition. You need to know why the frame cuts from neck to low waist, why the character must face the camera, where the window light comes from, and why a fixed mobile position shouldn't suddenly zoom. If the camera moves, the relative position of hands and clothing coverage changes, making frame-by-frame comparison meaningless.

The second is character design and lighting. Body proportions, skin continuity, chest direction, soft tissue, breathing, and hairstyle all belong to the character card. Users saying "it doesn't feel like a real person" won't be solved by adding "photorealistic." If skin is too flat, posture too stiff, left and right perfectly symmetrical, and there's no inertia in hair and soft tissue, the image still looks like a deforming poster.

The third is clothing patterns and fabric movement. Material only determines part of the result; where it covers, which edge the hand grabs, which part of the fabric moves first under force, and which layer is outside after flipping—these wearing relationships are more important. A bit more length, a bit tighter, or a higher hem can completely change the movement.

The fourth is action description. Actions shouldn't just be verbs; they need the actor, contact object, path, intermediate state, and end state. "Hands lift up" is too little information; "Hands hook the lowest left and right hem from below, first establishing outward tension, then rolling the entire lower half sequentially over the lower abdomen, navel, and upper abdomen" has checkable value.

The fifth is editing. Generated videos are often longer than the target, and action speeds won't match perfectly. Editing must involve finding action boundaries, segmented speed changes, fixing frame rates, and padding the final static frame, rather than uniform acceleration.

Finally, acceptance. A playable file only means technical success; passing the action table means internal checks are done; the user feeling "Yes, that's the feeling" is the final acceptance. These three things cannot be mixed.

Video Timeline and Script

After merging the main video's rhythm with the auxiliary video's complete upward roll action, the final script is six segments, exactly 218 frames:

If written as a standard script, it looks like this:

text
1Scene: Daytime, indoor light gray wall, natural light from right window. Fixed mobile horizontal close-up.
2
3Character: Duo Rou, fictional adult AI character. Full face always out of frame, end of high ponytail visible.
4
5Opening: Character stands facing forward, arms hanging naturally. Long blue and white print tube top covers to low waist.
6
7Main Action: Hands approach hem, hook lowest edges from below, pull outward first, then roll the entire lower half upward continuously. Fabric passes lower abdomen, navel, and upper abdomen sequentially. Abdomen fully exposed, long tube top becomes short tube top.
8
9Pause: Hands release and drop. Clothing maintains finished state. Character static for approx 0.9s.
10
11Closing Action: Right hand passes abdomen once, then exits frame. Left hand remains still.
12
13Ending: Character static facing forward, retaining only breathing and ponytail settling, until frame 218.

How to Describe Fabric

In AI videos, clothes easily look like plastic. A separate fabric description is needed: Cotton/Modal based, approx 8% Spandex, 170–190 g/m² lightweight single jersey; overall matte, fine knit particles and fiber micro-reflections visible up close; low bending stiffness, medium rebound. It can be form-fitting but not thick like neoprene or thin like stockings.

The physical process during rolling must also be written:

  1. Fingertips press into the bottom hem first, creating local depressions at contact points.
  2. Hands move up, tension spreads from fingertips upward and sideways, creating radial fine wrinkles.
  3. Fabric near the hands moves first, middle area lags slightly; no rigid planar translation.
  4. After passing the abdomen, the hem flips inward, forming two real fabric layers that can slide against each other, not merging into a thick tape.
  5. After hands release, fabric rebounds slightly before stabilizing over a few frames; no instant transition to a perfectly straight geometric band.

Also explicitly exclude latex, PVC, rubber, neoprene, glossy films, hard shells, and uniform mirror highlights. Negative prompts can't replace positive descriptions, but they prevent the model from using the easiest plastic texture to explain "tight" and "elastic."

Prompts

The role of prompts is to summarize what has already been determined, not to reinvent characters, clothes, and actions at the last minute. I now write in the following order, usable with any video tool:

text
1[Subject Declaration]
2This is our own designed fictional adult AI character, not a real person.
3
4[Input Responsibilities]
5Character image is responsible for proportions, skin tone, hairstyle, clothing, background, and camera.
6Action reference is responsible for action sequence, hand path, speed, pauses, and fabric movement.
7Do not copy the identity, face, or body of the person in the action reference.
8
9[Camera]
1016:9, frontal fixed mobile position, shot from below neck to low waist.
11Light cool gray wall, right-side window light. No zoom, no pan, no cuts.
12
13[Character]
14Include confirmed proportions, posture, skin, chest shape, and high ponytail from the character card.
15Shoulders, abdomen, and arms remain relaxed, not a fitness physique.
16
17[Clothing Start State]
18Clearly state pattern, length, coverage, material, hem position, and inner/outer layer colors.
19
20[Action Timeline]
21Write by time: approach, grab, establish tension, continuous roll, organize, release, pause, right hand passing abdomen, and final static.
22Write hand position and new clothing state for each segment.
23
24[Clothing End State]
25Entire lower half has entered chest area, abdomen fully exposed.
26No thick rings, hanging panels, or extra layers left under chest.
27
28[Secondary Motion]
29Retain breathing, soft tissue inertia, delayed ponytail swing, and stray hair settling.
30
31[Prohibitions]
32No new actions, no sequence changes, no repetition, no reverse play, no cuts.
33No plastic fabric, clothing teleportation, limb growth, finger clipping, or full face in frame.

"Keep character consistent," "replicate original video actions," and "make clothes more realistic" are sentences that look correct but are actually unverifiable. Change them to visible frame states: shoulders can't suddenly widen, chest volume can't shrink, skin can't change color, fabric must pass the navel, and no thick fabric ring can hang under the chest at the end.

Secondary Upgrades

After the first version was released, feedback was specific: chest looked smaller than the character card; didn't feel like a real person; body was too athletic, lacking visual appeal; clothes looked plastic after flipping; high ponytail lacked the effect of the long hair in the original video.

Later, the character card revealed three more issues: facial lighting was weird, like local white spots; chest shape was outward-expanding; legs were still not long enough. After lengthening, it was too much; the user paused generation, requesting a return to the middle of the previous two versions, finally re-confirming the 7.75 head-to-body ratio.

These feedbacks weren't all stuffed back into the video prompt but handled by problem category:

The first candidate to complete the full upward roll had passing clothing movement, but the character was still too thin and shoulders/arms too stiff. We didn't change the action table, just regenerated the scene character image and cleaned up conflicting terms in the character description.

刘多肉 - inline image

The second candidate continued using the same action set, only comparing character and material without re-discussing action sequence. Character, skin, and fabric all improved.

刘多肉 - inline image

Adding Background Music

This video needs a whistling song. We didn't reuse the original audio but used our own synthesized whistle melody.

The processing order: crop to target duration, reset time, apply 550 Hz high-pass and 3600 Hz low-pass, add 0.08s fade-in and 0.35s fade-out, then loudness normalization and limiting, finally mixing to AAC 48 kHz stereo.

Sound must be placed after action editing. If you lock the music to a 10s clip first, you'll have to reposition the music every time you adjust an action boundary—that's just making extra work for yourself.

刘多肉 - inline image

Final candidate is 1280×720, 30 fps, 218 frames, 7.266667 seconds. Video and audio have passed full decoding. Some differences remain: waist and collarbone are slightly thinner than the 61 kg character card, blue and white print color is lighter, ponytail secondary swing is weaker than the long hair in the benchmark, and individual hand trajectories aren't pixel-perfect.

Complete Production Process, Compressed into a Checklist

I will continue to follow this order for the next video:

  1. Identify the content structure used repeatedly by the benchmark account, then determine only one new creative point for this video.
  2. Fix the main reference video, confirming real video ID, duration, frame rate, and total frames.
  3. Export all frames and contact sheets, separately recording camera, body movement, and clothing state.
  4. If a local action is unclear in the main video, find an auxiliary reference just for that one action mechanism.
  5. Accept the formal character card first, then generate the scene character image specific to this video.
  6. Write clothing patterns, coverage, wearing relationships, static materials, and movement processes.
  7. Write the action as a timeline with frame numbers, then a video script readable by normal people.
  8. Prompts only summarize confirmed content, no new character or clothing settings at the last minute.
  9. Each generation only verifies one category of problem; keep failure contact sheets for single-item comparison with the previous round.
  10. After action, character, and fabric pass frame-by-frame inspection, perform segmented speed changes and add original sound.
  11. Finally check resolution, frame rate, frame count, duration, audio, and full decoding before handing over for user acceptance.

This process turns "replicate this" into a set of checkable creative files. Video tools can change, but creative descriptions, character cards, scene images, fabric descriptions, frame tables, scripts, and acceptance forms don't have to.

Common Errors

How to reduce failures caused by faces and content moderation

To be clear, these methods reduce misjudgment and input confusion, not bypass platform moderation.

  • The character must be a fictional adult AI character you have the right to use. State clearly in the prompt: "This is our own designed AI character, not a real person."
  • Neither the reference character image nor the final composition should show a full face; retain only the body areas and high ponytail needed for the action. If the face isn't part of the creativity, don't let it add identity judgment and consistency burdens.
  • Character images are for characters; action videos are for actions. Don't let a real-person reference carry the fourfold responsibility of face, body, clothes, and action.
  • When using a platform's own image functions, regenerate the scene character image once before handing it to the video stage. This makes the input source more consistent with the current platform.
  • If tasks containing real-person frames fail repeatedly, don't just mechanically replace a few synonyms. Remove unnecessary real-person references, use your own AI scene images and frame-by-frame action tables; if it still fails, accept the current tool's capability or moderation boundaries and change the production route.

How to avoid the lower half of clothing not rolling above the chest

Don't just write the final position. The model will easily jump from the start to short clothes, using clothing deformation or teleportation to cover the middle.

You must write four intermediate states: lower abdomen starts to show, fabric passes navel, fabric passes upper abdomen, entire lower half enters chest area. The final frame must also satisfy: abdomen fully exposed, no thick ring under chest, no hanging fabric panels, and clothing not falling back.

After generation, look at the frame images first; don't be fooled by smooth playback. Many videos look like they completed the roll when played, but stopping at key frames reveals the clothes just shortened, melted, or still have a ring of fabric hanging under the chest.

If the prompt already clearly states the whole process but it still stops under the chest after several rounds, it's not a matter of adding ten more "musts." Switch to a production method that accepts full action references, or break the main action into more controllable stages. Prompts have boundaries and cannot replace movement capabilities the tool itself lacks.

Переделать в YouMind

Превратите одну вирусную статью в полноценный рабочий процесс создания контента

Собирайте источники, расшифровывайте паттерны, создавайте активы, пишите черновики и публикуйте контент из одного рабочего пространства ИИ.

Исследовать YouMind
Для авторов

Превратите ваш Markdown в аккуратную статью для 𝕏

Когда вы публикуете длинные тексты, изображения, таблицы и блоки кода, форматирование в 𝕏 становится мучением. YouMind превращает полный черновик в Markdown в чистую статью, готовую к публикации в 𝕏.

Попробовать Markdown для 𝕏

Другие паттерны для анализа

Недавние виральные статьи

Смотреть другие виральные статьи