Step-by-Step Guide to Replicating Viral Douyin Videos with Codex + Hypit

@Pluvio9yte
الصينية16 سبتمبر 2026
204K
468
92
29
937

ليرة تركية؛ د

A comprehensive guide on using the open-source Hypit engine combined with Codex to replicate viral short-form videos. It details setting up AI agents for asset generation, editing, and compositing.

What is Hypit?

Hypit is an open-source project that enables AI Agents to create videos. You provide reference videos, character images, and requirements to Codex, and the Agent can use Hypit to organize asset generation, editing, subtitles, and picture-in-picture effects, finally exporting the video.

It also leaves behind an editable project file. If you want to change subtitles or adjust picture-in-picture later, you can reuse the generated assets and continue modifying them.

Open Source Address: hypit-ai/hypit

First, look at the finished product. On the left is the original video by "Chao Ge Shuo Che" (Brother Chao Talks Cars), and on the right is my replica using the character "Han Li".

雪踏乌云 - inline image

This was made using my own Minimax H3 API. If you use the official API or SeedDance, the results will be even better.

Below, we start from installation and walk through the entire operation completely. Model configuration has two paths: use built-in services if you have Hypit credits, or connect your own API if you don't. This example actually follows the self-owned API path.

  1. Clarify Roles: What do Hypit, Codex, and Generation Models each handle?

We need to first clarify the specific responsibilities of each tool in the pipeline to avoid confusion:

  • Codex's Role: As a code and engineering execution assistant, Codex parses natural language user requests, checks and configures the local development environment, reads and breaks down the storyboard structure of reference videos, generates and modifies project source files, and organizes calls to underlying engines to complete automated pipelines.
  • Image & Video Models' Role: Specifically responsible for generating static visual assets and dynamic scenes. In this example, the image model generates the first frame of an ancient-costume man in a modern livestream room, while the video model drives the first frame into a talking streamer animation and generates external traffic shots based on prompts.
  • Hypit Engine's Role: As the core orchestration center, Hypit records all generated asset dependencies, defines shot sequences and transition clocks, controls subtitle display timing and appearance, manages picture-in-picture position, layering, and rounded corner masks, and calls the underlying renderer to composite the final video once assets are ready.

To accommodate different readers' actual conditions, this tutorial explicitly splits generation model configuration into two independent paths:

  • Path A: You have available credits in your Hypit account and directly call built-in hosted models via the official HypiHub service.
  • Path B: You do not have Hypit platform credits and explicitly configure and connect your own third-party API interfaces.

You only need to choose one path (A or B) that fits your conditions to complete configuration, then merge into the common production and orchestration workflow.

  1. Preparation and Asset Specifications

Before starting, prepare Codex, a clear reference video, and a target character reference image you wish to substitute in.

Character reference images should ideally be single-person photos with distinct facial features, even lighting, and clear clothing/hair details. Since reference images contain no motion information, do not assume the model will automatically replicate the original host's shrugs, hand gestures, etc., from a static image alone.

If you plan to publish the result publicly and do not want to use the original host's voice, record replacement dialogue audio before formal production. Changing voiceover audio alters pronunciation pauses and segment durations, requiring re-alignment of cuts and subtitle timing.

  1. Installation

Execute the official Skill global installation command:

npx skills add hypit-ai/hypit -g

Refer to the Hypit Repository for installation entry points, and the Official Quickstart Manual for detailed process steps.

Path A: With Hypit Credits, Using Built-in Services

If you have credits, you can let Codex use Hypit's built-in services directly without configuring your own API.

Please use Hypit's built-in services to help me log in, check available models and credits. Tell me the estimated consumption first, then start generation.

Path B: Without Hypit Credits, Using Your Own API

I followed the self-owned API path this time: using Google for image generation and ZenMux's MiniMax H3 Max for video generation.

I first had Codex configure the interfaces, confirmed the models could be called, and then continued production.

Use Google API for image generation and ZenMux's MiniMax H3 Max for video. Please help me complete configuration and check availability. Explain costs first if paid testing is required.

雪踏乌云 - inline image

Give Original Video and Han Li Image to Codex

Next, I sent the link to "Chao Ge Shuo Che"'s video and the Han Li image together to Codex, specifying to replicate only the first 20 seconds.

Please replicate the first 20 seconds of this video, replacing the character with the Han Li image I provided. Preserve the original composition, cut rhythm, subtitles, and picture-in-picture. Keep the original audio. Complete all operations within the same Hypit project.

Be clear about the replication scope; otherwise, for a 10+ minute original video, Codex might plan for the entire length.

雪踏乌云 - inline image

Check First Frames Before Generating Video

I had Codex break down the shots and generate four first frames: Han Li in the livestream room, sunset city traffic, white Mercedes in tunnel, and daytime street traffic.

This step mainly checks if the character looks right, clothing is correct, and scenes haven't drifted. If first frames are unsatisfactory, modify them first before continuing to generate dynamic video.

Please show the first frames of the four scenes first. Preserve Han Li's face, long hair, and blue-gold robes. Car shots should not include the original host, subtitles, or picture-in-picture; these will be added uniformly by Hypit later.

[Insert Han Li Livestream Room First Frame Here]

Have Codex Generate Segments, Then Composite with Hypit

After confirming first frames, I had Codex generate four host segments and three car segments, requesting 5 seconds each, then used Hypit to composite the final 20 seconds.

Models generate visuals; Hypit handles cuts, subtitles, and picture-in-picture.

Please generate dynamic assets according to confirmed first frames and budget, then composite a 20-second video following the original rhythm. When car shots appear, place Han Li in a centered bottom picture-in-picture, with original audio and subtitles. Reuse already generated assets directly; do not regenerate.

雪踏乌云 - inline image

This time, from reference video to Han Li version final cut, my main tasks were giving requirements to Codex, reviewing results, and telling it what needed adjustment.

Hypit leaves behind not just a video, but an editable project. Next time you want to change characters, edit subtitles, or adjust picture-in-picture, you can continue from this project.

Of course, current actions and lip-syncing haven't reached 1:1 replication yet; using better models like the Seedance series would improve this significantly. If interested, try finding a short 10-20 second video first to see which steps it saves you.

Project Address: hypit-ai/hypit. If you find it useful, consider starring the project.

ريمكس في YouMind

قم بتحويل مقال سريع الانتشار إلى سير عمل كامل المحتوى

قم بتجميع المصدر وفك تشفير النمط وإنشاء الأصول وصياغة القصة وتوزيعها من مساحة عمل واحدة تعمل بالذكاء الاصطناعي.

اكتشف YouMind
للمبدعين

حول Markdown إلى مقالة 𝕏 نظيفة

عندما تنشر كتاباتك الطويلة، فإن الصور والجداول وكتل التعليمات البرمجية تجعل تنسيق 𝕏 مؤلمًا. YouMind يحول مسودة Markdown كاملة إلى مقالة نظيفة وجاهزة للنشر 𝕏.

حاول Markdown إلى 𝕏

المزيد من الأنماط لفك التشفير

المقالات الفيروسية الأخيرة

استكشاف المزيد من المقالات الفيروسية