Previously, making a video—scripting, voiceover, editing, and subtitling—each required a significant amount of time to learn, often leading people to give up before they even started.
Now, these tasks can be handed over directly to Codex.
You give it an article or just a theme, provide the corresponding video Skills, and Codex can complete the spoken script, voice, storyboard, subtitles, motion effects, and the final video. You don't need to learn editing first, nor do you need to understand code; as long as you can clearly describe the video you want, you can begin.
In this article, I will explain the entire process from start to finish. We'll start by making a video from an article without showing your face, and then add the editing method for real-person voiceovers. Even someone new to video can follow along to produce their first version.
My workflow is: Article or Theme → Spoken Script → Voiceover → Timeline → Storyboard → HyperFrames → Check Final Video.

This tutorial starts with the simplest path. You only need to know how to open Codex, create a folder, and put the article in; you can tell it what to do sentence by sentence for the rest.
First, Hand Over the Necessary Tools to Codex
A Skill can be understood as a job instruction manual for Codex. It tells Codex what tools to call, in what order to complete them, and what issues to check for when encountering voiceover, storyboard, or editing tasks.
The installation method can be unified into one sentence: Copy the GitHub link below to Codex, and let it read the project description, complete the installation, and check if it's usable. Leave the specific commands to Codex; you don't need to copy them yourself.
The tools used in this video workflow can be viewed by stage:

HyperFrames already includes several roles: /faceless-explainer is responsible for organizing articles or themes into non-face-showing explanations, /hyperframes-creative handles the voiceover rhythm and storyboarding, /media-use manages images, sound, and subtitles, and /hyperframes-animation handles visual changes. Usually, you just need to tell Codex "use /hyperframes to make this video," and it will select the necessary parts.
Visual styles can also be added as needed. If you want to make a collage-style knowledge video, you can give Codex the vox-explainer from HyperFrames Community Skills; if you want handwriting or drawing effects, choose p5-paint-animation from the same repository; for product introductions, interface demos, or promotional shots, let Codex refer to video-shotcraft. These are style extensions and don't all need to be installed for your first video.
After putting the tools into the workflow, Codex will first perform content editing, then handle sound, storyboarding, and final video checking. You only need to confirm if the results at each step match your intent.

Now create a new folder, for example, codex-video, and open it in Codex. Put your article, images, and the audio generated later into it, then give the HyperFrames repository link to Codex to complete the preparation for the first main line.
Why Some Article Videos Bring Continuous Traffic
Serena made a very specific attempt: instead of talking vaguely about "AI can make videos," she chose the "book breakdown" content niche and used Codex and HyperFrames to create videos suitable for Douyin.
Readers click in because they see a complete path: choose a niche people already watch, find a book, write a script, make a video, and repeat the same structure for the next book. Video production is just the execution method in the middle; what truly grabs people at the start is "can this method help me grow my account."
Her actual production sequence is also worth following: first determine the 30-second duration, vertical format, and publishing platform; write the script separately; split the script into segments; generate the final voice; get subtitle timing based on the voice; then let HyperFrames arrange the visuals; preview, modify, and finally render.
There is a very important order here: Finalize the words and the final audio first, then do the visuals. If you change one sentence of audio, the subtitles and shot timing will all change. Doing motion effects too early only leads to repeated rework later.
You don't have to make a book account. You can break down AI tools, explain work methods, talk about industry cases, or turn your long articles into video series. Fix the content scope first and use the same structure for every video so Codex can help you work faster and faster.
Step 1: Turn the Article or Theme into a Spoken Script
The spoken script is the skeleton of the entire video. No matter how beautiful the visuals are, if the beginning doesn't grab people or the method isn't clear, the video will be hard to finish.
There are three paths here.
Option 1: Only a Theme, Let Codex Write Directly
For example, if you plan to make a video on "How Codex Makes Videos," tell Codex four things: who the audience is, what problem it solves, how long the video is, and what you want the audience to do after watching.
1I want to make a 60-second vertical video with the theme "Using Codex for the first time to make a video."23The audience only knows how to use Codex and has never written a script or edited a video.4Start by telling them the final result, then explain clearly: generating the script, choosing a voice, using HyperFrames for visuals, and checking the final video.5Each sentence should be easy to read aloud, no formal subheadings, no jargon stacking.6End by asking them to run their first video with a theme.78Deliver only the spoken script for now; do not do storyboarding or video yet.
The more specific the theme, the easier the script is to write. Narrowing "making AI videos" to "turning a tool review into a 60-second vertical explanation" tells Codex what content to keep.
Option 2: Already Have an Article, Let Codex Adapt It into Spoken Language
Save the article as article.md and put it in the project folder along with the images. Then tell Codex:
1Please read article.md and adapt it into a 90-second spoken script.23Keep the core insights and most useful cases from the article; do not just summarize every paragraph.4Start by saying what the reader will get, keep only three steps in the middle, and explain each step in easy-to-understand sentences.5Do not read out links, footnotes, or formatting symbols from the article.6Please list the facts I need to verify separately.78Generate only narration.md for now.
When converting articles to scripts, the most common problem is that "the words are recognizable, but they don't sound like a person speaking." Give the link for humanizer-zh to Codex, let it install it, and process narration.md: keep facts and personal insights, but adjust translationese, mechanical parallelism, and overly formal sentences.
After adaptation, you still need to read it aloud yourself. Your mouth will find many problems your eyes missed.
Option 3: Let HyperFrames Take Over the Entire Faceless Video
If you only have materials and haven't thought of a script or visuals, you can enter /faceless-explainer via /hyperframes. It will organize themes, articles, or notes into an explanation direction, narration, and storyboard.
This path is easier and suitable for quickly making a first version. The first two options are better when you have clear viewpoints and want to strictly control the script content.
Regardless of which you choose, the script should ideally have these four parts:

- Hook: Present the result or a specific problem first.
- Problem: Tell the audience where they are currently stuck.
- Method: Explain the process in two or three steps.
- Next Step: Give an action they can take immediately after watching.
For example, the video intro for this tutorial could be: "Installed Codex but don't know what to do? You can start by using it to make a video. Give it an article, and it can handle the script, voice, subtitles, and visuals. Today, let's run through the shortest workflow."
Step 2: Choose Standard TTS or Clone Your Own Voice
Once the script is finalized, decide who will read it. For the first time, standard TTS is sufficient; once you can update content stably, then handle personal voice cloning.

Easiest: Use Existing TTS
TTS is inputting text to get an MP3 or WAV. You can use any text-to-speech function you have on hand, or give the open-source edge-tts to Codex, let it install it, select a Chinese voice, and generate audio and subtitles.
Listen to ten seconds first to confirm the voice, speed, and pauses before generating the full text. Standard TTS won't sound like you, but its advantage is speed, making it suitable for verifying content and visual workflows.
For Long-term Use of Your Own Voice: CosyVoice
CosyVoice is an open-source voice generation toolset. Give it a clean reference recording and the corresponding word-for-word text, and it can generate new narration with similar speaker characteristics.
Three things must not be mixed up here:
- reference.wav: Your own or an authorized reference recording.
- reference.txt: Every single word actually spoken in the reference recording.
- narration.md: The new script to be generated for this video.
The reference text must match the recording exactly. If you change a word on the fly, the text must be changed too. Background music, echoes, and heavy noise reduction will affect reference quality; try to record a clear, natural voice in a quiet environment.
Local installation of CosyVoice takes a few more steps than standard TTS. Beginners can give the official repository directly to Codex:
1Please help me install CosyVoice:2https://github.com/QwenAudio/CosyVoice34Please read the official README first, then check the current computer environment and complete the installation.5After installation, generate only one short test audio; do not generate the entire script directly.
It will prepare the running environment and download voice models locally, which takes longer than standard TTS. You don't need to execute commands one by one; if a step fails, leave the full error message in the Codex chat and let it check. After installation, you still need to listen to the test audio yourself to judge if the voice sounds like you and if the articulation is natural.
When generating the formal script, do not stuff the entire article in at once. Split it into several segments based on meaning, listen to each segment first to confirm the pronunciation of proper nouns, English, and numbers, then synthesize final.wav. If one segment has an issue, you only need to redo that segment.
If you don't want to clone your voice for now, just skip CosyVoice. The subsequent HyperFrames and Remotion only recognize the final audio file and don't care which TTS it comes from.
Step 3: Make the Visuals Follow the Final Voice
Now you should have two finalized files: narration.md and final.wav. From this moment on, final.wav is the true timeline of the video.

Let Codex transcribe the final audio to get the actual start and end times for each sentence, then generate subtitles. The script determines what the subtitles say, while the audio transcription is responsible for telling the subtitles when to appear.
Do not allocate time based on word count. Ten words might be spoken in half a second, or there might be two pauses in between. Subtitles, keywords, and images should all appear following the real voice.
You can simply instruct it like this:
1Please establish a timeline based on final.wav.2Use narration.md to verify subtitle text and use audio transcription results to get the real start and end times for each sentence.3Output captions.srt and timing.json.4If the script doesn't match the actual voice, please list the discrepancies; do not guess.
Step 4: Split the Script into a Storyboard
Storyboarding doesn't require learning professional terminology first. Cut the script into 4 to 8 segments, and for each segment, answer one question: What is the most important thing for the audience to see on screen when they hear this sentence?
If the article already has screenshots, photos, charts, or character cards, prioritize using these materials. Let the audience see the big picture first, then zoom in on the part being discussed. For sentences without existing images, you can use keywords, step cards, comparisons, timelines, and simple infographics.

The requirements for Codex can be very short:
1Please create a storyboard based on narration.md, timing.json, and the assets folder.2For each shot, specify: corresponding narration, start time, end time, main visual, required materials, and subtitle position.3Images already in the article must be scheduled for actual display; mark places where materials are missing first, do not fill them with irrelevant images.
If you don't know what motion effects look like, you can check two public effect libraries:
- HyperFrames Launches: Contains complete HyperFrames projects to see how final videos are broken into storyboards.
- video-shotcraft: Provides Remotion shot recipes and dynamic previews, better for product intros, interface demos, and promos.
video-shotcraft is an optional shot reference library, not a mandatory installation for article videos. Give the repository link to Codex, and it can read the shot recipes inside. Explain the information clearly first, then pick one or two suitable motion effects; the visuals will be more stable.
Step 5: Create the First Version of the Video with HyperFrames
By now, the script, voice, subtitle timing, and storyboard are all ready. Now hand these files over to HyperFrames:
1Using /hyperframes, turn the current project into a 9:16 vertical explanation video.23Please read narration.md, final.wav, captions.srt, timing.json, storyboard.md, and the assets folder.4The visuals should follow the real audio timing, and subtitles should not block main content.5Prioritize using original images and character cards from assets; use simple infographics for shots missing materials.6Complete the first 15 seconds as a sample and open the preview; I will continue with the full video after confirming the direction.
Watching the first 15 seconds allows you to catch three problems early: is the intro too slow, are the subtitles legible, and are the visuals actually explaining the narration. Once the direction is right, continue with the full video.
After the project is generated, tell Codex directly: "Open the HyperFrames local preview and tell me the preview URL and project folder."
When modifying, don't just say "make it more high-end" or "make it look better." Tell Codex specific objects and positions, such as: "The title at 6 seconds is too small," "This screenshot should stay for at least 3 seconds," "Subtitles are blocking the person," "The first shot enters too slowly." This feedback is easier for it to execute.
After checking, tell Codex: "Render the currently confirmed version as final.mp4, and check if the file plays normally after completion."
Seeing the preview page only means the project can open. The true sign of completion is that final.mp4 has been generated and you have watched it from start to finish, confirming the voice isn't cut off, subtitles aren't misaligned, and images aren't cropped.
How to Choose Between HyperFrames and Remotion
Both can turn text, sound, images, and timelines into videos. You just need to pick one as your final production tool.


HyperFrames is more like letting Codex reorganize visuals based on this specific content. Remotion is more like making a show template first and then filling content into fixed positions for every episode.
To go the Remotion route, give the link for Remotion Agent Skills to Codex and let it complete the installation. Then use /remotion-create to build the video, /remotion-studio to open the preview, and finally /remotion-render to render. The upstream narration.md, final.wav, subtitles, timeline, and images can all still be used; you don't need to redo them.
My suggestion is simple: run the first one with HyperFrames. After making a few episodes and finding the shot structure has stabilized, let Codex organize it into a fixed Remotion column.
If You Filmed a Real-Person Voiceover, Add video-use
The previous workflow can be completely faceless. When you are ready to face the camera, put narration.md next to the camera and film in segments. If you make a mistake, stop and restart from that sentence; you don't have to start from the beginning every time.
video-use handles already filmed materials. It first transcribes the video, finds out what was said in each segment, and confirms with you which content to keep before starting to organize editing, subtitles, and supplementary visuals.

For installation, you can give this directly to Codex:
1Please help me install https://github.com/browser-use/video-use.2Read install.md first, complete dependencies, FFmpeg, and Codex Skill registration.3Do not transcribe any videos after installation; wait until I put the materials in the folder.
video-use requires installing the entire repository because its editing assistant and scripts are located next to the Skill. It also needs FFmpeg, which is the underlying tool for processing video and audio. When using default voice transcription, you also need to configure an ElevenLabs API Key, which is the key that allows the tool to access the voice transcription service. Just hand the missing items to Codex to check.
After filming, put the raw video in a separate folder and tell Codex:
1Please use video-use to process the real-person voiceover in this folder.2Inventory the materials and transcribe them first, finding repetitions, pauses, mistakes, and restart positions.3Give me an editing plan first, listing segments to keep and delete; I will confirm before you cut.4Add Chinese subtitles after finalizing, and use corresponding images from assets to supplement visuals when steps are mentioned.5Finally, output a preview version, and output final.mp4 after checking for errors.
The complete workflow for this branch is: Spoken Script → Real-Person Filming → video-use Organizes Valid Segments → Subtitles and Supplementary Visuals → Final Video.
HyperFrames and video-use can also be used together. video-use first cuts the real-person voiceover smoothly, and then HyperFrames produces titles, step cards, screenshot demos, and intros/outros. One handles real-person materials, the other handles informational visuals; the division of labor will be clearer.
Finally, Turn Video into a Content Acquisition System
A single video only brings one test. The truly valuable part is continuing to address the same type of problem.
Serena's book breakdowns are memorable because the content object is very clear: a book, a viewpoint worth sharing, a short video, and the next book uses the same structure. You can replace this with your own field, such as breaking down an AI tool every week, solving an office problem every time, or explaining an industry case in each episode.
This content loop can be very simple:
1Target user's problem2→ Choose a specific theme3→ Write into a spoken script4→ Make into a video5→ Look at comments and completion feedback6→ Decide what to talk about next
The hook for customer acquisition should also land on specific results. "I made a video with AI" only shows you can use tools; "I turned a 3000-word tool review into a 60-second beginner tutorial" will make people who truly need that thing stop.
Don't rush to build a massive automated system. Choose a theme you can talk about for ten consecutive episodes and run the first one with Codex: finalize the script, generate the voice, align the timeline, and produce the video with HyperFrames. Reuse the file structure for the second one, and start keeping fixed intros and subtitle styles for the third.
By doing this, Codex truly transforms from "helping you make one video" to "accompanying you in continuously producing a category of content."
I am Miles, an AI algorithm expert who transitioned from a big tech company to an FDE. I have done algorithm R&D, optimization, and deployment, as well as corporate training and delivery. My X account reached 10,000 followers in 15 days and wrote three long articles with millions of exposures in one week, two of which later exceeded 2 million.
Follow me @miles_mazy, let's grow together and make money together.


![Ending Document Creation with AI: The Ultimate Prompt and 61 Consultant-Quality Samples [Must-Save Edition]](/cdn-cgi/image/width=1920,quality=90,format=auto,metadata=none/https%3A%2F%2Fcms-assets.youmind.com%2Fmedia%2F1788886306017_s470xy_HRr-Dl1boAE3UpK.jpg)



