AI edits the finished video: the whole pipeline, revealed
A viewer asked me yesterday: how exactly do you make your daily videos with AI? So I had CC generate a web-page workflow manual from my own project folder, and this episode walks through it.
Seven steps: topic selection, scripting with a voice DNA distilled from all my past scripts, voice-cloned narration, a digital human driven segment by segment, covers plus bottom-screen infographics, subtitle burn-in, and auto-filled copy for five platforms. The AI computes for one to two hours per video. A video costs tens of RMB—every step is a paid official API call. It can run fully automated, and it gets better with a human in it.
How it started: I lost my voice, so the workflow became a digital human
A viewer asked me yesterday how I use a digital human, CC and image-generation models to produce my daily short videos—especially the automated part. I’ve iterated this logic through many rounds, so this episode opens it all up. I had CC go into my video-production project folder, read the conversation history, the project memory and the delivery documents, and produce an up-to-date workflow manual as a visual web page in my own PPVI light style. The walkthrough below follows that manual.
The direct trigger: a cold plus an allergy wrecked my voice completely. I couldn’t record myself, so I rebuilt the whole workflow around a digital human presenter.
The prep list is short: record a clip of my own voice for cloning; find one continuous 60–90 second shot of my upper-screen presence from an earlier episode; and hand over every spoken script I’ve ever written. Then come the seven steps.
Step one: topics—I built a trending tool, and mostly don’t use it
Early on I built a local tool that picks topics from trending charts. Honestly, my videos rarely use its recommendations—not because they’re bad, but because every day I already have something I badly want to say, or something fans keep asking for. Sometimes the traffic is worse than an AI-picked topic would get. But it’s what I want to say and what they want to hear, and if I don’t make it, the group chat keeps asking.
Step two: scripting—voice DNA, plus a defense against the “AI flavor”
With the topic set, the script gets written against my voice DNA—the expression structure and style distilled from all my past spoken scripts.
Here I added an extension. I noticed that from Sonnet 4.6 onward, and the GPT-5-era models too, almost all of them think deeply in English and translate back into Chinese. That produces what I call the new generation of AI flavor—a touch of translationese. Natural Chinese says 流程循环; translationese says “the human is in the loop.” Feel the difference.
So the first draft is a dense stack of information—background material, cases, the whole logical thread written out in full—and then it goes to DeepSeek or Kimi K3 for a native-Chinese rewrite and review. Once that’s done I barely check it. Straight into audio generation.
Step three: voiceover—the verbatim script becomes a “TTS-generation script” first
The voice is already cloned; now TTS generates the audio. Since the whole line is automated, only API-callable TTS models qualify. By my testing, MiniMax’s 2.8 HD Speech is still the strongest machine-reading model at this point in time.
One key move here: I have CC convert the verbatim script into a “TTS-generation script.” What’s the difference? It pre-resolves the heteronym characters, pre-processes the numbers, and adds pauses and interjections in the right places—turning it into something a speech model understands and follows far more easily—then generates the full read-through audio.
Step four: the digital human—cut by pace and mood, stitched like it was edited
Same logic for the digital-human platform: API service only. I care about video clarity and lip-sync accuracy, so I chose the Shiliu digital-human API from Xiangliang Fangcheng.
CC cuts the voice track into segments: some parts want a faster pace, some want calm, some need a specific emotion mixed in. Each segment gets matched to different avatar footage of mine, generated one by one, then stitched. The result comes out looking like it was edited by hand—natural, with the expressions landing right.
Step five: covers and bottom-screen infographics—running in parallel
Covers come out in three aspect ratios for different platforms: 3:4, 4:3 and 16:9. After generation it automatically checks character consistency—whether the face still matches the reference photo I gave it.
Meanwhile, from the verbatim script it composes the polished bottom-screen visuals—slides or web pages. It slices time precisely: from second X to second Y, which visual should power the point I’m making. That’s HyperFrames (the HeyGen skill) for animated web-page graphics, or GPT’s Image 2 generating full-frame chalk-on-blackboard or clay-style B-roll illustrations.
One pitfall worth sharing: when you use Image 2, always have it generate the full frame—never a background with text overlaid afterward. That’s how the aesthetics stay consistent.
Steps six and seven: assembly and publishing
Once the material is rendered, it gets stitched, subtitles composed and burned in at a fixed position, and the date and my tagline stamped on like stickers.
Finally it writes every title, tag and description for all five platforms, opens the pages and fills them in automatically. I review, and I press the publish button myself.
Two honest numbers: time and cost
I won’t tell you the video is done in ten minutes. My own hands-on time in the whole process stays under ten minutes—but the AI needs one to two hours to finish all the actual computation. It has to think, write, generate images, composite. All of that takes time. And I won’t lie to you that one button press turns into hundreds of videos.
On cost: a video runs at tens of RMB, and every step is an official pay-as-you-go API call, paid for real. You can bring it down—swap in cost-effective models, like Gary in my group chat does at a few RMB per video, with decent results. Or you can spend freely: richer B-roll, higher image density, image-to-video B-roll—then one episode climbs to a few hundred RMB. So AI-made video spans a range from about ten RMB to a few hundred.
One more guardrail: before every step that calls a paid API, it runs a self-check first. Otherwise money gets wasted on doomed generations.
The full credits table
- Chief director: CC
- Voice: MiniMax
- Digital human: Shiliu
- Infographics: GPT Image 2
- Motion rendering: HyperFrames
- Subtitle alignment: Gemini 2.5 Flash
- Assembly and publishing: FFmpeg encoding + browser automation
The three questions fans ask most
Can it run fully automated, one click to a finished video? Yes. But for quality I still give feedback along the way: injecting my own views into certain passages, having it restructure the writing, suggesting layout tweaks on individual images after they’re generated. It works without a human. It works better with one.
Do viewers know it’s a digital human? Platforms now label AI-generated content. Unlabeled videos can still get traffic—some of my digital-human videos, including ones I made for other IPs, have gone viral at hundreds of thousands to millions of plays. But watch the review rules. On WeChat Channels, for example, when a video is about to break out—entering the 100k traffic pool—a human review kicks in. An unlabeled digital-human video gets its distribution restricted.
Can an ordinary person replicate this? Yes. Besides CC, Codex or WorkBuddy plus Kimi’s models can all do what I described. But every path requires serious time tuning the agent loop that executes it. I’ve been iterating this pipeline since two years ago. My own on-camera videos are also entirely AI-edited—that’s even simpler: no visuals to generate, just a rough cut, subtitles, publish.
Last thing: integrity, and money that’s clean
A lot of fans—good friends among them—forward me other bloggers’ videos and say: with standards that high you’ll never make big money; you have to fleece the audience, especially in paid knowledge. I could not disagree more.
Take the content I post on Channels—package it into a course, run paid traffic, sell it. You think I don’t know how? Of course I do. Internet advertising and marketing is literally my trade. But things that betray the delivery, I won’t do. I want to sleep soundly. Every yuan I earn stays clean.
Source: EP0072_audio.mp3 · ASR model gemini-2.5-pro (chunked parallel) · full text of the original video
[00:00] These days some AI bloggers, they like writing an obituary for their own account—maybe selling an ad alongside it too. I really don’t understand this kind of shameless behavior. These were the very first people to get their hands on the most advanced AI models and tools, and they gained a huge edge in knowledge and information from that. They were already in the paid knowledge business. And then they turn around and curse the very people who gave them an excellent product. And they even take that earlier case—a well-known company that got its free credits abused and ate the consequences—and use it to mock them. Standing on the abuser’s side, mocking the other party for being undignified. That’s like a convicted criminal—
[00:25] —at his own sentencing, telling the victim: that other girl never sued me, how undignified of you. Truly the textbook case of cursing your mother the moment you’re weaned. And it’s exactly this kind of fraud, this gray-market abuse, that got huge numbers of cross-border e-commerce sellers’ bank card ranges banned, their phone numbers unusable—it hurt a lot of people’s legitimate businesses. All I want to say is: can we hold ourselves to a higher standard? Honestly, speechless. Anyway, what I actually want to talk about today is the logic behind my video automation. A viewer asked me about it yesterday, and I’ve iterated this whole logic a lot. So today is really about the entire logic of making videos with AI. Once again, we’ll have Claude Code—
[00:50] —produce a web-page workflow manual based on my process. Here’s the prompt I gave it: here’s the situation. A fan of my short videos is asking me how I use a digital human and Claude Code, plus image-generation models, to produce my daily short videos—especially how the automated generation works. I’d like you to go into my short-video project folder, look through our conversation history, the project memory and the related delivery documents, and produce an up-to-date, comprehensive workflow summary.
[01:16] Present it visually and intuitively, using the PPVI light web style I designed myself. OK, we send the request off and wait for its reply. We’ll skip the part where the AI works, and I’ll take you straight to the deliverable in a moment. Alright, it’s done—let’s open it and take a look. Here’s how this started: a while back I had a cold plus an allergy, and my voice went completely. I couldn’t record videos myself anymore. So I simply iterated the whole workflow into one where a digital human does the presenting. Let’s walk through how the whole thing works—hopefully it gives you some inspiration.
[01:42] First, some material prep. For example, I record a clip of my own voice for voice cloning. Then I find a shot of my upper-screen presence from an earlier episode—one continuous take of roughly 60 to 90 seconds. And I give it all the spoken scripts from every one of my past episodes. With those materials ready, there are seven steps. Right at the start I actually built a local tool that picks topics automatically from trending charts. But to be honest, my videos rarely use the trends it recommends. It’s not that its picks are bad—
[02:08] —it’s that every day I already have something I badly want to say, or something the fans want to hear, so I just pick my own topic. Sometimes the traffic really is worse than an AI-picked topic would get. But it’s what I want to say, or what the fans want to hear—if I don’t make it, they just keep asking in the group chat, again and again, insisting I get it out as soon as possible. Once the topic is set—like I covered in a dedicated episode—my past spoken scripts get used to build a voice DNA of my expression structure and style. And here I actually made an extension. Because I noticed that from Sonnet 4.6 onward, and the GPT-5-era versions too—
[02:34] —almost all of these models do their deep thinking in English and translate the expression back into Chinese. This creates what I’ve called the new generation of AI flavor. It has a bit of a translationese feel. For example, natural Chinese would say 流程循环; translationese says “the human is in the loop”—in the loop. Feel the difference. The first draft is a very rich stack of information—background material, cases, the entire logical thread of the writing, all written out in full. It gets sent to models like DeepSeek—
[02:59] —or Kimi K3, for the Chinese-version writing and review. Once that’s written I basically don’t check it. Straight into the audio generation stage. As I mentioned, the voice is already cloned, and now TTS generates the audio. There are lots of options for this stage, but since ours is an automated pipeline, we choose the TTS models that can be called by API. By my testing so far—I said this in an earlier video too—MiniMax’s 2.8 HD Speech, at this point in time, is still the strongest machine-reading model.
[03:25] I have Claude Code turn that verbatim spoken script into a TTS-generation script. What’s the difference? It pre-processes the heteronym characters in there, pre-processes the numbers, adds pauses in the right places, adds interjections—turning it into a script that a speech-generation model understands more easily and follows more easily—and then generates the full read-through audio. Digital-human platforms, same logic: we pick the ones that offer API service. And within that, I also demand—
[03:50] —video clarity and lip-sync accuracy. So I chose the Shiliu digital-human API from Xiangliang Fangcheng to do this job. Then I have Claude Code cut the voice script into segments. Some parts suit a faster pace, some suit something calmer, and some parts even need a specific emotion mixed in. Each gets matched to my corresponding avatar footage for that kind of segment, generated piece by piece, then stitched together. What comes out looks like it’s been edited—very natural, and the expressions land right on point.
[04:16] Running at the same time: cover generation for different platforms in different ratios—3:4, 4:3 and 16:9, basically those three sizes. After generating, it also automatically checks character consistency—whether the person matches the reference photo I gave it. Meanwhile, based on the verbatim script of the video, it composes the bottom-half visuals—those polished images you see, whether slides or web pages. It slices the timing very precisely: from which second to which second, what kind of visual I need—
[04:41] —to power the point I’m making. This is where HyperFrames comes in—that skill from HeyGen—generating those animated web-page effects. Or GPT’s Image 2, directly generating full-frame chalk-on-blackboard looks, or clay-style looks, as companion B-roll illustrations. Once all the video material is rendered, it gets stitched together, plus subtitle composition and burn-in. The subtitles get placed automatically at the right position, and the date and my taglines—
[05:06] —get stamped on like stickers too. Finally, it writes all the titles, tags and copy for publishing on five platforms, and automatically opens those web pages and fills everything in for me. After I’ve checked it, I press the publish button myself. So there you go—that’s how my videos are made. I won’t lie to you that it’s done in 10 minutes. My own time spent across the whole process stays under 10 minutes, but for the AI to actually finish all the computation takes one to two hours. Because it has to think, it has to write, it has to generate images, it has to composite. All of that takes time.
[05:32] I won’t lie to you that one button press gives me hundreds of videos. After the video is generated, it also runs its own quality check. Let me quickly share some of the pitfalls too. For example, if you use Image 2 for images, always have it generate the full frame—not a background with text overlaid on top. That way the whole aesthetic stays consistent. And at every stage that needs an API call, needs payment, it runs a self-check first—otherwise money would get wasted on the generation. The total cost of one video is around a few tens of RMB. Every single stage is an official API call—
[05:57] —paid for real. If you want to lower the cost, that’s possible too—we just swap in some cost-effective models. Like Gary from my group chat: one of his episodes costs around a few RMB, and it turns out pretty decent. And if I’m free to spend, making richer B-roll images at higher image density—and if you also turn image-to-video into B-roll—one episode’s cost reaches a few hundred RMB. So AI-generated, AI-edited video really does span a price range from around ten RMB up to the hundreds. Now let’s look back at—
[06:22] —the whole pipeline and which model is responsible for what at each stage. The chief director is Claude Code. Voice is MiniMax. The digital human is Shiliu. Infographics are GPT Image 2. Motion rendering is HyperFrames. Subtitle alignment uses Gemini 2.5 Flash. And for assembly and publishing it’s FFmpeg encoding and browser automation. So what are the three questions fans ask most? The first: fully automated, one-click output—yes, it can be done. But throughout the process, for the sake of video quality, I still give some feedback and adjustments. For example, certain passages get—
[06:47] —my own views, my own expression added in. Or I have it restructure the writing. And after it generates images, there might be layout-tweak suggestions for individual ones. It can produce without a human involved—but with a human involved it comes out better. Next: do viewers know this is a digital human? Every platform now has a label for AI-generated content. If our content doesn’t carry that label, it can actually still get traffic. Some of my digital-human videos, including ones I made for other IPs, have reached—
[07:12] —hundreds of thousands, even millions of plays—real breakouts. But there’s one catch: the platforms’ current review rules. Take WeChat Channels. When a video is about to blow up—entering roughly the 100k traffic pool—there’s a so-called human review. If yours is a digital-human video without the label, its distribution gets restricted. And last: can an ordinary person replicate this whole pipeline? Actually yes. Besides Claude Code, Codex, or WorkBuddy plus Kimi’s models, can all do what I described. But every one of them—
[07:37] —takes a lot of time to tune: a whole agent, executing this thing as a loop, a cycle. This pipeline of mine has been iterating since two years ago. Including my own on-camera videos—those are also fully AI-edited. That’s even simpler: no visuals to generate at all, it just does a rough cut for me, adds subtitles, and out it goes. Now back to where we started. A lot of fans keep telling me—good friends especially—forwarding me other bloggers’ videos, saying if your standards are too high you’ll never make big money. You have to fleece the audience—
[08:02] —especially in paid knowledge. I could not disagree more. The content I post on Channels right now—package it into a course, run paid traffic, sell it. You think I don’t know how? Of course I know. What I do for a living is internet advertising, internet marketing. But things that are irresponsible about delivery—I won’t do them. I just want to sleep soundly, with every single yuan I earn staying clean. OK, that’s it for today’s video. Remember to follow me. See you in the next one. Bye-bye.