← Back to home

AI edits the finished video: the whole pipeline, revealed

Long-Form Video · EP0072 July 29, 2026 08:22
What this episode covers

A viewer asked me yesterday: how exactly do you make your daily videos with AI? So I had CC generate a web-page workflow manual from my own project folder, and this episode walks through it.

Seven steps: topic selection, scripting with a voice DNA distilled from all my past scripts, voice-cloned narration, a digital human driven segment by segment, covers plus bottom-screen infographics, subtitle burn-in, and auto-filled copy for five platforms. The AI computes for one to two hours per video. A video costs tens of RMB—every step is a paid official API call. It can run fully automated, and it gets better with a human in it.

How it started: I lost my voice, so the workflow became a digital human

A viewer asked me yesterday how I use a digital human, CC and image-generation models to produce my daily short videos—especially the automated part. I’ve iterated this logic through many rounds, so this episode opens it all up. I had CC go into my video-production project folder, read the conversation history, the project memory and the delivery documents, and produce an up-to-date workflow manual as a visual web page in my own PPVI light style. The walkthrough below follows that manual.

The direct trigger: a cold plus an allergy wrecked my voice completely. I couldn’t record myself, so I rebuilt the whole workflow around a digital human presenter.

The prep list is short: record a clip of my own voice for cloning; find one continuous 60–90 second shot of my upper-screen presence from an earlier episode; and hand over every spoken script I’ve ever written. Then come the seven steps.

Early on I built a local tool that picks topics from trending charts. Honestly, my videos rarely use its recommendations—not because they’re bad, but because every day I already have something I badly want to say, or something fans keep asking for. Sometimes the traffic is worse than an AI-picked topic would get. But it’s what I want to say and what they want to hear, and if I don’t make it, the group chat keeps asking.

Step two: scripting—voice DNA, plus a defense against the “AI flavor”

With the topic set, the script gets written against my voice DNA—the expression structure and style distilled from all my past spoken scripts.

Here I added an extension. I noticed that from Sonnet 4.6 onward, and the GPT-5-era models too, almost all of them think deeply in English and translate back into Chinese. That produces what I call the new generation of AI flavor—a touch of translationese. Natural Chinese says 流程循环; translationese says “the human is in the loop.” Feel the difference.

So the first draft is a dense stack of information—background material, cases, the whole logical thread written out in full—and then it goes to DeepSeek or Kimi K3 for a native-Chinese rewrite and review. Once that’s done I barely check it. Straight into audio generation.

Step three: voiceover—the verbatim script becomes a “TTS-generation script” first

The voice is already cloned; now TTS generates the audio. Since the whole line is automated, only API-callable TTS models qualify. By my testing, MiniMax’s 2.8 HD Speech is still the strongest machine-reading model at this point in time.

One key move here: I have CC convert the verbatim script into a “TTS-generation script.” What’s the difference? It pre-resolves the heteronym characters, pre-processes the numbers, and adds pauses and interjections in the right places—turning it into something a speech model understands and follows far more easily—then generates the full read-through audio.

Step four: the digital human—cut by pace and mood, stitched like it was edited

Same logic for the digital-human platform: API service only. I care about video clarity and lip-sync accuracy, so I chose the Shiliu digital-human API from Xiangliang Fangcheng.

CC cuts the voice track into segments: some parts want a faster pace, some want calm, some need a specific emotion mixed in. Each segment gets matched to different avatar footage of mine, generated one by one, then stitched. The result comes out looking like it was edited by hand—natural, with the expressions landing right.

Step five: covers and bottom-screen infographics—running in parallel

Covers come out in three aspect ratios for different platforms: 3:4, 4:3 and 16:9. After generation it automatically checks character consistency—whether the face still matches the reference photo I gave it.

Meanwhile, from the verbatim script it composes the polished bottom-screen visuals—slides or web pages. It slices time precisely: from second X to second Y, which visual should power the point I’m making. That’s HyperFrames (the HeyGen skill) for animated web-page graphics, or GPT’s Image 2 generating full-frame chalk-on-blackboard or clay-style B-roll illustrations.

One pitfall worth sharing: when you use Image 2, always have it generate the full frame—never a background with text overlaid afterward. That’s how the aesthetics stay consistent.

Steps six and seven: assembly and publishing

Once the material is rendered, it gets stitched, subtitles composed and burned in at a fixed position, and the date and my tagline stamped on like stickers.

Finally it writes every title, tag and description for all five platforms, opens the pages and fills them in automatically. I review, and I press the publish button myself.

Two honest numbers: time and cost

I won’t tell you the video is done in ten minutes. My own hands-on time in the whole process stays under ten minutes—but the AI needs one to two hours to finish all the actual computation. It has to think, write, generate images, composite. All of that takes time. And I won’t lie to you that one button press turns into hundreds of videos.

On cost: a video runs at tens of RMB, and every step is an official pay-as-you-go API call, paid for real. You can bring it down—swap in cost-effective models, like Gary in my group chat does at a few RMB per video, with decent results. Or you can spend freely: richer B-roll, higher image density, image-to-video B-roll—then one episode climbs to a few hundred RMB. So AI-made video spans a range from about ten RMB to a few hundred.

One more guardrail: before every step that calls a paid API, it runs a self-check first. Otherwise money gets wasted on doomed generations.

The full credits table

  • Chief director: CC
  • Voice: MiniMax
  • Digital human: Shiliu
  • Infographics: GPT Image 2
  • Motion rendering: HyperFrames
  • Subtitle alignment: Gemini 2.5 Flash
  • Assembly and publishing: FFmpeg encoding + browser automation

The three questions fans ask most

Can it run fully automated, one click to a finished video? Yes. But for quality I still give feedback along the way: injecting my own views into certain passages, having it restructure the writing, suggesting layout tweaks on individual images after they’re generated. It works without a human. It works better with one.

Do viewers know it’s a digital human? Platforms now label AI-generated content. Unlabeled videos can still get traffic—some of my digital-human videos, including ones I made for other IPs, have gone viral at hundreds of thousands to millions of plays. But watch the review rules. On WeChat Channels, for example, when a video is about to break out—entering the 100k traffic pool—a human review kicks in. An unlabeled digital-human video gets its distribution restricted.

Can an ordinary person replicate this? Yes. Besides CC, Codex or WorkBuddy plus Kimi’s models can all do what I described. But every path requires serious time tuning the agent loop that executes it. I’ve been iterating this pipeline since two years ago. My own on-camera videos are also entirely AI-edited—that’s even simpler: no visuals to generate, just a rough cut, subtitles, publish.

Last thing: integrity, and money that’s clean

A lot of fans—good friends among them—forward me other bloggers’ videos and say: with standards that high you’ll never make big money; you have to fleece the audience, especially in paid knowledge. I could not disagree more.

Take the content I post on Channels—package it into a course, run paid traffic, sell it. You think I don’t know how? Of course I do. Internet advertising and marketing is literally my trade. But things that betray the delivery, I won’t do. I want to sleep soundly. Every yuan I earn stays clean.

I won't lie to you that one button press turns into hundreds of videos.