A Master-Class Speech, Reproduced in Fifteen Minutes—Wiring Fable 5 into WorkBuddy to Build a Luo Zhenyu-Style Speech Expert
That "Luo Zhenyu speech" at the top of the video was AI-generated—script from the speech expert, voice from AI.
The first time I did this, in 2023, it took a sprawling agent architecture and a whole script library. This time, in an in-person class, I used WorkBuddy's built-in experts and skills: fifteen-odd minutes of concurrent research, four or five turns of conversation, and even the eight minutes of audio came out live. This episode takes the prompts apart turn by turn, plus the details of getting blocked by a safety guardrail on cloning a celebrity's voice, and where Chinese agents really differ from Claude Code and Codex.
I opened the video with a stretch of “a Luo Zhenyu speech”—AI anxiety, can’t study your way out and can’t lie down either, ending on the four-character line about the gentleman not being a mere vessel. Halfway through I couldn’t keep making it up: that speech was AI-generated. AI wrote the script, AI did the voice. Not bad, right?
You should know I’m a big fan of the way Luo Zhenyu speaks. I’ve been trying to imitate his delivery with AI since 2023, and since my business involved digital avatars, I’d been studying voice cloning tech the whole time anyway. Back then, getting to “Luo Zhenyu-style delivery” meant designing a sprawling agent architecture and pairing it with a corpus of his scripts, and even that barely reached the level I wanted.
Fifteen minutes, four or five turns of conversation
But on July 11, I really didn’t expect it to go this fast.
That day I was teaching Owl Academy’s 19th cohort in person. In class I only meant to run a quick example—wiring the Fable 5 model into Tencent’s WorkBuddy. Here’s what happened instead. Using nothing but WorkBuddy’s own “experts and skills,” roughly fifteen minutes of concurrent research and four or five turns of conversation gave me a Luo Zhenyu-style speech expert. It then generated that eight-minute speech audio from the opening, right there in the room.
This confirmed again something I say all the time: if you’re not learning AI right now, that’s fine—six months, a year from now, every technique we use today will have been replaced and there’ll be nothing left to learn. But the thinking behind collaborating with AI always iterates and always compounds.
What the prompts actually said
Let’s take those turns apart.
Turn one: multi-agent concurrent research. I told WorkBuddy to assemble a multi-agent concurrent research team. First, go collect Luo Zhenyu’s speaking style, core arguments, signature lines, delivery techniques, and the content of his major speeches over the years, all as the base corpus. Then have the agents build a speech expert system with a broad feature set: simulating the language style, integrating the knowledge structure, adapting to different settings, and taking feedback from the user to adjust.
Once the task went out, you could watch a pile of sub-agents spawn under WorkBuddy. A lot of students asked me how that happened. It’s because the prompt opened with “multi-agent concurrency, deep search across the web,” so it spun up several agents on its own, each searching from a different angle, then pulled the results back to the main agent for the next step.
In the end it pulled around 120,000 words of corpus off the web and extracted the speaking style and delivery techniques from it. Keep in mind I once did a whole episode on the “voice DNA language fingerprint” step alone. Work that used to eat a lot of time and effort now finishes in a turn or two, on the strength of the model and of an agent’s “experts and skills.”
Turn two: package it as an expert. WorkBuddy has a skill called Expert Manager. Select it in the lower left of the chat window and it packages the structured information from the previous two turns into a “logic and style speech expert” system, which you can call up in chat any time after that.
Turn three: add the hook system. This is where I put in some thinking of my own. A short-video script lives or dies on the opening hook. I trusted the speech expert to convert a script well, but a hook needs extra information behind it—a classical anecdote to cite, say, or a startling fact about the topic at hand. The expert system from the earlier turns had none of that, so I added a round: before writing the script, run one pass of content research just for the hook.
Voice cloning, and an interesting safety guardrail
For the audio, I found about 30 seconds of a Luo Zhenyu speech clip to use as the voice cloning reference, planning to do the rest with the MiniMax 2.8 HD Speech audio model.
And here I found out that WorkBuddy’s safety guardrails are the real thing: the moment it caught me trying to clone a celebrity’s voice, it refused to generate. Why do I say that’s WorkBuddy’s guardrail and not a model limitation? Because I finished that same audio generation task in Claude Code, on the same model. So I think WorkBuddy injects its safety rules at the system-prompt level: copyright violations, along with other potentially harmful behavior, don’t get through.
One more note: that final eight-minute speech uses a preset voice with a slight speed-up, not a clone of Luo Zhenyu’s voice. Cloning someone’s voice without their authorization is a red line that doesn’t move no matter which tool you switch to. Celebrity voices: nobody gets to clone them.
Chinese agents: where they lag, where they lead
This case also got me thinking: where’s the real difference between Chinese agents and the world’s best—Claude Code, Codex?
Any of these agents can connect to the same models, and the models keep getting stronger. I’m using all of them heavily right now, and my read is this: the difference is still in the prompting. An agent like WorkBuddy is in catch-up mode, filling in capability as it goes. But pick the right expert and skill for the job, give it a complete prompt, and it already covers most everyday needs. It has even pulled ahead of the foreign agent products on project team collaboration and on how easy MCP connectors are to wire up.
I’m also hoping Chinese models catch up.
Source: EP0060_audio.mp3 · ASR model gemini-2.5-pro (chunked parallel) · full text of the original recording
[00:00] Have a listen to this bit of a talk by Luo Zhenyu first. There’s something I’ve been holding in for a long time, and today I have to say it. Now that AI is here, what exactly are ordinary people like us supposed to do? You say, fine, I’ll learn—learn what? Learn to write code? AI writes it a hundred times faster than you. You say fine, I won’t learn, I’ll just coast—can you actually coast? You can’t. You scroll your phone at night and it’s another story about someone getting replaced by AI. There you go, that’s the anxiety. And that’s the real situation a lot of us are in today.
[00:25] Can’t move forward, can’t move back, stuck. So today I want to find, for all of us, and for my own jittery, panicky self, a phrase to stand on. Four characters: the gentleman is not a vessel. I’ll cut it there. So that talk was Luo Zhenyu—okay, I can’t keep the act up. That talk was AI-generated. Pretty good, right? You have to understand, I love Luo Zhenyu’s talks. Since 2023 I’ve been trying to use AI to imitate the way he speaks.
[00:50] And because of a business need to build digital avatars, I’ve also been researching voice cloning the whole time. But on the 11th, I honestly did not see this coming. That day I was teaching cohort 19 at Owl Academy in person, on WorkBuddy. Originally, to imitate the way Luo speaks plus the voice cloning, I’d designed this hugely elaborate agent structure and a matching corpus of his scripts to pull that off. But on that very day, in class, I meant to do a demo, and I hooked the Fable 5 model
[01:16] into WorkBuddy. And I didn’t expect that with WorkBuddy’s own—didn’t expect that with WorkBuddy’s own experts and skills, about fifteen minutes of parallel research and four or five rounds of conversation was all it took to produce a Luo Zhenyu-style speech expert. And we did the voice cloning live, generating that eight-minute audio clip you just heard. I think this proves that thing I keep saying: if you’re not learning AI right now, that’s fine—in six months, in a year,
[01:42] all of today’s techniques will have been swapped out anyway, and there’ll be nothing left to learn. But the way of thinking about collaborating with AI—that keeps iterating, that keeps growing. So let’s look at how the prompt was actually written. First I told WorkBuddy that I needed it to assemble a multi-agent parallel research team. First, have it collect Luo Zhenyu’s speaking style, core ideas, signature lines, delivery techniques, and the content of his major talks over the years, as a base corpus. Then have the agents build
[02:07] a speech expert system that could do a few different things: mimic his language style, pull his body of knowledge together, adapt to different settings, and take back-and-forth feedback and tweaks from whoever’s using the agent. Once a task like that goes out, you can see that a bunch of sub-agents get created underneath WorkBuddy. A lot of students ask me where those come from—it’s because the prompt opened with multi-agent, parallel, deep search across the whole web.
[02:32] So it automatically spins up several agents that handle the search from different angles, then pools the results back to the main agent for the next step. And right after that we see it found around 120,000 words of corpus across the web, and pulled his speaking style and delivery techniques out of that corpus. And bear in mind, on the Voice DNA thing alone—the language fingerprint—I’ve done a whole episode.
[02:57] That used to eat up a lot of time and effort. Now, because the models are this good, because the agent has these experts and skills, one or two rounds of conversation gets a job this complicated done. And then, in WorkBuddy, there’s a skill called Expert Manager. When you want to create an expert in WorkBuddy, you first pick that skill at the bottom left of the chat window, so it can take the structured information from the previous two rounds
[03:22] and package it into a speech expert system with its own logic and style, which you can then call up in later chats. And here I added some of my own thinking. As you know, a short-video voiceover script needs an opening hook to hold people. And I do believe the speech expert we’ve built can rewrite a script well. But the opening hook actually needs extra information to support it—citing a classic reference, telling an
[03:48] old anecdote, or leading with a startling fact about the topic at hand. That kind of hook system wasn’t part of the expert system we’d just built. So I tacked on one more round afterwards, telling it to go dig up hook material before it writes the script. I also went and found about a 30-second clip of Luo Zhenyu speaking as the reference for voice cloning, using MiniMax 2.8 HD Speech as the voice cloning model
[04:13] to do the rest. But here’s what I found: WorkBuddy’s safety guardrails are really well done. The moment it caught that I was trying to clone a public figure’s voice, it refused to generate. And why do I call that WorkBuddy’s guardrail? Because I ended up finishing the audio generation for this task in Claude Code with the exact same model. So I believe WorkBuddy has safety protections injected at the system prompt level.
[04:39] Copyright protection, plus other things that could do harm—it just won’t do them. And going through a case like this, I’ve been thinking about what the real difference is between domestic agents and the world’s top agents, including Claude Code, including Codex. Because any of these agents can hook up the same models, and the models keep getting stronger. And I’m using all of these agents hard, every day. In that situation, my personal read is that
[05:04] the difference is still in the prompt layer. Which is to say, an agent like WorkBuddy is still in catch-up mode—its capabilities are still being filled in. But as long as we pick the experts and skills that fit the job, and give it a complete prompt, it can already handle most of what you need day to day. And on project team collaboration and the convenience of MCP connectors, it has already pulled ahead of the overseas
[05:29] agent products. Here’s hoping the domestic models catch up too. That’s it for today’s video. See you tomorrow, bye.