The render finished at 3:47 AM. I know because my laptop fan spun down hard enough to wake me. I walked to the desk, hit play, and watched a version of myself I had never recorded deliver a nine-minute lesson I had never spoken. Clean lip sync, natural gestures, my voice. I had gone to bed at 11:30 PM after dropping a script into a Google Drive folder. The pipeline did the rest while I slept.
Here is the thesis I have landed on after two months of running this system: an AI video pipeline succeeds or fails on the seams, not the models. HeyGen's avatars are good enough. ElevenLabs' voices are good enough. The engineering problem, the actual problem, is chunking, sequencing, retrying, and stitching, and that is why Claude Code sits at the center of my stack rather than any single video tool. I built this because I produce lesson videos for my AI School courses, and the previous method cost roughly $300 per finished ten-minute video in editor fees plus about four hours of my own filming and review time. For a 40-lesson course, that math is brutal.

The Four Tools and the One Job Each Owns
HeyGen owns the face. My avatar is trained on my own webcam footage; HeyGen's Avatar 5 model accepts as little as 15 seconds of video, though closer to two minutes gives noticeably better motion. Avatar 5, launched in 2025, is the generation that finally crossed the uncanny valley for me. HeyGen has since shipped Avatar V as its successor, so the ceiling has only gone up since I built this. HeyGen also runs an MCP server, which is connected to my Claude environment; that matters later.
ElevenLabs owns the voice. I fed its professional voice cloning about two hours of clean audio, podcast recordings, tutorial voiceovers, narrated screencasts, well past the 30-minute minimum the pro cloning tier wants. Settings that survived my testing: stability 0.7, similarity 0.8, speed 1.0, style exaggeration low. Boring settings, consistent output. I did a deeper build with their conversational side in my ElevenLabs voice agent post.
Claude Code owns the seams. It watches a Drive folder, splits scripts into chunks, calls ElevenLabs per chunk, submits audio to HeyGen per chunk, polls renders, retries failures, and hands everything to the editor layer. This is the component you cannot buy, because it encodes decisions specific to my content.
Remotion owns the assembly. Transcription, on-screen captions synced to words, section titles, and stitching clips at the same sentence boundaries where the script was originally split. Free and self-hosted; my Remotion-by-prompt workflow covers that layer on its own.
The End-to-End Flow, Drive Folder to MP4
- Ingestion. I write a lesson script in markdown and drop it in a watched Drive folder.
- Semantic chunking. Claude Code splits it into 45 to 60 second chunks, always at sentence boundaries, never mid-thought.
- Voice per chunk. Each chunk goes to ElevenLabs with my cloned voice; MP3s come back, get duration-checked, and are re-queued if the timing looks wrong.
- Avatar per chunk. Each MP3 goes to HeyGen with my avatar ID; short chunks render in parallel, which is most of why the whole thing finishes overnight.
- Stitch. Remotion assembles clips in order, lays captions, renders the final MP4 into my output folder with a job summary: chunk count, runtime, retries.
One workaround worth admitting: when I built this, HeyGen's public API defaulted new renders to the older avatar generation, and Avatar 5 renders were reliable only through the web app. So Claude Code drove the web interface with Playwright for that step, browser automation as duct tape. It felt slightly criminal and worked flawlessly, and it is a useful reminder that an orchestrator that can only call APIs is weaker than one that can also click.
The Rule That Decides Whether Your Pipeline Ships
If I could send one sentence back to myself before the first failed overnight run: the quality ceiling of the entire pipeline is set by the chunker.
Not the avatar model. Not the voice. The chunking. Chunks that break mid-thought produce audible discontinuities at every stitch point. Chunks over 60 seconds degrade ElevenLabs' delivery. Chunks that open with a dangling conjunction lose their context and read flat. Every one of those flaws repeats at every boundary, and a ten-minute video has a dozen boundaries. Budget more engineering time for the chunker than for everything else combined; it is the difference between "impressive demo" and "wait, you didn't film this?"
What It Actually Costs
My current monthly stack, at the tiers I am on:
| Service | Cost | Covers |
|---|---|---|
| HeyGen Creator | ~$30/mo | Base avatar generations |
| HeyGen API credits | ~$4 per finished minute | Renders beyond the tier |
| ElevenLabs Creator | $22/mo | ~100 minutes of audio |
| Claude Code | $20-$200/mo | Orchestration, usage-dependent |
| Remotion | Free, self-hosted | Rendering on my machine |
Marginal cost per finished ten-minute video: right around $50, dominated by HeyGen render time. Against the $300 I used to pay an editor, that is a 6x reduction, and the four hours of filming and review became about 20 minutes of script review. Across a course catalog, the pipeline paid for its build time inside the first month.
Where I Refuse to Use It
Honesty section, because this is where most AI video content lies to you. The pipeline is for structured, instructional delivery: lessons, walkthroughs, product explainers, content where the script is the product and my face is packaging. I do not use it for anything where being personally present is the point. Viewers forgive an avatar teaching them Laravel; they do not forgive discovering that a personal story was delivered by a render. Disclosure is part of my setup, and I think it stops being optional the moment the uncanny valley closes.
Technical failure modes to expect: names and numbers are where lip sync still occasionally wobbles, chunk retries cluster around chunks with em-dash-heavy or bracketed text (clean your scripts), and the first avatar training attempt is usually your worst one. Re-record your training footage after you have seen what the model does with it.
The Bottleneck Moves, It Never Disappears
Before: production was the bottleneck, so I wrote fewer scripts than I had ideas. Now scripts are the bottleneck, which is a better problem, because writing is the part I actually want to be doing, and the part where the quality lives. The pipeline did not replace my work; it moved my work up the stack to the layer where judgment matters. The same shift, applied to b-roll instead of talking heads, is my Higgsfield timestamp workflow, and the two systems share the same orchestrator.
The fair test of a pipeline like this is not a demo reel — it is whether a student notices while they are trying to learn something. The lessons inside AI School are the output; some of them rendered while I was asleep, and I genuinely cannot always tell you which ones.