How to Actually Get Non-Uncanny Avatar Videos Out of HeyGen (Without Filming It Ten Times)
HeyGen's Avatar V is dramatically better than Avatar IV, and most people are still recording their 15-second reference like it's a passport photo. Eight habits that separate a stiff, robotic talking head from an avatar that actually looks like you.
Here's the deal with HeyGen in 2026: the model got dramatically better this spring, and most people didn't update their habits to match. Avatar V shipped in April, and it does something Avatar IV couldn't. It builds a fine-tuned model from your 15-second reference clip and uses that as the foundation for everything it generates. Motion, teeth, cadence, the little tilts of your head. All learned from you, not predicted from a single photo.
The problem? People are still recording that clip like they're taking a driver's license photo. Flat, stiff, arms glued to their sides, mumbling the consent code at a wall. Then they wonder why the finished video looks like a hostage tape. I've been putting Avatar V through real production work for three months, and the gap between a great avatar and a cursed one is almost never the platform. It's the 15 seconds you fed it and the script you handed it after. These eight habits are the ones that consistently move the needle from "uncanny" to "wait, is that actually you?"
1. Treat the 15-second reference like a screen test, not a mugshot
This is the single biggest lever you have, and it’s the one most people waste in the first two minutes of their HeyGen account. Everything downstream is built on this clip.
The key thing to understand: Avatar V builds a fine-tuned model based on your reference video. Instead of predicting your motion, teeth, facial expressions, and gestures from a single image alone, it watches how you actually move and uses that as the foundation for every video you generate. Translation: whatever you do in those 15 seconds, you’ll do forever. If you sit rigidly and mumble, your avatar will sit rigidly and mumble in every video you ever generate from that twin.
HeyGen’s own team is blunt about it: be expressive and natural, because your motion, energy, and gestures are all learned from this video. A flat, reserved delivery will produce a stiff avatar. Include natural hand gestures if you want your avatar to gesture. What you do in those 15 seconds is what it’ll replicate.
Record it like you’re pitching to an investor, not confirming your identity. Smile. Gesture. Move your hands within the frame. Actually perform those 15 seconds. Do it three or four times and pick the best take.
2. Shoot clean footage, and skip cinematic mode
The reference clip’s technical quality sets the ceiling for everything Avatar V can do with you. Junk in, junk out, and HeyGen has been consistent about what “clean” actually means.
If you record yourself to build a custom avatar, the input footage sets the quality ceiling. Clean footage in, clean avatar out. Shoot in 4K on a real camera or a modern smartphone, but skip cinematic or portrait mode. The generator wants a clean, evenly focused subject, not the artificial background blur those modes add.
A few more non-obvious rules from HeyGen’s own guide: look directly into the camera lens throughout the recording, avoid looking at the screen or scanning around, and if you’re using a script, try to spend the first 15 seconds looking straight at the lens before referring to it. That ensures your opening looks natural and direct.
One more: the recording must be one continuous clip. Don’t edit or splice clips together. You can adjust brightness, contrast, or saturation before uploading if needed. HDR footage is fine if that’s your camera’s default setting.
And light yourself properly. Soft, even, from the front. If it looks moody and cinematic, it’s wrong.
3. Write scripts for the ear, not the eye
This is where every marketer new to HeyGen faceplants. They paste in the exact copy from a landing page (dense, comma-choked, three ideas per sentence) and the avatar delivers it like a hostage reading a manifesto.
Write it out loud. Short sentences. One idea per line. Contractions on. The rule that actually works: write for speaking, not reading. Short sentences. Conversational tone. Imagine you’re explaining something to a colleague over coffee. Aim for 130-150 words per minute. That’s natural speaking pace. A 500-word script produces roughly a 3-minute video.
HeyGen also recommends keeping each script section under 2,000 characters for the best results , so if you’ve written a 4,000-character monologue, break it across scenes.
Front-load the payoff in the first ten seconds. If the reason to watch doesn’t land before the avatar’s second gesture, you’ve lost the room.
4. Use CAPS to steer emotion. The model is audio-driven
This is the Avatar V trick that isn’t obvious from the interface, and it’s genuinely powerful once you internalize it.
Avatar V is audio-driven. The energy and emotion in your voice directly controls how your avatar looks and moves. Excited voice = excited avatar. Monotone audio = flat, underperforming avatar. You can use CAPS in your script to emphasise words and prompt more energy.
This is the fix for the “why does my avatar look bored” problem. It’s not bored. Your synthesized voice is bored, and Avatar V is faithfully rendering that boredom. Capitalize the words you’d actually punch in a real read. “This is the ONE thing you need to know about Q3.” Sprinkle emphasis, don’t shotgun it. Two or three caps per paragraph, tops.
The same rule applies in reverse: a calm, measured delivery produces composed, precise gestures and neutral expressions. A passionate pitch produces energy, leaning forward, wider eyes, and more expressive hands. The avatar responds to what the voice is actually communicating.
5. Assign different motion references to different scenes
This is the trick nobody in the tutorial videos talks about, and it’s the one that made my avatars stop looking like a single mode on repeat.
In AI Studio, you can assign different motion reference videos to different scenes. For example, a more expressive motion style for your opening, and a calmer professional style for your main content. For side-angle shots, record a separate 15-second clip at that angle and use it as the motion reference for those scenes.
Record two or three motion references before you start producing anything real: one energetic (the “hook” opener), one calm and explanatory (the “main body”), and if you want angle variety, a third at a slight profile. Then assign them per scene in the editor. Your avatar stops looking like a single loop and starts looking like a presenter who actually has range.
6. Pick the avatar and voice as a pair, and then stop switching
The most common self-inflicted wound is voice-shopping mid-project. You pick avatar A, try voice 1, try voice 2, decide voice 2 is better, then swap the avatar. Now voice 2 doesn’t fit anymore. You’ve broken the vibe and burned an hour.
The avatar and voice set the tone before a single word lands, and HeyGen pairs them, so judge them together, not separately. Preview 3 to 4 avatars with your actual script, not the demo line. The same sentence reads warm on one face and corporate on another. Play the voice over the chosen avatar before committing. A voice that sounds fine alone can fall out of sync with a particular avatar’s face and default pacing. Lock one avatar and voice for a series so your channel feels like one presenter, not five.
Do the pairing test on your actual opening line, not “hi, welcome to my channel.” Then commit. If you’re building a series or a brand, one twin, one voice, forever. That’s how you get recognition.
7. Draft cheap, finalize expensive
This is the money habit, and it’s the one that separates people who use HeyGen sustainably from people who blow a month of credits in a week.
Avatar V is the premium engine, and it charges like one. Both Avatar V and Avatar IV use 20 credits per minute , which sounds fine until you realize you’ve re-rendered the same 90-second script eleven times because you keep tweaking commas. The move is to draft with a cheaper engine and only promote to V when the words are locked.
The cleanest workflow I’ve landed on, and it echoes what the reviewers who actually stress-test this platform recommend: the best HeyGen workflow is hybrid. Use Avatar III for drafts, Avatar V for identity-critical scenes, Video Agent for assembly, AI Studio for correction and external footage wherever physical presence isn’t needed.
In practice: get the script and pacing right on a cheap engine. Only then hit generate on Avatar V. Your monthly credit budget will last three times as long.
8. Use Seedance for the cinematic stuff, not Avatar V
This is the newest habit on the list, and most HeyGen users haven’t caught up to what shipped in April. HeyGen integrated Seedance 2.0 as a first-class engine, and it’s the right tool for a category of shot that Avatar V is genuinely bad at.
Seedance 2.0 is the model everyone has been talking about, and HeyGen is the first avatar video platform integrating it with identity-verified human faces. Every other implementation of Seedance is locked to fictional characters. With HeyGen, your real Digital Twin gets cinematic motion, dynamic camera work, and up to three avatars in a single scene.
The rule of thumb: Avatar V is for talking-head scenes, someone delivering a message straight to camera. If you want your twin walking through an office, gesturing across a stage, or interacting with another version of themselves, that’s a Seedance shot. The same Digital Twin you built for Avatar V is the one Seedance casts in those cinematic shots. Record once, walk through a coffee shop, gesture across a stage, hand off a line to two other Twins of yourself or your team. One identity, two engines, full coverage from talking head to cinematic scene.
You don’t have to pick. Compose the video from both. Talking-head coverage on Avatar V, cinematic beats on Seedance, cut between them. That’s the workflow that stopped my HeyGen output from feeling like a slideshow.
A bonus habit, because it’s the one everyone forgets
Use visible expressions, especially during pauses. Staying expressive during quiet moments helps the model create natural idling states between sentences. When you record your 15-second reference, don’t freeze between phrases. React, breathe, micro-nod. Otherwise your avatar’s between-sentence idle will be a dead stare, and dead stares are what the uncanny valley is made of.
The one habit that ties it all together: stop treating HeyGen like a text-to-video button and start treating it like directing an actor who happens to be you. Every knob does something specific. The 15-second reference is a performance, not a passport photo. The script is a read, not a paragraph. The engine choice is a budget decision, not a default. Nail those three and the avatar stops looking like AI and starts looking like the version of you who actually got enough sleep.
Sources
- https://www.heygen.com/
- https://community.heygen.com/public/resources/how-to-get-the-best-results-with-avatar-v-in-heygen
- https://community.heygen.com/public/resources/avatar-v-live-webinar-recap-top-questions-answered-2026-04-16
- https://help.heygen.com/en/articles/14602974-avatar-v-is-now-available-on-heygen
- https://www.heygen.com/blog/heygen-april-2026-release
- https://help.heygen.com/en/articles/11049837-create-your-first-video-in-our-studio