An AI talking avatar is a digital presenter that speaks your script with generated lip sync. Pick by the face: a photo avatar for a fictional persona, an AI twin for a real person who can record footage and consent, a stock avatar when the face does not matter. On October 2026 list prices a finished minute runs from about $0.19 to $3.75 depending on the avatar model, before retakes.
An AI talking avatar is a digital presenter that speaks your script with generated lip sync, and it comes in three forms: a stock presenter, a photo avatar made from one image, and an AI twin cloned from footage of a real person. For a fictional AI persona the photo avatar is the route that works, because twin features need a consent recording from a real human, and on October 2026 list prices it costs roughly $0.34 to $0.77 a finished minute on HeyGen's entry plan.
Features and prices checked October 2026 against HeyGen's pricing, credit table, Photo Avatar guide and Digital Twin FAQ, Synthesia's pricing, October plan changes and personal avatar docs, Hedra's talking avatar page and Character-3 page, D-ID's FAQ and Google's Flow credits page. Costs per minute are our arithmetic on those published rates, not a test.
Vendor pages rank for this topic, and each one describes its own product. This guide is the independent version: what the avatar types really are, which one a persona account can use, and what a minute of finished video costs. Avatars are one of several ways to put a persona on screen; our AI influencer video generators ranking covers the scene tools that sit beside them.
Who this is for
- Persona operators who have a consistent face in stills and now need it to talk to camera for tips, reactions and product videos.
- Real creators and founders who want to stop filming every script themselves.
- UGC sellers and small brands testing spokesperson ads without a shoot. Our AI UGC ads guide covers the ad side.
What an AI talking avatar is, and what it is not
An avatar tool holds a presenter still in frame and generates the performance: lips, expression, small head and shoulder movement. You supply a script or an audio file. That fixed framing is the point. It is why a one-minute monologue holds together in an avatar tool and falls apart in a scene model built for ten-second clips.
Two neighbours get confused with it. Lip sync tools change the mouth in a video that already exists; our best lip sync AI comparison covers those. Scene models such as Kling and Google Flow generate a whole moving shot and can include a spoken line. Kling's VIDEO 3.0 runs 3 to 15 seconds with native dialogue and Flow's Gemini Omni Flash tops out at 10 seconds, so they suit talking beats, not talking heads.
Photo avatar, custom avatar or AI twin: how to choose
Start with one question: is the face a real person who can sit in front of a camera? Everything else follows from the answer.
| Type | Built from | Consent step | Where the motion comes from | Fits |
|---|---|---|---|---|
| Stock avatar | A presenter from the tool's library | None | Pre-built by the vendor; nothing to set up | Training and explainer video where the face is not the brand |
| Photo avatar | One image you upload or design in the tool | HeyGen validates the look; Synthesia asks for a live consent recording | Generated from a still: lips, face and head | A fictional AI persona, or a fast version of yourself |
| AI twin (digital twin, personal avatar) | A few minutes of footage of a real person, in one take | A consent video recorded by that same person | Your own gestures and expressions, replayed | A real creator or founder cloning themselves for volume |
| Scene model with dialogue | Reference images plus a spoken line in the prompt | None | Whole scene, body and camera, in clips of 10 to 15 seconds | Short talking beats inside lifestyle scenes |
"Custom avatar" is the label that causes the confusion. Vendors use it for anything that is not from the stock library: HeyGen's pricing page counts custom avatars as video avatars, which are twins, while its photo avatars are listed separately. When a plan says "1 custom avatar", check which kind it means before you plan a persona around it.
The bottom-left cell is the one persona operators trip over. HeyGen's Digital Twin FAQ says the consent video must feature the same person as the footage and that content generated by third-party AI platforms is not supported. Synthesia's docs require a live consent recording to verify identity, even for a photo-based personal avatar, and limit photo avatars to paid plans. A fictional persona cannot pass either check, so its talking videos come from a photo avatar or a scene model.
The spokesperson video system, end to end
Whatever tool you choose, the pipeline has the same six parts. Most disappointing avatar video fails at the first or the third, not in the generation.
- 01Source image
One sharp, front-facing still of the locked face
- 02Avatar
Photo avatar or twin, created once and reused
- 03Script
Written to be spoken, with a hook and one idea
- 04Voice
One designed or cloned voice, the same every time
- 05Generate
Avatar speaks; b-roll and captions go on top
- 06Check, label, post
Face and voice check, then the AI label
How to make an AI talking avatar from a photo
This is the route a persona account takes. The steps below follow HeyGen's Photo Avatar guide; Hedra and D-ID follow the same shape with fewer menus.
- Pick the source image. Front-facing, sharp, with eyes and mouth clearly visible and the face filling a large part of the frame. HeyGen's guide warns against anything covering the mouth and against heavily stylised faces. Hedra asks for a clear, close-up portrait.
- Create the avatar. In HeyGen: Avatars, New Avatar, then Upload Photo or Design with AI. Name it after the persona and submit. HeyGen validates the look before you can use it.
- Add looks from your own stills. A look is another image of the same avatar in a new outfit or setting. Upload stills you made from your character sheet rather than prompting new ones inside the tool.
- Attach the voice. Pick or clone one voice and keep it. The free plan includes one voice clone; our guide to designing a persona voice covers a fictional voice, and cloning your own covers a real one.
- Paste the script and choose the engine. HeyGen's credit table lists a photo look at 7 credits a minute on Avatar III and 16 on Avatar IV.
- Generate a short test first. Fifteen seconds is enough to check the mouth, the eyes and the voice before you spend a full minute of credits.
That shared source image is the link between this guide and the scene side of the workflow, which is covered step by step in our grid-sheet method for AI video character consistency.
What an AI talking avatar costs per minute
Avatar tools sell credits, and the avatar model you pick changes the price of a minute more than tenfold on the same plan. The table divides each entry paid plan by its published rate, assuming monthly billing and every credit spent on avatar video.
| Tool and avatar type | Published rate | Minutes on the entry plan | Cost per finished minute |
|---|---|---|---|
| HeyGen Avatar III, video look | 4 credits a minute | 150 | about $0.19 |
| HeyGen Avatar III, photo look | 7 credits a minute | about 85 | about $0.34 |
| HeyGen Avatar IV, photo look | 16 credits a minute | 37.5 | about $0.77 |
| Synthesia Starter | about 100 credits a minute, after 10 included minutes | about 27 | about $1.05 |
| HeyGen Avatar IV, video look | 31 credits a minute | about 19 | about $1.50 |
| HeyGen Avatar V | 48 credits a minute | 12.5 | about $2.32 |
| Hedra Character-3, API list price | 2.5, 5 or 6.25 cents a second at 540p, 720p or 1080p | pay as you go | $1.50, $3.00 or $3.75 |
| Google Flow, Gemini Omni Flash 720p | 15 credits per 10-second clip | 50 free credits a day is about 30 seconds | 90 credits; plans are sold in credits |
How the maths works: HeyGen Creator is $29 a month for 600 credits, so a credit is about 4.8 cents. Synthesia Starter is $29 billed monthly; every Synthesia plan includes 10 video minutes and 500 credits, Starter adds 1,250 credits, and Synthesia puts video at about 100 credits a finished minute. Hedra's figures are the per-second API prices on its Character-3 page; its subscription plans are sold in credits whose per-second value the pricing page does not state. Checked October 2026.
Source: HeyGen pricing and credit table, Synthesia pricing and October plan note, Hedra Character-3 page; checked October 2026. Entry plans, monthly billing, first take.
Three things move your real number:
- Retakes. If you keep one take in two, double every figure. Test on a 15-second cut before generating the full script.
- Avatar seconds, not video seconds. HeyGen bills the avatar's generated time, and Synthesia notes its rate varies with avatar screen time. A video that cuts to b-roll for half its length costs about half.
- The model tier. Avatar V is three times the price of an Avatar IV photo look. Use the expensive model for the videos that carry the account, not for every daily clip.
If the avatar is replacing a human creator in ads, our free AI UGC ad cost calculator compares the two at your monthly volume.
Best AI talking avatar generator, by job
There is no single winner, because the tools are built for different faces. These picks rest on what each vendor documents, not on a scored test.
- A fictional persona, on a budget: HeyGen Photo Avatar on Avatar IV. It accepts an uploaded or AI-designed image, the free plan allows up to three photo avatars, and the photo look is one of the cheaper rates in the table.
- A long monologue from one image: Hedra. Its page says Hedra Avatar runs up to ten minutes from a single audio file and that you can use a photo of a presenter or a generated character. Paid plans carry commercial rights; free output is watermarked and non-commercial.
- Cloning yourself for volume: a twin. HeyGen's help center asks for at least two minutes of footage in one continuous take; Synthesia asks for one to five minutes. Both need the consent recording.
- Training and business video with stock presenters: Synthesia, whose paid plans list more than a hundred stock avatars and full HD downloads.
- A watermark you can live with: not D-ID if you need a clean frame. Its FAQ says trial and Lite videos carry a D-ID logo, Pro and Advanced carry a generic AI watermark, and Enterprise can customise the mark but not remove it. Videos are capped at five minutes.
- Talking beats inside real scenes: Kling or Google Flow, covered in how to make AI influencer videos.
- Local and free: InfiniteTalk, an Apache 2.0 model that drives an image or a video from audio. Its repository says results from a single image hold for about a minute before colour shifts set in, and it needs a serious GPU.
AI talking avatar free: what the free tiers allow
Free tiers are for testing a face and a format, not for running an account. The short version, checked October 2026:
- HeyGen Free: three videos a month of up to one minute (its help center says one to three, depending on region), up to three photo avatars and one voice clone. Watermark removal starts on Creator.
- Synthesia Basic: 10 minutes a month with nine stock avatars. Videos carry a watermark and are shared by link; downloads and personal avatars are paid.
- Hedra Free: no credit card; outputs are watermarked and for non-commercial use.
- Google Flow: 50 credits a day, enough for three 10-second 720p clips with Gemini Omni Flash.
None of them will carry a daily posting schedule, and watermarks or non-commercial terms rule out sponsored work on most. Treat the free month as a test of whether your persona still survives being animated. The full table, including Google Vids, D-ID, Vidnoz and the local route, is in our comparison of free talking avatar generators.
Keeping one persona consistent across avatar videos
An avatar fixes the face inside a single video. The drift happens between videos, in three places.
- A new look prompted inside the tool for every video
- A different voice, or the same voice at a different speed
- Avatar built from one still, scene clips from another
- Switching avatar engine between videos
- Wide framing that shrinks the face
- Looks uploaded from stills made off one character sheet
- One voice, one pace, saved as the avatar default
- One source image for the avatar and the scene element
- One engine per format, changed on purpose
- Chest-up framing; Synthesia notes lip sync works best framed closer
Looks deserve the most care. HeyGen lets you add a look by uploading reference photos, by describing it in a prompt, or by picking an inspiration image, at one credit per generated look. The prompt route is quick and it is also a fresh image generation each time, which is exactly how a face wanders. Build outfit and location stills from your own references first, then upload them.
The AI Influencers program goes deeper on this part: one character sheet feeding the avatar, the scene clips and the voice, and a weekly schedule that mixes all three.
A one-week plan for your first avatar video
- 1Day 1: choose the type
Fictional face: photo avatar. Real face with footage: twin. Write down which tool and which engine.
- 2Day 2: prepare the source
One front-facing still for a photo avatar, or a continuous take of a few minutes plus the consent video for a twin.
- 3Day 3: lock the voice
Design or clone one voice. Generate a 20-second sample and keep the settings.
- 4Day 4: write three scripts
About 130 to 150 spoken words fills a minute at an ordinary pace. One idea each, hook in the first line.
- 5Day 5: test at 15 seconds
Generate the first lines of each script. Check mouth, eyes and voice. Fix the source, not the edit.
- 6Day 6: generate and cut
Full takes, then cut to b-roll and captions so the avatar is not on screen for the whole minute.
- 7Day 7: label and post
Turn on the platform AI label, post one video and keep the other two for the week.
Mistakes that make avatar video look fake
- A script written to be read. Long sentences and written-English phrasing sound robotic in any voice. Write the way the persona would say it out loud.
- One unbroken minute of face. Viewers study a static presenter. Cut away every few seconds; it also lowers the bill.
- A source image with the mouth closed or hidden. Synthesia recommends a smiling photo with teeth visible for lip sync; HeyGen warns against anything covering the mouth.
- Posting free-tier output for a client. Hedra's free output is non-commercial, and watermarks stay on most free plans.
- Using a face you do not own. Only build avatars from a generated persona that is yours or a real person who has consented. Our AI likeness rights guide explains the risk.
Consent, likeness and labels
The consent recordings are not friction to route around; they are the reason these tools are usable for business at all. D-ID says outright that it keeps a watermark on every plan so viewers know the video is synthetic. Realistic avatar video also falls under the platforms' AI-label rules. Our Instagram AI label guide shows where the setting is and when it applies.
- Decide: photo avatar for a fictional face, twin for a real one
- Pick one source image: front-facing, mouth visible, face filling the frame
- Create the avatar once and name it after the persona
- Choose one voice and save it as the default
- Work out your cost per minute from the engine you will actually use
- Generate a 15-second test before any full script
- Upload looks from your own stills instead of prompting new ones
- Check the plan terms before anything sponsored goes out
- Turn the AI label on when you post
AI talking avatars: FAQ
What is an AI talking avatar?
An AI talking avatar is a digital presenter that speaks a script you type or an audio file you upload, with the mouth, face and head movement generated to match. It comes in three forms: a stock presenter from a tool's library, a photo avatar built from one image, and an AI twin cloned from recorded footage of a real person. The output is a spokesperson video with nobody on camera.
Can I make an AI talking avatar from a photo?
Yes. HeyGen builds a Photo Avatar from one uploaded image or from its Design with AI option, Hedra animates a single portrait with a script or audio file, and D-ID accepts an uploaded face image of up to 10 MB. Use a sharp, front-facing image where the eyes and mouth are clearly visible and the face fills much of the frame. Checked October 2026.
Is there a free AI talking avatar generator?
Several, each with a catch. HeyGen's free plan lists three videos a month of up to one minute. Synthesia's free Basic plan includes 10 minutes a month, watermarked and shared by link rather than downloaded. Hedra's free outputs are watermarked and for non-commercial use. Google Flow gives 50 credits a day, about 30 seconds of 720p talking clips. Checked October 2026.
How much does an AI talking avatar cost per minute?
On HeyGen's Creator plan, $29 for 600 credits, a photo avatar on Avatar IV uses 16 credits a minute, about $0.77 a finished minute, and Avatar V uses 48 credits, about $2.32. Synthesia Starter works out near $1.05 a minute if every credit goes to video. Hedra lists Character-3 from 2.5 cents a second through its API. Checked October 2026; retakes multiply all of it.
What is the difference between a photo avatar and an AI twin?
A photo avatar is generated from one still image, so the tool invents the motion; it takes minutes to make and works for a fictional character. An AI twin, also called a digital twin or personal avatar, is trained on a few minutes of footage of a real person, so it replays their real gestures and expressions. Twins need a consent recording from that same person.
Can I make a digital twin of an AI-generated persona?
Not through the twin features we checked. HeyGen's Digital Twin FAQ says the consent video must feature the same person as the footage and that content generated by third-party AI platforms is not supported. Synthesia requires a live consent recording for personal avatars. A fictional persona cannot record either, so use a photo avatar built from your persona still instead.
What is the best AI talking avatar generator?
It depends on the face. For a fictional persona, a photo-avatar tool that accepts a generated character, such as HeyGen or Hedra, fits best. For cloning yourself at volume, a twin feature such as HeyGen Digital Twin or a Synthesia personal avatar gives the most natural motion. For training video with stock presenters, Synthesia has the larger business feature set. This is a documentation-based view, not a scored test.
Your persona can talk. Now give it a system to talk every week.
AI Influencers, included in All Access, covers the pipeline behind the avatar: the character sheet, a persona voice, talking and scene video, labelling and monetization. One subscription covers all four programs, live coaching and the private community.
Price the video before you make it
Compare AI spokesperson video with human creators at your monthly volume, and get avatar model and pricing changes in the free Telegram channel.