Talking avatar makes a single still image speak. You give it one photo of a face and either create speech from a voice model and script or provide an audio track, and it animates the still so the mouth, expression, and head move in time with the voice. It's part of the Talking Video workstation.
What does Talking avatar do?
It takes one reference image (a portrait, character, product mascot - anything with a face) plus speech audio and produces a video of that image talking. You can create the speech in the page by choosing a Voice model and typing a Script, or switch to Select audio and provide an existing audio track.
How do I make a talking avatar?
Open the Talking Video workstation and make sure Talking Avatar is selected (the desktop mode toggle, or the mobile Choose a tool sheet).
Under Reference image, add your face image (see below).
In the Audio area, choose either Create audio or Select audio.
For Create audio, pick a Voice model, optionally an Overall mood, and type the Script it should speak; you can tag parts of the script with a mood chip and adjust the Speed slider. For Select audio, upload an audio file or generate a reusable voice clip.
Pick a Quality and Resolution. Best uses 720P (HD), Ultra uses FHD, and UltraS supports FHD, 720P, and QHD. Unavailable resolutions stay visible and show which quality unlocks them.
Check the per-second cost and estimate next to Generate, then tap Generate.
How do I add the reference image?
In the Reference image box you have two buttons:
Select content — pick an image you've already made or saved, from your own generations (followed by curated inspiration) or the community feed.
Upload image — upload a file from your device. HEIC photos (common on iPhone) are converted automatically.
Once added, the image shows in the box with a delete control so you can swap it. You need to be signed in to select or upload.
For the most natural result, start from a clear, well-lit, front-facing photo with the mouth closed or in a relaxed/neutral position (or use a model you created). The cleaner the face, the better the animation tracks it.
How do I add the audio?
The Audio area has two tabs:
Create audio - pick a Voice model and type a Script of up to 2,000 visible characters (the counter under the box counts your text, not mood-chip markup; going over turns the counter red and blocks Generate). While you type, the duration and total cost update immediately from a language-aware local approximation and carry an
≈marker. After you pause for about half a second, the app requests an exact backend quote and replaces the approximation when it arrives.
Next to the voice model, Overall mood (optional) sets one mood for the whole audio - None keeps the natural delivery, Neutral gives an even delivery, and the rest (Happy, Sad, Angry, Fearful, Disgusted, Surprised) color the entire read.
Inside the script you can also tag individual lines: the + Mood button under the box inserts an inline mood chip - pick its mood, type the line it should affect, and confirm with the enter icon (click a confirmed chip's text to edit it again, or its × to remove it). If you first select some script text, the button becomes + Mood to selected content and wraps exactly that selection (including the text of any chips inside it) into the new chip. The script box itself doesn't accept the <, >, or | characters - they are reserved by the chip format and are dropped as you type.
The Speed slider under the script sets the playback speed from 0.5× to 2× (default 1.0×). Speed changes the synthesized duration - and therefore the cost - so after you stop dragging, the app re-requests the exact quote for the new speed.
Select audio - provide an existing audio track. Use Generate audio to open the text-to-speech generator, type a script, and create a voice clip without leaving the page, or use Click to upload audio to upload your own audio file. See Generate audio (text-to-speech).
In Select audio, after an audio track is added it appears as a playable card with a Delete button if you want to replace it.
What are the audio length limits?
For Talking avatar, the speech audio must be at least 1 second and no longer than 3 minutes. UltraS has a stricter inclusive range of 4 to 180 seconds. If the exact synthesized duration or the selected audio's trimmed duration is outside that UltraS range, clicking Generate shows "UltraS Talking Avatar supports audio from 4 to 180 seconds." and does not send a generation request. In Select audio, the app also blocks uploaded or selected audio outside the general range and tells you the limit. The hint under the audio box reads "Please use audio shorter than 3 minutes." Because the output length follows the audio, the audio length is effectively your video length.
What does the Quality setting do?
Quality picks the render tier for your avatar. Talking avatar offers three tiers: Best, Ultra, and UltraS. The Resolution control is always visible: Best uses 720P (HD), Ultra uses FHD, and UltraS supports FHD, 720P, and QHD. Both quality and, for UltraS, resolution affect the per-second cost. You'll find these controls in the bottom bar next to Generate.
After generation, the finished video's detail page repeats the saved settings on both desktop and mobile: Quality (plus, for UltraS results, Resolution), and for Create audio results also the Voice model, the Duration, the Overall mood (shown only when a mood was chosen — records made before the field existed, or with None, simply omit the tile), and the Speed (e.g. 1.5×; omitted on records that predate the field). The Script box shows your text with any inline mood chips rendered as chips — the mood label in a small pill followed by the tagged line. This lets you confirm the original settings before using Reuse or Remix, both of which restore the script (chips included), the overall mood, and the speed into the workstation.
How much does a talking avatar cost?
It's billed per second of audio. The bottom bar shows the per-second rate, the duration of your audio, and the total estimate next to Generate before you generate. In Create audio, values marked with ≈ are local approximations while the exact script quote is being calculated, and the Speed slider scales the duration (2× speed roughly halves it); in Select audio, the duration comes from the selected audio track. The rate depends on the selected quality and, for UltraS, the selected resolution. Once the ≈ marker disappears, the displayed duration and total use the current backend quote.
Why can't I generate yet?
Generate is disabled until both inputs are ready. You'll be blocked if: there is no reference image, the image is still uploading or failed, Create audio has no voice model, no script, or a script over the 2,000 visible-character limit, Select audio has no audio, the selected audio is still uploading or failed, or non-UltraS audio/script is longer than the 3-minute limit. For UltraS, the button stays clickable once the inputs are ready; clicking it with audio shorter than 4 seconds or longer than 180 seconds shows the range warning instead of starting generation.
In Create audio, Generate remains clickable while the exact quote is missing, expired, or still being calculated. Clicking it in that state does not start generation; it starts or reuses the current quote request and shows "The duration is being calculated. Please try again shortly." When that request succeeds, the app shows "Duration calculated successfully." If it fails, the app shows the generation error instead of leaving the quote stuck in a calculating state. After the exact quote appears, click Generate again to submit with the same displayed quote.
What happens if an upload fails?
If the image fails, you'll see "Please try again later" under the image and Generate stays off until you re-add a working image. If the audio fails, you'll see "Failed to upload the audio. Please try again later." Just remove the failed item and add it again.
On mobile
The layout stacks vertically: Reference image first, then Audio with Create audio / Select audio tabs. The image and audio inputs work the same as desktop. Quality and Resolution are always shown in the mobile bottom bar with each label above its selector. Unavailable resolution entries are greyed out and include an info icon explaining which quality unlocks them. The 3-minute audio limit and per-second pricing are unchanged.
Related: Talking Video · Lip sync · Generate audio (text-to-speech) · Credits & billing