How the voice is generated from text
Every scene carries its own text, which is turned into audio using the voice chosen in the video settings: tone, gender, speaking pace, language and accent. The voice does not read mechanically — it respects punctuation, pauses at the end of a sentence and emphasizes whatever you marked as important. You can listen to the result on a single scene and adjust the pace or the pauses before the full render.
How lip sync works
Once the voice is generated, the platform matches the avatar’s lip movement to the sounds actually spoken, not to the raw text. That is why the sync stays correct when the same script is delivered in another language, with a different syllable structure. The result is a presenter that appears to be speaking, not an image with an audio file pasted over it.
Dubbing and translation in one click
An existing video can be dubbed into another language without rebuilding it: the platform translates the script, generates the voice in the new language, re-syncs the lips and rewrites the subtitles. The scenes, the edit and the brand elements stay identical, so you get the same information for another market in minutes rather than days. The translation stays editable — you can correct it against your internal terminology before rendering.
Pronunciation, proper names and technical terms
Brand names, abbreviations and technical terms are where generated voices most often get it wrong. That is what the account-level pronunciation dictionary is for: you write how a term should be read, and every future video says it correctly. You can also add targeted pauses or emphasis directly in the scene text.