AI Lip Sync Generator
Upload a portrait and an audio track, and get back video of that face saying those words, with the mouth and expression matched to the speech. The clip is as long as the audio you give it.
0/1 reference images
What makes a good source photo
The portrait does more for the result than the settings do.
- Front-facing and unobstructed — a face turned far to one side gives the model less to work with.
- Well lit and in focus. Softly-lit faces animate more convincingly than heavily shadowed ones.
- One face in frame. A group photo leaves the choice of subject ambiguous.
- Reasonable resolution. Upscaling a small crop tends to show once the face starts moving.
Where the audio comes from
Any speech track works — including one you generate here.
Upload a recording you already have, or write a script and generate narration with the voice generator first, then bring the file back here. Because the clip's length comes from the audio, timing the script before generating is the cheapest way to control how long the finished video runs.
How it is billed
Per second of output, and the audio sets the length.
Lip-sync is priced per second, like most of the video engines, so a longer audio track costs proportionally more — and because the audio decides the clip's length, trimming it is what controls the cost. Audio can run up to five minutes. The estimate on the form is the same figure the account is charged, calculated by the same code that bills it.
FAQ
Generator questions
Can't find what you're looking for? Visit help and support or email hello@usescenes.com and a human will reply.