AI Lip Sync Generator

Upload a portrait and an audio track, and get back video of that face saying those words, with the mouth and expression matched to the speech. The clip is as long as the audio you give it.

Make a talking video
Portrait (required)

0/1 reference images

A clearly visible, front-facing face works best.
Audio (required)
The clip is as long as this track. No voice yet? Generate one.
Estimating…

What makes a good source photo

The portrait does more for the result than the settings do.

  • Front-facing and unobstructed — a face turned far to one side gives the model less to work with.
  • Well lit and in focus. Softly-lit faces animate more convincingly than heavily shadowed ones.
  • One face in frame. A group photo leaves the choice of subject ambiguous.
  • Reasonable resolution. Upscaling a small crop tends to show once the face starts moving.

Where the audio comes from

Any speech track works — including one you generate here.

Upload a recording you already have, or write a script and generate narration with the voice generator first, then bring the file back here. Because the clip's length comes from the audio, timing the script before generating is the cheapest way to control how long the finished video runs.

How it is billed

Per second of output, and the audio sets the length.

Lip-sync is priced per second, like most of the video engines, so a longer audio track costs proportionally more — and because the audio decides the clip's length, trimming it is what controls the cost. Audio can run up to five minutes. The estimate on the form is the same figure the account is charged, calculated by the same code that bills it.

FAQ

Generator questions

Can't find what you're looking for? Visit help and support or email hello@usescenes.com and a human will reply.

Up to five minutes, and the clip matches the audio exactly. Longer audio costs proportionally more, so trimming the track is what controls the cost — and a shorter first attempt is the cheapest way to check the portrait works.