How to Make a Lip Sync Video (Photo or Footage)
Lip sync videos come from two starting points: a still photo you want to bring to life, or footage you want re-dubbed to a new track. The process is short either way — pick your source, pick your audio, generate — but a few material choices make the difference between a result that lands and one that looks off. This guide walks through both paths with the settings that matter.
Quick answer: To make a lip sync video: choose a clear, camera-facing photo (or a talking-head clip with one visible face), trim your audio to a strong 10-second moment, and run it through an AI lip sync tool. The result is an MP4 where the mouth matches the audio — best material is front-facing faces and vocal-forward audio.
Pick the right starting point
Photos sing (and speak) best when the mouth is clearly visible, the face is reasonably large in frame, and the lighting is even. Front-facing portraits beat profiles; sharp beats blurry; sunglasses and hands-over-mouth give the model nothing to animate. For footage, one speaker, face large in frame, steady camera — a webcam take works fine. Group shots and crowded scenes are the hardest case for any lip sync model, so if the result matters, start with a single subject.
Prepare the audio — 10 seconds is the sweet spot
Every lip sync model performs best on short, clear audio. 10 seconds captures a sentence, a hook, or a chorus line — and keeping it short means the model can spend its capacity on quality instead of length. Trim to the strongest moment: the punchline, the chorus, the one line that makes the clip. Use vocal-forward audio; heavy instrumentals or noisy backgrounds give the model a harder time finding the syllables. Any format works (MP3, WAV, M4A, OGG) — this site converts to a clean WAV in your browser before anything is uploaded.
Choose the tool for your case
If you have a photo and want it to talk or sing, use a photo lip sync tool (talking photo / singing photo). If you have footage with the wrong audio, use a video lip sync tool instead — it re-renders just the mouth region over the existing footage, so the scene and body stay as filmed. The two jobs use different engines, and picking the right one is most of the result.
The step-by-step (photo path)
- Pick a camera-facing portrait — mouth visible, even light
- Trim your audio to 10 seconds or less, vocal-forward
- Upload both to the talking photo maker on this site
- Generate — the render takes about a minute
- Download the MP4; post it anywhere
The step-by-step (footage path)
- Pick a clip up to 10MB with one clear speaker
- Trim the new audio to 10 seconds; match its energy to the scene
- Upload both to the video lip sync tool
- Generate — only the mouth region is re-rendered
- Download the re-dubbed MP4
What separates a good result from an uncanny one
| Choice | Works well | Works poorly |
|---|---|---|
| Face angle | Front-facing, slight angles | Extreme profiles, face turned away |
| Frame size | Face fills a good part of frame | Tiny face in a wide shot |
| Audio | Clear vocals, one speaker | Instrumental, crowd noise, overlapping voices |
| Length | 5–10 focused seconds | Long rambling takes |
| Lighting | Even, soft | Harsh shadows on the mouth, backlit |
| Obstructions | Nothing across the mouth | Hands, microphones, sunglasses, masks |
Rights and consent — the short version
Use photos of yourself, pets, people who've agreed, and content you made. Don't put words into someone else's mouth without their consent — it's both a terms-of-service line and, depending on where you live, a legal one. For public posts with music in them, check the platform's rules on copyrighted audio before you upload.
What this site doesn't do
- We don't lip sync material you don't have the rights to — that's in our terms, not just advice.
- We don't promise 'undetectable' results. This is a creative tool, not a deception tool.
- We don't store your photos, audio, or video — everything is processed for your render only.
Frequently Asked Questions
How long does generation take?
Can I lip sync to music?
Why does my audio get trimmed?
Can I lip sync a video that has no face in it?
More tools used in this guide
Upload a photo and 10 seconds of audio, and the face comes alive — mouth movements matched to the voice. 1 free video a day, no account, nothing installed.
Upload a video, drop in new audio, and the mouth in the footage is re-synced to the new soundtrack — dub a clip, fix a take, or swap the lines.
Make a photo sing: upload a portrait plus a song clip (up to 10 seconds) and the face performs it, mouth matched to the music.
More guides
Free tiers, watermarks, limits and pricing across the main AI lip sync tools — what each one actually gives you.