Preserve the source clip
Framing, body movement, wardrobe, and scene come from the video you upload.
Upload an existing video, add an MP3, WAV, recording, or generated voice, then match the visible mouth movement to the new speech for dubbing, localization, or a revised creator message.
Opens Lip Sync Video in Studio. Upload an existing video and a built-in voice or authorized audio track, then confirm usage rights.
1 · Source video
Authorized template · visuals to preserve
“Hola, hoy te enseño cómo convertir una idea sencilla en un video listo para compartir.”
2 · New audio
Kie Gemini TTS · generated voice
3 · Lip-synced result
Kie API output · reviewed 9:16
Existing video
The source clip supplies the original framing, body movement, expression, and visual timing.
Voice or owned audio
Use a built-in voice or an audio track you have permission to reproduce.
Aligned mouth movement
The workflow changes speech timing while aiming to preserve the source performance.
An AI lip sync generator is a finishing workflow. It starts with footage that already exists, then aligns visible mouth movement to a new voice track or authorized audio file.
Framing, body movement, wardrobe, and scene come from the video you upload.
Use clear audio with one primary speaker and minimal noise or overlapping voices.
A new language can reuse approved visuals, but translation and timing still need human review.
Inspect mouth closure, consonants, pauses, head turns, occlusion, and the end of the clip.
Keep the video and audio clean. The workflow aligns them; it does not verify your rights or the truth of the message.
Choose a clip with a visible mouth, one speaker, stable framing, and no heavy face occlusion.
Select a built-in voice or upload a clean track you own or may use.
Confirm rights, generate the result, then watch the full clip with sound before publishing.
This is one complete task, not two recycled lookalike clips. The source keeps the original scene and performance; the generated audio changes the message; the Kie result changes visible mouth timing to match it.
Spanish localization
The same source footage is intentionally preserved. Compare the mouth movement while listening to the new Spanish track; that before/audio/after relationship is the proof.
Source video
Original framing and movement
“Hola, hoy te enseño cómo convertir una idea sencilla en un video listo para compartir.”
New audio
Gemini TTS · Spanish
Synchronized output
New speech and mouth timing
Lip sync quality depends on a readable face and an understandable speech track. Improve the inputs before asking the model to reconcile them.
If you need to create the visual performance too, start in a different workflow.
Existing video + new speech. Best for dubbing, localization, and approved message changes.
Saved portrait + script or audio. Best when you need a new speaking performance.
Appearance image + reference video. Best when body movement—not speech—is the main control.
Each route starts from a different input or solves a different production task.
Use a clear single-speaker clip, clean audio, and the required rights confirmation before creating the result.
Lip sync an existing video