Lip Sync

Lip sync matches a character's mouth movements to spoken audio. Learn how AI video models generate synced speech natively or from a voice track.

Open the full generator

See it in action

One real prompt and the result it produced on Molyin. Remix it to start from here.

  • Kling 3.0
  • Image to video
  • 10s
  • 1080P

In a high-rise office building, the man leaned back, wearing a tired, disdainful expression, and said in Cantonese: 「其实……我真系唔系好 buy 你呢个 logic 啰。成个 proposal 根本 align 唔到我哋个 core value。你个 flow 咁乱,点样去 convince 个 client 呀?不如你返去 re-think 下个 angle,听朝早我要见到个 final version。」

Lip sync is the precise matching of a character's mouth movements to spoken audio. In AI video generation it comes in two flavors: native speech generation, where the model creates the voice and the matching lip motion together from dialogue you write in the prompt, and audio-driven animation, where you upload a voice recording and the model animates a face to speak it.

How it works

Speech is made of phonemes, and each phoneme maps to a visible mouth shape called a viseme. Modern video models learn this mapping directly from vast amounts of talking footage, so they don't just flap the mouth on the beat — they shape the lips, jaw, and facial muscles to match each syllable, along with natural blinks and head motion. Models with native audio generate the voice track and the frames in a single pass, which keeps timing perfectly aligned.

What it's best at

  • Talking-head content — spokesperson clips, UGC-style ads, and product explainers without a camera crew.
  • Dialogue scenes — characters that deliver written lines with matching expressions.
  • Multilingual delivery — the same character speaking different languages, each with correct mouth shapes.
  • Bringing stills to life — a portrait photo that speaks your script.

What a good input includes

Write dialogue in quotation marks inside the prompt and keep lines short — one or two sentences per shot reads most naturally. Name the tone ("warm, conversational") and, for audio-driven generation, use a clean voice recording without background music.

Questions about Lip Sync

Do I need to record audio for lip sync?

Not necessarily. Models with native speech generate the voice from dialogue you write in quotation marks. If you already have a recording, image-to-video with a driving voice track animates a face to speak it.

Which languages does lip sync work in?

Native-speech models handle multiple languages and regional accents — the sample above speaks Cantonese. Write the line in the language you want spoken and keep it to one or two sentences per shot.

Go deeper

What is Lip Sync? | Molyin