
Lip Sync
Lip sync means that mouth movements on screen exactly match the sounds being heard. In video editing and AI tools, it is one of the most difficult tasks, because the human eye notices even tiny discrepancies.
When a person speaks, their lips move in a fixed relationship to what can be heard. An “M” requires closed lips, an “O” requires a rounded mouth. Lip sync means: image and sound match exactly. It is missing, for example, when a video stutters and the sound gets ahead. We notice this immediately, even though we don’t consciously pay attention to the mouth. Even a shift of about a tenth of a second feels disruptive to many viewers.
Why even milliseconds stand out
When listening, the brain constantly evaluates the image as well. Anyone who sees the mouth of the person they’re talking to understands speech in noisy environments much better. This connection is so deeply ingrained that discrepancies are noticeably unpleasant. Experts speak of the “uncanny valley”: something looks almost real, and precisely because of that it appears wrong.
For broadcasting and streaming there are therefore fixed tolerances. Common recommendations allow the sound to lead the image by about 40 milliseconds or lag behind by around 60 milliseconds. If this is exceeded, the broadcast is considered faulty. In video conferences the tolerance is greater, because network delays are expected anyway.
The topic becomes economically interesting because of translations. A film dubbed into twenty languages has so far required twenty recording studios and many weeks of work. If software can adapt mouth movements to a new language, these costs drop sharply. That is exactly why streaming providers and advertising companies are investing in such tools.
From phoneme to mouth shape
The classic approach begins by breaking down sound into the smallest sound units, called phonemes. Each phoneme is assigned a typical mouth position, called a viseme. Several sounds share the same mouth shape: “P”, “B” and “M” look practically identical from the outside. An animation program then lines up these mouth shapes in the correct rhythm.
Modern AI systems work differently. They receive an audio track and a video of a person and generate new images of the lower face directly from that. In doing so, they learn from huge amounts of real video footage which mouth movement belongs to which sound. Nobody programs the rules by hand; the model derives them from the examples.
The difficulty lies in the details. Besides the lips, the jaw, cheeks, and tongue move as well, and shadows and wrinkles appear. The rhythm often doesn’t match either: an English sentence is usually shorter than its German translation. Good systems therefore stretch the speech slightly or adjust the translation to the available length.
From dubbed versions to deepfakes
The topic is encountered most often with dubbed films and series. Translators have always looked for phrasings that fit the content and at the same time match the mouth movements of the original. Video games and animated films also use automatic methods to make characters speak. The same applies to virtual assistants with a face.
In the news, lip sync mainly comes up in connection with deepfakes. These are videos in which a real person says things they never actually said. Convincing lip movement is the crucial ingredient for this. Conversely, detection programs look precisely for small errors in this area, such as teeth that blur unnaturally.
A common misconception is that this is only about image editing. In fact, voice, translation, and animation are closely interlinked. Anyone creating a foreign-language version usually has to solve all three areas at the same time. That is why many providers bundle voice cloning, translation, and lip adaptation into a single service.