
Auto-Dubbing
Auto-dubbing refers to the automatic translation of spoken language in videos, where software generates the new audio track. The voice often sounds like that of the original speaker.
When a film from the USA airs on German television, you hear German voices. In the past, actors always sat in a recording studio and re-recorded every sentence. With auto-dubbing, software takes over this task. It listens to the original audio, transcribes it, translates the text, and generates a new spoken audio track from it. What makes this special: the artificial voice can sound like the real person in the video, just in German, Spanish, or Japanese. This turns a single recording into versions in twenty languages within minutes.
Why videos are suddenly understood everywhere
Traditional dubbing is expensive. For a cinema film, translators, directors, sound engineers, and voice actors work together for weeks. Costs quickly reach five figures per hour of film. That’s why, until now, only content that made economic sense was dubbed: big films and series. A tutorial video from a small channel never got a French version.
Auto-dubbing shifts this boundary. When a translation costs just a few euros, it becomes worthwhile even for small productions. A video channel with 50,000 subscribers in Germany can reach a global audience with it. The same applies to companies: training videos, product presentations, and support material can be brought to many markets without major effort.
At the same time, a conflict arises. Voice actors are a distinct profession, especially in Germany with its long dubbing tradition. Many fear that their voices will be copied and reused without payment. In the USA, this was exactly a point of contention during the actors' union strikes. Contracts now exist that regulate the use of digitally recreated voices.
Four steps from the original track to the new language version
First, a program converts the spoken language into written text. This step is called speech recognition, and it works like dictating on a phone. It also records who speaks when and at which second each sentence begins. These timestamps become crucial later on.
In the second step, a language model translates the text. A language model is a program that has learned from huge amounts of text how language is structured. In the third step, the system analyzes the original voice and recreates it. Often just a few seconds of audio material are enough for this. The fourth step generates the new audio track and lays it under the video.
The hardest part is timing. A German sentence is usually longer than its English original, sometimes by thirty percent. If it doesn’t fit into the gap, the voice talks over the next scene. That’s why the software shortens phrasing or slightly changes the speaking pace. Some systems go further and alter lip movements through video editing so they match the new text. This is called lip-sync, and so far it only works reliably with calm recordings.
From YouTube channels to video conferencing
YouTube offers automatic audio tracks that you can switch between in the menu. Well-known channels like MrBeast’s have used this for years and reach a millions-strong audience outside the English-speaking world. Streaming providers are also testing the technology for documentaries and older titles, where an expensive studio production wouldn’t pay off.
In video conferences, translation now runs almost in real time. Providers like Zoom and Microsoft have built in corresponding features. The delay is a few seconds, because the system has to wait briefly until a sentence is finished being spoken. Only then can it translate meaningfully.
A common misconception is that auto-dubbing is the same as subtitles. Subtitles are written text that you read. Auto-dubbing generates actual audio that you hear. And so far, it doesn’t replace high-quality dubbing: irony, dialect, and dramatic emphasis are still rarely achieved convincingly by the technology. For news, explainer videos, and corporate communication, however, the quality is often sufficient.