
Zero-Talent Production
Zero-Talent Production refers to the making of films, music, advertising, or texts without any humans performing in front of the camera or at the microphone. Voices, faces, and bodies are generated entirely by computer.
Until now, making a commercial required people without exception. Someone had to stand in front of the camera, someone had to speak the lines, someone had to sing. The industry calls these people “talent” — meaning the performing actors, voice artists, and musicians. Zero-Talent Production refers to productions in which not a single one of these people appears anymore. Instead, faces, bodies, and voices are generated by computer programs that have learned from vast amounts of sample material. The term says nothing about quality — only that no one is being filmed or recorded anymore.
What’s at stake: fees, rights, credibility
The economic appeal is obvious. A commercial with a real shoot can quickly cost six figures. Camera crew, location, catering, fees, and travel expenses all add up. A purely synthetic version can be produced for a fraction of that. Above all, it can be modified endlessly — for instance, into twenty languages for twenty markets.
For actors and voice artists, this is a serious threat. This was precisely the issue at the heart of the major US actors' union strike in 2023. A key point of contention was whether studios should be allowed to scan a performer’s likeness once and then reuse it indefinitely afterward. The union succeeded in requiring consent and additional payment for such use. However, such rules only apply where collective bargaining agreements are in effect.
On top of that, there’s a trust problem. When a face in a video never actually existed, viewers lose an important point of reference. Until now, moving images were considered fairly reliable evidence. In the EU, the AI Act therefore requires that artificially generated content be labeled as such.
From text prompt to finished spot
It usually starts with a written command, known as a prompt. It describes in plain language what should be seen. A video model then generates image sequences from it. Such models have learned from millions of videos how fabric, hair, and light behave in motion. Well-known examples include Sora by OpenAI or Veo by Google.
The voice is generated separately. A language model turns written text into audible speech, including intonation and pauses. It can be trained either on an invented voice or on a recording of a real person. The second case is called voice cloning and is legally sensitive, because a person’s voice is part of their personality rights. In the end, image and sound are combined in an editing program, often by a single person.
A common misconception is that no one works on this anymore. In reality, the work simply shifts. Instead of camera operators and performers, you need people who craft prompts, sort through results, and fix errors. Often a hundred attempts are needed before five usable seconds emerge. Hands, on-screen text, and consistent faces across multiple shots remain typical weak points.
Where synthetic performers are already appearing
Development is furthest along with short advertising clips for social networks. There, quantity matters more than perfection, and the videos play on small screens anyway. Explainer videos for companies and training materials are also increasingly being produced this way. The provider Synthesia has specialized precisely in this area.
In music, there are platforms that generate entire songs, complete with vocals, on command. Such tracks are already landing on streaming service playlists. In the news sector, some broadcasters use virtual presenters, for example for weather reports. In big-budget feature films, however, use remains limited so far and is mostly confined to backgrounds and post-production.
In business news, you’ll typically encounter this term in connection with costs. Agencies and studios calculate how sharply production budgets can be reduced. It’s important to distinguish this from a deepfake: a deepfake deliberately recreates a real person, often without their knowledge. Zero-Talent Production can do without any real-life template at all and, as a technique, is initially legal.