Ablaufskizze in vier Schritten: Person im Markeranzug im Kameraring, daneben die erfassten Punktwolken der Kameras, daraus berechnetes Gelenkskelett, schließlich die fertig animierte digitale Figur.

Full-Body Capture

Full-body capture refers to techniques used to record the movements of an entire human body and translate them into digital data. This data is then used to drive a game character, a film character, or an avatar in a virtual environment.

Full-body capture roughly translates to “whole-body capture.” It refers to a process that measures the movements of a real person and converts them into numbers. Afterwards, a computer knows at every moment where the head, shoulders, elbows, hips, knees, and feet are located. These numbers can be transferred onto a drawn character, which then moves exactly like the person in front of the camera. The difference from a normal video is crucial: a video only shows images, whereas here measurement data is created that can be freely reused. From the same recording, one can later animate a dragon, a robot, or a cartoon character.

Why animation without real bodies remains tedious

Animating a character by hand is enormously labor-intensive. An animator has to determine, for every second, how each joint rotates. Just a few steps of a character can take hours to produce. And even then, the result often looks artificial, because human walking contains many small irregularities that no one consciously recreates.

Full-body capture reverses this relationship. An actor walks through the room once, and the movement is captured within seconds. This is precisely why major film productions and game studios have relied on it since the 2000s. Characters like Gollum from “The Lord of the Rings” or the Na’vi from “Avatar” are based on such recordings.

Economically, the term has since become interesting beyond the entertainment industry as well. Anyone wanting to build avatars for virtual meetings, fitness apps, or training simulations needs realistic body data. The market for the necessary hardware and software is growing accordingly, and many startups are working to make the process cheaper.

From markers on a suit to camera-based estimation

The classic method works with a tight suit covered in small reflective spheres, known as markers. Special cameras equipped with infrared light, invisible to humans, are positioned around the room. Each camera sees the spheres as bright points. From the viewing angles of multiple cameras, a program calculates the exact position of each sphere in space. The principle resembles spatial vision: two eyes viewing from slightly different angles provide depth, and many cameras provide very precise depth.

From these points, the software builds a skeleton. This is a simplified stick figure with typically 20 to 60 joints. At every moment, the system stores how far each joint is rotated. This data trace is then transferred onto a digital model, a process called retargeting. Because the character may have different proportions than the human, the software must adjust the movements accordingly.

Newer systems work without a suit. A neural network, that is, a program trained on examples for pattern recognition, estimates the joint positions directly from ordinary video footage. It was trained on millions of images in which the correct positions were known. This is significantly cheaper, but less accurate. Occluded body parts, such as an arm behind the back, remain a known problem.

From the film studio to the fitness app

The technique is most visible in feature films and major video games. The fighting and running movements of realistic game characters almost always originate from a capture studio. Sports games also use it, having real professional athletes record their characteristic movements once.

In everyday life, one encounters stripped-down variants. Some VR headsets capture not only the head and hands but also the legs via additional sensors. Fitness apps use a phone camera to analyze whether a squat is being performed correctly. In medicine, gait analysis helps detect abnormal loading patterns after an injury.

In the news, the term often appears alongside avatars, digital twins, and the debate over actors' rights. A common misconception: full-body capture only captures movement, not appearance. Facial expressions are usually recorded separately, a process called face capture. Rendering the skin and face of a character additionally requires a 3D model, which is created in a completely different way.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.