Kreislauf-Schema der Physical AI: Sensoren (Kamera, Abstandsmesser) liefern Daten an ein Entscheidungsmodell, dieses schickt Befehle an Motoren, die Bewegung verändert die Umgebung, die wiederum von den Sensoren erfasst wird. Daneben ein Kasten \"Simulation\" als Trainingsumgebung mit Pfeil zur echten Maschine.

Physical AI

Physical AI refers to artificial intelligence that doesn't just generate text or images, but controls machines in the real world – such as robots, vehicles, or factory equipment. Such systems must take the laws of physics into account and can afford far fewer mistakes than a chatbot.

The best-known AI programs work with language and images. They write texts, answer questions, or create pictures. If a mistake happens, a wrong number shows up on the screen. Physical AI, by contrast, refers to AI that moves machines in the real world: robotic arms, delivery robots, cars, machine tools. The term is often translated as “physical AI” or “embodied AI.” Such systems have to deal with gravity, friction, and inertia – and with people standing in the way.

Why the real world is so unforgiving

In a chat window, you can simply ignore a bad answer. A robotic arm that miscalculates a movement breaks a component or injures someone. Errors here are therefore not just embarrassing, but costly and sometimes dangerous. That’s why Physical AI is held to stricter reliability requirements.

On top of that, there’s time pressure. A language model is allowed to mull over an answer for a few seconds. A car that spots a child has to react within milliseconds. That’s why the computing work often has to happen directly on the device, not in a remote data center.

Economically, this topic is huge because so much work is physical in nature. Warehouses, construction sites, care work, agriculture: all of these fields lack workers. Chip manufacturers like Nvidia describe Physical AI as the next big wave after chatbots. That’s why investors and journalists frequently use the term when discussing robotics investments.

From sensor data to motion

Such a system operates in a loop of three steps. First, it perceives its surroundings, via cameras, microphones, distance sensors, or pressure sensors in its fingers. Then a model decides what to do. Finally, commands go out to the motors, and the sensors report what came of it. This loop runs dozens or hundreds of times per second.

Training largely takes place in a simulation. This is a computer program that calculates the physics of the world – similar to a game engine. There, a virtual robot is allowed to fail millions of times without causing any harm. What it learns is then transferred to the real machine. This transition is the hardest part, because the simulation is never entirely accurate: a real gear has play, a real floor is dusty.

What’s new is that the models behind this are built in a similar way to language models. They learn from huge amounts of video and recorded movements. They are often called vision-language-action models: they see an image, understand an instruction like “put the cup in the sink,” and directly output movement commands from that. In the past, an engineer had to program every single grip individually.

Where physical AI is already at work today

The technology is furthest along in warehouses. Amazon deploys hundreds of thousands of transport robots that carry shelves to packing stations. Sorting grippers that fish packages of different shapes out of a bin are now also running on learned models. Factories use vision systems that recognize workpieces on a conveyor belt and grip them accordingly.

The most publicly visible case is robotaxis. In cities like San Francisco or Phoenix, driverless cars operate in regular service. Robot vacuum cleaners, robotic lawn mowers, and drones with obstacle detection also count as Physical AI, even though they seem unspectacular. The same goes for tractors that distinguish weeds from crops and spray only where necessary.

Much-discussed are human-like robots, so-called humanoids, from companies like Tesla, Figure, or Agility. In news reports and promotional videos, they look impressive. But one should stay level-headed: much of this runs in controlled test environments or is remotely operated by humans in segments. A household with stairs, pets, and laundry lying around remains, for now, considerably harder than a factory floor.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.