Emotional AI

Emotional AI

Emotional AI refers to computer programs designed to infer feelings from faces, voices, or text. The technology is embedded in call center software, cars, and advertising analytics — though its scientific foundation is highly disputed.

Emotional AI is an umbrella term for computer programs designed to recognize, classify, or imitate human emotions. Such programs analyze what a person shows: facial expression on a camera image, the sound of a voice, the choice of words in a message. From this, they compute an assessment, such as “appears annoyed” or “appears content.” The distinction between showing and feeling is important: the software only measures outward signals, not inner experience. The older technical term for this is Affective Computing, meaning “computing with emotions.” It originates from the 1990s at the Massachusetts Institute of Technology, a well-known US university.

Why companies rely on emotion measurement

Emotions are valuable information for companies. A call center wants to know which customers are about to erupt in anger. An advertiser wants to know whether a commercial is boring or captivating. Surveying people is expensive and imprecise, because respondents often embellish their answers. Automatic measurement promises answers in real time and for thousands of people simultaneously.

At the same time, Emotional AI is one of the most controversial AI applications of all. Psychologist Lisa Feldman Barrett and colleagues evaluated over 1,000 studies in 2019. Their result: a facial expression does not reliably reveal an emotion. A furrowed brow can mean anger, concentration, or simply bright light. Cultural differences add to this, as do differences between individuals.

This is why politics is also stepping in. The European Union’s AI Act, in force since 2024, largely bans emotion recognition in the workplace and in educational institutions. Lawmakers fear surveillance and false assessments. In other areas, use remains permitted, but is classified as a high-risk application subject to strict requirements.

From pixels and audio tracks to emotion labels

Technically, Emotional AI is mostly a classification problem. The system receives inputs and assigns them to one of a few categories, often joy, anger, sadness, fear, disgust, surprise, neutral. It is trained with sample material that people have previously labeled by hand. Shown hundreds of thousands of images with such labels, the program learns statistical correlations between image features and label.

For faces, many systems use the Facial Action Coding System. This is a catalog of small muscle movements, such as “outer brow raiser.” For voices, the software analyzes pitch, volume, speaking rate, and pauses. For texts, it looks at words and sentence structure, which is called sentiment analysis. Some systems combine multiple channels because a single source is too unreliable.

A common misconception: the output sounds objective, but it is not. When the program reports “87 percent anger,” this only means that the input resembles the training examples that people labeled as anger. If these examples come predominantly from staged recordings from a single country, the system fails with genuine emotions elsewhere. Studies also found that some systems classified the faces of darker-skinned people as aggressive more often.

Where the technology stands today

Emotional AI is most widespread in customer service. Software analyzes ongoing phone calls and displays cues to the agent when a conversation threatens to turn sour. Providers such as Cogito or Uniphore sell exactly this. Market researchers also use cameras that observe test subjects while they watch advertising.

Cars are a second major market. In-cabin camera systems are meant to detect whether a driver is becoming tired or distracted. In the EU, drowsiness warning systems are now mandatory for new vehicles. However, these systems mainly measure blinking and gaze direction, not emotions in the strict sense.

In the news, the term usually appears in two contexts. First, in debates about surveillance, such as when authorities test emotion recognition at borders. Second, with voice assistants whose voices now imitate warmth or empathy. This imitation is the second half of the term: a system does not need to recognize emotions in order to simulate them.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.