Emotion Recognition

Emotion Recognition

Emotion recognition refers to computer programs that attempt to infer feelings from faces, voices, or texts. The technology is used in advertising research, call centers, and cars, is scientifically controversial, and is partially banned in the EU.

Emotion recognition is the attempt to use a computer program to determine how a person is currently feeling. To do this, the program receives recordings of a person: a photo, a video, an audio track, or written text. From this, it calculates an assessment, such as “angry”, “content”, or “bored”. Usually, it also provides a number that expresses how confident the program is in this assessment. There is an important distinction from facial recognition: facial recognition wants to know WHO is in an image, while emotion recognition wants to know HOW that person is feeling. The two methods are often confused, even though they answer different questions.

A billion-dollar market on shaky ground

For companies, the idea is extremely attractive. Whoever knows how a customer feels can treat them more effectively. That’s why advertising firms have test subjects filmed while watching commercials and then analyze the footage. Insurance companies and HR departments have experimented with software that evaluates applicants in video interviews. So this is not a niche topic, but rather one involving decisions that matter a great deal to individual people.

At the same time, the scientific foundation is weak. A major review study by a research group led by psychologist Lisa Feldman Barrett came to a clear conclusion in 2019: an inner feeling cannot be reliably read from a facial expression. A person may smile because they are happy, because they are being polite, or because they are nervous. The assumption that every emotion has the same expression worldwide is now considered outdated.

That is why lawmakers have taken action. The European Union’s AI Act largely bans emotion recognition in the workplace and in educational institutions. This means a boss is not allowed to have employees' moods checked via camera. In other areas, its use remains permitted, but it is considered particularly high-risk and must be strictly documented.

From muscle movements to emotion labels

Behind this technically lies a method called machine learning. A program is shown a very large number of examples with known answers so that it can find the pattern itself. For emotion recognition, tens of thousands of facial photos are collected that people have previously labeled as “sad” or “surprised”. The program learns which image features coincide with which label. In the end, it assigns new images to the same categories.

For faces, many systems work with measurement points on eyebrows, eyelids, and the corners of the mouth. For voices, pitch, volume, speech tempo, and pauses matter. For text, the software pays attention to word choice, for example words like “finally” or “disaster”. Some providers combine several sources and additionally measure pulse or skin conductivity.

The crucial point: the program never measures the feeling itself, only external signals. It guesses which label people would have assigned to similar signals. If the training images are mostly of actors exaggeratedly portraying emotions, the software fails with real, more restrained faces. And because certain ethnic groups are barely represented in many datasets, the systems measurably perform worse for these groups.

In the car, in the call center, in the exam

The technology is most commonly encountered in cars. New vehicles have an interior camera that monitors eyelid closure and gaze direction and warns in case of drowsiness. This is a stripped-down form of emotion recognition and is intentionally used in the EU as a safety feature. It is only concerned with the state of “inattentiveness”, not the full range of emotions.

In call centers, software runs in the background that assesses conversations in real time. If a caller’s tone is rated as agitated, a notification pops up on the screen or a supervisor joins the call. The same applies to exam proctoring software that looks for anomalies during online exams. Such programs regularly come under criticism because they classify normal behavior as suspicious.

In news reports, the term usually appears in connection with bans, fines, or studies. A related term is “sentiment analysis”: this refers to the analysis of texts, such as product reviews or posts on social media. It is considered less sensitive because it only evaluates the tone of a statement and not the inner life of an observed person.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.