Schematisches Gesicht mit eingezeichneten Landmark-Punkten: nummerierte Punkte entlang der Augenbrauen, Augen, Nase, Mundkontur und Kinnlinie; daneben eine kleine Heatmap, die die Wahrscheinlichkeitsverteilung für einen einzelnen Punkt zeigt.

Landmark Detection

Landmark Detection is a technique in which an AI model automatically identifies specific distinctive points on an object in images or video — such as the corners of the eyes, the corners of the mouth, or the tip of the nose on a face — and determines their precise position. These coordinates serve as the foundation for many applications, from face filters on social networks to medical image analysis.

Landmark Detection refers to the automatic identification of specific key points on an object within an image. A model calculates a precise position for each of these points — expressed as coordinates on the image. For a face, these are, for example, the two corners of the eyes, the tip of the nose, and the four corners of the mouth. The result is not a description like “there is a face,” but a precise map: point 1 is located at pixel 142/89, point 2 at 198/91 — and so on for all defined points. How many points are detected depends on the model and the use case: simple systems work with 5 points on the face, specialized ones with 68, 98, or even more.

Why the precise point position accomplishes so much

A coarse detection only says: “There is a face in this image.” Landmark Detection additionally says exactly where every part of that face is located. This is the difference between a zip code and a house number. Only with the precise coordinates can a system meaningfully process the image further.

One example: to virtually try on glasses, the app needs to know exactly where the eyebrows and the bridge of the nose are located — not just whether a face is present. If the head moves, the landmarks must be recalculated in every single video frame so that the glasses sit exactly in place. The same applies to emotion recognition: from the relative position of the corners of the mouth and the eyebrows, it can be inferred whether someone is smiling or frowning.

Landmark Detection is also a preliminary step for many more complex tasks. If a model is to check whether two photos show the same person, it first normalizes both faces based on the landmarks — rotating and scaling them so that the eyes are always in the same place. Only then does it compare the faces. Without this step, the comparison would be considerably less reliable.

How a model finds the points

Such a model is trained with thousands of images on which people have manually marked the key points. In doing so, the model does not learn what a “corner of the mouth” conceptually is — it learns to predict, from patterns in pixel colors and contrasts, where the next point is likely to be located. The result of the training is weights, i.e. internal numerical values of the model, that enable this prediction.

In practice, many systems work in two stages. First, a simple model roughly locates the face in the image. Then a second, specialized model takes over just the cropped facial region and determines the precise point positions. This division saves computing time, because the precise model only has to process a small image section instead of the entire image.

One well-known approach uses so-called heatmaps: instead of directly outputting a coordinate for each point, the model produces a probability map — a small image showing how likely the sought-after point is to be located at each position. The brightest spot on this map is then the predicted position. This method is more robust than direct coordinate prediction, because the model can better express uncertainty.

Landmark Detection in products and headlines

Landmark Detection is most visible in the camera filters of Instagram, Snapchat, or TikTok. When an app places dog ears on a face in real time or enlarges the mouth, Landmark Detection is at work behind the scenes, recalculating the facial points in every frame — that is, every single image of the video. The same applies to the automatic portrait focus of many smartphones: the camera software recognizes eyes as landmarks and focuses on them.

In medicine, the technique is used for analyzing X-ray and MRI scans. Radiologists mark specific anatomical points on such scans to measure angles or track growth changes. AI systems are increasingly taking over this task automatically and more reproducibly than by hand.

Landmark Detection also plays a role in whole-body analysis — for instance in sports science or physical therapy. Systems like Google's MediaPipe detect up to 33 body points in real time, from the shoulders over the hips to the tips of the toes. This makes it possible, for example, to analyze an athlete’s running technique for postural errors — without special suits or sensors, using just a regular camera.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.