Metadata

Metadata

Metadata is data about data: it describes when, where, by whom, and in what form something was created, without reproducing the actual content itself. It is precisely these additional details that make large amounts of data searchable – and they often reveal more about people than the content itself.

Metadata is information that says something about other information. In a photo on a mobile phone, the actual picture is the content. When it was taken, with which device, at which location, and how large the file is – that is metadata. It is usually embedded invisibly in the file and recorded automatically. You can picture it like the label on a moving box: the label reveals the contents, the room, and the packing date without anyone having to open the box. Almost every digital file and almost every digital action generates such labels.

What a label reveals about the sender

Without metadata, large amounts of data would be practically unusable. A hard drive with 40,000 photos is only searchable because each image carries a date, location, and format with it. Search engines, music services, and library catalogues work almost exclusively with such descriptions. They sort, filter, and find things without needing to understand the content itself.

Metadata is politically sensitive for a different reason. In the case of a phone call, it does not say what was spoken. But it does say who called whom, when, and for how long. A great deal can be inferred from a call to an addiction helpline at three in the morning, even without a single word of the conversation being known. That is why courts and parliaments have been arguing for years over so-called data retention – the blanket, suspicionless storage of connection data.

A common misconception is that metadata is harmless because it merely contains “framework information”. The opposite is closer to the truth. Content is unstructured and hard to evaluate, whereas metadata is cleanly structured and can be compared automatically millions of times over. That is precisely why it is so valuable to authorities, advertising companies, and data analysts.

Where the additional details come from

Some metadata is generated automatically when a file is created. Cameras write exposure time, lens, time of day, and often GPS coordinates into the image following a standard called EXIF. Word processing programs store author names, editing duration, and version history. This information is retained when the file is sent – often without the sender even thinking about it.

A second portion is assigned deliberately. A library enters title, author, year of publication, and keywords. Editorial offices tag articles with a department, author initials, and publication time. So that different systems can understand one another, there are fixed sets of rules, so-called schemas. They define which fields exist and how they are to be filled in.

Technically, metadata resides either directly within the file or separately in a database. Both approaches have drawbacks. If it sits inside the file, it is easily lost during conversion. If it sits in a database, the link to the file can break. Anyone who wants to get rid of metadata can delete it in the file properties; many social networks automatically strip location data upon upload.

Metadata in AI systems and in everyday life

When training AI models, metadata also has a say in determining quality. It indicates which source a text comes from, in which language it was written, and when it was created. This allows duplicates to be filtered out and licensing questions to be clarified. The debate over copyright in training data is almost always about this kind of provenance information.

Metadata is also being newly discussed as a way of labelling AI-generated content. A standard called C2PA attaches a signed note to images recording which program was used to create or edit them. The EU’s AI regulation likewise requires that generated content be marked in a machine-readable way. However, the method is not tamper-proof: anyone who removes the metadata also removes the notice.

In everyday life, people encounter it constantly, usually without noticing. The Photos app sorts pictures by travel destination, Spotify displays the album and release year, the mail client sorts by sender and date. Metadata is to be distinguished from payload data, i.e. the actual content. The boundary is not always sharp: an email’s subject line technically counts as header data, but in terms of content it often already reveals almost everything.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.