Safety Card

Safety Card

A safety card is an accompanying document in which a provider describes what its AI system can do, where it fails, and which risks it has tested. It functions like a package insert for software and now comes with nearly every major model release.

When a company releases a new computer program that writes texts or generates images on its own, no one on the outside can tell how reliably it works. This is exactly what the safety card is for. It is a public document that the manufacturer publishes alongside the program. It states what the program is intended for, which tests it has undergone, and in which situations it is known to make mistakes. The comparison with a package insert for medication fits quite well: effect, area of application, and side effects are all listed in one place. The length ranges from two pages to over a hundred pages for major releases.

What users and regulators gain from it

An AI model is a black box from the outside. You see the answer, but not how it came about. Anyone who has to decide whether to deploy such a system in a school, a bank, or a doctor’s office needs more than marketing promises. The safety card provides this foundation in written form.

The section on limitations is particularly important here. Manufacturers write, for example, that a model is not suitable as an advisor for medical questions or that it performs significantly worse with rare languages. Such notes are inconvenient, but they prevent misuse. Without them, a company might use a tool for tasks it was never intended for.

Add to that the legal pressure. The European Union’s AI Act requires technical documentation for high-risk applications. A safety card alone does not fulfill this obligation, but it is often the publicly visible part of it. One point, however, remains open: so far, the manufacturer itself writes what it reports on. There is no independent testing body for this, comparable to a TÜV inspection.

How such a document comes about

Before a model is released, the provider runs it through a series of standardized tests. The system is given thousands of test questions and the answers are evaluated. What is measured includes, for example, how often it fabricates factually incorrect statements or whether it treats people differently depending on their origin. The results end up as percentages and tables in the card.

A second component is called red teaming. In this process, experts deliberately try to induce malicious behavior in the model. They ask for instructions for weapons, ways to defraud someone, or they trick the built-in safeguards using role-play scenarios. Whatever succeeds is logged and usually fixed afterward. The safety card then describes which types of attacks were tested and what remained possible afterward.

The safety card should not be confused with the related model card. The model card mainly describes technical details and training data, i.e., architecture and origin. The safety card focuses on risks and misuse. In practice, the two formats blend together, and some companies call their document a system card instead.

Where these documents show up

You will most often find safety cards on the developer pages of major providers. OpenAI, Anthropic, and Google publish them at the same time as every new model. On the Hugging Face platform too, where freely available models are hosted for download, such a document is considered good practice. Anyone who uploads a model there without a card is quickly asked about it.

In business news, safety cards usually come up when their content is surprising. If it states that a model was able to reveal dangerous expert knowledge in a test, that becomes a headline. Criticism of their scope is also a recurring theme. Some cards are considered by experts to be too sparse or to be marketing dressed up as an evaluation.

For you personally, this has a practical benefit. If you want to know whether an AI tool is suitable for a task, it is worth taking a look at the section on limitations. It often states in just a few sentences what the advertising leaves out.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.