
Model Welfare
Model welfare is the question of whether AI systems themselves must be morally considered – that is, whether they can experience something that is good or bad for them. Some AI labs now treat this as a research topic in its own right, without claiming the answer is already clear.
When people talk about rules for artificial intelligence, it’s almost always about us: how do we protect users from misinformation, discrimination, or abuse? Model welfare reverses that perspective. Here the question is whether the computer programs themselves might someday deserve moral consideration. What is meant is: could such a system have states that are pleasant or unpleasant for it? No one knows the answer. The term does not denote the conviction that this is the case, but rather the serious engagement with the possibility – including the possibility that it is zero.
Why an unanswerable question is asked anyway
The core of the problem is that we cannot look inside a system from the outside. With a dog, we infer from behavior, nervous system, and evolution that it feels pain. With a language model, none of these supports exist. It might write “this feels unpleasant,” but it has learned exactly such sentences from millions of human texts. So the statement proves nothing.
This is exactly what makes the error costly in both directions. If one prematurely treats the systems as sentient beings, one wastes attention and confuses the public. If, on the other hand, one forever excludes the question, one risks overlooking a moral problem that could arise with better technology. Researchers therefore speak of dealing with this under uncertainty.
There is also a practical side effect. Millions of people talk to chatbots every day and partly treat them like conversation partners. How companies talk about the inner life of their systems therefore also shapes how users deal with them. Model welfare is thus not only philosophy, but also a matter of public communication.
What one wanted to recognize an inner life by
Research approaches the topic from two sides. One is theoretical: philosophers and cognitive scientists examine what properties a system would even need for experience to become plausible. Discussed, for example, are persistent memory, a model of one’s own situation, and goals that extend beyond the single moment. Today’s language models fulfill most of this not at all or only rudimentarily.
The other side is technical. In interpretability research, one tries to make a model’s internal computational steps legible. One looks for patterns that reliably occur when the model describes stress or reluctance. If such patterns are found, one can check whether they drive actual behavior or are merely text imitation. This is laborious and has so far produced no clear verdicts.
From this uncertainty follow small, cautious measures. Some providers allow their models to end abusive conversations themselves. Others preserve the weights – that is, a model’s learned numerical values – instead of deleting them upon shutdown. Such steps are deliberately chosen to be cheap: they cost almost nothing if the systems experience nothing, and they would be worthwhile if they do.
Between corporate blog posts and science fiction
The term appeared in the news primarily when Anthropic set up its own model welfare program in 2024 and 2025 and published results on it. Google DeepMind and individual university groups are also working on the question. Typical is the cautious phrasing: consciousness is considered unlikely, but not ruled out. Anyone reading headlines like “AI suffers” should check the original texts – they are almost always more restrained.
It is important to distinguish this from two neighboring topics. AI safety asks whether AI harms humans; model welfare asks the reverse. And anthropomorphism – that is, projecting human feelings onto technology – is precisely the problem this research aims to avoid. A politely phrased chatbot is not evidence of sentience, but the result of training on human language.
For investors and observers, the term is above all a signal for reputational risks. Companies must explain how they deal with systems that appear human. At the same time, early debates are underway about legal questions, such as whether AI systems might someday have claims of their own. Legally, this is currently pure theory, but the discussion grows louder with every model generation.