Responsible Scaling Policy

Responsible Scaling Policy

A Responsible Scaling Policy is a self-imposed rule set by an AI company that specifies which safety measures must apply once its systems reach a given capability level. If a system reaches a predefined threshold, it may only be further developed or released once the associated safeguards are in place.

Companies building very large AI systems have been publishing their own safety frameworks for a few years now. A Responsible Scaling Policy is one such framework. It describes how dangerous a system may become at most before additional safeguards kick in. To this end, the company defines tiers: each tier specifies what the system could be capable of and which precautions then become mandatory. The core is a promise: if a tier is reached before the matching precautions are ready, development or release is halted. The word “scaling” here refers to making systems larger and more powerful — precisely the process that is meant to be limited.

Why companies build in their own brakes

The driving force behind this is a race. Whoever has the stronger system first wins users, contracts, and investor money. In such a race, caution is a disadvantage because it costs time. A publicly committed policy is meant to ease this pressure: if all major providers commit to similar limits, no one loses out simply by being careful.

A second reason is the legal situation. Binding rules for very capable AI are only just emerging, for example in the European Union’s AI Act. Technology is developing faster than legislation. Self-commitments fill this gap provisionally and, at the same time, serve regulators as a template for later rules.

However, one should clearly recognize the difference from a law. A self-commitment can be changed, watered down, or reinterpreted by the company itself. There is no judge to enforce it, and usually no penalty for a violation. Critics therefore call it, at best, a first step.

Capability thresholds and red lines

Technically, such a policy consists of three parts. First, there are thresholds describing dangerous capabilities. A typical example: the system effectively helps a layperson build a bioweapon or carry out a major cyberattack. Second, there are tests used to check whether a threshold has been reached. Third, there are the measures that apply once that threshold is crossed.

Testing is done primarily by deliberately attacking one’s own system. Experts try to coax forbidden answers out of the model. This procedure is called red teaming, named after the attacker role in military exercises. In addition, standardized tests, so-called evaluations, are run that measure capabilities in numbers.

The measures cover two sides. One side protects usage: filters reject dangerous requests, and misuse is monitored. The other side protects the model itself from theft, for example through strict access controls in the data center. That’s because a stolen model can be operated without any restrictions at all.

Where these policies show up in the news

Every major AI lab now has such a document, though they go by different names. Anthropic uses the name Responsible Scaling Policy, with tiers called ASL-1 through ASL-4. OpenAI calls its framework the Preparedness Framework, and Google DeepMind refers to the Frontier Safety Framework. The underlying idea is the same in all cases.

You’ll most often encounter the term around product launches. Safety reports accompanying new models reference the respective policy. These state which tier the model has reached and which safeguards are therefore active. In 2025, Anthropic classified its Claude Opus 4 model under the stricter ASL-3 tier for the first time, as a precaution and without conclusive proof of danger.

These documents also play a role in politics. At international AI summits, companies present their frameworks to demonstrate to governments that they are capable of taking action. A common misconception is that this constitutes verified safety. In reality, companies generally assess for themselves whether they are complying with their own rules.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.