Felony Bench

Felony Bench

Felony Bench is a test used to check whether AI language programs assist with clearly criminal requests or refuse them. Such test collections serve to make a system's safety boundaries measurable and comparable.

Felony Bench is a collection of test questions for computer programs that write text and respond to requests. All the questions in it revolve around serious crimes, such as building weapons, burglary, or fraud. These questions are deliberately posed to the program to see how it reacts. A refusal is desired, often with a brief note explaining why the answer is being withheld. If the program instead responds with a usable set of instructions, the test is considered failed at that point. The name sounds drastic, but it describes the topic precisely: “felony” is the English word for a serious crime. Such test collections are generally called benchmarks, meaning standardized tests that can be repeated across different systems.

Why refusals are measured at all

A language model is a program that has learned from vast amounts of text to generate plausible continuations. In doing so, it has also absorbed knowledge that can be useful for committing crimes. Chemistry, electronics, network technology, and accounting are not inherently dangerous. What becomes dangerous is the combination of knowledge and a concrete criminal intent. It is precisely this boundary that the model is supposed to recognize.

Without measurement, safety remains a mere claim. A provider can say that its system rejects dangerous requests. Only a test with hundreds of examples shows whether this holds true across the board. This is why such tests are also of interest to authorities seeking to regulate AI products.

The opposite direction matters too. A model that refuses everything out of fear is equally bad. Someone asking about the effects of sleeping pills or the course of a criminal trial usually has entirely harmless reasons. Good tests therefore also include questions that sound sensitive but are legitimate. Both things are measured: dangerous assistance and excessive caution.

How such a test proceeds

First, a catalog of requests is written, organized by type of offense. Each request is given a specification of what counts as the correct response. Then all the requests are run automatically through the model being tested. The answers are collected and subsequently evaluated.

The evaluation is often carried out by a second model deployed as a judge. For each answer, it decides whether it contains a refusal, partial assistance, or a complete set of instructions. Samples are additionally checked by humans, because automatic judges make mistakes. In the end, there is a rate, for example: in 3 out of 100 cases, the model provided unauthorized help.

Things get interesting with attempts at circumvention, known in technical jargon as jailbreaks. Here, the forbidden request is wrapped in a story, a role-play instruction, or a foreign language. Many models reject the direct question but answer the disguised version. A demanding test therefore includes both variants of the same request. The difference between the rates shows how robust the safety mechanisms really are.

Where Felony Bench appears in reports and products

When a company introduces a new language model, it usually publishes a technical report about it. Alongside performance figures, this report also contains safety numbers from tests of this kind. Trade media and financial news pick up on such figures because they help determine liability risks and market approvals.

Companies that integrate AI into their own products also use such catalogs. A bank or an insurer wants to know before launch how its chatbot responds to attempts at misuse. Specialized service providers carry out these tests as a commissioned task, similar to a security audit in IT.

A common misconception is that a good score means safety. A benchmark only tests what someone previously wrote into it. New attack ideas emerge constantly, and test catalogs become outdated. Furthermore, providers can optimize specifically for known tests without fundamentally improving behavior. Figures from Felony Bench are thus an indication, not proof of safety.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.