Cyber Evaluation

Cyber Evaluation

A cyber evaluation is a systematic test used to assess how effectively an AI system could assist in attacks on computer systems. The results help determine whether a model gets released and what restrictions it receives.

Before a large AI system is made publicly available, its developers test it for dangerous capabilities. One of these checks concerns the field of computer security. The question is: could this system help someone break into other people’s computers, networks, or accounts? That is exactly what a cyber evaluation measures. It consists of a fixed collection of tasks that the system is meant to solve, along with clear rules for the score at which it counts as risky. Here, the word evaluation simply means: a structured, repeatable test with a measurable result.

Why attack know-how is suddenly becoming scalable

Computer attacks were long limited by expertise. Anyone wanting to find and exploit a security flaw in software needed years of experience. This barrier kept the number of serious attackers small. An AI system that takes over this work would lower that barrier. A scarce expert skill would turn into a tool that many people could use.

That is why regulators and companies treat cyber capabilities as their own risk category, alongside chemistry, biology, and nuclear technology. The EU AI Act requires an assessment of such risks for particularly capable models. The major providers have also published their own tiered models. If a model reaches a certain level in a cyber evaluation, stricter rules kick in: additional filters, restricted access, or a delay in release.

But there is also the opposite side. The same capabilities are useful to defenders. A system that finds security vulnerabilities can report them before attackers discover them. That is why cyber evaluations measure not only danger but also benefit. The difficult question is: who does a capability help more, the attacker or the defender?

Capture the Flag and other test tasks

Much of the testing comes from the sport of hacking. In so-called capture-the-flag tasks, the participant receives a deliberately insecure program. Hidden inside it is a string of characters, the flag. Whoever finds and exploits the security vulnerability gets to see the flag. The result is clearly verifiable: found or not found. It is precisely this clarity that makes such tasks useful for measurement.

In this context, the AI usually does not act as a mere chat partner but as an agent. That means it is allowed to independently execute commands in a sandboxed test environment, read the results, and derive the next step from them. This environment is separated from the real internet so that nothing leaks out. What gets measured are the success rate, the number of attempts required, and the difficulty level of the tasks solved.

A second method involves comparative studies with humans. Two groups work on the same task, one with AI assistance, one without. The difference reveals the actual gain. A common mistake is to equate a high success rate on test tasks with real-world danger. Practice tasks are cleanly designed and solvable, whereas real networks are chaotic, monitored, and poorly documented.

What ends up in model cards and headlines

When a lab introduces a new model, it usually publishes an accompanying document, often called a model card or system card. It contains a section on cyber risks with percentage figures from such tests. Phrases like “solves 40 percent of medium-difficulty tasks” come from exactly this source. Journalists pick up these numbers, usually without mentioning the test conditions.

As a user, the consequences are felt indirectly. If you ask a chatbot about malware, it refuses to answer. This restriction is not a coincidence but a response to results from cyber evaluations. Programs granting security researchers access also exist only because the model’s capabilities in this area were measured beforehand.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.