Legal Research Bench

Legal Research Bench

A Legal Research Bench is a standardized collection of legal research tasks used to test how well a computer program answers legal questions. The results show whether a program correctly finds and cites laws and rulings – or whether it invents sources.

Whoever sells software likes to claim it’s good. A testing procedure with fixed, always identical tasks makes such claims verifiable. A Legal Research Bench is exactly such a testing procedure, specifically for programs meant to answer legal questions. The tasks come from legal research: Which law applies here? Which court has already ruled on this? How can a contract text be interpreted? Every program gets the same tasks, and the answers are evaluated according to predetermined rules. This allows two programs to be compared fairly, instead of just believing advertising promises.

Why lawyers must remain suspicious

In law, it’s not just whether an answer sounds plausible that counts. It must be provable. A lawyer who cites a ruling must be able to produce the original. This is exactly where language models – programs that continue texts word by word – have a dangerous weakness. They sometimes generate case numbers and ruling names that never existed. Experts call this hallucination.

This is not a theoretical problem. In the USA, lawyers have been reprimanded by courts and fined because they submitted briefs containing invented rulings. A 2024 Stanford University study examined several commercial legal tools and found that a significant portion of their answers still contained false or unsubstantiated claims. The providers had previously advertised the opposite.

A benchmark forces such claims into measurable numbers. It does not answer the question of whether a program is useful. It answers the question of how often it gets things wrong. And that number is the decisive basis for law firms, legal departments, and regulatory authorities when deciding whether they are even allowed to use a tool.

Tasks, model solutions, and the pitfalls of evaluation

Such a test consists of three parts. First, a collection of questions, often several hundred to several thousand. Second, a model solution for each question, created by hand by legal experts. Third, an evaluation procedure that compares the answer with the model solution. Well-known examples from research are LegalBench and CaseHOLD, both of which are geared toward English-language US law.

The task types vary in difficulty. Simple tasks have a clear-cut answer: Which of five headnotes belongs to this ruling? This can be evaluated automatically. Difficult tasks require an entire legal opinion text. Here, either a human must evaluate it or a second AI model must step in as examiner – both are labor-intensive and prone to error themselves.

What is measured is usually more than just “right or wrong.” Typical metrics are: proportion of correct answers, proportion of fabricated citations, and proportion of cases in which the program admits it does not know the answer. The last criterion is more important than it sounds. A tool that stays silent when uncertain is considerably more useful in law than one that confidently guesses.

Between law firm software and regulation

You’ll encounter this term mainly in product announcements. Providers of legal software such as Harvey, Thomson Reuters, or LexisNexis advertise their research assistants with benchmark results. Anyone reading such numbers should check which test was used and who designed it. A self-built test that a company has optimized its own model against says very little.

Legal policy is also interested in this. The EU’s AI Act classifies certain applications in the judicial field as high-risk. For such systems, evidence of accuracy and robustness is legally required. Standardized tests are an obvious way to provide such evidence.

A common misconception is that a good benchmark score means reliability in everyday use. That is not true. Most of these tests are based on US law, barely cover German law, and often contain only closed, cleanly formulated cases. Real-world cases are messier. The benchmark is like a driving test on a practice course – not proof that someone can drive safely at night in the rain.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.