BrowseComp Benchmark

BrowseComp Benchmark

The BrowseComp benchmark is a difficulty test for AI systems that measures how well they can independently uncover complex facts through repeated web searches. Unlike simple knowledge questions, its tasks require many consecutive search steps – and can hardly be solved by guessing.

A benchmark is a standardized test used to measure how well a system solves a particular task. The BrowseComp benchmark was developed by OpenAI in 2025 and tests whether an AI can independently research difficult facts on the internet. The tasks are designed so that a single search is not enough. The model has to issue multiple search queries one after another, evaluate intermediate results, and assemble a precise answer from them. Right or wrong can be clearly checked for every task – cheating with vague answers doesn’t work. The name is made up of “Browse” (as in browsing the web) and “Comp” (short for “Competition” or “Comprehension”).

Why BrowseComp needs its own yardstick

Many older benchmarks have become too easy by now. Current AI models solve them almost flawlessly, making it hard to distinguish between them anymore. This is called saturation: the test no longer separates the good from the best.

BrowseComp therefore deliberately relies on tasks where even humans with free internet access often fail. In OpenAI’s own tests, human testers correctly solved only about 50% of the questions. Simple AI models without search capability scored below 2%. This shows that the test measures an ability that has genuinely not yet been mastered.

Behind this lies a fundamental problem. An AI that relies solely on stored knowledge fails on questions about current events or very specific facts. BrowseComp forces the model to actively search – and evaluates whether it does so systematically and reliably.

Structure and process of a BrowseComp test

Every task in the benchmark has exactly one correct answer – for example, a name, a year, or a place. The questions are deliberately phrased so that they cannot be answered from memory. A typical example is: “Which person won both Prize X and Prize Y and also lived in City Z?” Anyone who knows the answer has probably never seen it before.

The AI system then has to decide on its own: What do I search for first? What do I do with contradictory results? When am I confident enough to answer? This process – searching, evaluating, searching again – is referred to as “agentic browsing,” meaning self-directed browsing with a clear goal.

The evaluation is deliberately binary: an answer counts as either correct or incorrect. This prevents models from scoring points by delivering noncommittal paraphrases. Anyone who answers only half correctly gets zero points – just like in a math test, where the correct method without the correct result doesn’t count.

BrowseComp in practice and in the headlines

Since its release in spring 2025, BrowseComp has regularly appeared in comparison tests between major AI systems. OpenAI’s own model “o3 with browsing” scored around 51% in its first test – a figure no AI had previously achieved, while also showing how much room for improvement remains.

In tech media, the benchmark is cited above all when companies unveil new “research agents” – AI systems designed to independently research the web. BrowseComp provides a common metric with which claims like “our model researches better than the competition” can be verified.

For ordinary users, the benchmark is invisible, but it has direct effects. Anyone who today asks an AI assistant to gather current information on a topic benefits from improvements driven by tests like BrowseComp. The higher a model scores on this benchmark, the more reliably it handles research tasks in everyday use – without the user having to monitor every search step themselves.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.