Artificial Analysis

Artificial Analysis

Artificial Analysis is an independent platform that compares AI models from various providers and publishes the results in public rankings. It primarily measures quality, response speed, and price per request.

Artificial Analysis is a company that compares computer programs with artificial intelligence and publishes the results. This mainly refers to programs that respond to text inputs, i.e. chatbots like ChatGPT, Gemini, or Claude. Such programs are now offered by dozens of providers, and they differ greatly in quality and price. Artificial Analysis has them all solve the same tasks and measures the results. The outcome is a set of rankings showing how well each program performs, how fast it responds, and what it costs to use. The company is based in Australia and does not sell AI itself, so it is not dependent on any of the providers it tests.

Why a neutral ranking is needed

Providers of AI programs test their own products and publish the results themselves. In doing so, they naturally pick the tasks in which they look good. Anyone reading a press release rarely learns in which areas a model is weak. A third party that tests all providers using the same method creates a basis for comparison here.

For companies, this is about a lot of money. Anyone building software that sends millions of requests to an AI model every day pays for each individual response. A model that is only slightly worse but costs a tenth as much can be the better choice. It is precisely this trade-off between quality and cost that Artificial Analysis makes visible by putting both figures in the same chart.

That is why these numbers regularly appear in business news. When a Chinese lab releases a model that comes close to the most expensive US model in the ranking, that is news with consequences for stock prices. Here, rankings are less a sporting competition than market information.

What exactly is measured

Quality is determined using standardized collections of tasks, known as benchmarks. These are collections of questions with known correct answers, such as math problems, programming tasks, or university-level knowledge questions. Every model is given the same questions, and the proportion of correct answers yields a score. Artificial Analysis combines several such collections into one overall figure, the Intelligence Index.

In addition, it measures what a model feels like in real-world operation. Two figures are central here: the time until the first word of the response, and the number of words per second after that. Neither is checked just once but around the clock, because servers run at different speeds depending on load. The third figure is the price the provider charges per amount of text processed.

A common misconception is to take these scores as an objective measure of intelligence. They only measure how well a model solves certain test tasks. If tasks from a well-known benchmark appear in a model’s training material, it has essentially memorized the solutions. Experts call this contamination, and it is the reason why benchmarks must be regularly renewed.

Where these numbers show up

Artificial Analysis is most often encountered in reports about new model releases. Sentences like “is on par with GPT-5 in the Artificial Analysis Index” come from these rankings. Providers themselves also like to cite these figures in their announcements, though usually only when they turn out favorably.

The rankings are publicly available online and free of charge. Anyone building a project with AI can look there to find out which model fits their purpose. The company earns money through more detailed data packages for businesses and investors. Comparable offerings include LMArena, where people rate two anonymous responses against each other, and the rankings on the Hugging Face platform.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.