Cargo Cult Science

Cargo Cult Science

Cargo Cult Science refers to research that outwardly resembles genuine science but fails to fulfill its core principles — above all, honest scrutiny of one's own results for errors. In the AI industry, the term is used when experiments are indeed logged and published, but the methods are chosen such that a positive result is nearly unavoidable.

After World War II, anthropologists observed a strange phenomenon on islands in the Pacific. During the war, the inhabitants there had watched planes land and bring goods — canned food, clothing, tools. When the soldiers withdrew, the islands were left empty. In response, some communities built mock runways, wore headphones carved from wood, and waved at imaginary aircraft: they imitated the outward form without understanding what had actually brought the planes in the first place. The physicist Richard Feynman picked up this image in a famous 1974 speech. He called research that copies the rituals of science without fulfilling its core Cargo Cult Science. Since then, the term has become a common accusation — including in AI research.

The core of the problem: self-deception instead of self-criticism

According to Feynman, real science demands above all one thing: that the researcher does everything possible to prove themselves wrong. Anyone who sets up an experiment so that it must confirm their favorite result is not generating knowledge — they are merely staging it. That sounds like a rare edge case, but it isn’t.

In practice, this often happens unconsciously. One chooses the comparison method against which one’s own model looks best. One tests different settings until an impressive number comes out — and publishes only that number. One optimizes for a benchmark that is actually supposed to measure whether a model is good in general. What remains in the end is a number that looks good but says little about whether the system actually works in reality.

Why AI research is especially prone to this

No other field of research in recent years has published as many results as quickly as AI. On the platform arXiv, hundreds of new studies appear every day. The pressure to stand out is enormous — for researchers at universities just as much as for teams at companies trying to impress investors. This creates an incentive to present results in a way that makes them seem as impressive as possible.

One concrete structural problem is called p-hacking, or also HARKing (Hypothesizing After Results are Known): one tries out many variants and formulates the hypothesis only afterward, as if that had been exactly what one was looking for from the start. Statistically speaking, with enough attempts one will almost always find some positive result — whether it’s real or coincidental then becomes almost impossible to tell apart. Studies estimate that a substantial share of published AI results cannot be reproduced, meaning they do not hold up under independent replication.

Added to this is the problem of benchmark saturation. When a dataset serves as a yardstick for long enough, models become increasingly tailored precisely to it. The score rises — but the actual ability to solve new problems lags behind. The model has learned the ritual, not the substance.

Where the term appears in the AI debate

Cargo Cult Science comes up mainly in critical commentary on AI benchmarks. When a company announces that its model has, for the first time, outperformed a human on a particular test, the counter-question often follows: Was the model trained specifically on this exact test? Is the test even still meaningful if everyone is optimizing for it? This skepticism is not an attack on AI as such — it is a demand for genuine science.

The term also appears in debates about reproducible research. Several research groups have, in recent years, systematically attempted to replicate published AI results — with sobering outcomes. Some models that showed outstanding performance in papers performed barely better than older, simpler approaches under controlled conditions.

However, the term is also used as a warning to companies that deploy AI without understanding why a model works. Anyone who puts a system into production merely because it looked good in testing, without knowing its weaknesses, is, figuratively speaking, building a wooden runway — and waiting for results that will never come.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.