Preprint

Preprint

A preprint is a scholarly paper that its authors post publicly online before a journal has reviewed and accepted it. In AI research this is the norm: almost all major results appear first as a preprint, usually on the platform arXiv.

When researchers have a result, they write it up as a paper. In the past, they would first send this paper to a journal. There, other experts read it, criticize it, and demand changes. Only afterward is it published, often a year later. A preprint skips this route: the authors put the text online themselves immediately, freely readable by anyone. At that point, no one but the authors themselves has reviewed it.

Why AI research now appears almost exclusively in advance

In artificial intelligence, sometimes only weeks pass between two major innovations. A year of waiting for a journal would be fatal for one’s own career at this pace. Whoever waits risks another group publishing the same result first. That’s why a custom has taken hold in this field: you put the work online right away and submit it in parallel to a conference.

For outsiders, this has a big advantage. Almost all current AI research is readable for free, without an expensive journal subscription. Even students or journalists can read the original paper behind a headline. Many famous texts in the field were first only preprints, for instance the paper that introduced today’s architecture for language models in 2017.

The price for this is a lack of quality control. A preprint can be brilliant or simply wrong. There are papers with computational errors, with embellished measurements, or with claims that no one could verify. The term says nothing about quality, only about the timing of publication.

From text to entry on arXiv

The most important place for preprints is called arXiv, pronounced like the English word archive. The platform is run by Cornell University and is over thirty years old. Authors upload their finished PDF file there and assign it to a subject area. A small team takes a brief look to check whether it is even a scholarly paper at all. This screening is not a review of content; it takes hours rather than months.

Afterward, the text receives a fixed number, for example arXiv:1706.03762. Anyone can cite it permanently using this number. If an author later finds an error, they upload a new version. The old one, however, remains visible, so that changes can be tracked.

A preprint is therefore not the opposite of a peer-reviewed paper, but an earlier stage of the same paper. Many texts later do go through peer review by fellow experts and appear at a conference. The version on arXiv nevertheless remains in place. That’s why it’s worth checking, when reading, whether a paper has since been accepted.

Correctly assessing preprints in headlines

Reports about new AI capabilities almost always rely on preprints. Sentences like “a new study shows” often refer to a text that no one besides the authors has yet checked. Companies also use this format: Google, Meta, and other labs publish their model descriptions as preprints, sometimes on the same day as the product itself.

For you as a reader, this means above all one thing: a link to arXiv is not a seal of quality. Sensible questions are who wrote the paper, whether independent groups were able to replicate the result, and whether the raw data are publicly available. In the case of company publications, there’s the additional factor that they also serve as advertising for their own product.

A common misconception is that preprints are disreputable or inferior. That’s not true. Physics has worked this way since the nineties, and a large share of the most important AI papers were never anything else. It is a fast channel with an open outcome, nothing more.

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.