
arXiv
arXiv is a free online platform where researchers publish their papers before a journal has reviewed them. In AI research, most important papers appear there first.
arXiv is a public website where scientists upload their papers. Anyone can read and download these texts for free. The name is pronounced “archive,” since the capital X stands for the Greek letter Chi. The platform has existed since 1991, launched by a physicist in the USA. Today it holds over two million papers from physics, mathematics, computer science, and other fields. arXiv is operated by Cornell University, funded through donations and membership fees from universities.
Why AI news appears there first
Normally, research works like this: you write a paper and send it to a journal. There, other researchers read the text and look for errors. This review process is called peer review, meaning evaluation by fellow experts. It often takes six months to two years.
In AI research, that is far too slow. A method that is new today can already be outdated in a year. That’s why teams at Google, Meta, or OpenAI upload their results directly to arXiv. Such not-yet-reviewed texts are called preprints, meaning advance copies. Famous papers like “Attention Is All You Need,” the foundation of today’s language models, were published there.
For journalists and investors, arXiv is thus a kind of early-warning system. Anyone who wants to understand what a company is currently working on looks there. Conversely: a preprint is not established truth. Review by experts is, after all, still missing.
From upload to number
A researcher uploads their paper as a file, usually a PDF. They choose a category, for example “cs.LG” for machine learning within computer science. Afterward, volunteer moderators take a brief look at it. However, they do not check whether the results are correct. They merely filter out obvious nonsense and misclassified texts.
After usually one day, the text is online and receives a fixed number, such as 1706.03762. The first half indicates the year and month, the second is a running count. Using this number, anyone can permanently cite and find the paper again.
If the author changes something later, no new number is created. Instead, a version label is added, i.e. v1, v2, v3. All old versions remain visible. This makes it possible to trace what a team originally claimed. A paper may later also appear in a journal. The arXiv version remains in place nonetheless.
arXiv in the news and in academic study
Tech articles often contain a sentence like: “The team describes the method in a paper on arXiv.” Behind that is usually a link to the number. You can follow the link and look at the original text yourself. The opening paragraphs, called the abstract, summarize the result in a few sentences. This part is often reasonably understandable even without specialist knowledge.
Technical reports on well-known products can also be found there. For many language models, there is a so-called model card or technical report on arXiv. Anyone who wants to know what data a model was trained on will find the official details there.
A common misconception: because a text is on arXiv, it must be expert-reviewed. That is false. arXiv also hosts papers that no journal ever accepted. Related but different are platforms like bioRxiv for biology or SSRN for economics. They follow the same principle in other fields.