
Shadow Library
A shadow library is a large collection of books and academic papers that are available for free on the internet, even though the rights holders have not authorized this. Such collections have become a major point of contention because AI companies have used them to train their models.
A shadow library is a website that offers books, journal articles, and academic papers for free download. Unlike a regular online library, the authors and publishers have not consented to this. The files are therefore copies distributed without permission. Well-known examples include Library Genesis, Z-Library, Sci-Hub, and Anna’s Archive. Some of these collections comprise several million works, more titles than large university libraries hold. The name comes from the fact that they operate in the shadow of the law: technically openly accessible, legally prohibited.
Why publishers and AI companies are fighting over it
For publishers, shadow libraries are a direct attack on their business. A textbook often costs 60 euros in stores, and a single academic article can cost 30 dollars at some publishers. When the same work can be downloaded for free, a portion of revenue disappears. That is why publishers have been suing the operators for years and getting domains blocked.
Since the boom of AI language models, the dispute has taken on a new dimension. Such models learn from enormous amounts of text, and well-written books are especially valuable training material. Shadow libraries deliver exactly that, in large quantities and in one place. For a company, it is cheaper to download such an archive than to negotiate licenses individually with thousands of publishers.
This has happened repeatedly and has ended up in court. In lawsuits against major technology companies, internal messages came to light in which employees discussed downloading such archives. Authors' associations see this as a double harm: their works were first copied and then used to train systems that compete with them.
How the collections come into being and survive
The holdings mostly grow through user uploads. Students, researchers, and library staff scan books or upload PDF files they have access to through their university. At Sci-Hub, many articles entered the archive via shared university logins. The collections are therefore not the work of a single company, but the result of many individual contributions.
To remain accessible despite bans, the operators work with multiple layers of protection. They frequently change their internet address, run servers in various countries, and remain anonymous. Many holdings are also available as torrent files. Torrent means the data resides on many private computers simultaneously and is distributed from there. A single shutdown therefore does not delete the collection.
It is important to distinguish shadow libraries from legal offerings. Open access means that researchers voluntarily publish their own work for free. Project Gutenberg distributes works whose copyright protection has expired. Both of these are legal. A shadow library differs in that it distributes protected works without consent.
Where the term appears in the news
Shadow libraries are most often mentioned in reports on copyright lawsuits against AI companies. These reports cite dataset names such as Books3, a text collection of around 200,000 books that originated from such a source. It was used in several early language models and later taken offline after legal pressure. Damage claims now run into the billions.
The topic also comes up in everyday life, albeit under different names. Anyone searching for an expensive textbook and landing on a site with thousands of free PDFs is likely in a shadow library. Such downloads are not permitted under German copyright law and can result in cease-and-desist letters.
For investors and observers of the tech industry, the term is interesting for another reason. It marks an open legal question with significant financial risk. Should courts decide that training with such data is impermissible, companies face hefty payments and expensive licensing agreements. Several AI companies have therefore begun signing agreements with publishers and news organizations.