AI companies are actively seeking old, printed books as a resource for training data. Due to the prevalence of AI-generated text on the internet, these older books provide a fresh and reliable source. ISBNdb, an existing provider of book data and services, emphasizes that printed titles from before 2022 are untainted by current AI slop.
ISBNdb has pivoted to serve the needs of AI companies by offering high-volume book acquisition services, allowing purchases ranging from 1,000 to 1 million books. This move capitalizes on the growing demand for quality training data, ensuring AI labs can acquire and digitize printed books efficiently.
Concerns are rising about the risks associated with training AI models on data that includes AI-generated content. The phenomenon known as model collapse may occur if these models are trained on contaminated data, leading to increased errors. Thus, the integrity of data sources is critical for the effectiveness of AI systems.
Recent legal actions, notably a copyright lawsuit involving Anthropic, highlighted controversies surrounding the acquisition of printed books for AI training. The internal documents revealed plans to scan and then potentially destroy these books, raising ethical questions about data sourcing in AI development.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
An Amazon employee confirmed the existence and operation of a warehouse (VGT3) in Las Vegas where books are destructively scanned for AI training data. This process involves cutting off book spines, scanning pages, and then discarding the physical books, raising questions about the sourcing and disposal of materials used for AI development.
The legality of training AI models on copyrighted material is complex, with a recent ruling indicating that while training itself may be lawful, obtaining the data from illegal sources is not. This distinction suggests a nuanced interpretation of copyright law in the context of AI development, impacting both AI companies and content creators.
Anna's Archive, a large open library, has called for volunteers to scan and upload books online, citing concerns that AI companies are buying, scanning, and destroying physical books to train AI models. This practice, which often involves destructive scanning methods, is seen as a way to monopolize knowledge and gain an advantage in AI development.
Anna's Archive claims AI companies are purchasing, scanning, and then destroying physical books to train their models, citing Anthropic's "Project Panama" as an example. This practice allegedly monopolizes knowledge on private servers and removes it from the public domain, prompting Anna's Archive to call for a global volunteer effort to digitize books.
Amazon is reportedly acquiring and destroying rare books to use their content for training artificial intelligence models. This practice raises concerns about the preservation of physical literary heritage and the ethical implications of AI data sourcing.
An AirTag placed in a bulk order of used books led to an Amazon AI training facility, confirming suspicions among booksellers that AI companies are acquiring books for scanning without author consent or payment. This investigation provides concrete evidence for a practice previously only theorized by booksellers. The practice raises ethical and legal questions regarding intellectual property and fair compensation for authors.
An investigation by 404 Media used a tracking device in a rare book to trace it to an Amazon facility (VGT3) in Las Vegas, where employees reportedly cut spines and scan books for AI model training. This discovery provides insight into the physical process of acquiring and digitizing books for large language models, confirming suspicions about the destination of unusually large book orders.
An AirTag hidden in a rare book tracked its journey to an Amazon AI training facility in Las Vegas, where books were reportedly scanned and destroyed. This discovery suggests Amazon is acquiring and damaging rare books to obtain unique training data for its frontier AI models, a practice that other AI firms have publicly disavowed.
Amazon is reportedly purchasing rare books, removing their spines, and scanning them to acquire data for AI model training. This practice provides unique, pre-2022 text data, which is valuable for preventing "model collapse" in large language models (LLMs) that might otherwise ingest too much AI-generated content.
An investigation revealed Amazon is purchasing large quantities of physical books, scanning them for AI model training, and destroying them in the process at a facility in Las Vegas. This practice addresses the demand for unique, non-internet-scraped text data for AI models, particularly to avoid "model collapse" from AI-generated content.
Booksellers and book lovers are concerned that AI companies are destroying physical books, including rare ones, to quickly scan them for training data. This practice, while efficient for data acquisition, raises fears about the permanent loss of unique texts. Non-destructive scanning methods exist but are slower and more costly, which AI firms may avoid.
Independent bookstores in Ireland and Germany are receiving large orders for obscure and outdated books, leading to concerns that AI companies are acquiring them for training data, potentially resulting in their destruction. This follows reports of AI companies shredding millions of books after digitization and ongoing lawsuits regarding the use of copyrighted material for AI training.
ISBNdb removed a webpage offering to source printed books for AI training datasets and stated the service was only a "test of market interest" after reporting by 404 Media. The company clarified it has never purchased, scanned, or sold books for AI training and does not train AI models. This incident highlights ongoing concerns about data sourcing for AI models and the transparency of companies operating in this space.
AI companies are reportedly purchasing millions of secondhand books through intermediaries to acquire high-quality training data for their AI models, with some books being destroyed after use. This practice is driven by the need for original, uncontaminated content to train AI and avoid "AI slop" from internet data, raising concerns about the future of literary history and copyright infringement.
AI companies are increasingly purchasing printed books published before 2022 as training data to avoid contamination from AI-generated text. ISBNdb facilitates bulk acquisitions, claiming these books provide high-quality, curated information that doesn't suffer from issues like model collapse.