← All stories
● Covered by 6 sources · 15 reportsMedium impact10 negative4 neutral

AI Companies Source Old Books for Training Data, Avoiding AI-Generated Text

🔄 Updated 10d ago — new reporting from 404 Media
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • AI companies seek pre-2022 printed books for training data.
  • Printed books avoid AI-generated text contamination.
  • ISBNdb offers bulk acquisition services for AI firms.
  • Amazon is buying and scanning physical books for AI training.
  • Amazon destroys books after scanning them for AI training.
  • A tracking device revealed Amazon's book operation.
  • The final destination was an Amazon warehouse in Las Vegas, Nevada.
  • Amazon employees cut book bindings for faster scanning.
  • The Amazon team at the warehouse is called VGT3.
  • Amazon's statement to 404 Media did not mention AI training.
  • The VGT3 facility logo is a dinosaur holding a book.
  • 404 Media connected with a bookseller to plant the AirTag.
  • The rare book was part of a 1,000-book order.
  • VGT3 is part of a larger facility named LAS8.
  • The investigation started in July.
  • Booksellers receive massive bulk orders of used books with no apparent connection between titles.
  • The story was featured on a podcast.
  • Anna's Archive claims AI companies destroy books.
  • Anthropic's "Project Panama" is an example of book destruction.
  • Project Panama was exposed in a $1.5 billion copyright settlement.
  • Anthropic launched Project Panama in early 2024.
  • Anthropic spent tens of millions of dollars on books for Project Panama.
  • Anthropic used scanned books to train its Claude LLM.
  • Anna's Archive calls for a global volunteer effort to digitize books.
  • Anthropic's copyright settlement was for $1.5 billion.
  • The Anthropic settlement is the largest amount ever in a copyright case.
  • Anthropic paid $200 per title for its 7 million pirated books.
  • The ruling affirmed that using existing works to train AI models is fair use.
  • Independent bookstores across Europe report large, random book orders.
  • Judge William Alsup ruled Anthropic's AI training was lawful.
  • The ruling distinguished between lawful training and unlawful data acquisition.
  • An Amazon employee confirmed the existence and operation of the VGT3 warehouse.
  • The VGT3 warehouse is in the same facility as LAS8, Amazon's print-on-demand business.

AI Companies Target Old Books

AI companies are actively seeking old, printed books as a resource for training data. Due to the prevalence of AI-generated text on the internet, these older books provide a fresh and reliable source. ISBNdb, an existing provider of book data and services, emphasizes that printed titles from before 2022 are untainted by current AI slop.

ISBNdb's Role in Acquisitions

ISBNdb has pivoted to serve the needs of AI companies by offering high-volume book acquisition services, allowing purchases ranging from 1,000 to 1 million books. This move capitalizes on the growing demand for quality training data, ensuring AI labs can acquire and digitize printed books efficiently.

Risks of AI-Generated Data

Concerns are rising about the risks associated with training AI models on data that includes AI-generated content. The phenomenon known as model collapse may occur if these models are trained on contaminated data, leading to increased errors. Thus, the integrity of data sources is critical for the effectiveness of AI systems.

Copyright Issues and Scanning Controversies

Recent legal actions, notably a copyright lawsuit involving Anthropic, highlighted controversies surrounding the acquisition of printed books for AI training. The internal documents revealed plans to scan and then potentially destroy these books, raising ethical questions about data sourcing in AI development.

Updates

🕒 2026-08-26 · new reporting from 404 Media
  • An Amazon employee confirmed the existence and operation of the VGT3 warehouse.
  • The VGT3 warehouse is in the same facility as LAS8, Amazon's print-on-demand business.
🕒 2026-08-23 · new reporting from TechCrunch
  • Judge William Alsup ruled Anthropic's AI training was lawful.
  • The ruling distinguished between lawful training and unlawful data acquisition.
🕒 2026-08-21 · new reporting from Tom's Hardware
  • Anthropic's copyright settlement was for $1.5 billion.
  • The Anthropic settlement is the largest amount ever in a copyright case.
  • Anthropic paid $200 per title for its 7 million pirated books.
  • The ruling affirmed that using existing works to train AI models is fair use.
  • Independent bookstores across Europe report large, random book orders.
🕒 2026-08-21 · new reporting from Hacker News Front Page
  • Anna's Archive claims AI companies destroy books.
  • Anthropic's "Project Panama" is an example of book destruction.
  • Project Panama was exposed in a $1.5 billion copyright settlement.
  • Anthropic launched Project Panama in early 2024.
  • Anthropic spent tens of millions of dollars on books for Project Panama.
  • Anthropic used scanned books to train its Claude LLM.
  • Anna's Archive calls for a global volunteer effort to digitize books.
🕒 2026-08-19 · new reporting from 404 Media
  • The story was featured on a podcast.
🕒 2026-08-19 · new reporting from 9to5Mac
  • Booksellers receive massive bulk orders of used books with no apparent connection between titles.
🕒 2026-08-18 · new reporting from Tom's Hardware
  • The rare book was part of a 1,000-book order.
  • VGT3 is part of a larger facility named LAS8.
  • The investigation started in July.
🕒 2026-08-17 · new reporting from TechCrunch, Ars Technica
  • Amazon's statement to 404 Media did not mention AI training.
  • The VGT3 facility logo is a dinosaur holding a book.
  • 404 Media connected with a bookseller to plant the AirTag.
🕒 2026-08-17 · new reporting from 404 Media
  • Amazon is buying and scanning physical books for AI training.
  • Amazon destroys books after scanning them for AI training.
  • A tracking device revealed Amazon's book operation.
  • The final destination was an Amazon warehouse in Las Vegas, Nevada.
  • Amazon employees cut book bindings for faster scanning.
  • The Amazon team at the warehouse is called VGT3.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~15 min · 12 stories · Sep 05

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

How outlets covered it

An Amazon employee confirmed the existence and operation of a warehouse (VGT3) in Las Vegas where books are destructively scanned for AI training data. This process involves cutting off book spines, scanning pages, and then discarding the physical books, raising questions about the sourcing and disposal of materials used for AI development.

The legality of training AI models on copyrighted material is complex, with a recent ruling indicating that while training itself may be lawful, obtaining the data from illegal sources is not. This distinction suggests a nuanced interpretation of copyright law in the context of AI development, impacting both AI companies and content creators.

Anna's Archive, a large open library, has called for volunteers to scan and upload books online, citing concerns that AI companies are buying, scanning, and destroying physical books to train AI models. This practice, which often involves destructive scanning methods, is seen as a way to monopolize knowledge and gain an advantage in AI development.

Anna's Archive claims AI companies are purchasing, scanning, and then destroying physical books to train their models, citing Anthropic's "Project Panama" as an example. This practice allegedly monopolizes knowledge on private servers and removes it from the public domain, prompting Anna's Archive to call for a global volunteer effort to digitize books.

Amazon is reportedly acquiring and destroying rare books to use their content for training artificial intelligence models. This practice raises concerns about the preservation of physical literary heritage and the ethical implications of AI data sourcing.

An AirTag placed in a bulk order of used books led to an Amazon AI training facility, confirming suspicions among booksellers that AI companies are acquiring books for scanning without author consent or payment. This investigation provides concrete evidence for a practice previously only theorized by booksellers. The practice raises ethical and legal questions regarding intellectual property and fair compensation for authors.

An investigation by 404 Media used a tracking device in a rare book to trace it to an Amazon facility (VGT3) in Las Vegas, where employees reportedly cut spines and scan books for AI model training. This discovery provides insight into the physical process of acquiring and digitizing books for large language models, confirming suspicions about the destination of unusually large book orders.

An AirTag hidden in a rare book tracked its journey to an Amazon AI training facility in Las Vegas, where books were reportedly scanned and destroyed. This discovery suggests Amazon is acquiring and damaging rare books to obtain unique training data for its frontier AI models, a practice that other AI firms have publicly disavowed.

Amazon is reportedly purchasing rare books, removing their spines, and scanning them to acquire data for AI model training. This practice provides unique, pre-2022 text data, which is valuable for preventing "model collapse" in large language models (LLMs) that might otherwise ingest too much AI-generated content.

An investigation revealed Amazon is purchasing large quantities of physical books, scanning them for AI model training, and destroying them in the process at a facility in Las Vegas. This practice addresses the demand for unique, non-internet-scraped text data for AI models, particularly to avoid "model collapse" from AI-generated content.

Booksellers and book lovers are concerned that AI companies are destroying physical books, including rare ones, to quickly scan them for training data. This practice, while efficient for data acquisition, raises fears about the permanent loss of unique texts. Non-destructive scanning methods exist but are slower and more costly, which AI firms may avoid.

Independent bookstores in Ireland and Germany are receiving large orders for obscure and outdated books, leading to concerns that AI companies are acquiring them for training data, potentially resulting in their destruction. This follows reports of AI companies shredding millions of books after digitization and ongoing lawsuits regarding the use of copyrighted material for AI training.

ISBNdb removed a webpage offering to source printed books for AI training datasets and stated the service was only a "test of market interest" after reporting by 404 Media. The company clarified it has never purchased, scanned, or sold books for AI training and does not train AI models. This incident highlights ongoing concerns about data sourcing for AI models and the transparency of companies operating in this space.

AI companies are reportedly purchasing millions of secondhand books through intermediaries to acquire high-quality training data for their AI models, with some books being destroyed after use. This practice is driven by the need for original, uncontaminated content to train AI and avoid "AI slop" from internet data, raising concerns about the future of literary history and copyright infringement.

AI companies are increasingly purchasing printed books published before 2022 as training data to avoid contamination from AI-generated text. ISBNdb facilitates bulk acquisitions, claiming these books provide high-quality, curated information that doesn't suffer from issues like model collapse.