← All stories
● Covered by 1 source · 1 reportMedium impact

AI Companies Source Old Books for Training Data, Avoiding AI-Generated Text

New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

AI companies are increasingly purchasing printed books published before 2022 as training data to avoid contamination from AI-generated text. ISBNdb facilitates bulk acquisitions, claiming these books provide high-quality, curated information that doesn't suffer from issues like model collapse.

Key points

  • AI companies seek pre-2022 printed books for training data.
  • Printed books avoid AI-generated text contamination.
  • ISBNdb offers bulk acquisition services for AI firms.

AI Companies Target Old Books

AI companies are actively seeking old, printed books as a resource for training data. Due to the prevalence of AI-generated text on the internet, these older books provide a fresh and reliable source. ISBNdb, an existing provider of book data and services, emphasizes that printed titles from before 2022 are untainted by current AI slop.

ISBNdb's Role in Acquisitions

ISBNdb has pivoted to serve the needs of AI companies by offering high-volume book acquisition services, allowing purchases ranging from 1,000 to 1 million books. This move capitalizes on the growing demand for quality training data, ensuring AI labs can acquire and digitize printed books efficiently.

Risks of AI-Generated Data

Concerns are rising about the risks associated with training AI models on data that includes AI-generated content. The phenomenon known as model collapse may occur if these models are trained on contaminated data, leading to increased errors. Thus, the integrity of data sources is critical for the effectiveness of AI systems.

Copyright Issues and Scanning Controversies

Recent legal actions, notably a copyright lawsuit involving Anthropic, highlighted controversies surrounding the acquisition of printed books for AI training. The internal documents revealed plans to scan and then potentially destroy these books, raising ethical questions about data sourcing in AI development.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~39 min · 35 stories · Jul 22

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

AI companies are increasingly purchasing printed books published before 2022 as training data to avoid contamination from AI-generated text. ISBNdb facilitates bulk acquisitions, claiming these books provide high-quality, curated information that doesn't suffer from issues like model collapse.