← All stories
● Covered by 1 source · 1 reportMedium impact1 neutral

Oxford's Bodleian Library allows OpenAI to train AI models on historical texts

🔄 Updated 1h ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Oxford's Bodleian Library data is used to populate OpenAI's training set.
  • The partnership was announced in March 2025 for text digitization.
  • Staff raised concerns about reputational risk and environmental impact.
  • OpenAI has similar agreements with US research libraries under NextGenAI.

OpenAI Accesses Bodleian Library Data

The University of Oxford has granted OpenAI permission to utilize digitized historical texts from its Bodleian Library for training artificial intelligence models. Internal documents confirm that this material has been used to "populate the OpenAI training set."

This arrangement follows a partnership announced in March 2025, which initially focused on using OpenAI software to digitize library texts to increase accessibility for students and researchers. The initial announcement did not specify that the material would be used for AI model training.

Motivation for New Data Sources

Tech companies are actively seeking fresh data sources for AI training as existing online data becomes increasingly saturated with AI-generated content, reducing its utility. Physical and historical book collections, often not digitized online, represent valuable, untainted data for developing new AI models.

Reports from booksellers indicate a rise in orders for obscure historical titles, suggesting these unique texts are being acquired for digitization and subsequent AI training.

Internal Concerns and Broader Partnerships

Meeting minutes from the University of Oxford reveal concerns among staff, including members of the Bodleian governance committee, regarding the reputational risks associated with partnering with OpenAI. Staff also raised questions about the environmental impact of supporting energy-intensive AI technology.

Oxford is the sole UK institution participating in OpenAI's NextGenAI project, which includes similar agreements with US research libraries such as Boston Public Library, Caltech, MIT, and the University of Michigan. By June 2025, 125,000 scanned images, primarily 19th and 20th-century PhD theses, had been shared from the Bodleian collection.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~12 min · 9 stories · Sep 26

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

The University of Oxford has permitted OpenAI to use digitized historical texts from its Bodleian Library to train AI models. This partnership, initially announced for digitization purposes, now explicitly includes data for AI training, as tech companies seek new, non-AI-generated data sources.