The University of Oxford has granted OpenAI permission to utilize digitized historical texts from its Bodleian Library for training artificial intelligence models. Internal documents confirm that this material has been used to "populate the OpenAI training set."
This arrangement follows a partnership announced in March 2025, which initially focused on using OpenAI software to digitize library texts to increase accessibility for students and researchers. The initial announcement did not specify that the material would be used for AI model training.
Tech companies are actively seeking fresh data sources for AI training as existing online data becomes increasingly saturated with AI-generated content, reducing its utility. Physical and historical book collections, often not digitized online, represent valuable, untainted data for developing new AI models.
Reports from booksellers indicate a rise in orders for obscure historical titles, suggesting these unique texts are being acquired for digitization and subsequent AI training.
Meeting minutes from the University of Oxford reveal concerns among staff, including members of the Bodleian governance committee, regarding the reputational risks associated with partnering with OpenAI. Staff also raised questions about the environmental impact of supporting energy-intensive AI technology.
Oxford is the sole UK institution participating in OpenAI's NextGenAI project, which includes similar agreements with US research libraries such as Boston Public Library, Caltech, MIT, and the University of Michigan. By June 2025, 125,000 scanned images, primarily 19th and 20th-century PhD theses, had been shared from the Bodleian collection.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
The University of Oxford has permitted OpenAI to use digitized historical texts from its Bodleian Library to train AI models. This partnership, initially announced for digitization purposes, now explicitly includes data for AI training, as tech companies seek new, non-AI-generated data sources.