University of Oxford Partners with OpenAI to Train AI Models on Bodleian Library Texts
Key Takeaways
- •OpenAI will be able to train its AI models on historical texts from Oxford's Bodleian Libraries under a five-year partnership announced on March 4, 2025.
- •A February pilot project is digitizing roughly 3,500 public-domain dissertations dating from 1498 to 1884, aiming to make the centuries-old collections searchable and globally accessible.
- •The agreement is part of OpenAI's NextGenAI program, which has committed $50 million in research grants, computing resources, and API access to elite institutions including Harvard and MIT.
- •Oxford plans to roll out ChatGPT Edu across the university beginning in September 2025 after piloting it with about 750 users, and open-access research reports the digitization project are expected by early 2026.
- •Some university staff have raised concerns about the reputational risk of partnering with OpenAI over its copyright controversies and data sourcing practices, even though the pilot dissertations are public domain.

The University of Oxford has entered a five-year partnership with OpenAI that will allow the company behind ChatGPT to train its AI models on historical texts held in the Bodleian Libraries, one of the oldest and most renowned library systems in the world.
The collaboration, announced on March 4, 2025, began with a pilot project in February focused on digitizing roughly 3,500 public-domain dissertations dating from 1498 to 1884. These are centuries-old texts that had previously been largely inaccessible to anyone unable to physically visit Oxford, which meant that much of this material sat effectively out of reach for the global research community.
What the Partnership Involves
The agreement falls under OpenAI's NextGenAI program, which has committed $50 million in research grants, computing resources, and API access to a group of elite institutions that includes Harvard and MIT alongside Oxford.
For the Bodleian Libraries, the stated goal is straightforward: AI-powered metadata enrichment and improved transcription services will make the centuries-old collections searchable and globally accessible.
On the education side, Oxford plans to roll out ChatGPT Edu across the university beginning in September 2025. The tool has already been piloted with about 750 users. Research findings from the digitization project are expected in open-access reports by early 2026, giving observers concrete checkpoints for judging whether the partnership delivers on its stated goals.
Staff Concerns
Not everyone at Oxford is celebrating the arrangement. University staff have voiced concerns about the reputational risk of partnering with OpenAI, a company that has faced sustained criticism over copyright issues, data sourcing practices, and its rapid commercialization under CEO Sam Altman.
The dissertations covered by the pilot are public domain, meaning no copyright applies. Still, the optics of a roughly 900-year-old academic institution feeding its collections to a for-profit AI company have drawn scrutiny.
There is also the question of what each side gains. Oxford receives digitization support, compute resources, and educational tools. OpenAI receives high-quality, curated training data from a globally trusted source, along with the implicit endorsement that comes with an Oxford partnership — an endorsement whose value is heightened by the fact that the company's data practices remain contested.
The Bigger Picture
Oxford's deal fits a pattern that has been accelerating across higher education. Technology companies are increasingly turning to academic institutions as data sources, partly because web-scraped data is running into legal challenges and partly because academic corpora tend to be higher quality than the average internet.
Against that backdrop, how Oxford's ChatGPT Edu rollout and digitization reports are received will be watched by institutions weighing similar arrangements. The NextGenAI program's $50 million commitment, spread across multiple institutions, is substantial but modest relative to OpenAI's overall spending, which has been measured in billions of dollars annually.