BodleianLLM
Training open language models on cultural heritage data for humanities research.
Start Date: September 2026
End Date: August 2027
Funder: University of Oxford, John Fell Fund
Industry-built AI language models dominate research practice today, yet they are opaque, closed-source, and trained overwhelmingly on modern English web text, making them ill-equipped for the historical depth, linguistic range, and scholarly precision that humanities research demands.
BodleianLLM proposes to train an open language model on the cultural heritage collections that Oxford has progressively digitised over the past several decades. What distinguishes this material is its quality: these are editorially vetted, authoritative scholarly texts whose provenance and accuracy have been established through generations of curatorial and philological work that no commercial web crawl can replicate.
The project has four main goals.
- To assemble and document a humanities training corpus from Oxford's collections.
- To design the first-ever benchmark suite for testing how well AI models perform on humanities tasks such as manuscript transcription, historical named-entity recognition, and cross-lingual literary analysis.
- To train and rigorously evaluate a language model using Oxford's own computing infrastructure.
- To publish everything - model, code, data, and benchmarks - as fully open-source resources for use by other institutions worldwide.
The project also responds to a broader concern. Because industry now controls almost all LLM development, the priorities embedded in these models – what they are trained on, what they are optimised to do, and the standards by which they are evaluated – are determined by commercial interests, with little regard for the needs of scholarship. For the humanities, where research questions are shaped by languages, time periods, and forms of reasoning that often have little commercial value, this is a gap that market forces are unlikely to close. universities that do not participate in the development of language models will remain entirely dependent on tools built for other purposes. BodleianLLM is a deliberate attempt to establish an academic presence in this field. Alongside the model itself, it will release a fully open-source pipeline – training code, data documentation, and benchmarks – that other institutions can adapt to their own collections. This commitment to transparency and reproducibility sets the project apart from most current LLM development, in which even nominally ‘open’ models typically publish weights without the data or methodology needed for independent replication.
The project runs for twelve months from September 2026 and will be led by the PI alongside DiSc Senior Research Software Engineer Miguel Arana-Catania and a postdoctoral researcher with Megan Gooch (Head of the Centre for Digital Scholarship, Bodleian Libraries) coordinating engagement across the Humanities and GLAM divisions, and compute infrastructure and additional support provided by Digital Scholarship at Oxford (DiSc).
The project’s outputs will support major follow-on applications to the AHRC, UKRI, and ERC for a larger scale op en-source language model for humanities research.
Contact the project team at glenn.roe@humanities.ox.ac.uk.
Project team
Principal Investigator: Professor Glenn Roe
Co-Investigator: Dr Megan Gooch
Research Software Engineer: Dr Miguel Arana-Catania
Postdoctoral Researcher: to be appointed
Image credit: Bodleian Libraries: https://hdl.handle.net/2027/oxu1.604487234