Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
UAG in Oxford is offering a fixed-term 12-month role to deliver the BodleianLLM project: curating multilingual corpora, continual pre-training and fine-tuning on the University's infrastructure, and developing benchmarks for humanities tasks.
You will publish model weights, code and benchmarks as open-source resources, contribute to funding bids, collaborate on publications and workshops, and report to the PI with day-to-day supervision from DiSc.
Location:Humanities Divisional Office, Stephen A. Schwarzman Centre for the Humanities, Radcliffe Observatory Quarter, Woodstock Road, Oxford, OX2 6GG
Contract Fixed term for 12 months
Hours Full-time (37.5 per week)
The postholder is expected primarily to deliver the core objectives of the BodleianLLM project: curating, cleaning and documenting a multilingual humanities training corpus drawn from Oxford's GLAM collections; performing continual pre-training and fine-tuning of an open-weight foundation model on the University's own computing infrastructure; and, working with subject specialists and curators, designing and validating the first dedicated benchmark suite for evaluating language models on humanities research tasks.
You will also be expected to publish the model weights, code, corpus documentation and benchmarks as fully open-source resources; to contribute to the technical roadmap for a major follow-on funding bid; to collaborate on research publications and present at conferences; and to help deliver the project's one-day workshop for the GLAM, digital humanities and AI communities. You will report to the Principal Investigator, Professor Glenn Roe, with day-to-day technical supervision from a Senior Research Software Engineer at Digital Scholarship at Oxford (DiSc).
You will hold a relevant PhD/DPhil or have substantial experience of machine learning in a research or industry environment, and have the ability and willingness to combine machine learning research with sustained engagement with historical and multilingual source material. You will have a strong understanding of deep learning and natural language processing, with practical experience of training, adapting and deploying large language models.
You will also be able to process and transform historical, multilingual or otherwise non-standard textual data and to design and validate evaluation benchmarks; have strong Python and software engineering skills; have a strong publication record commensurate with experience; and have excellent communication and organisational skills. Experience of cultural heritage collections, reading knowledge of historical languages, and familiarity with HPC and MLOps practice would be desirable.
The duties and skills required are described in further detail in the job description.
Committed to equality and valuing diversity.
£39,424 to £41,636 (Grade 7.1 - 7.3)