Senior Researcher, Interpretability and AI Safety

Corehr

Oxford

Hybrid

GBP 50,000 - 70,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

The University of Oxford’s Department of Engineering Science seeks a full-time Senior Researcher to advance interpretability and AI safety for continually learning systems in the central Oxford area. The role focuses on evaluating dynamic models, designing experiments, and documenting results to ensure safety properties hold under changing data.

You will join the Technical Safety and Governance lab and work with the Oxford Martin AI Governance Initiative, contributing to collaboration,

Qualifications

  • PhD in ML, CS, or closely related field.
  • Several years of research experience.
  • Experience in interpretability, evaluation, lifelong learning, or AI safety.
  • Strong publication record in top venues (NeurIPS/ICML/ICLR/ACL/EMNLP).

Responsibilities

  • Research on interpretability and AI safety for continually learning systems.
  • Design, run, and analyze large-scale experiments on foundation models.
  • Collaborate with the TSG lab and Oxford Martin AI Governance Initiative.
  • Publish findings and contribute to governance-focused research.

Skills

Interpretable ML
Model evaluation
Continual learning
AI safety
Publication record

Education

PhD in ML/CS or related

Job description

Senior Researcher in Interpretability and AI Safety

central Oxford

We are seeking a full time Senior Researcher to joinTechnical Safety and Governance (TSG) labat the Department of Engineering Science in central Oxford. The post is funded by the Oxford Martin AI Governance Initiative and is fixed-term to 1 year, with the possibility of an extension for an additional year.

The Senior Researcher will work on Interpretability, evaluations and AI safety for continuously learning systems. They will evaluate systems whose behaviour keeps moving, and track, from the inside, whether the structures that carry their capabilities and safety properties hold up under the change.

You will be responsible for the follow : (full details of duties available from the Job Description)

Research
Collaboration and engagement

You will have completed a PhD in machine learning, computer science, or a closely related field and with several years of experience and have research experience in at least one of: mechanistic or representational interpretability; model evaluation and benchmarking; continual or lifelong learning; or AI safety (alignment, scalable oversight, adversarial robustness). A strong publication record at relevant venues (NeurIPS, ICML, ICLR, ACL, EMNLP, or established AI safety venues and workshops) together with the ability to design, run, and make sense of large-scale experiments on foundation models.

Informal enquiries may be addressed to Fazl Barez at fazl.barez@eng.ox.ac.uk.

For more information about working at the Department, see

www.eng.ox.ac.uk/about/work-with-us/

The Department holds an Athena Swan Bronze award, highlighting its commitment to promoting women in Science, Engineering and Technology.

AI, interpretability, continuous learning, evaluations chain-of-thought

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Researcher, Interpretability and AI Safety
Senior Researcher, Interpretability and AI Safety

AISafety • Oxford

Hybrid
GBP 49,000 - 58,000
Senior Researcher in Interpretability and AI Safety
Senior Researcher in Interpretability and AI Safety

Corehr • Oxford

On-site
GBP 60,000 - 80,000
Senior Researcher in Interpretability and AI Safety Closing date: Sep 22, 2026
Senior Researcher in Interpretability and AI Safety Closing date: Sep 22, 2026

University of Oxford, Department of Engineering Science • Oxford

On-site
GBP 49,000 - 58,000
Senior Researcher: AI Interpretability & Safety
Senior Researcher: AI Interpretability & Safety

University of Oxford, Department of Engineering Science • Oxford

On-site
GBP 49,000 - 58,000
Senior Researcher: Interpretability & AI Safety
Senior Researcher: Interpretability & AI Safety

Corehr • Oxford

Hybrid
GBP 50,000 - 70,000
Senior Researcher: AI Interpretability & Safety
Senior Researcher: AI Interpretability & Safety

AISafety • Oxford

Hybrid
GBP 49,000 - 58,000
Senior Researcher — AI Interpretability & Safety
Senior Researcher — AI Interpretability & Safety

Corehr • Oxford

On-site
GBP 60,000 - 80,000
Research Assistant on AI Safety
Research Assistant on AI Safety

Corehr • Oxford

On-site
GBP 36,000 - 39,000
AI Safety Research Scientist — Build Safe, Trustworthy AI
AI Safety Research Scientist — Build Safe, Trustworthy AI

Faculty • Greater London

Hybrid
GBP 90,000 - 130,000
Unlimited Annual Leave Policy
Private healthcare and dental
Enhanced parental leave
+3
[Expression of Interest] Research Engineer / Scientist, Alignment - London
[Expression of Interest] Research Engineer / Scientist, Alignment - London

Menlo Ventures • Greater London

On-site
GBP 260,000 - 370,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
+1