AI Safety and Alignment Researcher

AI Breaking Wire

Greater London

Hybrid

GBP 90,000 - 150,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity / stock options
Health, wellness & family care
World-class computing resources
Generous sabbatical and time-off

Job summary

Google DeepMind is seeking an AI Safety and Alignment Researcher to develop rigorous empirical and theoretical frameworks that keep advanced AI systems safe, robust, and aligned with human values.

The role focuses on foundational and applied research across model interpretability, robustness, and scalable oversight, with strong ties to product and engineering teams to embed safety into deployments.

Qualifications

  • PhD (preferred) in a quantitative field with strong research track record.
  • Strong background in ML safety, alignment, interpretability, or adversarial robustness.
  • Proficiency in Python and deep learning frameworks (TensorFlow or PyTorch).
  • Excellent communication skills with peer-reviewed publications.

Responsibilities

  • Conduct fundamental and applied research into model interpretability, robustness, and scalable oversight.
  • Design evaluation benchmarks to test model vulnerabilities against jailbreaks, prompt injection, and deceptive alignment.
  • Partner with product and engineering teams to integrate safety guardrails into deployment pipelines.
  • Collaborate with academic institutions and external research bodies on safety standards.

Skills

ML safety
Interpretability
Adversarial robustness
Academic publishing

Education

PhD in CS/Math/Physics/Philosophy

Tools

TensorFlow
PyTorch

Job description

About the Role

Google DeepMind is at the forefront of artificial intelligence research. We are seeking an AI Safety and Alignment Researcher to develop empirical and theoretical frameworks for ensuring advanced AI systems remain safe, robust, and aligned with human values.

Responsibilities
  • Conduct fundamental and applied research into model interpretability, robustness, and scalable oversight.
  • Design rigorous evaluation benchmarks to test model vulnerabilities against jailbreaks, prompt injection, and deceptive alignment.
  • Partner with product and engineering teams to integrate safety guardrails into deployment pipelines.
  • Collaborate with academic institutions and external research bodies on safety standards.
Requirements
  • Advanced degree (Ph.D. preferred) in Computer Science, Mathematics, Physics, or Philosophy with a heavy quantitative focus.
  • Strong background in machine learning safety, alignment, interpretability, or adversarial robustness.
  • Proficiency in Python and deep learning frameworks (TensorFlow or PyTorch).
  • Excellent communication skills with a proven track record of peer-reviewed publications.
Benefits
  • Industry-leading compensation including base salary and equity/stock options.
  • Comprehensive health, wellness, and family care benefits.
  • Access to world-class computing infrastructure and research resources.
  • Generous sabbatical and time-off policies.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Safety Researcher
Senior AI Safety Researcher

AI Breaking Wire • Greater London

Hybrid
GBP 110,000 - 170,000
Stock options
Health programs and wellness benefits
Learning and development allowances
+1
Lead AI Safety & Alignment Research Scientist
Lead AI Safety & Alignment Research Scientist

AI Breaking Wire • Greater London

Hybrid
GBP 90,000 - 150,000
Equity / stock options
Health, wellness & family care
World-class computing resources
+1
Senior AI Safety Researcher — Lead Alignment & Robustness
Senior AI Safety Researcher — Lead Alignment & Robustness

AI Breaking Wire • Greater London

Hybrid
GBP 110,000 - 170,000
Stock options
Health programs and wellness benefits
Learning and development allowances
+1
Research Scientist, AGI Safety and Alignment, DeepMind
Research Scientist, AGI Safety and Alignment, DeepMind

Google LLC • Greater London

On-site
GBP 70,000 - 110,000
Equity grants
AI Safety Lead
AI Safety Lead

Best AI Tools Wiki • Greater London

Hybrid
GBP 180,000 - 240,000
Relocation to London supported
Private healthcare for family
Generous research budget
+2
Research Engineer, Responsible Frontier AI Research, DeepMind
Research Engineer, Responsible Frontier AI Research, DeepMind

Google Inc. • Greater London

On-site
GBP 90,000 - 130,000
AGI Safety & Alignment Research Scientist
AGI Safety & Alignment Research Scientist

Google LLC • Greater London

On-site
GBP 70,000 - 110,000
Equity grants
Researcher, Frontier Risk Mitigations
Researcher, Frontier Risk Mitigations

OpenAI • Greater London

On-site
GBP 218,000 - 328,000
Researcher, Robustness & Safety Training
Researcher, Robustness & Safety Training

OpenAI • Greater London

Hybrid
GBP 283,000 - 372,000
Research Engineer: AI Safety & Alignment
Research Engineer: AI Safety & Alignment

Anthropic • Greater London

Hybrid
GBP 260,000 - 370,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
+1