Senior Machine Learning Engineer, Alignment and Safety

AI Breaking Wire

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 260,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity compensation
Competitive salary
Remote and hybrid options
Catered lunches and wellness stipends

Job summary

Anthropic is seeking a Senior Machine Learning Engineer to advance Constitutional AI and model alignment. You will scale alignment pipelines to keep Claude safe, helpful, and honest as capabilities grow.

Join a team partnering with researchers to translate alignment theory into production training loops, develop scalable RLHF systems, and build automated evaluations for biases and vulnerabilities.

The role offers flexible remote/hybrid work and competitive compensation.

Qualifications

  • 4+ years of industry experience building, training, and deploying large-scale ML models.
  • Strong experience with PyTorch and distributed training infrastructure (Megatron-LM, DeepSpeed, or similar).
  • Solid foundation in alignment methods, RLHF, and automated red-teaming.
  • Passion for AI safety and responsible development of AGI.

Responsibilities

  • Build and optimize scalable RLHF and constitutional AI training loops.
  • Develop automated evaluation frameworks to detect model vulnerabilities, bias, and undesirable behaviors.
  • Partner with research scientists to translate alignment techniques into production-grade training systems.

Skills

RLHF experience
PyTorch
Distributed training
Alignment methodologies

Tools

Megatron-LM
DeepSpeed

Job description

About the Role

At Anthropic, we build reliable, interpretable, and steerable AI systems. We are looking for a Senior Machine Learning Engineer to focus on Constitutional AI and model alignment. In this role, you will scale our alignment pipelines to ensure our Claude models remain safe, helpful, and honest as they grow more capable.

Responsibilities
  • Build and optimize scalable reinforcement learning from human feedback (RLHF) and constitutional AI training loops.
  • Develop automated evaluation frameworks to detect model vulnerabilities, bias, and undesirable behaviors.
  • Partner with research scientists to translate theoretical alignment techniques into production-grade training systems.
Requirements
  • 4+ years of industry experience building, training, and deploying large-scale machine learning models.
  • Extensive experience with PyTorch and distributed training infrastructure (Megatron-LM, DeepSpeed, or similar).
  • Strong foundation in alignment methodologies, RLHF, and automated red-teaming.
  • Passion for AI safety and responsible development of AGI.
Benefits
  • Top-tier medical, dental, and vision benefits.
  • Generous equity allocation and competitive base salaries.
  • Flexible work culture with remote and hybrid options.
  • Catered lunches and wellness stipends.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Engineer, Safety & Alignment
Senior Machine Learning Engineer, Safety & Alignment

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 400,000
Top-tier salary and equity grants
Comprehensive medical, dental, and eye
Senior ML Engineer - AI Alignment & RLHF (Remote)
Senior ML Engineer - AI Alignment & RLHF (Remote)

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Equity compensation
Competitive salary
Remote and hybrid options
+1
Senior ML Engineer: AI Safety & Alignment (RLHF)
Senior ML Engineer: AI Safety & Alignment (RLHF)

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 400,000
Top-tier salary and equity grants
Comprehensive medical, dental, and eye
Research Engineer / Scientist, Alignment
Research Engineer / Scientist, Alignment

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 500,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Research Scientist, Alignment
Research Scientist, Alignment

AI Breaking Wire • San Francisco (CA)

On-site
USD 300,000 - 450,000
Competitive salary and equity packages
Comprehensive health, dental, and visa
Flexible working arrangements and PTO
+1
Research Engineer / Scientist, Alignment
Research Engineer / Scientist, Alignment

Anthropic • California (MO)

Hybrid
USD 350,000 - 500,000
Senior Machine Learning Engineer, Core Systems
Senior Machine Learning Engineer, Core Systems

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Equity
Health benefits
Remote-friendly US culture
+1
Research Engineer, RL Engineering
Research Engineer, RL Engineering

Anthropic • Seattle (WA)

Hybrid
USD 520,000 - 850,000
Equity donation matching
Generous vacation
Parental leave
+2
Research Engineer, RL Engineering
Research Engineer, RL Engineering

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+1
Staff+ Site Reliability Engineer, Safeguards ML Infra
Staff+ Site Reliability Engineer, Safeguards ML Infra

AI Chopping Block • New York (NY)

Hybrid
USD 405,000 - 485,000