Senior ML Engineer: AI Safety & Alignment (RLHF)

AI Breaking Wire

San Francisco, Northern (CA, KY)

Hybrid

USD 250,000 - 400,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Top-tier salary and equity grants
Comprehensive medical, dental, and eye

Job summary

Anthropic, based in San Francisco, is seeking a Senior Machine Learning Engineer to advance AI safety and alignment as our models scale. You will help build reliable, interpretable, and steerable systems that behave safely and beneficially.

You will implement and scale alignment techniques such as RLHF and Constitutional AI, develop automated evaluation pipelines for vulnerabilities, and partner with policy and research teams to embed safety standards into model training.

Qualifications

  • 4+ years of industry experience building and deploying ML models at scale.
  • Strong proficiency in Python, PyTorch, and distributed computing frameworks (e.g., Ray, DeepSpeed).
  • Deep familiarity with safety challenges in generative AI and large language models.

Responsibilities

  • Implement and scale alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI.
  • Develop automated evaluation frameworks to detect vulnerabilities, hallucinations, and harmful outputs.
  • Partner with policy and research teams to operationalize safety standards into model training pipelines.
  • Optimize training infrastructure for safety-focused objective functions.

Skills

Python
PyTorch
Reinforcement learning
AI safety

Education

Ph.D. in Computer Science
M.S. in Computer Science
B.S. in Computer Science

Tools

Ray
DeepSpeed

Job description

Anthropic, based in San Francisco, is seeking a Senior Machine Learning Engineer to advance AI safety and alignment as our models scale. You will help build reliable, interpretable, and steerable systems that behave safely and beneficially.

You will implement and scale alignment techniques such as RLHF and Constitutional AI, develop automated evaluation pipelines for vulnerabilities, and partner with policy and research teams to embed safety standards into model training.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer - AI Alignment & RLHF (Remote)
Senior ML Engineer - AI Alignment & RLHF (Remote)

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Equity compensation
Competitive salary
Remote and hybrid options
+1
Senior Machine Learning Engineer, Safety & Alignment
Senior Machine Learning Engineer, Safety & Alignment

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 400,000
Top-tier salary and equity grants
Comprehensive medical, dental, and eye
Senior Machine Learning Engineer, Alignment and Safety
Senior Machine Learning Engineer, Alignment and Safety

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Equity compensation
Competitive salary
Remote and hybrid options
+1
Research Scientist, AI Alignment & Safety
Research Scientist, AI Alignment & Safety

AI Breaking Wire • San Francisco (CA)

On-site
USD 300,000 - 450,000
Competitive salary and equity packages
Comprehensive health, dental, and visa
Flexible working arrangements and PTO
+1
Research Engineer: AI Safety & Alignment
Research Engineer: AI Safety & Alignment

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 500,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
AI Alignment Research Scientist — Safety & Reasoning
AI Alignment Research Scientist — Safety & Reasoning

Safetytalent • San Francisco (CA)

On-site
USD 100,000 - 150,000
Research Scientist, Alignment
Research Scientist, Alignment

AI Breaking Wire • San Francisco (CA)

On-site
USD 300,000 - 450,000
Competitive salary and equity packages
Comprehensive health, dental, and visa
Flexible working arrangements and PTO
+1
AI Alignment Research Engineer — Evaluation & Safety
AI Alignment Research Engineer — Evaluation & Safety

W3 Sourcing • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
AI Safety Evaluator & Model Alignment Specialist
AI Safety Evaluator & Model Alignment Specialist

Great Value Hiring • United States

On-site
RL Systems Engineer: Build Safe, Steerable AI
RL Systems Engineer: Build Safe, Steerable AI

Anthropic • New York (NY)

Hybrid
USD 500,000 - 850,000
Equity donation matching
Generous vacation
Parental leave
+2