Senior Machine Learning Engineer, Safety & Alignment

AI Breaking Wire

San Francisco, Northern (CA, KY)

Hybrid

USD 250,000 - 400,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Top-tier salary and equity grants
Comprehensive medical, dental, and eye

Job summary

Anthropic, based in San Francisco, is seeking a Senior Machine Learning Engineer to advance AI safety and alignment as our models scale. You will help build reliable, interpretable, and steerable systems that behave safely and beneficially.

You will implement and scale alignment techniques such as RLHF and Constitutional AI, develop automated evaluation pipelines for vulnerabilities, and partner with policy and research teams to embed safety standards into model training.

Qualifications

  • 4+ years of industry experience building and deploying ML models at scale.
  • Strong proficiency in Python, PyTorch, and distributed computing frameworks (e.g., Ray, DeepSpeed).
  • Deep familiarity with safety challenges in generative AI and large language models.

Responsibilities

  • Implement and scale alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI.
  • Develop automated evaluation frameworks to detect vulnerabilities, hallucinations, and harmful outputs.
  • Partner with policy and research teams to operationalize safety standards into model training pipelines.
  • Optimize training infrastructure for safety-focused objective functions.

Skills

Python
PyTorch
Reinforcement learning
AI safety

Education

Ph.D. in Computer Science
M.S. in Computer Science
B.S. in Computer Science

Tools

Ray
DeepSpeed

Job description

# Senior Machine Learning Engineer, Safety & AlignmentAnthropic## Job DescriptionAnthropic is looking for a Senior Machine Learning Engineer to focus on AI Safety and Alignment. You will help build reliable, interpretable, and steerable AI systems, ensuring our models behave safely and beneficially as they scale.### Responsibilities- Implement and scale alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI.- Develop automated evaluation frameworks to detect vulnerabilities, hallucinations, and harmful outputs.- Partner with policy and research teams to operationalize safety standards into model training pipelines.- Optimize training infrastructure for safety-focused objective functions.### Requirements- 4+ years of industry experience building and deploying machine learning models at scale.- Strong proficiency in Python, PyTorch, and distributed computing frameworks (e.g., Ray, DeepSpeed).- Deep familiarity with safety challenges in generative AI and large language models.- B.S., M.S., or Ph.D. in Computer Science or a related quantitative discipline.### Benefits- Top-tier compensation including competitive salary and equity grants.- Comprehensive medical, dental, and vision plans.- Parental leave benefits and flexible PTO.- Wellness stipends and professional development budgets.## Skills & Tagspythonpytorchrlhfsafetyllm## Job DetailsFull-timeSan Francisco, CA$250k – $400k USDPosted August 10, 2026Expires October 9, 2026
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer: AI Safety & Alignment (RLHF)
Senior ML Engineer: AI Safety & Alignment (RLHF)

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 400,000
Top-tier salary and equity grants
Comprehensive medical, dental, and eye
Research Scientist, Alignment
Research Scientist, Alignment

AI Breaking Wire • San Francisco (CA)

On-site
USD 300,000 - 450,000
Competitive salary and equity packages
Comprehensive health, dental, and visa
Flexible working arrangements and PTO
+1
Machine Learning Engineer, Trust & Safety
Machine Learning Engineer, Trust & Safety

Aibreakingwire • San Francisco (CA)

On-site
USD 160,000 - 260,000
Mission-driven culture
Competitive compensation
Machine Learning Engineer, Trust & Safety
Machine Learning Engineer, Trust & Safety

AI Breaking Wire • San Francisco (CA)

On-site
USD 160,000 - 260,000
Research Engineer / Scientist, Alignment
Research Engineer / Scientist, Alignment

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 500,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Research Engineer: AI Safety & Alignment
Research Engineer: AI Safety & Alignment

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 500,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
AI Safety Research Engineer
AI Safety Research Engineer

Neura Market • San Francisco (CA)

Hybrid
USD 350,000 - 500,000
Research Scientist, AI Alignment & Safety
Research Scientist, AI Alignment & Safety

AI Breaking Wire • San Francisco (CA)

On-site
USD 300,000 - 450,000
Competitive salary and equity packages
Comprehensive health, dental, and visa
Flexible working arrangements and PTO
+1
Research Engineer / Scientist, Alignment Anthropic San Francisco, CA
Research Engineer / Scientist, Alignment Anthropic San Francisco, CA

Neura Market • San Francisco (CA)

Hybrid
USD 350,000 - 500,000
Research Engineer, AI Safety & Alignment
Research Engineer, AI Safety & Alignment

character • Redwood City (CA)

On-site
USD 120,000 - 150,000