Research Scientist, Alignment

AI Breaking Wire

San Francisco (CA)

On-site

USD 300,000 - 450,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary and equity packages
Comprehensive health, dental, and visa
Flexible working arrangements and PTO
Catered lunches and state-of-the-art办公

Job summary

Anthropic in San Francisco is seeking a Research Scientist to join the Alignment team, focusing on keeping frontier AI systems safe, helpful, and honest as capabilities scale.

You will design empirical experiments, develop novel alignment methods like RLHF and Constitutional AI, and collaborate with infrastructure and engineering to scale training runs safely. Outstanding researchers with a PhD or equivalent and strong communication are encouraged to apply.

Qualifications

  • Ph.D. or equivalent practical experience in ML, CS, statistics or related field.
  • Strong background in deep learning, PyTorch, and large-scale model training.
  • Demonstrated track record of impactful research in AI alignment, safety, or robustness.
  • Excellent communication skills and commitment to responsible AI development.

Responsibilities

  • Design and execute empirical experiments to understand safety properties of large language models.
  • Develop alignment techniques, including RLHF and Constitutional AI.
  • Collaborate with infrastructure and engineering teams to scale training runs safely.
  • Publish research findings and contribute to the broader scientific understanding of AI safety.

Skills

Python
PyTorch
ML alignment
RLHF
Research

Education

Ph.D. or equivalent in ML/CS/Statistics

Job description

# Research Scientist, AlignmentAnthropic## Job Description### About the RoleAnthropic is looking for a Research Scientist to join our Alignment team. You will play a crucial role in ensuring that frontier AI systems remain safe, helpful, and honest as capabilities scale.### Responsibilities- Design and execute empirical experiments to understand the safety properties of large language models.- Develop novel alignment techniques, including Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI.- Collaborate closely with infrastructure and engineering teams to scale training runs safely.- Publish research findings and contribute to the broader scientific understanding of AI safety.### Requirements- Ph.D. or equivalent practical experience in Machine Learning, Computer Science, Statistics, or a related field.- Strong background in deep learning, PyTorch, and large-scale model training.- Demonstrated track record of impactful research in AI alignment, safety, or robustness.- Excellent communication skills and a deep commitment to responsible AI development.### Benefits- Competitive salary and equity packages.- Comprehensive health, dental, and vision insurance.- Flexible working arrangements and generous PTO.- Catered lunches and state-of-the-art office facilities.## Skills & Tagspythonpytorchllmalignmentrlhfresearch## Job DetailsFull-timeSan Francisco, CA$300k – $450k USDPosted July 30, 2026Expires September 28, 2026
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist, AI Alignment & Safety
Research Scientist, AI Alignment & Safety

AI Breaking Wire • San Francisco (CA)

On-site
USD 300,000 - 450,000
Competitive salary and equity packages
Comprehensive health, dental, and visa
Flexible working arrangements and PTO
+1
Senior Machine Learning Engineer, Safety & Alignment
Senior Machine Learning Engineer, Safety & Alignment

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 400,000
Top-tier salary and equity grants
Comprehensive medical, dental, and eye
Research Engineer: AI Safety & Alignment
Research Engineer: AI Safety & Alignment

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 500,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Research Engineer / Scientist, Alignment
Research Engineer / Scientist, Alignment

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 500,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
AI Safety Research Engineer
AI Safety Research Engineer

Neura Market • San Francisco (CA)

Hybrid
USD 350,000 - 500,000
Researcher, Alignment
Researcher, Alignment

OpenAI • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Researcher, Alignment
Researcher, Alignment

OpenAI • Los Angeles (CA)

Hybrid
USD 120,000 - 150,000
Relocation assistance
Hybrid work model
Inclusive workplace
Research Engineer / Scientist, Alignment Anthropic San Francisco, CA
Research Engineer / Scientist, Alignment Anthropic San Francisco, CA

Neura Market • San Francisco (CA)

Hybrid
USD 350,000 - 500,000
Research Engineer / Scientist, Alignment Science
Research Engineer / Scientist, Alignment Science

Safetytalent • San Francisco (CA)

On-site
USD 100,000 - 150,000
AI Alignment Research Scientist — Safety & Reasoning
AI Alignment Research Scientist — Safety & Reasoning

Safetytalent • San Francisco (CA)

On-site
USD 100,000 - 150,000