AI Safety & Post-Training Scientist

Scale

San Francisco, New York (CA, NY)

On-site

USD 216,000 - 270,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Scale Labs, a leader in AI safety research, seeks a Research Scientist focused on Safety Post Training to advance post-training methods and interpretability techniques for safer frontier AI systems.

You will design post-training pipelines, study how training choices affect safety and alignment, and collaborate with policymakers, engineers, and researchers to translate findings into safety standards and evaluation benchmarks.

Qualifications

  • At least three years of experience addressing sophisticated ML problems, whether in a research setting or in product development.
  • Experience with post-training and RL techniques such as RLHF, DPO, GRPO, and similar approaches.
  • Strong written and verbal communication skills to operate in a cross-functional team.

Responsibilities

  • Develop and apply post-training methods and interpretability techniques to make frontier AI systems safer, and better understood by researchers and policymakers.
  • Design and run post-training pipelines to study how training choices affect model safety, robustness, and alignment properties.
  • Collaborate with policymakers, engineers, and other researchers to translate post-training and interpretability findings into actionable safety standards, evaluation benchmarks, and best practices.

Skills

Post-training techniques
RL techniques (RLHF, DPO, GRPO)
Generative AI research
3+ years ML experience
Cross-functional communication

Job description

Scale Labs, a leader in AI safety research, seeks a Research Scientist focused on Safety Post Training to advance post-training methods and interpretability techniques for safer frontier AI systems.

You will design post-training pipelines, study how training choices affect safety and alignment, and collaborate with policymakers, engineers, and researchers to translate findings into safety standards and evaluation benchmarks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Safety Scientist (Post-Training & Interpretability)
AI Safety Scientist (Post-Training & Interpretability)

United States Digital Space LLC • New York (NY), San Francisco (CA)

On-site
USD 216,000 - 270,000
Health, dental and vision coverage
Equity
Learning and development stipend
+1
AI Safety Post-Training Scientist
AI Safety Post-Training Scientist

Scale AI, Inc. • New York (NY)

On-site
USD 216,000 - 270,000
Comprehensive health, dental, and vision coverage
Retirement benefits
Learning and development stipend
+2
Research Scientist, Safety Post Training
Research Scientist, Safety Post Training

Scale • San Francisco (CA), New York (NY)

On-site
USD 216,000 - 270,000
Research Lead: AI Safety & Pre-Training at Scale
Research Lead: AI Safety & Pre-Training at Scale

FAR.AI • Berkeley (CA)

On-site
USD 180,000 - 240,000
Research Scientist, Safety Post Training San Francisco, CA Apply →
Research Scientist, Safety Post Training San Francisco, CA Apply →

Scale AI, Inc. • New York (NY)

On-site
USD 216,000 - 270,000
Comprehensive health, dental, and vision coverage
Retirement benefits
Learning and development stipend
+2
Research Scientist, Safety Post Training
Research Scientist, Safety Post Training

United States Digital Space LLC • New York (NY), San Francisco (CA)

On-site
USD 216,000 - 270,000
Health, dental and vision coverage
Equity
Learning and development stipend
+1
Research Lead: Pre-Training Safety & Safe AI
Research Lead: Pre-Training Safety & Safe AI

FAR.AI • United States

On-site
USD 160,000 - 230,000
AI Safety Training Researcher – National Security (Remote)
AI Safety Training Researcher – National Security (Remote)

OpenAI • San Francisco (CA)

On-site
USD 380,000 - 500,000
Equity
Remote work option
Research Scientist - AI Safety & Scalable ML
Research Scientist - AI Safety & Scalable ML

Center for AI Safety (CAIS) • San Francisco (CA)

On-site
USD 140,000 - 200,000
Health insurance for you and dependets
401K plan + 4% matching
Unlimited PTO
+2
AI Safety Researcher: Training & Robustness
AI Safety Researcher: Training & Robustness

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000