AI Safety Scientist (Post-Training & Interpretability)

United States Digital Space LLC

New York, San Francisco (NY, CA)

On-site

USD 216,000 - 270,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental and vision coverage
Equity
Learning and development stipend
Generous PTO

Job summary

Scale Labs in New York is seeking a Research Scientist focused on Safety Post-Training to develop methods and interpretability techniques that make frontier AI systems safer and better understood by researchers and policymakers.

You will design post-training pipelines, develop interpretability-informed evaluations, and collaborate with policymakers, engineers, and researchers to translate findings into tangible safety standards and benchmarks.

Qualifications

  • Commitment to safe, secure and trustworthy AI deployments as frontier capabilities advance.
  • Experience with post-training and RL techniques such as RLHF, DPO, GRPO, and similar approaches.
  • Track record of published research in machine learning, particularly in generative AI.
  • At least three years of experience addressing sophisticated ML problems, whether in a research setting or in product development.
  • Strong written and verbal communication skills to operate in a cross-functional team.

Responsibilities

  • Design and run post-training pipelines to study how training choices affect model safety, robustness, and alignment properties.
  • Develop interpretability-informed evaluations that reveal how and why models produce unsafe or undesirable behaviors, guiding mitigations.
  • Collaborate with policymakers, engineers, and researchers to translate post-training and interpretability findings into actionable safety standards, evaluation benchmarks, and best practices.

Skills

Post-training methods
Interpretability techniques
Policy research
RLHF
DPO
GRPO
Written and verbal communication

Job description

Scale Labs in New York is seeking a Research Scientist focused on Safety Post-Training to develop methods and interpretability techniques that make frontier AI systems safer and better understood by researchers and policymakers.

You will design post-training pipelines, develop interpretability-informed evaluations, and collaborate with policymakers, engineers, and researchers to translate findings into tangible safety standards and benchmarks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Safety & Post-Training Scientist
AI Safety & Post-Training Scientist

Scale • San Francisco (CA), New York (NY)

On-site
USD 216,000 - 270,000
AI Safety Post-Training Scientist
AI Safety Post-Training Scientist

Scale AI, Inc. • New York (NY)

On-site
USD 216,000 - 270,000
Comprehensive health, dental, and vision coverage
Retirement benefits
Learning and development stipend
+2
AI Safety Research Scientist — Post-Training
AI Safety Research Scientist — Post-Training

Scale AI, Inc. • New York (NY)

On-site
USD 216,000 - 270,000
Equity-based compensation
Comprehensive health, dental, vision
Retirement benefits
+3
Safety Post-Training AI Research Scientist
Safety Post-Training AI Research Scientist

Scale AI, Inc. • San Francisco (CA)

On-site
USD 216,000 - 270,000
Equity
Health insurance
Dental insurance
+5
Research Scientist, Safety Post Training
Research Scientist, Safety Post Training

Scale AI, Inc. • San Francisco (CA)

On-site
USD 216,000 - 270,000
Equity
Health insurance
Dental insurance
+5
Research Scientist, Safety Post Training
Research Scientist, Safety Post Training

Scale AI, Inc. • New York (NY)

On-site
USD 216,000 - 270,000
Equity-based compensation
Comprehensive health, dental, vision
Retirement benefits
+3
Research Scientist, Safety Post Training
Research Scientist, Safety Post Training

Scale • San Francisco (CA), New York (NY)

On-site
USD 216,000 - 270,000
Research Scientist, Safety Post Training San Francisco, CA Apply →
Research Scientist, Safety Post Training San Francisco, CA Apply →

Scale AI, Inc. • New York (NY)

On-site
USD 216,000 - 270,000
Comprehensive health, dental, and vision coverage
Retirement benefits
Learning and development stipend
+2
Research Scientist, Safety Post Training
Research Scientist, Safety Post Training

United States Digital Space LLC • New York (NY), San Francisco (CA)

On-site
USD 216,000 - 270,000
Health, dental and vision coverage
Equity
Learning and development stipend
+1
AI Safety Controls & Monitoring Research Scientist
AI Safety Controls & Monitoring Research Scientist

Scale AI, Inc. • San Francisco (CA)

On-site
USD 216,000 - 270,000
Equity
Health, dental, vision coverage
Retirement benefits
+3