AI Safety & Agent Robustness Scientist

AISafety

New York, Northern (NY, KY)

Hybrid

USD 216,000 - 270,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental coverage
Vision coverage
Generous PTO

Job summary

Scale Labs is seeking a Research Scientist focused on agent robustness to advance safe and aligned AI systems. You will study agent capabilities, develop evaluation harnesses, and prototype mitigations for complex failure modes across interacting AI agents.

This full-time role offers opportunities to publish, collaborate across industry and academia, and contribute to safety protocols as frontier AI capabilities evolve.

Qualifications

  • Experience with post-training and RL techniques such as RLHF, DPO, GRPO, and similar approaches.
  • Track record of published research in machine learning, particularly in generative AI.
  • At least three years of experience addressing sophisticated ML problems, whether in a research setting or in product development.
  • Strong written and verbal communication skills to operate in a cross-functional team.

Responsibilities

  • Research the science of AI agent capabilities with a focus on safety, risk factors, and benchmarking methodologies.
  • Design and build harnesses to test AI agents’ tendency to take harmful actions when pressured or manipulated by the environment.
  • Design and build exploits and mitigations for new failure modes as AI agents gain capabilities like coding, web browsing, and computer use.
  • Characterize and design mitigations for potential failure modes or broader risks in systems with multiple interacting AI agents.

Skills

RL techniques
Agent scaffolding
Evaluation harnesses
Cross-functional collaboration
Publication experience

Tools

SWE-bench
WebArena
OSWorld
Inspect

Job description

Scale Labs is seeking a Research Scientist focused on agent robustness to advance safe and aligned AI systems. You will study agent capabilities, develop evaluation harnesses, and prototype mitigations for complex failure modes across interacting AI agents.

This full-time role offers opportunities to publish, collaborate across industry and academia, and contribute to safety protocols as frontier AI capabilities evolve.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Agent Safety & Robustness Research Scientist
AI Agent Safety & Robustness Research Scientist

United States Digital Space LLC • New York (NY), San Francisco (CA)

On-site
USD 216,000 - 270,000
Research Scientist, Agent Robustness
Research Scientist, Agent Robustness

Scale AI • San Francisco (CA)

On-site
USD 197,000 - 247,000
Comprehensive health coverage
Retirement benefits
Generous PTO
+1
Research Scientist, Agent Robustness
Research Scientist, Agent Robustness

AISafety • New York (NY), Northern (KY)

Hybrid
USD 216,000 - 270,000
Health insurance
Dental coverage
Vision coverage
+1
Research Scientist, Agent Robustness
Research Scientist, Agent Robustness

United States Digital Space LLC • New York (NY), San Francisco (CA)

On-site
USD 216,000 - 270,000
AI Agent Safety & Robustness Research Scientist
AI Agent Safety & Robustness Research Scientist

Scale AI • San Francisco (CA)

On-site
USD 197,000 - 247,000
Comprehensive health coverage
Retirement benefits
Generous PTO
+1
Safety & Systems Researcher for Autonomous AI Agents
Safety & Systems Researcher for Autonomous AI Agents

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Relocation assistance
Health insurance
AI Safety & Post-Training Scientist
AI Safety & Post-Training Scientist

Scale • San Francisco (CA), New York (NY)

On-site
USD 216,000 - 270,000
AI Agent Engineer — Scale Safe, Smart Agents
AI Agent Engineer — Scale Safe, Smart Agents

TRM Labs • United States

Hybrid
USD 200,000 - 275,000
Equity plan
AI Safety Researcher — RLHF & Adversarial Robustness
AI Safety Researcher — RLHF & Adversarial Robustness

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
AI Safety Training Researcher — Robustness Focus
AI Safety Training Researcher — Robustness Focus

OpenAI • United States

On-site
USD 380,000 - 500,000