Policy Lead for Agentic AI Safety

United States Digital Space LLC

United States

On-site

USD 140,000 - 180,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

United States Digital Space LLC is seeking a role within its Safety Systems team to address real-world risks from model misalignment as models operate over extended horizons. You will investigate how misaligned behavior emerges across trajectories and translate insights into behavioral policies, evaluations, monitoring, and safeguards.

The position emphasizes empirical safety work, policy development, and collaboration with research, engineering, security, and product teams to balance safety

Qualifications

  • Background in AI agent safety, privacy, security, or adjacent fields.
  • Strong understanding of AI alignment and misaligned model behavior.
  • Ability to work directly with evaluation data and identify limitations, patterns, and opportunities for deeper investigation.

Responsibilities

  • Identify vulnerabilities that emerge as models interact with tools, data, and external systems, and translate them into model- and system-level safeguards.
  • Develop threat models and empirical frameworks for understanding harmful outcomes from misaligned behavior.
  • Build frameworks for understanding harmful outcomes arising from model misalignment.
  • Identify the underlying behaviors and system conditions that drive those outcomes.
  • Turn findings into policy frameworks, evaluation criteria, online measurement and safeguards.
  • Develop human data campaigns and gold sets to ground measurement and evaluation of emerging behaviors and risks.
  • Partner with research, engineering, security, and product teams to shape model and system safety, balancing safety, utility, and business risk.
  • Inform deployment decisions, system cards, safeguards reports, and the company’s broader approach to agentic safety.
  • Build monitoring approaches that detect regressions and emerging risks after deployment.

Job description

United States Digital Space LLC is seeking a role within its Safety Systems team to address real-world risks from model misalignment as models operate over extended horizons. You will investigate how misaligned behavior emerges across trajectories and translate insights into behavioral policies, evaluations, monitoring, and safeguards.

The position emphasizes empirical safety work, policy development, and collaboration with research, engineering, security, and product teams to balance safety

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Model Policy & Agentic Safety Lead
Model Policy & Agentic Safety Lead

OpenAI • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Relocation support
Hybrid work model
Agentic Safety Policy Lead for Frontier AI
Agentic Safety Policy Lead for Frontier AI

Showcify • United States

On-site
USD 180,000 - 240,000
Model Policy Manager, Agentic Safety
Model Policy Manager, Agentic Safety

United States Digital Space LLC • United States

Hybrid
USD 140,000 - 180,000
Model Policy Manager, Agentic Safety
Model Policy Manager, Agentic Safety

Showcify • United States

Hybrid
USD 180,000 - 240,000
Model Policy Manager, Agentic Safety
Model Policy Manager, Agentic Safety

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Relocation support
Hybrid work model
Safety Systems Engineer — Full-Stack Tools for Safe AI
Safety Systems Engineer — Full-Stack Tools for Safe AI

United States Digital Space LLC • United States

Remote
USD 120,000 - 190,000
Safety & AI Policy Projects Associate
Safety & AI Policy Projects Associate

Handshake • Seattle (WA)

Hybrid
USD 108,000 - 132,000
Equity
401(k) match
Parental leave
+4
Frontier AI Safety Researcher: Mitigations & Alignment
Frontier AI Safety Researcher: Mitigations & Alignment

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 280,000
Chief AI Safety & Adversarial Evaluation Lead
Chief AI Safety & Adversarial Evaluation Lead

Moonshotteam • Washington, Denver (CO), Atlanta (GA)

Hybrid
USD 110,000 - 145,000
Healthcare package
Dental & Vision Insurance
Life & Disability Insurance
+1
Researcher, Agent Safety, Training and Evaluations
Researcher, Agent Safety, Training and Evaluations

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Hybrid work model
Relocation assistance