Model Policy & Agentic Safety Lead

OpenAI

San Francisco (CA)

Hybrid

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Relocation support
Hybrid work model

Job summary

OpenAI is seeking a Safety Systems role focused on policy and model misalignment as AI models become more capable and autonomous. You will investigate real-world harms, translate insights into behavioral policies, and develop safeguards.

The role is based in San Francisco with relocation support and a hybrid model: three days in the office per week with optional work from home on Thursdays and Fridays. You will work with research, engineering, security, product, and policy teams to shape model

Qualifications

  • Background in AI agent safety, privacy, security, or adjacent fields.
  • Strong interest in AI alignment and understanding of misaligned model behavior.
  • Sufficient technical fluency to work with evaluation data and training data.
  • Hands-on experience with model data and evaluation results.

Responsibilities

  • Identify vulnerabilities as models interact with tools, data, and external systems.
  • Develop threat models and empirical frameworks for harmful outcomes from misalignment.
  • Build frameworks to understand harmful outcomes from misalignment.
  • Identify underlying behaviors and system conditions driving those outcomes.
  • Turn findings into policy frameworks, evaluation criteria, and safeguards.
  • Develop human data campaigns and gold sets for measurement and evaluation.
  • Collaborate with research, engineering, security, and product teams to balance safety, utility, and risk.
  • Inform deployment decisions and OpenAI's safety approach.
  • Build monitoring approaches to detect regressions and emerging risks after deployment.

Skills

AI safety
Model evaluation
Policy development
Threat modeling
Data analysis

Job description

OpenAI is seeking a Safety Systems role focused on policy and model misalignment as AI models become more capable and autonomous. You will investigate real-world harms, translate insights into behavioral policies, and develop safeguards.

The role is based in San Francisco with relocation support and a hybrid model: three days in the office per week with optional work from home on Thursdays and Fridays. You will work with research, engineering, security, product, and policy teams to shape model

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Model Policy Manager, Agentic Safety
Model Policy Manager, Agentic Safety

OpenAI • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Relocation support
Hybrid work model
Frontier-Model Safety Researcher & Evaluations
Frontier-Model Safety Researcher & Evaluations

OpenAI • San Francisco (CA)

Hybrid
USD 380,000 - 500,000
Relocation assistance
Hybrid work model
AI Safety & Oversight Systems Researcher
AI Safety & Oversight Systems Researcher

OpenAI • San Francisco (CA)

Hybrid
USD 195,000 - 230,000
Relocation assistance
Hybrid work model
Adversarial Model Research PM: Safety & Evaluation Lead
Adversarial Model Research PM: Safety & Evaluation Lead

OpenAI • San Francisco (CA)

Hybrid
USD 239,000 - 328,000
Relocation assistance
Hybrid work model
Technical Program Manager, Safety & Model Evaluation (Hybrid)
Technical Program Manager, Safety & Model Evaluation (Hybrid)

OpenAI • San Francisco (CA)

Hybrid
USD 207,000 - 285,000
Relocation assistance
Hybrid work model
AI Safety Policy Lead - Biosecurity
AI Safety Policy Lead - Biosecurity

Slope • San Francisco (CA)

Hybrid
USD 180,000 - 280,000
Relocation assistance
Model Policy Manager
Model Policy Manager

Slope • San Francisco (CA)

Hybrid
USD 180,000 - 280,000
Relocation assistance
Researcher, Agent Safety, Training and Evaluations
Researcher, Agent Safety, Training and Evaluations

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Hybrid work model
Relocation assistance
Researcher, Agent Safety, Training and Evaluations
Researcher, Agent Safety, Training and Evaluations

OpenAI • San Francisco (CA)

Hybrid
USD 380,000 - 500,000
Relocation assistance
Hybrid work model
Researcher, Agent Safety, Training and Evaluations
Researcher, Agent Safety, Training and Evaluations

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 190,000 - 230,000
Relocation assistance