Model Policy

Neura Market

San Francisco (CA)

Hybrid

USD 180,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenAI seeks a Policy-focused engineer to shape model behavior in high-risk contexts. You will design and maintain policies across safety domains, translate risk into behavioral specifications, and develop scalable safeguards for deployment.

You’ll collaborate with research, engineering, product, and operations to turn risk insights into measurable policy. This role is based in San Francisco with a hybrid schedule and relocation support.

Qualifications

  • Experience building or applying policies, taxonomies, harm models, threat models, or risk frameworks for complex technical or societal systems.
  • Ability to translate risk and harm models into behavioral specifications and evaluation criteria.
  • Ability to work across research, engineering, product, and policy teams to operationalize policy.
  • Comfort using empirical evidence, including evaluations and deployment observations, to inform policy decisions.
  • Strong written and verbal communication about complex safety tradeoffs.

Responsibilities

  • Design and maintain model policies across safety-relevant domains, including dual-use, agentic, and emerging frontier-risk areas.
  • Translate risk and harm models into clear behavioral specifications, evaluation criteria, grading guidance, and system-level safeguards.
  • Define practical boundaries between beneficial uses of AI and assistance that could materially enable harm, exploitation, misuse, or unsafe outcomes.
  • Build policy artifacts that support model training, evaluation, and deployment. Partner with safety researchers, engineers, product teams, and other stakeholders to operationalize policy into scalable model behavior and measurable safeguards.
  • Use red-teaming results, deployment data, model failures, over-refusals, under-refusals, and ambiguous edge cases to improve policy and evaluation quality over time.
  • Identify emerging capability areas where frontier AI systems could create new safety challenges or lower barriers to harm.
  • Study real-world deployments to identify where model behavior succeeds, fails, or drifts from the intended safety posture.
  • Combine longer-horizon safety research with hands-on launch and deployment work.
  • Contribute to system cards, safety reports, policy documentation, launch reviews, and external communications on OpenAI's approach to model safety and risk mitigation.
  • Design and run human data campaigns, including gold set construction, labeling guidance, calibration, adjudication, and eval coverage analysis, to ensure policies can be reliably measured and improved.

Skills

Policy development
Risk modeling
Red-teaming
Cross-functional collaboration
Evaluation methods
System design thinking

Job description

About the Team

Our Safety Systems team is at the forefront of OpenAI's mission to build and deploy safe AGI, driving our commitment to AI safety and fostering a culture of trust and transparency.

Within Safety Systems, the Model Policy team aligns model behavior with desired human values and norms. We co-design policy with models and for models by driving rapid policy taxonomy iteration based on data and defining evaluation criteria for foundational models’ ability to reason about safety.

About the Role

If you have a specific expertise or speciality related to this work, please note it in your application via your resume, cover letter or application note.

Frontier AI systems are expanding what people can do across domains, creating both enormous opportunities and difficult safety questions: when should a model help, when should it refuse, and how do we make those boundaries clear enough to train, evaluate, and enforce?

In this role, you will help define how OpenAI’s models should behave in high-risk or high-ambiguity contexts, such as agentic systems, multimodal systems, user safety, privacy, and other emerging risk domains.

This is an ideal role for someone who can move across unfamiliar topics, reason from first principles, and turn ambiguity into practical model behavior. You will work closely with research, engineering, product, preparedness, and operations teams to build policies that are technically grounded, measurable, and responsive to real-world risk.

In this role, you will:
  • Design and maintain model policies across safety-relevant domains, including dual-use, agentic, and emerging frontier-risk areas.
  • Translate risk and harm models into clear behavioral specifications, evaluation criteria, grading guidance, and system-level safeguards.
  • Define practical boundaries between beneficial uses of AI and assistance that could materially enable harm, exploitation, misuse, or unsafe outcomes.
  • Build policy artifacts that support model training, evaluation, and deployment. Partner with safety researchers, engineers, product teams, and other stakeholders to operationalize policy into scalable model behavior and measurable safeguards.
  • Use red-teaming results, deployment data, model failures, over-refusals, under-refusals, and ambiguous edge cases to improve policy and evaluation quality over time.
  • Identify emerging capability areas where frontier AI systems could create new safety challenges or lower barriers to harm.
  • Study real-world deployments to identify where model behavior succeeds, fails, or drifts from the intended safety posture.
  • Combine longer-horizon safety research with hands-on launch and deployment work.
  • Contribute to system cards, safety reports, policy documentation, launch reviews, and external communications on OpenAI's approach to model safety and risk mitigation.
  • Design and run human data campaigns, including gold set construction, labeling guidance, calibration, adjudication, and eval coverage analysis, to ensure policies can be reliably measured and improved.
You might thrive in this role if you:
  • Have strong judgment about how advanced AI systems may affect real-world risk, especially in ambiguous, fast-moving, or high-impact areas.
  • Have experience building or applying policies, taxonomies, harm models, threat models, or risk frameworks for complex technical, social, or adversarial systems.
  • Can move across domains without needing to be the deepest subject-matter expert in every area, while knowing when to seek expert input.
  • Can turn fuzzy questions into structured policy frameworks, evaluation criteria, operational guidance, and enforceable model behavior.
  • Are comfortable using empirical evidence, including evaluations, red-teaming results, deployment observations, and model failure modes, to inform policy decisions.
  • Think in systems across policy, data, graders, classifiers, training, deployment safeguards, measurement, monitoring, and escalation workflows.
  • Have technical judgment about what model behavior can realistically be trained, measured, evaluated, and enforced at scale.
  • Work well across research, engineering, product, policy, domain experts, and operational teams.
  • Write clearly about complex tradeoffs where safety, user value, and implementation constraints all matter.
  • Take a pragmatic approach to safety, focused on reducing real-world risk while preserving legitimate, beneficial, and socially valuable uses of AI.
  • Enjoy fast-paced, collaborative research environments where priorities shift as models, evidence, and risks change.
  • Stay grounded in implementation details, empirical results, and what can actually be trained or measured.
Workplace & Location

This role is based in our San Francisco office. We do encourage you to apply even if you prefer a different work location as factors may change over time.

We offer relocation support to new employees, and we use a hybrid model: three days in the office per week with optional work from home on Thursdays and Fridays.

Our open-plan offices have height-adjustable desks, conference rooms, phone booths, well-stocked kitchens full of snacks and drinks, three in-house prepared meals daily, a private outdoor space for working in the sun or socializing, nap rooms, private bike storage, and more.

Equal Employment Opportunity Statement

We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Model Policy
Model Policy

Slope • San Francisco (CA)

On-site
USD 207,000 - 295,000
Medical, dental, and vision insurance
401(k) retirement plan with employer match
Paid parental leave
+3
Model Policy, Frontier Cyber Risk
Model Policy, Frontier Cyber Risk

OpenAI • Los Angeles (CA)

Hybrid
USD 207,000 - 295,000
Model Policy, Frontier Cyber Risk
Model Policy, Frontier Cyber Risk

Slope • San Francisco (CA)

On-site
USD 207,000 - 295,000
Medical, dental, and vision insurance
401(k) retirement plan with employer match
Paid parental leave
+4
Model Policy, Frontier Cyber Risk
Model Policy, Frontier Cyber Risk

OpenAI • San Francisco (CA)

Hybrid
USD 207,000 - 295,000
Model Policy (Rodrigo)
Model Policy (Rodrigo)

Triwill Group • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation assistance
Hybrid work model
Fullstack Engineer, Safety Engineering
Fullstack Engineer, Safety Engineering

OpenAI • San Francisco (CA)

On-site
USD 210,000 - 325,000
Model Policy Manager
Model Policy Manager

Slope • San Francisco (CA)

Hybrid
USD 180,000 - 280,000
Relocation assistance
Fullstack Engineer, Safety Engineering OpenAI San Francisco
Fullstack Engineer, Safety Engineering OpenAI San Francisco

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Researcher, Trustworthy AI
Researcher, Trustworthy AI

OpenAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Relocation assistance
Researcher, Safety Oversight
Researcher, Safety Oversight

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 150,000