Head of AI Safety Policy & Harm Mitigation

Anthropic

San Francisco (CA)

Hybrid

USD 330,000 - 395,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Equity donation matching
Generous vacation and parental leave
Flexible working hours
Office space in SF

Job summary

Anthropic is seeking a leader to direct the Safeguards organization, shaping policy, evaluations, and enforcement for Claude usage across multiple harm areas. You will mentor managers, drive strategy, and collaborate with product, research, and engineering to implement effective mitigations.

The role demands deep domain knowledge in AI safety and strong cross-team collaboration, ensuring policies scale with frontier models and align with company objectives.

Qualifications

  • Experience leading teams including managers or senior specialists in AI safety, product policy, or related field.
  • Deep familiarity with consumer harm areas (child safety, wellbeing, manipulation, election integrity).
  • Strong cross-team collaboration with distributed ownership decisions.
  • Understanding of frontier model development, training/fine-tuning cycles, and deployment contexts.
  • Ability to translate policy into enforceable, measurable mechanisms and communicate with diverse audiences.
  • Sound judgment in ambiguous, high-consequence decisions.

Responsibilities

  • Lead and grow managers and teams responsible for consumer harms portfolio (child safety, user well-being, manipulation, election integrity).
  • Coordinate policy decisions across portfolio and maintain clear documentation of decisions and ownership.
  • Set strategy for mitigations on top of models, in partnership with alignment teams and product surfaces.
  • Prioritize harms areas, justify tradeoffs to leadership.
  • Serve as escalation point for high-severity decisions and rapid risk responses.
  • Collaborate with engineering, data science, product, legal, and research across development cycles.

Skills

Team leadership
AI safety
Policy design
Cross-functional collaboration
Executive communication

Education

Bachelor's degree or higher

Job description

Anthropic is seeking a leader to direct the Safeguards organization, shaping policy, evaluations, and enforcement for Claude usage across multiple harm areas. You will mentor managers, drive strategy, and collaborate with product, research, and engineering to implement effective mitigations.

The role demands deep domain knowledge in AI safety and strong cross-team collaboration, ensuring policies scale with frontier models and align with company objectives.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Director of AI Safety Policy & Governance
Director of AI Safety Policy & Governance

Alex Loftus • San Francisco (CA), Northern (KY)

Hybrid
USD 330,000 - 395,000
Product Manager, AI Safety & Safeguards
Product Manager, AI Safety & Safeguards

Anthropic • New York (NY)

Hybrid
USD 385,000 - 460,000
Office space and amenities
Flexible working hours
Equity donation matching
+1
Product Manager, AI Safeguards
Product Manager, AI Safeguards

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Product Manager, AI Safety & Safeguards
Product Manager, AI Safety & Safeguards

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
AI Safety & Safeguards Product Manager
AI Safety & Safeguards Product Manager

Alex Loftus • San Francisco (CA), Northern (KY)

Hybrid
USD 305,000 - 385,000
Head of Policy Design, Societal Harms
Head of Policy Design, Societal Harms

Anthropic • San Francisco (CA)

Hybrid
USD 330,000 - 395,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+2
Product Manager, Safeguards (Account Integrity & Abuse)
Product Manager, Safeguards (Account Integrity & Abuse)

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Product Manager, Safeguards (Generalist)
Product Manager, Safeguards (Generalist)

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Cybersecurity Product Lead for AI Security
Cybersecurity Product Lead for AI Security

Anthropic • New York (NY)

Hybrid
USD 385,000 - 460,000
Cyber Safeguards Enforcement Lead
Cyber Safeguards Enforcement Lead

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000