Head of AI Safety & Adversarial Evaluation

Moonshot

Denver (CO)

On-site

USD 110,000 - 120,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

15 days paid vacation
Federal holidays
Healthcare package
Dental & Vision Insurance
Life & Disability Insurance
Employee Assistance Program
401k matching
Share options
Maternity/Paternity leave

Job summary

Moonshot is seeking a Head of AI Safety to lead the delivery, development, and growth of our AI Safety portfolio, combining violence prevention, safeguarding, and behavioral risk with evaluating AI system safety for frontier AI contexts.

You will collaborate with model, policy, trust & safety, product, research, and engineering teams, manage client relationships, lead staff and ensure ethical, compliant delivery while pursuing strategic partnerships with governments and regulators.

Qualifications

  • Experience in trust & safety, online harms, or violence prevention with applicability to AI systems.
  • Ability to translate complex safety concepts for technical and government audiences.
  • Experience designing evaluation frameworks or interventions for harm categories such as extremism or child safety.
  • Proven project leadership, team and budget management, and client-facing work.
  • Excellent written communication for government, foundation, or enterprise audiences.

Responsibilities

  • Lead and quality-assure applied AI safety work across harm categories using red teaming and adversarial evaluation.
  • Advise frontier AI companies on improving model safety across products, policies, and interventions.
  • Translate insights from subject-matter experts into practical guidance for model safety, policy, and engineering teams.
  • Set methodological approach for the portfolio and develop structured evaluation frameworks.
  • Oversee risk, resources, and partnerships, ensuring ethical and compliant delivery.

Skills

Leadership
AI Safety
Red Teaming
Adversarial Evaluation
Client Management
Policy & Government Engagement
Security Clearance
Travel flexibility

Job description

Moonshot is seeking a Head of AI Safety to lead the delivery, development, and growth of our AI Safety portfolio, combining violence prevention, safeguarding, and behavioral risk with evaluating AI system safety for frontier AI contexts.

You will collaborate with model, policy, trust & safety, product, research, and engineering teams, manage client relationships, lead staff and ensure ethical, compliant delivery while pursuing strategic partnerships with governments and regulators.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Chief AI Safety & Adversarial Evaluation Lead
Chief AI Safety & Adversarial Evaluation Lead

Moonshotteam • Washington, Denver (CO), Atlanta (GA)

Hybrid
USD 110,000 - 145,000
Healthcare package
Dental & Vision Insurance
Life & Disability Insurance
+1
Head of AI Safety & Risk Governance with Equity Options
Head of AI Safety & Risk Governance with Equity Options

Moonshot • Washington

On-site
USD 110,000 - 145,000
Private healthcare package
Dental & Vision Insurance
Life & Disability Insurance
+2
Director, AI Safety & Risk — Equity Options
Director, AI Safety & Risk — Equity Options

Moonshot • Washington

On-site
USD 110,000 - 120,000
15 days paid vacation
Private healthcare
Dental & Vision Insurance
+4
Head of AI Safety & Risk - Equity Options
Head of AI Safety & Risk - Equity Options

Moonshot • Atlanta (GA)

On-site
USD 110,000 - 120,000
15 days vacation leave
Federal holidays and additional leave
Private healthcare package
+6
Head of AI Safety
Head of AI Safety

Moonshot • Washington

On-site
USD 110,000 - 120,000
15 days paid vacation
Private healthcare
Dental & Vision Insurance
+4
Head of AI Safety
Head of AI Safety

Moonshot • Denver (CO)

On-site
USD 110,000 - 120,000
15 days paid vacation
Federal holidays
Healthcare package
+6
Head of AI Safety
Head of AI Safety

Moonshotteam • Washington, Denver (CO), Atlanta (GA)

Hybrid
USD 110,000 - 145,000
Healthcare package
Dental & Vision Insurance
Life & Disability Insurance
+1
Head of AI Safety
Head of AI Safety

Moonshot • Washington

On-site
USD 110,000 - 145,000
Private healthcare package
Dental & Vision Insurance
Life & Disability Insurance
+2
Head of AI Safety
Head of AI Safety

Moonshot • Atlanta (GA)

On-site
USD 110,000 - 120,000
15 days vacation leave
Federal holidays and additional leave
Private healthcare package
+6
Head of AI Safety Policy & Harm Mitigation
Head of AI Safety Policy & Harm Mitigation

Anthropic • San Francisco (CA)

Hybrid
USD 330,000 - 395,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+2