AI Safeguards & Integrity Enforcement Analyst

Anthropic

Washington (District of Columbia)

Hybrid

USD 285,000 - 330,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is seeking a Safeguards Analyst focused on Integrity & Authenticity to design automated enforcement workflows and review processes for detecting and mitigating misuse of AI systems.

You will partner with Engineering and Data Science to optimize detection models, review flagged content, and enforce policies addressing coordinated inauthentic behavior, election interference, and privacy harms.

Qualifications

  • Experience in trust & safety, policy enforcement, threat intelligence or related field focusing on influence operations, disinformation, or privacy harms.
  • Experience standing up and scaling policy enforcement or content review workflows.
  • Proficiency in SQL and/or other data analysis tools to draw insights from large datasets.
  • Experience identifying emerging risks and communicating findings to cross‑functional teams.
  • Experience with generative AI products and writing prompts for content review.
  • Understanding challenges of implementing product policies at scale in content moderation.

Responsibilities

  • Design and architect automated enforcement systems and review workflows at scale.
  • Partner with Engineering and Data Science to optimize detection models for policy violations.
  • Review flagged content to drive enforcement and policy improvements.
  • Enforce usage policies to mitigate AI‑enabled inauthentic behavior and election interference.
  • Support Safeguards policy design with feedback from real enforcement scenarios.
  • Stay updated on AI policy enforcement best practices and regulatory developments.

Skills

Trust & Safety
Policy enforcement
Threat intelligence
AI product experience
Data analysis

Education

Bachelor's degree

Tools

SQL
Python
OSINT Tools

Job description

Anthropic is seeking a Safeguards Analyst focused on Integrity & Authenticity to design automated enforcement workflows and review processes for detecting and mitigating misuse of AI systems.

You will partner with Engineering and Data Science to optimize detection models, review flagged content, and enforce policies addressing coordinated inauthentic behavior, election interference, and privacy harms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Safeguards Enforcement Architect
AI Safeguards Enforcement Architect

Anthropic • New York (NY)

Hybrid
USD 285,000 - 330,000
AI Integrity & Safeguards Enforcement Analyst
AI Integrity & Safeguards Enforcement Analyst

Anthropic • San Francisco (CA)

Hybrid
USD 285,000 - 330,000
AI Safeguards & Integrity Enforcement Analyst
AI Safeguards & Integrity Enforcement Analyst

United States Digital Space LLC • United States

Hybrid
USD 285,000 - 330,000
AI Safeguards Enforcement Specialist
AI Safeguards Enforcement Specialist

Anthropic • Washington

Hybrid
USD 285,000 - 330,000
AI Safeguards Enforcement Analyst
AI Safeguards Enforcement Analyst

Anthropic • New York (NY)

Hybrid
USD 285,000 - 330,000
Remote AI Safeguards & Cyber Enforcement Analyst
Remote AI Safeguards & Cyber Enforcement Analyst

Anthropic • San Francisco (CA)

Hybrid
USD 285,000 - 330,000
Equity donations
Flexible hours
Vacation policy
+2
Cyber Harm Safeguards Analyst — Remote
Cyber Harm Safeguards Analyst — Remote

Appliedmethods • San Francisco (CA), Northern (KY)

Hybrid
USD 285,000 - 330,000
AI Safeguards Analyst – Violence & Extremism
AI Safeguards Analyst – Violence & Extremism

Anthropic • San Francisco (CA)

On-site
USD 285,000 - 330,000
Identity & Access Safeguards Analyst
Identity & Access Safeguards Analyst

Anthropic • San Francisco (CA)

Hybrid
USD 285,000 - 330,000
Equity donation matching
Generous vacation
Parental leave
+2
Identity & Access Safeguards Analyst
Identity & Access Safeguards Analyst

Anthropic • New York (NY)

Hybrid
USD 285,000 - 330,000