AI Safeguards Enforcement Architect

Anthropic

New York (NY)

Hybrid

USD 285,000 - 330,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Anthropic is seeking a Safeguards Analyst focused on Integrity & Authenticity to build enforcement workflows for preventing misuse of AI systems in coordinated inauthentic behavior, election manipulation, and targeting individuals.

You will work across harm areas including AI-enabled influence operations and privacy-related surveillance, shaping policy enforcement for safe user interaction with our products.

Qualifications

  • Experience in trust & safety, policy enforcement, or threat intelligence with focus on influence operations, disinformation, or privacy harms.
  • Experience standing up and scaling policy enforcement or content review workflows.
  • Proficiency in SQL and/or other data analysis tools to draw insights from large datasets.
  • Experience identifying emerging risks and communicating findings to cross-functional teams including Product, Policy, Engineering, and Legal.
  • Experience working with generative AI products, including writing effective prompts for content review and enforcement.
  • Understanding of the challenges involved in implementing product policies at scale, including in the content moderation space.

Responsibilities

  • Design and architect automated enforcement systems and review workflows that scale effectively while maintaining high accuracy.
  • Partner with Engineering and Data Science teams to optimize detection models for policy violations and automated enforcement systems.
  • Review flagged content to drive enforcement and policy improvements.
  • Enforce usage policies with a focus on detecting and mitigating AI-enabled influence operations, coordinated inauthentic behavior, election interference, and targeting, tracking, or surveillance of individuals and groups.
  • Support the Safeguards policy design team by providing detailed feedback on policy gaps based on real enforcement scenarios.
  • Keep up to date with emerging AI policy enforcement best practices, evolving threat actor tactics, and the regulatory landscape around elections, privacy, and surveillance, using these to inform our decision-making and workflows.

Skills

Trust & safety
Policy enforcement
Threat intelligence
Data analysis
Generative AI
AI policy at scale

Education

Bachelor’s degree or equivalent

Tools

Python

Job description

Anthropic is seeking a Safeguards Analyst focused on Integrity & Authenticity to build enforcement workflows for preventing misuse of AI systems in coordinated inauthentic behavior, election manipulation, and targeting individuals.

You will work across harm areas including AI-enabled influence operations and privacy-related surveillance, shaping policy enforcement for safe user interaction with our products.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Safeguards & Integrity Enforcement Analyst
AI Safeguards & Integrity Enforcement Analyst

Anthropic • Washington

Hybrid
USD 285,000 - 330,000
AI Integrity & Safeguards Enforcement Analyst
AI Integrity & Safeguards Enforcement Analyst

Anthropic • San Francisco (CA)

Hybrid
USD 285,000 - 330,000
AI Safeguards Enforcement Analyst
AI Safeguards Enforcement Analyst

Anthropic • New York (NY)

Hybrid
USD 285,000 - 330,000
AI Safeguards Enforcement Specialist
AI Safeguards Enforcement Specialist

Anthropic • Washington

Hybrid
USD 285,000 - 330,000
Remote AI Safeguards & Cyber Enforcement Analyst
Remote AI Safeguards & Cyber Enforcement Analyst

Anthropic • San Francisco (CA)

Hybrid
USD 285,000 - 330,000
Equity donations
Flexible hours
Vacation policy
+2
AI Safeguards Specialist — Contract (Remote Possible)
AI Safeguards Specialist — Contract (Remote Possible)

Employer.com • San Francisco (CA)

On-site
USD 165,000 - 235,000
AI Safeguards Analyst – Violence & Extremism
AI Safeguards Analyst – Violence & Extremism

Anthropic • San Francisco (CA)

On-site
USD 285,000 - 330,000
Identity & Access Safeguards Analyst
Identity & Access Safeguards Analyst

Anthropic • New York (NY)

Hybrid
USD 285,000 - 330,000
Cyber Safeguards Enforcement Lead
Cyber Safeguards Enforcement Lead

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Cyber Harm Safeguards Analyst — Remote
Cyber Harm Safeguards Analyst — Remote

Appliedmethods • San Francisco (CA), Northern (KY)

Hybrid
USD 285,000 - 330,000