Remote AI Safeguards & Cyber Enforcement Analyst

Anthropic

San Francisco (CA)

Hybrid

USD 285,000 - 330,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity donations
Flexible hours
Vacation policy
Parental leave
Office space

Job summary

Anthropic is seeking a Safeguards Enforcement Analyst, Cyber Harm, to review content and apply enforcement actions across our products, focusing on detecting and mitigating cyber misuse of our AI systems. The role involves analyzing flagged activity related to cyberattacks, malware, and exploitation, with weekend escalations possible.

You will work with Engineering and Data Science teams, maintain high accuracy, and stay current on threat trends and policy gaps to inform enforcement decisions

Qualifications

  • Experience in cybersecurity, including offensive techniques, exploit development, malware analysis, or vulnerability research.
  • Experience performing content review, abuse investigations, or policy enforcement at volume.
  • Proficiency in SQL and/or Python for data analysis and threat detection.
  • Experience identifying emerging risks and communicating findings to cross-functional teams.
  • Experience working with generative AI products, including writing effective prompts for content review and enforcement.

Responsibilities

  • Review flagged content and accounts to make accurate, well-documented enforcement decisions in line with our usage policies.
  • Detect and mitigate potential misuse of AI systems to facilitate cyberattacks, malware creation, exploitation tooling, and related harmful cyber operations.
  • Triage and elevate novel, ambiguous, or high-severity cases to appropriate stakeholders.
  • Provide detailed feedback to policy design teams on gaps surfaced through enforcement scenarios.
  • Partner with Engineering and Data Science to surface detection model errors and quality signals to improve precision and recall.
  • Maintain high accuracy and consistency standards across review queues.
  • Keep up to date with emerging AI policy enforcement best practices and evolving cyber threat landscape to inform decisions.

Skills

SQL
Python
Cybersecurity fundamentals
Content review
Stakeholder communication
Generative AI

Education

Bachelor’s degree

Job description

Anthropic is seeking a Safeguards Enforcement Analyst, Cyber Harm, to review content and apply enforcement actions across our products, focusing on detecting and mitigating cyber misuse of our AI systems. The role involves analyzing flagged activity related to cyberattacks, malware, and exploitation, with weekend escalations possible.

You will work with Engineering and Data Science teams, maintain high accuracy, and stay current on threat trends and policy gaps to inform enforcement decisions

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cyber Harm Safeguards Analyst — Remote
Cyber Harm Safeguards Analyst — Remote

Appliedmethods • San Francisco (CA), Northern (KY)

Hybrid
USD 285,000 - 330,000
AI Safeguards Enforcement Specialist
AI Safeguards Enforcement Specialist

Anthropic • Washington

Hybrid
USD 285,000 - 330,000
AI Safeguards Enforcement Analyst
AI Safeguards Enforcement Analyst

Anthropic • New York (NY)

Hybrid
USD 285,000 - 330,000
Cyber Safeguards Enforcement Lead
Cyber Safeguards Enforcement Lead

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
AI Integrity & Safeguards Enforcement Analyst
AI Integrity & Safeguards Enforcement Analyst

Anthropic • San Francisco (CA)

Hybrid
USD 285,000 - 330,000
Safeguards Enforcement Analyst, Cyber Harm
Safeguards Enforcement Analyst, Cyber Harm

Anthropic • Washington

Hybrid
USD 285,000 - 330,000
Safeguards Enforcement Analyst, Cyber Harm
Safeguards Enforcement Analyst, Cyber Harm

Anthropic • New York (NY)

On-site
USD 285,000 - 330,000
Cyber Policy Analyst — AI Safeguards & Compliance
Cyber Policy Analyst — AI Safeguards & Compliance

Anthropic • San Francisco (CA)

On-site
USD 190,000 - 285,000
Safeguards Enforcement Analyst, Cyber Harm
Safeguards Enforcement Analyst, Cyber Harm

Doist • New York (NY), Washington, San Francisco (CA)

Hybrid
USD 285,000 - 330,000
Safeguards Enforcement Lead, Cyber Harms
Safeguards Enforcement Lead, Cyber Harms

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000