Safeguards Enforcement Analyst, Safety Evaluations Remote-Friendly (Travel-Required) | San Fran[...]

Anthropic

San Francisco (CA)

Remote

USD 230,000 - 270,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Anthropic is seeking a Safeguards Enforcement Analyst in San Francisco to ensure our models meet safety standards. This role is remote-friendly but requires occasional travel. You will collaborate with various teams to manage evaluations and drive improvements.

The ideal candidate has experience in trust and safety and is comfortable in fast-paced, ambiguous environments. Strong program management skills and a willingness to expand technical tools are essential.

Qualifications

  • Experience in trust and safety, content operations, or policy enforcement.
  • Ability to thrive in ambiguous, fast-moving environments.
  • Experience building processes from scratch.

Responsibilities

  • Support model launch readiness through evaluations.
  • Partner with policy and domain experts throughout evaluations.
  • Manage evaluation outcomes and drive mitigations when needed.

Skills

Trust and safety
Program management
Process building
Technical toolkit expansion
Data tools proficiency

Education

Bachelor's degree or equivalent

Tools

SQL
Dashboards
Spreadsheets

Job description

Safeguards Enforcement Analyst, Safety Evaluations

Remote-Friendly (Travel-Required) | San Francisco, CA | Washington, DC; San Francisco, CA | New York City, NY

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the Role

Anthropic's Safeguards team is responsible for enforcing our policies, protecting users, and ensuring our platform is not misused. As a Safeguards Enforcement Analyst focused on Safety Evaluations, you will play a central role in ensuring our models meet safety and policy standards before and after launch. You will run and monitor evaluations, drive mitigations when issues surface, coordinate the creation of new evals, and help build the processes and documentation that allow the team to scale this work over time.

This role requires someone who is detail-oriented, comfortable navigating ambiguity, and capable of coordinating across teams to break new ground and drive work to completion. This work is deeply cross‑functional — you’ll partner closely with policy experts, Safeguards engineering teams, and many other stakeholders throughout the organization to ensure our evaluations are comprehensive and current, and that findings translate into meaningful improvements to model behavior.

Responsibilities
  • Support model launch readiness by running evaluations, monitoring and interpreting results, and surfacing regressions or unexpected behavior changes to relevant stakeholders
  • Partner closely with policy and domain experts throughout the evaluation lifecycle — from identifying risks and scoping the right evaluation approach, to coordinating creation of new evals and ensuring existing ones remain current with evolving policies, threat vectors, and model capabilities
  • Work with cross-functional stakeholders to help manage evaluation outcomes, including interpreting results and driving mitigations where needed
  • Think strategically about eval quality to build processes and eval paradigms that keep evaluations unsaturated, high‑signal, and insightful as models improve
  • Build out processes and frameworks for creating product‑specific evaluations as Anthropic’s product surface area expands
  • Help design and scope tooling improvements that accommodate evolving eval needs and expand self‑serve eval creation and iteration for non‑technical users
  • Write and maintain rigorous documentation for evaluation creation, execution, and interpretation as the team builds out eval tooling and processes
You may be a good fit if you
  • Have experience in trust and safety, content operations, policy enforcement, or a related operational role at a technology company
  • Thrive in ambiguous, fast‑moving environments — you’re energized rather than frustrated when the path forward isn’t clearly defined and you need to figure it out as you go
  • Have experience building processes, workflows, or programs from scratch (zero‑to‑one work), not just maintaining existing ones
  • Have strong program‑management instincts, naturally creating structure around complex, multi‑stakeholder efforts by tracking timelines, dependencies, and deliverables to keep work on track
  • Are eager to expand your technical toolkit, including adopting internal tools and AI‑assisted workflows (e.g., Claude Code) to accelerate your work
  • Can manage multiple concurrent workstreams across different domain areas without losing track of details — strong prioritization and context‑switching are essential when deadlines and priorities shift quickly
  • Are a strong generalist comfortable moving fluidly across different types of work and switching contexts throughout the day
  • Are comfortable making judgment calls with incomplete information and escalating appropriately when needed
  • Communicate clearly and concisely, both in writing and cross‑functionally
Strong candidates may also have
  • Experience operating under tight, high‑stakes timelines — such as product launch cycles, incident response, or regulatory deadlines — where information and priorities can shift with little notice
  • Experience coordinating across engineering, policy, and product teams to translate findings into concrete action
  • Experience building and maintaining SOPs, runbooks, and operational documentation in fast‑changing environments
  • Proficiency with data tools (SQL, dashboards, spreadsheets) sufficient to maintain and improve workflows
  • Comfort working with sensitive content areas as part of eval creation or enforcement review responsibilities

The annual compensation range for this role is $230,000 – $270,000 USD.

Logistics

Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience

Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience

Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position

Location‑based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.

Visa sponsorship: We do sponsor visas. We will make every reasonable effort to obtain a visa if an offer is made.

We encourage you to apply even if you do not believe you meet every single qualification. Please do not exclude yourself prematurely.

Equal Employment Opportunity

As set forth in Anthropic’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law. If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection. We request this information for the purpose of measuring the effectiveness of outreach and recruitment efforts.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Safeguards Enforcement Analyst, Safety Evaluations
Safeguards Enforcement Analyst, Safety Evaluations

Anthropic • New York (NY)

On-site
USD 230,000 - 270,000
Competitive compensation
Equity donation matching
Generous vacation
+3
Technical Program Manager, Reliability Engineering San Francisco, CA | New York City, NY | Seat[...]
Technical Program Manager, Reliability Engineering San Francisco, CA | New York City, NY | Seat[...]

Anthropic • New York (NY)

On-site
USD 290,000 - 365,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Safeguards Enforcement Analyst, Cyber Harm
Safeguards Enforcement Analyst, Cyber Harm

Anthropic • San Francisco (CA)

On-site
USD 285,000 - 330,000
Equity donations
Flexible hours
Vacation policy
+2
Technical Program Manager, Safeguards (Infrastructure & Evals)
Technical Program Manager, Safeguards (Infrastructure & Evals)

Anthropic • New York (NY)

On-site
USD 290,000 - 365,000
Competitive salary
Flexible working hours
Generous vacation and parental leave
+1
Safeguards Enforcement Analyst, Conventional Weapons
Safeguards Enforcement Analyst, Conventional Weapons

anthropic • New York (NY), San Francisco (CA), Washington

On-site
USD 245,000 - 330,000
Staff+ Site Reliability Engineer, Safeguards ML Infra
Staff+ Site Reliability Engineer, Safeguards ML Infra

Anthropic • Seattle (WA), New York (NY), San Francisco (CA)

On-site
USD 405,000 - 485,000
Staff+ Software Engineer, Safeguards
Staff+ Software Engineer, Safeguards

Menlo Ventures • San Francisco (CA)

On-site
USD 320,000 - 485,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Data Scientist, Safeguards New York City, NY; San Francisco, CA | New York City, NY; Seattle, WA
Data Scientist, Safeguards New York City, NY; San Francisco, CA | New York City, NY; Seattle, WA

Anthropic • New York (NY)

On-site
USD 275,000 - 370,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Safeguards Enforcement Analyst, Integrity & Authenticity
Safeguards Enforcement Analyst, Integrity & Authenticity

Anthropic • San Francisco (CA)

On-site
USD 285,000 - 330,000
Safeguards Enforcement Analyst, Access Controls & Identity
Safeguards Enforcement Analyst, Access Controls & Identity

Anthropic • San Francisco (CA)

Hybrid
USD 285,000 - 330,000