Software Engineer, Safeguards

Anthropic

San Francisco (CA)

On-site

USD 320,000 - 485,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is looking for software engineers in San Francisco to help build safety and oversight mechanisms for AI systems. You will monitor models, prevent misuse, and ensure user well-being by developing systems for detecting unwanted model behaviors.

Candidates are expected to have a Bachelor’s degree in Computer Science or Software Engineering and proficiency in Python and Typescript. The annual compensation range is $320,000 - $485,000.

Qualifications

  • Bachelor’s degree in Computer Science, Software Engineering or comparable experience.
  • Proficiency in Python and Typescript.
  • Ability to work across the stack.
  • Strong communication skills and ability to explain complex technical concepts to non-technical stakeholders.

Responsibilities

  • Develop monitoring systems to detect unwanted behaviors from our API partners.
  • Build abuse detection mechanisms and infrastructure.
  • Surface abuse patterns to research teams for model hardening.
  • Build multi-layered defenses for real-time safety mechanism improvements.

Skills

Proficiency in Python
Proficiency in Typescript
Strong communication skills
Ability to work across the stack

Education

Bachelor's degree in Computer Science or Software Engineering

Job description

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the role

We are looking for software engineers to help build safety and oversight mechanisms for our AI systems. As a software engineer on the Safeguards team, you will work to monitor models, prevent misuse, and ensure user well-being. This role will focus on building systems to detect unwanted model behaviors and prevent disallowed use of models. You will apply your technical skills to uphold our principles of safety, transparency, and oversight while enforcing our terms of service and acceptable use policies.

Responsibilities
  • Develop monitoring systems to detect unwanted behaviors from our API partners and potentially take automated enforcement actions; surface these in internal dashboards to analysts for manual review
  • Build abuse detection mechanisms and infrastructure
  • Surface abuse patterns to our research teams to harden models at the training stage
  • Build robust and reliable multi-layered defenses for real-time improvement of safety mechanisms that work at scale
Qualifications (You may be a good fit if you)
  • Bachelor’s degree in Computer Science, Software Engineering or comparable experience
  • Proficiency in Python and Typescript
  • Ability to work across the stack
  • Strong communication skills and ability to explain complex technical concepts to non-technical stakeholders
Strong Candidates (may also have)
  • 8+ years of experience in a software engineering position
  • Experience with integrity, spam, fraud, or abuse detection and mitigation
  • Experience building trust and safety detection mechanisms and intervention for AI/ML systems
  • Experience with prompt engineering, jailbreak attacks, and other adversarial inputs
  • Experience working closely with operational teams to build custom internal tooling
Compensation

Annual compensation range for this role: $320,000 - $485,000 USD.

Equal Opportunity

As set forth in Anthropic’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff+ Software Engineer, Safeguards
Staff+ Software Engineer, Safeguards

Menlo Ventures • New York (NY)

Hybrid
USD 320,000 - 485,000
Equity donation matching
Flexible hours
Office space
+1
Staff+ Software Engineer, Safeguards
Staff+ Software Engineer, Safeguards

Menlo Ventures • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Staff+ Software Engineer, Safeguards Data
Staff+ Software Engineer, Safeguards Data

Visa Hunt • New York (NY), San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Staff+ Software Engineer, Safeguards Data
Staff+ Software Engineer, Safeguards Data

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 485,000
Competitive compensation
Benefits
Equity donation matching
+4
Machine Learning Infrastructure Engineer, Safeguards Research
Machine Learning Infrastructure Engineer, Safeguards Research

Jobzhr • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 500,000
ML/Research Engineer, Safeguards
ML/Research Engineer, Safeguards

Anthropic • New York (NY)

Hybrid
USD 350,000 - 500,000
Staff+ Software Engineer, Safeguards Data
Staff+ Software Engineer, Safeguards Data

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
Machine Learning Infrastructure Engineer, Safeguards Research Anthropic San Francisco, CA | New[...]
Machine Learning Infrastructure Engineer, Safeguards Research Anthropic San Francisco, CA | New[...]

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 500,000
Office in San Francisco
ML/Research Engineer, Safeguards
ML/Research Engineer, Safeguards

Anthropic • San Francisco (CA)

On-site
USD 350,000 - 500,000
Data Engineer, Safeguards
Data Engineer, Safeguards

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 405,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+2