Staff Software Engineer, AI Safety & Abuse Detection

EngineersOfAI

San Francisco, Northern (CA, KY)

Hybrid

USD 320,000 - 485,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Anthropic in San Francisco is seeking a software engineer for the Safeguards team to build safety and oversight mechanisms for our AI systems. You will monitor models, prevent misuse, and ensure user wellbeing across the stack.

This role focuses on detecting unwanted model behaviors, enforcing policies, and surfacing abuse patterns to researchers for hardening models at the training stage. You will contribute to scalable defenses and real-time safety improvements.

Qualifications

  • Bachelor's degree in Computer Science, Software Engineering or comparable experience.
  • Proficiency in Python and Typescript.
  • Ability to work across the stack.
  • Strong communication skills and ability to explain complex technical concepts to non-technical stakeholders.

Responsibilities

  • Develop monitoring systems to detect unwanted behaviors from API partners and surface results in dashboards for analysts.
  • Build abuse detection mechanisms and infrastructure.
  • Surface abuse patterns to research teams to harden models at the training stage.
  • Build multi-layered defenses for real-time safety improvements at scale.

Skills

Python
Typescript
Cross-stack development
Communication

Education

Bachelor's degree in Computer Science or related field

Job description

Anthropic in San Francisco is seeking a software engineer for the Safeguards team to build safety and oversight mechanisms for our AI systems. You will monitor models, prevent misuse, and ensure user wellbeing across the stack.

This role focuses on detecting unwanted model behaviors, enforcing policies, and surfacing abuse patterns to researchers for hardening models at the training stage. You will contribute to scalable defenses and real-time safety improvements.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer, AI Safety & Safeguards
Staff Software Engineer, AI Safety & Safeguards

Jobs in JS • Seattle (WA)

Hybrid
USD 320,000 - 485,000
Staff Software Engineer - AI Safety & Safeguards
Staff Software Engineer - AI Safety & Safeguards

Menlo Ventures • New York (NY)

Hybrid
USD 320,000 - 485,000
Equity donation matching
Flexible hours
Office space
+1
Staff Software Engineer, AI Safety & Abuse Detection
Staff Software Engineer, AI Safety & Abuse Detection

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Staff Software Engineer, Distributed Systems for Safe AI
Staff Software Engineer, Distributed Systems for Safe AI

Anthropic • San Francisco (CA)

On-site
USD 170,000 - 250,000
Staff Software Engineer — AI Safety & Distributed Systems
Staff Software Engineer — AI Safety & Distributed Systems

Anthropic • New York (NY)

Hybrid
USD 320,000 - 485,000
Staff Software Engineer, AI Safety & Safeguards
Staff Software Engineer, AI Safety & Safeguards

Anthropic Limited • New York (NY), Northern (KY)

Hybrid
USD 320,000 - 485,000
AI Safety & Oversight Engineer
AI Safety & Oversight Engineer

Anthropic • San Francisco (CA)

On-site
USD 320,000 - 485,000
Senior AI Safeguards Engineer — Distributed Systems
Senior AI Safeguards Engineer — Distributed Systems

Anthropic • San Francisco (CA)

On-site
USD 190,000 - 270,000
Health insurance
Parental leave
Flexible PTO
+7
Software Engineer, Safeguards
Software Engineer, Safeguards

Anthropic • San Francisco (CA)

On-site
USD 320,000 - 485,000
Staff+ Software Engineer, Safeguards
Staff+ Software Engineer, Safeguards

Menlo Ventures • New York (NY)

On-site
USD 320,000 - 485,000
Equity donation matching
Flexible hours
Office space
+1