AI Safety & Oversight Engineer

Anthropic

San Francisco (CA)

On-site

USD 320,000 - 485,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is looking for software engineers in San Francisco to help build safety and oversight mechanisms for AI systems. You will monitor models, prevent misuse, and ensure user well-being by developing systems for detecting unwanted model behaviors.

Candidates are expected to have a Bachelor’s degree in Computer Science or Software Engineering and proficiency in Python and Typescript. The annual compensation range is $320,000 - $485,000.

Qualifications

  • Bachelor’s degree in Computer Science, Software Engineering or comparable experience.
  • Proficiency in Python and Typescript.
  • Ability to work across the stack.
  • Strong communication skills and ability to explain complex technical concepts to non-technical stakeholders.

Responsibilities

  • Develop monitoring systems to detect unwanted behaviors from our API partners.
  • Build abuse detection mechanisms and infrastructure.
  • Surface abuse patterns to research teams for model hardening.
  • Build multi-layered defenses for real-time safety mechanism improvements.

Skills

Proficiency in Python
Proficiency in Typescript
Strong communication skills
Ability to work across the stack

Education

Bachelor's degree in Computer Science or Software Engineering

Job description

Anthropic is looking for software engineers in San Francisco to help build safety and oversight mechanisms for AI systems. You will monitor models, prevent misuse, and ensure user well-being by developing systems for detecting unwanted model behaviors.

Candidates are expected to have a Bachelor’s degree in Computer Science or Software Engineering and proficiency in Python and Typescript. The annual compensation range is $320,000 - $485,000.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer - AI Safety & Safeguards
Staff Software Engineer - AI Safety & Safeguards

Menlo Ventures • New York (NY)

Hybrid
USD 320,000 - 485,000
Equity donation matching
Flexible hours
Office space
+1
Software Engineer, Safeguards
Software Engineer, Safeguards

Anthropic • San Francisco (CA)

On-site
USD 320,000 - 485,000
Staff Software Engineer, AI Safeguards & Monitoring
Staff Software Engineer, AI Safeguards & Monitoring

Menlo Ventures • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Staff Software Engineer, Data Safeguards & Governance
Staff Software Engineer, Data Safeguards & Governance

Visa Hunt • New York (NY), San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Staff+ Software Engineer, Safeguards
Staff+ Software Engineer, Safeguards

Menlo Ventures • New York (NY)

Hybrid
USD 320,000 - 485,000
Equity donation matching
Flexible hours
Office space
+1
Software Engineer — AI Safeguards & Oversight
Software Engineer — AI Safeguards & Oversight

SignalAI • New York (NY)

Hybrid
USD 320,000 - 485,000
Staff+ Software Engineer, Safeguards
Staff+ Software Engineer, Safeguards

Menlo Ventures • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Staff Software Engineer, Safeguards Data Platforms
Staff Software Engineer, Safeguards Data Platforms

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 485,000
Competitive compensation
Benefits
Equity donation matching
+4
Staff+ Software Engineer, Safeguards Data
Staff+ Software Engineer, Safeguards Data

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 485,000
Competitive compensation
Benefits
Equity donation matching
+4
Staff+ Software Engineer, Safeguards Data
Staff+ Software Engineer, Safeguards Data

Visa Hunt • New York (NY), San Francisco (CA)

Hybrid
USD 320,000 - 485,000