Staff Software Engineer - AI Safety & Safeguards

Menlo Ventures

New York (NY)

Hybrid

USD 320,000 - 485,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity donation matching
Flexible hours
Office space
Vacation & parental leave

Job summary

Anthropic is hiring software engineers to build safety and oversight mechanisms for AI systems. As a member of the Safeguards team, you will monitor models, prevent misuse, and ensure user wellbeing with scalable defenses and automated enforcement actions.

You will contribute to detecting unwanted model behaviors, surface abuse patterns to research teams, and strengthen safeguards across the stack using Python and TypeScript.

Qualifications

  • Bachelor’s degree in Computer Science, Software Engineering or comparable experience.
  • Proficiency in Python and Typescript.
  • Ability to work across the stack.
  • Strong communication skills and ability to explain complex technical concepts to non‑technical stakeholders.

Responsibilities

  • Develop monitoring systems to detect unwanted behaviors from API partners and take automated enforcement actions, surfacing results on internal dashboards for analysts.
  • Build abuse detection mechanisms and infrastructure.
  • Surface abuse patterns to research teams to harden models during training.
  • Build robust and reliable multi-layered defenses for real‑time improvement of safety mechanisms that work at scale.

Skills

Python
Typescript
Cross-stack development
Communication

Education

Bachelor’s degree in Computer Science or Software Engineering

Job description

Anthropic is hiring software engineers to build safety and oversight mechanisms for AI systems. As a member of the Safeguards team, you will monitor models, prevent misuse, and ensure user wellbeing with scalable defenses and automated enforcement actions.

You will contribute to detecting unwanted model behaviors, surface abuse patterns to research teams, and strengthen safeguards across the stack using Python and TypeScript.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Data Safeguards & Governance
Staff Software Engineer, Data Safeguards & Governance

Visa Hunt • New York (NY), San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Staff Software Engineer, Safeguards Data Platforms
Staff Software Engineer, Safeguards Data Platforms

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 485,000
Competitive compensation
Benefits
Equity donation matching
+4
AI Safety & Oversight Engineer
AI Safety & Oversight Engineer

Anthropic • San Francisco (CA)

On-site
USD 320,000 - 485,000
Software Engineer, Safeguards
Software Engineer, Safeguards

Anthropic • San Francisco (CA)

On-site
USD 320,000 - 485,000
Staff Data Platform Engineer, Safeguards
Staff Data Platform Engineer, Safeguards

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
Staff+ Software Engineer, Safeguards
Staff+ Software Engineer, Safeguards

Menlo Ventures • New York (NY)

Hybrid
USD 320,000 - 485,000
Equity donation matching
Flexible hours
Office space
+1
Staff Data Platform Engineer, AI Safeguards & Governance
Staff Data Platform Engineer, AI Safeguards & Governance

Anthropic • New York (NY)

Hybrid
USD 320,000 - 485,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+2
Staff Software Engineer — AI Safety Evaluation Systems
Staff Software Engineer — AI Safety Evaluation Systems

Menlo Ventures • New York (NY)

Hybrid
USD 320,000 - 485,000
Product Manager — AI Safety & Safeguards
Product Manager — AI Safety & Safeguards

Anthropic • San Francisco (CA)

Hybrid
USD 305,000 - 385,000
AI Integrity & Safeguards Specialist
AI Integrity & Safeguards Specialist

Anthropic • United States

Hybrid
USD 285,000 - 330,000