Staff Software Engineer — AI Safety Evaluation Systems

Menlo Ventures

New York (NY)

Hybrid

USD 320,000 - 485,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is seeking an experienced engineer to build evaluation infrastructure for its safety and abuse-detection systems. You will design experiments, build datasets from real-world traffic, and ship evaluation methods into pipelines that gate changes to Claude's safety stack.

You will work at the intersection of ML research and engineering, applying tests across harm areas, measuring detection and investigation quality, and contributing to robust, trustworthy AI systems.

Qualifications

  • Proficiency in Python and data pipelines experience.
  • Experience with LLMs and agentic systems, including multi-step reasoning.
  • Strong data analysis skills and the ability to draw reliable insights from large datasets.
  • Ability to move fluidly between research prototypes and production-grade code.
  • Translate ambiguous problems into concrete, testable experiments.

Responsibilities

  • Build and own the evaluation harness for an agentic investigation system, defining metrics and tests.
  • Construct datasets representing real-world misuse across harm areas.
  • Measure agent performance end-to-end and drive improvements in hard harm areas.
  • Analyze coverage to close measurement gaps and keep evals high-signal.
  • Productionize research into regression and release pipelines for every agent change.
  • Build tooling for policy experts to author, run, and iterate evaluations without heavy engineering support.
  • Construct RL environments to improve Claude’s safety investigation capabilities.

Skills

Python
Data analysis
LLMs & agentic systems
Production-quality code
Experiment design
Cross-stack development

Education

Bachelor's degree

Tools

LLMs
Prompt engineering
Distributed systems
Data processing frameworks

Job description

Anthropic is seeking an experienced engineer to build evaluation infrastructure for its safety and abuse-detection systems. You will design experiments, build datasets from real-world traffic, and ship evaluation methods into pipelines that gate changes to Claude's safety stack.

You will work at the intersection of ML research and engineering, applying tests across harm areas, measuring detection and investigation quality, and contributing to robust, trustworthy AI systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hybrid Software Engineer — AI Safety Evaluations
Hybrid Software Engineer — AI Safety Evaluations

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
ML Engineer - AI Safety & Evaluation Pipelines
ML Engineer - AI Safety & Evaluation Pipelines

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Generous vacation and parental leave
Flexible working hours
Lovely office space
+1
Staff Software Engineer - AI Safety & Safeguards
Staff Software Engineer - AI Safety & Safeguards

Menlo Ventures • New York (NY)

Hybrid
USD 320,000 - 485,000
Equity donation matching
Flexible hours
Office space
+1
Research Engineer, AI Evaluation & Metrics
Research Engineer, AI Evaluation & Metrics

Menlo Ventures • New York (NY)

On-site
USD 500,000 - 850,000
Staff+ Software Engineer, Safeguards Evals
Staff+ Software Engineer, Safeguards Evals

Menlo Ventures • New York (NY)

Hybrid
USD 320,000 - 485,000
Staff Software Engineer, Safeguards Evals
Staff Software Engineer, Safeguards Evals

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Software Engineer, Safeguards Evals San Francisco, CA | New York City, NY
Software Engineer, Safeguards Evals San Francisco, CA | New York City, NY

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Generous vacation and parental leave
Flexible working hours
Lovely office space
+1
Staff Software Engineer, Safeguards Data Platforms
Staff Software Engineer, Safeguards Data Platforms

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 485,000
Competitive compensation
Benefits
Equity donation matching
+4
Staff Software Engineer, Data Safeguards & Governance
Staff Software Engineer, Data Safeguards & Governance

Visa Hunt • New York (NY), San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Staff Software Engineer, Claude Code
Staff Software Engineer, Claude Code

Anthropic • Seattle (WA)

On-site
USD 120,000 - 180,000