Hybrid Software Engineer — AI Safety Evaluations

Anthropic

San Francisco (CA)

Hybrid

USD 320,000 - 485,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is seeking a skilled engineer to build evaluation infrastructure for detecting AI misuse. In this role, you'll craft experiments and datasets to enhance the reliability of Claude's safety systems.

Ideal candidates should have a strong Python background, experience with data pipelines, and a comprehensive understanding of LLM capabilities and their implications in safety mechanisms.

The role emphasizes trust in automated abuse detection methods, contributing to AI's beneficial development.

Qualifications

  • 6+ years of industry software engineering experience preferred.
  • Experience in trust and safety or abuse detection systems is a plus.
  • Strong communication skills and an understanding of AI system risks.

Responsibilities

  • Own the evaluation harness for an investigation system.
  • Construct high-quality eval datasets representing real-world misuse.
  • Measure agent performance and identify measurement gaps.
  • Productionize research for agent changes and model upgrades.
  • Build tools for policy experts to author evaluations.

Skills

Proficiency in Python
Experience in building data pipelines
Working understanding of LLMs
Strong data analysis skills
Ability to translate ambiguous problems

Education

Bachelor’s degree or equivalent

Tools

Large-scale data processing

Job description

Anthropic is seeking a skilled engineer to build evaluation infrastructure for detecting AI misuse. In this role, you'll craft experiments and datasets to enhance the reliability of Claude's safety systems.

Ideal candidates should have a strong Python background, experience with data pipelines, and a comprehensive understanding of LLM capabilities and their implications in safety mechanisms.

The role emphasizes trust in automated abuse detection methods, contributing to AI's beneficial development.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer — AI Safety Evaluation Systems
Staff Software Engineer — AI Safety Evaluation Systems

Menlo Ventures • New York (NY)

Hybrid
USD 320,000 - 485,000
ML Engineer - AI Safety & Evaluation Pipelines
ML Engineer - AI Safety & Evaluation Pipelines

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Generous vacation and parental leave
Flexible working hours
Lovely office space
+1
Software Engineer — AI Safeguards & Oversight
Software Engineer — AI Safeguards & Oversight

SignalAI • New York (NY)

Hybrid
USD 320,000 - 485,000
Staff+ Software Engineer, Safeguards Evals
Staff+ Software Engineer, Safeguards Evals

Menlo Ventures • New York (NY)

Hybrid
USD 320,000 - 485,000
Research Engineer, AI Evaluation & Metrics
Research Engineer, AI Evaluation & Metrics

Menlo Ventures • New York (NY)

On-site
USD 500,000 - 850,000
Staff Software Engineer, Safeguards Evals
Staff Software Engineer, Safeguards Evals

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Staff Software Engineer - AI Safety & Safeguards
Staff Software Engineer - AI Safety & Safeguards

Menlo Ventures • New York (NY)

Hybrid
USD 320,000 - 485,000
Equity donation matching
Flexible hours
Office space
+1
Staff Software Engineer, Claude Code
Staff Software Engineer, Claude Code

Anthropic • Seattle (WA)

On-site
USD 120,000 - 180,000
Software Engineer, Safeguards Evals San Francisco, CA | New York City, NY
Software Engineer, Safeguards Evals San Francisco, CA | New York City, NY

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Generous vacation and parental leave
Flexible working hours
Lovely office space
+1
Senior AI Security Engineer: Claude-Driven Apps
Senior AI Security Engineer: Claude-Driven Apps

SignalAI • New York (NY)

Hybrid
USD 320,000 - 485,000