Cybersecurity Evaluations Engineer (AI Safety & Robustness)

Anthropic

United States

Remote

USD 150,000 - 230,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Anthropic is hiring Cyber Evaluations Engineers to build and run evaluations that measure cyber-relevant capabilities and safeguard robustness in our models. You will design new evals, run per-release robustness testing, and analyze data on jailbreaks and prompt bypasses to understand safeguards performance.

You will design detection probes and shape the layered abuse-detection architecture with the policy team, partnering with engineering to translate findings into improvements across models.

Qualifications

  • Experience building or running evaluations, benchmarks, or test suites for software or ML systems.
  • Hands-on cybersecurity experience (e.g., CTF participation, vulnerability research, exploit development, or security research).
  • Proficiency in Python.
  • Strong ability to communicate evaluation results with cross-functional stakeholders.

Responsibilities

  • Design and run evaluations to assess cyber-relevant risk in new models.
  • Execute per-release safeguard-robustness testing before launches.
  • Analyze results and clearly communicate findings to stakeholders.
  • Design probes for cyber misuse and detection architecture with policy team.
  • Build and maintain tooling to run and score evaluations.
  • Collaborate with policy and engineering to translate findings into improvements.

Skills

Python
Cybersecurity experience
Communication

Job description

Anthropic is hiring Cyber Evaluations Engineers to build and run evaluations that measure cyber-relevant capabilities and safeguard robustness in our models. You will design new evals, run per-release robustness testing, and analyze data on jailbreaks and prompt bypasses to understand safeguards performance.

You will design detection probes and shape the layered abuse-detection architecture with the policy team, partnering with engineering to translate findings into improvements across models.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Safety & Cyber Evaluations Engineer
AI Safety & Cyber Evaluations Engineer

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 300,000 - 405,000
Cyber Evaluations Engineer for Safe AI & Robustness
Cyber Evaluations Engineer for Safe AI & Robustness

Anthropic • San Francisco (CA), Washington

Hybrid
USD 300,000 - 405,000
Cyber Evaluations Engineer
Cyber Evaluations Engineer

Anthropic • United States

Remote
USD 150,000 - 230,000
Cyber Evaluations Engineer
Cyber Evaluations Engineer

EngineersOfAI • San Francisco (CA), Northern (KY)

On-site
USD 300,000 - 405,000
AI Security Research Engineer: Evaluation & Defense
AI Security Research Engineer: Evaluation & Defense

General Analysis • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Cyber Red Team Specialist: AI Safety & Adversary Testing
Cyber Red Team Specialist: AI Safety & Adversary Testing

OpenAI • Washington

Hybrid
USD 180,000 - 280,000
Relocation assistance
Hybrid work model
AI Safety Evaluations Analyst
AI Safety Evaluations Analyst

Anthropic • New York (NY)

Hybrid
USD 230,000 - 270,000
Competitive compensation
Equity donation matching
Generous vacation
+3
Cybersecurity Expert for Safe AI Evaluation
Cybersecurity Expert for Safe AI Evaluation

Next Frontier Capital • United States

On-site
USD 103,000 - 165,000
Evals Engineer, Offensive Cyber
Evals Engineer, Offensive Cyber

Zealot Labs • New York (NY)

On-site
USD 140,000 - 200,000
Cyber Evaluations Engineer
Cyber Evaluations Engineer

Anthropic • San Francisco (CA), Washington

On-site
USD 300,000 - 405,000