AI Safety Red Team Engineer

Anthropic

California (MO)

Hybrid

USD 320,000 - 405,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is seeking a Red Team Engineer to help ensure the safety of our deployed AI systems and products. You will take an adversarial approach to uncover vulnerabilities across our product ecosystem before they can be exploited by malicious actors.

Your work will blend traditional security practices with AI safety concerns, including jailbreaking, abuse vectors, and novel attack chaining. You will collaborate with Product, Engineering, and Policy teams to translate findings into concrete

Qualifications

  • Experience with penetration testing, red teaming, or application security.
  • Experience with model jailbreaking and testing large-scale agentic workflows for non-obvious prompt injection vectors.
  • Strong web application security skills and hands-on experience with security testing tools such as Burp Suite or Metasploit.
  • Experience building custom automation, including LLM-specific testing frameworks.
  • A track record of discovering novel attack vectors and chaining vulnerabilities in creative ways.

Responsibilities

  • Conduct comprehensive adversarial testing across Anthropic's product surfaces, developing creative attack scenarios that combine multiple exploitation techniques
  • Research and implement novel testing approaches for emerging capabilities, including agent systems, tool use, and new interaction paradigms
  • Design and execute "full kill chain" attacks that emulate real-world threat actors attempting to achieve specific malicious objectives
  • Build and maintain systematic testing methodologies that evaluate every aspect of our systems
  • Develop automated testing frameworks to enable continuous assessment at scale
  • Collaborate with Product, Engineering, and Policy teams to translate findings into concrete improvements
  • Help establish metrics for measuring detection effectiveness of novel abuse

Skills

Penetration testing
Red teaming
Application security
Model jailbreaking
Agentic workflows testing
Communication skills

Education

Bachelor’s degree or equivalent

Tools

Burp Suite
Metasploit
Custom scripting frameworks

Job description

Anthropic is seeking a Red Team Engineer to help ensure the safety of our deployed AI systems and products. You will take an adversarial approach to uncover vulnerabilities across our product ecosystem before they can be exploited by malicious actors.

Your work will blend traditional security practices with AI safety concerns, including jailbreaking, abuse vectors, and novel attack chaining. You will collaborate with Product, Engineering, and Policy teams to translate findings into concrete

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Safety Red Team Engineer
AI Safety Red Team Engineer

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 405,000
Red Team Engineer — AI Safety & Adversarial Testing (Flexible Hours)
Red Team Engineer — AI Safety & Adversarial Testing (Flexible Hours)

Doist • San Francisco (CA)

Hybrid
USD 320,000 - 405,000
Equity donation matching
Generous vacation
Parental leave
+2
AI Safety Red Team Engineer
AI Safety Red Team Engineer

Crossing Hurdles • United States

On-site
USD 120,000 - 180,000
AI Safety Red Team Specialist (Remote)
AI Safety Red Team Specialist (Remote)

Obsidian • New York (NY)

On-site
USD 90,000 - 130,000
Experience in human data-driven AI red teaming
Role in enhancing AI safety and trustworthiness
Frontier AI Safety Red Team Expert
Frontier AI Safety Red Team Expert

Obsidian • New York (NY)

On-site
USD 170,000 - 260,000
Cyber Red Team Specialist: AI Safety & Adversary Testing
Cyber Red Team Specialist: AI Safety & Adversary Testing

OpenAI • Washington

Hybrid
USD 180,000 - 280,000
Relocation assistance
Hybrid work model
AI Safety Red Team Engineer - Remote
AI Safety Red Team Engineer - Remote

Obsidian • New York (NY)

Remote
USD 85,000 - 120,000
Experience in AI red teaming
Role in enhancing AI robustness
Red Team Engineer, Safeguards
Red Team Engineer, Safeguards

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 405,000
Staff AI Safety Engineer — Red Team & Guardrails
Staff AI Safety Engineer — Red Team & Guardrails

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Top-tier compensation
Stock options
Health & wellness
+3
AI Safety Red Team Engineer (Remote • EN/DA)
AI Safety Red Team Engineer (Remote • EN/DA)

Neon • United States

Remote
USD 120,000 - 180,000