AI Safety Red Team Engineer

Anthropic

San Francisco (CA)

Hybrid

USD 320,000 - 405,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Anthropic’s Safeguards team is seeking a Red Team Engineer to help ensure the safety of our deployed AI systems and products. You will take an adversarial approach to uncover vulnerabilities across our product ecosystem before they can be exploited by malicious actors.

You’ll simulate threat actors who chain multiple attack vectors to achieve objectives, spanning product features, accounts, payments, and novel exploit paths.

Qualifications

  • Experience in penetration testing, red teaming, or application security.
  • Experience in model jailbreaking and testing large-scale agentic workflows for non-obvious prompt injection vectors.
  • Strong technical skills in web application security, including hands-on expertise with security testing tools (e.g., Burp Suite, Metasploit, custom scripting frameworks).
  • Experience building custom automation, including LLM-specific testing frameworks.
  • A track record of discovering novel attack vectors and chaining vulnerabilities in creative ways.
  • A public body of work such as CVEs, blog posts, or disclosed bug bounty reports.
  • Strong written and verbal communication skills, with the ability to explain technical concepts to varied audiences.

Responsibilities

  • Conduct comprehensive adversarial testing across Anthropic's product surfaces, developing creative attack scenarios that combine multiple exploitation techniques
  • Research and implement novel testing approaches for emerging capabilities, including agent systems, tool use, and new interaction paradigms
  • Design and execute "full kill chain" attacks that emulate real‑world threat actors attempting to achieve specific malicious objectives
  • Build and maintain systematic testing methodologies that evaluate every aspect of our systems
  • Develop automated testing frameworks to enable continuous assessment at scale
  • Collaborate with Product, Engineering, and Policy teams to translate findings into concrete improvements
  • Help establish metrics for measuring detection effectiveness of novel abuse

Skills

Penetration testing
Red teaming
Web security
Burp Suite
Metasploit
LLM testing
Automation
Agentic workflows
Threat modeling

Education

Bachelor's degree

Tools

Burp Suite
Metasploit
Custom scripting frameworks

Job description

Anthropic’s Safeguards team is seeking a Red Team Engineer to help ensure the safety of our deployed AI systems and products. You will take an adversarial approach to uncover vulnerabilities across our product ecosystem before they can be exploited by malicious actors.

You’ll simulate threat actors who chain multiple attack vectors to achieve objectives, spanning product features, accounts, payments, and novel exploit paths.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Safety Red Team Engineer
AI Safety Red Team Engineer

Anthropic • California (MO)

Hybrid
USD 320,000 - 405,000
Red Team Engineer — AI Safety & Adversarial Testing (Flexible Hours)
Red Team Engineer — AI Safety & Adversarial Testing (Flexible Hours)

Doist • San Francisco (CA)

Hybrid
USD 320,000 - 405,000
Equity donation matching
Generous vacation
Parental leave
+2
AI Safety Red Team Specialist (Remote)
AI Safety Red Team Specialist (Remote)

Obsidian • New York (NY)

On-site
USD 90,000 - 130,000
Experience in human data-driven AI red teaming
Role in enhancing AI safety and trustworthiness
AI Safety Red Team Engineer
AI Safety Red Team Engineer

Crossing Hurdles • United States

On-site
USD 120,000 - 180,000
Red Team Engineer, Safeguards
Red Team Engineer, Safeguards

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 405,000
AI Safety Red Team Engineer (Remote • EN/DA)
AI Safety Red Team Engineer (Remote • EN/DA)

Neon • United States

Remote
USD 120,000 - 180,000
AI Red Teamer: Offensive Security for AI Systems (Remote)
AI Red Teamer: Offensive Security for AI Systems (Remote)

Handshake • United States

Remote
USD 150,000 - 210,000
Cyber Red Team Specialist: AI Safety & Adversary Testing
Cyber Red Team Specialist: AI Safety & Adversary Testing

OpenAI • Washington

Hybrid
USD 180,000 - 280,000
Relocation assistance
Hybrid work model
Remote AI Safety Red Team Engineer
Remote AI Safety Red Team Engineer

Neon • United States

Remote
USD 120,000 - 190,000
Staff AI Safety Engineer — Red Team & Guardrails
Staff AI Safety Engineer — Red Team & Guardrails

Visa Hunt • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Top-tier compensation
Stock options
Health & wellness
+3