AI Safety Red Team Specialist

Neon

United States

Remote

USD 120,000 - 170,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Mercor is building a red team to probe AI models with adversarial inputs and surface vulnerabilities. This role focuses on reviewing outputs touching sensitive topics and creating high-quality data to strengthen AI safety.

You will conduct jailbreaking, prompt injection, and bias-exploitation tests across multi-turn conversations, documenting reproducible attack cases and delivering artifacts for customers.

Responsibilities

  • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation
  • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks
  • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent
  • Document reproducibly: produce reports, datasets, and attack cases customers can act on
  • You bring strong judgment about language and content: you can tell whether an AI response is accurate, complete, and appropriate, and explain why
  • You’re rigorous: you notice subtle errors, inconsistencies, and gaps that others skim past
  • You’re structured: you work to guidelines and quality standards consistently, not ad hoc
  • You’re communicative: you explain your reasoning clearly to technical and non-technical audiences
  • You’re adaptable: you thrive moving across projects, task types, and customers
  • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction
  • Cybersecurity: penetration testing, exploit development, reverse engineering
  • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing
  • Creative probing: psychology, acting, writing for unconventional adversarial thinking

Job description

Mercor is building a red team to probe AI models with adversarial inputs and surface vulnerabilities. This role focuses on reviewing outputs touching sensitive topics and creating high-quality data to strengthen AI safety.

You will conduct jailbreaking, prompt injection, and bias-exploitation tests across multi-turn conversations, documenting reproducible attack cases and delivering artifacts for customers.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • New York (NY)

On-site
USD 110,000 - 180,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • San Francisco (CA)

Remote
USD 120,000 - 170,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • New York (NY)

On-site
USD 120,000 - 160,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • San Francisco (CA)

Remote
USD 90,000 - 140,000
Remote AI Red Team Engineer
Remote AI Red Team Engineer

Mercor • New York (NY)

Remote
USD 90,000 - 160,000
Remote Adversarial ML Specialist — AI Safety Red Team
Remote Adversarial ML Specialist — AI Safety Red Team

Obsidian • San Francisco (CA)

Remote
USD 120,000 - 180,000
Remote work
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • San Francisco (CA)

On-site
USD 120,000 - 180,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • San Francisco (CA)

Remote
USD 120,000 - 180,000
Remote AI Safety Red Team Expert
Remote AI Safety Red Team Expert

Mercor • New York (NY)

On-site
USD 120,000 - 180,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Obsidian • New York (NY)

Remote
USD 120,000 - 170,000