AI Safety Expert - Red Team

Mercor

San Francisco (CA)

Remote

USD 23,000 - 34,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Mercor, headquartered in San Francisco, is seeking AI Safety Experts to work remotely on English and Malay language safety testing for AI models. This contract role focuses on red-teaming models, data generation, and reproducible reporting to help clients understand and mitigate risks.

Applicants should have strong English and Malay proficiency, prior AI adversarial testing experience, and excellent communication skills for explaining risks to diverse stakeholders.

Qualifications

  • Fluent in English and Malay.
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
  • Strong communication skills to explain risks clearly to technical and non-technical stakeholders.

Responsibilities

  • Red team conversational AI models and agents. Focus on jailbreaks, prompt injections, misuse cases, and bias exploitation.
  • Generate high-quality human data. Annotate failures, classify vulnerabilities, and flag systemic risks.
  • Apply structure by following taxonomies, benchmarks, and playbooks to maintain testing consistency.
  • Document reproducibly. Produce reports, datasets, and attack cases that customers can act on.
  • Work independently and asynchronously to meet deadlines while improving AI model performance.

Skills

English
Malay
Red teaming
Cybersecurity
Communication

Job description

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.

Position: AI Safety Experts — English & Malay
Type: Contract
Compensation: $17–$25/hour
Location: Remote

Role Responsibilities
  • Red team conversational AI models and agents. Focus on jailbreaks, prompt injections, misuse cases, and bias exploitation.
  • Generate high-quality human data. Annotate failures, classify vulnerabilities, and flag systemic risks.
  • Apply structure by following taxonomies, benchmarks, and playbooks to maintain testing consistency.
  • Document reproducibly. Produce reports, datasets, and attack cases that customers can act on.
  • Work independently and asynchronously to meet deadlines while improving AI model performance.
Qualifications
Must-Have
  • Fluent in English and Malay.
  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
  • Strong communication skills to explain risks clearly to technical and non-technical stakeholders.
Preferred
  • Experience in Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction.
  • Background in Cybersecurity: penetration testing, exploit development, reverse engineering.
  • Knowledge in socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing.
  • Creative probing skills: psychology, acting, writing for unconventional adversarial thinking.
Resources & Support
  • For details about the interview process and platform information, please check: https://talent.docs.mercor.com/welcome
  • For any help or support, reach out to: support@mercor.com
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Safety Expert - Red Team
AI Safety Expert - Red Team

Mercor • New York (NY)

On-site
USD 120,000 - 180,000
AI Safety Specialist - Bilingual
AI Safety Specialist - Bilingual

Mercor • San Francisco (CA)

Remote
AUD 31,000 - 43,000
AI Safety Expert - Adversarial ML
AI Safety Expert - Adversarial ML

Mercor • New York (NY)

Remote
USD 22,000 - 30,000
AI Safety Specialist - Bilingual
AI Safety Specialist - Bilingual

Mercor • New York (NY)

Remote
USD 90,000 - 130,000
AI Safety Experts - English & Malayalam
AI Safety Experts - English & Malayalam

Neon • United States

Remote
USD 90,000 - 130,000
AI Safety Experts — English & Malayalam
AI Safety Experts — English & Malayalam

Neon • United States

Remote
USD 90,000 - 160,000
AI Safety Expert - Red Teaming
AI Safety Expert - Red Teaming

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 140,000
AI Safety Experts — English & Tamil
AI Safety Experts — English & Tamil

Neon • United States

Remote
USD 90,000 - 140,000
AI Safety Experts - English & Tamil
AI Safety Experts - English & Tamil

Neon • United States

Remote
USD 120,000 - 180,000
AI Safety Expert - Adversarial ML
AI Safety Expert - Adversarial ML

Mercor • San Francisco (CA)

Remote
USD 120,000 - 170,000