AI Safety Expert - Adversarial ML

Mercor

New York (NY)

Remote

USD 22,000 - 30,000

Full time

12 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Mercor is seeking AI Safety Experts to join our fully remote team. In this contract role, you will red team conversational AI models to identify jailbreaks, prompt injections, and misuse cases, and generate high-quality data by annotating failures and flagging risks.

You must be fluent in English and Marathi, with native proficiency, and capable of clear communication with both technical and non-technical audiences. You will follow taxonomies and playbooks to maintain testing quality.

Qualifications

  • English and Marathi fluency with native proficiency in both languages.
  • Excellent judgment about language and content accuracy.
  • Strong attention to subtle errors and inconsistencies.
  • Structured approach to guidelines and quality standards.
  • Clear communication with technical and non-technical audiences.
  • Adaptability across projects, tasks, and customers.

Responsibilities

  • Red team conversational AI models and agents to identify jailbreaks and misuse cases.
  • Generate high-quality human data by annotating failures and flagging systemic risks.
  • Apply structure by following taxonomies, benchmarks, and playbooks for consistent testing.
  • Document reproducibly by producing reports, datasets, and attack cases for customer action.

Skills

English & Marathi fluency
Attention to detail
Strong judgment
Structured approach
Clear communication
Adaptability
Adversarial ML
Cybersecurity
Socio-technical risk
Creative probing

Job description

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.

Position

Position: AI Safety Experts — English & Marathi
Type: Contract
Compensation: $16–$22/hour
Location: Remote

Role Responsibilities
  • Red team conversational AI models and agents to identify jailbreaks, prompt injections, and misuse cases.
  • Generate high-quality human data by annotating failures, classifying vulnerabilities, and flagging systemic risks.
  • Apply structure by following taxonomies, benchmarks, and playbooks to maintain consistent testing.
  • Document reproducibly by producing reports, datasets, and attack cases for customer action.
Qualifications
Must-Have
  • Fluent Language Skills Required: English & Marathi. Native fluency in English and Marathi is required.
  • Strong judgment about language and content accuracy.
  • Rigorous attention to subtle errors and inconsistencies.
  • Structured approach to guidelines and quality standards.
  • Clear communication with technical and non-technical audiences.
  • Adaptability across projects, task types, and customers.
Preferred
  • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction.
  • Cybersecurity: penetration testing, exploit development, reverse engineering.
  • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing.
  • Creative probing: psychology, acting, writing for unconventional adversarial thinking.
Resources & Support
  • For details about the interview process and platform information, please check: https://talent.docs.mercor.com/welcome
  • For any help or support, reach out to: support@mercor.com

PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Safety Experts — English & Marathi
AI Safety Experts — English & Marathi

Neon • United States

Remote
USD 90,000 - 150,000
AI Safety Experts - English & Marathi
AI Safety Experts - English & Marathi

Neon • United States

Remote
USD 120,000 - 180,000
AI Safety Expert - Red Team
AI Safety Expert - Red Team

Mercor • San Francisco (CA)

Remote
USD 23,000 - 34,000
AI Safety Experts - English & Malayalam
AI Safety Experts - English & Malayalam

Neon • United States

Remote
USD 90,000 - 130,000
AI Safety Experts — English & Tamil
AI Safety Experts — English & Tamil

Neon • United States

Remote
USD 90,000 - 140,000
AI Safety Expert - Adversarial ML
AI Safety Expert - Adversarial ML

Mercor • San Francisco (CA)

Remote
USD 120,000 - 170,000
AI Safety Experts - English & Tamil
AI Safety Experts - English & Tamil

Neon • United States

Remote
USD 120,000 - 180,000
AI Safety Experts - English & Telugu
AI Safety Experts - English & Telugu

Neon • United States

Remote
USD 120,000 - 170,000
AI Safety Experts — English & Malayalam
AI Safety Experts — English & Malayalam

Neon • United States

Remote
USD 90,000 - 160,000
AI Safety Experts — English & Assamese
AI Safety Experts — English & Assamese

Neon • United States

Remote
USD 120,000 - 180,000
Remote work
Wellness resources
Work on safety-critical AI