AI Safety Expert - Adversarial ML

Mercor

Greater London

Remote

GBP 17,000 - 23,000

Full time

13 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Mercor is seeking AI Safety Experts (English & Odia) for a remote contract role. You will test and audit conversational AI models, create high-quality data, and document findings to help customers improve model behavior.

The position offers flexible hours, independent work, and the chance to influence safety testing across projects. Strong linguistic judgment and attention to detail are essential for success in this fast-paced, globally remote environment.

Qualifications

  • Native English and Odia proficiency.
  • Strong judgment about language and content.
  • Meticulous attention to detail and consistency.
  • Structured approach to guidelines and quality standards.
  • Clear communication with technical and non-technical audiences.
  • Adaptability across projects, task types and customers.

Responsibilities

  • Red team conversational AI models and agents, including jailbreaks, prompt injections, misuse cases and bias exploitation.
  • Generate high-quality human data; annotate failures, classify vulnerabilities and flag systemic risks.
  • Apply structure using taxonomies, benchmarks and playbooks for consistent testing.
  • Document reproducibly; produce reports, datasets and attack cases for customers.
  • Work independently and asynchronously; thrive in flexible hours while improving AI model performance.

Skills

English (native)
Odia
Judgment about language & content
Attention to detail
Structured guidelines & standards
Clear communication
Adaptability across projects

Job description

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.


Position

AI Safety Experts - English & Odia
Type: Contract
Compensation: $16-$22/hour
Location: Remote


Role Responsibilities


  • Red team conversational AI models and agents. Conduct jailbreaks, prompt injections, misuse cases, and bias exploitation.

  • Generate high-quality human data. Annotate failures, classify vulnerabilities, and flag systemic risks.

  • Apply structure. Follow taxonomies, benchmarks, and playbooks to ensure consistent testing.

  • Document reproducibly. Produce reports, datasets, and attack cases that customers can act on.

  • Work independently and asynchronously. Thrive in flexible hours while improving AI model performance.


Qualifications

Must-Have


  • Native fluency in English and Odia.

  • Strong judgment about language and content.

  • Rigorous attention to detail and consistency.

  • Structured approach to guidelines and quality standards.

  • Clear communication with technical and non-technical audiences.

  • Adaptability across projects, task types, and customers.


Preferred


  • Experience in Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction.

  • Background in Cybersecurity: penetration testing, exploit development, reverse engineering.

  • Knowledge of socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing.

  • Creative probing skills: psychology, acting, writing for unconventional adversarial thinking.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Safety Expert - Red Team
AI Safety Expert - Red Team

Mercor • Greater London

On-site
GBP 50,000 - 65,000
AI Safety Specialist - Fully Remote | Upto $22/hr
AI Safety Specialist - Fully Remote | Upto $22/hr

Mercor • Greater London

Remote
GBP 21,000 - 23,000
AI Safety Specialist - Remote
AI Safety Specialist - Remote

Mercor • Greater London

Remote
GBP 72,000 - 86,000
AI Safety Red Team Engineer - Adversarial ML (Remote)
AI Safety Red Team Engineer - Adversarial ML (Remote)

Mercor • Greater London

Remote
GBP 17,000 - 23,000
AI Adversarial Specialist - Fully Remote
AI Adversarial Specialist - Fully Remote

Mercor • Greater London

Remote
GBP 52,000 - 81,000
Remote AI Safety Specialist: Adversarial Testing (EN/AS)
Remote AI Safety Specialist: Adversarial Testing (EN/AS)

Mercor • Greater London

Remote
GBP 21,000 - 23,000
AI Safety Specialist - Fully Remote
AI Safety Specialist - Fully Remote

Mercor • Greater London

Remote
GBP 70,000 - 110,000
AI Safety Red Team Specialist (Remote)
AI Safety Red Team Specialist (Remote)

Mercor • Greater London

On-site
GBP 60,000 - 90,000
Remote AI Adversarial Specialist — Red Team & Safety
Remote AI Adversarial Specialist — Red Team & Safety

Mercor • Greater London

Remote
GBP 80,000 - 110,000
AI Safety Red Teamer Expert
AI Safety Red Teamer Expert

Mercor • Greater London

On-site
GBP 90,000 - 130,000