AI Safety Specialist - Bilingual

Mercor

Dubai

Remote

AED 150,000 - 190,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Mercor is building a dedicated red team to probe AI models with adversarial inputs and surface vulnerabilities. This role involves reviewing AI outputs on sensitive topics and generating red team data to strengthen safety for customers.

As a remote position, you will annotate failures, classify vulnerabilities, and document reproducible reports and attack cases, adhering to strict guidelines and quality standards.

Qualifications

  • Strong judgment about language and content; ability to judge AI responses for accuracy, completeness, and safety.
  • Rigorous attention to subtle errors, inconsistencies and gaps.
  • Structured approach following guidelines and quality standards, not ad hoc.
  • Excellent written/spoken communication to explain reasoning to technical and non-technical audiences.

Responsibilities

  • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation
  • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks
  • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent
  • Document reproducibly: produce reports, datasets, and attack cases customers can act on

Skills

Judgment and reasoning
Rigorous attention to detail
Structured approach
Clear communication
Adaptability

Job description

Location: Remote

Fluent Language Skills Required: English & Malayalam. Native fluency in English and Malayalam is required for this position.

Why This Role Exists

At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers.

This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated.

What You’ll Do
  • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation

  • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks

  • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent

  • Document reproducibly: produce reports, datasets, and attack cases customers can act on

Who You Are
  • You bring strong judgment about language and content: you can tell whether an AI response is accurate, complete, and appropriate, and explain why

  • You’re rigorous: you notice subtle errors, inconsistencies, and gaps that others skim past

  • You’re structured: you work to guidelines and quality standards consistently, not ad hoc

  • You’re communicative: you explain your reasoning clearly to technical and non-technical audiences

  • You’re adaptable: you thrive moving across projects, task types, and customers

Nice-to-Have Specialties
  • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction

  • Cybersecurity: penetration testing, exploit development, reverse engineering

  • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing

  • Creative probing: psychology, acting, writing for unconventional adversarial thinking

What Success Looks Like
  • You uncover vulnerabilities automated tests miss

  • You deliver reproducible artifacts that strengthen customer AI systems

  • Evaluation coverage expands: more scenarios tested, fewer surprises in production

  • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary

Why Join Mercor
  • Build experience in human data-driven AI red teaming at the frontier of safety

  • Play a direct role in making AI systems more robust, safe, and trustworthy

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Safety Specialist (English & Malayalam)
Remote AI Safety Specialist (English & Malayalam)

Tanqeeb • Dubai

On-site
AED 81,000 - 111,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • Dubai

Remote
AED 150,000 - 190,000
AI Red Teamer Expert
AI Red Teamer Expert

Tanqeeb • Dubai

On-site
AED 300,000 - 420,000
AI Engineer - Security
AI Engineer - Security

TechBiz Global GmbH • Abu Dhabi

On-site
AED 180,000 - 300,000
AI Engineer - Security
AI Engineer - Security

Client of TechBiz Global • Abu Dhabi

On-site
AED 240,000 - 360,000
Senior Offensive Security Engineer
Senior Offensive Security Engineer

Deriv.com • Dubai

Hybrid
AED 400,000 - 720,000
Principal AI Security Expert
Principal AI Security Expert

Uney GmbH • Dubai

On-site
AED 400,000 - 600,000
Music Producer - Malayalam Specialist
Music Producer - Malayalam Specialist

Mercor • Dubai

On-site
AED 67,000 - 112,000
AI Safety Engineer: Free-Tier Abuse Mitigation (Remote)
AI Safety Engineer: Free-Tier Abuse Mitigation (Remote)

Tanqeeb • Abu Dhabi

On-site
AED 514,000 - 698,000
Travel stipend
Offsite trips
Learning stipend
+2
Lead AI Scientist / Head of AI Solutions
Lead AI Scientist / Head of AI Solutions

EstateSight AI • Abu Dhabi

On-site
AED 450,000 - 900,000
Opportunity to lead AI innovation with societal impact
Research-driven environment
Leadership exposure across teams