AI Safety Expert - Adversarial ML

mercor

India

À distance

INR 2 136 000 - 2 937 000

Plein temps

Il y a 9 jours
Générateur de candidature

Une candidature complète en une minute — un CV et une lettre de motivation personnalisés, prêts à envoyer.

Passez les filtres ATS

Résumé du poste

Mercor connects elite creative and technical talent with leading AI research labs. This remote contract role focuses on safety testing of conversational AI, including red-teaming and data annotation, to reveal vulnerabilities and risks.

You will work asynchronously, fluent in English and Bengali, with strong judgment and clear communication to technical and non-technical audiences. Preferred experience includes adversarial ML and cybersecurity concepts.

Qualifications

  • Native fluency in English and Bengali.
  • Strong judgment about language and content accuracy.
  • Rigorous attention to subtle errors and inconsistencies.
  • Ability to work to guidelines and quality standards consistently.
  • Clear communication with technical and non-technical audiences.
  • Adaptability across projects, task types, and customers.

Responsabilités

  • Red team conversational AI models and agents. Perform jailbreaks, prompt injections, misuse cases, and bias exploitation.
  • Generate high-quality human data. Annotate failures, classify vulnerabilities, and flag systemic risks.
  • Apply structure by following taxonomies, benchmarks, and playbooks to ensure consistent testing.
  • Document reproducibly to produce reports, datasets, and attack cases that customers can act on.
  • Work independently and asynchronously to meet deadlines while improving AI model performance.

Connaissances

English fluency
Bengali fluency
Attention to detail
Clear communication

Description du poste

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: AI Safety Experts — English & Bengali Type: Contract Compensation: $16–$22/hour Location: Remote

Responsibilities
  • Red team conversational AI models and agents. Perform jailbreaks, prompt injections, misuse cases, and bias exploitation.
  • Generate high-quality human data. Annotate failures, classify vulnerabilities, and flag systemic risks.
  • Apply structure by following taxonomies, benchmarks, and playbooks to ensure consistent testing.
  • Document reproducibly to produce reports, datasets, and attack cases that customers can act on.
  • Work independently and asynchronously to meet deadlines while improving AI model performance.
Must-Have
  • Native fluency in English and Bengali.
  • Strong judgment about language and content accuracy.
  • Rigorous attention to subtle errors and inconsistencies.
  • Ability to work to guidelines and quality standards consistently.
  • Clear communication with technical and non-technical audiences.
  • Adaptability across projects, task types, and customers.
Preferred
  • Experience in Adversarial ML : jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction.
  • Cybersecurity skills: penetration testing, exploit development, reverse engineering.
  • Socio-technical risk expertise: harassment/disinfo probing, abuse analysis, conversational AI testing.
  • Creative probing skills: psychology, acting, writing for unconventional adversarial thinking.
Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

AI Safety Expert - Red Teamer
AI Safety Expert - Red Teamer

mercor • Inde

À distance
INR 2 670 000 - 2 937 000
AI Safety Expert - Fully Remote | Upto $22/hr
AI Safety Expert - Fully Remote | Upto $22/hr

mercor • Inde

À distance
INR 2 136 000 - 2 937 000
AI Safety Expert - Red Teaming
AI Safety Expert - Red Teaming

mercor • Inde

À distance
INR 2 670 000 - 2 937 000
AI Red Team Specialist - Remote | Upto $22/hr
AI Red Team Specialist - Remote | Upto $22/hr

mercor • Inde

À distance
INR 2 670 000 - 2 937 000
AI Adversarial Specialist - Fully Remote | Upto $22/hr
AI Adversarial Specialist - Fully Remote | Upto $22/hr

mercor • Inde

À distance
INR 2 670 000 - 2 937 000
AI Safety Expert - Red Team
AI Safety Expert - Red Team

Mercor • Mumbai

À distance
INR 2 136 000 - 2 937 000
AI Safety Expert - Adversarial ML
AI Safety Expert - Adversarial ML

Mercor • Bengaluru

À distance
INR 1 200 000 - 2 400 000
AI Safety Expert - Fully Remote | Upto $22/hr
AI Safety Expert - Fully Remote | Upto $22/hr

Mercor • Mumbai

À distance
INR 1 000 000 - 2 000 000
AI Safety Specialist - Bilingual
AI Safety Specialist - Bilingual

Mercor • Mumbai

À distance
INR 900 000 - 1 300 000
AI Red Team Tester
AI Red Team Tester

Alignerr Corp. • Bengaluru

À distance
INR 1 378 000 - 2 755 000