AI Safety Expert - Red Team

Mercor

Johannesburg

Presencial

ZAR 1 453 000 - 2 098 000

Tempo integral

14 dias+
Gerador de candidaturas

Destaca-te nesta função — gera um currículo e uma carta de apresentação personalizados em cerca de um minuto.

Ultrapassa os filtros ATS

Resumo da oferta

Mercor is building a remote AI safety red team to rigorously test conversational models. You will expose vulnerabilities, annotate failures, and contribute to datasets and reports that help customers strengthen their AI systems.

The role demands prior red-teaming experience and a structured, communicative approach to documenting risks for diverse stakeholders across multiple projects.

Qualificações

  • Prior red teaming experience in AI adversarial work, cybersecurity, or socio-technical probing.
  • Structured thinker who uses frameworks or benchmarks, not ad-hoc hacks.
  • Clear communicator who explains risks to both technical and non-technical stakeholders.
  • Adaptable and able to move across projects and customers.

Responsabilidades

  • Red team conversational AI models and agents: jailbreaking, prompt injections, misuse cases, bias exploitation, multi-turn manipulation.
  • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks.
  • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent.
  • Document reproducibly: produce reports, datasets, and attack cases customers can act on.

Conhecimentos

Red teaming
AI security
Cybersecurity
Adversarial thinking

Descrição da oferta de emprego

Location: Remote

Fluent Language Skills Required: English & Portuguese (global, excluding Brazilian Portuguese). Native fluency in English and Portuguese (global, excluding Brazilian Portuguese) is required for this position.

Why This Role Exists

At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers.

This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated.

What You’ll Do
  • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation

  • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks

  • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent

  • Document reproducibly: produce reports, datasets, and attack cases customers can act on

Who You Are
  • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing)

  • You’re curious and adversarial: you instinctively push systems to breaking points

  • You’re structured: you use frameworks or benchmarks, not just random hacks

  • You’re communicative: you explain risks clearly to technical and non-technical stakeholders

  • You’re adaptable: thrive on moving across projects and customers

Nice-to-Have Specialties
  • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction

  • Cybersecurity: penetration testing, exploit development, reverse engineering

  • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing

  • Creative probing: psychology, acting, writing for unconventional adversarial thinking

What Success Looks Like
  • You uncover vulnerabilities automated tests miss

  • You deliver reproducible artifacts that strengthen customer AI systems

  • Evaluation coverage expands: more scenarios tested, fewer surprises in production

  • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary

Why Join Mercor
  • Build experience in human data-driven AI red teaming at the frontier of safety

  • Play a direct role in making AI systems more robust, safe, and trustworthy

Obtém a tua avaliação gratuita e confidencial do currículo.

ou arrasta e larga o ficheiro aqui.

Similar jobs

Ofertas semelhantes que vale a pena comparar

Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • Johannesburg

Presencial
ZAR 350 000 - 520 000
AI Safety Red Team Engineer - Remote
AI Safety Red Team Engineer - Remote

Mercor • Johannesburg

Presencial
ZAR 1 453 000 - 2 098 000
Remote AI Safety Red Team Engineer
Remote AI Safety Red Team Engineer

Obsidian • Johannesburg

Presencial
ZAR 1 457 000 - 2 104 000
Freelance AI Red Team Engineer
Freelance AI Red Team Engineer

Mindrift • África do Sul

Presencial
ZAR 462 078 - 658 111
Competitive hourly rates
Remote work flexibility
Experience in cutting-edge AI projects
AI / Emerging Tech Security Analyst
AI / Emerging Tech Security Analyst

Alignerr Corp. • Johannesburg

Teletrabalho
ZAR 917 000 - 2 063 000
Fully remote
Flexible schedule
Contract-based freelance work (hourly)
Freelance Cybersecurity Analyst - AI Trainer
Freelance Cybersecurity Analyst - AI Trainer

Mindrift • África do Sul

Presencial
ZAR 462 078 - 658 111
Flexible working hours
Remote work
Competitive pay up to $24/hour
+1
Vulnerability Management Analyst
Vulnerability Management Analyst

Alignerr Corp. • Johannesburg

Teletrabalho
ZAR 917 000 - 1 605 000
Threat Intelligence Analyst
Threat Intelligence Analyst

Alignerr Corp. • Johannesburg

Teletrabalho
ZAR 1 146 000 - 2 292 000
Autonomy
Flexibility
Global collaboration
Network & Infrastructure Security Analyst
Network & Infrastructure Security Analyst

Alignerr Corp. • Johannesburg

Teletrabalho
ZAR 917 000 - 1 605 000
Autonomy
Global collaboration
Flexible schedule
Security Operations Analyst
Security Operations Analyst

Alignerr Corp. • Johannesburg

Teletrabalho
ZAR 138 000 - 276 000
Autonomy
Variety
Global collaboration