AI Safety Specialist - Fully Remote

Mercor

Berlin

Remote

EUR 70.000 - 110.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Mercor is building a red team for AI safety, focusing on adversarial testing of conversational models. You’ll explore jailbreaks, prompt injections, and misuse scenarios while generating high-quality data and reproducible artifacts to reinforce client AI systems.

This remote role offers flexible engagement across projects and clients, with clear guidelines and wellness resources for sensitive content. English and Dutch fluency are required.

Qualifikationen

  • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing).
  • You’re curious and adversarial: you instinctively push systems to breaking points.
  • You’re structured: you use frameworks or benchmarks, not just random hacks.
  • You’re communicative: you explain risks clearly to technical and non-technical stakeholders.
  • You’re adaptable: thrive on moving across projects and customers.

Aufgaben

  • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation.
  • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks.
  • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent.
  • Document reproducibly: produce reports, datasets, and attack cases customers can act on.

Kenntnisse

Red team testing
Adversarial AI
Prompt injection
Bias exploitation
Multi-turn manipulation

Tools

Taxonomies
Benchmarks
Playbooks

Jobbeschreibung

Location

Remote

Fluent Language Skills Required

English & Dutch. Native fluency in English and Dutch is required for this position.

Why This Role Exists

At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers.

This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated.

What You’ll Do
  • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation
  • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks
  • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent
  • Document reproducibly: produce reports, datasets, and attack cases customers can act on
Who You Are
  • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing)
  • You’re curious and adversarial: you instinctively push systems to breaking points
  • You’re structured: you use frameworks or benchmarks, not just random hacks
  • You’re communicative: you explain risks clearly to technical and non-technical stakeholders
  • You’re adaptable: thrive on moving across projects and customers
Nice-to-Have Specialties
  • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction
  • Cybersecurity: penetration testing, exploit development, reverse engineering
  • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing
  • Creative probing: psychology, acting, writing for unconventional adversarial thinking
What Success Looks Like
  • You uncover vulnerabilities automated tests miss
  • You deliver reproducible artifacts that strengthen customer AI systemsEvaluation coverage expands: more scenarios tested, fewer surprises in production
  • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary
Why Join Mercor
  • Build experience in human data-driven AI red teaming at the frontier of safety
  • Play a direct role in making AI systems more robust, safe, and trustworthy
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

AI Safety Expert - Red Teaming
AI Safety Expert - Red Teaming

Mercor • Berlin

Vor Ort
EUR 70.000 - 110.000
AI Safety Red Teamer
AI Safety Red Teamer

AIUC • Deutschland

Remote
EUR 90.000 - 130.000
Senior Agentic AI Engineer (m/f/d)
Senior Agentic AI Engineer (m/f/d)

Mercanis • Berlin

Remote
EUR 90.000 - 140.000
Flexible remote setup
Career growth opportunities
Autonomy in AI components
Freelance AI Red Team Engineer
Freelance AI Red Team Engineer

Mindrift • Berlin

Vor Ort
EUR 54.611 - 77.780
Competitive hourly rates
Flexible working hours
Experience with advanced AI projects
Risk Engineer - Fully Remote | Upto $80/hr
Risk Engineer - Fully Remote | Upto $80/hr

Obsidian • Berlin

Remote
EUR 242.493 - 484.986
Generalist Expert
Generalist Expert

Mercor • Deutschland

Vor Ort
EUR 27.552 - 41.328
Competitive pay
Flexible schedule
Opportunity to collaborate with experts in AI
Technical Sales Consultant - Fully Remote | Upto $150/hr
Technical Sales Consultant - Fully Remote | Upto $150/hr

Obsidian • Berlin

Remote
EUR 60.000 - 90.000
Information Security Analyst
Information Security Analyst

Obsidian • Berlin

Vor Ort
EUR 16.907 - 20.772
Performance bonuses
Task completion pay based on quality
Freelance Cybersecurity Analyst - AI Trainer
Freelance Cybersecurity Analyst - AI Trainer

Mindrift • Hamburg

Vor Ort
EUR 54.611 - 77.780
Competitive pay rates
Work on advanced AI projects
Flexibility around other commitments
Freelance Cybersecurity Analyst - AI Trainer
Freelance Cybersecurity Analyst - AI Trainer

Mindrift • Hamburg

Vor Ort
EUR 46.860 - 67.948
Competitive hourly rates up to $58
Flexibility to work on your own schedule
Valuable experience on advanced AI projects