AI Safety Expert - Red Teaming

Mercor

Berlin

Vor Ort

EUR 70.000 - 110.000

Teilzeit

14 Tage+
Bewerbungsgenerator

Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Mercor is building a remote red team to probe AI models with adversarial inputs. The role focuses on jailbreaks, prompt injections, bias testing, and safety testing across client projects.

This work is text-based and conducted under clear guidelines to ensure safety. You will generate high-quality human data, annotate vulnerabilities, and document reproducible reports and attack cases that customers can act on.

Qualifikationen

  • Prior red-teaming experience focusing on AI systems.
  • Experience with adversarial testing, cybersecurity, and socio-technical probing.
  • Ability to structure tests using frameworks or benchmarks.
  • Strong written and verbal communication for diverse stakeholders.

Aufgaben

  • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias testing.
  • Generate high-quality annotated data: identify vulnerabilities and surface risks.
  • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent.
  • Document reproducibly: produce reports, datasets, and attack cases for customers.

Kenntnisse

Red team experience
AI adversarial work
Cybersecurity
Socio-technical probing
Structured methodology
Communication
Adaptability

Jobbeschreibung

Location: Remote

Fluent Language Skills Required: English & Swedish. Native fluency in English and Swedish is required for this position.

Why This Role Exists

At Mercor, we believe the safest AI is the one that’s already been attacked — by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers.

This project involves reviewing AI outputs that touch on sensitive topics such as bias, misinformation, or harmful behaviors. All work is text-based, and participation in higher-sensitivity projects is optional and supported by clear guidelines and wellness resources. Before being exposed to any content, the topics will be clearly communicated.

What You’ll Do
  • Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation

  • Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks

  • Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent

  • Document reproducibly: produce reports, datasets, and attack cases customers can act on

Who You Are
  • You bring prior red teaming experience (AI adversarial work, cybersecurity, socio-technical probing)

  • You’re curious and adversarial: you instinctively push systems to breaking points

  • You’re structured: you use frameworks or benchmarks, not just random hacks

  • You’re communicative: you explain risks clearly to technical and non-technical stakeholders

  • You’re adaptable: thrive on moving across projects and customers

Nice-to-Have Specialties
  • Adversarial ML: jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction

  • Cybersecurity: penetration testing, exploit development, reverse engineering

  • Socio-technical risk: harassment/disinfo probing, abuse analysis, conversational AI testing

  • Creative probing: psychology, acting, writing for unconventional adversarial thinking

What Success Looks Like
  • You uncover vulnerabilities automated tests miss

  • You deliver reproducible artifacts that strengthen customer AI systems

  • Evaluation coverage expands: more scenarios tested, fewer surprises in production

  • Mercor customers trust the safety of their AI because you’ve already probed it like an adversary

Why Join Mercor
  • Build experience in human data-driven AI red teaming at the frontier of safety

  • Play a direct role in making AI systems more robust, safe, and trustworthy

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

AI Safety Red Teamer Expert
AI Safety Red Teamer Expert

Mercor • Berlin

Vor Ort
EUR 90.000 - 130.000
Information Security Analyst
Information Security Analyst

Obsidian • Berlin

Vor Ort
Performance bonuses
Task completion pay based on quality
Freelance AI Red Team Engineer
Freelance AI Red Team Engineer

Mindrift • Berlin

Remote
Competitive hourly rates
Flexible working hours
Experience with advanced AI projects
Risk Engineer - Fully Remote | Upto $80/hr
Risk Engineer - Fully Remote | Upto $80/hr

Obsidian • Berlin

Remote
Generalist Expert
Generalist Expert

Mercor • Deutschland

Remote
Competitive pay
Flexible schedule
Opportunity to collaborate with experts in AI
Senior Security Agent / AI Red Team Engineer
Senior Security Agent / AI Red Team Engineer

Jobtailor • Deutschland

Hybrid
EUR 103.000 - 142.000
Senior SWE - AI Specialist
Senior SWE - AI Specialist

Mercor • Berlin

Vor Ort
EUR 178.000 - 249.000
Freelance Cybersecurity Analyst - AI Trainer
Freelance Cybersecurity Analyst - AI Trainer

Mindrift • Hamburg

Remote
Competitive pay rates
Work on advanced AI projects
Flexibility around other commitments
First-Line Supervisors of Non-Retail Sales Workers
First-Line Supervisors of Non-Retail Sales Workers

StudySmarter • Berlin

Remote
Technical Sales Consultant - Fully Remote | Upto $150/hr
Technical Sales Consultant - Fully Remote | Upto $150/hr

Obsidian • Berlin

Remote
EUR 60.000 - 90.000