AI Safety Red Teamer Expert

Mercor

Berlin

Vor Ort

EUR 90.000 - 130.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Eine zielgenaue Bewerbung für diesen Job — ein maßgeschneiderter Lebenslauf und ein Anschreiben, die genau zur Stellenanzeige passen.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Mercor is seeking experienced AI Safety Red Teamer to identify vulnerabilities in frontier AI systems through adversarial testing. You will design challenging prompts, uncover model weaknesses, and evaluate AI behavior across complex, high-risk, and ambiguous topics.

Responsibilities include designing adversarial prompts to stress-test frontier AI models, identifying jailbreaks, unsafe behaviours, hallucinations, and policy failures, and evaluating model robustness across misinformation, cyber,

Qualifikationen

  • Bachelor's degree or higher in Computer Science, Cybersecurity, Journalism, Communications, Psychology, Biology, Chemistry, Public Policy, or a related discipline.
  • 5+ years of professional experience in AI Safety, AI Red Teaming, Trust & Safety, cybersecurity, investigative journalism, life sciences, or a related field.
  • Strong analytical reasoning, prompt design, and written communication skills.
  • Experience designing adversarial prompts or evaluating frontier AI systems.

Aufgaben

  • Design adversarial prompts to stress-test frontier AI models.
  • Identify jailbreaks, unsafe behaviours, hallucinations, and policy failures.
  • Evaluate model robustness across misinformation, cyber, biosecurity, fraud, political content, and other sensitive domains.
  • Document vulnerabilities and contribute to safety benchmarking and red-teaming reports.
  • Collaborate with AI researchers to improve model alignment, robustness, and safety.

Kenntnisse

Analytical reasoning
Prompt design
Written communication
Adversarial testing
Frontier AI familiarity

Ausbildung

Bachelor's degree or higher

Jobbeschreibung

We are seeking experienced AI Safety Red Teamers to identify vulnerabilities in frontier AI systems through adversarial testing. You will design challenging prompts, uncover model weaknesses, and evaluate AI behavior across complex, high-risk, and ambiguous ("grey-area") topics.

Responsibilities
  • Design adversarial prompts to stress-test frontier AI models.

  • Identify jailbreaks, unsafe behaviours, hallucinations, and policy failures.

  • Evaluate model robustness across misinformation, cyber, biosecurity, fraud, political content, and other sensitive domains.

  • Document vulnerabilities and contribute to safety benchmarking and red-teaming reports.

  • Collaborate with AI researchers to improve model alignment, robustness, and safety.

Required Qualifications
  • Bachelor's degree or higher in Computer Science, Cybersecurity, Journalism, Communications, Psychology, Biology, Chemistry, Public Policy, or a related discipline.

  • 5+ years of professional experience in AI Safety, AI Red Teaming, Trust & Safety, cybersecurity, investigative journalism, life sciences, or a related field.

  • Strong analytical reasoning, prompt design, and written communication skills.

  • Experience designing adversarial prompts or evaluating frontier AI systems.

Preferred Qualifications
  • Experience with AI Red Teaming, RLHF, SFT, AI Alignment, or Trust & Safety.

  • Familiarity with jailbreak testing, prompt engineering, or adversarial evaluation methodologies.

  • Expertise in one or more grey-area domains, including cyber, biosecurity, political content, misinformation, or scientific safety.

Why Join?
  • Help secure and strengthen the next generation of frontier AI models.

  • Work on cutting-edge adversarial testing alongside leading AI researchers and safety teams.

  • Influence how AI systems respond to complex, real-world safety challenges.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Obsidian • Berlin

Vor Ort
EUR 90.000 - 140.000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Mercor • Berlin

Vor Ort
EUR 75.000 - 110.000
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Mercor • Berlin

Vor Ort
EUR 70.000 - 120.000
AI Safety Expert - Red Teaming
AI Safety Expert - Red Teaming

Mercor • Berlin

Vor Ort
EUR 70.000 - 110.000
Senior Security Agent / AI Red Team Engineer
Senior Security Agent / AI Red Team Engineer

Jobtailor • Deutschland

Hybrid
EUR 103.000 - 142.000
Senior AI Engineer, Security Infrastructure
Senior AI Engineer, Security Infrastructure

Air • Deutschland

Hybrid
EUR 120.000 - 190.000
Freelance AI Red Team Engineer
Freelance AI Red Team Engineer

Mindrift • Berlin

Remote
Competitive hourly rates
Flexible working hours
Experience with advanced AI projects
Cybersecurity Expert - Offensive Security
Cybersecurity Expert - Offensive Security

Mercor • Berlin

Vor Ort
EUR 90.000 - 130.000
InfoSec Expert
InfoSec Expert

aitrainer • Deutschland

Remote
EUR 72.638 - 121.064
Freelance Cybersecurity Analyst - AI Trainer
Freelance Cybersecurity Analyst - AI Trainer

Mindrift • Hamburg

Remote
Competitive pay rates
Work on advanced AI projects
Flexibility around other commitments