AI Safety Red Teamer Expert

Mercor

Milano

In loco

EUR 90.000 - 130.000

Tempo pieno

14 giorni+

Ricevi più risposte dai datori di lavoro

Invia un CV specifico per questa offerta in pochi minuti.

Descrizione del lavoro

Mercor seeks experienced AI Safety Red Teamers to identify vulnerabilities in frontier AI systems through adversarial testing. You will design prompts, uncover model weaknesses, and evaluate AI behavior across high-risk, ambiguous topics.

Join a team of researchers to strengthen model alignment and safety, document findings, and contribute to benchmarking in a fast-moving field.

Competenze

  • Bachelor's degree or higher in a relevant field.
  • 5+ years of professional experience in AI Safety or related field.
  • Experience designing adversarial prompts or evaluating frontier AI systems.
  • Strong analytical reasoning, prompt design, and written communication.

Mansioni

  • Design adversarial prompts to stress-test frontier AI models.
  • Identify jailbreaks, unsafe behaviours, hallucinations, and policy failures.
  • Evaluate model robustness across misinformation, cyber, biosecurity, fraud, political content, and other sensitive domains.
  • Document vulnerabilities and contribute to safety benchmarking and red-teaming reports.
  • Collaborate with AI researchers to improve model alignment, robustness, and safety.

Conoscenze

Analytical reasoning
Prompt design
Written communication
Adversarial testing
AI Safety
Red Teaming

Formazione

Bachelor's degree or higher
Related disciplines

Descrizione del lavoro

We are seeking experienced AI Safety Red Teamers to identify vulnerabilities in frontier AI systems through adversarial testing. You will design challenging prompts, uncover model weaknesses, and evaluate AI behavior across complex, high-risk, and ambiguous ("grey-area") topics.

Responsibilities
  • Design adversarial prompts to stress-test frontier AI models.

  • Identify jailbreaks, unsafe behaviours, hallucinations, and policy failures.

  • Evaluate model robustness across misinformation, cyber, biosecurity, fraud, political content, and other sensitive domains.

  • Document vulnerabilities and contribute to safety benchmarking and red-teaming reports.

  • Collaborate with AI researchers to improve model alignment, robustness, and safety.

Required Qualifications
  • Bachelor's degree or higher in Computer Science, Cybersecurity, Journalism, Communications, Psychology, Biology, Chemistry, Public Policy, or a related discipline.

  • 5+ years of professional experience in AI Safety, AI Red Teaming, Trust & Safety, cybersecurity, investigative journalism, life sciences, or a related field.

  • Strong analytical reasoning, prompt design, and written communication skills.

  • Experience designing adversarial prompts or evaluating frontier AI systems.

Preferred Qualifications
  • Experience with AI Red Teaming, RLHF, SFT, AI Alignment, or Trust & Safety.

  • Familiarity with jailbreak testing, prompt engineering, or adversarial evaluation methodologies.

  • Expertise in one or more grey-area domains, including cyber, biosecurity, political content, misinformation, or scientific safety.

Why Join?
  • Help secure and strengthen the next generation of frontier AI models.

  • Work on cutting-edge adversarial testing alongside leading AI researchers and safety teams.

  • Influence how AI systems respond to complex, real-world safety challenges.

Ottieni la revisione del curriculum gratis e riservata.
o trascina qui il file.
Similar jobs

Offerte di lavoro simili che vale la pena confrontare

AI Safety Red Teamer Expert
AI Safety Red Teamer Expert

Mercor • Roma

In loco
EUR 90.000 - 130.000
Adversarial AI Safety Architect for Frontier Models
Adversarial AI Safety Architect for Frontier Models

Mercor • Roma

In loco
EUR 90.000 - 130.000
Remote AI Safety Red Teamer - Adversarial Prompting
Remote AI Safety Red Teamer - Adversarial Prompting

Mercor • Roma

In loco
EUR 84.000 - 100.000
Frontier AI Safety Evaluator
Frontier AI Safety Evaluator

Mercor • Roma

Remoto
EUR 60.000 - 90.000
Frontier AI Safety Evaluator: Expert Feedback
Frontier AI Safety Evaluator: Expert Feedback

Mercor • Roma

In loco
EUR 65.000 - 95.000
Freelance Cybersecurity Analyst - AI Trainer
Freelance Cybersecurity Analyst - AI Trainer

Mindrift • Roma

Remoto
Competitive pay rates up to $34/hour
Flexible freelance work hours
Experience on advanced AI projects
Remote AI Training Lead - Data & Model Quality
Remote AI Training Lead - Data & Model Quality

YO IT Consulting • Roma

In loco
EUR 34.440 - 55.104
Research Engineer Position on Secure Agentic AI Systems
Research Engineer Position on Secure Agentic AI Systems

AI4I • Collegno

In loco
EUR 45.000 - 75.000
Competitive compensation
Full support for conference travel
Professional development opportunities
+2
Applied AI Architect, Industries
Applied AI Architect, Industries

Anthropic • Milano

Ibrido
EUR 110.000 - 170.000
Competitive compensation
Equity donation matching
Generous vacation
+3
AI Solutions Engineer
AI Solutions Engineer

N-iX • Lombardia

In loco
EUR 60.000 - 90.000