Frontier AI Safety Red Teamer - Adversarial Testing Lead

Mercor

Milano

In loco

EUR 90.000 - 130.000

Tempo pieno

14 giorni+
Generatore di candidature

Non inviare un curriculum generico — genera un curriculum e una lettera di presentazione personalizzati per questo specifico impiego.

Supera i filtri ATS

Descrizione del lavoro

Mercor seeks experienced AI Safety Red Teamers to identify vulnerabilities in frontier AI systems through adversarial testing. You will design prompts, uncover model weaknesses, and evaluate AI behavior across high-risk, ambiguous topics.

Join a team of researchers to strengthen model alignment and safety, document findings, and contribute to benchmarking in a fast-moving field.

Competenze

  • Bachelor's degree or higher in a relevant field.
  • 5+ years of professional experience in AI Safety or related field.
  • Experience designing adversarial prompts or evaluating frontier AI systems.
  • Strong analytical reasoning, prompt design, and written communication.

Mansioni

  • Design adversarial prompts to stress-test frontier AI models.
  • Identify jailbreaks, unsafe behaviours, hallucinations, and policy failures.
  • Evaluate model robustness across misinformation, cyber, biosecurity, fraud, political content, and other sensitive domains.
  • Document vulnerabilities and contribute to safety benchmarking and red-teaming reports.
  • Collaborate with AI researchers to improve model alignment, robustness, and safety.

Conoscenze

Analytical reasoning
Prompt design
Written communication
Adversarial testing
AI Safety
Red Teaming

Formazione

Bachelor's degree or higher
Related disciplines

Descrizione del lavoro

We are seeking experienced AI Safety Red Teamers to identify vulnerabilities in frontier AI systems through adversarial testing. You will design challenging prompts, uncover model weaknesses, and evaluate AI behavior across complex, high-risk, and ambiguous ("grey-area") topics.

Responsibilities
  • Design adversarial prompts to stress-test frontier AI models.

  • Identify jailbreaks, unsafe behaviours, hallucinations, and policy failures.

  • Evaluate model robustness across misinformation, cyber, biosecurity, fraud, political content, and other sensitive domains.

  • Document vulnerabilities and contribute to safety benchmarking and red-teaming reports.

  • Collaborate with AI researchers to improve model alignment, robustness, and safety.

Required Qualifications
  • Bachelor's degree or higher in Computer Science, Cybersecurity, Journalism, Communications, Psychology, Biology, Chemistry, Public Policy, or a related discipline.

  • 5+ years of professional experience in AI Safety, AI Red Teaming, Trust & Safety, cybersecurity, investigative journalism, life sciences, or a related field.

  • Strong analytical reasoning, prompt design, and written communication skills.

  • Experience designing adversarial prompts or evaluating frontier AI systems.

Preferred Qualifications
  • Experience with AI Red Teaming, RLHF, SFT, AI Alignment, or Trust & Safety.

  • Familiarity with jailbreak testing, prompt engineering, or adversarial evaluation methodologies.

  • Expertise in one or more grey-area domains, including cyber, biosecurity, political content, misinformation, or scientific safety.

Why Join?
  • Help secure and strengthen the next generation of frontier AI models.

  • Work on cutting-edge adversarial testing alongside leading AI researchers and safety teams.

  • Influence how AI systems respond to complex, real-world safety challenges.

Ottieni la revisione del curriculum gratis e riservata.
o trascina qui il file.
Similar jobs

Offerte di lavoro simili che vale la pena confrontare

AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Obsidian • Roma

In loco
EUR 60.000 - 90.000
Remote AI Safety Red Teamer - Adversarial Prompting
Remote AI Safety Red Teamer - Adversarial Prompting

Mercor • Roma

In loco
EUR 84.000 - 100.000
Frontier AI Safety Evaluator
Frontier AI Safety Evaluator

Mercor • Roma

Remoto
EUR 70.000 - 100.000
Explosives & Energetic Materials Expert for AI Safety Audit
Explosives & Energetic Materials Expert for AI Safety Audit

SME Careers • Italia

In loco
EUR 70.000 - 120.000
Freelance Cybersecurity Analyst - AI Trainer
Freelance Cybersecurity Analyst - AI Trainer

Mindrift • Roma

Remoto
Competitive pay rates up to $34/hour
Flexible freelance work hours
Experience on advanced AI projects
Remote AI Ethics & Safety Analyst
Remote AI Ethics & Safety Analyst

Alignerr • Italia

In loco
EUR 24.000 - 72.000
Fully remote
Flexible schedule
Freelance autonomy
+2
Research Engineer Position on Secure Agentic AI Systems
Research Engineer Position on Secure Agentic AI Systems

AI4I • Collegno

In loco
EUR 45.000 - 75.000
Competitive compensation
Full support for conference travel
Professional development opportunities
+2
AI Security Advisor (Secure-by-Design) — Remote
AI Security Advisor (Secure-by-Design) — Remote

World Food Programme • Roma

Ibrido
EUR 78.000 - 123.000
Remote C++ AI Backend Engineer (Contractor)
Remote C++ AI Backend Engineer (Contractor)

YO IT Consulting • Roma

In loco
EUR 45.000 - 65.000
Remote AI Security Advisor — Secure by Design
Remote AI Security Advisor — Secure by Design

Logistics Cluster • Roma

Ibrido
EUR 65.000 - 90.000