AI Safety Red Teamer Expert

Mercor

Dublin

On-site

EUR 83,000 - 100,000

Full time

10 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Mercor is seeking an AI Safety Red Teamer for a contract engagement. You design adversarial prompts to stress-test frontier AI models, identify jailbreaks, unsafe behaviors, and policy failures, and document vulnerabilities for safety benchmarking.

In collaboration with AI researchers, you will help improve model alignment, robustness, and safety through rigorous evaluation across sensitive domains such as misinformation, cyber, biosecurity, and political content.

Qualifications

  • Bachelor-level or higher in a related field and 5+ years in AI safety or red teaming.
  • Proven ability to design and analyze adversarial prompts for frontier AI systems.
  • Strong written and verbal communication skills.

Responsibilities

  • Design adversarial prompts to stress-test frontier AI models.
  • Identify jailbreaks, unsafe behaviors, hallucinations, and policy failures.
  • Evaluate model robustness across misinformation, cyber, biosecurity, fraud, political content, and other domains.
  • Document vulnerabilities and contribute to safety benchmarking and red-teaming reports; collaborate with researchers.

Skills

Analytical reasoning
Prompt design
Written communication
Adversarial prompt design

Education

Bachelor's degree in Computer Science
Bachelor's degree in Cybersecurity
Bachelor's degree in Journalism
Bachelor's degree in Communications
Bachelor's degree in Psychology
Bachelor's degree in Biology
Bachelor's degree in Chemistry
Bachelor's degree in Public Policy

Job description

About the job

Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey .

Position: AI Safety Red Teamer
Type: Contract
Compensation: $70–$84/hour
Location: Remote

Role Responsibilities

Design adversarial prompts to stress-test frontier AI models . Identify jailbreaks, unsafe behaviors, hallucinations, and policy failures. Evaluate model robustness across misinformation, cyber, biosecurity, fraud, political content, and other sensitive domains. Document vulnerabilities and contribute to safety benchmarking and red-teaming reports. Collaborate with AI researchers to improve model alignment, robustness, and safety.

Qualifications
Must-Have
  • Bachelor's degree or higher in Computer Science , Cybersecurity , Journalism , Communications , Psychology , Biology , Chemistry , Public Policy , or a related discipline.
  • 5+ years of professional experience in AI Safety , AI Red Teaming , Trust & Safety , cybersecurity, investigative journalism, life sciences, or a related field.
  • Strong analytical reasoning, prompt design, and written communication skills.
  • Experience designing adversarial prompts or evaluating frontier AI systems .
Preferred
  • Experience with AI Red Teaming , RLHF , SFT , AI Alignment , or Trust & Safety .
  • Familiarity with jailbreak testing, prompt engineering, or adversarial evaluation methodologies.
  • Expertise in one or more grey‑area domains, including cyber, biosecurity, political content, misinformation, or scientific safety.
Get your free, confidential resume review.
or drag and drop your file here.