GenAI Red Team Engineer — Remote, 35h/wk

Mercor

San Francisco (CA)

On-site

USD 100,000 - 140,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Mercor is hiring for a full-time remote role within the United States to probe frontier AI models, find weaknesses, and design robust evaluation tasks. You'll work with researchers to produce reproducible findings and strengthen benchmarks in a collaborative team.

The role requires 1+ years in an AI evaluation or related research field and advanced credentials (MSc/PhD in a STEM field). Strong Python and Git skills, plus knowledge of LLM capabilities, are essential.

Qualifications

  • MSc or PhD in a STEM field or equivalent research experience.
  • 1+ year in a research, research-engineering, security, or AI-evaluation role.
  • Proven ability to identify vulnerabilities, edge cases, or failure modes in LLMs/ML systems.

Responsibilities

  • Probe frontier AI models to identify weaknesses in coding, ML, and analysis tasks.
  • Design challenging, fair-to-grade tasks that expose model gaps.
  • Document findings with clear evidence and reproducible steps.

Skills

Python
Git
LLM evaluation
Adversarial testing
Data analysis
Written communication

Education

MSc or PhD in STEM

Job description

Mercor is hiring for a full-time remote role within the United States to probe frontier AI models, find weaknesses, and design robust evaluation tasks. You'll work with researchers to produce reproducible findings and strengthen benchmarks in a collaborative team.

The role requires 1+ years in an AI evaluation or related research field and advanced credentials (MSc/PhD in a STEM field). Strong Python and Git skills, plus knowledge of LLM capabilities, are essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Red-Team Specialist for Adversarial Evaluation
LLM Red-Team Specialist for Adversarial Evaluation

Mercor • New York (NY)

On-site
USD 90,000 - 150,000
Remote work within the United States
GenAI Vulnerability Engineer (Remote, 35h/wk)
GenAI Vulnerability Engineer (Remote, 35h/wk)

Mercor • New York (NY)

Remote
USD 120,000 - 170,000
GenAI Benchmark Research Scientist - Remote (35h/wk)
GenAI Benchmark Research Scientist - Remote (35h/wk)

Mercor • San Francisco (CA)

Remote
USD 120,000 - 180,000
Remote GenAI Benchmark Architect — Data Science
Remote GenAI Benchmark Architect — Data Science

Mercor • New York (NY)

On-site
USD 120,000 - 170,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • San Francisco (CA)

Remote
USD 120,000 - 180,000
AI Safety Red Team Specialist (Remote)
AI Safety Red Team Specialist (Remote)

Mercor • San Francisco (CA)

On-site
USD 90,000 - 130,000
AI Safety Red Team Engineer (Remote • EN/DA)
AI Safety Red Team Engineer (Remote • EN/DA)

Neon • United States

Remote
USD 120,000 - 180,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • San Francisco (CA)

On-site
USD 120,000 - 180,000
Remote AI Safety Red Team Engineer
Remote AI Safety Red Team Engineer

Neon • United States

Remote
USD 120,000 - 190,000
Remote AI Safety Red Team Expert
Remote AI Safety Red Team Expert

Mercor • New York (NY)

On-site
USD 120,000 - 180,000