LLM Red-Team Specialist for Adversarial Evaluation

Mercor

New York (NY)

On-site

USD 90,000 - 150,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Remote work within the United States

Job summary

Mercor is seeking a research-focused professional to join a GenAI red-team within a leading AI lab network. The role involves probing frontier models, designing robust evaluation tasks, and documenting findings for reproducibility.

This full-time W-2 position is remote within the United States, requiring about 35 hours per week and close collaboration with researchers to strengthen benchmark tasks and defenses against failure modes in AI systems.

Qualifications

  • Master's or PhD in a STEM field, or equivalent research-focused experience requiring data analysis and coding.
  • 1+ years of experience in a research, research-engineering, security, or AI-evaluation role.
  • Demonstrated ability to identify vulnerabilities, edge cases, or failure modes in LLMs or ML systems via red teaming or evaluation.
  • Proficiency in Python and Git, with ability to script probes and analyses.
  • Strong familiarity with LLM capabilities and evaluation techniques.
  • Experience in AI training, model evaluation, or benchmark/task authoring is preferred.
  • Attention to detail, creative problem-solving, strong written communication, able to work independently in ambiguity.
  • Ability to commit approx 35 hours per week.

Responsibilities

  • Probe frontier AI models on coding, ML, and analysis tasks to find failure modes.
  • Design challenging, fair tasks to grade model weaknesses.
  • Document findings clearly with reproducible steps.
  • Collaborate with researchers to improve benchmark tasks.

Skills

Python
Git
LLM evaluation
Adversarial testing
Data analysis

Education

MSc or PhD in a STEM field or equivalent

Job description

Mercor is seeking a research-focused professional to join a GenAI red-team within a leading AI lab network. The role involves probing frontier models, designing robust evaluation tasks, and documenting findings for reproducibility.

This full-time W-2 position is remote within the United States, requiring about 35 hours per week and close collaboration with researchers to strengthen benchmark tasks and defenses against failure modes in AI systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Red Team Engineer — Remote, 35h/wk
GenAI Red Team Engineer — Remote, 35h/wk

Mercor • San Francisco (CA)

On-site
USD 100,000 - 140,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • San Francisco (CA)

Remote
USD 120,000 - 180,000
Remote Adversarial ML Specialist — AI Safety Red Team
Remote Adversarial ML Specialist — AI Safety Red Team

Obsidian • San Francisco (CA)

Remote
USD 120,000 - 180,000
Remote work
AI Safety Red Team Specialist (Remote)
AI Safety Red Team Specialist (Remote)

Mercor • San Francisco (CA)

On-site
USD 90,000 - 130,000
Remote AI Adversarial Red Team Specialist
Remote AI Adversarial Red Team Specialist

Mercor • New York (NY)

Remote
USD 90,000 - 140,000
Remote AI Safety Red Team Engineer
Remote AI Safety Red Team Engineer

Neon • United States

Remote
USD 120,000 - 190,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • San Francisco (CA)

On-site
USD 120,000 - 180,000
Remote AI Red Team Safety Specialist: Adversarial Testing
Remote AI Red Team Safety Specialist: Adversarial Testing

Obsidian • San Francisco (CA)

Remote
USD 90,000 - 130,000
Remote AI Safety Red Team Expert
Remote AI Safety Red Team Expert

Mercor • New York (NY)

On-site
USD 120,000 - 180,000
AI Red Team Specialist — Adversarial Testing (Remote)
AI Red Team Specialist — Adversarial Testing (Remote)

Mercor • San Francisco (CA)

Remote
USD 130,000 - 170,000