Remote LLM Red Team Engineer: Edge Cases & Failures

Weekday 1

United States

Remote

USD 83,000 - 124,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Fully remote
Weekly payments

Job summary

Weekday 1 is seeking analytical and technically skilled professionals for a fully remote, 35-hour-per-week role focused on red-teaming frontier AI models. You will design challenging, multi-step tasks to expose hidden vulnerabilities and evaluation gaps in real-world scenarios.

Collaborate with AI researchers to convert failure modes into high-quality benchmark tasks, contributing to robust, safe, and reliable model evaluation.

Qualifications

  • Master's degree or PhD in a STEM field or equivalent practical experience.
  • Minimum 1 year of AI research or related technical work experience.
  • Experience identifying vulnerabilities and edge cases in LLMs or ML systems.
  • Proficient in Python and Git with ability to script experiments.

Responsibilities

  • Investigate frontier AI model performance across coding, ML, reasoning, and complex tasks.
  • Identify hidden failure modes and edge cases that standard tests miss.
  • Design evaluation tasks that are objective and reproducible.
  • Document findings with clear explanations and reproducible methods.
  • Collaborate with benchmark designers to refine tasks and grading criteria.
  • Share insights to improve AI evaluation quality and coverage.

Skills

Analytical thinking
Attention to detail
Independent problem solving
Strong written communication

Education

Master's degree or PhD in STEM

Tools

Python
Git

Job description

Weekday 1 is seeking analytical and technically skilled professionals for a fully remote, 35-hour-per-week role focused on red-teaming frontier AI models. You will design challenging, multi-step tasks to expose hidden vulnerabilities and evaluation gaps in real-world scenarios.

Collaborate with AI researchers to convert failure modes into high-quality benchmark tasks, contributing to robust, safe, and reliable model evaluation.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote LLM Safety & Red Teaming Expert
Remote LLM Safety & Red Teaming Expert

Weekday 1 • United States

Remote
USD 69,000 - 124,000
Remote LLM Red Team Specialist - Frontier Model Vetting
Remote LLM Red Team Specialist - Frontier Model Vetting

Obsidian • New York (NY)

On-site
USD 110,000 - 170,000
LLM Red Team Specialist - Failure Modes & Edge Cases
LLM Red Team Specialist - Failure Modes & Edge Cases

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments
GenAI Red Team Engineer — Remote, 35h/wk
GenAI Red Team Engineer — Remote, 35h/wk

Mercor • San Francisco (CA)

On-site
USD 100,000 - 140,000
LLM Red-Team Specialist for Adversarial Evaluation
LLM Red-Team Specialist for Adversarial Evaluation

Mercor • New York (NY)

On-site
USD 90,000 - 150,000
Remote work within the United States
Remote AI Safety & Red Teaming Expert
Remote AI Safety & Red Teaming Expert

Weekday AI (YC W21) • United States

On-site
USD 68,880 - 123,984
Remote AI Red Teamer - LLM Generalist (Stress-Test)
Remote AI Red Teamer - LLM Generalist (Stress-Test)

Apply • Northern (KY)

Hybrid
USD 120,000 - 180,000
Remote AI Safety Red Team Engineer
Remote AI Safety Red Team Engineer

Neon • United States

Remote
USD 120,000 - 190,000
Remote Part-Time LLM Security Red Team Engineer
Remote Part-Time LLM Security Red Team Engineer

OpenTrain AI, Inc. • United States

Remote
USD 47,000 - 63,000
Remote work
Worldwide eligibility
AI Safety Red Team Engineer (Remote • EN/DA)
AI Safety Red Team Engineer (Remote • EN/DA)

Neon • United States

Remote
USD 120,000 - 180,000