LLM Red Team Specialist - Failure Modes & Edge Cases

Weekday 1

United States

Remote

USD 83,000 - 124,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Fully remote
Weekly payments

Job summary

Weekday 1 is seeking analytical and technically skilled professionals for a fully remote, 35-hour-per-week role focused on red-teaming frontier AI models. You will design challenging, multi-step tasks to expose hidden vulnerabilities and evaluation gaps in real-world scenarios.

Collaborate with AI researchers to convert failure modes into high-quality benchmark tasks, contributing to robust, safe, and reliable model evaluation.

Qualifications

  • Master's degree or PhD in a STEM field or equivalent practical experience.
  • Minimum 1 year of AI research or related technical work experience.
  • Experience identifying vulnerabilities and edge cases in LLMs or ML systems.
  • Proficient in Python and Git with ability to script experiments.

Responsibilities

  • Investigate frontier AI model performance across coding, ML, reasoning, and complex tasks.
  • Identify hidden failure modes and edge cases that standard tests miss.
  • Design evaluation tasks that are objective and reproducible.
  • Document findings with clear explanations and reproducible methods.
  • Collaborate with benchmark designers to refine tasks and grading criteria.
  • Share insights to improve AI evaluation quality and coverage.

Skills

Analytical thinking
Attention to detail
Independent problem solving
Strong written communication

Education

Master's degree or PhD in STEM

Tools

Python
Git

Job description

This role is for one of our clients

Compensation: $60-$90 per hour

Join a pioneering AI initiative focused on building next-generation evaluation benchmarks for frontier AI models. We are seeking analytical and technically skilled professionals to identify where advanced AI systems fail in subtle, real-world scenarios. Working in a red-teaming environment, you will design challenging, multi-step tasks that expose hidden vulnerabilities, reasoning gaps, and edge cases that traditional evaluations often miss.

In this role, you'll collaborate closely with AI researchers to transform discovered failure modes into high-quality benchmark tasks that improve the robustness, safety, and reasoning capabilities of state-of-the-art AI systems.

This is a fully remote, full-time engagement requiring approximately 35 hours per week.

Requirements

Key Responsibilities
  • Investigate how frontier AI models perform across coding, machine learning, analytical reasoning, and complex problem-solving tasks.
  • Identify hidden failure modes, edge cases, reasoning errors, and vulnerabilities that may not be apparent through standard testing.
  • Design challenging evaluation tasks that accurately measure AI capabilities while remaining objective and reproducible.
  • Document findings with clear technical explanations, supporting evidence, and reproducible methodologies.
  • Collaborate with benchmark designers and AI researchers to refine evaluation tasks, eliminate loopholes, and strengthen grading criteria.
  • Share insights and recommendations with cross-functional teams to continuously improve AI evaluation quality and benchmark coverage.
Required Qualifications
  • Master's degree, PhD, or equivalent practical experience in a STEM discipline involving research, coding, or advanced data analysis.
  • Minimum 1 year of experience in AI research, research engineering, security research, AI evaluation, or a related technical field.
  • Demonstrated experience identifying vulnerabilities, adversarial behaviors, edge cases, or failure modes in Large Language Models or other machine learning systems.
  • Strong proficiency in Python and Git, with the ability to build custom scripts for experimentation, testing, and analysis.
  • Solid understanding of modern Large Language Models, their strengths, limitations, and evaluation methodologies.
  • Experience with AI benchmarking, model evaluation, adversarial testing, prompt engineering, or dataset creation is highly desirable.
  • Excellent analytical thinking, creativity, and attention to detail, with the ability to solve ambiguous, open-ended problems independently.
  • Outstanding written communication skills for documenting technical findings clearly and accurately.
  • Ability to commit approximately 35 hours per week on a consistent basis.
Preferred Qualifications
  • Experience with AI safety, red teaming, adversarial machine learning, or security research.
  • Background in benchmark design, evaluation framework development, or AI quality assurance.
  • Experience creating reproducible technical experiments and documenting complex failure analyses.
  • Familiarity with frontier AI research methodologies and model capability assessments.
Why Join
  • Help shape the future of AI evaluation by identifying critical weaknesses before they reach production.
  • Work on cutting-edge AI systems alongside researchers developing next-generation language models.
  • Apply your technical expertise to improve AI reliability, reasoning, and robustness.
  • Contribute directly to benchmark development that influences the evolution of advanced AI technologies.
  • Enjoy the flexibility of a fully remote engagement while working on impactful research initiatives.
Equal Opportunity

We are committed to fostering an inclusive and diverse environment where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.

Contract & Engagement Details
  • Independent contractor engagement.
  • Fully remote with flexible working hours.
  • Expected commitment of approximately 35 hours per week.
  • Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
  • Work does not require access to confidential or proprietary information from any current or former employer.
  • Payments are issued weekly based on approved work completed.
  • At this time, we are unable to support H1-B or STEM OPT candidates.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineering Expert
Software Engineering Expert

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Fully remote
Weekly payments
STEM Researcher - Computational Fields
STEM Researcher - Computational Fields

Weekday 1 • United States

Remote
USD 83,000 - 124,000
Machine Learning & NLP Expert
Machine Learning & NLP Expert

Weekday 1 • United States

Remote
USD 110,000 - 152,000
AI Evaluation Specialist
AI Evaluation Specialist

Weekday 1 • United States

Remote
USD 80,000 - 113,000
Data Science & Quantitative Analysis Expert
Data Science & Quantitative Analysis Expert

Weekday 1 • United States

Remote
USD 109,000 - 164,000
Fully remote
Weekly payments
AI Safety & Red Teaming Specialist
AI Safety & Red Teaming Specialist

Weekday 1 • United States

Remote
USD 69,000 - 124,000
Remote | Research Engineer - Code Generation & Model Evaluation — $50–$100/hour
Remote | Research Engineer - Code Generation & Model Evaluation — $50–$100/hour

24-MAG • United States

Remote
USD 69,000 - 138,000
Remote work
Flexible hours
Contractor engagement
Technology AI Evaluation Expert
Technology AI Evaluation Expert

Weekday 1 • United States

Remote
USD 21,000 - 28,000
Fully remote
Flexible hours
Weekly payments
AI Rater Guidelines Writer (Linguist / Instructional Designer)
AI Rater Guidelines Writer (Linguist / Instructional Designer)

Weekday 1 • United States

Remote
USD 62,000 - 90,000
Fully remote engagement
AI Red Team Engineer for LLMs
AI Red Team Engineer for LLMs

OpenTrain AI, Inc. • United States

Remote
USD 47,000 - 63,000
Remote work
Worldwide eligibility