LLM Red Team Specialist

OpenTrain AI

Northern (KY)

Hybrid

USD 109,000 - 163,000

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

OpenTrain AI is seeking an LLM Red Team Specialist to support the development of benchmark challenges. You will probe large language models for coding, ML, and analysis tasks, turning weaknesses into difficult but fair tasks with clear reproducible steps.

You will work independently on ambiguous, open-ended problems while documenting evidence and contributing to a feedback loop that strengthens benchmark quality.

Qualifications

  • MSc or PhD in STEM (or equivalent practical research experience).
  • At least 1 year of AI evaluation, security, research engineering, or related role.
  • Demonstrated ability to surface vulnerabilities, edge cases, or failure modes in LLMs or ML systems.
  • Working proficiency in Python and Git.
  • Excellent written communication and ability to work independently on ambiguous problems.

Responsibilities

  • Probe LLMs on coding, ML, and analysis tasks to identify failures.
  • Convert weaknesses into well-crafted tasks that are fair to grade.
  • Document findings with clear evidence and reproducible steps.
  • Collaborate with task authors to close loopholes and grading gaps.
  • Contribute to a feedback loop for improving benchmark rigor.
  • Identify vulnerabilities and edge cases through red teaming or adversarial testing.

Skills

Python
Git
LLM evaluation
Adversarial testing
Written communication
Independent work

Education

MSc or PhD in STEM

Tools

Git
Python tooling

Job description

About OpenTrain

OpenTrain is the #1 platform for finding and building careers in AI training and data labeling. We help people discover specialized projects, build an AI training profile, and apply in minutes while contributing to the fast-growing field of human-led AI development.

OpenTrain AI is hiring and contracting for this role. Creating an OpenTrain account is free.

About AI Training and Model Evaluation

AI training is the human side of building artificial intelligence. Specialists evaluate model behavior, identify weaknesses, and create high-quality examples and feedback that help AI systems become more accurate, reliable, and useful.

Red teaming is a particularly challenging form of model evaluation: you deliberately probe systems for vulnerabilities, edge cases, shortcuts, and failures that ordinary testing may miss.

The Role

OpenTrain is seeking an LLM Red Team Specialist to support the development of next-generation agentic evaluation benchmarks. You will probe large language models on coding, machine learning, and analysis tasks, then turn discovered weaknesses into difficult but fair benchmark challenges.

You will work independently on ambiguous, open-ended problems while documenting evidence clearly and contributing to an ongoing feedback loop that strengthens benchmark quality.

  • Fully remote within the United States
  • Approximately 35 hours per week
  • Listing parameters indicate a commitment of 20+ hours per week
  • Pay: $60–$90 per hour, based on experience
  • English-language work
What You\'ll Do

Your work will combine adversarial testing, benchmark authoring, technical investigation, and precise written communication. Findings should be clear enough for researchers and collaborators to reproduce and act on.

  • Probe LLMs on coding, machine learning, and analysis tasks to identify subtle failures.
  • Convert model weaknesses into well-crafted tasks that are challenging for models but fair to grade.
  • Document findings with clear evidence and reproducible steps.
  • Collaborate with task authors to close loopholes, shortcuts, and grading gaps.
  • Contribute to a continuous feedback loop for improving benchmark rigor.
  • Identify vulnerabilities, edge cases, and failure modes through red teaming or adversarial testing.
Requirements

This role requires strong familiarity with how large language models work, where they fail, and how their performance can be evaluated. Equivalent practical research experience may substitute for the stated academic background.

  • MSc or PhD in a STEM field, or equivalent practical research experience
  • At least one year of experience in AI evaluation, security, research engineering, research, or a related role
  • Demonstrated ability to surface vulnerabilities, edge cases, or failure modes in LLMs or machine-learning systems
  • Working proficiency in Python and Git
  • Strong knowledge of LLM capabilities, limitations, and evaluation techniques
  • Excellent written communication
  • Ability to work independently on ambiguous, open-ended problems
Helpful Background

Experience in AI training, model evaluation, or benchmark and task authoring is helpful. The listing identifies the experience level as entry level, while the required skills call for demonstrated research, security, or AI-evaluation capability.

  • AI training experience
  • Model evaluation experience
  • Benchmark or task-authoring experience
  • Research-engineering or security experience
Why Join AI Training Work

AI training and evaluation work lets specialists help shape how cutting-edge models behave. Projects are often remote and flexible, making it possible to build experience in a rapidly expanding technology field while applying skills in research, coding, security, and analysis.

  • Contribute directly to the quality and reliability of advanced AI systems
  • Work remotely with a flexible weekly commitment
  • Apply research, programming, and adversarial-testing skills to frontier models
  • Build experience in a growing AI training and evaluation career path
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote LLM Red Team Specialist - Adversarial AI Benchmarking
Remote LLM Red Team Specialist - Adversarial AI Benchmarking

OpenTrain AI • Northern (KY)

Hybrid
USD 109,000 - 163,000
Market Research AI Training Expert
Market Research AI Training Expert

OpenTrain AI • Northern (KY)

Hybrid
USD 41,000 - 69,000
LLM Red Team Specialist — Failure Modes & Edge Cases
LLM Red Team Specialist — Failure Modes & Edge Cases

Dorado • United States

Remote
USD 120,000 - 180,000
Technical Sales AI Training Expert
Technical Sales AI Training Expert

OpenTrain AI • Northern (KY)

Hybrid
USD 34,000 - 55,000
AI QA Trainer - LLM Evaluation - Freelance Project
AI QA Trainer - LLM Evaluation - Freelance Project

Meridial • United States

Remote
Secure computer and high-speed internet required
Software Engineering AI Code Evaluator
Software Engineering AI Code Evaluator

OpenTrain AI • Northern (KY)

Hybrid
USD 55,000 - 83,000
Remote work
Flexible hours
AI training portfolio
Financial Analysis AI Evaluation Expert
Financial Analysis AI Evaluation Expert

OpenTrain AI • Northern (KY)

Hybrid
USD 100,289,000 - 157,597,000
Remote contract work
Part-time engagement
AI Vulnerability Expert - Fully Remote | Upto $90/hr
AI Vulnerability Expert - Fully Remote | Upto $90/hr

United States Digital Space LLC • United States

Remote
USD 83,000 - 124,000
Computer Systems Analyst, AI Training
Computer Systems Analyst, AI Training

OpenTrain AI • Northern (KY)

Hybrid
USD 83,000 - 165,000
Remote work
Part-time schedule
US-based eligibility
ML Research Engineer - PhD - AI Trainer
ML Research Engineer - PhD - AI Trainer

Obsidian • Seattle (WA)

Hybrid
USD 100,000 - 150,000