Remote LLM Red Team Specialist - Frontier Model Vetting

Obsidian

New York (NY)

On-site

USD 110,000 - 170,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Cincinnatus LLC is building capabilities to stress-test frontier AI models. You will work in a red-teaming setup to design and probe multi-step tasks that reveal vulnerabilities, edge cases, and failure modes in cutting-edge AI systems, with approximately 35 hours per week and fully remote work within the United States.

This is a full-time W-2 role, where you collaborate with researchers to turn findings into stronger benchmark tasks, document evidence and reproducible steps, and help the

Qualifications

  • MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy domain requiring data analysis and coding.
  • 1+ years of experience in a research, research-engineering, security, or AI-evaluation role.
  • Demonstrated ability to identify vulnerabilities, edge cases, or failure modes in LLMs or ML systems.
  • Working proficiency in Python and Git, with the ability to script your own probes and analyses.
  • Strong familiarity with LLM capabilities, limitations, and evaluation techniques.
  • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.
  • A perfectionist mindset: high attention to detail, creativity in finding what others missed, strong written communication, and the ability to work independently through ambiguous, open-ended problems.
  • Ability to engage reliably for approximately 35 hours per week.

Responsibilities

  • Probe models: Explore frontier AI models on coding, ML, and analysis tasks, finding spots where they quietly err.
  • Design challenges: Turn weaknesses into well-crafted tasks that are hard for models but fair to grade.
  • Document findings: Write up evidence and steps that others can reproduce.
  • Strengthen tasks: Collaborate with authors to close loopholes, shortcuts, and grading gaps.
  • Work as a team: Share insights with researchers and experts to improve benchmarks.

Skills

Python
Git
LLM evaluation
Red-teaming
Research proficiency

Education

MSc or PhD in STEM

Tools

Python tooling

Job description

Cincinnatus LLC is building capabilities to stress-test frontier AI models. You will work in a red-teaming setup to design and probe multi-step tasks that reveal vulnerabilities, edge cases, and failure modes in cutting-edge AI systems, with approximately 35 hours per week and fully remote work within the United States.

This is a full-time W-2 role, where you collaborate with researchers to turn findings into stronger benchmark tasks, document evidence and reproducible steps, and help the

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Frontier AI Red Team Specialist
Frontier AI Red Team Specialist

Obsidian • New York (NY)

Remote
USD 90,000 - 130,000
Remote AI Safety Red Teamer for Frontier Models
Remote AI Safety Red Teamer for Frontier Models

Mercor • San Francisco (CA)

Remote
USD 170,000 - 260,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • San Francisco (CA)

On-site
USD 120,000 - 180,000
LLM Red-Team Specialist for Adversarial Evaluation
LLM Red-Team Specialist for Adversarial Evaluation

Mercor • New York (NY)

On-site
USD 90,000 - 150,000
Remote work within the United States
Lead Frontier AI Safety Red Teamer
Lead Frontier AI Safety Red Teamer

Obsidian • San Francisco (CA)

On-site
USD 150,000 - 230,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • New York (NY)

On-site
USD 90,000 - 140,000
Remote AI Red Team - Safety & Adversarial Testing
Remote AI Red Team - Safety & Adversarial Testing

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 140,000
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • San Francisco (CA)

Remote
USD 120,000 - 180,000
GenAI Benchmark Research Scientist — Remote, Part-Time
GenAI Benchmark Research Scientist — Remote, Part-Time

Obsidian • New York (NY)

Remote
USD 120,000 - 150,000
GenAI Red Team Engineer — Remote, 35h/wk
GenAI Red Team Engineer — Remote, 35h/wk

Mercor • San Francisco (CA)

On-site
USD 100,000 - 140,000