Frontier AI Red Team Specialist

Obsidian

New York (NY)

Remote

USD 90,000 - 130,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Cincinnatus LLC is seeking specialists to red-team frontier AI models, crafting multi-step tasks and benchmarks to reveal vulnerabilities and edge cases. You will work in a collaborative, research-focused environment, designing, testing, and documenting high-quality evaluation tasks for leading AI systems.

The role supports remote, full-time engagement within the United States, with a flexible weekly cadence around 35 hours and close coordination with researchers to iteratively improve

Qualifications

  • MSc or PhD in a STEM field or equivalent research experience.
  • 1+ years in research, research-engineering, security, or AI evaluation.
  • Ability to identify vulnerabilities or failure modes in LLMs/ML systems via red teaming or evaluation.
  • Proficiency in Python and Git for scripting probes and analyses.
  • Strong knowledge of LLM capabilities and evaluation techniques.
  • Experience in AI training, model evaluation, or benchmark authoring preferred.
  • Attention to detail and independent problem solving with clear written communication.
  • Available to commit ~35 hours per week.

Responsibilities

  • Probe frontier AI models to identify how they behave on coding, ML, and analysis tasks and where they quietly fail.
  • Design challenging tasks that are hard for models but fair to grade.
  • Document findings with clear evidence and reproducible steps.
  • Strengthen tasks by closing loopholes, shortcuts, and grading gaps in collaboration with authors.
  • Share insights with researchers and experts to improve benchmarks.

Skills

Python
Git
LLM evaluation
Research analysis
Adversarial testing
Written communication

Education

MSc/PhD in STEM

Job description

Cincinnatus LLC is seeking specialists to red-team frontier AI models, crafting multi-step tasks and benchmarks to reveal vulnerabilities and edge cases. You will work in a collaborative, research-focused environment, designing, testing, and documenting high-quality evaluation tasks for leading AI systems.

The role supports remote, full-time engagement within the United States, with a flexible weekly cadence around 35 hours and close coordination with researchers to iteratively improve

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote LLM Red Team Specialist - Frontier Model Vetting
Remote LLM Red Team Specialist - Frontier Model Vetting

Obsidian • New York (NY)

On-site
USD 110,000 - 170,000
Frontier AI Safety Red Team Expert
Frontier AI Safety Red Team Expert

Obsidian • New York (NY)

On-site
USD 170,000 - 260,000
Remote AI Safety Red Teamer for Frontier Models
Remote AI Safety Red Teamer for Frontier Models

Mercor • San Francisco (CA)

Remote
USD 170,000 - 260,000
Lead Frontier AI Safety Red Teamer
Lead Frontier AI Safety Red Teamer

Obsidian • San Francisco (CA)

On-site
USD 150,000 - 230,000
Frontier AI Safety Red Team Lead
Frontier AI Safety Red Team Lead

Mercor • New York (NY)

On-site
USD 150,000 - 190,000
Remote AI Red Team - Safety & Adversarial Testing
Remote AI Red Team - Safety & Adversarial Testing

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 140,000
Remote AI Red Team Specialist
Remote AI Red Team Specialist

Obsidian • San Francisco (CA)

Remote
USD 120,000 - 150,000
GenAI Benchmark Research Scientist (Remote, Part-Time)
GenAI Benchmark Research Scientist (Remote, Part-Time)

Obsidian • San Francisco (CA)

Remote
USD 120,000 - 160,000
AI Safety Red Team Specialist (Remote)
AI Safety Red Team Specialist (Remote)

Obsidian • New York (NY)

On-site
USD 90,000 - 130,000
Experience in human data-driven AI red teaming
Role in enhancing AI safety and trustworthiness
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • San Francisco (CA)

Remote
USD 120,000 - 180,000