GenAI Vulnerability Researcher (Remote, 35h/w)

Obsidian

San Francisco (CA)

Remote

USD 130,000 - 160,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Obsidian is seeking a researcher to join its GenAI evaluation team, focusing on probing frontier models, identifying failure modes, and converting findings into stronger benchmark tasks. This remote role (about 35 hours/week) operates under Cincinnatus LLC as the employer of record.

Responsibilities include designing tasks, documenting results with reproducible steps, and collaborating with researchers to enhance benchmark quality.

Qualifications

  • MSc or PhD in a STEM field or equivalent research experience requiring data analysis and coding.
  • 1+ years of experience in a research, research-engineering, security, or AI-evaluation role.
  • Demonstrated ability to identify vulnerabilities or failure modes in LLMs or ML systems through red teaming, adversarial testing, security research, or rigorous model evaluation.
  • Proficiency in Python and Git, with the ability to script probes and analyses.
  • Strong familiarity with LLM capabilities, limitations, and evaluation techniques.

Responsibilities

  • Probe frontier AI models to reveal how they behave on coding, ML, and analysis tasks and find where they quietly fail.
  • Design challenging tasks that are fair to grade and hard for models to exploit.
  • Document findings with clear evidence and reproducible steps for others to follow.
  • Collaborate with task authors to close loopholes, improve grading, and strengthen benchmarks.
  • Share insights with researchers to continuously improve benchmark quality.

Skills

Red teaming
Adversarial testing
Model evaluation
Data analysis
Experiment scripting

Education

MSc or PhD in STEM

Tools

Python
Git

Job description

Obsidian is seeking a researcher to join its GenAI evaluation team, focusing on probing frontier models, identifying failure modes, and converting findings into stronger benchmark tasks. This remote role (about 35 hours/week) operates under Cincinnatus LLC as the employer of record.

Responsibilities include designing tasks, documenting results with reproducible steps, and collaborating with researchers to enhance benchmark quality.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GenAI Vulnerability Engineer (Remote, 35h/wk)
GenAI Vulnerability Engineer (Remote, 35h/wk)

Mercor • New York (NY)

Remote
USD 120,000 - 170,000
GenAI Benchmark Research Scientist (Remote, Part-Time)
GenAI Benchmark Research Scientist (Remote, Part-Time)

Obsidian • San Francisco (CA)

Remote
USD 120,000 - 160,000
GenAI Benchmark Research Scientist (Remote, 35h/wk)
GenAI Benchmark Research Scientist (Remote, 35h/wk)

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 160,000
GenAI Benchmark Research Scientist - Remote (35h/wk)
GenAI Benchmark Research Scientist - Remote (35h/wk)

Mercor • San Francisco (CA)

Remote
USD 120,000 - 180,000
GenAI Evaluation Scientist (Remote, 35h/wk)
GenAI Evaluation Scientist (Remote, 35h/wk)

Mercor • New York (NY)

Remote
USD 90,000 - 120,000
Remote Data Scientist – GenAI Benchmark & Task Design
Remote Data Scientist – GenAI Benchmark & Task Design

Obsidian • New York (NY)

On-site
USD 90,000 - 150,000
GenAI Research & Evaluation Specialist
GenAI Research & Evaluation Specialist

Obsidian • New York (NY)

Remote
USD 100,000 - 130,000
GenAI Benchmark Research Scientist — Remote, Part-Time
GenAI Benchmark Research Scientist — Remote, Part-Time

Obsidian • New York (NY)

Remote
USD 120,000 - 150,000
GenAI Red Team Engineer — Remote, 35h/wk
GenAI Red Team Engineer — Remote, 35h/wk

Mercor • San Francisco (CA)

On-site
USD 100,000 - 140,000
Remote GenAI Benchmark Architect — Data Science
Remote GenAI Benchmark Architect — Data Science

Mercor • New York (NY)

On-site
USD 120,000 - 170,000