Open-Model AI Safety Lead: Red Team & Validation | Equity

Reflection AI

New York (NY)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Top-tier compensation
Comprehensive medical, dental, and vision insurance
Fully paid parental leave
Paid time off
Daily lunch and dinner provided

Job summary

A leading AI research firm in New York, NY, is looking for a specialist to oversee the red-teaming and adversarial evaluation pipeline for their models. The ideal candidate should possess a graduate degree in Computer Science or a related field and have a deep understanding of LLM safety and adversarial techniques. A strong software engineering background is essential, along with the ability to make high-stakes decisions regarding model safety. This role offers top-tier compensation and comprehensive benefits.

Qualifications

  • Strong technical understanding of LLM safety and adversarial attacks.
  • Experience in building automated evaluation pipelines.
  • Ability to make high-stakes decisions regarding model safety.

Responsibilities

  • Own the red-teaming evaluation pipeline for models.
  • Work with Alignment team on safety findings.
  • Validate model releases meet risk thresholds.
  • Develop automated safety benchmarks.
  • Implement state-of-the-art jailbreaking techniques.

Skills

LLM safety expertise
Software engineering skills
Adversarial attack knowledge
Reinforcement Learning (RLHF/RLAIF) understanding

Education

Graduate degree in Computer Science or related discipline

Job description

A leading AI research firm in New York, NY, is looking for a specialist to oversee the red-teaming and adversarial evaluation pipeline for their models. The ideal candidate should possess a graduate degree in Computer Science or a related field and have a deep understanding of LLM safety and adversarial techniques. A strong software engineering background is essential, along with the ability to make high-stakes decisions regarding model safety. This role offers top-tier compensation and comprehensive benefits.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff AI Safety Engineer - Red Team & Guardrails
Staff AI Safety Engineer - Red Team & Guardrails

B Capital • San Francisco (CA)

On-site
USD 120,000 - 160,000
Top-tier compensation
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave
+2
Cyber Red Team Specialist: AI Safety & Adversary Testing
Cyber Red Team Specialist: AI Safety & Adversary Testing

OpenAI • Washington

Hybrid
USD 180,000 - 280,000
Relocation assistance
Hybrid work model
Remote Frontier AI Red Team Analyst – Adversarial Safety
Remote Frontier AI Red Team Analyst – Adversarial Safety

Crossing Hurdles • United States

On-site
USD 70,000 - 90,000
AI Safety Red Team Specialist (Remote)
AI Safety Red Team Specialist (Remote)

Obsidian • New York (NY)

On-site
USD 90,000 - 130,000
Experience in human data-driven AI red teaming
Role in enhancing AI safety and trustworthiness
Remote AI Red Team - Safety & Adversarial Testing
Remote AI Red Team - Safety & Adversarial Testing

Obsidian • San Francisco (CA)

On-site
USD 100,000 - 140,000
Remote AI Safety & Red Teaming Expert
Remote AI Safety & Red Teaming Expert

Weekday AI (YC W21) • United States

On-site
Senior AI Red Team Scientist for Safety & Evaluation
Senior AI Red Team Scientist for Safety & Evaluation

SupportFinity™ • New York (NY)

On-site
USD 190,000 - 211,000
401(k) plan
Bonus program
Equity opportunity
Remote AI Safety Red Team Specialist
Remote AI Safety Red Team Specialist

Mercor • New York (NY)

On-site
USD 110,000 - 180,000
Remote AI Safety Red Team Specialist (EN/SE)
Remote AI Safety Red Team Specialist (EN/SE)

Mercor • New York (NY)

On-site
USD 120,000 - 180,000
Remote AI Safety Red Team Expert (Adversarial ML)
Remote AI Safety Red Team Expert (Adversarial ML)

Mercor • New York (NY)

Remote
USD 110,000 - 170,000