Senior Member of Technical Staff - Model Safety

Xcede

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Xcede partners with a frontier AI research company to recruit a Member of Technical Staff focused on AI Safety. You will lead red-teaming, design scalable safety evaluation pipelines, and work at the intersection of model evaluation, alignment, and software engineering.

You will collaborate with alignment researchers, set safety thresholds for releases, and develop state-of-the-art jailbreaking defenses to stay ahead of threats.

Qualifications

  • Deep expertise in LLM safety, red teaming, adversarial attacks, interpretability, or model evaluation.
  • Strong software engineering skills and experience building large-scale ML systems or automated evaluation infrastructure.
  • Experience with AI alignment techniques, RLHF, RLAIF, or related safety methodologies is highly desirable.
  • MS, PhD, or equivalent practical experience in Machine Learning, AI Safety, Computer Science, or a related field.

Responsibilities

  • Lead red-teaming and adversarial evaluation efforts to uncover model vulnerabilities, misuse risks, and alignment gaps.
  • Design and build scalable safety evaluation frameworks and automated testing pipelines.
  • Partner closely with alignment researchers to translate safety findings into production guardrails.
  • Act as a key decision-maker in determining whether model releases meet safety thresholds.
  • Research and implement state-of-the-art jailbreaking techniques and defenses to stay ahead of emerging threats.
  • Develop dynamic safety benchmarks that evolve alongside increasingly capable AI systems.

Skills

LLM safety
Red teaming
Adversarial attacks
Model evaluation
Large-scale ML systems
AI alignment methods (RLHF/RLAIF)

Education

MS/PhD in ML/CS

Tools

Python
ML infra
CI/CD

Job description

We're partnering with a frontier AI research company on a search for a Member of Technical Staff focused on AI Safety.

The company is building next-generation open-weight foundation models with a mission to make advanced AI broadly accessible. Their team includes researchers, engineers, and operators from some of the world's leading AI labs and technology companies, working on the frontier of model capabilities, alignment, and deployment.

This is an opportunity to help define how advanced AI systems are evaluated, stress-tested, and safely deployed. You'll be working at the intersection of adversarial research, red teaming, model evaluation, alignment, and software engineering—helping ensure next-generation foundation models are robust, reliable, and secure before they reach the world.

Overview:
  • Lead red-teaming and adversarial evaluation efforts to uncover model vulnerabilities, misuse risks, and alignment gaps
  • Design and build scalable safety evaluation frameworks and automated testing pipelines
  • Partner closely with alignment researchers to translate safety findings into production guardrails
  • Act as a key decision-maker in determining whether model releases meet safety thresholds
  • Research and implement state-of-the‑art jailbreaking techniques and defenses to stay ahead of emerging threats
  • Develop dynamic safety benchmarks that evolve alongside increasingly capable AI systems
Skills required:
  • Deep expertise in LLM safety, red teaming, adversarial attacks, interpretability, or model evaluation
  • Strong software engineering skills and experience building large-scale ML systems or automated evaluation infrastructure
  • Experience with AI alignment techniques, RLHF, RLAIF, or related safety methodologies is highly desirable
  • MS, PhD, or equivalent practical experience in Machine Learning, AI Safety, Computer Science, or a related field
  • Ability to operate in a high-agency, fast-moving environment and make high-stakes technical decisions
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Safety Red Teamer Expert
AI Safety Red Teamer Expert

Obsidian • San Francisco (CA)

On-site
USD 150,000 - 230,000
AI Safety Red Teamer Expert
AI Safety Red Teamer Expert

Mercor • New York (NY)

On-site
USD 150,000 - 190,000
AI Safety Red Teamer Expert
AI Safety Red Teamer Expert

Obsidian • New York (NY)

On-site
USD 170,000 - 260,000
AI Safety & Red Team Lead - Senior Technical Staff
AI Safety & Red Team Lead - Senior Technical Staff

Xcede • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer, Safety
Machine Learning Engineer, Safety

Harrison Clarke • San Francisco (CA)

On-site
USD 190,000 - 275,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Obsidian • New York (NY)

On-site
USD 120,000 - 180,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 160,000
null
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Mercor • San Francisco (CA)

On-site
USD 150,000 - 190,000
AI Safety Specialist - Evaluation Expert
AI Safety Specialist - Evaluation Expert

Mercor • New York (NY)

On-site
USD 120,000 - 190,000
AI Safety Specialist - Remote | Upto $84/hr
AI Safety Specialist - Remote | Upto $84/hr

Obsidian • New York (NY)

Remote
USD 140,000 - 210,000