Research Manager, Biological Safety

EngineersOfAI

San Francisco, Northern (CA, KY)

Hybrid

USD 320,000 - 480,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Anthropic is seeking a hands-on manager for its Safeguards organization to lead research engineering focused on biological safety evaluations and classifiers. You will grow a team of researchers and engineers, set technical direction, and ensure rigorous evaluation of models in production traffic.

You will review eval designs, understand classifier failures, and communicate progress to research, product, and policy partners while maintaining technical depth.

Qualifications

  • Experience managing a technical team, including hiring, coaching, and performance management.
  • A record of setting technical direction for a team and making prioritization calls under uncertainty.
  • Proficiency in Python, with a background in scientific programming and data analysis.
  • Knowledge of modern biology across both measurement and engineering: high-throughput assay.

Responsibilities

  • Manage, coach, and grow a team of research scientists and engineers working on biological safety evaluations and classifiers, including hiring, onboarding, performance, and career development.
  • Set the technical direction and roadmap for the biological safety research agenda, and make the calls about what the team builds, what it deprioritizes, and when a safeguard is ready to ship.
  • Own the quality of capability evaluations that assess what new models can do in the biological domain, and turn results into deployment recommendations that leadership can act on.
  • Guide the development of training and evaluation datasets for our safety classifiers, working with internal and external threat modeling experts to ground them in realistic risk.
  • Oversee the training and iteration of safety classifiers alongside ML engineers, optimizing jointly for adversarial robustness and low false-positive rates.
  • Ensure the team invests in the tooling and pipelines that make evaluation and classifier development fast and repeatable.
  • Establish how the team measures classifier and eval performance against production traffic, identifies gaps, and prioritizes improvements.
  • Direct red-teaming and stress-testing of safeguards as threats, models, and product surfaces evolve.
  • Partner with Research, Product, Policy, and government affairs colleagues to embed biological safety throughout the model development lifecycle, and serve as an escalation point for biological content.
  • Represent the team's work in external communications including model cards, blog posts, and policy documents
  • Track developments in biology, machine learning, and biosecurity for their potential to create new risks or enable new mitigations

Skills

Python
ML fundamentals
Biology knowledge

Job description

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the role

Anthropic's Safeguards organization builds the policies, evaluations, and enforcement systems that keep our models from contributing to catastrophic harm. We are hiring a manager to lead the research engineering team responsible for biological safety: the evaluations, datasets, and classifiers that govern how our models handle biological knowledge.

You will lead a team of research scientists and engineers who design and run capability evaluations against frontier models, curate training data for our safety classifiers, train and iterate on those classifiers alongside our ML engineers, and measure how they hold up against adversarial pressure in production traffic. You will set the technical direction for that work, decide where the team invests, and own the results.

This is a hands-on management role. Most of your time goes to growing and directing the team, but you will keep enough technical depth to review an eval design, interrogate a classifier's failure modes, and represent the work credibly to Research, Product, and Policy partners.

The core tension your team owns is precision: safeguards need to be robust against sophisticated actors while staying out of the way of the far larger population of legitimate researchers using Claude to accelerate life sciences work. Getting that tradeoff right is an empirical problem, and your team is the one measuring it.

Key responsibilities
  • Manage, coach, and grow a team of research scientists and engineers working on biological safety evaluations and classifiers, including hiring, onboarding, performance, and career development

  • Set the technical direction and roadmap for the biological safety research agenda, and make the calls about what the team builds, what it deprioritizes, and when a safeguard is ready to ship

  • Own the quality of capability evaluations that assess what new models can do in the biological domain, and turn results into deployment recommendations that leadership can act on

  • Guide the development of training and evaluation datasets for our safety classifiers, working with internal and external threat modeling experts to ground them in realistic risk

  • Oversee the training and iteration of safety classifiers alongside ML engineers, optimizing jointly for adversarial robustness and low false-positive rates

  • Ensure the team invests in the tooling and pipelines that make evaluation and classifier development fast and repeatable

  • Establish how the team measures classifier and eval performance against production traffic, identifies gaps, and prioritizes improvements

  • Direct red-teaming and stress-testing of safeguards as threats, models, and product surfaces evolve

  • Partner with Research, Product, Policy, and government affairs colleagues to embed biological safety throughout the model development lifecycle, and serve as an escalation point for biological content

  • Represent the team's work in external communications including model cards, blog posts, and policy documents

  • Track developments in biology, machine learning, and biosecurity for their potential to create new risks or enable new mitigations

Minimum qualifications
  • Experience managing a technical team, including hiring, coaching, and performance management

  • A record of setting technical direction for a team and making prioritization calls under uncertainty

  • Proficiency in Python, with a background in scientific programming and data analysis

  • A solid grasp of ML fundamentals, sufficient to critically review evaluation design and classifier development

  • Knowledge of modern biology across both measurement and engineering: high-throughput assay

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Manager, Biological Safety
Research Manager, Biological Safety

anthropic • San Francisco (CA)

On-site
USD 405,000 - 485,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Biological Safety Research Manager
Biological Safety Research Manager

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 480,000
Biological Safety Research Manager
Biological Safety Research Manager

anthropic • San Francisco (CA)

Hybrid
USD 405,000 - 485,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Staff+ Site Reliability Engineer, Safeguards ML Infra
Staff+ Site Reliability Engineer, Safeguards ML Infra

Socket.dev • San Francisco (CA), Northern (KY)

Hybrid
USD 291,000 - 485,000
Staff+ Software Engineer, ML Inference Path
Staff+ Software Engineer, ML Inference Path

Socket.dev • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Equity donation matching
Generous vacation
Parental leave
+2
Product Manager, Safeguards (Cyber)
Product Manager, Safeguards (Cyber)

Alex Loftus • San Francisco (CA), Northern (KY)

Hybrid
USD 305,000 - 385,000
Safeguards Enforcement Analyst, Safety Evaluations
Safeguards Enforcement Analyst, Safety Evaluations

Anthropic • New York (NY)

Hybrid
USD 230,000 - 270,000
Competitive compensation
Equity donation matching
Generous vacation
+3
Research Scientist, Life Sciences
Research Scientist, Life Sciences

Menlo Ventures • San Francisco (CA)

On-site
USD 300,000 - 320,000
Staff+ Site Reliability Engineer, Safeguards ML Infra
Staff+ Site Reliability Engineer, Safeguards ML Infra

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
ML/Research Engineer, Safeguards
ML/Research Engineer, Safeguards

Anthropic • New York (NY)

Hybrid
USD 350,000 - 500,000