Research Lead, Pre-Training Safety

AISafety

Northern (KY)

Hybrid

USD 150,000 - 230,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

FAR.AI is a non-profit AI research institute focused on safety and beneficial AI. The Research Lead will develop and steer work on pre-training safety, shaping models’ capabilities and internal representations at their source.

Initial focus includes capability control to remove harmful capabilities while preserving benign ones, with validation at scale and potential applications to prevent misuse of open-weight models.

Responsibilities

  • You’ll build and lead the team, set its research direction, mentor Members of Technical Staff to scale your vision, and remain hands-on enough to write code and run experiments yourself.
  • This role offers high autonomy in an impact-driven environment, pursuing empirically grounded, scalable ML safety research.
  • Focus on pre-training safety, data filtering at scale, and responsible deployment strategies.

Job description

FAR.AI is hiring a Research Lead to develop and lead our work on pre-training safety, shaping models’ capabilities and internal representations at their source, rather than trying to fix them after the fact.

Our initial focus is capability control: removing harmful capabilities while preserving benign ones. We see this as a promising way to prevent misuse of open-weight models in areas such as CBRN and cyber by removing offensive capabilities, and reducing loss-of-control risks by removing knowledge of oversight mechanisms. We will validate approaches like pre-training data filtering at scale, drive adoption of successful methods, and explore techniques such as gradient routing and unlearning..

We are scaling methods like Deep Ignorance by over an order of magnitude (>100B parameter models with >1T tokens). You will direct this work, partner with our red team to stress-test the resulting models, and analyze how well the methods scale to frontier systems.

Our research directions include:

  • Improved data filtering methods, such as using data attribution (e.g. influence-based selection) or more sophisticated classifiers

  • Using methods like gradient routing to isolate dual-use capabilities in components of the model (e.g. specific MoE experts)

  • Training to actively remove harmful capabilities, such as interleaving next-token prediction with unlearning, as opposed to simply filtering data

  • Adding synthetic data to pre-training or mid-training to shape the representations and behavior of the model

You’ll build and lead the team, set its research direction, mentor Members of Technical Staff to scale your vision, and remain hands-on enough to write code and run experiments yourself. This role offers high autonomy in an impact-driven environment, pursuing empirically grounded, scalable ML safety research.

About Us

FAR.AI is a non-profit AI research institute working to ensure advanced AI is safe and beneficial for everyone. Our mission is to facilitate breakthrough AI safety research, advance global understanding of AI risks and solutions, and foster a coordinated global response.

We’re structured to support that work from early research through real-world adoption:

Independent by design. We can pursue what's most impactful based on our theory of change and share what we find publicly.

A portfolio approach. Rather than focus on one single direction, we run diverse bets across the safety stack. We take promising ideas from initial experiments to deployment, informed by red-team partnerships with frontier labs and governments.

Serious infrastructure for ambitious research. A dedicated engineering team runs our compute

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Lead - Pre-training Safety
Research Lead - Pre-training Safety

FAR.AI • United States

On-site
USD 160,000 - 230,000
Research Lead, AI Safety & Pre-Training
Research Lead, AI Safety & Pre-Training

AISafety • Northern (KY)

Hybrid
USD 150,000 - 230,000
Research Lead: AI Safety & Pre-Training at Scale
Research Lead: AI Safety & Pre-Training at Scale

FAR.AI • Berkeley (CA)

On-site
USD 180,000 - 240,000
Research Lead: Pre-Training Safety & Safe AI
Research Lead: Pre-Training Safety & Safe AI

FAR.AI • United States

On-site
USD 160,000 - 230,000
Research Lead - Pre-training Safety
Research Lead - Pre-training Safety

FAR.AI • Berkeley (CA)

On-site
USD 180,000 - 240,000
Senior Research Engineer
Senior Research Engineer

Aisafety • Berkeley (CA)

Hybrid
USD 150,000 - 250,000
Catered lunch and dinner
Visa sponsorship for in-person employees
Work-related travel expenses covered
Research Lead
Research Lead

Aisafety • Berkeley (CA)

Hybrid
USD 170,000 - 270,000
Catered lunch and dinner
Visa sponsorship
Research Scientist
Research Scientist

Aisafety • Berkeley (CA)

Hybrid
USD 120,000 - 190,000
Catered lunch and dinner
Visa sponsorship for in-person employees
Reimbursement for work-related travel
Research Engineer
Research Engineer

Aisafety • Berkeley (CA)

Hybrid
USD 100,000 - 190,000
Catered lunch and dinner
Potential for remote work
Sponsorship for work-related travel
Technical Program Manager, Research
Technical Program Manager, Research

Aisafety • Berkeley (CA)

Hybrid
USD 125,000 - 190,000