Research Lead - Pre-training Safety

FAR.AI

Berkeley (CA)

On-site

USD 180,000 - 240,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

FAR.AI, a Berkeley-based AI safety research institute, seeks a Research Lead to define and own a research program on pre-training safety and capability control. You will lead a growing team, mentor researchers, and push for empirical, scalable safety methods from data filtering to model governance.

You will collaborate with red-teaming partners and governments, drive research roadmaps, and publish results to advance the field while maintaining a strong hands-on coding and experimentation

Qualifications

  • PhD or equivalent in ML/CS.
  • Strong track record in security/safety research.
  • Experience leading research teams and mentoring staff.

Responsibilities

  • Articulate a clear research agenda with a theory of change for AI safety.
  • Grow and lead a team of technical staff in pursuit of the agenda.
  • Lead novel research projects with ambiguous milestones.
  • Share findings via publications, blogs, and presentations to drive adoption.
  • Mentor junior researchers and contribute to the research culture at FAR.AI.

Skills

Team Leadership
ML Safety Research
Mentor Staff
Experimentation
Code Proficiency

Education

PhD in ML/CS

Tools

Python
PyTorch

Job description

FAR.AI is hiring a Research Lead to develop and lead our work on pre-training safety, shaping models' capabilities and internal representations at their source, rather than trying to fix them after the fact.

Our initial focus is capability control: removing harmful capabilities while preserving benign ones. We see this as a promising way to prevent misuse of open-weight models in areas such as CBRN and cyber by removing offensive capabilities, and reducing loss-of-control risks by removing knowledge of oversight mechanisms. We will validate approaches like pre-training data filtering at scale, drive adoption of successful methods, and explore techniques such as gradient routing and unlearning..

We are scaling methods like ++Deep Ignorance++ by over an order of magnitude (>100B parameter models with >1T tokens). You will direct this work, partner with our red team to stress-test the resulting models, and analyze how well the methods scale to frontier systems.

Our research directions include:

  • Improved data filtering methods, such as using data attribution (e.g. influence-based selection) or more sophisticated classifiers

  • Using methods like gradient routing to isolate dual-use capabilities in components of the model (e.g. specific MoE experts)

  • Training to actively remove harmful capabilities, such as interleaving next-token prediction with unlearning, as opposed to simply filtering data

  • Adding synthetic data to pre-training or mid-training to shape the representations and behavior of the model

You’ll build and lead the team, set its research direction, mentor Members of Technical Staff to scale your vision, and remain hand-on enough to write code and run experiments yourself. This role offers high autonomy in an impact-driven environment, pursuing empirically grounded, scalable ML safety research.

About Us

FAR.AI is a non-profit AI research institute working to ensure advanced AI is safe and beneficial for everyone. Our mission is to facilitate breakthrough AI safety research, advance global understanding of AI risks and solutions, and foster a coordinated global response.

We’re structured to support that work from early research through real-world adoption:

Independent by design. We can pursue what’s most impactful based on our theory of change and share what we find publicly.

A portfolio approach. Rather than focus on one single direction, we run diverse bets across the safety stack. We take promising ideas from initial experiments to deployment, informed by red-team partnerships with frontier labs and governments.

Serious infrastructure for ambitious research. A dedicated engineering team runs our compute cluster and experiment-scaling stack, so researchers spend their time on research instead of on infra.

Setting the standard. Our events convene key decision makers; our red-team works with frontier developers and governments; and our communications inform the public. Together, this drives adoption and sets the new standard in safety.

Since our founding in July 2022, we’ve grown to ++50 staff++, published ++40 academic papers++, and convened leading ++AI safety events++. Our work is recognized globally, with publications at premier venues such as NeurIPS, ICML including a ++Best Paper Honorable Mention in 2026++, and ICLR, and features in the ++Financial Times++, ++Nature News++, ++Wired Magazine++ and ++MIT Technology Review++. We conduct pre-deployment testing on behalf of frontier developers such as OpenAI and independent evaluations for governments ++including the EU AI Office++ and publish the ++AI Security Leaderboard++ based on our red-teaming expertise. We help steer and grow the AI safety field through ++developing++ ++research++ ++roadmaps++ with renowned researchers such as Yoshua Bengio; running ++FAR.Labs++, an AI safety-focused co-working space in Berkeley housing 40 members; and supporting the community through ++targeted grants++ to technical researchers.

About FAR.Research

We explore promising research directions in AI safety and scale up only those showing a high potential for impact. When an approach proves effective, we develop it into a minimum viable demonstration and work with AI developers and governments to support real-world adoption.

Our recent and ongoing research includes:
  • Adversarial Robustness: working to rigorously solve security problems through building a science of security and robustness for AI, from ++demonstrating superhuman systems can be vulnerable++, to ++scaling laws for robustness++ and ++jailbreaking constitutional classifiers++.

  • Mechanistic Interpretability: ++finding++ ++issues++ ++with++ Sparse Autoencoders, probing deception using ++AmongUs++, understanding ++learned planning++ in SokoBan, and interpretable data attribution.

  • Red-teaming: conducting pre- and post-release adversarial evaluations of frontier models (e.g. ++Claude 4 Opus++, ++ChatGPT Agent++, ++GPT-5++); developing ++novel attacks++ to support this work.

  • Evals: developing evaluations for new threat models, e.g. ++persuasion++ and ++tampering risks++, and launching a new research agenda on eval awareness

  • Mitigating AI deception: studying when ++lie detectors induce honesty or evasion++, and developing ++approaches++ to deception and sandbagging.

  • Applied Interpretability: using interpretability to tackle concrete safety problems (better probes, backdoor detection, deception monitoring), aiming for fast feedback loops, often in collaboration with our other pods.

About the Role

Research Leads define and own a research workstream end-to-end. Day-to-day, that means:

  • Articulate a research agenda with a clear theory of change for mitigating catastrophic risks from human-level or superhuman AI systems, and/or vastly increasing the upside of such systems.

  • Grow and lead a team of technical staff in pursuit of this agenda, either directly or in partnership with an engineering co-lead.

  • Lead novel research projects where there may be unclear markers of progress or success.

  • Share your research findings through written content (e.g. academic publications, blog posts) and presentations (e.g. ML conferences, policymaker briefings) to drive adoption and change.

  • Mentor and coach junior team members in research skills and ML engineering.

  • Contribute to the FAR.AI intellectual environment and research culture, for example by giving feedback on early-stage proposals.

  • Build a research field around your agenda through FAR.AI’s grantmaking and events, and connect it to real-world deployments through our independent testing and government advising.

This role would be a great fit if you:
  • Want to work on the most impactful research directions, alongside mission-driven colleagues who’ll push them forward with you.

  • Wish to pursue empirically grounded, scalable research directions that lean, technically strong teams can drive forward.

  • Value the ability to speak freely. We don't censor our researchers. We just ask that you protect confidential information and make clear when you're speaking personally or on behalf of the organization.

  • Want to advise and collaborate with governments, leading AI companies, and academics. We’re a small organization that punches above its weight by working closely with these partners: through red-testing, technical standards work, and research collaborations.

This role would be a poor fit if you:
  • Prefer solo IC research to leading a team toward a shared agenda. Some people can do great research that way, but in this role we’re looking for someone whose research direction is strong enough that other excellent researchers want to build it with them.

  • Prioritize novelty and intellectual elegance over impact. We care about both --- a mathematically elegant solution to AI safety would be wonderful --- but when we have to choose, we choose what makes AI safer in practice.

  • Can only work with the largest compute clusters available at industry labs or need to be compensated with equity in a rapidly growing startup. We offer competitive salaries and sizable compute budgets on a cluster that we manage, but if you value these things over having a positive impact on the future, then you may be more suited to a for-profit lab.

About You

To be a strong candidate for the Research Lead -

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Lead
Research Lead

Aisafety • Berkeley (CA)

Hybrid
USD 170,000 - 270,000
Catered lunch and dinner
Visa sponsorship
Research Lead - Pre-training Safety
Research Lead - Pre-training Safety

FAR.AI • United States

On-site
USD 160,000 - 230,000
Research Lead, Pre-Training Safety
Research Lead, Pre-Training Safety

AISafety • Northern (KY)

Hybrid
USD 150,000 - 230,000
Technical Program Manager, Research
Technical Program Manager, Research

Aisafety • Berkeley (CA)

Hybrid
USD 125,000 - 190,000
Research Scientist
Research Scientist

Aisafety • Berkeley (CA)

Hybrid
USD 120,000 - 190,000
Catered lunch and dinner
Visa sponsorship for in-person employees
Reimbursement for work-related travel
Technical Program Manager, Research
Technical Program Manager, Research

FAR.AI • Berkeley (CA)

Hybrid
USD 125,000 - 190,000
Visa sponsorship for US work
Remote/U.S. location flexibility
Competitive compensation USD 125,000–$
Research Engineer
Research Engineer

Aisafety • Berkeley (CA)

Hybrid
USD 100,000 - 190,000
Catered lunch and dinner
Potential for remote work
Sponsorship for work-related travel
Senior Research Engineer
Senior Research Engineer

Aisafety • Berkeley (CA)

Hybrid
USD 150,000 - 250,000
Catered lunch and dinner
Visa sponsorship for in-person employees
Work-related travel expenses covered
Research Manager
Research Manager

Aisafety • San Francisco (CA)

On-site
USD 170,000 - 260,000
Health insurance
401K plan + 4% matching
Unlimited PTO
+2
Research Lead: AI Safety & Pre-Training at Scale
Research Lead: AI Safety & Pre-Training at Scale

FAR.AI • Berkeley (CA)

On-site
USD 180,000 - 240,000