Anthropic Fellows Program — AI Safety

Socket.dev

San Francisco (CA)

Remote

CAD 200,000 - 224,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Weekly stipend and research funding
Mentorship from Anthropic researchers
Access to shared workspace (Berkeley/L

Job summary

Anthropic invites applicants to join the Fellows program, a 4-month, full-time research fellowship designed to advance AI safety research. Fellows will work on empirical projects using external infrastructure, mentored by distinguished researchers, with a weekly stipend and funds for compute.

This role is open to candidates located in the US, UK, or Canada, including remote arrangements. The program emphasizes collaboration, rigorous research, and producing a public output such as a paper

Qualifications

  • Fluent in Python programming.
  • Strong technical background in computer science, mathematics, or physics.
  • Ability to communicate ideas clearly and work well in a team.
  • Capable of moving quickly to implement ideas in empirical AI research.

Responsibilities

  • Work on empirical AI research using external infrastructure to produce a public output (e.g. a paper).
  • Receive direct mentorship from Anthropic researchers.
  • Participate in a 4-month, full-time fellowship with a weekly stipend and compute funding.

Skills

Python programming
Technical background CS/Math/Physics
Clear communication
Collaborative work

Job description

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

Anthropic Fellows Program overview

The Anthropic Fellows Program is designed to foster AI research and engineering talent. We provide funding and mentorship to promising technical talent - regardless of previous experience.

Fellows will primarily use external infrastructure (e.g. open-source models, public APIs) to work on an empirical project aligned with our research priorities, with the goal of producing a public output (e.g. a paper submission). In one of our earlier cohorts, over 80% of fellows produced papers.

We run multiple cohorts of Fellows each year and review applications on a rolling basis. This application is for cohorts starting in July 2026 and beyond.

What to expect
  • 4 months of full-time research
  • Direct mentorship from Anthropic researchers
  • Access to a shared workspace (in either Berkeley, California or London, UK)
  • Connection to the broader AI safety and security research community
  • Weekly stipend of 3,850 USD / 2,310 GBP / 4,300 CAD + benefits (these vary by country)
  • Funding for compute (~$15k/month) and other research expenses
Interview process

The interview process will include an initial application & reference check, technical assessments & interviews, and a research discussion.

We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team.

Compensation

The expected base stipend for this role is 3,850 USD / 2,310 GBP / 4,300 CAD per week, with an expectation of 40 hours per week for 4 months (with possible extension).

Fellows workstreams

Due to the success of the Anthropic Fellows for AI Safety Research program, we are now expanding it across teams at Anthropic. We expect there to be significant overlap in the types of skills and responsibilities across the roles and will by default consider candidates for all the workstreams.

Some of the workstreams may include unique assessment steps; we therefore ask you for workstream preferences in the application. You can see an overview of the current workstreams below:

  1. AI Safety Fellows
  2. AI Security Fellows
  3. ML Systems & Performance Fellows
  4. Reinforcement Learning Fellows
  5. Economics & Societal Impacts Fellows

This page is specific to one of the Anthropic Fellows Workstreams, see also the main Anthropic Fellows posting.

Across the workstreams, you may be a good fit if you:
  • Are motivated by making sure AI is safe and beneficial for society as a whole
  • Are excited to transition into empirical AI research and would be interested in a full-time role at Anthropic
  • Have a strong technical background in computer science, mathematics, or physics
  • Thrive in fast-paced, collaborative environments
  • Can implement ideas quickly and communicate clearly
Strong candidates may also have:
  • Strong background in a discipline relevant to a specific Fellows workstream (e.g. economics, social sciences, or cybersecurity)
  • Experience in areas of research or engineering related to their workstream
Candidates must be:
  • Fluent in Python programming
  • Available to work full-time on the Fellows program
AI Safety Fellows
Mentors, research areas, & past projects

Fellows will undergo a project selection & mentor matching process. Potential mentors include:

  • Sam Bowman
  • Sara Price
  • Alex Tamkin
  • Nina Panickssery
  • Trenton Bricken
  • Logan Graham
  • Jascha Sohl-Dickstein
  • Joe Benton
  • Collin Burns
  • Fabien Roger
  • Samuel Marks
  • Kyle Fish
  • Ethan Perez

Our mentors will lead projects in select AI safety research areas, such as:

  • Scalable Oversight: Developing techniques to keep highly capable models helpful and honest, even as they surpass human-level intelligence in various domains.
  • Adversarial Robustness and AI Control: Creating methods to ensure advanced AI systems remain safe and harmless in unfamiliar or adversarial scenarios.
  • Model Organisms: Creating model organisms of misalignment to improve our empirical understanding of how alignment failures might arise.
  • Model Internals / Mechanistic Interpretability: Advancing our understanding of the internal workings of large language models to enable more targeted interventions and safety measures.
  • AI Welfare: Improving our understanding of potential AI welfare and developing related evaluations and mitigations.

On our Alignment Science and Frontier Red Team blogs, you can read about past projects, including:

  • Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data: Alex Cloud and Minh Le, et al., mentors including Samuel Marks and Owain Evans
  • Open-source circuits: Michael Hanna and Mateusz Piotrowski with mentorship from Emmanuel Ameisen and Jack Lindsey

For a full list of representative projects for each area, please see these blog posts: Introducing the Anthropic Fellows Program for AI Safety Research, Recommendations for Technical AI Safety Research Directions.

Unique candidate criteria

You might be a particularly great fit for this workstream if you:

  • Are motivated by reducing catastrophic risks from advanced AI systems
  • Have experience with empirical ML research projects
  • Have experience working with large language models
  • Have experience in one of the research areas mentioned above
  • Have a track record of open-source contributions
Logistics

Logistics Requirements: To participate in the Fellows program, you must have work authorization in the US, UK, or Canada and be located in that country during the program.

Workspace Locations: We have designated shared workspaces in London and Berkeley where fellows will work from and mentors will visit. We are also open to remote fellows in the UK, US, or Canada. We will ask you about your availability to work from Berkeley or London (full- or part-time) during the program.

Visa Sponsorship: We are not currently able to sponsor visas for fellows. To participate in the Fellows program, you need to have or independently obtain full‑time work authorization in the UK, the US, or Canada.

Program Duration: The program runs for 4 months, full‑time. If you can't commit to the full duration, please still apply and note your constraints in the application. We review these requests on a case‑by‑case basis.

Please note: We do not guarantee that we will make any full‑time offers to fellows. However, strong performance during the program may indicate that a Fellow would be a good fit for full‑time roles at Anthropic. In previous cohorts, 25-50% of fellows received a full‑time offer, and we’ve supported many more to go on to do great work on AI safety and security at other organizations.

How we're different

We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long‑term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We’re an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest‑impact work at any given time. As such, we greatly value communication skills.

The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.

Come work with us!

Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Anthropic Fellows Program, ML Systems & Reinforcement Learning
Anthropic Fellows Program, ML Systems & Reinforcement Learning

Nerdleveltech • Northern (KY)

Hybrid
USD 200,000 - 210,000
Workspace in Berkeley/London
Remote-friendly
Anthropic Fellows Program — ML Systems & Performance London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA
Anthropic Fellows Program — ML Systems & Performance London, UK; Ontario, CAN; Remote-Friendly, United States; San Francisco, CA

Anthropic Limited • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 220,000
Mentorship from Anthropic researchers
Weekly stipend
Shared workspace
+1
Anthropic Fellows Program
Anthropic Fellows Program

Heelsandtech • Berkeley (CA)

On-site
Weekly stipend
Funding for research expenses
Mentorship from AI experts
Technical Program Manager, Safeguards (Infrastructure & Evals)
Technical Program Manager, Safeguards (Infrastructure & Evals)

Anthropic • San Francisco (CA)

On-site
USD 290,000 - 365,000
Competitive salary
Generous vacation
Parental leave
+1
Technical Program Manager (Infrastructure)
Technical Program Manager (Infrastructure)

Anthropic • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
Generous vacation
Flexible working hours
+1
Staff+ Software Security Engineer
Staff+ Software Security Engineer

Anthropic • San Francisco (CA)

On-site
USD 405,000 - 485,000
Competitive salary
Flexible working hours
Generous vacation and parental leave
Software Engineer, Research Tools
Software Engineer, Research Tools

Anthropic • San Francisco (CA)

Hybrid
USD 300,000 - 405,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+1
Machine Learning Infrastructure Engineer, Safeguards Research
Machine Learning Infrastructure Engineer, Safeguards Research

Anthropic • New York (NY)

Hybrid
USD 350,000 - 500,000
Equity donation matching
Vacation and parental leave
Flexible working hours
+1
Staff Research Engineer, Discovery Team
Staff Research Engineer, Discovery Team

Anthropic • New York (NY)

Hybrid
USD 350,000 - 850,000
Equity donation matching
Generous vacation
Parental leave
+2
Research Engineer, Machine Learning (Reinforcement Learning)
Research Engineer, Machine Learning (Reinforcement Learning)

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+2