AI Safety Researcher: Safety, Evaluation & Red-Teaming

Thinking Machines Lab Inc.

San Francisco (CA)

On-site

USD 350,000 - 475,000

Full time

27 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Parental leave
Relocation support

Job summary

Thinking Machines Lab Inc. in San Francisco is seeking a safety researcher to bridge research and hands-on engineering, focusing on making models safe and trustworthy.

You will explore how training shapes refusals and how to evaluate boundaries, designing experiments to inform model training and evaluation strategies.

Join a team that values rigorous analysis, data-driven evaluation, and responsible AI practices, with opportunities to influence real-world deployments.

Qualifications

  • Bachelor’s degree or equivalent in Computer Science, Machine Learning, Physics, Mathematics, or related field with strong theoretical and empirical grounding.
  • Background in AI safety research with hands-on experience in at least one safety area (RLHF/RLAIF, alignment, red-teaming, etc.).
  • Proficiency in Python and familiarity with DL frameworks; comfortable debugging distributed training.

Responsibilities

  • Work across data curation, safety-focused fine-tuning, evaluations, and red-teaming across the development stack.
  • Design experiments to understand how training shapes model refusals and safe behavior.
  • Develop safety evaluations and synthetic data to test boundaries and risk surfaces.
  • Collaborate with teams to implement mitigations for safety failures and jailbreaks.

Skills

Python programming
Deep learning frameworks
Communication clarity
AI safety research background

Education

Bachelor’s degree or equivalent in CS/ML/Physics/Math
PhD in CS/ML/Physics/Math optional

Tools

PyTorch
TensorFlow
JAX

Job description

Thinking Machines Lab Inc. in San Francisco is seeking a safety researcher to bridge research and hands-on engineering, focusing on making models safe and trustworthy.

You will explore how training shapes refusals and how to evaluate boundaries, designing experiments to inform model training and evaluation strategies.

Join a team that values rigorous analysis, data-driven evaluation, and responsible AI practices, with opportunities to influence real-world deployments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Safety Researcher — Investigate Model Behavior & Impact
AI Safety Researcher — Investigate Model Behavior & Impact

Typesafe AI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research, Safety
Research, Safety

Thinking Machines Lab Inc. • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
AI Safety Researcher (6-Week) — Evaluate & Document
AI Safety Researcher (6-Week) — Evaluate & Document

TypeSafe AI • San Francisco (CA)

On-site
USD 150,000 - 190,000
AI Safety Research Scientist
AI Safety Research Scientist

Center for AI Safety • San Francisco (CA)

On-site
USD 140,000 - 200,000
Health Insurance for you and dependets
401K with 4% matching
Unlimited PTO
+3
Frontier-Model Safety Researcher & Evaluations
Frontier-Model Safety Researcher & Evaluations

OpenAI • San Francisco (CA)

Hybrid
USD 380,000 - 500,000
Relocation assistance
Hybrid work model
AI Safety Engineer — Research to Production Leader
AI Safety Engineer — Research to Production Leader

Abundant • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 450,000
Health insurance
Dental insurance
Vision insurance
+1
Senior Member of Technical Staff - Model Safety
Senior Member of Technical Staff - Model Safety

Xcede • San Francisco (CA)

On-site
USD 180,000 - 240,000
Agent Safety Researcher: Training & Evaluation
Agent Safety Researcher: Training & Evaluation

CHEManager International • San Francisco (CA)

Hybrid
USD 190,000 - 240,000
Researcher, AI Agent Safety & Controls
Researcher, AI Agent Safety & Controls

CHEManager International • San Francisco (CA)

Hybrid
USD 180,000 - 260,000
Relocation assistance
Hybrid work model
Senior AI Red Team Scientist for Safety & Evaluation
Senior AI Red Team Scientist for Safety & Evaluation

SupportFinity™ • New York (NY)

On-site
USD 190,000 - 211,000
401(k) plan
Bonus program
Equity opportunity