AI Safety Researcher: Alignment, Evaluation & Red Teaming

Thinking Machines Lab Inc.

San Francisco (CA)

On-site

USD 350,000 - 475,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Visa sponsorship
Unlimited PTO
Parental leave
Relocation assistance
Comprehensive benefits

Job summary

Thinking Machines Lab Inc. seeks a safety researcher in San Francisco to advance trustworthy AI through research and hands-on work. You’ll explore how models refuse harmful requests, design experiments, and guide training and evaluation strategies.

Join a multidisciplinary team working across data filtering, post-training techniques, safety evaluations, synthetic data generation, and red-team activities to surface risks and inform mitigations, with strong emphasis on empirical grounding.

Qualifications

  • Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.
  • Background in AI safety research, with hands‑on experience in at least one area of safety, such as RLHF/RLAIF, alignment and preference modeling, deliberative alignment, safety evaluations, or red‑team­ing.

Responsibilities

  • Build data filtering pipelines and quality classifiers to shape what models learn from pre‑training corpora, and study how those early interventions affect downstream safety behavior.
  • Apply post‑training techniques, including RL from human and AI feedback and policy‑based reasoning approaches, to shape how models handle harmful, sensitive, and dual‑use requests.
  • Design, build, and maintain safety evaluations, with particular focus on measuring model behavior on long‑horizon and agentic tasks.
  • Generate and curate synthetic data to train and evaluate models on refusal boundaries and safety‑relevant behaviors.
  • Red‑team our models and products to surface failure modes, jailbreaks, and emergent risks before deployment, and design mitigations for what you find.

Skills

Python
PyTorch
TensorFlow
JAX
Communication

Education

Bachelor's degree or equivalent
PhD in related field

Job description

Thinking Machines Lab Inc. seeks a safety researcher in San Francisco to advance trustworthy AI through research and hands-on work. You’ll explore how models refuse harmful requests, design experiments, and guide training and evaluation strategies.

Join a multidisciplinary team working across data filtering, post-training techniques, safety evaluations, synthetic data generation, and red-team activities to surface risks and inform mitigations, with strong emphasis on empirical grounding.

Get your free, confidential resume review.

or drag and drop your file here.