Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Dormont Manufacturing Co is looking for a researcher skilled in empirical machine learning and dedicated to improving model behavior and interpretability. You will design experiments, investigate model monitorability, and collaborate with teams to enhance training interventions.
This role requires flexibility in tackling theoretical questions and translating them into measurable experiments. Ideal candidates will have a background in large ML models, and you'll enjoy a hybrid work model based in the United States with relocation assistance.
The CoT Monitorability team at OpenAI studies whether and when the chain-of-thought of frontier reasoning models is monitorable enough to support scalable oversight. We study how to measure monitorability, which training mechanisms affect monitorability, and speculative methods to improve monitorability. While we mostly focus on CoT monitorability at the moment, we care more generally about any form of monitorability, auditing methods, and improving alignment.
We were the first to show that chain-of-thought monitoring can be a practical additional safety mechanism, and today our monitoring systems are actively used on OpenAI’s largest RL training runs to detect misbehavior. The issues we surface are then used to help improve our reward functions, environments, etc (without directly training against a CoT monitor).
Our work sits in Alignment and intersects with model training, alignment evaluations, monitoring, and frontier-risk research. We care most about monitorability where the stakes are high, and about preserving useful oversight signals as models become more capable.
We’re looking for a researcher with strong empirical ML expertise and a deep interest in model behavior, alignment, or interpretability. Direct chain-of-thought interpretability experience is welcome but not required; strong candidates may come from broader interpretability, alignment, model training, or investigative model-behavior work.
As a researcher on the Alignment team, you will design and run experiments that improve our understanding of model monitorability. You will investigate how training interventions across the model-development pipeline influence whether reasoning remains legible, build evaluations that make those questions measurable, and help translate findings into practical oversight and training recommendations. You may also help develop new monitoring models or methods and apply them to OpenAI’s largest training runs.
This role is especially well suited for someone who can move from an ambiguous model-behavior question to a concrete experimental setup: formulate the hypothesis, build the evaluation or intervention, run the experiment, analyze the result, and decide what the evidence supports. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees.
Design and run empirical studies of chain-of-thought monitorability across frontier reasoning models and training settings.
Build evaluations that measure whether monitors can reliably predict properties of interest, including high-stakes forms of misbehavior.
Investigate how pre-training, synthetic data, mid-training, post-training, reinforcement learning, and other interventions improve or degrade monitorability.
Analyze model behavior and turn observations from monitoring into hypotheses, experiments, and recommendations.
Translate research findings into practical monitoring and oversight approaches that can inform real training runs.
Collaborate with researchers and engineers across model training, alignment evaluations, monitoring, and frontier-risk work.
Produce externally publishable research when results advance the broader science of alignment.
Have strong hands-on experience training, evaluating, or debugging large ML models, especially LLMs.
Have deep curiosity, interest in alignment, and high agency.
Bring depth in alignment, interpretability, model behavior, empirical ML, or adjacent research.
Are excited to investigate chain-of-thought monitorability, monitoring methods, and scalable oversight.
Can turn ambiguous research questions into measurable experiments and follow the evidence when results are subtle or noisy.
Move comfortably between research ideation and engineering execution.
Are curious about multiple approaches to understanding model behavior and are not committed to only one methodological lens.
Operate with high independence while collaborating closely across research and engineering teams.
Care about making increasingly capable AI systems more monitorable, trustworthy, and safe.
We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic.
Requests for reasonable accommodations can be made through this link.
Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates.
To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form. No response will be provided to inquiries unrelated to job posting compliance.
$250K - $445K Offers Equity
The base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience. If the role is non-exempt, overtime pay will be provided consistent with applicable laws. In addition to the salary range listed above, total compensation also includes generous equity, performance-related bonus(es) for eligible employees, and the following benefits.
Medical, dental, and vision insurance for you and your family, with employer contributions to Health Savings Accounts
Pre-tax accounts for Health FSA, Dependent Care FSA, and commuter expenses (parking and transit)
401(k) retirement plan with employer match
Paid parental leave (up to 24 weeks for birth parents and 20 weeks for non-birthing parents), plus paid medical and caregiver leave (up to 8 weeks)
Paid time off: flexible PTO for exempt employees and up to 15 days annually for non-exempt employees
13+ paid company holidays, and multiple paid coordinated company office closures throughout the year for focus and recharge, plus paid sick or safe time (1 hour per 30 hours worked, or more, as required by applicable state or local law)
Mental health and wellness support
Employer-paid basic life and disability coverage
Annual learning and development stipend to fuel your professional growth
Daily meals in our offices, and meal delivery credits as eligible
Relocation support for eligible employees
Additional taxable fringe benefits, such as charitable donation matching and wellness stipends, may also be provided.
More details about our benefits are available to candidates during the hiring process.
This role is at-will and OpenAI reserves the right to modify base pay and other compensation components at any time based on individual performance, team or company results, or market conditions.