Post-Training ML Engineer: Alignment & Evaluation

Doist

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Doist is seeking an accomplished ML engineer to join Devin's post-training and alignment efforts. This role blends research and hands-on engineering to shape how ambitious AI agents learn, evaluate, and interact with humans in long-horizon tasks.

You will develop post-training recipes, design meaningful evaluations, and advance techniques like RLHF, RLAIF, and constitutional approaches. The team values rigorous experimentation, systems thinking, and rapid iteration in a world-class AI lab

Qualifications

  • Post-training experience in ML systems and alignment
  • Ability to design and run end-to-end evaluation pipelines
  • Experience with large-scale distributed training is a plus
  • Strong fundamentals in probability, statistics, and ML theory
  • Ability to interpret experimental data and distinguish signals from noise

Responsibilities

  • Develop post-training recipes across datasets, training stages, and hyperparameters.
  • Design evaluations that capture meaningful signals and progress.
  • Investigate results and diagnose why training behaves unexpectedly.
  • Apply RLHF, RLAIF, and constitutional methods to agent behavior.
  • Scale experiments with data and compute, exploring new methodologies.

Skills

RLHF
RLAIF
Reward learning
Probability
Statistics
ML theory

Job description

Doist is seeking an accomplished ML engineer to join Devin's post-training and alignment efforts. This role blends research and hands-on engineering to shape how ambitious AI agents learn, evaluate, and interact with humans in long-horizon tasks.

You will develop post-training recipes, design meaningful evaluations, and advance techniques like RLHF, RLAIF, and constitutional approaches. The team values rigorous experimentation, systems thinking, and rapid iteration in a world-class AI lab

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, Post-Training
Research Engineer, Post-Training

cognition • San Francisco (CA)

On-site
USD 150,000 - 210,000
Mid-Training Research Engineer — Scientific LLMs
Mid-Training Research Engineer — Scientific LLMs

Doist • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Staff Engineer - AI Systems & Post-Training RL
Staff Engineer - AI Systems & Post-Training RL

Cohere • New York (NY)

On-site
USD 150,000 - 230,000
Lunch stipend and health benefits
Dental benefits
RRSP/401K/Pension
+6
Post-Training AI Researcher — Enterprise LLM Alignment
Post-Training AI Researcher — Enterprise LLM Alignment

Distyl • New York (NY), San Francisco (CA)

Hybrid
USD 150,000 - 250,000
Equity
Medical insurance
Flexible time off
+6
Senior ML Engineer: AI Safety & Alignment (RLHF)
Senior ML Engineer: AI Safety & Alignment (RLHF)

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 400,000
Top-tier salary and equity grants
Comprehensive medical, dental, and eye
Research Engineer – Experimental ML Systems
Research Engineer – Experimental ML Systems

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 180,000
Research Engineer: RL & Post-Training LLM Systems
Research Engineer: RL & Post-Training LLM Systems

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
ML Systems Engineer — RL Training & Finetuning
ML Systems Engineer — RL Training & Finetuning

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+1
Research, Post-Training
Research, Post-Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
ML Research Engineer - PhD - AI Trainer
ML Research Engineer - PhD - AI Trainer

Obsidian • Seattle (WA)

Hybrid
USD 100,000 - 150,000