Mid-Training Research Engineer — Scientific LLMs

Doist

Menlo Park (CA)

On-site

USD 250,000 - 350,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Doist is seeking a Midtraining Research Engineer to advance scientific reasoning in frontier models. You will curate data, generate synthetic datasets, and build robust evals while enabling scalable, GPU-driven training experiments.

The role supports pre-training groundwork and requires hands-on distillation techniques and scalable tooling. You will collaborate with RL researchers, scientists, and engineers to push data-driven breakthroughs and improve model intelligence in science-focused

Qualifications

  • Bachelor's degree or equivalent experience required.
  • Experience training LLMs at scale is essential.
  • Ability to design and evaluate scientific reasoning data.

Responsibilities

  • Identify, process, and curate novel sources of scientific data for large-scale model training.
  • Generate high-quality synthetic data to fill gaps in scientific knowledge.
  • Build evaluations that correlate with downstream scientific task performance.
  • Develop and apply techniques such as self-distillation and on-policy distillation.
  • Design and run large-scale training experiments, partnering with supercompute engineers to scale across thousands of GPUs.
  • Build tools to study how data choices shape model intelligence.

Skills

LLM training
Data curation
Large-scale training
Eval design
Distributed training

Education

Bachelor's degree or equivalent

Tools

PyTorch
TensorFlow

Job description

Doist is seeking a Midtraining Research Engineer to advance scientific reasoning in frontier models. You will curate data, generate synthetic datasets, and build robust evals while enabling scalable, GPU-driven training experiments.

The role supports pre-training groundwork and requires hands-on distillation techniques and scalable tooling. You will collaborate with RL researchers, scientists, and engineers to push data-driven breakthroughs and improve model intelligence in science-focused

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer - Midtraining
Research Engineer - Midtraining

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Research Engineer - Midtraining
Research Engineer - Midtraining

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Research Engineer - Frontier Models for Scientific Discovery
Research Engineer - Frontier Models for Scientific Discovery

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Research Engineer: RL & Post-Training LLM Systems
Research Engineer: RL & Post-Training LLM Systems

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
Founding Engineer - Post-training LLM's
Founding Engineer - Post-training LLM's

CT19 • New York (NY)

On-site
USD 180,000 - 260,000
Member of Technical Staff - Post-Training Research
Member of Technical Staff - Post-Training Research

United States Digital Space LLC • New York (NY)

On-site
USD 180,000 - 240,000
Applied Post-Training LLM Research Scientist
Applied Post-Training LLM Research Scientist

Modal Labs • New York (NY)

On-site
USD 180,000 - 240,000
RL Research Scientist for LLM Post-Training & Code Models
RL Research Scientist for LLM Post-Training & Code Models

AMD • Santa Clara (CA)

On-site
USD 180,000 - 260,000
AMD benefits
Machine Learning Researcher – LLM
Machine Learning Researcher – LLM

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 180,000 - 280,000
Post-Training ML Engineer: Alignment & Evaluation
Post-Training ML Engineer: Alignment & Evaluation

Doist • San Francisco (CA)

On-site
USD 150,000 - 210,000