Research Engineer, Post-training & Reasoning

Cobalt

San Francisco (CA)

Hybrid

USD 150,000 - 230,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Founding-team equity
Salary discussion

Job summary

Cobalt is seeking a Research Engineer to advance post-training optimization for expert reasoning. This role welcomes candidates pursuing a PhD or Master’s and involves designing SFT, DPO, and RL experiments with cross-functional teams.

You will work with healthcare domain data to improve AI performance and interpretability, publish where appropriate, and engage directly with frontier labs and healthcare AI partners.

Qualifications

  • Strong ML engineering fundamentals.
  • Real exposure to post-training methods (SFT, preference optimization, RL fine-tuning).
  • A track record of shipping research or research-grade engineering.
  • Comfortable working with a part-time research lead.
  • Excited by applied work in a domain with real-world consequences (healthcare).

Responsibilities

  • Designing and running SFT, DPO, and RL experiments on reasoning traces.
  • Building benchmarks and evals for clinical and adjudication reasoning.
  • Turning raw expert outputs into high-quality training data and pipelines.
  • Collaborating with frontier labs and healthcare AI partners on bespoke data/eval engagements.
  • Publishing results where it makes sense.

Skills

ML engineering
Post-training methods
Research contributions
Team collaboration
Healthcare impact interest

Education

Pursuing PhD or Master’s

Tools

PyTorch
HuggingFace
vLLM
Deepspeed/FSDP

Job description

Company Description

Cobalt builds expert reasoning data infrastructure for AI. We work with credentialed domain experts, physicians, nurses, surgeons, payer Medical Directors, to capture how they actually reason through high-stakes decisions, and we turn them into training data, benchmarks, and evals for frontier AI labs and applied AI companies.

Role Description

This is a Research Engineer role focused on post-training and reasoning. Full-time or part-time; we’re open to candidates currently pursuing a PhD or Master’s. The responsibilities include conducting research in post-training optimization and reasoning techniques, developing innovative algorithms, and collaborating with cross-functional teams to apply findings to advanced AI systems. The role also involves analyzing complex datasets, enhancing AI models, and contributing to cutting-edge R&D projects aimed at optimizing AI performance and interpretability.

What we’re looking for:
  • Strong ML engineering fundamentals. Comfortable training and fine-tuning LLMs end-to-end (PyTorch, HF, vLLM, deepspeed/FSDP, or similar)
  • Real exposure to post-training methods (SFT, preference optimization, RL fine-tuning), not just having read the papers
  • A track record of shipping research or research-grade engineering: publications, strong open-source contributions, or production ML systems at a lab/frontier company
  • Comfortable working with a part-time research lead. You can take a direction and run, surface tradeoffs early, and don’t need someone in the room every day
  • Excited by applied work in a domain with real-world consequences (you don’t have to come from healthcare; you do have to care about it)
The work spans the full post-training stack as applied to expert reasoning:
  • Designing and running SFT, DPO, and RL (GRPO/PPO and successors) experiments on reasoning traces from our expert network
  • Building benchmarks and evals that meaningfully measure clinical and adjudication reasoning — not just final-answer accuracy, but the reasoning path
  • Turning raw expert outputs into high-quality training datasets: schema design, quality controls, scaling pipelines
  • Working directly with customers (frontier labs, healthcare AI companies) on bespoke data and eval engagements
  • Publishing where it makes sense
What we offer
  • Founding-team equity
  • Competitive salary (band depends on level; let’s talk)
  • Flexible work environment (hybrid in-person SF/NY WFH options)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Post-Training & Reasoning AI Engineer — Hybrid (Equity)
Post-Training & Reasoning AI Engineer — Hybrid (Equity)

Cobalt • San Francisco (CA)

Hybrid
USD 150,000 - 230,000
Founding-team equity
Salary discussion
Research Engineer - Midtraining
Research Engineer - Midtraining

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
AI Researcher - Frontiers & Reasoning
AI Researcher - Frontiers & Reasoning

Logical Intelligence • San Francisco (CA)

On-site
USD 150,000 - 210,000
Research, Post-Training
Research, Post-Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

On-site
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Research Scientist/Research Engineer, Midtraining
Research Scientist/Research Engineer, Midtraining

Periodic • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Research Scientist/Engineer
Research Scientist/Engineer

Periodic Labs • Menlo Park (CA)

On-site
USD 250,000 - 350,000
Research Engineer/Research Scientist, RL/Reasoning
Research Engineer/Research Scientist, RL/Reasoning

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 160,000
Relocation assistance
Hybrid work model
Commitment to accommodation for disabilities
AI Research Engineer
AI Research Engineer

TTN Talent • San Francisco (CA)

On-site
USD 200,000 - 400,000
RESEARCHER, POST-TRAINING
RESEARCHER, POST-TRAINING

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Research Scientist LLM
Senior Research Scientist LLM

techire ai • San Francisco (CA)

On-site
USD 350,000 - 500,000
Stock options
Remote work worldwide
Competitive compensation