RESEARCHER, POST-TRAINING

MakerMaker.AI

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MakerMaker.AI in San Francisco is seeking an experienced research lead to focus on model post-training efforts. The ideal candidate will drive research efforts on supervised fine-tuning, reinforcement learning, and evaluation models, while collaborating closely with engineering teams.

This role requires a strong foundation in machine learning research, with a minimum of 5 years of experience, especially in post-training methodologies. Familiarity with PyTorch and excellent data curation skills are essential.

Qualifications

  • 5+ years of hands-on ML research experience.
  • Strong track record of post-training research.
  • Experience designing evaluation suites.

Responsibilities

  • Lead post-training research and design curation data.
  • Build and maintain evaluation suites.
  • Run rigorous experiments and write internal findings.
  • Identify and address failure modes in models.

Skills

Post-training research
Statistical analysis
Large-scale data curation
Communication skills
Experience with PyTorch

Education

PhD in ML, statistics, CS, or adjacent

Job description

ABOUT THE COMPANY

We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site.

ABOUT THE ROLE

You'll lead our work on model post-training: supervised fine-tuning, preference data, reinforcement learning from human and AI feedback, reward modeling, and the evaluation suites that tell us what's actually working. You'll own a research area that meaningfully shapes our model behavior and capability.

This is a hands‑on senior research role. You'll set direction, run experiments, and ship into production. You'll partner with the data, infrastructure, and engineering teams to make the post‑training pipeline reliable and fast: improvements there compound into every model we ship.

WHAT YOU'LL DO
  • Lead post‑training research: SFT, RLHF/RLAIF, RLVR, DPO and successor methods, reward modeling, preference data design
  • Design and curate the data that goes into post‑training (from sourcing, to filtering, to quality assessment)
  • Build and maintain the evaluation suites that measure what matters; resist Goodharting your own benchmarks
  • Run rigorous experiments (controls, ablations, statistical significance) and write up internal findings clearly
  • Scale data pipelines and the infrastructure team to scale training
  • Identify and characterize failure modes (reward hacking, distribution drift, eval saturation) and design experiments to address them
  • Stay current on the post‑training literature; bring useful methods in, ignore the noise
WHAT WE'RE LOOKING FOR
  • Strong track record of post‑training research (SFT, RL, reward modeling) at a frontier‑model lab or equivalent
  • 5+ years of hands‑on ML research experience
  • Comfort with large‑scale data curation and preference‑data pipelines
  • Experience designing evaluation suites for capabilities that aren't easily benchmarked
  • Fluent in PyTorch or equivalent; comfortable at the scale of distributed training
  • Strong statistical instincts: you'd notice a flawed comparison before someone else points it out
  • Strong written communication
NICE TO HAVE
  • PhD in ML, statistics, CS, or adjacent
  • Published research at NeurIPS, ICML, ICLR, COLM, RLC, or comparable venues
  • Experience with reward hacking detection, scaling reward models, or RLHF infrastructure
  • Synthetic data generation experience
  • Background in RL math (policy gradients, importance sampling, off‑policy methods)
  • Open‑source contributions to post‑training infrastructure
THIS ROLE IS PROBABLY NOT FOR YOU IF
  • You are primarily interested in pretraining (that's a different role)
  • You would rather invent novel methods in isolation than ship them into a model that real users run
  • You prefer benchmarks that are stable to evaluation work where the right answer isn't yet defined
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Competitive cash and equity compensation (>90th percentile)
Ownership and autonomy
Health, vision, dental benefits
+4
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
Research, Post-Training Data
Research, Post-Training Data

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
RESEARCHER (GENERAL)
RESEARCHER (GENERAL)

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Research Scientist, Post-Training
Research Scientist, Post-Training

David Joseph & Company • San Francisco (CA)

On-site
USD 150,000 - 450,000
Research, Post-Training
Research, Post-Training

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental and vision benefits
Unlimited PTO
+2
Research Engineer
Research Engineer

ThirdLayer, Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
Research Scientist: Post-Training
Research Scientist: Post-Training

Generalist • Somerville (MA), San Mateo (CA)

On-site
USD 100,000 - 130,000
Member of Technical Staff, Post-training
Member of Technical Staff, Post-training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
Member of Technical Staff, Research — Early Career(PHD)
Member of Technical Staff, Research — Early Career(PHD)

Goaly • Menlo Park (CA)

Hybrid
USD 140,000 - 230,000
Meals and office benefits
Visa sponsorship