Senior ML Infrastructure Engineer — Frontier RL & LLM Training

Preference Model

San Francisco (CA)

On-site

USD 200,000 - 350,000

Full time

46 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health, vision, dental benefits
401K match
Lunch provided onsite
Weekly snack orders
Visa sponsorship & relocation support

Job summary

Preference Model in San Francisco seeks a Senior ML Infrastructure Engineer to design, build, and scale the infrastructure that powers post-training research on in-house RL environments. You will develop core ML framework primitives and tooling to accelerate reproducible experimentation and reduce time from idea to result.

Work closely with Research Engineers to translate research needs into scalable infra while tackling distributed systems challenges, cloud platforms, and high-throughput

Qualifications

  • Proven experience building production-grade ML infrastructure.
  • Hands-on with distributed training and high-throughput systems.
  • Solid experience with cloud platforms and container orchestration.

Responsibilities

  • Design, build, and scale compute, scheduling, and data infra for post-training research.
  • Develop core ML framework primitives and internal tooling for reproducible experiments.
  • Create evaluation, logging, and deployment tooling to catch failures early.
  • Collaborate with Research Engineers to translate needs into scalable infra

Skills

LLM inference infrastructure
Distributed systems
Kubernetes
AWS/GCP
PyTorch/JAX
RL training frameworks

Tools

vLLM
Megatron
SGLang
Slime
veRL
Ray Train
SkyRL

Job description

Preference Model in San Francisco seeks a Senior ML Infrastructure Engineer to design, build, and scale the infrastructure that powers post-training research on in-house RL environments. You will develop core ML framework primitives and tooling to accelerate reproducible experimentation and reduce time from idea to result.

Work closely with Research Engineers to translate research needs into scalable infra while tackling distributed systems challenges, cloud platforms, and high-throughput

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Platform Engineer — Scale Research ML Infra
Senior ML Platform Engineer — Scale Research ML Infra

techire ai • San Francisco (CA)

On-site
USD 270,000 - 330,000
Stock options
Senior ML Engineer: RL Environments for Frontier Models
Senior ML Engineer: RL Environments for Frontier Models

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive cash and equity compensation
Health, vision, and dental benefits
401K match
+2
Member of Technical Staff - ML Infrastructure Engineer, Post-training
Member of Technical Staff - ML Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 200,000 - 350,000
Health, vision, dental benefits
401K match
Lunch provided onsite
+2
Staff ML Engineer — RL & Production Systems
Staff ML Engineer — RL & Production Systems

People In AI • San Francisco (CA)

Hybrid
USD 270,000 - 280,000
ML Systems Engineer — RL Training & Finetuning
ML Systems Engineer — RL Training & Finetuning

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+1
RL Infrastructure Engineer - Scale Distributed Training
RL Infrastructure Engineer - Scale Distributed Training

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Engineering Manager, ML Infrastructure & Systems
Engineering Manager, ML Infrastructure & Systems

Cursor • San Francisco (CA)

On-site
USD 180,000 - 260,000
ML Infrastructure Engineer: Scalable Training and Deployment
ML Infrastructure Engineer: Scalable Training and Deployment

Epsilon • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Production ML Engineer — Scale & Train LLMs
Production ML Engineer — Scale & Train LLMs

Anthropic • San Francisco (CA)

On-site
USD 350,000 - 850,000
Generous vacation and parental leave
Flexible working hours
Office space for collaboration
Research Infra Engineer for AI & RL Systems
Research Infra Engineer for AI & RL Systems

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 300,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Paid parental leave
+1