Senior ML Infrastructure Engineer — High-Throughput AI Research

Preference Model

San Francisco (CA)

On-site

USD 180,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Cash and equity
Ownership & autonomy
Visa sponsorship
Relocation support
Lunch onsite
Weekly snacks
401K match
Health benefits

Job summary

Preference Model is seeking a Senior ML Infrastructure Engineer to build the systems powering post-training research on large language models. You will design and scale the compute, scheduling, and data infrastructure for in‑house RL environments and develop core ML framework primitives used by researchers daily.

You will partner with Research Engineers to translate research needs into infrastructure requirements, ensuring reproducible experiments and fast iteration.

Qualifications

  • Strong software engineering fundamentals and production‑grade infrastructure experience.
  • Experience with ML frameworks and distributed systems.
  • Hands‑on with AWS/GCP and Kubernetes.
  • Familiarity with LLM training/inference internals is a plus.

Responsibilities

  • Design, build, and scale the compute, scheduling, and data infrastructure that powers post-training research on our in-house RL environments.
  • Develop and maintain core ML framework primitives and internal tooling that researchers rely on daily, accelerating reproducible experimentation and reducing time from idea to result.
  • Build evaluation and benchmarking infrastructure, monitoring, logging, and debugging tooling, and automated testing and deployment systems, so failures are caught early and infrastructure stays reliable as it scales.
  • Partner directly with Research Engineers to translate research needs into infrastructure requirements, and ship fast in response to their feedback.

Skills

PyTorch
JAX
Distributed systems
Cloud platforms (AWS/GCP)
Kubernetes
Data pipelines
LLM training familiarity

Job description

Preference Model is seeking a Senior ML Infrastructure Engineer to build the systems powering post-training research on large language models. You will design and scale the compute, scheduling, and data infrastructure for in‑house RL environments and develop core ML framework primitives used by researchers daily.

You will partner with Research Engineers to translate research needs into infrastructure requirements, ensuring reproducible experiments and fast iteration.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infra Engineer: Scalable AI Training Systems
Senior ML Infra Engineer: Scalable AI Training Systems

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Engineering Manager, ML Infrastructure & Systems
Engineering Manager, ML Infrastructure & Systems

Cursor • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Senior ML Systems Engineer — Scalable AI Infra
Senior ML Systems Engineer — Scalable AI Infra

Meta • Menlo Park (CA)

On-site
USD 347,000 - 403,000
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5
Infrastructure Kernel Engineer for Scalable AI Training
Infrastructure Kernel Engineer for Scalable AI Training

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Lead ML Infrastructure & Evaluation
Lead ML Infrastructure & Evaluation

Cursor • New York (NY)

On-site
USD 180,000 - 260,000
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
Senior AI Infrastructure Engineer — Real-Time ML Systems
Senior AI Infrastructure Engineer — Real-Time ML Systems

Ambient.ai • San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Stock options
Comprehensive benefits
Flexible time off
Senior ML Systems Engineer: Scalable Training Frameworks
Senior ML Systems Engineer: Scalable Training Frameworks

Cohere • San Francisco (CA)

Hybrid
USD 150,000 - 180,000